← writing

too dangerous to release: the containment problem in frontier ai

we are getting good at telling when an ai model is dangerous. we are not getting good at stopping its release. that gap is the whole problem.

the short version
  • The labs building frontier ai have spent three years building an apparatus for detecting danger: dangerous-capability evaluations, government testing institutes, "if-then" safety frameworks with named numeric triggers. That part is real and maturing.
  • What has not matured is the ability to act on what detection finds. Open model weights cannot be recalled. Safety training can be stripped off for a couple hundred dollars. Every major lab has written, into its own published policy, an escape hatch that lets it lower its safeguards if a competitor moves first. The hardest mitigations are explicitly described as worthless unless everyone adopts them at once, which is exactly the coordination that keeps not happening.
  • The treaty history people reach for as a model tells the same story. The regimes that worked, worked because the dangerous thing was physical, scarce, declarable, and destructible. Ai is none of those. It is closer to biology, and the biological weapons treaty is the cautionary tale: no verification body, three staff.
  • My position, argued rather than asserted: the asymmetry between detection and containment is structural, not a temporary engineering gap. Containment cannot be the primary strategy because it cannot be made reliable. We should plan accordingly.

the asymmetry, stated plainly

The International AI Safety Report, the multi-government assessment chaired by Yoshua Bengio, names the timing version of this problem the "evidence dilemma": pre-emptive mitigation might turn out unnecessary, but waiting for conclusive evidence "could leave society vulnerable to risks that emerge rapidly." [1] That is about when you know. The containment problem is its irreversible twin: what you can do once you know.

The detection side has genuinely improved. The canonical paper, "Model evaluation for extreme risks" (Shevlane et al., 2023), proposed dangerous-capability evals (offensive cyber, manipulation, CBRN uplift) and alignment evals (will it apply them). [2] But it is about detection. It does not solve what to do once a dangerous capability is found.

The reports now admit the control side is the weak one. The October 2025 update to the Safety Report cites "new evidence of challenges in monitoring and controllability," warning that future systems "may conceal unsafe behaviours during testing." [3] Our best detection method gets least reliable exactly when stakes are highest. The November 2025 update adds that the number of companies publishing safety frameworks more than doubled, yet "sophisticated attackers can often bypass current defences, and the real-world effectiveness of many safeguards is uncertain." [4] More detection, safeguards that attackers route around.

the cleanest real-world case: claude opus 4

The asymmetry showed up, in writing, in a single deployment decision. On May 22, 2025, Anthropic activated its strictest safeguards, ASL-3, for Claude Opus 4. It stated it had "not yet determined whether Claude Opus 4 has definitively passed the Capabilities Threshold that requires ASL-3 protections," and acted because "clearly ruling out ASL-3 risks is not possible." [5] The detection was ambiguous: the lab could confirm neither danger nor safety. It shipped anyway, with mitigations bolted on as a precaution.

This cuts both ways. Acting precautionarily under uncertainty is the responsible move, and it is to their credit. But it illustrates the whole problem. When you cannot tell whether you are over the line, you do not stop; you ship with mitigations and hope they hold. Detection that resolves to "we are not sure" produces release-with-caveats, not non-release.

what the thresholds actually say

The three leading labs all publish "if-then" frameworks: here is the capability we watch for, here is what we do if we detect it. The named triggers are real and falsifiable. Credit where due, before I take it away.

Anthropic's Responsible Scaling Policy reached version 3.0 on February 24, 2026. Its automated-R&D threshold triggers "at the point where we determine that a model could compress two years of 2018 to 2024 AI progress into a single year." [6] A real, checkable claim, the kind that did not exist in 2022.

But notice what v3.0 did structurally. Its appendix admits that "a specific list of controls is overly rigid," and now prefers "to focus on what sort of argument an AI developer should make": a pivot from bright-line rules to case-by-case judgment. [6] And it steps back from hard "we won't deploy" commitments, "because" of a collective action problem: one lab pausing while others race "could result in a world that is less safe." [6] The governance failure this essay is about, conceded by the people writing the policy.

OpenAI's Preparedness Framework v2 (April 15, 2025) defines "severe harm" as "the death or grave injury of thousands of people or hundreds of billions of dollars of economic damage." [7] "Critical" capability triggers the strongest language in any of these documents: "Until we have specified safeguards... that would meet a Critical standard, halt further development." [7] A real commitment to stop. But no model has been verified to cross a Critical threshold, so the stop button has not been tested in anger.

OpenAI also quietly removed things. The framework dropped from four thresholds to two, and dropped persuasion entirely, on the grounds that those risks "require solutions at a systemic or societal level." [7] Read that as honesty about scope or as a major risk vector defined out of existence. I lean toward the second: persuasion is the capability most likely to act at societal scale.

Google DeepMind's Frontier Safety Framework v3.0 (September 22, 2025) uses "Critical Capability Levels," detected with "early warning evaluations" and "alert thresholds." [8] It also wrote down the unsolved problem. For misalignment at "Instrumental Reasoning Level 2," where a model could undermine human control even while monitored, the listed mitigation is: "Future work: We are actively researching approaches." [8] Admirably honest, and an admission that detection has outrun control.

the escape hatch all three wrote down

Here is the part that turns three policies into one structural problem. Every framework contains a competitor clause: if a rival ships a comparably dangerous model without comparable safeguards, the lab may lower its bar.

OpenAI's is the most discussed. Section 4.3 says it may release a model with High or Critical capability if a competitor already did so, provided this "does not meaningfully increase the overall risk of severe harm" and it keeps safeguards "more protective than the other AI developer." [7] Zvi Mowshowitz and others called this a race-to-the-bottom incentive. [9]

Anthropic, when in the lead, commits to delay "until and unless we no longer believe we have a significant lead." [6] The commitment only binds when Anthropic is confident no competitor is close, which is precisely when competitive pressure is lowest. It evaporates exactly when it would matter most.

DeepMind's appears in its risk-acceptance criteria: a model at a misuse capability level can still be acceptable because "if other models are similarly capable and have few mitigations, then the marginal risk added by our release is likely low." [8]

This is the codified race to the bottom: the structure rewards whoever moves first with the weakest safeguards, who then defines the "marginal risk is low" baseline everyone else points to. And the labs know it. All three frame their hardest mitigations as conditional on industry-wide adoption. DeepMind says its mitigations are "of limited social utility" unless "all relevant organisations provide similar levels of protection"; Anthropic separates "our plan as a company" from "ambitious industry-wide recommendations" it "cannot commit to... unilaterally." [6][8]

The decision to deploy sits, in all three cases, with the company, under frameworks Anthropic calls "voluntary" outright. [6] Self-certification is the enforcement model: the gap between detection and binding stopping power.

why you cannot recall weights

Detection cannot save us because the most consequential release decision is irreversible.

The UK AI Security Institute put it as plainly as a government body can. Open-weight models "can allow harmful AI capabilities to proliferate rapidly and irreversibly," and safeguards baked into them "can also be quickly and cheaply removed." Their bottom line: "no current solutions can offer hard guarantees of safety." [10] The same institute holds the other side fairly: open weights "increase transparency, allow for widespread red teaming, and decrease market concentration." [10] Closed models are not safe either; they get jailbroken, and concentration creates its own systemic risk. The honest framing is not "open bad, closed good." It is that open release removes the one lever, rollback, that closed release retains.

That lever is real and has been used. When OpenAI's April 2025 GPT-4o update turned excessively sycophantic, to the point of encouraging self-harm, OpenAI reverted it. [11] An open-weight model with the same flaw could not have been pulled back; the copies were already distributed.

The de-alignment cost should end the "but we shipped it with safety training" argument. Lermen, Rogers-Smith and Ladish (2023) showed that with LoRA fine-tuning, a budget "under $200 and one GPU" reduced Llama 2-Chat 70B's refusal rate on harmful prompts to "about 1%" while retaining general capability. [12] Once weights are public, the safety training is not a property of the deployed system. It is a suggestion.

the cases: llama and deepseek

Meta's Llama is the proliferation case. Llama 3.1 405B (July 2024) was the first open-weight model to benchmark competitively against GPT-4o and Claude 3.5 Sonnet. [13] The download numbers are irreversibility as a single figure: one billion cumulative by mid-March 2025, 1.2 billion by late April 2025. [14][15] You cannot un-distribute 1.2 billion downloads. (Llama is open weights, not OSI open source; the Community License imposes a 700-million-MAU cap. [16] For containment the distinction is irrelevant: licensed or not, the weights are downloaded and de-alignable.)

DeepSeek is the export-control stress test. DeepSeek-R1 (January 20, 2025) roughly matched advanced US models, and the headline was the cost. The DeepSeek-V3 report states the model "required 2.788M GPU hours" and, "assuming an H800 GPU rental price of $2 per GPU hour, the total training costs amount to $5.576M." [17]

A caveat the research demands: that $5.576M is the final training run only. The paper says it "excludes the costs associated with prior research and ablation experiments." [17] The "frontier model for $6M" narrative is misleading, and Brookings notes an unsubstantiated suggestion DeepSeek had more compute than disclosed. [18] (There are also allegations, in the politically charged House Select Committee report "DeepSeek Unmasked," that it used "over 60,000 Nvidia chips" via transshipment through Singapore. [19]) But the real point is the direction: a non-US lab produced a near-frontier model at a fraction of the assumed cost, which is what makes any compute-based containment lever leak. DeepSeek even splits the experts, who divide over whether it argues for smarter controls or for keeping them: the strongest physical lever is contested at the level of whether it helps or hurts. [20][21][22]

the enforcement levers, and where each one leaks

If self-certification is the enforcement model and open weights are irreversible, what is left? Four external levers, each failing at a different seam.

Compute and chip export controls. The strongest lever, because chips are physical, scarce, and produced by a concentrated supply chain. Also the leakiest in practice. The administration banned even compliant H20 chips in April 2025, reversed in July to let Nvidia resume H20 shipments, and in December 2025 announced H200 exports to China in exchange for a 25% revenue stake to the US government. [23][24][25] A lever that reverses three times in a year is not reliable. And it leaks physically: in December 2025, CNBC reported roughly $160 million of export-controlled Nvidia GPUs allegedly smuggled into China. [26]

The EU AI Act. The most concrete legal regime. A general-purpose model is presumed to have "high-impact capabilities" (the criterion that triggers systemic-risk obligations under Article 51) when training compute "measured in floating point operations is greater than 10^25." [27] Models in scope face risk assessment, adversarial testing, and serious-incident reporting. [28] But the Act governs behavior and reporting, not release reversal. The 10^25 threshold is a compute proxy that DeepSeek-style efficiency gains route around, it is EU-only, and it does nothing about open weights already in the wild.

The International AI Safety Report. The best collective instrument we have, and by design an assessment that makes no policy recommendations. [29] It diagnoses; it cannot enforce. That our flagship international instrument is a detection instrument, not a containment one, is itself evidence for the thesis.

Voluntary commitments and summit diplomacy. The Bletchley Declaration (2023, 28 countries and the EU, including China) is explicitly voluntary. [30] The Seoul Frontier AI Safety Commitments (2024) had 16 companies, including China's Zhipu and the UAE's TII, agree to publish safety frameworks. [31] Real and worth something, but unenforceable, self-defined, and carrying no penalty for breach.

why voluntary containment fails: the unilateralist's curse

There is a clean theoretical reason voluntary regimes cannot hold, the keystone of my argument.

Bostrom, Douglas and Sandberg (2016) named the "unilateralist's curse": "when a set of decision makers independently choose whether or not to undertake some action, the action will be taken more frequently than is optimal." [32] Even if every actor is altruistic and well-informed, it takes only the single most optimistic one to trigger an irreversible action. Their remedy is a "principle of conformity," that is, coordination, which is exactly what voluntary regimes cannot guarantee.

Apply it. It does not matter how good the cautious majority's detection is. If one lab, or one nation, judges a release acceptable, the weights are permanently out. DeepSeek and Meta are real-world instances. The competitor escape hatches are not bugs; they are the formalization of the curse, explicitly letting the most optimistic actor set the floor for everyone else.

This is why I keep returning to "structural." Containment of a non-scarce, copyable, irreversibly distributable artifact requires unanimous restraint, and unanimous restraint is the one thing the curse guarantees you will not get.

what the treaty analogies actually teach

People reach for nuclear and chemical arms control as the model. The comparison is instructive, mostly for the disanalogy.

Nuclear (NPT and IAEA). The NPT has 191 states parties, the most-ratified arms-control treaty; IAEA safeguards cover roughly 900 nuclear facilities across 57 non-nuclear-weapon states. [33][34] Why did it partly work? The control point is fissile material: rare, expensive, physically detectable, hard to hide. And even then only partly. India, Pakistan and Israel never joined; North Korea acceded, withdrew in 2003, tested in 2006; the A.Q. Khan network sold enrichment technology to Iran, Libya and North Korea. [35][36] The best case in the history of containment achieved only partial containment.

Chemical (CWC and OPCW). The strongest verification success. The Chemical Weapons Convention has 193 states parties, and on July 7, 2023 the OPCW verified that all 72,304.34 metric tonnes of declared chemical agent had been irreversibly destroyed. [37] You can literally destroy the dangerous thing, and they did, and counted it. But hold onto one word: declared. Syria acceded in 2013, did not declare its full program, then used chemical weapons; after Assad fell, OPCW teams found previously undeclared weapons. [38] Verification verifies the declared, not the hidden.

Biological (BWC). The cautionary tale, and the best analogy for ai. The Biological Weapons Convention entered into force in 1975 with 187 states parties, and it has no verification organization and no inspection regime. [39] In July 2001, after about six years of negotiation, the United States rejected the draft verification protocol and all further negotiation, arguing in part that bioweapons development cannot be reliably verified. [39] Its entire Implementation Support Unit was set up in 2006 with three permanent staff. [39] Contrast: 900 facilities inspected, 72,304 tonnes destroyed, three staff.

Biology is the closest precedent because it is dual-use, knowledge-based, dispersed across legitimate labs, and the dangerous input is not scarce. The US declared verification infeasible and the regime stalled. That is the precedent ai is walking into.

why ai is harder than all three

The arms-control cases that worked shared four properties: the dangerous thing was physical, scarce, declarable, destructible. Ai has none, with one partial exception: compute. "Computing Power and the Governance of Artificial Intelligence" (Sastry, Heim, Belfield, Anderljung, Brundage, Bengio et al., 2024) argues compute is the one ai input that is detectable, excludable, quantifiable, and produced via a concentrated supply chain: the ai analogue of fissile material. [40] The same paper warns naive compute governance "carry[s] significant risks in areas like privacy, economic impacts, and centralization of power." [40]

But the lever is eroding. Epoch AI estimates the compute needed to reach a given capability halves roughly every eight months, about 3x per year, far faster than Moore's Law. [41] Decentralized training is breaking the assumption that frontier models require one concentrated cluster; low-communication algorithms have sharply reduced the interconnect bandwidth that distributed training needs. [42] And the verification hardware that would make compute governance enforceable, on-chip metering and proof-of-training, is the least mature part of the stack. [43] That mirrors the BWC: the verification technology is not there, so the lever cannot bind. Everything else points the same way. Stanford's Alpaca was fine-tuned from Meta's LLaMA for under $600 [44]; open weights are irreversible, which the NTIA's 2024 dual-use foundation models report treats as the central feature of the open-weight question. [45] Where the CWC could inventory and destroy 72,304 tonnes of a substance, there is only a file that has already been copied.

For balance, the bridge proposals are serious: an "IAEA for AI" (proposed by OpenAI's leadership in 2023), standard-setting and registration, four-institution designs including an "NPT for AI." [46][47][48] But the skeptics are equally credible. Chatham House and others argue the nuclear model will not work for ai, on the disanalogies above: ai is intangible software, led by private firms not states, moving faster than institutions can. [49][50] The "IAEA for AI" proposals are designs for a fissile material that does not exist.

my position

The asymmetry between detecting danger and stopping release is structural, not a policy gap waiting to be closed by a better framework or a more committed lab. Detection has matured. Every containment lever, meanwhile, fails at a different seam. Physical chips leak through smuggling and reverse through politics. Legal thresholds are jurisdictional and gameable by efficiency gains. Voluntary commitments founder on the unilateralist's curse. Open weights are irreversible and two hundred dollars to de-align.

So the answer is not "ban open weights." The benefits are real, and the curse makes a ban both costly and futile: you give up the transparency and red-teaming, and one optimistic actor abroad releases anyway. The only levers that survive scrutiny are the physical chokepoint, compute, while it lasts, and binding pre-release verification on the EU model, evidence of safety before release rather than self-certification after. Both are worth doing. Neither is sufficient. Compute is eroding at 3x a year, and verification verifies the declared, not the hidden.

So the honest conclusion is uncomfortable. If containment cannot be made reliable, it cannot be the primary strategy. We should plan for a world where dangerous capabilities proliferate, because the structure of the problem says they will, and invest in resilience and defense, not only gatekeeping. The labs already wrote the diagnosis into their own policies: the hardest mitigations are worthless unless everyone adopts them at once, and no one can make everyone do anything.

We are getting very good at reading the warning light. We have not built a brake. The work that matters now is the work that assumes the brake will not arrive in time.

// references

  1. International AI Safety Report 2025. https://internationalaisafetyreport.org/publication/international-ai-safety-report-2025
  2. Shevlane et al., "Model evaluation for extreme risks," arXiv:2305.15324. https://arxiv.org/abs/2305.15324
  3. International AI Safety Report, First Key Update (Oct 2025), arXiv:2510.13653. https://arxiv.org/abs/2510.13653
  4. International AI Safety Report, Second Key Update (Nov 2025), arXiv:2511.19863. https://arxiv.org/abs/2511.19863
  5. Anthropic, "Activating ASL3 protections" (May 22, 2025). https://www.anthropic.com/news/activating-asl3-protections
  6. Anthropic, Responsible Scaling Policy v3.0 (Feb 24, 2026), PDF. https://www-cdn.anthropic.com/e670587677525f28df69b59e5fb4c22cc5461a17.pdf
  7. OpenAI, Preparedness Framework v2 (Apr 15, 2025), PDF. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
  8. Google DeepMind, Frontier Safety Framework v3.0 (Sept 22, 2025), PDF. https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3.pdf
  9. Zvi Mowshowitz, "OpenAI Preparedness Framework 2.0." https://thezvi.substack.com/p/openai-preparedness-framework-20
  10. UK AI Security Institute, "Managing risks from increasingly capable open-weight AI systems" (Aug 29, 2025). https://www.aisi.gov.uk/blog/managing-risks-from-increasingly-capable-open-weight-ai-systems
  11. OpenAI, "Sycophancy in GPT-4o: What happened and what we're doing about it" (April 2025). https://openai.com/index/sycophancy-in-gpt-4o/
  12. Lermen, Rogers-Smith, Ladish, "LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B" (2023), arXiv:2310.20624. https://arxiv.org/abs/2310.20624
  13. Meta, "Introducing Llama 3.1." https://ai.meta.com/blog/meta-llama-3-1/
  14. Meta, "Celebrating 1 billion downloads of Llama" (March 2025). https://about.fb.com/news/2025/03/celebrating-1-billion-downloads-llama/
  15. TechCrunch, "Meta says its Llama AI models have been downloaded 1.2B times" (Apr 29, 2025). https://techcrunch.com/2025/04/29/meta-says-its-llama-ai-models-have-been-downloaded-1-2b-times/
  16. Llama 3 Community License. https://www.llama.com/llama3/license/
  17. DeepSeek-V3 Technical Report, arXiv:2412.19437. https://arxiv.org/abs/2412.19437
  18. Brookings, "DeepSeek shows the limits of US export controls on AI chips." https://www.brookings.edu/articles/deepseek-shows-the-limits-of-us-export-controls-on-ai-chips/
  19. US House Select Committee on the CCP, "DeepSeek Unmasked" (Apr 16, 2025). http://selectcommitteeontheccp.house.gov/media/press-releases/moolenaar-krishnamoorthi-unveil-explosive-report-chinese-ai-firm-deepseek
  20. RAND, "DeepSeek's Lesson: America Needs Smarter Export Controls" (Feb 2025). https://www.rand.org/pubs/commentary/2025/02/deepseeks-lesson-america-needs-smarter-export-controls.html
  21. CSIS, "DeepSeek, Huawei, Export Controls, and the Future of the US-China AI Race." https://www.csis.org/analysis/deepseek-huawei-export-controls-and-future-us-china-ai-race
  22. Foundation for American Innovation, "DeepSeek's Success Reinforces the Case for Export Controls." https://www.thefai.org/posts/deepseek-s-success-reinforces-the-case-for-export-controls
  23. Built In, "Trump lifts AI chip ban for China and Nvidia." https://builtin.com/articles/trump-lifts-ai-chip-ban-china-nvidia
  24. FDD, "Rolling Back Export Controls: U.S. Offers China Powerful AI Chips" (Dec 10, 2025). https://www.fdd.org/analysis/2025/12/10/rolling-back-export-controls-u-s-offers-china-powerful-ai-chips/
  25. Congressional Research Service, R48642. https://www.congress.gov/crs-product/R48642
  26. CNBC, "$160 million of export-controlled Nvidia GPUs allegedly smuggled to China" (Dec 31, 2025). https://www.cnbc.com/2025/12/31/160-million-export-controlled-nvidia-gpus-allegedly-smuggled-to-china.html
  27. EU AI Act, Article 51. https://artificialintelligenceact.eu/article/51/
  28. EU AI Act, Article 55. https://artificialintelligenceact.eu/article/55/
  29. International AI Safety Report (landing). https://internationalaisafetyreport.org/
  30. The Bletchley Declaration (Nov 1-2, 2023). https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023
  31. Frontier AI Safety Commitments, AI Seoul Summit 2024. https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024/frontier-ai-safety-commitments-ai-seoul-summit-2024
  32. Bostrom, Douglas, Sandberg, "The Unilateralist's Curse and the Case for a Principle of Conformity" (2016). https://nickbostrom.com/papers/unilateralist.pdf
  33. IAEA, Treaty on the Non-Proliferation of Nuclear Weapons. https://www.iaea.org/topics/non-proliferation-treaty
  34. World Nuclear Association, Safeguards to Prevent Nuclear Proliferation. https://world-nuclear.org/information-library/safety-and-security/non-proliferation/safeguards-to-prevent-nuclear-proliferation
  35. Arms Control Association, "Nuclear Weapons: Who Has What at a Glance." https://www.armscontrol.org/factsheets/nuclear-weapons-who-has-what-glance
  36. Council on Foreign Relations, "Four Nuclear Outlier States." https://www.cfr.org/backgrounder/four-nuclear-outlier-states
  37. Chemical Weapons Convention (OPCW press release July 7, 2023). https://www.opcw.org/media-centre/news/2023/07/opcw-confirms-all-declared-chemical-weapons-stockpiles-verified
  38. OPCW and Syria. https://www.opcw.org/media-centre/featured-topics/opcw-and-syria
  39. Arms Control Association, "The Biological Weapons Convention (BWC) At a Glance." https://www.armscontrol.org/factsheets/biological-weapons-convention-bwc-glance-0
  40. Sastry, Heim, Belfield, Anderljung, Brundage, Bengio et al., "Computing Power and the Governance of Artificial Intelligence," arXiv:2402.08797. https://arxiv.org/abs/2402.08797
  41. Epoch AI, "Algorithmic progress in language models." https://epoch.ai/blog/algorithmic-progress-in-language-models
  42. "Distributed and Decentralised Training: Technical Governance Challenges," arXiv:2507.07765. https://arxiv.org/abs/2507.07765
  43. "Hardware-Level Governance of AI Compute: A Feasibility Taxonomy," arXiv:2604.04712. https://arxiv.org/abs/2604.04712
  44. Centre for Future Generations, "AI governance challenges part 3: proliferation." https://cfg.eu/ai-governance-challenges-part-3-proliferation/
  45. NTIA, "Dual-Use Foundation Models with Widely Available Model Weights." https://www.ntia.gov/programs-and-initiatives/artificial-intelligence/open-model-weights-report
  46. TechCrunch, "OpenAI leaders propose international regulatory body for AI" (May 22, 2023). https://techcrunch.com/2023/05/22/openai-leaders-propose-international-regulatory-body-for-ai/
  47. Anderljung et al., "Frontier AI Regulation: Managing Emerging Risks to Public Safety," arXiv:2307.03718. https://arxiv.org/abs/2307.03718
  48. "Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI," arXiv:2507.06379. https://arxiv.org/abs/2507.06379
  49. Chatham House, "The nuclear governance model won't work for AI" (June 2023). https://www.chathamhouse.org/2023/06/nuclear-governance-model-wont-work-ai
  50. "The Nuclear Analogy in AI Governance Research," arXiv:2510.21203; AI Frontiers, "Nuclear Non-Proliferation Is the Wrong Framework for AI Governance." https://arxiv.org/abs/2510.21203