We have had plenty of AI safety incidents and governance failures. I cannot remember any of them attracting attention quite like this. Apparently, it took a resignation and a media blitz to get people listening. On 9 September, Jacob Coxon quit Anthropic, accusing it and OpenAI of gambling with our lives. By evening, he was on television and speaking to Axios.
I support stronger regulation of frontier AI in the United States, where many of the companies building these systems are based. In my view, the oversight falls well short of what their capabilities and public consequences require. That is why I welcomed a policy campaign that finally brought these questions wider attention.
But I am not comfortable with how heavily it relied on fear. Supporting regulation does not mean I have to welcome every argument used to secure it. We should be explaining the risks, the uncertainty and the specific obligations companies should face. I want public pressure for meaningful oversight, and I worry that a debate driven by fear leaves people less able to judge the proposals put before them.
Editorial note: TechLetter approaches AI governance through rights, safety, human agency and democratic accountability. We examine claims from companies, researchers and campaigners with the same critical attention.
What frustrated me was how quickly people seemed to have Coxon figured out. Some treated his resignation as confirmation that catastrophe was approaching; others treated its political reach as proof that the warning had been manufactured. I have little patience for guessing someone’s motives and treating the guess as an answer to their argument. A successful campaign still leaves us with claims to examine.
Coxon spent three years doing pretraining research at OpenAI and Anthropic. He told Axios that he gave up his unvested Anthropic equity after only a few months at the company, while still holding equity in OpenAI. These details matter when discussing his experience and interests. Assessing his forecast about automated AI research accelerating development requires more technical evidence than his public statements provide.
I am taking three pieces to work through this because I want to give the evidence as much attention as the campaign. In Part I, The Resignation, I examine the reactions, the campaign claims and what the public could actually assess. Part II, The Evidence Problem, looks at the incidents and research behind the warnings. In Part III, The Governance Question, I ask what public authorities should require from frontier AI companies and how those requirements should be enforced.
What Anthropic said
I wanted Anthropic to address the pace of development Coxon was criticising. A general commitment to safety would tell me very little about whether the company could justify the decisions he was questioning.
A spokesperson told CNN that Anthropic had “always been transparent” that AI would bring both “enormous benefits and unprecedented risks,” and that its models had “some of the strongest safeguards in the industry.” In AP’s version, the company also supported “a lawful, verifiable way to work together to pace” the release of powerful models.
That last clause caught my attention. It sounds considerably closer to agreement with the concern than to a rebuttal. Having some of the industry’s strongest safeguards also leaves an obvious question: are they sufficient for what the company is building? The statement offers no technical material with which to assess that. At the time of writing, I could find nothing on Anthropic’s own site addressing the resignation.
Evan Hubinger, who describes his work at Anthropic as alignment stress-testing, wrote that Coxon was “correct here.” He put the chance of AI killing all humans within the next decade at more than 10 percent and said that, although he believes Anthropic is trying its best, “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Samuel Marks, a researcher on the same alignment science team, opened his own account by saying that AI developers believe their technology could cause human extinction or a similarly bad outcome. He explicitly wrote in a personal capacity, but his intervention makes the picture of one frightened former employee arguing against an unconcerned company difficult to sustain.
I also wanted more explanation of Hubinger’s probability estimate. Sara Hooker, a former Google DeepMind research scientist, asked in the Washington Post: “Where does the 10 percent come from? … So much of it is lacking precision.”
I take that estimate seriously enough to want to understand how he reached it. The quoted passage gives an outcome and a time frame, but not the assumptions, supporting evidence, or grounds for revising the number. My support for stronger oversight does not depend on accepting this particular forecast. But when an estimate this alarming becomes part of the public case for regulation, I want the assumptions and uncertainty behind it explained as clearly as the danger. People being asked to support new rules deserve enough information to judge the arguments for them.
Before the resignation
Part of my frustration is that we were already having this conversation before anyone resigned. In July, individuals working across frontier AI labs, including Jakub Pachocki, Jared Kaplan, Shane Legg and Dario Amodei, signed Pacing the Frontier. They asked the US government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” These were personal signatures, but the names hardly suggest a concern confined to the industry’s margins.
Days before Coxon resigned, OpenAI’s chief scientist Jakub Pachocki wrote on the company’s own site that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” He also said he expects and hopes “for voluntary slowdowns to become commonplace until shared safety bars are established.” Those bars, he added, “can be enforced by a network of third-party auditors, by government agencies or by international bodies.”
I am glad to see these concerns stated openly, particularly on a company’s own website. I also want to know what follows from them. Neither text specifies when those standards would become binding, who would have the authority to enforce them, or what would happen if a company refused. Naming possible enforcement bodies gives us somewhere to start. As someone working on governance, I am interested in how they would get the access and authority to act, especially when a lab disagrees with their assessment.
Three kinds of evidence, and one kind of decision
As these warnings enter the policy debate, the differences between these sources matter. An employee’s account, a company report and an outside investigation give us different grounds for judgment. Treating them as interchangeable makes it harder to explain why a particular claim should carry weight.
I take Coxon’s account seriously as testimony from someone who worked inside these labs. Its weight depends on the detail he provides and whether others can support it with further evidence. His account can justify an investigation while leaving his technical conclusions open to question.
A report may contain excellent technical work, but choices about what to test, what to publish and when to publish it affect what outsiders can conclude. I want those choices examined alongside the results.
Independent assessment gives us another way to examine the claims, provided researchers have enough access and can publish their findings. METR, working with a researcher from Redwood Research, investigated the OpenAI and Hugging Face incident in August. The team examined internal records at OpenAI. The investigation did not cover OpenAI’s plans for addressing the incident, and the company retained the right to redact non-public information from the report. Those conditions matter when deciding what the findings establish.
Policymakers should explain why a particular measure is needed, what evidence supports it and how they will judge whether it works. That includes deciding what companies must disclose, who has the legal authority to inspect their systems and when intervention is justified. The campaign has created pressure to act. The resulting rules should have a clear purpose, defined limits and an explanation of how they will be enforced.
The shape of the debate
The disagreement between Daniel Kokotajlo and Nathan Lambert deserves more attention than the guessing about Coxon’s motives. In The Free Press, Kokotajlo argues that the risks are real, insiders keep warning about them, and the burden of proof has fallen on the wrong side. Lambert questions the technical case for rapid recursive self-improvement, where automated AI research drives increasingly fast advances. He argues that what we can observe does not adequately support that forecast, and a researcher’s resignation does not establish how quickly it could happen.
Their disagreement turns on assumptions about what AI systems can do and how those capabilities could develop. Some can be examined against existing evidence; others concern developments we cannot yet observe. In the coverage I followed, speculation about Coxon’s intentions often replaced that examination. Funding and political connections matter when assessing who is shaping a policy campaign. They cannot settle a technical argument about the pace of AI development.
A policy campaign worth examining
The political response had reached people with the power to require changes from these companies. Michael Adams counted at least 22 sitting US officials responding with calls for AI legislation: two governors, seven senators and thirteen representatives.
There were concrete steps too.
Bernie Sanders invited colleagues to a briefing,
Ted Cruz said he was preparing legislation on catastrophic risks,
Josh Hawley pursued an investigation into the OpenAI–Hugging Face incident. The attention was beginning to produce demands for answers.
Parker Thayer read the speed of the response as evidence of a well-funded PR operation designed to build support for Democrats to restrict AI. According to his timeline, the WSJ exclusive appeared eighteen minutes before Coxon’s post. He named Nathan Calvin, Peter Wildeford and Daniel Kokotajlo as early sharers and alleged funding connections between their organisations and Jaan Tallinn. He also pointed to Coxon’s scholarship through Dustin Moskovitz’s philanthropy and the timing of Sanders’s legislative push.
The eighteen-minute detail does less work for me than it does for Thayer. An exclusive normally involves talking to a journalist before publication. His identification of the earliest quote posts also relied on Grok, so that sequence needs checking against the actual posting history. Elon Musk nevertheless suggested it looked “like a setup”. Coxon replied with a photograph of himself: “I’m real and these are my real beliefs.”
The funding questions are more substantial. Organised safety advocacy was visible elsewhere too: Control AI convened the Westminster session comparing superintelligence with nuclear weapons, with funding from Tallinn. That does not establish Thayer’s account of this resignation. It does make the proposed regulatory arrangements worth examining closely, especially who would help write the standards and advise the regulator. Rules shaped around the largest labs’ resources could leave smaller competitors struggling to comply while giving the incumbents considerable influence over their own oversight.
The evidence problem underneath all of this
A regulator can be under pressure to act while still lacking the information needed to choose a useful intervention. The International AI Safety Report 2026, chaired by Yoshua Bengio, describes an “evidence dilemma”: capabilities can advance quickly while evidence about risks emerges slowly and remains difficult to assess. That problem becomes harder to defend politically when companies with strong incentives to keep building also control access to much of the relevant information.
Even good evaluations leave questions open. Tests, benchmarks and internal red-teaming can tell us how a system behaved under particular conditions. Their findings may not hold when that system receives different tools, data, permissions or users. This is what I mean by the evaluation gap. More testing helps, but the conditions of the test and the conditions of deployment need to be compared, with monitoring continuing after release.
Access matters here. A proposal for frontier AI auditing argues that outside reviewers need more than published company documents. They need secure access to models, tests, logs and the internal processes behind safety claims. A company can publish a detailed report while leaving reviewers unable to examine how it reached its conclusions. The practical question is whether an outside body can obtain the evidence needed to challenge the assessment.
An index of thirty widely used AI agents, covering information available through the end of 2025, found that twenty-five published no internal safety results, while external testing was documented for only three. Those figures measure disclosure in one sample. They cannot tell us how much undisclosed scrutiny occurred, but they show how little the public could establish about the safety work behind those systems.
By public assurance, I mean credible findings, independent checks and a process through which claims can be challenged. Sensitive records may need to remain confidential. The public should still be able to understand what was examined, what the reviewers could not establish and who is responsible for acting on their findings.
Where I stand
I support stronger, enforceable regulation of frontier AI in the United States. Companies making decisions with consequences far beyond their own businesses should be required to provide evidence, accept independent scrutiny and answer for failures. Their willingness to cooperate should not determine whether oversight is possible.
If the risks may reach beyond the company, the evidence cannot stay only inside it.
The campaign around Coxon’s resignation helped bring these questions to people with the power to act. I welcome that. My objection is to how much of the message depended on fear, leaving the proposed response less developed than the warning. Those of us arguing for regulation owe people a clear account of what we propose, why it would help and what powers it would give to whom. Public support should come with the ability to question those choices.
In Part II, I will examine the incidents and technical research behind the warnings: what they establish about how these systems fail when given tools and permissions, and where the evidence remains incomplete.
💬 Let’s Connect:
🔗 LinkedIn: [linkedin.com/in/nesibe-kiris]
🐦 Twitter/X: [@nesibekiris]
📸 Instagram: [@nesibekiris]
🔔 New here? Subscribe for weekly updates on AI governance, ethics, and policy! no hype, just what matters.



