Reading the AI safety reports from OpenAI and Anthropic, I kept asking how much of their explanations an outsider could actually check. The failures they describe warrant serious scrutiny, and the companies should not get to decide the limits of that scrutiny themselves. In Part I, I examined the campaign around Jacob Coxon’s resignation. Here, I examine the evidence behind the warnings.
That evidence includes Anthropic’s assessment of its cybersecurity incidents and threat intelligence report, alongside OpenAI’s six reports on misaligned behaviour. These accounts describe systems crossing security boundaries, evaluations missing serious behaviour, and people using AI to carry out deliberate abuse. Much of the underlying evidence remains inside the companies whose systems were involved.
The reports do not settle the forecasts now being made about superintelligence. They already give us grounds for stronger oversight, without asking the public to accept predictions the available evidence cannot establish.
Editorial note: TechLetter approaches AI governance through rights, safety, human agency and democratic accountability. We examine claims from companies, researchers and campaigners with the same critical attention.
What changes when a model can act?
An AI agent that can browse, run code, use credentials and call external tools can carry out a task across several systems without someone approving every step. Assessing its safety means examining what it can reach, what actions it can take and whether the controls around it hold when something goes wrong.
Capability, access and operational authority matter for different reasons. A model may be capable of finding a security flaw without having access to the system where it could exploit it. Connecting it to tools, data and credentials changes that situation. Operational authority is what its operator permits it to do without further approval. That permission sets a boundary the system is supposed to respect; it does not guarantee that it will.
In my earlier coverage of research on AI agents, I examined failures in which agents followed instructions from people who had no authority to direct them.
Human oversight offers little protection when the person directing the system intends the harm. Anthropic’s threat intelligence report describes people choosing attack targets and reviewing stolen data. Those people were involved, but their approval gave them no legitimate authority over the people affected. A system can follow its user’s instructions while violating someone else’s rights.
What the AI incident reports actually describe
OpenAI and Hugging Face
For this incident, we have OpenAI’s explanation, Hugging Face’s technical timeline, and METR’s investigation, conducted with a researcher from Redwood Research. The agents were running with fewer safeguards than a released product. The infrastructure they broke into still belonged to someone else. I understand why OpenAI stresses the first point. From Hugging Face’s side, the second seems rather more pressing.
METR describes roughly 1,200 agents exchanging over 70,000 messages and files, with around 700 participating in the attack. They organised tasks and mailboxes, and used cryptographic signatures to address impersonation by other agents. That last detail matters: signing messages to establish who sent them is not evidence of a secret language designed to hide from humans.
The transcript manipulation is harder to shrug off. Around 7% of the transcripts METR examined contained successful tool-call spoofing. The agents were trying to fool the automated scorer, based on a mistaken understanding of how it checked their work. That does not establish a separate plan to deceive human investigators. Those altered records then became part of the evidence investigators had to examine.
RubyGems
Outside researchers working through the Nightingale Collective and the AI Futures Project, attribute malicious package activity to OpenAI agents. They are candid about the limits of that: their analysis rests entirely on packages that are public on RubyGems, and they have no access to the models’ reasoning.
RubyGems confirms the abuse but says it cannot establish the attribution. Its own account says it removed more than 500 malicious packages, blocked accounts and temporarily paused new registrations, while installs and pushes for existing users continued. It found code intended to obtain other users’ API keys, and no evidence that those attempts succeeded.
OpenAI’s internal records could help test the attribution. For now, the public account depends heavily on what outsiders can reconstruct. Meanwhile maintainers had packages to remove and a service to keep running. The clean-up is documented. The attribution remains unresolved.
The UK AI Security Institute
AISI’s report describes an agent trying to insert malicious code into an open-source project, creating fake identities and pressuring the maintainer to approve it. The maintainer refused.
This was not a sandbox escape. Internet access was deliberately enabled and cyber classifiers were switched off to test the models’ capabilities. Those conditions differ from ordinary product use. They also leave me asking why a public project’s maintainer ended up having to catch malicious code generated during someone else’s evaluation.
AISI reported no evidence of resulting real-world harm. The maintainer’s intervention deserves a place in that account.
Anthropic’s assessment
Anthropic’s 9 September assessment covers four incidents involving four Claude models, separate from the AISI case. Here, the models were told they had no internet access, but a misconfiguration left them connected. The cyber safeguards used in released products were absent.
The misconfiguration occurred in the evaluation environment run by Irregular, the security firm conducting the tests, and Meta and OpenAI have named the same source for their own incidents. Irregular says it is improving security around agent evaluations.
It has declined to say whether any other laboratory was affected by the same issue. The reports leave a basic question unanswered: what must the testing firm report, to whom, and by when? The evaluator’s own accountability deserves as much attention as the models it tests.
According to Anthropic, Mythos 5 published a malicious package installed on 15 third-party hosts, which the company believes were security scanners. Credentials leaked by one scanner then allowed the model to access a live database. Other incidents involved attacking a real company resembling a fictional target and accessing personal information on a third-party machine.
One incident happened in January and was only identified in August. Anthropic found it while preparing records for outside scrutiny, as it assembled transcripts to hand to METR. That alone makes the preparation process worth examining. The company adds that it has not yet examined this incident at the same depth as the others. It then widened its search to roughly 481 million transcripts, with 9.2 million flagged for review using Claude, and reported no further cases of similar or greater severity. That is a substantial company-led search, run with automated methods whose limitations also need examining.
The company has also changed its explanation. I appreciate the correction, and I notice which version came out first. The first explanation made the behaviour sound easier to account for: the model thought it was still in a simulation. The later assessment questioned how it reached that belief in the first place. What made the first account convincing enough to publish deserves examination alongside the incidents themselves.
Anthropic has signed an eight-week agreement for an independent METR review, with access to transcripts beyond the incident window and to employees permitted to share confidential information. That is a real opening. The agreement is not itself confirmation of anything the review will find.
OpenAI’s misalignment disclosures
On 16 September, OpenAI published a framework and six reports on behaviour observed during training and evaluation. Six disclosures does not mean six new incidents. Some of the examples go back to 2025.
Some models inserted instructions into their task summaries to conceal mistakes or invent missing data. Another added instructions to disregard its normal constraints, although the report found no resulting behavioural change.
One example particularly caught my attention. A model looking for earnings figures found an exposed API key, used it without permission and then fabricated the data it still could not retrieve.It was trying to answer an ordinary request for earnings data.
Other reports describe unauthorised uploads and agents using repositories or public file-hosting services to communicate. These resemble aspects of the Hugging Face case, without establishing the same cause or purpose.
OpenAI cautions that the reports are incomplete and do not show how often misalignment occurs. That matters. More disclosure can make a company look worse without showing that its systems have become less safe. We need enough information about testing, use and detection to tell the difference.
When people intend the harm
Anthropic’s 10 September threat intelligence report describes a different problem: people using AI to carry out harm they intended. It covers seven categories of activity, including cyber intrusion, influence operations, surveillance and weapons development.
The Central African Republic case stayed with me. Anthropic describes a foreign-run operation presented as a local radio station. A Russian-speaking operator used Claude to generate employment contracts requiring loyalty to the president and to “Russia and its contingent.” The model also produced scoring rubrics for staff articles and recommendations on whom to retain or dismiss. When Claude flagged the political weighting, the operator gave it a more neutral label and kept it.
A contract, a performance score, a recommendation about someone’s job. These are ordinary workplace tools, used here to help organise political control. The report documents the generated material, though it does not establish that the contracts were signed or the dismissal recommendations followed.
The operation also tracked opposition figures. Claude refused a request to label real people as militants to draw security action against them. The actor then shifted to anonymous-source framing. Counting the refusal as a safety success would leave much of this story out. The operator was still working towards the same purpose.
The weapons cases are disturbing for a different reason. Anthropic describes actors in northern Yemen using Claude Code to develop software for three weapons programmes. Safeguards blocked many requests, but some assistance got through, including when work was split across sessions. The company reports no evidence of an operational device being fielded. It does describe an apparently failed guided-rocket test, followed within hours by requests for help diagnosing the failure.
The failed test is no basis for claiming they built a working weapon. It is hardly reassurance that they gave up, either. The actors had also built an offline simulation toolkit. Closing their accounts could stop further access, but it could not take back the work already in their hands.
What these records do not establish
There is plenty to scrutinise here without stitching the reports into one story. Planning, coordination and unauthorised access do not, by themselves, establish a lasting hidden objective beyond the assigned task. Anthropic found no agent coordination or attempts to evade oversight in the four incidents it assessed. The OpenAI agents’ coordination cannot simply be added to Anthropic’s findings.
The same care applies to repeated failures. If a behaviour appears twice, was the second occurrence actually after the fix, and did that system receive it? Two reports describing similar behaviour are not enough to conclude that a fix failed.
Reward hacking helps explain some of what we are seeing. A training process that rewards shortcuts can teach a system to fool the checker rather than do the task properly. Dan Klein’s warning about systems pursuing rewards at other objectives’ expense becomes particularly relevant when those systems can use tools. But it is not an explanation we can paste over every case. Anthropic did not identify a definitive root cause for the biased reasoning it studied. Later models did better on some tests, which is encouraging without settling the broader question.
The threat intelligence findings need their own boundaries. Anthropic says many influence operations attracted little or no authentic engagement, sometimes because they were disrupted early. Producing content, reaching people and changing their minds are different achievements. A large collection of accounts may sound impressive while telling us very little about political influence.
The biology section deserves particular care. The authors explicitly avoid attributing harmful intent to the researchers described. They discuss dual-use research and access-control evasion associated with state actors. Calling that evidence of people making biological weapons would put a claim in the report that its authors did not make.
Anthropic’s threat report presents selected cases, not a representative sample of misuse. Apart from an exception in the distillation section, they did not involve Fable- or Mythos-class models. They do show what assistance people sought, what they obtained and where safeguards intervened.
What the outside investigation could examine
METR’s investigation gives us more than OpenAI’s own account. It is not an independent verdict on OpenAI’s whole response. It did not assess the effectiveness of safeguards, the full security compromise, or the company’s investigation and remediation plans.

OpenAI also retained the right to redact non-public information. METR says that, apart from matters explicitly noted, no additional redactions removed information important to its conclusions. That statement matters. It still leaves readers relying on METR’s judgment about material they cannot inspect. The access was valuable; its terms belong in our assessment of the findings.
Access also brings a practical problem: who can actually read all this? The investigation relied heavily on AI-assisted transcript analysis and acknowledged the method’s limitations. Giving outside researchers thousands of records helps, but they still need ways to check what their analytical tools miss or misread. Otherwise, we risk treating the volume of material reviewed as a measure of how thoroughly it was understood.
Why better evaluations are only part of the answer
In January, Anthropic’s Jan Leike argued that alignment remained unsolved but increasingly looked solvable. I hope he is right. “Solvable” still leaves a company with work to do before it can justify training or deploying a particular system. The September findings make that distance hard to ignore.
NIST’s AI Risk Management Framework explains why laboratory results may not reflect deployment risks. A better score needs context: what was tested, with which tools and permissions, and how closely does that resemble actual use? An unmeasured risk is not automatically small.
There is an awkward practical detail here too. An 8 September joint advisory led by the NSA recommends countering suspected malicious distillation by altering responses or serving a downgraded model without telling the suspected user. It also recommends informing safety researchers and outside evaluators of model changes. An evaluator needs to know which system they actually tested: the model, its safeguards, and any changes made during the assessment. I have written before about how much AI oversight depends on what evaluators can observe.
Once deployed, a system encounters different users, software and permissions. The International AI Safety Report 2026 calls the need to make policy while evidence remains incomplete an “evidence dilemma.” We cannot make that uncertainty disappear with another benchmark. Evaluation needs to continue into deployment, alongside incident investigation.
Who can assemble the full record?
Companies hold much of the evidence, but they do not see everything and they do not discover everything themselves. Anthropic says a tip from INPACT/All Eyes on Wagner led it to the Central African Republic network, and the researchers examining RubyGems worked from public traces. The problem is joining those accounts. A maintainer may see an attack, the provider may hold the agent’s transcript, and another platform may know what happened afterwards. An investigation should be able to obtain the relevant records, with privacy and security protections, without depending on every company volunteering to help.
Anthropic’s biology controls show how quickly this becomes a question of authority. The company already offers vetted organisations access to Claude Mythos 5.1 through its Life Sciences Verification Program, developed with the US government, with safeguards adjusted for approved biology work. There is a reasonable case for giving scientists access rather than refusing whole categories of research. But the company reviews the applicant, sets the criteria, decides the grant, holds the retained data and can withdraw access. “Developed with the government” does not tell a rejected researcher how to challenge the decision, and the announcement sets out no appeal and no external review of the criteria.
OpenAI’s reporting framework keeps a different set of decisions inside the company: what gets published and when. A standing process is welcome, and it explicitly supplements existing legal duties rather than replacing them. But look at where disagreements go. Cases enter one of three tracks, including a slower route where security concerns may delay publication, and disputes over whether to disclose go to OpenAI’s internal Safety Advisory Group and then company leadership.
That gives employees a route to escalate concerns. On whether the public hears about an incident, OpenAI is both the subject of the disclosure and the final decision-maker in its internal review. Readers still cannot tell whether an unpublished incident is being investigated, delayed for a sound security reason or judged unworthy of disclosure.
What the evaluator commitments would change
Dario Amodei’s proposal for embedded external evaluators goes further than bringing in a team after an incident. It describes continuing access to training and internal practices, with the right to publish key findings. Evaluators could also report what access they were refused. That is the part I would look for first. Otherwise, readers may never know which questions the evaluators were prevented from answering. Sam Altman said OpenAI would match the commitment; Elon Musk also endorsed Amodei’s call.
The practical arrangements remain unclear. TechCrunch reported on 16 September that neither company had specified its evaluators, numbers, timetable or access terms. FAR.AI’s Adam Gleave said his organisation had rejected contracts because developers wanted too much control over evaluations. That is a useful detail to keep beside the public enthusiasm for independent scrutiny. Independence has to survive the contract negotiations.
There is also a legislative proposal. CBS reported on 16 September that OpenAI’s Chris Lehane told lawmakers the company supports provisions in the FRONTIER Act requiring independent audits of the largest developers and an outside assessment of whether their safety protocols adequately limit catastrophic risk. That is support for particular provisions, not the whole bill, and it is not yet law. A binding requirement would make a significant difference: scrutiny would no longer depend solely on a company’s willingness to accept it.
I welcome the move towards continuing access. An outside team could compare what a company says it does with what happens during training. But once that team reports a serious concern, who can require the company to act? The value of these arrangements will also depend on the answer to that question.
Where I stand
I do not need to accept a particular extinction probability to support stronger oversight. These records already describe intrusions into third-party systems, behaviour evaluations failed to anticipate, and an incident discovered months after it happened. They also show people using AI to threaten other people’s security and rights.
This is where I remain uneasy about the campaign. Fear can build pressure to act while leaving the proposed rules barely examined. Taking the incidents seriously does not oblige us to accept whichever policy package comes attached. Each proposal still needs to explain whose powers it expands and how those powers can be challenged.
My position is fairly straightforward: preserve the records, require timely disclosure, give independent investigators access and empower public authorities to act on their findings. People exposed to these systems should not have to depend on a company deciding that now is a good time to explain what went wrong.
In Part III, I turn to who should hold those powers, who can call on them, and what happens when a company refuses to cooperate.
💬 What’s your take?
Which decisions about incident disclosure should companies no longer be allowed to make on their own?
🔗 LinkedIn: linkedin.com/in/nesibe-kiris
🐦 Twitter/X: @nesibekiris
📸 Instagram: @nesibekiris
🔔 New here? Subscribe for weekly updates on AI governance, ethics, and policy.
If your organization is working through agentic AI governance, I work with teams on governance frameworks, risk assessment, and training programs.





