This month, the AI safety debate kept returning to who should decide how fast the technology develops. Lab leaders offered oversight proposals. Politicians argued over regulation. Critics asked whether the companies calling for rules were trying to protect their market position. Following these exchanges, I was less interested in who had changed sides than in how much control the companies were prepared to give up.
I am not giving AI companies credit for welcoming independent oversight until they explain who chooses the evaluator, who pays, and what the public gets to see. Can an evaluator publish an uncomfortable finding without losing access or funding? Who can require the company to act on it? Those details matter when a company that welcomed scrutiny starts disagreeing with the results.
I read this month’s proposals with those questions in mind. I also looked for something that gets less attention in negotiations between labs and evaluators: whether people harmed by these systems could request an investigation, obtain relevant evidence or challenge a decision. From where I sit, working on AI governance, an oversight arrangement has to explain what those people can do when the company’s answer is not good enough.
Part III of a three-part series: Fear, Evidence, and Frontier AI Governance.
Part II asked who controls the evidence about AI safety incidents. This piece asks who gets to decide what the evidence means, and what follows when the answer is inconvenient.
Who Chooses and Funds AI Safety Evaluators?
An evaluator chosen and funded by the company it examines can still do serious work. The harder question is whether its contract and funding let it continue after publishing findings the company disputes. On 12 September, Dario Amodei committed Anthropic to bringing external evaluators inside the company. The company would still choose whom to invite and negotiate the terms.
Those negotiations already cause problems. FAR.AI’s Adam Gleave told TechCrunch his organisation had rejected contracts because developers wanted too much control over evaluations. Walking away protects an evaluator’s independence, but does not secure the access needed to investigate.
On 18 September, the AI Evaluator Forum published its minimum conditions. These include freedom from company ownership or governance, no payments tied to findings, protection against retaliatory lawsuits, and funding arrangements that survive an unwelcome result. That funding protection is the condition I would check first. An evaluator may have the right to publish but little confidence that it can continue the work afterwards.
Demis Hassabis’s proposal would create a standards body under federal oversight. The proposed FRONTIER Act would require the largest covered developers to retain publicly licensed evaluators. In Hassabis’s model, the appointment process needs explaining. Under the bill, public licensing would limit whom a developer could hire, while leaving the developer to choose its evaluator.
Support can also come without a payment. METR’s August funding statement says it does not accept funding from frontier AI companies or donations made by or at the direction of their staff, but receives substantial free model tokens for its work. Could an investigation continue if that support disappeared? Nothing in the statement suggests it has been withdrawn to influence findings, but it does identify a dependency that funding and access arrangements should address. Independence requirements should cover donated access as well as direct funding.
Money also leaves part of the problem untouched. An evaluator can be financially independent while sharing a company’s assumptions about which risks matter and whose experience counts as evidence. The Forum calls for multiple evaluators with different expertise and room to disagree. That condition belongs in the agreements too. Otherwise, we could end up funding independent organisations to keep asking the same narrow set of questions.
What Can Independent AI Evaluators Access?
Amodei’s proposal would give evaluators access to internal tools, staff and training processes, and the right to publish key findings. Anthropic could redact certain protected information, but not suppress findings simply because they were unfavourable. Evaluators could report significant redactions and denied access. I would require them to disclose material restrictions. Readers should not have to guess what an assessment could not examine.
Section 5(e) of the proposed FRONTIER Act would require access to reasonably necessary unredacted records, personnel and systems, with narrowly tailored security and confidentiality conditions. Material limits on scope or access would have to appear in the assessment report. That disclosure requirement is worth keeping through any negotiations over the bill.
Scope matters too. The AI Evaluator Forum’s letter also puts company practices within the investigation’s scope. This matters because, as Part II showed, permissions, disabled safeguards and test environments helped shape the incidents. An investigation that stops at what the model did lets the company’s decisions escape scrutiny.
The EU AI Act gives the Commission powers over general-purpose AI models to demand information, conduct evaluations and require corrective measures under specified conditions. But these are powers exercised by a public authority. They give neither researchers a standing right to inspect models nor affected people the ability to compel a particular evaluation or obtain its records.
An 8 September joint advisory led by the NSA recommends altering responses or serving downgraded models to counter malicious distillation without alerting suspected users, while informing safety researchers and evaluators of model changes. That notification should be an enforceable duty. I raised the same problem in enterprise AI assessments: a provider can change the model behind a service while the customer keeps relying on the old assessment.
Amodei proposes continuing access. FRONTIER Act sections 5(h) and 5(i) would allow additional assessments and require supplemental reporting within seven days of an evaluator determining that specified earlier findings or assurances no longer hold. These arrangements need to connect: evaluators must hear about changes, retain access to investigate them and update their findings when necessary.
What Would Make OpenAI, Anthropic, Google DeepMind or Meta Stop a Release?
An evaluator can identify a dangerous model. Who can stop its release? This is where I would put pressure on the proposals.
Demis Hassabis describes a clear stopping point: voluntary prerelease assessments could become a requirement for deployment in the US once the assessment process proved effective.
Dario Amodei supports regulation covering all US frontier developers, alongside permanent evaluators and safety requirements tied to particular capabilities. That goes beyond voluntary cooperation. I would judge those rules by who could overrule Anthropic, on what grounds, and whether Anthropic would have to comply.
OpenAI’s 21 September paper proposes shared standards for capability measurement, risk assessment, automated AI research and incident reporting. Governments would decide how to incorporate them into law. The paper links to Hassabis’s framework without adopting his deployment condition. OpenAI is specific about what these standards would not create: licensing or mandatory prerelease approval. I would like equal clarity about which consequences it would support when a model fails.
Elon Musk proposes competitors testing one another’s models, with Chinese participation and reputational and legal pressure on companies that ignore warnings. Rivals may be excellent at spotting weaknesses. The proposal still needs to explain what happens when a company disputes a rival’s finding and proceeds anyway.
Mark Zuckerberg points to Meta’s decision to delay Muse as evidence that labs can set a safer pace themselves. I am not prepared to treat a delay announced by Meta as proof that Meta’s judgment is sufficient. Outsiders still need the evidence behind the decision to proceed. Alexandr Wang supports external evaluators and independent oversight of launches. Meta should explain whether that oversight could prevent a launch its leadership wanted.
Microsoft’s draft Code of Conduct sets specific behavioural limits, including no resistance to shutdown or tampering with reasoning records. Microsoft says it is not yet training models on the draft; the revised version is intended to guide development from 2027. These are commitments to test, not safety results to credit. The draft describes evaluation work and outside expertise, but leaves unanswered who could require a response when a model broke the rules.
Satya Nadella calls for broader representation across countries, disciplines and institutions. I agree with him. The next question is what those representatives would be able to decide. Microsoft’s public consultation and OpenAI’s proposed consultations offer opportunities to comment; neither, by itself, gives participants decision-making power. As I asked in AGI, Built to Benefit Whom?, which decisions would these companies accept they could no longer make alone?
The proposed FRONTIER Act supplies a more concrete route. Its IVO requirement covers developers whose combined figures with affiliates over the preceding 36 months exceed $5 billion in gross revenue and reach at least $10 billion in AI-related development spending. Assessments would be required at least every six months. If an IVO determined that inadequate catastrophic-risk mitigation posed an imminent catastrophic risk, it would have to refer the matter to the Commerce Secretary as soon as practicable, within 72 hours.
The referral would be mandatory; an intervention would not be automatic. The Secretary could restrict or suspend activities under the emergency powers. That creates a route from an evaluator’s finding to government action, but leaves the decision to act with the Secretary. I would watch that discretion closely, including what explanation the public would receive if a serious referral produced no intervention.
Who Can See and Challenge AI Safety Findings?
This is where my support for mandatory evaluation runs into the bill’s limits. Under the proposed FRONTIER Act, the regulator would receive the full assessment. The public would receive a company-redacted version within 30 days. Redactions would need stated reasons, but only designated officials could trigger the bill’s review process. Reports and supporting materials would also be exempt from federal freedom-of-information disclosure.
The bill would give licensed evaluators broad immunity for losses arising from catastrophic risks in models they assessed. Its narrow exception covers death or serious physical injury caused by willful misconduct, with a demanding proof standard. Someone bringing a claim could struggle to obtain the evidence while the evaluator benefits from broad immunity. That combination needs defending, not burying in the legal detail.
There is a separate question about whether the evidence survives. The bill would preserve assessment reports and supporting materials for at least five years. That does not necessarily preserve everything needed to reconstruct an incident.
Christopher David LaRoche makes this distinction in Lawfare: investigators need the relevant model versions, logs, inputs, outputs and operating conditions, alongside the expertise to interpret them. I would require incident-triggered preservation. A timely report is little help if nobody can reconstruct what happened.
The Nightingale Collective’s DseWiki investigation shows what outsiders can contribute. Researchers reconstructed roughly 18,000 agent posts from public edit histories and attributed the activity to OpenAI agents. Their analysis could be examined because those traces were publicly available. Access to evidence should not depend on an operation leaving records on someone else’s server.
Commercial confidentiality, privacy and security justify protecting some information. They do not settle who should be able to challenge a redaction or obtain evidence of harm. A law that compels independent scrutiny should also give people a usable route to question the resulting assessment.
Who Must Respond When People Report AI Harm?
The proposed FRONTIER Act would let members of the public report critical safety incidents. But the Under Secretary would have to review developers’ reports and could choose whether to review everyone else’s. Submissions would need an incident date and an explanation of why the event qualifies. The person outside the company may have the least information and would receive the weaker guarantee of attention.
The channel’s scope creates another problem. Its incident categories concern model security, loss of control, catastrophic harm and certain deceptive behaviour. Political surveillance or manipulation would not automatically qualify.
Consider the cases in Anthropic’s threat report: employment contracts demanding political loyalty in the Central African Republic, and an influence operation targeting Malaysian constituencies through fake accounts. These records do not prove that every recommendation was implemented or that voters were persuaded. They do show uses that deserve investigation without an extinction forecast attached. People should not have to translate a rights violation into a catastrophic-risk claim to get an authority’s attention.
Existing privacy, civil-rights and consumer-protection routes still matter, wherever their jurisdiction and scope apply. The question is how this new regime would work alongside them. Section 9 would restrict new state obligations on developers in specified transparency, auditing and incident-reporting areas, while preserving listed exceptions. I would not accept limits on alternative oversight without a convincing account of what people could demand from the federal system instead.
I raised a similar problem in Why Your AI Policy Fails Before Anyone Enforces It: responsibility spreads across organisations, and the person affected struggles to find anyone required to respond. A public reporting channel should come with a duty to assess credible complaints and explain the decision, including where to take a concern that falls outside its remit.
What Will Governments Require of Frontier AI Companies?
Türkiye’s foreign minister signed the 21 September call for control of frontier AI models, alongside leaders from countries including Kenya, Singapore and South Africa, and the European Commission president. It calls for mandatory testing, independent evaluation and sufficient access for evaluators, and proposes exploring an international institution. The United States and China were absent from the published list.
My government has now endorsed demands I support. I would like to see what it intends to do at home. The statement gives someone in Türkiye no new way to obtain evidence or challenge an unsafe deployment. Its value will depend partly on whether signatories turn their diplomatic position into domestic obligations.
Trump’s 14 September response presented a strong president as the only AI guardrail needed. Personal confidence is a poor basis for public oversight. The arrangements discussed here need to work when political leaders dismiss a warning, as well as when companies do.
Beijing’s response to Amodei criticised fear-mongering and confrontation. That should not be confused with rejecting safety governance: China published its AI Safety Governance Framework 3.0 the same week. Treating China as a reason to abandon safeguards avoids examining what either government is actually prepared to enforce.
China’s rules for human-like AI interaction services, for example, require options to copy and delete interaction data, including chat histories. Those are concrete protections worth examining. They do not, by themselves, give users access to the technical evidence behind a disputed safety claim.
Countries outside the leading labs’ home markets have reasons to demand a say. Their workers, researchers and voters will live with decisions made elsewhere. From Türkiye, I expect our participation to include demands for evidence and remedies, and an explanation of which domestic authority will pursue them. A signature gives us a commitment against which to judge that work.
What I Expect From AI Oversight
I support binding public oversight of frontier AI, and I will judge it by what happens when a company refuses to cooperate.
The FRONTIER Act would give officials powers that company promises cannot. Its limits on public access, evaluator liability and state regulation still need challenging. Supporting regulation does not oblige us to accept those terms.
I began this series welcoming the attention AI safety was finally receiving and questioning how much of it depended on fear. The evidence gives us enough reason to act without accepting every policy offered in its name. Before I call an oversight regime a success, I need to see an evaluator retain access after an unwelcome finding, an authority require action over a company’s objection, and an affected person get a complaint examined.
I do not expect a statute to solve alignment. I expect it to stop leaving companies with the final say over risks the rest of us have to live with.
💬 What’s your take?
Which decisions about incident disclosure should companies no longer be allowed to make on their own?
🔗 LinkedIn: linkedin.com/in/nesibe-kiris
🐦 Twitter/X: @nesibekiris
📸 Instagram: @nesibekiris
🔔 New here? Subscribe for weekly updates on AI governance, ethics, and policy.
If your organization is working through agentic AI governance, I work with teams on governance frameworks, risk assessment, and training programs.
What makes an AI safety evaluation independent?
An outside organisation is not automatically independent. Its funding, access and publication rights must survive findings the company dislikes. I would also examine who sets the questions: financial independence means little if evaluators cannot investigate the decisions and risks they consider important.
What are embedded AI safety evaluators?
Embedded evaluators work with continuing access to a developer’s internal systems, staff and training processes. That can reveal problems a finished-model test misses. The agreement still matters: what can they inspect, what can they publish, and can the company end their access after an unwelcome finding?
Can an AI safety evaluator stop a model’s release?
An evaluator’s finding does not, by itself, prevent a release. That requires a binding arrangement or legal authority to restrict deployment. Every oversight proposal should explain who can act on a failed assessment and what happens if the developer disputes it.
Should AI safety reports be public?
The public should be able to examine findings and understand significant limits on an assessment. Privacy and security can justify withholding some details. They do not justify leaving the company with an unchallengeable decision about what everyone else gets to see.
How should people affected by AI challenge a safety decision?
They need a route to submit evidence, have a credible complaint assessed and receive a reasoned response. Where a concern falls outside a safety body’s remit, there should be a clear referral route. Publishing a report does not give an affected person any of those rights by itself.





