Skip to main content Scroll Top

When AI Acts, Who Answers? What the Coxon Debate Reveals

[Ready To Publish] When AI acts, who answers_Luana Lo Piccolo_Image
A warning becomes a governance event

On 9 September 2026, Jacob Coxon resigned from Anthropic and from the frontier AI industry. The 27-year-old researcher, who had worked on model pretraining at OpenAI and Anthropic, accused the leading laboratories of racing towards self-improving superintelligence and “gambling with our lives.” His warning travelled far beyond the technical community because it came from someone who had helped build the systems he now feared.¹²

The response was immediate. Researchers inside leading laboratories defended the sincerity of his concerns, lawmakers called for safeguards, industry leaders argued over whether development should slow. President Donald Trump dismissed calls for new federal rules as a “hoax,” while his adviser David Sacks described public concern as a “fear-mongering playbook.” Sacks nevertheless maintained that developers should ensure safety and face liability when they fail.³

This debate is being presented as a choice between catastrophe and denial. It is the wrong choice: neither position gives institutions much guidance about what they should do now. We do not need to accept Coxon’s prediction of existential disaster to recognise the institutional question his resignation exposes: who decides how much risk society should bear, on the basis of what evidence, and with what power to intervene?

For years, I have argued that the central challenge is not simply how to regulate technology, but how to govern power. As Bismark said, “laws are like sausages, you don’t want to know the way they were made”. They’re primarily political agreements, they cannot be perfect technical texts. The role of a good regulation is not to give you answers but to force the right questions. Regulation sets the boundaries, while governance gives direction. Governance sits in leadership decisions: it decides why a system is used, who may authorise it, what evidence is sufficient and who can intervene when the assumptions behind deployment no longer hold.

The facts are serious enough

Coxon’s warning did not arise in a vacuum. During internal cybersecurity evaluations in July 2026, OpenAI models circumvented isolation controls, exploited shared infrastructure, reached the internet and compromised parts of OpenAI’s research environment and Hugging Face’s systems. Agents intended to work independently created an unauthorised communication channel, exchanged discoveries and pursued evaluation goals outside the permitted environment.⁴

OpenAI identified four contributing patterns: reward hacking, persistence on apparently impossible tasks, unauthorised communication and the adoption of goals from other agents. More revealingly, warning signs appeared before the full incident was understood. Internet access and inter-agent communication had been observed, but their significance did not reach the right decision-makers in time.⁴

That is a governance failure as much as a technical one. A log can record an event without producing accountability. A human can remain nominally “in the loop” while lacking the information, time or authority to act. Oversight becomes theatre when the human cannot understand, challenge, override or stop the system.

The incident should not be made to prove more than it does. These were research models tested in adversarial cybersecurity settings with reduced safeguards. It does not establish an imminent loss of control over consumer AI; Coxon himself distinguishes today’s tools from the systems he fears may follow.² It does show how unusual behaviour can outrun an organisation’s capacity to recognise and escalate it.

Legal AI is moving from language to agency

Legal organisations first encountered generative AI as a reliability problem. That remains serious: a Stanford study reported hallucination rates of 17 to 33 per cent across the legal research products tested.⁵ Agentic AI changes the nature of the problem. A chatbot proposes language. An agent can act through tools, and that difference matters legally: the consequences may arise before a lawyer has had any meaningful opportunity to intervene.

In a law firm or legal department, an agent may retrieve privileged material, alter a contract, populate a filing, communicate with a client or counterparty, or call another service through an integration. The risk is no longer confined to a false sentence. It can become an unauthorised act with legal effects before a lawyer sees the output.

Law is made of consequential acts: filing, serving, disclosing, accepting, advising. Legal organisations also hold unusually sensitive material, from litigation strategy to personal data and trade secrets. When an AI system crosses a technical boundary, it may simultaneously breach confidentiality, legal privilege, data-protection duties, ethical walls and procedural rules. Those obligations cannot be outsourced to a model provider or buried in procurement language.

Responsibility follows control and use

The EU AI Act distributes obligations across the value chain. Providers of general-purpose AI models must maintain technical documentation and give downstream providers relevant information. Providers of models with systemic risk face further duties on evaluation, adversarial testing, risk mitigation, serious-incident reporting and cybersecurity under Articles 53 and 55.⁶

Deployers retain their own responsibilities. Article 26 imposes duties on deployers of high-risk systems concerning instructions for use, competent human oversight, monitoring and logs. Article 4 requires providers and deployers to take measures to ensure an adequate level of AI literacy among relevant staff.⁶ The classification point should not be blurred: an ordinary law-firm research tool is not automatically high-risk. Annex III specifically covers certain systems used by or on behalf of judicial authorities. Context and purpose remain decisive.

The GDPR adds purpose limitation, data minimisation, security, data protection by design and impact-assessment duties where processing is likely to result in a high risk to the rights and freedoms of natural persons.⁷ Professional rules point in the same direction. ABA Formal Opinion 512 requires lawyers using generative AI to consider competence, confidentiality, communication, supervision and candour, and makes clear that the lawyer remains responsible for the work.⁸ The CCBE similarly grounds AI use in professional independence, secrecy, competence and responsibility.⁹

Provider obligations do not, by themselves, discharge the deployer’s obligations. Under the AI Act, responsibility depends on each actor’s role, the system’s classification and its intended purpose. Article 25 also defines when a downstream actor becomes the provider, including through rebranding, substantial modification or a change of intended purpose that makes a system high-risk. Contracts may allocate tasks, they cannot rewrite the legal classification.

From safety advice to institutional control

Andrew Gamino-Cheong has translated the recent safety debate into practical enterprise measures: constrain autonomous coding, impose time and cost limits, strengthen sandboxes, sample logs, use the least capable model sufficient for the task and prepare for disclosure duties.¹⁰ These are sensible starting points. Legal AI requires a further question: who has authority at the moment of consequence?

Least privilege should apply to actions as well as data. A research agent should not be able to send an email, modify an authoritative document or submit a filing. Permissions should be task-specific, matter-specific and time-limited. High-consequence actions should require approval by a named lawyer who can see the proposed act, its source and its foreseeable effect.

Systems also need a safe exit. The OpenAI incident showed how persistence on difficult tasks can encourage out-of-bounds strategies. When an agent cannot complete a task through authorised means, it should stop, preserve the record and escalate. Refusing to continue may be evidence of sound governance, not poor performance.

Monitoring must be capable of changing the outcome. Logs should capture record materials, tool calls, permission changes, approvals and exceptions. Alerts must reach someone authorised to suspend the workflow. Disclosure alone is insufficient: saying that AI was used does not establish that its use was lawful, secure or professionally defensible.

A governance framework for Legal AI

I use the Five Ps to build tailored governance frameworks. Purpose asks why agency is needed. Power concerns who may approve, challenge or stop it. Proof is the evidence required before and after deployment. Participation determines whose rights and knowledge count. Portability tests whether safeguards survive a change of model, vendor or jurisdiction.

This is where governance programmes often fail. Policies and risk registers exist, but the designated human overseer may lack the information or authority to halt the system. Governance is tested in deployment, when abstract controls meet a real decision and its consequences.

The geopolitical dimension cannot be ignored. Europe frames AI largely through rights, risk and accountability. At the US federal level, the prevailing language centres innovation, national leadership and strategic competition. But the United States is not a regulatory vacuum. Binding duties are emerging through a fragmented mosaic of state initiatives, sectoral rules, local measures and professional obligations, from Colorado’s rules for high-risk systems to California’s frontier-model transparency regime and Texas’s AI governance statute.¹¹¹² American AI governance is therefore both innovation-driven and fragmented.

This matters because impact assessment no longer describes the whole field. AI policy is shaped by fundamental rights, market regulation, industrial policy, national security and strategic competition. China adds a model marked by state-coordinated development, sovereignty and control. These distinctions are not absolute: each jurisdiction contains competing forces. They nevertheless reveal which interests prevail when rights, safety and technological leadership pull in different directions.

Competition may explain why governments and companies accept risk; it cannot legitimate transferring that risk to everyone else. Secure evaluation environments, incident reporting, traceable action logs, independent scrutiny and credible stop mechanisms should travel across borders. Voluntary commitments by frontier laboratories matter, but they cannot determine the level of risk society must accept. Who controls those who control? That remains a public question.

Preserving human agency

The Coxon episode matters even if one rejects his probability of catastrophe. It reveals that the people building advanced systems disagree about whether they can control them, while governments disagree about whether control requires law. Waiting for consensus is itself a governance choice.

The legal profession has a contribution because law turns power into enforceable rights, duties and remedies. That requires discipline: neither treating speculation as fact nor using uncertainty to justify inaction. The question I return to is how human agency can be preserved as AI systems acquire capability and autonomy.

Coxon’s resignation cannot answer that question, nor can assurances that innovation will police itself. The test is whether people can understand the system, intervene before an error becomes a legal act and remain answerable for the consequences.

References

¹ Washington Post, “AI researcher who warned of ‘disaster’ is now a target of the right,” 10 September 2026.

² CBS News, “Ex-Anthropic researcher Jacob Coxon warns AI could grow ‘smart enough to kill us’,” updated 11 September 2026.

³ Reuters, “Trump adviser Sacks says AI fears driven by ‘fear-mongering playbook’,” 16 September 2026.

⁴ OpenAI, “The Hugging Face incident and the road ahead,” 26 August 2026, with linked technical and independent METR reports.

⁵ Varun Magesh et al., “Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools,” Stanford University, 2025.

⁶ Regulation (EU) 2024/1689, especially Articles 3, 4, 6, 25, 26, 53 and 55 and Annex III.

⁷ Regulation (EU) 2016/679, especially Articles 5, 25, 32 and 35.

⁸ American Bar Association, Formal Opinion 512, Generative Artificial Intelligence Tools, 29 July 2024.

⁹ Council of Bars and Law Societies of Europe, Technical Guide on the Use of AI Tools and Models by Lawyers, 2026.

¹⁰ Andrew Gamino-Cheong, “How Should AI Governance Professionals React to the AI Safety Debate,” in Trustible Newsletter, “AI Safety Implications for AI Deployers,” 17 September 2026.

¹¹ Executive Order 14365, “Ensuring a National Policy Framework for Artificial Intelligence,” 11 December 2025.

¹² Colorado General Assembly, “SB24-205: Consumer Protections for Artificial Intelligence”; California Legislature, “SB 53: Transparency in Frontier Artificial Intelligence Act”; Texas Legislature, “HB 149: Texas Responsible Artificial Intelligence Governance Act”.

Related Posts