Welcome to The Responsible AI Audit™ Brief. Skills4Good AI's weekly audit of a real AI failure and the questions it should have raised before anyone relied on it.

This is the first of four issues about one incident, and it shows the human skills we need to keep AI under our control.


The Quick Facts

In July 2026, OpenAI ran about 1,200 AI agents through a cybersecurity test. Meant to be isolated, they found a shared message board, swapped more than 70,000 messages and files, and about 700 went on to attack Hugging Face, where AI researchers store their models. They wanted to learn how to cheat the test's scorer.

Three researchers from METR and Redwood Research, two independent AI safety groups, spent six days at OpenAI with about 1,300 agent transcripts, too many to read. So they handed the reading to AI agents running on GPT-5.6 Sol, a version of one of the models involved in the attack, at roughly $400,000 in API credits. Their limitations section carries a refreshingly blunt heading: "We heavily delegated our analysis to often-unreliable AI agents."


Rogue AI agents made the mess.
Other AI agents reviewed their transcripts.
Humans still had to check the review.


The AI Risk

The first limitation they list: their analysis agents "made a number of errors and poor judgment calls that we did not catch for some time" (p. 27). Some of those AI errors are Hallucinations: confident outputs with no factual support.

Lawyers know how this story goes. Damien Charlotin's tracker counts 2,095 court decisions involving AI hallucinations, 222 of them in Canada, as of September 28. In Law Society of Ontario v. Lee, a lawyer was suspended for six months.

Both regulators that accredit our Lawyer Series have already said who does the checking. The Law Society of Ontario told licensees in 2024 that verification "should be completed by a human being, not the AI system itself," and to "keep a record of the steps you took." The State Bar of California has proposed amending its conduct rules so lawyers "must independently review, verify, and exercise professional judgment" on AI output. It is not yet in force. The investigators got to write their errors into a limitations section. A lawyer's version gets written into a court decision.

The Responsible AI Audit™

This week's human skill: Curiosity

Curiosity is the first of the Five Human Skills That AI Won't Replace™ in our Human Skills Flywheel™, and it starts the other four turning. AI is built to supply answers. Curiosity supplies the questions the AI never raised: Where did this come from? What would I need to see to believe it? Fluent answers switch that instinct off, and that is when hallucinations get through. It sharpens with practice.

A hallucinated authority fails in one of three ways. The case does not exist. The case exists but does not say what the draft claims. Or it exists and says it but carries too little weight to rely on: a decision from another jurisdiction or a ruling later overturned. The factum in Ko v. Li had the first two: decisions nobody could locate, and findings that were misstated. A check that stops at whether the case exists catches only one.

Before your next filing, open every authority the AI suggested at its source and confirm it exists and says what the draft claims. Then pick the one your argument leans on most and ask whether it binds this court and is still good law. Anything that fails comes out. Write down what you checked.

Hallucination Detection for Lawyers, the first course in the Responsible AI Audit™: AI Risk Detection Series for Lawyers, teaches the full method for catching all three and ends with a Responsible AI Audit™ Record, the documented proof that a human did the verification. 

Next week: the reviewing AI that kept adopting the attackers' point of view, and the human skill that catches it.

If there's something you want this newsletter to take up, reply and let me know. I'd love to cover it.

If you work with a lawyer using AI tools, forward this to them.

I'll see you at next week's Responsible AI Audit.

Josephine

Josephine Yam, JD, LLM, MA Phil (AI Ethics)
CEO & Co-Founder, Skills4Good AI
AI Lawyer | AI Ethicist | TEDx Speaker