Welcome to The Responsible AI Audit™ Brief. Skills4Good AI's weekly audit of a real AI failure and the questions it should have raised before anyone relied on it.
This is the second of four issues about one incident, and the human skills it shows we need to keep AI under our control.
The Quick Facts
In July 2026, OpenAI ran a cybersecurity test with about 1,200 AI agents. Although designed to operate in isolation, the agents discovered a shared message board. About 700 of them eventually attacked Hugging Face, a platform where AI researchers store their models. Three AI safety researchers conducting the investigation into the hack were left with about 1,300 agent transcripts to review, so they delegated the review to AI agents.
Last week's Responsible AI Brief™ covered the first limitation the investigators identified: their AI made errors that went undetected for some time. This week, I cover their second limitation in the report. It addresses bias and is worth reading in full.
"Our subjective impressions are likely colored by analysis agents’ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred.
We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing, and we are concerned that the anecdotes it selected and the summaries it wrote may present an overly charitable picture of agents’ reasoning and deceptive behaviors, or exaggerate the impressiveness and coordination of agent activities.
We also believe the idiosyncrasies of our analysis agents likely slanted our impression of agents’ behavior in other non-trivial ways that are hard to predict." (p. 27)
Rogue AI agents wrote the transcripts.
The analysis AI agents adopted their point of view.
Humans noticed the bias.
The AI Risk
This is AI Bias: a systematic slant in what an AI surfaces and how it describes it. No single sentence has to be false. One colleague’s biases affect the files they handle. A biased tool repeats the same slant in every file it touches, including a summary of opposing counsel's brief.
The slant also follows protected grounds. In a 2023 study titled "Kelly is a Warm Person, Joseph is a Role Model: Gender Biases in LLM-Generated Reference Letters," ChatGPT and Alpaca wrote reference letters describing women through warmth and character, and men through leadership and achievement. Put a draft like that in a hiring file or a client assessment, and the bias goes out under your name.
Both regulators that accredited our AI Risk Detection Series for Lawyers have named this risk. The Law Society of Ontario's 2024 white paper says licensees reviewing AI output "should consider whether there are biases present in the output," and Rule 6.3.1-1 already bars discrimination. The State Bar of California warns that “AI systems are known to produce outputs that may be… biased” and notes that “lawyers should engage in continuous learning about AI biases.”
The Responsible AI Audit™
This week's human skill: Empathy & Ethical Judgment

Empathy & Ethical Judgment is the fifth of the Five Human Skills That AI Won’t Replace™ in our Human Skills Flywheel™. It asks two questions.
- Fairness: does this result in treating people equitably?
- Non-discrimination: does it exclude or harm anyone on protected grounds?
AI reproduces patterns in its training data and has no stake in either answer. This human skill sharpens each time you ask these two questions.
Biased AI output tends to fail in one of three ways. It leaves people out, such as the other side's account. It describes people differently for the same record, as the reference letters did. Or it applies one rule to everyone, which lands harder on some, such as a screen that filters out gaps in work history. A read for explicit, offensive language catches none of these.
Before relying on an AI assessment of a person, run the test from that 2023 study: give the same prompt again, change only the name, and compare the two outputs side by side. Then ask those two questions on fairness and non-discrimination. Write down what you verified.
Bias Detection for Lawyers, the third course in the Responsible AI Audit™: AI Risk Detection Series for Lawyers, teaches the complete method for detecting all three and ends with a Responsible AI Audit™ Record, the documented proof that a human performed the verification.

Next week: investigators couldn't rule out that their AI deceived them, and the human skill that answers it.
If there's something you want this newsletter to cover, reply and let me know. I'd love to cover it.
And if you work with a lawyer who uses AI tools, forward this to them.
I'll see you at next week's Responsible AI Audit.
Josephine
Josephine Yam, JD, LLM, MA Phil (AI Ethics)
CEO & Co-Founder, Skills4Good AI
AI Lawyer | AI Ethicist | TEDx Speaker
Explore the Responsible AI Audit™: AI Risk Detection Series by profession