AI Detectors Are Mostly Junk. This Time the Detector Was Right and the Auditor Was Wrong.
I’ve spent two years arguing that AI detection tools are close to worthless. I still think the studies showing they flag non-native English speakers at higher rates are damning, and I still think any university disciplining a student on the strength of a probability score is committing an injustice it will eventually have to apologize for. None of that has changed.
But there’s a category of case where the objection collapses, and we just watched a clean example of it. GPTZero found apparent AI-generated hallucinations in four PwC Middle East reports. Its earlier work led EY and KPMG to retract reports containing the same kind of fabrication. That’s three of the Big Four, found by the same outside party, using the tool everyone in my camp has been calling snake oil.
The detector was right. Not partly right. Not accidentally right. Right, three firms running, on documents that carried a firm’s name and a partner’s sign-off.
Ask who was harmed and the objection disappears
The case against these tools has always been about the cost of a false positive. A student flagged wrongly loses a grade, a scholarship, sometimes a degree, and has no way to prove a negative. That asymmetry is why I’ve been hostile. The person accused is powerless and the accusation is unfalsifiable.
Now run the same test on PwC. What happens to a global firm when a detection startup wrongly claims its report contains a fabricated citation? Its quality control team pulls the source, finds the source, publishes the source, and the startup’s credibility takes the damage. The firm has the documents. The firm has the working papers. The firm has a professional obligation to have kept them.
There’s no asymmetry to protect here. The accused party is the one holding all the evidence, which means a false positive gets refuted in an afternoon.
And the true positives keep landing.
The client paid for judgment and received autocomplete
Here’s what makes this worse than any student’s essay. A machine shop in Sharjah commissioning a market entry study, or a sovereign fund taking a sector report into an investment committee, isn’t buying prose. It’s buying the assertion that a named professional stands behind the numbers. That’s the entire product. The billing rate exists because a person with a license and a liability exposure reviewed the work.
A fabricated citation in that document means nobody performed the review. Not that a machine helped — that no human checked the machine. There’s no version of professional skepticism, the standard the firms themselves wrote and teach and bill against, that survives a hallucinated source making it into a client deliverable.
The firms’ own standards convict them. I don’t need a detection tool’s confidence score to make that argument. The retractions at EY and KPMG already did.
What I still won’t accept
I’m not converting. I don’t want these tools anywhere near a classroom, a disciplinary hearing, or a hiring pipeline, and I don’t believe the accuracy claims that vendors put in press releases, including the 99.5% figure Pangram published this week for image detection. Detection remains a probabilistic guess dressed as a finding, and in the hands of an institution with power over an individual it’s a weapon.
The difference is who’s holding it. Pointed at a student, it’s an accusation nobody can rebut. Pointed at a firm that charges for assurance, it’s a request to produce the working papers.
One of those is an injustice. The other is an audit.
The Big Four just failed theirs.