GPTZero Found Hallucinations in Four PwC Reports After EY and KPMG Retractions
GPTZero has identified apparent AI-generated hallucinations in four PwC Middle East reports, following earlier investigations by the same firm that led EY and KPMG to retract reports containing similar fabrications. Three of the four largest professional services firms have now had published client-facing work called into question by an outside detection vendor.
The pattern matters more than any individual document. These are not internal drafts or marketing pieces. They are deliverables carrying a firm’s name, produced under engagement letters, and in some cases used as inputs to investment and policy decisions. A fabricated citation in that context is a quality control failure at the point where the firm’s entire value proposition sits, which is the assertion that a qualified professional reviewed the work before it went out.
Detection is becoming a compliance function
Until recently AI detection was discussed almost entirely as an education problem, and the criticism of it was largely justified. Classifier outputs are probabilistic, false positives fall disproportionately on non-native English writers, and no vendor has produced accuracy claims that survive independent replication at the individual-document level.
The professional services cases invert the asymmetry that made detection dangerous in classrooms. A firm accused of publishing a fabricated source holds the working papers, the source database subscriptions, and the review trail, and can refute a false positive in an afternoon. When it cannot refute it, the finding stands on the missing evidence rather than on the classifier’s confidence score. That is a materially stronger evidentiary position than any academic misconduct case, and it is why these findings have produced retractions rather than disputes.
The commercial consequence is a new buyer. Detection vendors built for universities are being pulled toward risk, compliance, and assurance teams inside the organizations that generate the documents. New York-based Pangram, which raised $9 million led by Menlo this week alongside the release of its 4.0 model and a first image detection model claiming 99.5% accuracy, is positioned for exactly that market. So is the broader category of provenance and content-authenticity tooling.
The unresolved part
Nothing about a detection finding tells you how the fabrication got into the document. A hallucinated citation can come from an unsupervised drafting tool, from a junior analyst who pasted output without checking it, or from a research process where the model was used for discovery and its suggestions were never verified against a database. The remedy differs in each case, and no classifier distinguishes between them.
Which leaves the firms with a straightforward internal problem and a harder external one. The internal problem is a control: verify every citation against a source of record before publication, a step that predates AI and was always supposed to happen. The external problem is that an outside vendor with no access to their systems found the failure three times before they did.