All postsForward-Deployed Model

KPMG, EY and PwC all published AI-hallucinated reports in 2026. What it means for who builds your intelligence layer.

The short answer

Three of the four largest professional services firms published research in 2026 that turned out to contain fabricated citations, and the interesting part is not that it happened but that it could not be caught before publication. KPMG withdrew an agentic AI report in June 2026 after an outside review found that only a handful of its citations pointed correctly to the source they named, and organisations including UBS, the NHS, Swiss Federal Railways and Transport for London said claims about their AI usage were untrue or misleading (TechCrunch, GPTZero).

The structural point for a buyer is this: when the thing you purchase is a finished document produced somewhere you cannot see, there is no mechanism by which you could have caught the error either.

Key takeaways

  • Three separate 2026 incidents, three different firms. KPMG in June, EY Canada in May, PwC Middle East in July. Deloitte’s own case, a partial refund to the Australian government over invented citations, is from 2025 and is not a fourth 2026 event.
  • The firms involved sell AI governance. Each of these reports was research about responsible or agentic AI, published by a firm that advises enterprises on exactly that.
  • The failure is in the delivery model, not the firm. A deck produced away from the client, handed over finished, has no check available to the client at the point it matters.
  • Regulation does not close this gap. EU AI Act Article 50 transparency duties apply from 2 August 2026 and govern disclosure, not accuracy. A labelled report with fabricated footnotes is compliant and still wrong.
  • Carry one question into the next vendor meeting. Who, or what system, verified every citation and case study in this, and can I see how.

What happened, in order

In May 2026, EY Canada withdrew a cybersecurity report on fraud in loyalty programmes after researchers found a large share of its cited sources were fabricated, misattributed, or pointed at pages that no longer resolve. Among them was a market figure attributed to a McKinsey report that does not appear in any McKinsey publication. EY Canada removed the document and said it was reviewing how it had been published (Computing, Information Age).

In June 2026, KPMG pulled an October 2025 report on agentic AI. A forensic review found that only five of its forty-five citations pointed correctly at the source named; the rest ranged from mangled to partially invented. Separately, several of the organisations whose AI programmes the report described told the press that the descriptions were wrong. KPMG removed the report from its sites pending an internal investigation.

In July 2026, four PwC Middle East thought-leadership reports were flagged in the same way. The footnotes included citations to pages that do not contain the evidence claimed, an academic paper on air quality in Riyadh with no trace in the journal or under the authors named, and a source URL carrying a utm_source=chatgpt.com parameter, which is the artefact of a link copied out of a chat session rather than out of the publication (The Irish Times, GPTZero).

Deloitte belongs in the timeline but not in the 2026 count. In October 2025 it agreed a partial refund to Australia’s Department of Employment and Workplace Relations over a report on the welfare compliance system that contained a fabricated quotation from a court judgment and references to research papers that do not exist (CFO Dive). It is real, it is the earliest of the four, and it is why the PwC finding was written up as completing a sweep. It is not new evidence.

The part that should bother a buyer

Each of these documents was research about AI. Two were specifically about how enterprises should adopt and govern it. The firms publishing them sell AI governance advisory to the same audience the reports were written for.

That is not a competence story and it is not worth treating as one. Every one of these firms employs people who are extremely good at this work. The interesting question is how an organisation with that much review capacity ships a document in which a footnote resolves to a chat URL, and the answer is not carelessness. It is that nothing in the production path was positioned to catch it.

A research report is produced by a team, checked by reviewers who read the prose rather than re-derive the sources, approved by a partner accountable for the argument rather than the citation list, and published. At no point does a system hold the claims against the material they came from, because there is no such system in the loop. The document is the artefact. The evidence behind it exists as links in a PDF, and a link in a PDF is a claim about a source, not a connection to one.

Then the same shape of artefact gets handed to a client. A deck, a model, a set of recommendations, billed by the day, with the risk on the client’s side. The client reads the argument. The client cannot re-derive the evidence either, for the same reason the reviewers could not: the material lives somewhere the client has no access to, and checking it would cost more than the engagement.

Why an owned system fails differently

A context engine built inside a company’s own stack fails in plenty of ways, and we are not claiming otherwise. It fails on bad source data, on reconciliation between systems that disagree about the same customer, on freshness when an upstream feed goes quiet. Those are real and they are what the work consists of.

What it does not do is fail silently in this particular way, because provenance is not a footnote in it. When an agent produces an answer from an owned intelligence layer, the answer carries where it came from: which system, which record, when it was last updated. The check is mechanical and it runs on the client’s own infrastructure, on the client’s own data, which is the only place the check can actually be performed. You are not trusting an assurance about the verification process. You are looking at the verification.

That difference is the whole argument for building inside rather than buying finished. It is not that engineers embedded in your stack are more careful than partners at a global firm. It is that the artefact they leave behind is a system you can interrogate, and a document is not.

There is a second-order version of the same point. A consulting output is produced once and then decays, and nobody updates it. An owned layer is continuously wrong in ways you can see and fix, which is a better position than being periodically wrong in ways nobody surfaces until an outside researcher runs a citation check on your vendor.

What the EU AI Act does and does not fix

European buyers reaching for regulation here will be disappointed, and it is worth being precise about why.

EU AI Act Article 50 transparency obligations become applicable on 2 August 2026. They require providers of generative systems to mark synthetic output in a machine-readable form and make it detectable as AI-generated, and they require deployers to disclose in defined situations, including certain published text on matters of public interest (EU AI Act, Article 50).

Read the obligation carefully and it is a labelling duty, not an accuracy duty. A consultancy that generated a report with a model, marked it correctly and disclosed the fact would have satisfied Article 50 while publishing every fabricated citation the researchers found. Transparency about how a document was made tells you nothing about whether the sources in it exist.

This is worth saying plainly because the compliance conversation and the diligence conversation are being conflated, and a buyer who treats an AI-use disclosure as an accuracy guarantee has bought nothing. The transparency regime is a floor for the market. It is not a check on your vendor.

One question for the next vendor conversation

Ask it of us too:

Who, or what system, verified every citation, case study and number in what you are about to hand me, and can I see how.

A vendor with a real answer describes a mechanism: the sources are resolved automatically, the claims are held against the systems they came from, here is the trace, here is what it looks like when it fails. A vendor without one describes a policy: rigorous quality assurance, senior review, a commitment to the responsible use of AI. Three firms issued statements in the second register this year, after the fact.

The reason this question works as diligence is that it cannot be answered well by a firm whose deliverable is a document. The mechanism does not exist there, not because anyone chose to skip it, but because the delivery model has nowhere to put it.

What this does not mean

It does not mean these firms are finished, that consulting has no place, or that anyone should feel good about it. Strategy work, organisational change and regulatory interpretation are genuinely well served by a firm with deep bench and real institutional memory, and a company that needs a defensible external view on a board decision should buy one.

It means something narrower. When the work is building the intelligence a company will run on, the deliverable has to be a system inside that company, auditable by the people who own it, because that is the only arrangement in which the buyer can check anything at all. Three separate incidents in five months, at three firms that sell AI governance, is a reasonable amount of evidence for a claim that otherwise sounds like positioning.

FAQ

Which Big Four firms published AI-hallucinated reports in 2026? KPMG, EY Canada and PwC Middle East. KPMG withdrew an agentic AI report in June 2026, EY Canada withdrew a cybersecurity report in May 2026, and four PwC Middle East thought-leadership reports were flagged in July 2026. Deloitte’s comparable incident, a partial refund to the Australian government, dates from 2025 (GPTZero).

Does this mean the Big Four are bad at AI? No, and that is the wrong conclusion. The failure is in the delivery model. When the deliverable is a document produced away from the client and handed over finished, nobody at the client can hold the claims against the material they came from, so an unverified citation survives to publication and then survives the client’s review too.

How is a forward-deployed model different? The engineers build inside the client’s own stack, against the client’s own systems. The check on a claim is the system that produced it. Provenance is a property of the build, not an assurance in a cover letter.

Does the EU AI Act stop this? Not in the way buyers hope. Article 50 applies from 2 August 2026 and governs disclosure and machine-readable marking of AI-generated content, not the accuracy of research. A correctly labelled report with invented citations is compliant and still wrong.

What should I ask a vendor because of this? Who, or what system, verified every citation, case study and number in what you are about to hand me, and can I see how. A firm with a mechanism will show you the mechanism. A firm without one will read you its quality-assurance policy.



Nucleo builds the intelligence layer inside your stack, model-agnostic and governed to EU rules, so the evidence behind an answer is something you can inspect rather than something you are told about. Talk to our engineers.

The future belongs to those who see it before it is obvious, and build it while everyone else is still out there scouting for the best AI tools.