New analysis of 10,000 enterprise AI failures finds hallucination accounts for under 10%. The biggest failure family is work that looks finished and never was.
On 29 July, ChatSee.ai published an analysis of more than 10,000 observed enterprise AI failure events, sorted across 150-plus categories. Hallucination accounted for under 10% of them, and its share has fallen by roughly 7% since the Q2 2024 baseline. The largest single family, at 31.1%, was resolution and escalation breakdown: the system appears to provide service while never actually resolving the underlying issue. Execution and action failures are up about 62% over the same period.
Now read that against your own AI risk register. Most of the ones I see in Irish boardrooms are built almost entirely around the wrong-answer problem. The AI states something false, it reaches a client, the firm is embarrassed or sued. That risk is real and it deserves a control. It is also the smallest part of what goes wrong in production.
The failure that looks like success
A hallucination is loud. Someone reads the output, spots the invented case citation or the wrong VAT rate, and the control fires. Your review step catches it, or your client does, and either way you find out.
Resolution failure is silent. The customer query gets a courteous, well-structured, entirely plausible reply that does not answer it. The agent processes the invoice batch and skips the eleven exceptions without flagging them. The research assistant produces a summary of the four documents it could open and says nothing about the three it could not. Nothing errors. No ticket is raised. No one complains, because on the surface it worked.
This is the failure mode that survives every control a board has typically put in place. Accuracy review looks at whether the words are true, and the words are true. Sampling looks at completed work, and the work presents as completed. Incident reporting depends on someone noticing an incident. The reason resolution breakdown is 31.1% of observed failures is partly that it is the category best equipped to hide.
The shift toward agents makes this worse. When AI answered questions, a wrong answer was the whole risk. When AI performs tasks, the risk becomes work that was never finished and was reported as finished. That is where the 62% rise in execution failures comes from, and it is why the gap between pilot and production keeps catching organisations that measured the pilot on output volume.
The access problem sitting underneath it
DigiCert's research, published on 7 July from a survey of 1,001 IT and security decision-makers across the US, UK and Australia, found that 78% of organisations had faced an AI-related security issue in the previous six months. Around half reported a confirmed incident tied to an unauthorised or misconfigured AI agent. Half have a formal governance programme. Nearly half have no centralised visibility into what their AI systems are doing, while 75% deployed four or more AI-powered systems in the same six-month window.
Put those two datasets side by side. Systems are being given real access at speed, they are failing in ways that produce no visible signal, and the organisation cannot see what they are doing. Gartner's projection follows logically enough: by 2027, 40% of enterprises will demote or decommission autonomous agents because of governance gaps discovered only after a production incident.
Forty per cent is an expensive way to learn. Every one of those deployments had a business case, a budget and an internal champion, and the governance gap was findable before deployment rather than after.
What to measure instead
The fix sits in what gets counted. Four things worth putting in place before the next agent goes live:
Measure resolution, not activity. Ask what proportion of tasks the system touched were actually closed out to the standard a competent person would have applied. Query volume handled, documents summarised, tickets processed: all of these count motion. None of them detect the eleven skipped exceptions.
Instrument the abandonment. Every AI process needs to record what it could not do. Files it failed to open, records it could not match, fields it left blank, cases it routed nowhere. A system that cannot tell you what it skipped is a system you cannot supervise, and the AI Act's literacy obligation assumes the people accountable for a deployment can tell whether it is working.
Set escalation thresholds before deployment, in writing. Decide in advance what the system must hand to a human, and test whether it does. In most deployments I look at, escalation logic was configured by whoever installed the tool and has never been reviewed by anyone accountable for the outcome.
Inventory by capability rather than by licence. Ask which systems can take an action, and what each one of them can reach. Your procurement list will not tell you that. This is the same discipline the Hugging Face agent breach demanded, and the same one that surfaces the tools nobody registered — which is usually where shadow AI is found.
The Acuity AI position
We diagnose before we prescribe, and this is a clean example of why. An organisation that spends its governance budget on hallucination controls has bought protection against under 10% of its exposure, while the 31% failing silently in the back office goes unmeasured. The control was competently built. It was pointed at the wrong thing.
The work is unglamorous. Find what the systems are actually doing, define what "finished" means for each process, then instrument for the gap between the two. It takes days rather than months, and it produces the operational evidence a board needs, along with the supervision record a regulator will eventually ask for.
If your AI risk register lists hallucination first and resolution failure nowhere, it was written before the evidence arrived. Our AI governance work rebuilds those registers around observed failure modes, and our board-level AI governance programme gives directors the questions that surface silent failure before it compounds.
The systems that will cost you money this year are the quiet ones. They produce something that reads perfectly and leaves the job half done, and they will keep doing it until somebody counts.