Acuity AI Advisory
← Insights
·5 min read

Four AI Services Failed Together. Your Fallback Was One of Them.

G

Ger Perdisatt

Founder, Acuity AI Advisory

On 3 September four major AI services degraded within ninety minutes of each other. The overlap exposed a resilience assumption most Irish firms have never actually tested.

Last Thursday, 3 September, between roughly 2:26pm and 6pm Irish time, four of the largest AI services degraded at once. Anthropic opened an incident on Claude at 9:26am Eastern covering the web application, the API and multiple models, and closed it just under three hours later. Grok started failing four minutes after that, across web, mobile, plugins and API endpoints in two US regions, with traffic back to normal after 1pm Eastern. OpenAI opened at 10:43am Eastern across fifteen ChatGPT components and four Codex ones, applied a mitigation thirty-four minutes later and declared resolution at 12:55pm. Independent monitoring picked up elevated failures on Gemini between about 10:45 and 11:15. Google never posted an incident report at all.

Taken one at a time, these were unremarkable. Services go down. What matters is that they went down together, inside a ninety-minute window, in the middle of the Irish working afternoon. Most organisations answer the AI vendor risk question with some version of "we have more than one provider". Thursday was the live test of that answer, and it came back badly.

A second vendor is not a second system

Four separate companies wobbling inside ninety minutes is a straightforward consequence of how the stack is built. Model providers run on a small number of hyperscale clouds. They resolve through a handful of DNS operators, sit behind two or three content delivery networks, and authenticate through the same identity providers. You can buy from four vendors and still be buying one dependency chain in four wrappers.

A fallback only earns the name if the alternative fails independently. That is the entire point of having one. In practice almost nobody checks. The procurement question that gets asked is "who else could supply this?" The question that determines whether you stay operational is "what do they both sit on?" Those have different answers, and only the second one shows up during an outage.

Firms that spent 2019 mapping their cloud dependencies have mostly not repeated the exercise for the AI layer, because the AI layer arrived as a feature inside software they already owned rather than as a procurement event with a risk assessment attached.

Financial entities already signed up to this

For anyone in scope of DORA, this is not an abstract resilience discussion. It is an existing obligation with a filing date.

The European Supervisory Authorities published the first list of designated critical ICT third-party providers on 18 November 2025. Nineteen names, including Amazon Web Services EMEA, Google Cloud EMEA and Microsoft Ireland Operations. DORA requires a register of information covering every ICT third-party arrangement, active monitoring of concentration risk, and a demonstrable exit strategy for each material dependency. The register goes to the competent authority by 31 March.

The gap I see repeatedly in Irish registers is that the AI layer is recorded as a feature of an existing contract rather than as an ICT service in its own right. A copilot bundled into a licence renewal, a summarisation tool inside a case management platform, a model API called by an internal agent that somebody built in a fortnight. None of it went through the third-party risk process, because none of it looked like a supplier.

Test it against Thursday. If a drafting assistant, a client-facing chatbot or a model call sitting inside a claims workflow is unavailable for three hours, that is an ICT-related disruption. Whether it appears in your register is now a supervisory question rather than an internal one.

For everyone else, the cost is real and unmeasured

Outside regulated financial services the picture is quieter and more expensive. A Dublin law firm or a mid-sized accountancy practice does not stop when the AI stops. People sigh, open a blank document and work the way they worked in 2023. Nobody logs it, nobody costs it, and the loss disappears into a slightly slower afternoon.

That holds right up until AI moves from assistance into the critical path. Client onboarding checks that now depend on automated document extraction. A first-line query handler that answers clients directly. Bundle review the day before a hearing. Once the tool is load-bearing, a three-hour outage stops being an inconvenience and becomes a service failure with a client on the other end of it.

The second-order problem is worse. Manual fallback paths decay. Six months after a process is automated, the person who knew how to do it by hand has moved on, the checklist is out of date, and the answer to "can we do this the old way for an afternoon?" turns out to be no. That degradation is silent and nobody owns it.

Four things to check this month

Write down which business processes stop when the AI stops. Processes, not systems. If nobody can answer that in a meeting, you have found the real gap and it is not a technical one.

Map the layer underneath your providers. Which cloud, which CDN, which identity provider. If your primary and your fallback share two of the three, you have a single supplier with extra invoices.

Test the fallback for real, on a working day. Not a tabletop exercise. Cut over, run live work through the alternative for an hour, and see what breaks. Untested failover has a poor record of working first time.

Keep one manual path alive per critical process, and date it. Someone should walk it end to end twice a year. If the last person who did it has left, that control is already gone.

The Acuity AI position

We diagnose before we prescribe, and outages are a useful diagnostic because they are honest. A vendor questionnaire tells you what a supplier says about its resilience. Three hours of downtime on a Thursday afternoon tells you what your own organisation actually does without the tool, which is a different and far more useful piece of information.

The organisations that came through 3 September well were the ones who knew which of their processes were genuinely dependent, and had a way of working for an afternoon without them. That knowledge costs very little to acquire and it does not require a platform decision or a new framework.

We work with Irish boards and leadership teams on exactly that mapping: what is deployed, what depends on it, what sits underneath it, and what happens on the afternoon it is unavailable. Vendor-neutral, evidence first.

operational resiliencedorathird party riskboard advisoryireland