When AI gets withdrawn: a design problem, not an AI problem
When an AI system is deployed and then quietly withdrawn, the headline says “AI failed”. I think that is almost always the wrong conclusion. In most cases the technology did what it was built to do. What failed was the design, and the expectations set around it.
We have seen this before. A business rushes into a new technology, believing it will solve every problem, without thinking enough about design, change management or the experience of the person on the other end. Then the disappointment arrives, and the technology takes the blame.
This article looks at why AI customer service and ordering systems are being scaled back, what that has in common with the paperless office and cloud bill shock, and what to do differently – especially, always giving the customer a human option.
I’m sure everyone has their own war stories, like the automated phone service or bot that gets stuck in a loop of directing you to the website for help and information, or mis-interprets your request. I have seen projects implemented, with good intention, to ‘streamline’ customer service, but end up only being able to service the most common issues and complaints.

We have been here before
The paperless office. In 1975 Business Week predicted that by 1990 most record-handling would be electronic. Instead, paper use climbed for the next two decades. Sellen and Harper’s research for The Myth of the Paperless Office found the cause was not the technology. People kept paper because it suited how they actually worked. The designers had imagined the office as a filing system, not as people reading, annotating and passing documents around. Paper use did eventually fall, but only once tools began to be enhanced to match the way people work.
Cloud bill shock. Cloud delivered what it promised: speed and flexibility. What nobody designed was how the spending would be owned and controlled. In Flexera’s 2025 survey, 84% of respondents said managing cloud spend was their top cloud challenge. The technology worked. The operating model around it did not exist yet, and people just used the same paradigms in cloud that they had been using on-premises.
The pattern is the same each time. A technology arrives with a big promise. The business rushes in, believing it will fix everything, without designing the service, preparing the people or testing the experience. When the result disappoints, the technology takes the blame.
READ ARTICLE: Digital Transformation Roadmap: From Legacy to Innovation
Four cases, four different lessons
These are not all the same failure, and not all of them ended in withdrawal. That is the point: each one shows a design or expectations gap, not a flaw in AI itself.
| Case | What happened | The gap |
| Klarna (2024-25) | Claimed its assistant did the work of 700 agents. In 2025 its CEO said cost had been too dominant a factor, quality fell, and it began hiring human agents again. | Optimised for cost, not for the customer. The CEO later said customers should always be able to reach a person. |
| McDonald’s (2021-24) | Ended a test of AI drive-thru order-taking with IBM, after about two and a half years in around 100 restaurants. It still expects voice ordering to be part of its future. | Reports of wrong orders, and a viral clip of a runaway nugget order, show a system that met the real, messy drive-thru before it was ready. |
| Commonwealth Bank (2025) | Announced 45 customer service job cuts after introducing a voice bot, then reversed them, saying its assessment had not considered all relevant business considerations. The union said call volumes were rising, not falling. | The business case assumed a result that the operational data did not support. An expectations problem. |
| Air Canada (2024) | A tribunal held the airline liable when its chatbot gave wrong advice on bereavement fares, and rejected the idea that the chatbot was a separate entity. | Nobody owned the accuracy of what the AI said. You are responsible for your AI’s words. |
In none of these did the technology “stop working”. In each, a business expected a result it had not designed for, and a customer or an employee paid for the gap.
Note also that none of these companies has given up on AI. Klarna, McDonald’s and CBA all say they will keep using it. They are adjusting the design.

Why the edge case breaks the script
Most AI service deployments are designed and demonstrated on the happy path: the common, simple request that fits the script. Real customers do not stay on the happy path. They complain, they are upset, their request is unusual, the policy has an exception, they are speaking over a noisy road, or their situation is simply not in the training material. That is why they are calling for help.
The cruel part is that the edge cases are where service matters most. A grieving passenger asking about a fare. A customer disputing a charge. Someone in financial hardship. These are the moments when a customer is most likely to judge your business, and the moments when a script is most likely to fail.
When there is no easy way to a human, the AI becomes a barrier rather than a service. One early tester of Klarna’s assistant described it as basically a filter on the way to a person, an opinion that resonates with me. The customer is trapped, the frustration builds, and when a human finally answers, they inherit an angrier customer and no context.
The business pays too: repeat contacts, complaints, staff overtime, lost customers, and in Air Canada’s case, a legal liability. The saving on the spreadsheet disappears into the cost of cleaning up.
The design fix is simple to say and takes discipline to do. Test against your worst cases, not your demo script. Collect real complaints, exceptions and difficult calls, and run them through the system before launch. Then decide in advance what the AI is never allowed to decide alone, such as complaints, refunds above a limit, safety issues and vulnerable customers.

Always design the human path
If you take one rule from this article, make it this: a customer must always be able to reach a person. Klarna’s CEO made the same point after its reversal, saying a business should be clear to customers that a human is always available. A good escalation path has six features:
- It is easy to find. One clear step, such as a button or a spoken “agent”, not a maze of menus.
- It triggers automatically. Repeated failure, a negative tone, words like “complaint” or “bereavement”, a high-value or high-risk request, or low confidence from the AI should all hand the customer to a person without being asked.
- It passes on context. The person receives the conversation and what was tried, so the customer does not start again.
- It is properly staffed. A human path with a two-hour queue is a trap with extra steps. Staff need the time and training for the harder cases, and should be measured on resolution, not on call speed.
- It is honest. Tell customers they are dealing with AI, and what it can and cannot do.
- It has a fallback. When the AI is wrong, slow or offline, the service must still work.
For a mid-sized business or not-for-profit, the human path can be small. It can be one phone number answered during business hours or a named inbox with a promised response time. What matters is that it exists, it works, and the customer can find it.

Set honest expectations and honest measures
Most withdrawals start with an inflated promise. “AI will replace the call centre” is a very different commitment from “AI will handle simple, routine requests and free staff for harder ones”. Decide which one you are making to the board, to staff and to customers, and be clear about what the system will not handle in its first year.
Then measure what matters. Many dashboards flatter the AI because they count the wrong thing. A customer who gives up is counted as “contained”. A conversation that sends the customer round in circles is counted as “handled”.
| A measure that can mislead | A better measure |
| Containment (conversations that did not reach a person) | Resolution: the issue was fixed and the customer did not contact you again within seven days |
| Cost per conversation | Cost per resolved issue, including the human work that follows a failed conversation |
| Average handle time | Time to resolution, and time to reach a person when one is wanted |
| Overall satisfaction | Satisfaction for the AI path and the human path, reported separately |
Three more habits keep expectations honest:
- Take a baseline before you start. You cannot show improvement without a record of where you were. This is the same idea as treating doing nothing as Option 0.
- Pilot with stop conditions. Agree in advance what result would make you pause or roll back, such as a rise in complaints or repeat contacts, and who has the authority to call it.
- Own the output. Check what the AI says against your policies, as the Air Canada ruling made clear. If you would not publish it on your website, do not let a chatbot say it.
READ ARTICLE: The need for Responsible AI
Does this taint what AI can do?
It can, and that is the real cost. Every customer trapped in a bot, every wrong order and every denied refund teaches people to distrust AI. It also teaches some leaders to write it off, just as a failed first attempt can set back a good idea for years. The paperless office was labelled a myth, yet office paper use did decline once tools matched the way people work.
The lesson from the cases is a hopeful one. Klarna, McDonald’s and CBA are all still pursuing AI. What changed was the design: protecting human access, testing against real conditions and being honest about results. The technology was never the weak link.
So when you hear “AI failed”, ask three questions. Who designed the experience? What happens when the script runs out? And what was promised, to whom, and measured how? Those answers will tell you far more than the headline.
Have you been on the receiving end of an AI system that fell over? I would like to hear what happened.
