AI in Customer Support: What 'Doing Great' Actually Means
AI & Automation September 8, 2026 5 min read

AI in Customer Support: What 'Doing Great' Actually Means

Sam Altman says AI is crushing it in customer support. He's not wrong — but he's also not telling the whole story. Here's what's really happening.

Sam Altman dropped a casual line recently that got a lot of nods in boardrooms and a lot of eye-rolls in contact centers: customer support, he said, is 'doing great' with AI. And honestly? He's right. He's also leaving out about half the picture.

The thing is, 'doing great' is doing a lot of heavy lifting in that sentence. It's the kind of phrase that sounds like a verdict when it's really just a partial score. Customer support has become one of the clearest proving grounds for AI in the enterprise — not because it's glamorous, but because it's measurable, high-volume, and frankly desperate for efficiency. Ticket queues don't care about hype cycles. But the way AI is actually performing inside support operations is more complicated, more conditional, and more dependent on human judgment than the headline suggests.

Why Support Became AI's First Real Test

Think about what customer support actually looks like at scale. Thousands of conversations a day, most of them asking variations of the same ten questions. Years of documented chat logs, macro responses, and knowledge articles just sitting there. Defined workflows with predictable outcomes. If you were designing a sandbox for AI to prove itself in, you'd build something that looks a lot like a mid-size SaaS company's support queue.

That's the honest reason AI moved into support faster than almost any other business function. The conditions were favorable. The data was there. The ROI case wrote itself: deflect more tickets, reduce handle time, staff fewer overnight shifts. And in those narrow, well-defined lanes, AI has genuinely delivered. Password resets, order tracking, return policy questions, appointment reschedules — these are interactions where intent is clear, the rules are explicit, and the acceptable outcomes fit in a short list. AI handles them well. Not perfectly, but well enough that the economics make sense.

Klarna's widely cited AI assistant is the kind of example that gets quoted in every vendor deck for a reason. Cutting response times from minutes to seconds across multiple languages and markets is a real achievement. Intercom's Fin agent resolving a large share of inbound questions — and crucially, handing off unresolved ones with full context rather than dropping the customer into a void — shows that the handoff problem, which used to be a disaster, is getting solved. Vagaro, the booking platform for salons and spas, auto-resolves nearly half its incoming requests with AI while keeping satisfaction scores in the low nineties. That's not a pilot. That's operational infrastructure.

So yes, AI is doing great in customer support — if you measure it the way most vendors measure it.

The Metrics That Make AI Look Good (And What They Miss)

Here's where the story gets more interesting. Deflection rates and handle time are seductive metrics. They're clean, they're fast, and they map directly to cost. When a VP of Support walks into a budget meeting and says 'we deflected 40% more tickets this quarter,' that lands. Nobody asks a follow-up question about the quality of those deflections.

But deflection isn't the same as resolution. A customer who gets a bot response, gives up, and churns three months later didn't 'deflect' — they left. And the support team never knew it happened. The handle time metric has the same problem. You can absolutely cut average handle time by having AI draft the first response, but if that draft is confidently wrong and the agent sends it anyway because they're juggling six conversations, you've just made things faster and worse simultaneously.

The gap between 'the bot closed the ticket' and 'the customer's problem was solved' is where a lot of the real trouble lives. And it's a gap that traditional support metrics were never designed to catch. Churn risk, relationship strength, lifetime value — these move slowly, they're hard to attribute, and they don't show up in the weekly dashboard. By the time they show up in the revenue numbers, the connection to a degraded support experience is almost impossible to prove.

This isn't an argument against AI in support. It's an argument for measuring it more honestly.

The Unglamorous Work Where AI Actually Shines

Some of the most valuable AI work in customer support is invisible to customers entirely. Mark Waks, a senior managing director of customer experience at Slalom, put it well: the real value isn't in automating conversations, it's in everything around them. Triage. Summarization. Routing. Knowledge retrieval. The work that agents hate, that slows everything down, and that nobody outside the contact center ever thinks about.

Consider what happens when a customer calls about a billing issue that touches three different systems. Before AI, the agent spent the first four minutes pulling up account history, reading through the last five interaction notes, and figuring out which team handled it before. With AI-assisted summarization, they walk in already knowing. The customer doesn't have to repeat themselves. The agent isn't reactive — they're ready. That's not a headline feature. It doesn't demo well. But it changes the texture of every single interaction.

Nik Kale, a principal engineer at Cisco, described something similar: AI-driven intent classification that gets customers to the right agent faster, eliminating the maddening loop of being transferred between queues before landing somewhere useful. The win isn't replacing agents. It's cutting the wasted cycles that make customers furious before the actual conversation even starts.

And for new agents, AI-generated call summaries and suggested responses aren't just efficiency tools — they're training wheels that actually work. A newer agent who starts every interaction with a strong AI-drafted response learns faster, makes fewer mistakes, and ramps to full productivity in a fraction of the time. That's a real operational advantage that rarely gets counted in the deflection metrics.

Where AI Breaks — And Why It Breaks Confidently

The failure modes of AI in customer support are specific and worth understanding, because they're not random. They cluster around a few predictable conditions.

The first is emotional complexity. A customer who's been charged incorrectly three times and is threatening to cancel is not a routing problem. They're not looking for a policy explanation. They need to feel heard by someone with the authority and empathy to actually fix it. AI is genuinely bad at this — not because it can't generate sympathetic language, but because generated sympathy reads as hollow almost immediately. 'I understand your frustration' is the fourth time in a row the bot has said that while doing nothing. Customers know the difference.

The second failure mode is multi-system complexity. When a problem spans billing, product, and logistics simultaneously — a delayed order that was also charged twice and involved a product that's been discontinued — the interaction requires judgment calls that don't fit neatly into any workflow. AI can handle the pieces, but assembling them into a coherent resolution requires the kind of contextual reasoning that current systems don't reliably provide.

The third, and arguably the most dangerous, is knowledge base decay. This one doesn't get enough attention. AI systems in customer support are only as good as the data they're grounded in. When that data is stale — old pricing, deprecated features, outdated policies — the AI doesn't know it's wrong. It answers confidently with information that's six months out of date. The customer gets a definitive-sounding response that's completely incorrect. They act on it. Something goes wrong. And the trust damage is worse than if the system had just said 'I'm not sure, let me connect you with someone.'

Knowledge bases decay faster than most teams can update them. Products change. Policies shift. Regulatory requirements evolve. Every gap in that grounding data is a potential confident wrong answer waiting to happen. Before any organization scales AI to first contact, the state of the knowledge base deserves at least as much attention as the model selection.

The Augmentation Model: Why Copilot Beats Autopilot

Michael Hutchison, who leads customer experience at eClerx, described the pattern that seems to be working best across organizations that have gotten past the pilot stage: augmenting agents rather than replacing them. Automated call summarization. Real-time prompts. Next-best-action guidance. These tools reduce after-call work, help newer agents perform at a higher level, and keep humans in the decision seat for anything that actually matters.

This is the copilot model, and it's outperforming the autopilot model for a simple reason: it preserves human judgment exactly where it's needed while removing friction everywhere else. The agent still makes the call on whether to escalate, how much goodwill credit to offer, and whether this customer sounds like they're about to leave. The AI handles the cognitive overhead of remembering policy details, drafting the initial response, and logging the interaction afterward.

Full automation works in the lanes where it works — the narrow, well-defined, low-stakes interactions. But the instinct to push automation as far as possible, as fast as possible, is where organizations run into trouble. Every interaction that gets automated is an interaction where a human is no longer checking the output. And as the complexity of those interactions increases, the error rate goes up while the visibility goes down.

Governance Isn't a Compliance Checkbox Anymore

There's a governance conversation happening in customer support right now that wasn't happening two years ago. When AI handles first contact at scale, questions about auditability, escalation rules, and accountability stop being back-office concerns and become core operational design decisions.

Who's responsible when an AI gives a customer incorrect information about their rights under a return policy? What happens when a vulnerable customer — someone in financial distress, someone dealing with a health crisis — gets routed into an automated flow that has no mechanism to recognize they need a human? What's the paper trail when a compliance-sensitive interaction is handled entirely by a model?

These aren't hypothetical edge cases. They're the kinds of situations that end up in regulatory reviews and news stories. Building in escalation triggers, maintaining logs of AI-driven interactions, and establishing clear rules for when automation must hand off to a human — these are no longer optional features. They're the price of deploying AI at the front of the customer experience.

The organizations that are doing this well aren't treating governance as a constraint on AI deployment. They're treating it as part of the design. Escalation rules built in from the start. Human review triggered automatically when confidence scores drop below a threshold. Clear ownership for auditing AI outputs on a regular cadence. It takes more work upfront, but it's the difference between AI that builds trust over time and AI that creates liability.

The Question Nobody's Really Asking Yet

Here's what the current conversation about AI in customer support mostly ignores: what does it do to the relationship over time?

Support interactions aren't just problem-solving events. They're moments of contact between a customer and a brand. Done well, they build loyalty. Done badly, they erode it. The question of whether AI-handled interactions build or erode that relationship over months and years is genuinely unanswered right now. The data doesn't exist yet, or if it does, it's not being shared publicly.

Short-term satisfaction scores can stay high even as something more subtle shifts. A customer who gets fast, accurate answers from a bot for twelve months might still feel less connected to the brand than one who talked to a human occasionally. Or they might not. We don't actually know. And the organizations making aggressive bets on full automation are essentially running that experiment on their customer base without knowing the outcome.

That's not an argument for caution above all else. It's an argument for keeping humans genuinely in the loop — not as a compliance measure, but as a hedge against a risk that the current metrics can't see.

Sam Altman isn't wrong. Customer support is one of the places AI is delivering real, measurable value. The deflection numbers are real. The handle time improvements are real. The agent augmentation benefits are real. But 'doing great' is a starting point, not a finish line. The organizations that will actually win with AI in support are the ones that measure more than deflection, maintain their knowledge bases like they matter, keep humans in the decisions that matter most, and stay honest about what they don't yet know.

#AI & Automation#GZOO#BusinessAutomation
AI in Customer Support: What 'Doing Great' Actually Means | GZOO