A Moroccan e-commerce and logistics client deployed an AI chatbot trained on their entire help center and instructed to answer anything. Within six weeks, CSAT on chatbot-handled tickets had dropped to 41%, against 88% for human-handled tickets, and complaint volume about the chatbot itself was rising. The bot was not incompetent, it was unscoped. It attempted to answer refund disputes, shipping damage claims, and account security issues with the same generic confidence it used for tracking number lookups, and customers could tell the difference immediately when it got something consequential wrong.
The Tiered Resolution Architecture
Start by classifying every historical support ticket into intents and measuring two things per intent: volume share and resolution complexity. Intents that are high-volume and low-complexity, order status, tracking, return policy explanation, password reset, store hours, basic product specs, belong in tier one: fully automated, no human review needed, because a wrong answer here is low-stakes and easily corrected. Tier two covers moderate-complexity intents like initiating a return or modifying an order, the AI can execute the workflow but a human reviews or the system requires explicit customer confirmation before finalizing anything irreversible. Tier three is anything touching money disputes, safety, legal complaints, or emotionally charged situations (a customer explicitly angry or repeating a request), these route to a human immediately, with the AI's role limited to gathering context and pre-filling the ticket, never attempting resolution.
The mechanism that makes this work in practice is a confidence threshold tied to intent classification, not just to the AI's own self-reported certainty. If the system cannot classify the incoming message into a tier-one intent with high confidence, it should escalate silently, handing the conversation to a human with full context already gathered, rather than the bot guessing and forcing the customer to repeat themselves after a wrong answer. This single design choice, silent escalation on low confidence rather than confident wrong answers, is usually the difference between a chatbot that raises CSAT and one that tanks it. We rebuilt the Moroccan client's bot around this tiered model, cut its scope to 14 tier-one intents covering 58% of ticket volume, and CSAT on automated resolutions rose to 79% within two months, while overall support cost per ticket dropped because agents stopped having to clean up after failed bot attempts.
- Classify historical tickets by intent and rank each by volume share and resolution complexity before scoping the bot.
- Restrict full automation to a closed list of high-volume, low-complexity, low-stakes intents (tier one).
- Require explicit customer confirmation or human review before the AI finalizes anything irreversible (tier two).
- Route money disputes, safety, legal, and emotionally charged conversations straight to a human, with the AI only pre-gathering context (tier three).
- Escalate silently on low classification confidence instead of letting the bot guess and force a repeat request.
- Measure CSAT and resolution time by tier, not a single blended deflection-rate number.
INSIGHT
Deflection rate as a headline automation metric is actively misleading: a bot can deflect 80% of tickets from human agents while destroying customer trust on the 20% it should never have touched. The metric that matters is CSAT per tier, tracked separately, because it is the only number that tells you whether you scoped the automation correctly.
Focus Point designs and develops tiered AI support systems scoped around what should be automated, not what could be. Let's map your ticket intents and build the right escalation architecture.
Build a support automation that protects your CSAT¿Listo para ponerlo en práctica?
Empecemos un proyecto juntos.
Cuéntanos sobre tu marca. Te respondemos con una lectura estratégica en 48h.