ServicesSupport agents

Deflection is a grounding problem

An agent that answers from your documentation, your ticket history and your live systems, escalates when it should, and can prove where every answer came from.

What it actually is

Not a chat widget with your FAQ pasted into a prompt. A retrieval system over your real content, wrapped in rules about what it may and may not do, connected to the systems that hold the answer, and measured against a test set built from your own tickets.

01Answers from your documentation, changelog and resolved tickets
02Reads live state from billing, orders or the product before replying
03Cites the source, so an agent can verify the answer in one click
04Escalates on uncertainty, sentiment or account value, not just keywords
05Writes the outcome back to the helpdesk and CRM

Why the off the shelf version does not hold up

01

Vendor deflection numbers do not survive contact with reality

Marketing implies 30 to 50 percent. The average team sees 10 to 15 in year one. The variable that closes the gap is not the model, it is whether the system is anchored in real documentation. Grounded implementations reach 85 percent and above.

02

Ungrounded answers are worse than no answer

A confident wrong reply about a refund, a policy or a deadline costs more than a slow human one. Once a customer catches the bot inventing, they stop trusting the channel and every future deflection is lost too.

03

The answer usually lives outside the knowledge base

Where is my order, why was I charged this, did my payment go through. None of these are in a help doc. They are in the order system, the billing system and the product, so an agent that can only read documentation cannot resolve them.

04

Nobody defines what it is not allowed to do

Most deployments specify what the bot should answer and never specify what it must refuse. Refunds, legal commitments, medical or financial advice and anything touching a regulated obligation need an explicit boundary, not a hopeful prompt.

The engineering behind it

What makes it work in production

The parts that decide whether this survives contact with real data, real volume and real edge cases.

Answers are pinned to your documents, not the model's memory

Content is chunked, embedded and retrieved per question, and the answer is constrained to the retrieved passages. If nothing relevant comes back, the system says so and escalates rather than filling the gap.

Retrieval augmented generation

Hard limits the model cannot talk its way past

Refusal rules, topic boundaries and action permissions are enforced outside the prompt, so a persuasive customer cannot argue the agent into issuing a refund or making a commitment it should not make.

Guardrails and policy enforcement

We measure it against your real tickets before launch

We build a test set from your resolved conversations and score every change against it. That is how you get a defensible deflection number rather than a vendor claim, and how a prompt change stops silently regressing quality.

Evaluation harness

Uncertainty routes to a person with the context attached

Low retrieval confidence, negative sentiment, high account value or a restricted topic all trigger handoff, and the human receives the conversation, the sources consulted and the reason for escalation.

Human in the loop escalation

The bill stays proportional to the value

Caching, routing simple questions to smaller models, trimming retrieved context and, where the volume justifies it, fine tuning a smaller model to do the repetitive work. Token spend is tracked per conversation so cost per resolution is a number you can see.

Inference cost optimisation

You can see what it did and why

Every conversation is logged with the sources retrieved, the decision taken and the confidence. When something goes wrong you can replay it rather than guess.

Observability and tracing

What you end up owning

A live agent in your helpdesk, not a demo environment
Integrations to the systems that hold the answers
Escalation and refusal rules written down and enforced
An evaluation set built from your resolved tickets
A dashboard for deflection, escalation and cost per resolution
Complete source code and prompts, owned by you

Built into the tools you already run

Helpdesk

  • Zendesk
  • Intercom
  • Gorgias
  • Freshdesk
  • Help Scout
  • Front

CRM

  • HubSpot
  • Salesforce
  • Pipedrive
  • Zoho CRM
  • Close

Knowledge

  • Notion
  • Confluence
  • Zendesk Guide
  • GitBook
  • Intercom Articles

Systems of record

  • Stripe
  • Shopify
  • Chargebee
  • Custom APIs
  • Webhooks

Start with one workflow

A 30 minute audit, no sales pitch. We map where this fits in your stack, what it would take to build, and whether it is worth doing at all.