Skip to content

A support copilot that shipped because the evals said it could

A retrieval-grounded support assistant taken from an untrusted prototype to production behind an evaluation set, cutting first-response time.

Financial Services · Growth-stage financial services company

-43%
First-response time
94%
Answer accuracy

measured on the eval set

1 in 3
Tickets resolved without a person

The challenge

Support volume was outpacing the team. An earlier in-house AI prototype gave confident wrong answers on regulated questions, so nobody was willing to put it in front of customers, and it had sat unshipped for months.

What we did

We built the eval set first, from real historical tickets with known-correct answers, before writing any of the assistant. Then a retrieval pipeline grounded strictly in the product and policy documentation, answering with citations, refusing rather than guessing outside its sources, and handing off to a person on anything touching a regulated topic. Every change was measured against the evals before release, and the eval suite runs on every deploy.

The difference between the prototype that sat unshipped and the system that went live was not the model. It was a number the team could watch.

Once quality was measurable, shipping stopped being a matter of nerve and became a matter of evidence: the eval score cleared the threshold the support lead had set, on their own historical tickets, and the decision made itself. The eval set belongs to the client and outlived our engagement, which is the point of building it first.

Stack

PythonTypeScriptVector databasePostgreSQL

Have something to build?

Tell us the problem. We'll come back with a plan, a price, and who'd actually build it.

  • Free scoping call
  • Reply within 1 business day
  • No lock-in