Skip to content

A support copilot that shipped because the evals said it could

A support assistant that answers from the company's own documentation. It went from a prototype nobody trusted to production, gated by a test set built from past tickets, and cut first-response time.

Financial Services · Growth-stage financial services company

-43%
First-response time
94%
Answer accuracy

measured on the eval set

1 in 3
Tickets resolved without a person

The challenge

Ticket volume was growing faster than the support team. An earlier in-house AI prototype gave confident wrong answers to regulated questions. Nobody was willing to put it in front of customers, and it had sat unreleased for months.

What we did

Before writing any of the assistant, we built the eval set: historical tickets whose correct answers were already known. The assistant looks up the product and policy documentation and answers strictly from it, with citations. Outside those sources it refuses to answer and does not guess. Anything touching a regulated topic is handed to a person. Every change was measured against the eval set before release, and the eval suite runs on every deploy.

What took this from a shelved prototype to a live system was a number the team could watch. Changing the model was not what did it.

Once quality could be measured, going live stopped being a matter of nerve. The support lead set a threshold. The eval score cleared it on the team's own historical tickets, and the decision made itself. The eval set belongs to the client and outlived our engagement. That is the reason to build it first.

Stack

PythonTypeScriptVector databasePostgreSQL

Have something to build?

Tell us the problem. We'll come back with a plan, a price, and who would build it.

  • Free scoping call
  • Reply within 1 business day
  • No lock-in