Skip to main content

Research

An AI score is only useful if you can defend it.

This only works if people accept the evidence. So it is on us to show how the scoring works, where it can be trusted, and where it cannot.

We publish how it works, rather than claims about how accurate it is.

Lines of work

Four questions we keep asking.

Scoring reliability

Can an AI score be trusted?

We compare AI scores with scores given by people on the same calls, and track how closely they match. When the scoring guide is vague, they match less often. That tells us to tighten the guide, not to hide the result.

Rubric design

What makes a scoring point usable

A point that just says "discussed budget" cannot be scored fairly. Every point says what it is for, what passing and failing sound like, and what to listen for. Vague wording is the main cause of inconsistent scores.

Model selection

Which model for which job

Different jobs need different models, and they vary in accuracy and cost. We test and compare rather than always reaching for the biggest model.

Guardrails

Stopping made-up evidence

We work in three steps. First we pull out only what was said. Then we look for patterns in those facts. Then we write it up, adding nothing new.

Calibration

We measure how often we agree with people.

The honest answer on AI scoring is that it depends on how clearly the scoring guide is written. A clear point gives consistent scores. A vague one starts arguments.

So we keep measuring how often the AI and a person give the same score, and we use the disagreements to improve the scoring guide rather than explain them away.

Calibration viewPreview

Grader spread

4 pts

Within target
QA lead
84
Team manager
88
Client reviewer
86

Same call, three graders - calibrated so scoring stays consistent across the floor.

What we will not claim

The limits, stated plainly.

A model cannot see what a manager knows. The account history, the office politics, the thing said off the call. It should never be the final word on how good someone is, and here it is not. It finds the evidence. A person decides what it means.

We do not publish a single accuracy figure. Accuracy depends on your framework, the type of call and the sound quality, so one number would be misleading.

How it works matters as much as what it says.

If you want to pick apart how the scoring works before you trust it, that is exactly the right instinct.