Research
An AI score is only useful if you can defend it.
This only works if people accept the evidence. So it is on us to show how the scoring works, where it can be trusted, and where it cannot.
We publish how it works, rather than claims about how accurate it is.
Lines of work
Four questions we keep asking.
Scoring reliability
Can an AI score be trusted?
Rubric design
What makes a scoring point usable
Model selection
Which model for which job
Guardrails
Stopping made-up evidence
Calibration
We measure how often we agree with people.
The honest answer on AI scoring is that it depends on how clearly the scoring guide is written. A clear point gives consistent scores. A vague one starts arguments.
So we keep measuring how often the AI and a person give the same score, and we use the disagreements to improve the scoring guide rather than explain them away.
Grader spread
4 pts
Same call, three graders - calibrated so scoring stays consistent across the floor.
What we will not claim
The limits, stated plainly.
A model cannot see what a manager knows. The account history, the office politics, the thing said off the call. It should never be the final word on how good someone is, and here it is not. It finds the evidence. A person decides what it means.
We do not publish a single accuracy figure. Accuracy depends on your framework, the type of call and the sound quality, so one number would be misleading.
How it works matters as much as what it says.
If you want to pick apart how the scoring works before you trust it, that is exactly the right instinct.