A QA Calibration Playbook: Getting Your Scores to Actually Agree
If five reviewers score the same call five different ways, your QA program is measuring the reviewer, not the agent. A step-by-step routine to make your scores mean the same thing every time.

Here's a test for any QA program: take one interaction, give it to five reviewers, and have them score it independently. If you get five different scores, you don't have a quality program — you have five opinions wearing the same rubric. The number an agent receives depends on who happened to review them, and everyone on the floor knows it.
Calibration is the routine that fixes this. It's the ongoing work of getting every scorer — human or automated — to reach the same verdict on the same interaction. It's unglamorous, it never really ends, and it's the difference between QA scores that drive coaching and QA scores nobody believes. Here's how to run it.
Why scores drift
Reviewers diverge for predictable reasons, and naming them tells you what to fix:
- Vague criteria. "Displayed empathy" means something different to everyone. Anything open to interpretation will be interpreted differently.
- Different weighting. Two reviewers agree the agent skipped a step but disagree on whether it's a minor deduction or an auto-fail.
- Severity creep. A reviewer who's seen ten bad calls in a row scores the eleventh harder. Context bleeds into the score.
- The halo effect. A warm, articulate agent gets the benefit of the doubt on a step they actually missed; a flat-sounding one doesn't.
None of these are about reviewer competence. They're about a rubric and a routine that leave too much room. Calibration closes the room.
Run a calibration session
The core routine is simple and should run on a fixed cadence — weekly or biweekly, not "when we get to it":
- Pick one real interaction ahead of time. Choose something with a few genuine judgment calls in it, not an obvious pass or fail. The gray-area calls are where drift lives.
- Everyone scores it independently first. No discussion, no peeking. You need each reviewer's honest, uninfluenced score to see where you actually stand.
- Reveal the spread. Put all the scores up together. The disagreements — especially where reviewers split on the same criterion — are the entire point of the session.
- Discuss only the divergences. Don't re-litigate the criteria everyone agreed on. Spend the time where reviewers split, and have each side explain their reasoning.
- Decide the right answer, out loud. Reach an agreed score and, more importantly, an agreed rule: "Going forward, if the agent does this, it's a full deduction." Write the rule down.
- Update the rubric with what you learned. Every divergence that traces back to an ambiguous criterion is a rubric fix, not just a conversation. Capture it so the same argument doesn't recur next month.
The output of a good calibration session isn't the agreed score on one call. It's a sharper rubric and a shared set of rules that make the next thousand scores more consistent.
Measure your agreement
You can't manage drift you don't measure. You don't need heavy statistics — a simple, repeatable agreement check is enough for most teams:
- After each session, record how many reviewers landed within an agreed tolerance of the consensus score (say, within five points on a hundred-point scale — pick a tolerance and hold it steady).
- Track that agreement rate over time. If it climbs, calibration is working. If it stalls, your rubric still has soft spots.
- Watch per-criterion agreement, not just the total. A solid overall score can hide one criterion reviewers never agree on — and that's usually the one quietly poisoning trust.
These numbers are illustrative, not benchmarks: the point isn't hitting a magic agreement percentage, it's watching your own number move in the right direction.
Fix the rubric, not just the reviewers
The instinct after a bad calibration session is to tell reviewers to try harder to agree. That treats the symptom. Persistent disagreement on a criterion is almost always the criterion's fault: it's vague, it's subjective, or it's asking reviewers to judge something they can't reliably observe.
For every criterion reviewers keep splitting on, either define it in observable terms ("confirmed the account number before discussing the balance") or cut it. A rubric full of behaviors you can actually point to calibrates itself; a rubric full of vibes never will.
Calibrating an automated scorer
If you've added AI-assisted evaluation, it's just another reviewer — and it needs calibrating the same way. Include the tool's scores in your sessions:
- Score the calibration interaction with the tool alongside your humans and put it in the same spread.
- Where the tool diverges from the human consensus, treat it exactly as a human outlier: understand the reasoning, and either correct the criterion, correct the tool's configuration, or decide the tool was right and your humans drifted.
- Watch for the tool being consistently off in one direction on a specific criterion — that's a configuration fix that improves every future score at once.
An automated scorer's superpower is that it never suffers severity creep or the halo effect. But it will faithfully apply a vague criterion in its own consistent-but-wrong way. Calibration is how you catch that.
Keep it a habit
Calibration isn't a project you finish — it's a routine you keep. The month you stop running sessions is the month your scores quietly start drifting apart again, and you won't notice until an agent challenges a score you can't defend.
Teams treat calibration as a launch activity: tighten everything up before go-live, then move on. Drift doesn't respect your launch date. New reviewers join, policies change, edge cases accumulate. A standing session on the calendar is cheap insurance against the slow decay of your scores back into opinion.
What to do next
- Run one calibration session this week on a genuinely gray-area interaction, with independent scoring first.
- Record your agreement rate — however rough — so you have a baseline to move.
- Turn the top two disagreements into rewritten, observable criteria before the next session.
- If you use an automated scorer, put it in the next spread as just another reviewer.
Scores only matter if they mean the same thing regardless of who assigned them. Calibration is the work that makes that true — and it's the foundation everything else in QA, from coaching to accountability, quietly depends on.