The CX Operator
Operational
Subscribe
← Briefing index

A QA Calibration Playbook: Getting Your Scores to Actually Agree

If five reviewers score the same call five different ways, your QA program is measuring the reviewer, not the agent. A step-by-step routine to make your scores mean the same thing every time.

Desk
Playbooks
Filed by
The CX Operator Desk
Date
Jul 2, 2026
Read time
5 min
A QA Calibration Playbook: Getting Your Scores to Actually Agree

Here's a test for any QA program: take one interaction, give it to five reviewers, and have them score it independently. If you get five different scores, you don't have a quality program — you have five opinions wearing the same rubric. The number an agent receives depends on who happened to review them, and everyone on the floor knows it.

Calibration is the routine that fixes this. It's the ongoing work of getting every scorer — human or automated — to reach the same verdict on the same interaction. It's unglamorous, it never really ends, and it's the difference between QA scores that drive coaching and QA scores nobody believes. Here's how to run it.

Why scores drift

Reviewers diverge for predictable reasons, and naming them tells you what to fix:

None of these are about reviewer competence. They're about a rubric and a routine that leave too much room. Calibration closes the room.

Run a calibration session

The core routine is simple and should run on a fixed cadence — weekly or biweekly, not "when we get to it":

  1. Pick one real interaction ahead of time. Choose something with a few genuine judgment calls in it, not an obvious pass or fail. The gray-area calls are where drift lives.
  2. Everyone scores it independently first. No discussion, no peeking. You need each reviewer's honest, uninfluenced score to see where you actually stand.
  3. Reveal the spread. Put all the scores up together. The disagreements — especially where reviewers split on the same criterion — are the entire point of the session.
  4. Discuss only the divergences. Don't re-litigate the criteria everyone agreed on. Spend the time where reviewers split, and have each side explain their reasoning.
  5. Decide the right answer, out loud. Reach an agreed score and, more importantly, an agreed rule: "Going forward, if the agent does this, it's a full deduction." Write the rule down.
  6. Update the rubric with what you learned. Every divergence that traces back to an ambiguous criterion is a rubric fix, not just a conversation. Capture it so the same argument doesn't recur next month.

The output of a good calibration session isn't the agreed score on one call. It's a sharper rubric and a shared set of rules that make the next thousand scores more consistent.

Measure your agreement

You can't manage drift you don't measure. You don't need heavy statistics — a simple, repeatable agreement check is enough for most teams:

These numbers are illustrative, not benchmarks: the point isn't hitting a magic agreement percentage, it's watching your own number move in the right direction.

Fix the rubric, not just the reviewers

The instinct after a bad calibration session is to tell reviewers to try harder to agree. That treats the symptom. Persistent disagreement on a criterion is almost always the criterion's fault: it's vague, it's subjective, or it's asking reviewers to judge something they can't reliably observe.

For every criterion reviewers keep splitting on, either define it in observable terms ("confirmed the account number before discussing the balance") or cut it. A rubric full of behaviors you can actually point to calibrates itself; a rubric full of vibes never will.

Calibrating an automated scorer

If you've added AI-assisted evaluation, it's just another reviewer — and it needs calibrating the same way. Include the tool's scores in your sessions:

An automated scorer's superpower is that it never suffers severity creep or the halo effect. But it will faithfully apply a vague criterion in its own consistent-but-wrong way. Calibration is how you catch that.

Keep it a habit

Calibration isn't a project you finish — it's a routine you keep. The month you stop running sessions is the month your scores quietly start drifting apart again, and you won't notice until an agent challenges a score you can't defend.

Teams treat calibration as a launch activity: tighten everything up before go-live, then move on. Drift doesn't respect your launch date. New reviewers join, policies change, edge cases accumulate. A standing session on the calendar is cheap insurance against the slow decay of your scores back into opinion.

What to do next

Scores only matter if they mean the same thing regardless of who assigned them. Calibration is the work that makes that true — and it's the foundation everything else in QA, from coaching to accountability, quietly depends on.