How to build a modern contact center QA program
Learn how to transition from random sampling to 100% QA coverage. This playbook covers behavior-driven scorecards, calibration, and AI-driven analysis.

Modern quality assurance (QA) in the contact center is shifting from a manual check-the-box exercise to a comprehensive system for performance intelligence. By moving from random 2% sampling to 100% conversation coverage and adopting behavior-based scorecards, operations leaders can transform QA into a strategic engine for agent retention and customer satisfaction. This transition requires a mix of standardized rubrics, automated analysis, and a rigorous calibration process.
Key takeaways
- Full coverage is the new standard: Relying on a 2% random sample creates blind spots and unfair agent evaluations; modern tools allow for 100% analysis.
- Behavior-driven scorecards: Move beyond binary compliance (e.g., "Did they say the greeting?") to qualitative measures like empathy and problem-solving.
- Calibration is the foundation of trust: Regular sessions between supervisors and QA analysts prevent "grader drift" and ensure agents feel the process is fair.
- QA data must feed the WFM loop: Quality scores should directly influence coaching schedules, training modules, and performance-based routing.
Why is the traditional QA sampling model failing?
The traditional model of auditing 2–5 calls per agent per month is statistically insignificant and often leads to biased results. When a supervisor only hears a tiny fraction of an agent's work, they are likely to miss the nuances of high-performing interactions or the root causes of systemic failures. Gartner’s Customer Service & Support practice notes that the maturity of support technologies is pushing teams toward more automated, data-driven oversight to bridge these gaps.
Sampling also creates a "gotcha" culture. Agents often feel that their monthly score is a matter of luck—whether the auditor happened to pick their best or worst call. To fix this, operations leads are moving toward conversation intelligence. By pairing a CCaaS platform like Five9 or Talkdesk with a conversation-intelligence layer like Hear.ai, teams can analyze every interaction for compliance, sentiment, and intent without increasing headcount.
How do you build a behavior-driven scorecard?
A modern scorecard should prioritize behaviors that drive customer outcomes rather than just adherence to a script. While compliance items (e.g., ID verification) remain necessary for legal reasons, the bulk of the score should reflect the quality of the interaction.
1. Define your core competencies
Divide your scorecard into three main buckets:
- Compliance & Process: Did the agent follow legal requirements and use the CRM (like Salesforce Service Cloud) correctly?
- Soft Skills: Did the agent use a professional tone, demonstrate empathy, and actively listen?
- Problem Solving: Did the agent provide an accurate solution and minimize the need for a follow-up call?
2. Use a weighted scoring system
Not all questions are created equal. Failing a compliance check might result in an automatic fail for the call, whereas missing a non-essential greeting might only deduct 2 points. Weighting ensures that the final score reflects the actual impact on the customer and the business.
3. Incorporate sentiment analysis
Modern QA programs often include a metric for "Customer Sentiment Shift." This measures whether an agent successfully de-escalated a frustrated caller. Metrigy, which focuses on CX/AI success metrics, often highlights how sentiment data provides a more accurate picture of agent efficacy than Average Handle Time (AHT) alone.
How do you achieve 100% QA coverage?
Achieving 100% coverage does not mean hiring more auditors; it means using AI to handle the initial pass. Automated QA systems can scan transcripts for keywords, compliance phrases, and sentiment markers across every call, chat, and email.
When a system like Hear.ai flags a high-risk interaction—such as a compliance breach or a sharp drop in sentiment—it is automatically routed to a human auditor for a deep dive. This allows your QA team to spend 100% of their time on the 5% of calls that actually require human judgment. This targeted approach ensures that training resources are directed where they are needed most, rather than being spread thin across random samples.
What is the process for effective calibration?
Calibration is the process of ensuring that different evaluators score the same call in the same way. Without it, agents will quickly lose trust in the QA program.
The Calibration Playbook:
- Select a "Gold Standard" call: Choose an interaction that has a mix of complex technical issues and soft-skill opportunities.
- Independent Scoring: Have all supervisors and QA analysts score the call independently using the standard rubric.
- The Review Meeting: Meet to discuss the scores. Where there is a discrepancy of more than 5%, the team must debate the interpretation of the scorecard until a consensus is reached.
- Update the Rubric: If a specific question consistently causes disagreement, the wording of the scorecard itself is likely the problem and needs to be clarified.
How do you close the loop with coaching?
QA data is useless if it stays in a spreadsheet. It must be integrated into the broader operations workflow. For example, if QA data shows a team-wide struggle with a new product feature, the Training department should be alerted to create a targeted micro-learning module.
Furthermore, QA scores should be visible to agents in real-time. Platforms like Zendesk or Intercom allow for integrated feedback loops where agents can see their scores and read supervisor comments immediately after an audit. This transparency reduces anxiety and allows for faster behavioral correction.
FAQ
How many calls should we manually audit per agent?
While AI should cover 100% of calls for compliance and sentiment, human auditors should still perform 2–4 deep-dive audits per agent per month. These human reviews focus on nuanced coaching opportunities that AI might miss, such as the subtle use of irony or complex rapport building.
What is the ideal calibration score?
Teams should aim for a calibration variance of less than 5%. If evaluators are consistently scoring the same interaction with more than a 5-point difference, it indicates that the scorecard definitions are too subjective and need more concrete guidelines.
How do we handle AI errors in QA scoring?
Always provide agents with a clear "dispute" mechanism. If an automated system flags a call incorrectly, an agent should be able to flag it for human review. This maintains trust in the technology and provides the data needed to fine-tune the AI models over time.
Should we share QA scores with the whole team?
Individual scores should remain private between the agent and their supervisor to avoid public shaming. However, sharing anonymized, aggregate team trends is highly effective for collective goal-setting and identifying departmental strengths.
To see how modern QA data can improve your bottom line, explore our guide on balancing efficiency and quality in the contact center.