The CX Operator
Operational
Subscribe
← Briefing index

Building a modern QA program: A playbook for total coverage

Transition from manual sampling to 100% coverage. This playbook covers scorecard design, calibration, and using conversation intelligence for better coaching.

Desk
QA
Filed by
The CX Operator Desk
Date
Aug 10, 2026
Read time
4 min
Building a modern QA program: A playbook for total coverage

The modern quality assurance (QA) program identifies systemic service gaps by analyzing 100% of customer interactions rather than relying on a small, random sample of calls. By shifting from a 'policing' mindset to a data-driven coaching model, operations leaders can use QA to drive specific business outcomes like retention and resolution speed. This transition requires a combination of outcome-based scorecards, automated conversation intelligence, and a tight feedback loop with training teams.\n\nKey takeaways\n- Full coverage is the new standard: Manual sampling of 1-2% of calls leaves a large share of blind spots; modern programs use automation to monitor every interaction.\n- Outcome-over-checklist: Effective scorecards prioritize problem resolution and sentiment over rigid script adherence.\n- Calibration is non-negotiable: Regular sessions between QA and leadership ensure that 'good' looks the same across the entire floor.\n- QA as a coaching engine: Data from QA should directly inform personalized agent training plans rather than just serving as a performance grade.\n\n## What does a modern QA program look like?\n\nIn the traditional model, a QA analyst listens to a handful of calls per agent each month. This method is statistically insignificant and often catches 'bad' calls by pure luck, leading to agent frustration and a lack of trust in the data. According to Metrigy (https://www.metrigy.com), which tracks CX and AI success metrics, top-performing organizations are increasingly moving toward automated analysis to capture a more accurate picture of agent performance.\n\nA modern program treats QA as a business intelligence function. Instead of just checking if an agent said the required greeting, the program looks for markers of customer frustration, compliance risks, and the root causes of long handle times. This allows managers to see if a spike in call volume is due to a product defect or a training gap.\n\n## Step 1: Designing an outcome-based scorecard\n\nMost legacy scorecards are too long. When an agent has to hit 25 different points on a checklist, they stop focusing on the customer and start focusing on the list. A modern scorecard should be lean, focusing on three core areas: Compliance, Process, and Outcome.\n\n1. Compliance: Non-negotiable items like ID verification and privacy disclosures. This is where a conversation-intelligence layer like Hear.ai (https://hear.ai) is particularly useful, as it can automatically flag every call that misses a required legal statement.\n2. Process: Did the agent use the right tools? Did they follow the troubleshooting flow? This measures the efficiency of your internal playbooks.\n3. Outcome: Did the customer get what they needed? This is the most important metric. If an agent breaks a minor script rule but solves a complex problem that prevents a churn event, the QA score should reflect that success.\n\n## Step 2: Moving from 2% sampling to 100% coverage\n\nTo get a true sense of what is happening on the floor, you need to see everything. This is no longer a manual task. Organizations typically pair a CCaaS platform like Five9 (https://www.five9.com) or Genesys (https://www.genesys.com) with specialized conversation intelligence tools. \n\nBy automating the 'first pass' of QA, your analysts stop being data entry clerks and start being investigators. Automation can surface the 'outliers'—the 5% of calls where a customer was shouting or where an agent went silent for 60 seconds. This allows the human QA team to spend their limited time on the interactions that actually require a human's nuanced judgment. Gartner's Hype Cycle for Customer Service & Support (https://www.gartner.com/en/customer-service-support) highlights how domain-specific AI is maturing to handle these types of complex analysis tasks, allowing for deeper insights into customer intent.\n\n## Step 3: The calibration ritual\n\nQA data is only useful if it is consistent. Calibration is the process where multiple stakeholders—QA analysts, team leads, and even department heads—score the same call independently and then compare their results. \n\nIf one supervisor gives a call an 85 and another gives it a 60, your program has a 'drift' problem. This inconsistency destroys agent morale. You should hold calibration sessions at least bi-weekly. The goal is not to find a 'perfect' score, but to align on the 'why' behind the grade. When the leadership team is aligned, the coaching that follows is much more effective. For more on how to manage these feedback loops, see our guide on [agent-coaching-frameworks.html].\n\n## Step 4: Closing the loop with training and WFM\n\nQA should never exist in a vacuum. The data generated from scorecards should be a primary input for your Workforce Management (WFM) and training strategies. If QA data shows that a large share of agents are struggling with a new product launch, that is a signal for the training team to develop a targeted module, not a reason to penalize individual agents.\n\nWhen QA is integrated with platforms like Salesforce Service Cloud (https://www.salesforce.com/service/) or Zendesk (https://www.zendesk.com), you can start to see the correlation between QA scores and Customer Satisfaction (CSAT). This helps prove the ROI of your QA program to executive leadership. If you are also managing automated support, you may need to adjust these processes; see our playbook on [how-to-audit-ai-agents.html].\n\n## FAQ\n\nHow many calls should a human still listen to?\nWhile AI should handle 100% of the initial screening and compliance checks, humans should still perform deep-dive audits on 2-5 high-value or high-risk calls per agent per month to provide nuanced coaching that automation might miss.\n\nHow do we prevent agents from feeling 'spied on' with 100% coverage?\nTransparency is key. Explain that 100% coverage protects agents from being judged on one 'bad' call that happened to be sampled. It ensures their overall performance and hard work are seen fairly across thousands of interactions.\n\nWhat is the most important metric in a modern QA program?\nWhile every center is different, 'External Resolution'—whether the customer's issue was solved from their perspective—is generally the strongest indicator of a high-quality interaction.\n\nTo learn more about optimizing your floor's performance, explore our related guide on building an effective agent coaching framework.