The CX Operator
Operational
Subscribe
← Briefing index

Stop sampling calls: A migration plan for 100% QA coverage

Transition from manual 2% sampling to 100% QA coverage with this step-by-step migration plan. Learn to map scorecards to AI and manage the data firehose.

Desk
QA
Filed by
The CX Operator Desk
Date
Sep 22, 2026
Read time
5 min
Stop sampling calls: A migration plan for 100% QA coverage

Moving from 2% sampling to 100% QA coverage requires transitioning from manual spot-checks to automated conversation intelligence. This migration involves mapping legacy scorecards to machine-readable prompts, validating AI accuracy against human benchmarks, and shifting the QA team's focus from scoring to high-level analysis. By reviewing every interaction, operators eliminate the 'lottery effect' where agents are penalized or rewarded based on a non-representative slice of their work.

Key Takeaways

Why the 2% sampling model is a liability

For decades, the industry standard has been to manually review 2 to 5 calls per agent per month. This model is statistically insignificant. It fails to capture the rare but catastrophic compliance errors and does not provide enough data to identify genuine behavioral patterns.

When a QA lead at a high-volume center uses a platform like Zendesk (https://www.zendesk.com) to manage tickets, they see the volume but only a fraction of the context. If an agent has one bad day in a month of excellence, and those are the two calls pulled for review, their performance metrics are ruined. Conversely, a poor performer might 'win' the QA lottery by having their only two good calls reviewed. This inconsistency destroys agent trust. According to Metrigy (https://www.metrigy.com), companies that integrate AI into their QA processes see more consistent performance metrics because the data set is comprehensive, not anecdotal.

Phase 1: Establish the human baseline

You cannot automate what you cannot define. Before introducing AI, you must ensure your human QA team is aligned. Pick 100 calls and have three different QA evaluators score them using your existing scorecard.

If the evaluators do not agree on at least 90% of the scores, your scorecard is too subjective. AI requires clear, logical instructions. Terms like 'demonstrated empathy' need to be broken down into observable behaviors, such as 'acknowledged the customer's frustration' or 'used the customer's name.' For a deeper dive into making these adjustments, see Practical QA scorecard designs that agents actually respect.

Phase 2: Mapping the scorecard to machine logic

Once your manual baseline is solid, the next step is 'prompt engineering.' This is the process of translating your scorecard questions into instructions for a Large Language Model (LLM) or a conversation intelligence tool.

Modern CCaaS platforms like Five9 (https://www.five9.com) or Genesys (https://www.genesys.com) provide the raw transcripts, but you need an intelligence layer like Hear.ai (https://hear.ai) to analyze them.

Instead of asking 'Was the agent professional?', you program the system to look for specific triggers:

This shift from 'feeling' to 'features' is what allows the machine to process thousands of calls per minute with high reliability.

Phase 3: The Shadow Run and Calibration

Do not switch off manual QA immediately. Run the automated system in 'shadow mode' for 30 days. During this period, the AI scores 100% of the calls, while the human team continues their 2% sampling.

Compare the results. Where the AI and the human disagree, investigate why. Often, the AI is more accurate because it doesn't get tired or bored. However, AI can miss sarcasm or complex cultural nuances. Use these discrepancies to refine your prompts. Gartner's Hype Cycle for Customer Service & Support (https://www.gartner.com/en/customer-service-support) notes that while speech analytics is maturing, human calibration remains the safety net for complex sentiment analysis.

Phase 4: Operationalizing the data firehose

The biggest challenge of 100% QA coverage is not getting the data—it is acting on it. If you move from scoring 200 calls a month to 20,000, your old reporting methods will break.

You must move toward 'Management by Exception.' Instead of looking at every score, the QA manager looks at:

  1. Outliers: Which agents are consistently falling below the compliance threshold?
  2. Trends: Is there a sudden spike in 'product defect' mentions across the entire floor?
  3. Compliance Flags: Immediate alerts for high-risk violations. For handling these moments, refer to When the flag drops: A runbook for real-time compliance escalations.

Phase 5: Re-skilling the QA Team

In a 100% coverage model, the QA team's job changes. They stop being 'scorecard fillers' and start being 'performance strategists.'

They spend their time:

This transition elevates the QA function from a back-office administrative task to a core driver of operational intelligence. For more on this structural shift, see Modernizing Contact Center QA: A Playbook for Full Coverage.

FAQ

How do agents react to being monitored 100% of the time?

Initially, there is often pushback, but this usually fades when agents realize the system is fairer. It captures their 'saves' and great moments that manual sampling misses, providing a more balanced view of their hard work.

Is 100% QA coverage more expensive than sampling?

While the software costs for tools like OpenAI (https://openai.com) or specialized QA platforms are higher than manual labor alone, the ROI comes from reduced compliance fines and more efficient coaching. You are trading variable labor costs for scalable software costs.

Can AI handle 'soft skills' like empathy?

AI is excellent at identifying the markers of empathy, such as specific phrases or tone shifts. However, for high-stakes emotional interactions, human oversight is still required to validate the machine's interpretation of the customer's sentiment.

What happens to the manual QA staff?

They move into higher-value roles. Instead of listening to random calls, they focus on auditing the AI's accuracy, handling complex grievances, and developing high-level coaching strategies that the AI identifies as necessary.

Transitioning to full coverage is the only way to gain a true baseline of your floor's performance and protect the business from systemic risk. Explore our guide on Why your QA scorecards feel like a trap—and how to fix them to ensure your automation is built on a fair foundation.