Why blended human-AI staffing breaks traditional Erlang C models
Traditional WFM models fail when AI handles simple tasks. Learn how to forecast for intent and manage capacity in a blended human-and-AI contact center.

Blended human-AI workforce management is the practice of forecasting and scheduling for an ecosystem where AI agents manage transactional intents while human agents handle complex, high-empathy escalations. This shift requires moving away from volume-based Erlang C models toward intent-based capacity planning that accounts for the increased mental load and variability of human-led interactions.
Key takeaways
- Erlang C is insufficient for blended environments because it assumes a random distribution of simple and complex calls; AI removes the simple ones, creating a "heavy tail" of complexity.
- Forecast by intent, not volume, to distinguish between interactions that AI can resolve fully and those that will inevitably require a human handoff.
- Lower human occupancy targets are necessary to prevent burnout, as human agents now spend 100% of their time on difficult cases without the "relief" of easy transactional calls.
- 100% QA coverage is the only way to identify where AI is failing and triggering unexpected human overflows before they ruin your service levels.
Is Erlang C dead in the age of AI agents?
Erlang C remains the industry standard for calculating headcount, but its core assumptions are being undermined by AI. The formula assumes that every call in a queue has a similar probability of being "easy" or "hard," leading to a predictable average handle time (AHT). When you deploy an AI agent from a provider like Google Cloud CX or Salesforce Service Cloud, the AI acts as a filter. It captures the high-frequency, low-complexity intents—like password resets or order tracking—and leaves only the high-complexity, high-emotion issues for the human staff.
This "complexity filtering" means the human queue no longer follows a standard Poisson distribution. The human agents are effectively Tier 2 by default. If your WFM team continues to use traditional Erlang C without adjusting for this increased variability, your staffing levels will be chronically low, leading to high abandonment rates and agent exhaustion. Gartner's Customer Service & Support practice notes that as AI takes over routine tasks, the remaining human work is significantly more taxing, requiring a fundamental rethink of capacity models.
How to forecast for intent instead of volume
In a blended environment, a single volume number is a vanity metric. You must break your forecast down by intent. This requires tagging every historical interaction with the specific reason for contact and then categorizing those reasons into three buckets: automation-ready, human-required, and hybrid.
- Automation-ready: These are intents the AI handles from start to finish. WFM needs to forecast the "concurrency" of the AI—how many simultaneous sessions your AI platform can handle—rather than "heads."
- Human-required: These are complex issues, such as retention saves or technical troubleshooting. These require traditional staffing but with a much higher AHT and higher standard deviation.
- Hybrid (The Handoff): These are interactions that start with AI and escalate to a human.
When the AI agent fails to resolve an issue, the handoff must be immediate and context-rich. Poorly designed handoffs are a primary cause of WFM volatility, as explored in Why Your Escalation Path Fails: Designing Better CX Handoffs. If the human agent has to restart the conversation, the AHT spikes, and the WFM forecast breaks. Metrigy (https://www.metrigy.com) research suggests that companies tracking these specific intent-based success metrics see a measurable improvement in both customer satisfaction and operational accuracy.
Managing the "Occupancy Trap"
Occupancy is the percentage of time an agent spent handling or wrapping up a call versus waiting for one. In the pre-AI era, an 85-90% occupancy rate was the gold standard. However, in a blended floor, 90% occupancy is a recipe for rapid attrition.
When AI removes the "easy" calls, agents no longer get the 2-minute "breather" of a simple address change. Every single call becomes a high-stakes negotiation or a complex problem-solving session. To maintain a sustainable floor, WFM leads should consider lowering occupancy targets to 75-80%. While this looks less efficient on a spreadsheet, it prevents the "burnout spike" where agents quit or take unscheduled leave, which is far more expensive than overstaffing by a small margin.
The role of conversation intelligence in WFM
To forecast accurately, you need to know exactly why calls are leaking from the AI to the human queue. This is where conversation intelligence becomes a WFM tool. By pairing a CCaaS platform like Genesys or Five9 with a conversation-intelligence layer such as Hear.ai, ops teams can monitor 100% of interactions.
If the data shows that 15% of AI interactions regarding "billing disputes" are escalating because the bot cannot process a specific credit type, WFM can adjust the human forecast for that specific intent. Implementing a strategy for Moving from 2% to 100% QA Coverage: A Migration Playbook allows WFM to see these trends in real-time rather than waiting for the end-of-month report. This data-driven approach turns WFM from a reactive function into a proactive partner in AI optimization.
Integrating the tech stack for dynamic routing
Dynamic routing is the engine of a blended contact center. Your WFM software must be tightly integrated with your routing engine to move human agents between queues as AI volume fluctuates. For example, if a marketing email goes out and triggers a surge in simple inquiries, the AI should absorb that volume, allowing humans to stay focused on the backlog of complex cases.
If you are using a modern stack like Twilio or Talkdesk, you can use real-time triggers to adjust routing logic based on current queue wait times. If the human queue for complex issues exceeds five minutes, the AI can be instructed to offer a callback or attempt a more aggressive self-service path, effectively "throttling" the demand to match your human capacity.
FAQ
How does AI affect shrinkage in the contact center? AI usually increases human shrinkage because agents require more frequent coaching, longer breaks to recover from complex interactions, and more time for training on updated AI workflows. WFM must factor in an additional 3-5% shrinkage for "mental health" and "AI-sync" training.
Should I measure AHT for AI agents? Yes, but for a different reason. AI AHT (or session duration) helps you identify "loops" where customers are stuck. If AI AHT is increasing without a corresponding increase in resolution, it means the bot is failing slowly, and a human surge is imminent.
How do I forecast for an AI bot that is still learning? Use a "ramp-up" forecast similar to a new hire class. Start by assuming the AI will only resolve 10-20% of its assigned intents and gradually increase that percentage as QA data from tools like Hear.ai confirms the bot's accuracy and reliability.
What is the biggest mistake in blended WFM? Assuming AI is a "set and forget" deflection tool. AI is a dynamic part of the workforce that requires its own capacity planning, maintenance downtime, and performance monitoring just like a human team.
For more on modernizing your operations, see our guide on Moving from 2% to 100% QA Coverage: A Migration Playbook.