If you run QA in a contact centre, you already know the maths. A senior reviewer takes 15 to 30 minutes to score one call against your scorecard. They review maybe 20 calls a week. Your floor takes 50,000 calls a week. The reviewed sample is 0.04%. Adding more QA reviewers shifts the rate to 0.08%, then 0.12%, but it never gets close to a number anyone in operations would call coverage. This piece is about the architecture that gets you to 100%.
The QA bottleneck, from first principles#
The bottleneck is not lack of will. Most QA teams would love to review more calls. The bottleneck is unit cost. A senior reviewer scores about four calls an hour, so reviewing everything on a floor taking tens of thousands of calls a week would mean a review team several times the size of the one you have. Nobody approves that budget, so you sample.
Sampling looked acceptable when "auditable evidence" meant a binder of reviewed-call summaries on a shelf. It is much harder to defend now that PRIN 2A.9.8R asks firms to regularly monitor the outcomes their customers are experiencing, and clients expect programme-level visibility from BPOs in real time. The rule sets no sample size. But a binder of 5% tells you what those calls contained, not what the other 95% did, and that gap is the thing the question is really about.
Why hiring more QA reviewers does not scale#
The intuitive response to a QA bottleneck is "hire more reviewers". It does not work, for three reasons.
Linear cost, sublinear value. Doubling reviewers doubles cost. It also doubles the variance between reviewers, since two senior reviewers reading the same call will score it differently. Inter-rater reliability is rarely measured because it usually exposes a depressing number. The more reviewers you hire, the more important calibration becomes, and the time you spend calibrating cuts into the time available to review.
The next call is not the bottleneck call. Reviewers cannot pick the calls that matter most because they cannot see what is on a call before they listen to it. They sample randomly or by length, hoping to catch breaches. The breaches you actually need to find are on the calls outside the sample.
Reviewers burn out. Listening to 30 calls a day for compliance failures is grim work. Senior reviewers leave for coaching roles, agent ops roles, or out of the industry. The team you build erodes faster than you can rebuild it.
Why coverage stops being a budget decision#
AI call scoring removes the unit-cost barrier that made sampling necessary. Scoring becomes machine work billed by the minute of audio rather than an hour of a senior reviewer's attention, so how much you review stops being a budget decision and becomes a configuration one.
The cost is not the headline though. The headline is what you can do with 100% coverage that you cannot do with 5%. We covered the foundations of how AI call QA works in a separate post, but here are the operational consequences specific to contact centres.
What 100% coverage actually unlocks#
Breach detection during the call, not weeks later. A critical compliance failure (urgency language, an ignored objection to marketing, a missed mini-Miranda) is currently caught when a customer complains weeks later. With live mid-call breach detection, a high-confidence breach fires a webhook to your CRM or agent desktop while the call is still happening. Coaching can be in the next call rather than the next month.
Real coaching, on every agent, every week. Per-call coaching drafts feed into per-agent coaching memory. The coaching the agent receives next time builds on what was said last time. If they have improved on the flagged area, the AI acknowledges it. If they have not, the language escalates. Manual QA cannot do this because no human has the time to remember every coaching note for every agent.
Agent risk profiles, accurate to the call. A heatmap of which agents are driving which breach types becomes possible because every call or sale is scored. Today, agent performance reviews are anchored on three or four reviewed calls and a wall of subjective impressions. With AI scoring, agent risk is a number computed across hundreds of conversations, not a folk theory.
Client reporting. BPOs running several client programmes can give each client its own scorecard, then pull that client's scores, breaches and pass rates over the REST API into the reporting they already send, rather than assembling monthly PDFs by hand. A single scored call can also be shared with a client through a read-only link that expires. More on contact centre and BPO QA.
Capacity decoupled from QA headcount. If the QA bottleneck disappears, you can grow agent headcount without first hiring more reviewers. The unit economics of adding seats stops including a QA staffing tax.
The architecture of 100% coverage#
What does the system actually look like? Five components, each of which is built into the platform today.
1. Audio ingestion
Calls have to get into the system. There are three patterns:
- Live streaming over WebSocket from your dialer. Twilio Media Streams and Amazon Connect live media streaming via a customer-run Lambda bridge are the most common. Our Twilio integration and AWS Connect integration walk through the wiring. Other dialers connect via a generic WebSocket protocol.
- Batch upload of recordings via REST API or the dashboard. Useful for legacy stacks where streaming is not available, or for archived recordings being scored retrospectively.
- Drag-and-drop for one-off calls being reviewed in the dashboard.
The streaming path adds the ability to do mid-call breach detection. The batch path is otherwise functionally identical, just without the live alerts.
2. Speech recognition with diarisation
Production-grade ASR converts audio to a transcript with speaker separation. Word-error rate matters less than diarisation accuracy and keyterm boosting (so product names and acronyms specific to your business get recognised correctly). The output is a structured transcript with agent and customer turns, plus timestamps.
3. LLM scoring against your scorecard
The transcript goes to a large language model with your scorecard. The model scores each criterion, returns a verdict, and returns the transcript evidence that justifies it, or records that no relevant evidence was found. The evidence is the part that matters: a score with no evidence is unauditable.
4. Per-tenant calibration
Your compliance team corrects AI verdicts they disagree with, marks gold-standard calls as exemplars, and provides written guidance through a knowledge base. All three feed into the prompt the next time the AI scores: up to five of the most recent corrections on a criterion are shown to it as examples. They are examples rather than rules, so the AI can still disagree — but it has your senior reviewer's reading of the scorecard to draw on, not only the scorecard's words.
5. Output integration
Results need to land where the people who act on them are working. That means three integration paths:
- Dashboard for QA leads, ops directors and compliance officers who want to review individual calls.
- Webhooks, HMAC-signed once you set a signing secret, for live breach alerts to your CRM, agent desktop or Slack.
- REST API to pull scores back per call so you can render them inside your own client portal, agent desktop or case-management system.
What QA staff do once full coverage is the default#
The honest answer: they do better work. The 1% sample reviews disappear because the AI does them all. What remains is high-leverage human work that AI cannot do.
Calibration of the AI. Senior reviewers correct edge cases the AI got wrong. Each correction becomes one of the examples the AI is shown the next time it scores that same criterion. This is the highest-leverage QA work in the new world: one reviewer's judgement on a hard case becomes an example other calls on that criterion are checked against. One person's judgement scales across the floor.
Coaching delivery. The AI generates the coaching draft. A human delivers the coaching conversation. The shape of coaching shifts from "we listened to your call yesterday and noticed" to "we have noticed across your last 40 calls that". Coaching becomes data-driven rather than anecdote-driven, and that lands harder with agents.
Compliance review of flagged calls. The AI flags critical breaches; humans review them, decide on severity, escalate where needed, and document the resolution. The breach register becomes the operational artefact, not the spreadsheet of reviewed calls.
Programme-level analysis. Patterns across the floor: which scorecard items are getting harder to pass, which agents are improving, which campaigns have rising breach rates. AI insights briefs, generated on request, give compliance leads a strategic view that nobody had time to build manually.
Working out the case for your own floor#
You can do this with numbers you already have, and we would rather you used yours than ours.
Take a reviewer's fully loaded hourly cost and how many calls one reviewer scores in an hour: that gives you what a reviewed call costs you today. Multiply by the calls you review each month, and you have the cost of your current sample. Then ask what the calls outside it are worth: a missed breach that reaches a client or a regulator, the coaching that never happened, the complaint that could have been caught. Ask us for a quote sized to your call volume and set the two against each other.
That said, ROI on a tool the regulator now expects is the wrong metric. The right metric is whether your firm can credibly answer "how do you systematically monitor outcomes" without flinching. Sampling does not.
Where to start#
If your contact centre or BPO is at the point where adding more QA staff has diminishing returns, the right next step is to see what AI scoring actually catches that your sample misses. In a short demo we score synthetic calls against a scorecard like yours, so you see your scorecard, your breaches, and your coaching drafts on the call. When you want to try your own recordings, we put a DPA in place first. Email hello@callguardai.co.uk and we will book it in.