Most QA teams would check every conversation if they could. The reason they don’t is time: a careful review of one call takes ten minutes or more, and nobody has the budget to multiply the team by fifty.
Automatic checking changes that maths. But switching from a small sample to every conversation isn’t only a tooling change. It changes what your reviewers do, what your reports mean and what your agents experience. This is a practical plan for making the move without losing trust along the way.
Step 1: Decide what “checked” means
Before any tool touches a conversation, write down what a check actually is. For most teams that is the scorecard, but scorecards written for human reviewers often rely on judgement that isn’t written anywhere.
Go through each line and ask: could two reviewers read this line and the same conversation, and reach the same answer? If not, rewrite it until they would. We cover how in How to write a QA scorecard that can be checked automatically.
Expect to split a few lines into two and drop a few that are really coaching topics rather than checks. “Showed empathy” is worth coaching; “acknowledged the customer’s problem before offering a fix” is something you can check.
Step 2: Start with the checks that carry risk
Don’t automate the whole scorecard on day one. Start with three to five checks where a miss is costly and the rule is clear:
- Identity confirmed before account details are discussed
- Call recording disclosed near the start
- Complaints acknowledged and logged
- Prices and fees stated before the customer agrees
- Signs of financial difficulty or vulnerability spotted
These are the checks where a small sample hurts most, because a single miss matters and misses are rare.
Step 3: Run both side by side
For the first few weeks, keep your manual reviews going and compare them with the automatic results on the same conversations. You are looking for two things:
- Agreement. How often does the automatic result match your reviewer? Track this per check, not overall. A check that agrees 97% of the time is ready. One that agrees 70% of the time needs rewriting.
- Disagreements worth reading. When they differ, who was right? Often the reviewer missed something on a long call. Sometimes the rule was ambiguous. Both are useful.
Don’t aim for perfect agreement. Human reviewers don’t agree with each other perfectly either. Aim for agreement at least as good as your reviewers manage with each other.
Step 4: Let uncertainty go to a person
The single most important design choice is what happens when the automatic check isn’t sure. A system that always gives an answer will sometimes give a confident wrong one, and people stop trusting it the first time they notice.
A better arrangement is three outcomes per check: yes, no, or not sure. Clear results are recorded. Unclear ones land in a review queue. Your reviewers now spend their time only where a judgement is needed.
Keep a small random share of passed conversations in the queue too. Those spot checks tell you whether the automatic passes can be trusted, and they keep the agreement figure honest.
Step 5: Change what reviewers do
Once coverage is automatic, your reviewers’ job shifts:
| Before | After |
|---|---|
| Listening to routine conversations | Deciding the unclear ones |
| Scoring a few calls per agent | Reviewing patterns across every call |
| Finding problems by luck | Following up on every failure |
| Arguing about sample fairness | Tightening rules that cause disagreement |
Tell the team early. Reviewers who fear being replaced will quietly look for reasons the system is wrong. Reviewers who see it as removing the dull part of the job will help you tune it.
Step 6: Show agents the evidence
Agents accept feedback more easily when they can see exactly what it’s based on. Every result should point at the line in the conversation that decided it, not just a score. “Failed: fee not stated” with the relevant lines highlighted is a conversation starter. “Scored 72%” is an argument.
Coach on patterns rather than single calls. With every conversation checked, you can say “this happened on 14 of your calls this month” instead of “on the one call I listened to”.
Step 7: Report differently
With full coverage, a pass rate is a real rate, not an estimate from five calls. That makes trends meaningful week to week. It also means numbers may look worse at first, because you are now seeing the problems a sample used to miss. Warn leadership before the first report, not after.
Where Assay fits
Assay is built around this plan. You write your scorecard as plain English rules, connect your helpdesk or send transcripts, and every conversation is checked. Each rule is answered yes, no or not sure, each answer points at the line behind it, and anything uncertain goes to your review queue with a random share of passes as spot checks. You can see how often it agrees with your reviewers, rule by rule.