Most QA teams review a handful of conversations per agent each month. It feels thorough. You listen closely, score carefully and write good notes. But a small sample can only tell you so much, and it helps to know exactly how much before you make decisions on it.
What a sample can tell you
Say an agent handles 600 conversations a month and you review 5 of them. They pass 4. Is their pass rate 80%?
Probably not exactly. With 5 conversations, the honest answer is “somewhere between about 38% and 96%”, at the usual 95% confidence. The range is that wide because 5 is a small number.
Here is how the margin shrinks as the sample grows, for a pass rate around 50%, where the uncertainty is largest:
| Conversations reviewed | Margin of error (95% confidence) |
|---|---|
| 5 | ± 44 points |
| 10 | ± 31 points |
| 20 | ± 22 points |
| 50 | ± 14 points |
| 100 | ± 10 points |
| 400 | ± 5 points |
Two things stand out. First, a few reviews per agent can’t separate a good agent from an average one. Their ranges overlap almost completely. Second, precision gets expensive fast. Halving the margin of error means reviewing four times as many conversations.
What a sample will miss
Pass rates are only half the story. QA also exists to catch things that shouldn’t happen: a missed disclosure, a refund promised without approval, a customer mentioning they’re struggling to pay.
These are usually rare. Suppose a problem appears in 2% of an agent’s conversations. The chance that your sample contains even one of them:
| Conversations reviewed | Chance of seeing it at least once |
|---|---|
| 5 | 10% |
| 10 | 18% |
| 20 | 33% |
| 50 | 64% |
With 5 reviews a month, nine months out of ten you won’t see it at all. And when you do, it will look like a one-off, because you have nothing to compare it with.
For problems that matter to a regulator or an ombudsman, “we would probably have caught it eventually” is a hard position to defend.
How to sample better, if you sample
Sampling isn’t wrong. It is the only option when every review takes a person ten minutes. If you rely on it, you can make it work harder:
- Pick at random. Letting reviewers choose conversations, or taking the first few of the month, skews results towards the easy or the memorable.
- Sample more from new starters and new processes. That is where problems are most likely.
- Target known risk. Add every conversation that mentions a complaint, a cancellation or financial difficulty, on top of the random sample. Keep the two apart when you report, or the pass rate will look worse than it is.
- Report ranges, not single numbers. “Between 60% and 95%” is less satisfying than “80%”, but it stops people drawing conclusions the data can’t support.
- Calibrate reviewers. If two reviewers would score the same conversation differently, the sample size hardly matters. Score a few shared conversations together every month.
When to stop sampling
The case for checking every conversation gets stronger when:
- You have more than a few agents, so the sample per agent is tiny.
- You operate under rules where a single miss is costly, such as disclosures, vulnerable customers or complaint handling.
- Your reviewers spend most of their time listening to conversations that turn out to be fine.
Checking everything doesn’t mean a person reads everything. It means every conversation is checked against the same rules, and people spend their time on the ones that failed or are unclear. The sample becomes a spot check on the automatic results, not the whole picture.
That is how Assay works: every conversation is checked against your scorecard, each result points at the line that decided it, and anything uncertain goes to a person. A small random share of passed conversations goes to a person too, so you can see how often the automatic results agree with your team.
The short version
- A few reviews per agent can’t tell you their real pass rate. Expect ranges of 20 to 40 points either way.
- A problem in 2% of conversations will go unseen most months with a small sample.
- If you sample, sample at random, add risky conversations separately and report ranges.
- If single misses are costly, check every conversation and use people for the ones that need judgement.