Skip to content

QA sampling: how many conversations do you need to review?

What a sample of 5, 20 or 50 conversations can and can’t tell you, with the numbers behind it, and when to check every conversation instead.

Blog· QA programmes· · 3 min read

Most QA teams review a handful of conversations per agent each month. It feels thorough. You listen closely, score carefully and write good notes. But a small sample can only tell you so much, and it helps to know exactly how much before you make decisions on it.

What a sample can tell you

Say an agent handles 600 conversations a month and you review 5 of them. They pass 4. Is their pass rate 80%?

Probably not exactly. With 5 conversations, the honest answer is “somewhere between about 38% and 96%”, at the usual 95% confidence. The range is that wide because 5 is a small number.

Here is how the margin shrinks as the sample grows, for a pass rate around 50%, where the uncertainty is largest:

Conversations reviewed Margin of error (95% confidence)
5 ± 44 points
10 ± 31 points
20 ± 22 points
50 ± 14 points
100 ± 10 points
400 ± 5 points

Two things stand out. First, a few reviews per agent can’t separate a good agent from an average one. Their ranges overlap almost completely. Second, precision gets expensive fast. Halving the margin of error means reviewing four times as many conversations.

What a sample will miss

Pass rates are only half the story. QA also exists to catch things that shouldn’t happen: a missed disclosure, a refund promised without approval, a customer mentioning they’re struggling to pay.

These are usually rare. Suppose a problem appears in 2% of an agent’s conversations. The chance that your sample contains even one of them:

Conversations reviewed Chance of seeing it at least once
5 10%
10 18%
20 33%
50 64%

With 5 reviews a month, nine months out of ten you won’t see it at all. And when you do, it will look like a one-off, because you have nothing to compare it with.

For problems that matter to a regulator or an ombudsman, “we would probably have caught it eventually” is a hard position to defend.

How to sample better, if you sample

Sampling isn’t wrong. It is the only option when every review takes a person ten minutes. If you rely on it, you can make it work harder:

When to stop sampling

The case for checking every conversation gets stronger when:

Checking everything doesn’t mean a person reads everything. It means every conversation is checked against the same rules, and people spend their time on the ones that failed or are unclear. The sample becomes a spot check on the automatic results, not the whole picture.

That is how Assay works: every conversation is checked against your scorecard, each result points at the line that decided it, and anything uncertain goes to a person. A small random share of passed conversations goes to a person too, so you can see how often the automatic results agree with your team.

The short version

Try it on your own calls

Tell us about your team. We’ll reply within two working days to set up a trial with 50 of your transcripts and your current scorecard.

  • No card or contract for the trial
  • Your conversations are never used to train models
  • We go through the results with you