Autonomous Agent Workflows & Operations AutomationPlaybook3 min readUpdated September 2026

Quality Assurance Sampling: How Much to Check, and How

Reviewing every single interaction for quality is rarely realistic once volume grows past a small team, and reviewing whatever happens to be convenient produces a biased picture that tells you less than no review at all, because it creates false confidence. Sampling done well sits between those two failure modes: enough coverage to catch real patterns, selected in a way that doesn't quietly favor easy cases.

The method matters as much as the size. A random ten percent sample and a targeted ten percent sample looking specifically at flagged or escalated interactions answer different questions, and conflating the two is a common source of QA programs that miss the problems that actually matter.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you pick a QA sample size?

A flat rule like 'review five percent of everything' is easy to explain but arbitrary, and it often reviews too little of a low-volume, high-risk category while reviewing far more than necessary of a high-volume, low-risk one. Instead, size the sample to what you're actually trying to catch: if a defect happens roughly one in twenty times and you want a reasonable chance of catching at least one instance per week, the sample needs to be large enough to make that statistically plausible, not just a tidy percentage.

For most small operations teams, this means sampling more heavily on newer processes or newer team members until quality stabilizes, then tapering off as confidence builds, rather than a flat rate applied uniformly forever.

What is the difference between random and targeted QA sampling?

Random sampling tells you the baseline quality level across everything, which is what you need for a general health check. Targeted sampling, reviewing every escalation, every low customer rating, or every interaction flagged by a specific rule, tells you about your worst cases specifically, which is what you need to find and fix root causes. Run both, and don't let one substitute for the other; a program that only reviews escalations will systematically miss quiet, undetected problems in the interactions that never got flagged.

Build a rubric before you start scoring anything

Score against a written rubric with specific, observable criteria, not a general impression of whether an interaction 'felt right.' A rubric with three or four concrete dimensions, accuracy, completeness, tone, and whether the resolution actually addressed the underlying issue, produces scores that different reviewers apply consistently. Without a written rubric, two reviewers scoring the same interaction routinely disagree, which makes trend data across a team meaningless.

Score each interaction on these dimensions:

  • Accuracy: whether the information given or action taken in the interaction was actually correct.
  • Completeness: whether every part of the customer's request was addressed before the interaction ended.
  • Tone: whether the language and manner fit the situation and the customer.
  • Resolution: whether the outcome addressed the underlying issue, not just the question as first asked.

Calibrate reviewers against each other periodically

Even with a good rubric, reviewers drift apart over time in how strictly they apply it. Have two or three reviewers independently score the same handful of interactions every month and compare results, discussing any meaningful gap until they converge on a shared interpretation. Skipping calibration is how a QA program slowly loses credibility: a team member scored harshly by one reviewer and leniently by another has a legitimate reason to distrust the whole process.

Turn findings into a workflow, not just a score

A QA score that lives in a spreadsheet and never reaches the person whose work it describes doesn't improve anything. Route flagged interactions into a workflow tool like Process Street or ClickUp so the reviewer's notes, the specific example, and a required follow-up conversation are all attached to one record, not scattered across a scorecard and a separate coaching conversation nobody documented.

Watch for sampling bias hiding in convenience

The easiest interactions to review are often the shortest, best-documented, or most recently completed ones, and a sample built from whatever's easiest to grab systematically underrepresents the messy, long, or older cases where problems are more likely to hide. Pull your sample from a genuinely randomized list, not the top of whatever queue happens to be open, and periodically check whether your sample's average length and complexity actually still match the full population it's supposed to represent.

A quick gut check: if every reviewed case this month finished in under ten minutes, ask where the longer, messier cases went instead.

Executive Capability Standard

What Good Looks Like

A working QA sampling program uses both random and targeted samples against a written rubric, calibrates reviewers against each other regularly, and routes every finding into a tracked follow-up, not just a score.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review a sample of recent interactions by hand against a draft rubric to see how consistent scoring feels before formalizing anything.
2. Do Manually:Run manual random and targeted sampling for a month to establish a baseline before building any automated routing.
3. Delegate:Assign an independent reviewer, not the direct manager, to score interactions and flag patterns for a coaching conversation.
4. Automate:Route flagged interactions automatically into a workflow tool like Process Street or ClickUp so nothing depends on someone remembering to follow up.
5. Buy:Bring in outside QA expertise to build the initial rubric and sampling method if nobody internally has designed one before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

What sample size is enough for a small operations team?

It depends on volume and defect rarity: review enough each week to expect a couple of instances of a defect that occurs about five percent of the time. Adjust the sample up for newer processes and down once quality has stabilized, rather than holding one flat rate forever.

Should managers review their own team's QA scores?

They should see the scores, but an independent reviewer, not the direct manager, should do the initial scoring to avoid an unconscious bias toward leniency or harshness based on the working relationship. The manager's role is acting on the findings in a coaching conversation, not producing the score itself.

How often should the QA rubric be updated?

Review it whenever the underlying process it's scoring changes meaningfully, and otherwise leave it stable for at least a couple of quarters so trend data stays comparable. A rubric that changes every month makes it impossible to tell whether quality actually improved or the measuring stick just moved.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides