Quality Assurance Sampling: How Much to Check, and How
Reviewing every single interaction for quality is rarely realistic once volume grows past a small team, and reviewing whatever happens to be convenient produces a biased picture that tells you less than no review at all, because it creates false confidence. Sampling done well sits between those two failure modes: enough coverage to catch real patterns, selected in a way that doesn't quietly favor easy cases.
The method matters as much as the size. A random ten percent sample and a targeted ten percent sample looking specifically at flagged or escalated interactions answer different questions, and conflating the two is a common source of QA programs that miss the problems that actually matter.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you pick a QA sample size?
A flat rule like 'review five percent of everything' is easy to explain but arbitrary, and it often reviews too little of a low-volume, high-risk category while reviewing far more than necessary of a high-volume, low-risk one. Instead, size the sample to what you're actually trying to catch: if a defect happens roughly one in twenty times and you want a reasonable chance of catching at least one instance per week, the sample needs to be large enough to make that statistically plausible, not just a tidy percentage.
For most small operations teams, this means sampling more heavily on newer processes or newer team members until quality stabilizes, then tapering off as confidence builds, rather than a flat rate applied uniformly forever.
What is the difference between random and targeted QA sampling?
Random sampling tells you the baseline quality level across everything, which is what you need for a general health check. Targeted sampling, reviewing every escalation, every low customer rating, or every interaction flagged by a specific rule, tells you about your worst cases specifically, which is what you need to find and fix root causes. Run both, and don't let one substitute for the other; a program that only reviews escalations will systematically miss quiet, undetected problems in the interactions that never got flagged.
Build a rubric before you start scoring anything
Score against a written rubric with specific, observable criteria, not a general impression of whether an interaction 'felt right.' A rubric with three or four concrete dimensions, accuracy, completeness, tone, and whether the resolution actually addressed the underlying issue, produces scores that different reviewers apply consistently. Without a written rubric, two reviewers scoring the same interaction routinely disagree, which makes trend data across a team meaningless.
Score each interaction on these dimensions:
- Accuracy: whether the information given or action taken in the interaction was actually correct.
- Completeness: whether every part of the customer's request was addressed before the interaction ended.
- Tone: whether the language and manner fit the situation and the customer.
- Resolution: whether the outcome addressed the underlying issue, not just the question as first asked.
Calibrate reviewers against each other periodically
Even with a good rubric, reviewers drift apart over time in how strictly they apply it. Have two or three reviewers independently score the same handful of interactions every month and compare results, discussing any meaningful gap until they converge on a shared interpretation. Skipping calibration is how a QA program slowly loses credibility: a team member scored harshly by one reviewer and leniently by another has a legitimate reason to distrust the whole process.
Turn findings into a workflow, not just a score
A QA score that lives in a spreadsheet and never reaches the person whose work it describes doesn't improve anything. Route flagged interactions into a workflow tool like Process Street or ClickUp so the reviewer's notes, the specific example, and a required follow-up conversation are all attached to one record, not scattered across a scorecard and a separate coaching conversation nobody documented.
Watch for sampling bias hiding in convenience
The easiest interactions to review are often the shortest, best-documented, or most recently completed ones, and a sample built from whatever's easiest to grab systematically underrepresents the messy, long, or older cases where problems are more likely to hide. Pull your sample from a genuinely randomized list, not the top of whatever queue happens to be open, and periodically check whether your sample's average length and complexity actually still match the full population it's supposed to represent.
A quick gut check: if every reviewed case this month finished in under ten minutes, ask where the longer, messier cases went instead.
What Good Looks Like
A working QA sampling program uses both random and targeted samples against a written rubric, calibrates reviewers against each other regularly, and routes every finding into a tracked follow-up, not just a score.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Process Street works well for running the scoring checklist itself as a repeatable, structured step rather than a freeform note.
ClickUp is a reasonable place to route flagged interactions and required follow-up conversations so nothing gets scored and then forgotten.
Frequently Asked Questions
What sample size is enough for a small operations team?
It depends on volume and defect rarity: review enough each week to expect a couple of instances of a defect that occurs about five percent of the time. Adjust the sample up for newer processes and down once quality has stabilized, rather than holding one flat rate forever.
Should managers review their own team's QA scores?
They should see the scores, but an independent reviewer, not the direct manager, should do the initial scoring to avoid an unconscious bias toward leniency or harshness based on the working relationship. The manager's role is acting on the findings in a coaching conversation, not producing the score itself.
How often should the QA rubric be updated?
Review it whenever the underlying process it's scoring changes meaningfully, and otherwise leave it stable for at least a couple of quarters so trend data stays comparable. A rubric that changes every month makes it impossible to tell whether quality actually improved or the measuring stick just moved.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Cutting Status Meetings by Writing Better Async Updates
How to replace recurring status meetings with written updates that actually give people the information they need, across time zones and without a live call.
Automating the Boring Parts of a New Hire's First Day
A runbook for automating the repetitive account, hardware, and workspace steps of onboarding, while keeping the parts that need a human touch untouched.
Automating Customer Onboarding Without Losing the Human Touch
Which parts of onboarding to automate first, which to keep human, and how to sequence the rollout so time-to-value actually improves.
Governing AI Agents Before They Touch Your Operations
A practical way to decide which operational tasks an AI agent can run unsupervised, which need a human check, and how to document the difference.
Metrics Worth Putting on an Operations Dashboard (and Ones That Aren't)
How to pick operations metrics that actually predict problems, instead of ones that just look active, and how often each one is worth checking.
Zapier or Make for AI Agent Handoffs: A COO's Buying Guide
A practical comparison of Zapier and Make for routing AI agent tasks between tools, with the criteria that actually decide which one fits your team.