Summarize with AI
Quality assurance in a call center is the routine of sampling real customer interactions, scoring them against a published standard, and feeding what you find back into coaching and process change. It is a measurement system with a coaching loop attached, and the loop is the part most programs skip.
A scoring form on its own produces a number nobody acts on. Agents learn to game it, team leads treat it as paperwork, and the score drifts up while customer satisfaction does not move. That gap between a rising internal score and a flat external one is the clearest sign a program has stopped working.
This guide covers what belongs in a quality assurance program, which metrics earn a place on the scorecard, and 21 practical tips grouped by the stage you are at.

TL;DR
Quality assurance in a call center means sampling interactions, scoring them against a written standard, and turning the findings into coaching. A workable starting cadence is four to six scored interactions per agent per month, with a calibration session every two weeks so scores mean the same thing across every evaluator.
The failure mode to watch for is a quality score climbing while customer satisfaction stays flat, which means you are measuring compliance with a form rather than the experience. If your floor has fewer than about 15 agents, a formal program is overhead. Score a handful of calls a week yourself and spend the saved time on coaching instead.
Key takeaways
- Quality assurance is a coaching loop, not a scoring exercise, and the loop is what produces the improvement.
- Calibrate evaluators every two weeks or your scores measure the evaluator rather than the agent.
- A rising internal score with flat customer satisfaction means the form is wrong, not that the agents improved.
- Score four to six interactions per agent per month as a starting cadence and adjust from the variance you see.
- Compliance items belong on the same scorecard as tone and accuracy so they get audited weekly.
- Automation should handle scoring volume and transcript search, and leave the judgment calls to a person.
- Below roughly 15 agents, a formal QA program costs more than the coaching it enables.
Table of contents
- What quality assurance in a call center is
- What a quality assurance program is for
- The four parts of a QA program
- Metrics a quality assurance scorecard should track
- Seven tips to build the foundation
- Seven tips to monitor and measure
- Seven tips to coach and improve
- Compliance inside quality assurance
- How to roll out a program in 90 days
- Quality assurance mistakes to avoid
- Call center quality assurance FAQ
- The bottom line
What quality assurance in a call center is
Quality assurance in a call center is the systematic evaluation of customer interactions against a documented standard, followed by coaching that closes the gaps the evaluation finds. It applies to calls, chats, emails, and messaging, and the standard should be written down where agents can read it before they are scored against it.
Three things distinguish a real program from a scoring habit. The standard is published in advance. Evaluators are calibrated against each other on a fixed schedule. Every score produces either a coaching conversation or a process change, and the program tracks which.
Everything else is implementation detail. The 21 tips below are ordered by the stage most centers are at, so start at the section that matches yours rather than at tip one.
What a quality assurance program is for
The point is to find the small number of repeatable behaviors that move outcomes, then get everyone doing them. That is narrower than it sounds. Most scorecards carry twenty line items when four or five actually predict whether the customer got what they needed.
A working QA program produces four outputs. It tells you which agents need coaching and on what. It tells you which processes fail regardless of who is handling them. It gives you evidence for compliance audits. It shows you which of your standards no longer match what customers care about.
If your program is not producing all four, the most common cause is that nobody owns the process-change output. Agent coaching gets an owner by default. Process findings pile up in a spreadsheet because they belong to a department that does not sit in the call center.
The four parts of a QA program
Every functioning QA program has the same four parts. Missing any one of them explains most of the reasons programs stall.
Clear objectives
Tie the program to two or three business outcomes and write them down. Higher first contact resolution, fewer compliance findings, and better retention on a specific journey are examples. Without stated objectives, the scorecard drifts toward whatever is easy to observe, which is usually greeting wording and call closing rather than whether the problem got solved.
A written evaluation framework
The framework defines what a good interaction looks like, how each item is weighted, who evaluates, how often, and how disputes are resolved. Publish it. An agent who can read the standard before being scored against it will argue with the score far less than one who sees it for the first time in a coaching session.
Regular monitoring
Sample consistently rather than heavily. Four to six interactions per agent per month is a workable start, weighted toward the contact reasons that carry the most risk or volume. Mix random samples with targeted ones. Random tells you the baseline, targeted tells you whether a known problem is getting better.
A feedback loop that closes
Every evaluation ends in one of three places. A coaching conversation, a process change request, or a note that the standard itself needs updating. Track how many of each you generate and how many actually get resolved. A program with a hundred evaluations and no closed process changes is measuring, not improving.
Metrics a quality assurance scorecard should track
Pick a small set and understand what each one distorts when it is over-weighted. Every call center metric can be gamed, and agents will optimize for whatever you actually reward.
| Metric | What it measures | How quality assurance uses it | Failure mode when over-weighted |
|---|---|---|---|
| First contact resolution | Issues closed without a repeat contact | Targets coaching at the causes of repeats | Agents discourage callbacks instead of resolving |
| Customer satisfaction | The customer view of the interaction | Validates whether the scorecard reflects reality | Survey gaming and selective solicitation |
| Average handle time | Time spent per interaction | Flags outliers worth listening to in both directions | Rushed calls and higher repeat contact rates |
| Quality score | Adherence to the written standard | Drives individual coaching plans | Score inflation with no change in outcomes |
| Compliance adherence | Required disclosures and consent handling | Produces the audit trail | Scripted recitation the customer ignores |
| Transfer rate | Contacts moved to another team | Separates routing faults from skill gaps | Agents hold calls they cannot resolve |
Read the pairs together rather than individually. Quality score against customer satisfaction tells you whether your standard is right. Average handle time against first contact resolution tells you whether speed is costing you resolutions.
Seven tips to build the foundation
These seven tips apply whether you are starting a QA program from scratch or rebuilding one that stopped producing results.
1. Train continuously rather than at onboarding
Onboarding training decays within a quarter. Replace the annual refresher with short lessons tied to actual findings, so that when evaluations show three agents missing the same disclosure, the lesson goes out that week rather than in the next training cycle.
2. Define KPIs you can act on
A metric earns its place only if you can name the coaching action it triggers. First contact resolution triggers repeat-cause analysis. Compliance adherence triggers a script review. If a number has no action attached, it belongs on a report and not on the scorecard.
3. Review the evaluation criteria quarterly
Standards go stale faster than most teams expect. Every quarter, check each line item against the last quarter of customer comments and remove anything that no longer correlates with a good outcome. A shorter form scored honestly beats a long one scored quickly.
4. Build criteria per channel
Voice, chat, and email fail in different ways. Voice QA looks at tone, pacing, and interruption. Chat looks at response gaps and copy-paste overuse. Email looks at completeness and whether the reply actually closes the loop. One shared form across all three will miss what matters in each.
5. Give agents a knowledge base worth using
A large share of quality failures are information failures rather than behavior failures. Before you coach someone for giving a wrong answer, check whether the right answer was findable in under thirty seconds. If it was not, that is a content ticket rather than a coaching note.
6. Involve agents in writing the standard
Agents who helped write the scorecard defend it. Agents who received it argue with it. Run a working session with three or four experienced agents before you publish any new form, and change at least one thing based on what they say.
7. Publish scorecards agents can see
Show each agent their own scores, the team median, and the trend over the last six evaluations. Visibility does more for consistency than any incentive scheme, and it removes the suspicion that scoring is arbitrary. Keep individual comparisons private to the agent and their lead.
Fewer manual reviews
Let AI handle the calls that never needed a person
Bigly Sales AI agents take routine and after-hours contacts so your team handles the conversations that need judgment. A working setup takes days rather than a quarter.
Seven tips to monitor and measure
This group covers how you gather evidence. Everything here exists to make QA findings defensible rather than anecdotal.
8. Use monitoring tools to widen the sample
Manual review reaches a small percentage of interactions. Monitoring tools that flag keywords, silence, overtalk, and sentiment shifts let you review the interactions most likely to be interesting rather than the ones that happened to come up in a random sample.
9. Use speech and text analytics for themes, not verdicts
Analytics is reliable at surfacing patterns across thousands of interactions and unreliable as a final judgment on any single one. Use it to find the top five recurring problems this month, then have a person listen to a handful of examples before you act.
10. Score a balanced set of items
A scorecard that only measures efficiency produces fast, unhelpful calls. One that only measures warmth produces long, pleasant calls that resolve nothing. Weight resolution, accuracy, compliance, and manner deliberately, and write the weightings on the published form.
11. Calibrate every two weeks
Have every evaluator score the same interaction independently, then compare. Variance above roughly one point on a five-point item means your standard is ambiguous rather than that an evaluator is wrong. Fix the wording. Calibration is the single highest-return habit in call center QA.
12. Bring customer feedback into the scorecard
Match your quality scores against survey responses for the same interactions. Where they disagree consistently, the scorecard is measuring the wrong thing. This comparison is the only reliable check that your internal standard reflects the external experience.
13. Use mystery shopping sparingly
A quarterly round of mystery contacts catches things internal review misses, particularly on the paths customers use that you never test. Keep it occasional. Used monthly it becomes theatre, and agents start recognizing the pattern.
14. Automate scoring volume, not judgment
Automated scoring handles objective items well, including disclosure presence, hold-time adherence, and required verification steps. Leave empathy, judgment, and de-escalation to human review. Automating the objective half typically frees enough evaluator time to double the sample on the half that matters.
Seven tips to coach and improve
Everything to this point produces evidence. This group turns evidence into behavior change, which is the part that determines whether the QA program was worth running.
15. Coach toward first contact resolution
Take the repeat contacts from last month, code why the first attempt failed, and coach against those causes specifically. Most centers find two or three dominant causes, usually a missing authority to resolve, a knowledge gap on one product, or a handoff that drops context.
16. Close the loop with the agent
Deliver feedback within a few days of the interaction while the agent still remembers it, and end every session with one specific thing to change before the next evaluation. Feedback delivered a month later is a performance review, not coaching.
17. Train soft skills with real recordings
Empathy, patience, and active listening improve faster when agents hear real examples than when they read a definition. Build a small library of anonymized recordings that show a difficult contact handled well, and use those instead of role-play scripts.
18. Run peer reviews alongside formal scoring
Agents reviewing each other spot things evaluators miss, particularly phrasing that technically passes the form and still frustrates the customer. Keep peer reviews out of formal ratings so they stay useful rather than becoming a second appraisal.
19. Ask agents to self-score first
Have the agent score their own interaction before the evaluation session. Where their score and yours differ, you learn exactly which part of the standard is unclear. Self-scoring turns the session from a verdict into a comparison.
20. Keep the standard current with practice
Customer expectations and the tools available to your team both move. Review what has changed once a quarter and update the form, particularly around automation, since expectations for what a bot should handle before a human steps in have shifted quickly. The Bigly Sales AI calling glossary is a useful shared vocabulary when updating standards that involve automated handling.
21. Study wins as carefully as failures
Most programs analyze bad interactions and celebrate good ones. Reverse half of that. Take the three highest-scoring interactions each month and break down what the agent actually did, then put those specifics into the next training cycle. Recognition works better when it is precise about what was good.
Compliance inside quality assurance
Compliance items belong on the same scorecard as everything else, audited weekly rather than in an annual review. That is the difference between finding a problem in a sample of six calls and finding it in a regulator letter.
For outbound teams the baseline is the FTC guidance on complying with the Telemarketing Sales Rule, which covers calling windows, required disclosures, and Do Not Call obligations. Data handling has its own baseline in the FTC guide to protecting personal information, which is worth reading before you decide who can access recordings.
Two consent points are commonly stated wrong. The FCC one-to-one consent rule was vacated in January 2025 and never took effect, so prior express written consent under the TCPA remains the operative standard for automated and artificial voice marketing calls. Separately, the cross-channel revocation provision, under which a revocation received on one channel applies to all channels from that sender, takes effect on January 31, 2027. Build revocation handling into your QA checks now rather than close to that date.
Recording consent varies by state, and several states require all parties to consent. The Bigly Sales legal and compliance overview covers how these obligations apply to automated calling. Confirm your own program with counsel.
How to roll out a program in 90 days
Days 1 to 30. Write the standard and publish it. Keep it to no more than ten scored items. Run a working session with three experienced agents and change at least one item based on their input. Score a pilot set of interactions without sharing individual results, purely to see where the form is ambiguous.
Days 31 to 60. Start formal scoring at four interactions per agent per month. Run your first calibration session in week five and a second in week seven. Expect meaningful disagreement in the first two sessions. That disagreement is the form talking, not the evaluators.
Days 61 to 90. Begin coaching conversations tied to the scores and start logging process-change requests separately from coaching notes. At day 90, compare quality scores against customer satisfaction for the same interactions. If they disagree, revise the form before you scale the program.
Quality assurance mistakes to avoid
Scoring more interactions than you can coach on. Evaluations without follow-up produce paperwork and resentment in equal measure.
Letting the quality score become the goal. When the score rises and customer satisfaction does not, the program has started measuring itself.
Skipping calibration because everyone is busy. Uncalibrated scores measure the evaluator, and agents work that out within a month.
Using QA findings in disciplinary action without warning. Do that once and honest self-assessment ends permanently.
Building one scorecard for every channel. Voice, chat, and email fail differently, and a shared form hides all three sets of failures.
Owning coaching findings but not process findings. Process problems that belong to another department will sit unresolved unless someone in the call center is accountable for chasing them.
Call center quality assurance FAQ
What is quality assurance in a call center?
It is the practice of sampling customer interactions, scoring them against a written and published standard, and using the findings to coach agents and fix processes. It covers calls, chat, email, and messaging. The scoring half is straightforward to set up. The half that determines whether the program works is the loop that turns each finding into either a coaching conversation or a process change request.
How many calls should you score per agent per month?
Four to six is a workable starting point for most centers, weighted toward the contact reasons carrying the most volume or risk. Increase the sample where you see high variance between an agent’s individual evaluations and reduce it where scores are consistently stable. The practical limit is not how many you can score, it is how many coaching conversations you can actually hold afterward.
What is calibration and why does it matter?
Calibration means having every evaluator independently score the same interaction, then comparing results. Variance above roughly one point on a five-point item usually indicates the standard is worded ambiguously rather than that an evaluator is mistaken. Without calibration, an agent’s score depends on who happened to review them, agents notice that quickly, and the credibility of the whole program goes with it.
Which metrics belong on a QA scorecard?
First contact resolution, customer satisfaction, quality score, compliance adherence, and transfer rate cover most needs. Average handle time is worth watching as a diagnostic but is dangerous as a target, because rewarding it produces rushed interactions and more repeat contacts. Read metrics in pairs. Quality score against customer satisfaction shows whether your standard reflects what customers actually experienced.
Can AI replace manual call scoring?
Partly. Automated scoring is reliable on objective items such as whether a required disclosure was given, whether verification steps were completed, and whether hold-time rules were followed. It is unreliable on empathy, judgment, and de-escalation. The productive split is to automate the objective half and reinvest the freed evaluator time into reviewing the subjective half across a larger sample.
How do you stop agents from gaming the scorecard?
Assume they will optimize for whatever you reward, then design accordingly. Pair every efficiency measure with a quality measure, check quality scores against customer survey responses for the same interactions, and revise any line item that can be satisfied without helping the customer. Involving agents in writing the standard also reduces gaming, because a standard people helped build is harder to treat cynically.
Does a small call center need a QA program?
Below roughly 15 agents, a formal QA program usually costs more than it returns. A team lead scoring a handful of interactions a week and coaching directly gets most of the benefit without the overhead of forms, calibration schedules, and reporting. Formalize when you can no longer keep the whole floor’s performance in your head, which for most managers happens somewhere between 15 and 25 agents.
How should compliance be handled in quality assurance?
Put compliance items on the same weekly scorecard as tone and accuracy rather than reviewing them annually. For outbound teams that means calling windows, required disclosures, consent capture, and Do Not Call handling. Recording consent rules vary by state and several require all parties to consent. Treat the scorecard as your audit trail and confirm your specific obligations with counsel.
How long before a QA program shows results?
Expect the first useful findings within 30 days and measurable behavior change closer to 90. The first month tells you mostly about your form rather than your agents, since early disagreement between evaluators is a wording problem. Real improvement starts once calibration has settled and coaching conversations are happening within a few days of the interaction being scored.
The bottom line
Quality assurance in a call center works when the loop closes. Sample consistently, calibrate the people doing the scoring, and make sure every finding produces either a coaching conversation or a process change that somebody owns. A program that generates scores and nothing else will raise its own numbers and move nothing that customers notice.
Start smaller than you think you need to. Ten scored items, four interactions per agent per month, and a calibration session every two weeks will teach you more in a quarter than a comprehensive framework you cannot sustain. If the volume of routine contacts is what makes the sample unmanageable, look at what automation can absorb on the Bigly Sales pricing page before you add evaluator headcount.
Cut the routine volume
Give your reviewers fewer calls worth reviewing
See which contact types an AI agent can take off your queue so your quality work concentrates on the conversations that matter. The walkthrough runs about 20 minutes.







