
Most collections scorecards are inherited. Somebody adapted a contact-centre QA sheet years ago, added a couple of compliance lines, and the floor has been scoring against it ever since. It produces a number every month, the number is usually in the nineties, and nobody has asked what the number would do under challenge.
The test that matters is narrow and specific: a consumer disputes what was said on a call, and you have to show what actually happened. A scorecard that survives that moment is built differently from one that produces a tidy monthly average.
This is how to build the first kind. It is written for a floor that reviews calls by hand as much as for one running automated scoring, because the design principles are the same either way.
The blended score is the original sin
Nearly every inherited scorecard ends in a single percentage. Ninety-four percent. It feels like information and it is not, for one reason: it mixes categories that are not commensurable.
A collector who was warm, efficient, hit every talk-off beat, and omitted the disclosure does not score 94 percent. They had one compliance failure and a lot of good salesmanship. Averaging those together produces a number that is high, comforting and useless — and worse, it hides the one thing you needed to see behind nine things you did not.
Report per dimension, never blended. A useful scorecard produces something like:
| Dimension | What it answers | How it fails |
|---|---|---|
| Required disclosures | Was the required language actually said, in the opening | Binary — said or not said |
| Prohibited content | Did anything said fall outside what is permitted | Binary — present or absent |
| Consumer requests honoured | Was a stated request acted on within the call | Binary, with the timestamp |
| Identity and third-party handling | Was the right party reached and confirmed before disclosure | Binary |
| Conduct and tone | Was the interaction professional under pressure | Graded — coaching signal |
| Process adherence | Was the account handled per the client's own requirements | Graded — coaching signal |
The first four are pass or fail and each failure is an event that needs a name and a timestamp. The last two are graded because they are genuinely a spectrum, and they are the coaching half of the programme. Reporting them together as one figure destroys both.
That split is also the split between the two audiences for the scorecard. Compliance needs the binary rows. Team leaders need the graded ones. One number serves neither.
Every finding needs evidence attached to it
The difference between a scorecard that survives a dispute and one that does not usually comes down to whether the finding is a claim or a citation.
"Missing disclosure — 0 points" is a claim. Somebody asserted it, a month ago, and to check it you have to go and find the call.
"Missing disclosure at 00:04 — [quoted line]" is a citation. It names the moment, reproduces what was said, and lets anybody form their own view in ten seconds.
Build the second kind. Concretely, that means every compliance finding carries:
- The timestamp, to the second.
- The quoted line the finding refers to.
- A direct link to that moment in the audio, not to the top of the recording.
- Who reviewed it and when, and whether they confirmed or dismissed it.
That last item is the one people leave out and later regret. A finding with no reviewer attached is an assertion with no accountability, and a dismissed finding with no record of who dismissed it is worse than no finding at all. The audit trail is not bureaucracy — it is the thing that makes the rest of the record credible.
Categories should come from the rules, not from the vendor
Here is where inherited scorecards go wrong most often. A generic contact-centre QA tool ships with categories built for service and sales — greeting, empathy, resolution, closing. Those are real things, and none of them is what governs a collection call.
The category set for collections should be traceable to the actual requirements. The Fair Debt Collection Practices Act and its implementing regulation are public — the regulation text sits at the Electronic Code of Federal Regulations, the rulemaking history is on the Federal Register, and the Consumer Financial Protection Bureau publishes examination materials describing what its examiners look at. The Federal Trade Commission publishes annual reporting on collection complaints, which is a useful read for where consumer friction actually clusters.
Calling practices are a separate body of rules again, administered by the Federal Communications Commission, and worth keeping distinct in the scorecard rather than folding into a general compliance bucket.
Two design notes that follow from reading the sources rather than a template:
Your client requirements are a separate category from the regulatory ones. Placement clients frequently impose stricter rules than the law does — specific language, prohibited hours tighter than the regulation, escalation requirements. Those belong on the scorecard, and they belong in their own row, because a client-rule failure and a regulatory failure have different consequences and should not be averaged into a common "compliance" score.
Nothing here is legal advice, including this article. A scorecard is an operational artefact built to reflect obligations your own counsel has identified. Any vendor telling you their category set confers compliance is selling you something that does not exist.
Calibration: the step everyone skips
A scorecard is only as consistent as the people applying it, and two reviewers scoring the same call will disagree more than either expects. Without calibration, a collector's score partly measures which supervisor drew their calls that month — which is both unfair and, if the scores drive any employment decision, a genuine exposure.
Calibration is unglamorous and it works:
- Everyone who scores calls reviews the same three or four calls independently.
- They score without conferring.
- The group compares, and every disagreement gets discussed to a resolution.
- The scorecard definition is amended wherever the disagreement was caused by ambiguous wording rather than by reviewer error.
That fourth step is the one that gets dropped, and it is the only one that compounds. A disagreement is usually evidence that a category is loosely worded. Fixing the wording removes that disagreement permanently, for everyone, on every future call. Fixing the reviewer removes it once.
Monthly is a reasonable cadence. Quarterly is the minimum at which drift stays visible.
Where automation helps, and where it should stay out
Automated scoring is very good at exactly the categories that are binary and language-anchored: was a required phrase present, did prohibited content appear, was a stated request followed by a change in behaviour. Those are detectable and they are also the ones a sample is worst at catching, which is a fortunate overlap.
Automation is much weaker at the graded categories. Whether a collector handled an angry consumer with professionalism is a judgement about a whole interaction, and a model producing a number for it is producing a number, not an assessment.
So the useful division is:
- Machine: transcribe everything, score the binary categories on every call, surface the flags with timestamps and quotes.
- Human: confirm or dismiss each flag, score the graded categories on a sample, run calibration, decide what any of it means.
Our approach follows that split deliberately — findings are reported per dimension and never as one blended score, and every flag plays the few seconds it refers to. The features page sets out what each capability does, and how it works covers the flow from call to morning report.
There are things it does not do, and they are worth stating plainly because they affect scorecard design. Review is next-morning by design, not live, so nothing here supports a supervisor whispering during a call. Calls tie to collectors and times rather than to the account in your platform, so account-level scoring needs an integration that may or may not exist for your dialler. And audio comes off the collector machine, so a desk phone produces nothing to score.
The scorecards that fall apart are the ones where the finding is a checkbox. Somebody ticked "disclosure missing" in March and by June nobody can tell you which call, which second, or whether the reviewer was right. Attach the quote or do not bother scoring the category.
— QA lead, agency collections, 9 years, name withheld by request
Retention, and the question nobody asks until it matters
A scorecard is a record, and records have a retention question attached. Decide three things explicitly rather than by default:
How long recordings are kept. Long enough to cover the dispute and examination windows that apply to your business, and no longer than you have a reason for. Both directions carry risk.
How long the scores are kept, separately. Scores are far smaller than audio and frequently worth keeping longer, because the trend is the thing that shows whether a corrective action worked. Keeping the score after the audio ages out is a reasonable and common design.
Who can see what. Roles and permissions on the QA record are not a nice-to-have. A collector being able to read their own scores is usually right; being able to edit them never is. The National Institute of Standards and Technology publishes access-control and records guidance that is a sensible starting frame even though it is not written for this industry.
A minimum viable scorecard
If you are rebuilding from scratch, this is a defensible starting point. It is deliberately short — long scorecards do not get applied consistently.
Binary, every call, evidence required:
- Required disclosure present in the opening.
- No prohibited content.
- Right party confirmed before any account detail disclosed.
- Stated consumer requests honoured within the call.
- Client-specific mandatory requirements met.
Graded, on a sample, coaching only:
- Professionalism under pressure.
- Clarity and accuracy of what was communicated.
- Process and documentation.
Report 1 to 5 as counts of failures with timestamps. Report 6 to 8 as a trend per collector. Never add them together.
As of August 2026, if you want the binary half of that running on every call rather than a sample, that is what our Voice tier does at $99 per monitored collector per month, with supervisors free; the Essentials tier at $39 covers desktop activity and the daily report without voice capture. The pricing page has the detail, and talk to us if you want to know honestly whether your dialler and your handsets are compatible before you consider it.
Frequently asked questions
What should a collections QA scorecard actually contain?
Two separate groups. A short set of binary compliance categories scored on every call with evidence attached — required disclosures, prohibited content, right-party confirmation, consumer requests honoured, and any client-specific mandates. And a small set of graded coaching categories — professionalism, clarity, process — scored on a sample. Keep them separate in the reporting, because averaging them produces a comfortable number that hides the failures you built the scorecard to find.
Why is a single blended QA score a problem?
Because it mixes categories that are not commensurable. A collector who performed well on nine dimensions and omitted a required disclosure on the tenth still scores in the nineties, so the one finding that carried consequence is buried by nine that did not. Per-dimension reporting costs nothing extra and makes the compliance rows visible on their own terms.
How often should we run calibration sessions?
Monthly is a good cadence and quarterly is the practical minimum. Everyone who scores calls reviews the same three or four calls independently, then the group resolves every disagreement. The step people skip is amending the scorecard wording when a disagreement was caused by an ambiguous definition rather than reviewer error, and that is the only step that stops the same disagreement recurring.
What evidence should a QA finding carry?
The timestamp to the second, the quoted line it refers to, a link that plays that moment rather than the start of the call, and the identity of whoever confirmed or dismissed it. A finding without those is an assertion, and assertions are what fall apart when a consumer disputes what was said a month later.
Can automated scoring handle the whole scorecard?
No, and a system claiming otherwise is overreaching. Automation is genuinely strong on the binary, language-anchored categories — was a phrase present, did prohibited content appear — which happens to be exactly what sampling detects worst. It is weak on graded judgements like professionalism under pressure. Let the machine decide what a human looks at, and let the human decide what it means.
How long should we keep call recordings and QA scores?
Decide it explicitly rather than by default, and decide the two separately. Recordings should cover the dispute and examination windows that apply to your business and no longer than you have a reason for. Scores are far smaller and frequently worth keeping longer, because the trend is what shows whether a corrective action actually worked. Your own counsel should set the windows.
Does a QA scorecard make us compliant?
No. A scorecard is an operational artefact that reflects obligations your counsel has identified; it does not create compliance and no vendor category set confers it. What a well-built scorecard does is make your own conduct visible to you, with evidence, quickly enough to act on. That is worth a great deal, and it is a different claim from compliance.
Reviewing two percent of your calls?
CollectionsQA records, transcribes and scores every call your collectors take, and sends one supervisor report each morning naming the calls that need a human.