
Every collections floor runs quality assurance. Almost every one of them runs it on a sample, because that is what the tooling and the headcount allow: a supervisor pulls a handful of calls per collector per month, scores them against a sheet, and files the result. The programme exists, the binder is current, and everybody moves on.
The uncomfortable part is arithmetic. A sample that small is not a measurement of your floor. It is a measurement of the sample. And the specific thing a compliance programme most needs to catch — the rare, serious breach — is close to invisible at those rates.
This is the maths, honestly, with the parts that argue against us left in.
What a 2 percent sample actually sees
Start with a floor that is doing well. Say a collector makes 800 dials a month, of which 200 are real conversations, and say they get something meaningfully wrong on 1 in 100 of those conversations. That is a good collector: a 99 percent clean rate.
Two things follow, and the second one surprises people.
First, that collector produces about two problem calls a month. Not zero. Two.
Second, if you review 2 percent of their conversations — four calls — the chance that your sample contains either of those two problem calls is small. You are drawing four cards from a deck of two hundred that contains two cards you care about. Most months you draw neither, and your QA record says the collector is clean. Over a year you might catch one.
The programme is not lying. It is answering a different question from the one you asked. You asked "is my floor compliant?" and the sample answered "were these four calls compliant?"
| Conversations per collector per month | Reviewed at 2% | Reviewed at 5% | Reviewed at 100% |
|---|---|---|---|
| 200 | 4 calls | 10 calls | 200 calls |
| Chance of catching a 1-in-100 issue that month | roughly 1 in 25 | roughly 1 in 10 | certain |
| Supervisor hours at ~6 min per review | ~0.4 hr | ~1 hr | not humanly possible |
| What the record supports | "these four were fine" | "these ten were fine" | "the floor was fine" |
That last row is the one that matters when somebody asks you to prove something.
The bottom-right cell is also why sampling exists at all. Reviewing two hundred calls per collector per month by hand is not a staffing problem you can solve by hiring — at twenty collectors it is thousands of supervisor hours a month. Sampling was never a bad idea. It was the only idea available.
The rare event is the whole point
General contact-centre QA is mostly about coaching: tone, efficiency, whether the agent followed the flow. For that purpose a sample is genuinely fine. Coaching signals are common, so a small sample finds them.
Collections is different, because the thing you are watching for is rare and carries consequence. A missing disclosure, a third-party disclosure, a call placed outside permitted hours, a continued call after a cease request — these are not everyday events on a decent floor. They are exactly the low-frequency, high-cost events that sampling is structurally worst at detecting.
The regulatory environment assumes you are watching. The Consumer Financial Protection Bureau supervises debt collection under the Fair Debt Collection Practices Act and its implementing regulation, and its published examination materials describe expectations around monitoring and corrective action rather than around sampling percentages. The regulation itself is public at the Electronic Code of Federal Regulations, and the Federal Trade Commission has enforced in this space for decades and publishes annual reporting on collection complaints.
None of those sources mandates a review rate. That is the point people miss: nobody will tell you 2 percent is too low. You will find out it was too low the same way everyone else does, after the fact, when a complaint arrives and you go looking for the call.
The second failure: you cannot find the call afterwards
Detection is only half of it. The other half is retrieval.
When a complaint lands — a consumer disputes what was said, a client asks for the recording, a regulator asks for a period — the question is not "did QA score this call?" It is "can you produce it, and can you show what was said?" A sampling programme creates a record of the calls it happened to review. It creates nothing at all about the other 98 percent beyond the raw audio, if the raw audio was even kept.
That gap shows up in a predictable way. Somebody spends a day listening to recordings trying to find the one the complaint refers to, working from a date and a rough time. If the recordings are searchable only by collector and timestamp, that is genuinely a day of work. If the calls were transcribed and scored, it is a search.
This is the practical argument for reviewing everything that has nothing to do with catching anyone. A complete, searchable record is an operational asset even on a floor where nothing ever goes wrong. It answers questions in minutes instead of days.
What changes when the review is automated
The reason census review is now possible and was not ten years ago is that transcription and scoring stopped being manual. Speech recognition got good enough and cheap enough that transcribing every call is a line item rather than a project.
That changes the supervisor's job rather than eliminating it. The work moves from finding problem calls to judging them:
- Every call is transcribed and scored against a rubric built for collections rather than a generic contact-centre sheet.
- Calls with nothing on them generate nothing. No report, no queue, no supervisor time.
- Calls with something on them produce a flag that names the moment and quotes the line.
- A human confirms or dismisses each flag. The system does not decide — it decides what to put in front of a person.
The daily output of that is small. On a floor of twenty, a typical morning report is a handful of flagged moments, most of which take under a minute each to dismiss. That is a fundamentally different shape of work from listening to four calls per collector in the hope that they are representative.
Our own how it works page walks the mechanics, and the features page lists what is switched on per machine rather than globally, which matters more than it sounds — see below.
The floors that get burned are almost never the ones with no QA programme. They are the ones with a QA programme that produced a clean binder every month for two years, because nobody ever asked what percentage of reality the binder covered.
— Compliance manager, third-party collections, 12 years, name withheld by request
The honest objections
There are real arguments against reviewing everything, and a page that pretended otherwise would not be worth reading.
"Scoring by machine will flag things that are fine." Yes. Any automated rubric produces false positives, and the answer is not to pretend otherwise but to keep the human in the loop and to report per dimension rather than as one blended score. A flag is a pointer to a timestamp, not a verdict. If a system tells you a collector failed without showing you the four seconds it is talking about, it is asking for trust it has not earned.
"Recording everything raises its own legal questions." It does, and they are not the same in every state. Roughly fifteen states require all-party consent to record a call, and a floor dialling across state lines usually finds the strictest rule on the call governs it. We set out what we know and where we stop on the coverage page, and the compliance page describes the attestation that gates recording in the software. This is orientation, not legal advice, and a recording programme should be reviewed by your own counsel.
"We take card payments on these calls." Then say so before you switch anything on. Card numbers spoken aloud on a recorded call are a PCI question, and the honest answer is that redaction on the call path is not something we do. That changes what you should enable. The Federal Communications Commission rules on calling practices and the National Institute of Standards and Technology guidance on protecting sensitive data are both worth reading alongside your own card-brand obligations.
"Our collectors use desk phones." Then voice capture will not see those calls at all. Audio comes off the collector machine, which is what makes it dialler-agnostic and also what makes a handset invisible to it. Desktop activity still works; voice does not. That is a real limitation and it should be settled before anyone buys anything.
What this costs, and why the pricing is shaped this way
As of August 2026 our pricing is per monitored collector: $39 a seat for activity and daily reporting without voice capture, and $99 a seat once every call is recorded, transcribed and scored. Supervisors and administrators are free, because a supervisor reading a report generates no transcription minutes and no analysis, so charging for them would be charging for nothing.
The comparison worth making is not against your current QA software licence. It is against the supervisor hours currently spent listening to a sample that cannot answer the question, plus the cost of the day somebody spends hunting for one recording after a complaint. Whether that comparison favours us depends on your floor size and your current process, and if it does not, we would rather you knew that before you bought.
The pricing page has the full breakdown by seat count.
What to do with this if you change nothing else
Even if you never buy monitoring software, two things are worth doing this quarter.
Work out your actual review rate. Not the target in the policy — the real one. Calls reviewed last month divided by conversations held last month. Most floors have never computed it and are surprised by the number.
Then compute what that rate can detect. If you review 4 of 200 calls, you are sampling 2 percent, and a behaviour occurring on 1 percent of calls will escape you in roughly 24 months out of 25 for that collector. Write that sentence down next to your review rate. It is the honest description of what your programme currently proves.
If the answer is comfortable, you have a well-calibrated programme and this article was not for you. If it is not, you now have a number to take to whoever controls the budget, which is more useful than a vendor claim.
Frequently asked questions
What percentage of calls should a collections floor review?
There is no mandated percentage, and any vendor quoting one is inventing it. The regulators describe expectations around monitoring and corrective action rather than a review rate. The useful way to answer it for your own floor is to work backwards: decide the rarest behaviour you genuinely need to catch, then check whether your sample size can detect something occurring at that frequency. On most floors the honest answer is that the current rate cannot.
Is reviewing 100 percent of calls actually realistic?
It is realistic when the transcription and scoring are automated and a human reviews only what gets flagged. It is not realistic if it means people listening to every call, which is why sampling became standard in the first place. The distinction matters: census review is automated, census listening is not, and nobody is proposing the second one.
Does more QA coverage mean more discipline for collectors?
Not in our experience, and floors that use it that way get worse results. Full coverage mostly surfaces process problems rather than people problems — a disclosure that gets skipped when the consumer interrupts, a script step that does not survive a hostile opening. Those are fixable by changing the script or the training. Using complete monitoring primarily as a disciplinary instrument tends to produce collectors who sound like they are reading a legal notice, which is its own compliance risk.
Can automated scoring replace a compliance analyst?
No, and it should not try. Automated scoring decides what a human looks at; the human decides what it means. A flagged moment is a pointer to a timestamp with a quote attached, and confirming or dismissing it takes seconds precisely because the analyst is not hunting for it. The analyst role shifts from searching to judging, which is the part that actually needs their expertise.
What about states that require all-party consent?
Roughly fifteen states require all-party consent to record, and the rest generally permit recording with one party consenting, which your collector supplies. A floor dialling across state lines will usually find the strictest rule that applies to a given call governs it. Treat that as orientation rather than advice: statutes change, and how they apply turns on facts about your specific programme that only your counsel can assess.
We already record every call. Is that the same thing?
Recording is storage; QA is review. Most platforms a collections floor already runs will store every call, and that is genuinely valuable for retrieval. What they generally do not ship is a collections scorecard — an FDCPA-aware category set, disclosure detection, and a per-dimension result. If your platform stores the audio and your review still happens on a sample, you have the retrieval half solved and the detection half open.
What if our collectors work from home?
That is the case that broke the old model, and it is the reason this product exists in the shape it does. The platform your accounts live in stores recordings and stops there; generic contact-centre QA tools score calls but ship no collections rubric; and neither category watches a collector working from a spare bedroom. Desktop-level oversight of a remote collector, with the call content read against the rules that govern the call, is the seam this sits in.
Reviewing two percent of your calls?
CollectionsQA records, transcribes and scores every call your collectors take, and sends one supervisor report each morning naming the calls that need a human.