The 2% Problem
Here is how call quality has been managed in Australian businesses for thirty years. A team leader picks a few calls per agent per month, listens to them, fills in a scorecard, and has a conversation about the result some weeks later. Everyone involved knows it is unsatisfying. Almost nobody says so, because there was no alternative.
It is worth naming exactly why it fails, because the failures explain what AI is actually solving.
Statistically meaningless
An agent handling a few hundred calls a month gets three or four scored. The result describes which calls were picked more than it describes the person. Two agents of identical ability can score very differently on chance alone.
Far too late to coach
Nobody learns from feedback on a conversation they cannot remember. Coaching works close to the event, and a monthly cycle guarantees it never is.
Feels like a verdict
A rare score on a rare sample reads as judgement rather than help. That is why quality programmes generate defensiveness instead of improvement — the format teaches people to dispute the sample.
Expensive in the wrong currency
It consumes senior time. A team leader spends hours a week listening to calls instead of leading, which is the most costly labour in the room being spent on the least leveraged task.
The honest summary
Most quality assurance programmes are performed rather than useful. They exist so that the organisation can say it has one. That is not a criticism of the people running them — it is a description of what is possible when your only instrument is a human ear and a spreadsheet.
What Actually Changed
Two things arrived at roughly the same time, and their combination is what makes this a 2026 conversation rather than a 2019 one.
Speech-to-text became accurate and cheap enough to run on every call. Not perfect — we will come back to that — but good enough that transcribing everything is now a routine platform function rather than a budget line.
Language models became able to evaluate against a rubric. This is the genuinely new capability. Earlier systems could only search transcripts for keywords, which produced the well-known absurdity of an agent scoring well for saying "I understand your frustration" in a robotic monotone. A model can now be given a written description of what a good call looks like in your business and asked to assess a conversation against it — including things keyword search could never catch, like whether a question was actually answered.
| Manual QA | Keyword spotting | AI evaluation | |
|---|---|---|---|
| Coverage | ~2% of calls | 100%, but only for phrases | ✓ 100% of calls |
| Latency | Days to weeks | Immediate | ✓ Immediate |
| Understands meaning | ✓ Fully | ✗ Not at all | Substantially |
| Consistent between calls | ✗ Varies by assessor and by mood | ✓ Perfectly | ✓ Highly |
| Cost per call scored | High — senior labour | Negligible | ✓ Low |
| Trusted by the team | Sometimes — depends on the leader | ✗ Openly gamed | Depends entirely on rollout |
That last row is the whole game, and section six is about it. But first, precision about capability — because the fastest way to lose a team's trust is to over-claim.
What AI Can and Cannot Score
Any vendor unwilling to draw these lines is selling you a future argument. Here is the honest version.
| Reliability | What it covers | How to use it |
|---|---|---|
| ✓ Reliable | Was it said? Greeting, recording disclosure, required statements, identity checks, mandated next steps, the offer being made at all. Conversation mechanics. Talk-over ratio, dead air, longest monologue, interruptions, hold time, who spoke more. Classification. Call reason, topic, product mentioned, outcome, escalation. |
Automate fully. Report on it. Alert on it. This is the layer that earns the licence fee. |
| Partially reliable | Sentiment and customer effort. Genuinely useful across thousands of calls and over time; weak on any individual call, because tone is ambiguous and Australians in particular say "no worries" while being extremely unimpressed. | Use in aggregate and for trend detection. Never quote a single call's sentiment score at a person. |
| ✗ Not reliable | Was the answer correct for this customer? Requires knowledge of their account, history and circumstances. Were they actually satisfied? Politeness is not satisfaction. Relationship context. The fourth call in a bad week reads completely differently to the first. |
Keep human. These are the calls a manager should be listening to — and AI is excellent at finding them for you. |
The rule that keeps you honest: automate the objective and observable, and reserve human judgement for the interpretive. Used that way, AI does not replace your team leaders — it stops them wasting time on the 95% of calls that were fine, so they can spend it on the ones that were not.
It is also worth knowing where the errors come from, because they are not random. Accuracy drops with strong accents, background noise, poor line quality, crosstalk, and specialist vocabulary — trade terms, drug names, part numbers, legal phrases. Which means the errors concentrate in exactly the industries and workforces most likely to be treated unfairly by a naive scoring system, and that is a reason to design carefully rather than a reason to avoid the technology. Our guide to AI transcription, summaries and CRM auto-notes covers what drives accuracy in practice.
Compliance: From Sampling to Coverage
If you need one business case that survives scrutiny, this is it — because it replaces a statistical argument with a complete record.
Plenty of Australian businesses have obligations that attach to what is said on a call: a disclosure before taking payment, a statement that the call is recorded, identifying yourself in a prescribed way, a defined process for customers in vulnerable circumstances, a required warning before a product is sold. Traditionally you assured that by sampling and hoping the sample was representative.
2%
Typical share of calls a manual programme can review
100%
Share an automated check can cover
Day 2
When a new starter's missing disclosure surfaces — instead of a quarterly audit
Two consequences, both valuable. Systemic problems surface early. A new starter who has never said a required disclosure shows up almost immediately, which is when it is cheap to fix, rather than at a quarterly review when there are three months of calls behind it. And you can demonstrate coverage rather than describing a methodology — a fundamentally stronger position in front of a regulator, an auditor or an insurer.
The obligations are also growing. Australian businesses now sit inside an expanding set of duties around scams, consumer communication and outage transparency — see the Scams Prevention Framework and the new transparency rules. Every one of them makes complete coverage worth more than it was last year.
Two compliance points that cut the other way
First, recording and monitoring calls has its own rules, which vary by state and territory on top of federal law. Get your recording notification and your policy right before you build a scoring programme on top of it — our guide to call recording law, benefits and setup is the starting point. Second, recordings and transcripts are personal information under the Privacy Act 1988, so where they are stored and processed is part of your obligation. AI features are the most common way that data quietly leaves the country — see where your calls actually live.
Coaching That Finally Lands
The scoring is the boring half. The interesting half is what complete coverage lets a good team leader do, and it is a genuine change in kind rather than degree.
| Before | After |
|---|---|
| "You scored 3 out of 5 on this call from three weeks ago." | "Here are the four calls this week where the customer went quiet straight after you quoted the price. Let's listen to two of them." |
| "Try to build more rapport." | "On calls where you asked a question in the first thirty seconds, your resolution rate is noticeably higher. Do more of that." |
| New starters shadow whoever is free. | New starters get the five best real calls for their most common enquiry type, chosen from actual data. |
| Nobody knows why Tuesdays are bad. | Tuesday's dominant call reason is visible, and staffing or self-service can answer it. |
| Best practice is whatever the loudest senior agent does. | Best practice is identified from calls that actually resolved, and taught deliberately. |
Notice the pattern: every "after" row is specific, recent and evidenced. That is the difference between coaching that changes behaviour and coaching that produces a nod. The biggest single win in most businesses is new-starter ramp time, because a new agent can be shown what good looks like on the exact enquiries they will face tomorrow instead of learning by absorption over three months.
The Trust Problem — the Real Risk
Everything in this article is achievable. The reason most of these projects underperform has nothing to do with the technology and everything to do with how it arrives.
Put yourself on the other side of it. You take calls all day. Somebody announces that from Monday, an AI will listen to every one of them and give you a score. No matter how carefully that is worded, the first thought is not "excellent, better coaching". It is "they are looking for a reason".
The four commitments that decide whether this works
- Tell them before you switch it on. Finding out afterwards is the thing that causes lasting damage, and it is entirely avoidable.
- Show each person their own data first. Let the team see their scores before any manager acts on them. This single step converts more scepticism than any explanation.
- Put it in writing that scores alone never lead to discipline. A human listens to the actual call before any consequence. Scoring a conversation is not the same as understanding it.
- Start with something that helps rather than judges. Automatic call notes are the ideal first use case: the AI does the paperwork nobody wants to do, and the team's first experience of it is relief.
That last point is worth more than it sounds. Teams that first encounter this technology removing admin accept scoring far more readily than teams who meet it as a new way of being marked. It costs nothing to sequence it that way, and it changes the outcome.
There is a related point about fairness. Because transcription accuracy varies with accent and line quality, a naive system can systematically disadvantage particular staff. Check your scores for that pattern deliberately — compare across teams and individuals and ask whether differences track performance or track something else entirely. If you find it, fix the process rather than defending the tool.
Scoring What Matters, Not What's Measurable
The second failure mode is subtler and hits well-run operations hardest: anything you measure becomes a target, and anything that becomes a target gets optimised at the expense of the thing you actually wanted.
Score script adherence, and you get agents who follow the script while helping the customer less. Score call duration, and you get calls that end quickly and customers who ring back tomorrow — which looks like two efficient calls and is actually one failure. Score positive language, and you get relentless cheerfulness in situations that call for straightforwardness.
| Tempting to score | What you'll actually get | Better |
|---|---|---|
| Script adherence | Scripts followed, problems unsolved | Was the customer's actual question answered? |
| Average call duration | Fast calls, repeat calls | Resolved without a follow-up contact |
| Positive language count | Performative cheerfulness | Customer effort — how hard was this for them? |
| Calls handled per hour | Rushed calls and burnt-out staff | Outcomes per hour, including quality |
| Sentiment score per call | Disputes about individual calls | Sentiment trend across a team over weeks |
The practical defence is to score outcomes rather than behaviours wherever you possibly can, and to revisit the rubric every quarter with the question: what is this measure quietly encouraging? Being able to score everything makes rubric design far more consequential than it was when you only scored three calls a month, because now the incentive reaches every conversation.
Your AI Agent Needs QA More Than Your Team Does
Here is the argument almost nobody makes, and it may be the most important section in this article.
If you have deployed an AI voice agent to answer calls, you have created a worker who handles high volume with no supervisor, no team leader walking past, and no colleague overhearing something wrong. Consider the asymmetry:
Human error is self-limiting
One person misunderstands a policy and gets it wrong a handful of times before somebody corrects them. The damage is bounded by one person's shift and one person's caseload.
AI error is systematic
An AI agent with a wrong understanding applies it identically to every single call until somebody notices. Four hundred calls, four hundred identical errors, no variation to trigger anyone's suspicion.
So the same pipeline that reviews your team should review your AI, against questions that are specific to it:
- Did it answer accurately, or confidently invent something?
- Did it stay inside its defined scope, or drift into territory it should have escalated?
- Did it hand over to a human at the right moment — and inside the same conversation, without making the customer start again?
- Did it make a commitment on your behalf that it had no business making?
- What did callers ask that it could not handle? (This is the single most valuable output — it is your product and process roadmap, written by your customers.)
A fair test for any AI answering vendor
"Show me how I audit what the AI said." If the answer is a dashboard of call counts and containment rates, that is metrics, not assurance. You want transcripts, the ability to search them, and scoring against your own criteria. A provider selling AI answering without the means to audit it is selling you half a product — and it is the risky half they kept.
Which calls to give an AI in the first place is a separate and equally important decision, covered in which calls to automate and which to keep human.
A Six-Step Rollout
| Step | What to do | Why this order |
|---|---|---|
| 1. Write down what a good call is | Two pages, in plain language, agreed by the people who actually manage the team. Not aspirational — descriptive. | AI cannot score a standard you have never articulated. Most organisations discover at this step that they disagree internally, which is itself worth finding out. |
| 2. Start with notes, not scores | Turn on transcription and automatic call summaries first. Let the team experience the technology as removing paperwork. | Buys goodwill, proves accuracy on your actual calls and accents, and costs you nothing in trust. |
| 3. Calibrate against humans | Have your best assessor score fifty calls independently, then compare with the AI. Investigate every disagreement. | You will find both AI errors and rubric ambiguity. Fixing them now is far cheaper than arguing about them later. |
| 4. Tell the team, then show them their own data | Explain what is measured, what it is used for and what it will never be used for. Then give each person their own scores before any manager sees them in a review. | This is the step that decides whether the programme succeeds. Do not compress it. |
| 5. Coach for a full quarter before anything else | No scorecards in reviews, no league tables, no consequences. Coaching only. | Lets the data prove itself as help rather than judgement, and gives you a baseline to measure improvement against. |
| 6. Review the rubric, then extend | Ask what the measures are quietly encouraging. Adjust. Only then extend to compliance alerting and formal reporting. | Rubric design is now consequential on every call rather than three a month. Revisit it every quarter, permanently. |
Most failed implementations skipped straight to step six. The sequence is the intervention.
What to Look For in a Platform
Buying criteria, in the order we would weight them.
Built in, not bolted on
Recording, transcription, AI summaries and reporting as part of the platform rather than three vendors and an integration project. Uniden Voice includes them in one simple per-user rate.
Your rubric, not theirs
You must be able to define the criteria in your own words. A fixed vendor scorecard describes a generic contact centre, and you do not run one.
Australian data handling
Ask specifically where audio and transcripts are processed for AI features, whether they are retained by a third party, and whether they train models. Uniden Voice keeps this data on Australian infrastructure by default.
Searchable transcripts, not just dashboards
You need to find the actual conversation behind any number. A dashboard you cannot drill into is decoration.
Pushes into the systems you use
Summaries and outcomes landing in your CRM automatically — 1,000+ integrations and open APIs — because a transcript nobody sees is a transcript nobody uses.
Every channel, one view
Quality is not a voice-only question. Calls, SMS and chat should be assessable in one place — see one queue, every channel.
For the wider platform decision this sits inside, the best contact centre software in Australia covers the full evaluation, and why Uniden Voice leads Australian CCaaS covers where we sit in it.
Frequently Asked Questions
What to Read Next
Scoring every call is one capability. These are the ones it depends on and enables.