Why the Feature Lists All Look the Same
There is a structural reason every 2026 quote reads alike, and understanding it makes the rest of this article obvious.
Speech recognition, speech synthesis and large language models are all now available as commodity services. A phone platform can add "AI transcription" by sending audio to a third-party speech service and displaying what comes back. It can add an "AI receptionist" by connecting a text-generation service between a caller and a call flow. Neither requires the provider to have built anything in the field, and both can be shipped in weeks. That is not a criticism β plenty of excellent products are assembled from good components β but it means the presence of a feature tells you nothing about its quality, its cost stability, its data path or who can fix it.
The consequence for a buyer. The differentiating questions have moved down a level. Not "do you have AI transcription" β everyone does β but "whose speech model, running where, trained on what kind of audio, at what accuracy on Australian proper nouns, with what happens to the audio afterwards". Those questions have different answers from different providers, and the answers are knowable. The rest of this article is how to find them without needing to be a machine learning engineer.
The Four Layers Behind Any AI Phone Feature
Every AI capability on a business phone system sits on four layers. A provider may own all four, or one, and the difference decides what they can do for you when something breaks.
| Layer | What it does | What it decides for you | The question to ask |
|---|---|---|---|
| 1. Network | Carries the call. Interconnects with Australian carriers, holds the numbers, delivers the audio. | Call quality, jitter, the audio the model actually hears, and whether a fault can be traced end to end. | "Do you operate your own network and interconnects, or buy wholesale from someone who does?" |
| 2. Platform | Controls the call. Queues, routing, hours, recording, the app, the admin console. | What you can configure, how quickly a change takes effect, and whether AI is genuinely in the call path or bolted to the side. | "Is the AI a step inside your call flow, or a separate product with a forwarded call?" |
| 3. Model | Understands and generates. Speech to text, text to speech, reasoning, summarisation. | Accuracy, latency, tone, cost stability, and what happens when the upstream model changes. | "Whose models, in which region, and what happens to my configuration when they update?" |
| 4. Integration | Acts on the result. Writes to the CRM, books the calendar, raises the ticket, sends the SMS. | Whether the AI produces work or merely produces text. | "Show me the record it writes, in my system, on a live call." |
Why the layer question is not academic
Consider a real failure: outbound calls from your main number begin arriving on mobiles labelled as suspected spam, and the AI receptionist starts getting hung up on. Is that a network problem, a platform configuration problem, or a model problem? A provider who owns the network and the platform can look at both, correlate them, and answer. A provider who resells both can only raise a ticket with two other companies and relay the replies. The number of companies between you and the fix is a real product attribute, and it is invisible on a feature list. Our note on what to ask about AI infrastructure goes further into this.
Nine Tests You Can Run in a Trial
All nine take under two hours in total and need nothing but a phone. Run the identical script against every provider on your shortlist, in the same week, from the same handsets and networks, and write down what happens. The differences are usually dramatic and completely absent from the sales material.
1. Interrupt it
Start speaking while the AI is mid-sentence, the way an impatient caller does. A capable system stops, listens, and picks up your point. A weak one talks over you, finishes its script, then asks you to repeat. This single test predicts satisfaction better than any other because real callers interrupt constantly.
2. Feed it local proper nouns
Say three suburb names, two surnames and one product name from your actual customer base. Then read the transcript. Australian place names, Aboriginal and Torres Strait Islander names, and multicultural surnames are where generic models fail hardest, and where your bookings actually live.
3. Time the gap
Count the silence between the end of your sentence and the start of its reply. Under about a second reads as conversation. Beyond roughly two seconds, callers assume the line has dropped and start saying "hello?" β which then gets transcribed as their answer.
4. Check a transcript you know
Take a call you personally had, read the transcript, and mark the errors. Then read the AI summary of that transcript. A summary built on a flawed transcript is confidently wrong, and confidently wrong is worse than absent β because people stop checking.
5. Ask for a human
Say "I need to speak to a person" at three different points: at the greeting, mid-flow, and after giving details. All three should work immediately, without the caller having to guess a phrase, and the details already given should travel with them.
6. Call at 9pm
Most demos happen at 11am on a Tuesday. Your hardest calls do not. Ring the trial number after hours, on a Saturday, and on a public holiday if the trial spans one. Check that hours, greetings and escalation behave as configured rather than as assumed.
7. Follow the result into your systems
The test is not "did it produce a summary" but "did a record appear in the CRM, against the right contact, with the right fields, within a minute". Ask them to demonstrate it writing into your system during the trial, not a generic demo tenant.
8. Ask what happens when it is down
Models and the services around them have outages. The right answer is a defined fallback β the call rings a group, or a queue, or a mailbox with a proper greeting β not an undefined one. Ask to see the fallback configured, and ask how you are notified.
9. Ask it something it cannot know
Ask about a service you do not offer, or a price that is not published. The correct behaviour is to say it does not have that information and offer a person. A system that invents a plausible answer will one day invent a plausible price, and you will be the one holding it.
Run the tests before the discount conversation, not after
Once a discount is on the table, it becomes psychologically expensive to disqualify a provider on test 1 or test 9. Do the testing while everything is still hypothetical. The two hours are the cheapest part of the entire procurement, and they are the only part that observes the product rather than the pitch.
Australian Accents, Place Names and Proper Nouns
This deserves its own section because it is the single most common gap between a demo that impresses and a deployment that annoys.
General speech models are trained overwhelmingly on North American English. They handle a broad Australian accent better than they used to, and the marketing claims about accent support are usually not false. But accent is the easy part. The hard part is proper nouns, and proper nouns are what a business phone call is largely made of: a suburb, a street, a surname, a company, a product code, a policy number.
| Category | Why it breaks | What it costs you |
|---|---|---|
| Suburb and place names | Woy Woy, Kariong, Ngunnawal, Coolangatta, Warrnambool and thousands more appear rarely or never in general training data. | A booking at the wrong address, or a callback that never happens because the address was nonsense and nobody trusted the note. |
| Surnames | Australia's surname distribution is unusually diverse. Greek, Vietnamese, Lebanese, Italian, Indian, Chinese and Aboriginal and Torres Strait Islander names are common and frequently mangled. | Customers whose names are consistently misspelled notice, and it reads as carelessness rather than as a technical limitation. |
| Business and product names | Your own product codes, model numbers and internal terms are unique to you and unknown to any general model. | Every enquiry about your flagship product is logged under a name nobody can search for. |
| Numbers spoken naturally | "Double oh seven", "two four for", "oh eight" β Australians speak digits in patterns that trip literal transcription. | Wrong callback numbers, which is the most expensive small error in the entire system. |
The question that actually matters
Not "does it understand Australian accents" β they will say yes, and they will be broadly right. Ask instead: "can I give the system a list of my suburbs, my product names and my staff names, so it recognises them, and how long does that take to apply?" A system that accepts a custom vocabulary and applies it in minutes is a different class of product from one that cannot. Then test it with your own list. This is also the single highest-return configuration step after go-live, and it is frequently never done because nobody mentions it.
Latency: The Number Nobody Puts on a Quote
Conversational latency is the delay between a caller finishing a sentence and the system beginning its reply. It never appears on a quote, it is trivially measurable, and it changes the entire character of the interaction.
Under ~1s
Reads as a conversation. Callers behave normally, interrupt normally, and rarely mention that it is automated.
~1β2s
Noticeable but tolerable. Callers slow down and over-articulate, which paradoxically reduces recognition accuracy.
Over ~2s
Callers assume the call has dropped. They say "hello?", talk over the reply, or hang up. The transcript fills with fragments.
Latency is a property of the whole chain, not of the model alone: the network path the audio takes, whether the audio is processed in-region or shipped offshore and back, how the platform hands the call to the model, and how the reply is synthesised. This is exactly why the layer question in the previous section matters β a provider who controls the network and the platform can do something about latency, and a provider who resells both cannot.
Test latency on a mobile, not on the office wi-fi
Demos are given from good connections. Most of your callers are on mobile networks, often moving, sometimes in a car park. Ring the trial number from a mobile with two bars, and again from a moving vehicle if you can do it safely as a passenger. That is the honest condition, and the difference between providers widens considerably under it. Our VoIP and SIP troubleshooting guide covers how to tell a network problem from a platform one.
Where the Recording Lives and Who Can Get It
An AI phone system creates three artefacts that did not exist before: an audio recording, a text transcript, and a generated summary. All three are records about identifiable people, and all three are usually created by default. This is where the compliance questions live, and they are worth settling before go-live rather than during an incident.
| Question | Why it matters | What a good answer sounds like |
|---|---|---|
| In which country is the audio processed? | Processing and storage can happen in different places. Audio may be sent offshore for recognition even when the recording is stored locally. | A specific region, and a willingness to put it in writing rather than a reassurance that it is "secure". |
| Where are recordings and transcripts stored? | Determines which jurisdiction's law reaches them, and who you are relying on. | Named location, named jurisdiction, and a statement about subprocessors. |
| Is my audio used to train models? | Customer conversations are commercially sensitive and contain other people's personal information. | A clear no by default, or an explicit opt-in you control β and the ability to see the setting. |
| How long is everything kept, and can I change it? | Retention is a decision, not a default. Too short and you lose your evidence; too long and you are holding personal information without a reason. | Configurable retention per record type, with deletion that actually deletes. |
| Who inside my organisation can retrieve a call? | Recordings of customer conversations are not general staff reading material. | Role-based access, and an access log you can inspect. |
| Can I export and leave? | Two years of transcripts becomes a genuine asset and a genuine lock-in. | A documented bulk export, tested before you need it. |
Two Australian obligations to have on the page
Recording notification is state-based. Listening-device and surveillance-device legislation differs across the states and territories, so the practical standard is to notify at the start of every recorded call on every line β including the AI-answered path, which people forget because it does not feel like a "recorded call". A single organisation-wide notification standard is simpler and errs in the right direction. Automated decisions carry a disclosure obligation from 10 December 2026: where personal information is used in automated decision-making capable of significantly affecting a person's rights or interests, privacy policy disclosure requirements apply. Most phone-system uses of AI are capture and triage rather than decision, but if yours prioritises, screens or declines, examine it before that date. See our note on automated decisions and the December 2026 obligation. This is general guidance, not legal advice.
What the AI Is Allowed to Decide
The most consequential setting in an AI phone system is not a model choice. It is the boundary between what the system captures and what the system decides. Providers vary enormously in how much control they give you over that boundary, and the good ones make it explicit.
| Automate freely | Automate with a human check | Never automate |
|---|---|---|
| Answering every call so nothing rings out | Prioritising a queue by stated reason | Whether a situation is an emergency |
| Capturing name, number, reason and callback time | Quoting a published price | Approving credit, refunds or expenditure |
| Answering the same twenty published questions | Booking into a calendar with capacity rules | Anything affecting a person's rights, housing, care or employment |
| Transcribing, summarising and filing | Routing by detected topic | A caller in distress, or a vulnerable caller |
| Sending a confirmation SMS | Detecting sentiment for review | Declining or screening someone out |
The test for the middle column is reversibility. If getting it wrong produces an inconvenience that a person can fix within the hour, automate it and review the exceptions. If getting it wrong produces a consequence the customer carries β a missed emergency, a wrong price they relied on, a refused service β a person decides and the AI prepares the decision. Ask each provider to show you where in their console that boundary is set. If it cannot be set, it has been set for you. Our piece on which calls to automate works through the categories in detail.
The Feature List, Decoded
Six phrases that appear on nearly every 2026 quote, what each one can mean, and the follow-up that resolves the ambiguity.
| What the quote says | What it might mean | Ask this |
|---|---|---|
| "AI receptionist" | Anything from a menu with speech input, to a scripted bot, to a genuine conversational agent that can be interrupted and can book. | "Can it handle a caller who changes their mind mid-sentence, and can it write a booking into my calendar?" |
| "AI transcription" | Sometimes near-verbatim; sometimes a rough gist. Accuracy on your vocabulary is untested until you test it. | "Can I add a custom vocabulary, and can I see a transcript of a call I choose?" |
| "AI call summaries" | A summary of the transcript, inheriting every transcript error, with no signal about confidence. | "Can I see the transcript beside the summary, and does the summary flag what it could not hear?" |
| "AI call scoring" | Consistent, reviewable criteria β or an opaque number staff cannot appeal. | "Can I see and edit the criteria, and can an agent see why they scored what they scored?" |
| "Sentiment analysis" | Useful as a trend across thousands of calls; unreliable as a judgement about one call or one person. | "Is this reported at cohort level or used to assess individuals?" |
| "AI-powered routing" | Either genuine topic detection, or a keyword rule with a new label. | "What happens when it routes wrongly, and how does the caller get out of it?" |
A Scorecard for Four Providers
Copy this into a spreadsheet, one column per provider, and fill it in during the trial week. Score each row 0, 1 or 2 β absent, present, genuinely good. Fourteen rows, twenty-eight points available.
| # | Criterion | 0 | 1 | 2 |
|---|---|---|---|---|
| 1 | Handles interruption | Talks over you | Stops, needs a repeat | Stops and follows the new point |
| 2 | Local proper nouns | Frequent errors | Mostly right | Right, and custom vocabulary supported |
| 3 | Response latency | Over 2s | 1β2s | Under ~1s, holds up on mobile |
| 4 | Transcript accuracy on a known call | Gist only | Good, some errors | Near-verbatim, uncertainty marked |
| 5 | Escalation to a human | Hard or scripted | Works at the greeting | Works anywhere, context travels |
| 6 | After-hours behaviour | Undefined | Configurable hours | Hours, holidays and a real after-hours path |
| 7 | Writes into your systems | Text only | Generic integration | Demonstrated in your CRM, live |
| 8 | Behaviour when unavailable | Unknown | Falls to voicemail | Defined fallback plus alerting |
| 9 | Says "I don't know" | Invents answers | Sometimes defers | Reliably defers and offers a person |
| 10 | Data residency | Won't specify | Storage stated | Processing and storage stated in writing |
| 11 | Retention control | Fixed | One global setting | Per record type, deletion verified |
| 12 | Scope boundary control | Vendor decides | Some toggles | You set what it may decide |
| 13 | Layers owned | Resells all four | Owns platform | Owns network and platform |
| 14 | Support that can see the call | Offshore ticket queue | Local business hours | Local, 24/7, can inspect the actual call |
What the scores usually reveal
Two things, in our experience of watching businesses run this exercise. First, the spread is much wider than the price spread β providers within ten per cent of each other on cost are routinely twenty points apart on this sheet. Second, the rows that separate providers are almost never the rows in the sales deck. Rows 1, 9, 10 and 13 do most of the discriminating, and none of them appear on a feature comparison table anywhere.
Five Ways Buyers Get This Wrong
Judging by the demo
A demo is a rehearsed conversation on a good connection with a cooperative caller. Every system looks competent. Insist on a trial number you can ring yourself, at times you choose, from your own mobile.
Comparing feature lists
Feature lists converged eighteen months ago. Comparing them in 2026 is comparing marketing departments. Compare behaviour on the same nine tests instead.
Automating the wrong end
The instinct is to automate the hardest calls, because they hurt most. The return is in the highest-volume, lowest-judgement calls β opening hours, address, status, booking. Start there and the hard calls get a better human, sooner.
Ignoring the integration layer
An AI that produces beautiful summaries nobody reads has produced nothing. If it does not write into the system where work happens, it is a transcript archive with a subscription.
No baseline before switching on
Record answer rate, average speed to answer, unanswered calls and after-hours volume for two weeks before go-live. Without it, every later claim about improvement is an argument rather than a measurement.
Not asking how AI is charged
Per minute, per call, per seat, or bundled β the four models produce very different bills at your volume. Ask for a worked example at your actual monthly call count, in writing, and ask what happens if volume doubles.
Turning the Scores Into a Shortlist
A simple, defensible way to move from fourteen scored rows to a decision.
| Step | What to do | Why |
|---|---|---|
| 1. Apply the hard filters first | Eliminate anyone scoring 0 on rows 9, 10 or 11 β invents answers, will not specify where data is processed, or fixed retention. | These are not preferences. They are the ones that become somebody else's problem to explain later. |
| 2. Rank the rest on rows 1β9 | These are the behavioural rows. They decide what callers experience every day for the next three years. | Daily experience compounds. A two-second latency gap is felt on every single call. |
| 3. Use rows 12β14 as the tie-break | Control, layers owned, and whether support can actually see your call. | These decide what happens on the worst day, which is the day the decision gets judged. |
| 4. Normalise price last | Five-year total: seats Γ months, plus numbers, call spend, hardware, setup and porting, integration, and any per-minute AI charges. | Price is the easiest thing to compare and the least likely to be the thing that goes wrong. Compare it last, on one page, in one unit. |
| 5. Ask the two closing questions | "What does the first ninety days look like, week by week?" and "Who do I ring at 4pm on a Thursday, and what can they see?" | Implementation quality and support reachability are the two variables that most often decide whether a good product becomes a good outcome. |
If you want the ranked market view alongside this framework, our comparison of the best business phone systems in Australia covers the major providers on price, hosting, support and AI, and our contact centre software guide does the same for higher-volume teams.
How We Answer These Questions
It would be a strange article that set fourteen criteria and then declined to answer them, so here is where Uniden Voice over Cloud sits on the ones that matter most.
Layers 1 and 2
Australian network and Australian-hosted platform, with the AI as a step inside the call flow rather than a forwarded call to a separate product. When something is wrong, the people looking at it can see the call, the routing and the model behaviour together.
Australian voices and vocabulary
AI phone agents with natural Australian voices, and custom vocabulary so your suburbs, product names and staff names are recognised rather than guessed. That list is the first thing we set up, because it is the highest-return ten minutes in the whole deployment.
Australian data, Australian support
Australian owned, Australian hosted, with Australian people answering the phone. On the retention and access questions, we would rather write the answer into your account than reassure you about it.
You set the boundary
What the AI answers, what it books, what it must hand to a person, and what happens when it does not know β all configurable, all visible, and all reviewable after go-live when you have real calls to look at.