Why the Demo Number Is Not Your Number
Every STT vendor quotes a word error rate the percentage of words transcribed wrong. The number in the pitch deck is almost always measured on a public benchmark: clean audio, native accents, one speaker at a time, no background noise. It is a real number. It is also close to meaningless for your evaluation, because none of your calls look like that.
Your audio is telephony-compressed narrower bandwidth than a podcast mic, typically around 8kHz. It has hold music bleeding through, a customer in a car, an agent in an open-plan BPO floor, and — if you run a multilingual operation — speakers who switch languages mid-sentence without warning. Every one of those conditions degrades accuracy, and none of them show up in a vendor benchmark.
The only number worth trusting is WER measured on your own recordings. Pull fifty real calls — not your cleanest ones, your normal ones and run them through the vendor’s system before you sign anything. If a vendor resists giving you a trial for this, that itself is the answer.
Real-Time vs Post-Call: You Are Buying Two Different Things
This is the split that matters most and gets blurred most often in vendor pitches. Real-time and post-call transcription solve different problems, on different budgets, for different teams.
| Post-call transcription | Real-time transcription |
When text is available | After the call ends | While the call is happening |
Primary use | QA scoring, analytics, compliance archive, coaching | Live agent assist, in-call prompts, voice bots |
Value to the agent | None during the call | Immediate — suggestions while they’re talking |
Infrastructure demand | Lower — batch processing | Higher — low-latency streaming pipeline |
Typical buyer | QA and compliance teams | Ops teams building agent assist or a voice bot |
If your goal is 100% QA coverage instead of random sampling, post-call is enough and considerably cheaper. If your goal is reducing average handle time by surfacing an answer to the agent while the customer is still talking, you need real-time and real-time is a genuinely different engineering commitment, not a toggle on the same product. Buy the one you actually need; a lot of budget gets wasted on real-time infrastructure for a QA use case that never needed it.
The Mono-Channel Diarization Problem
Diarization is working out who said what — separating the agent’s speech from the customer’s in a transcript. It sounds like a solved problem. It is, if your calls are recorded in stereo, with the agent on one channel and the customer on the other. Most legacy call recording systems are not set up that way — they record in mono, both speakers mixed into a single track.
Separating two voices from one mixed audio stream is a materially harder problem than reading two clean channels, and accuracy on mono-channel diarization is meaningfully worse across the industry. If your recording infrastructure is mono today, ask every vendor to demonstrate diarization on your mono recordings specifically — not on a stereo sample file. This single check eliminates a category of vendors who quietly assumed stereo input and never mentioned it.
What STT and TTS Actually Change Operationally
The pitch usually leads with the feature. The business case should lead with what changes on your dashboards.
Layer | What it enables | What it should move |
Real-time STT | Live prompts, next-best-action, in-call knowledge retrieval | Average handle time, first-contact resolution |
Post-call STT | Full-coverage QA instead of sampling, searchable call archive | QA coverage %, coaching cycle time |
TTS | Voice bot responses, automated outbound (reminders, status updates) | Containment rate, agent hours freed from repetitive calls |
Sentiment / compliance layer | Flags risk phrases, tone, missed disclosures automatically | Compliance detection rate, escalation quality |
The honest caveat: none of this changes anything if the transcript underneath it is wrong. A misheard product name or a garbled account number doesn’t stay a transcription error — it becomes a bad CRM entry, a skewed sentiment score, and a QA evaluation built on a sentence nobody actually said. Accuracy on your audio is not a nice-to-have line item; it’s the floor everything else sits on.
Where the Pricing Hides Extra Cost
A common vendor tactic: quote a low headline rate per minute, then charge separately for diarization, sentiment analysis, entity extraction, and language detection — the features you actually need to make the transcript useful. The base rate looks competitive in a spreadsheet; the effective cost with everything switched on is a different number entirely.
●Ask every shortlisted vendor for one all-in price with diarization, sentiment, and entity extraction enabled — not a base rate with an asterisk.
● Ask about processing time, not just accuracy. A transcript that takes twenty minutes to return on a one-hour call is useless for same-shift summaries or same-day QA, no matter how accurate it is.
● Ask what happens to volume pricing at your actual call volume, not the trial-tier example in the sales deck.
The Compliance Question Most Teams Skip
A voice recording of a customer is personal data under most modern privacy regimes — the EU’s GDPR is explicit that a recorded voice qualifies, and equivalent principles apply under other regional data-protection laws. That has a direct, practical consequence for procurement: when you route call audio through a transcription vendor, their certifications and data-handling practices become part of your compliance exposure, not a separate concern you can outsource and forget.
Before signing, get clear answers on where recordings and transcripts are stored, how long they’re retained, whether the vendor is certified for the regions you operate in, and what happens to the data if you terminate the contract. In regulated sectors — banking, insurance, healthcare — this is not a legal-team formality; it’s the difference between a vendor decision and a liability.
A Practical Evaluation Checklist
What to actually do before you sign, in order:
1. Decide real-time or post-call first. They are different purchases with different infrastructure demands — don’t let a vendor sell you real-time for a QA-only need.
2. Test WER on fifty of your own hardest calls, not the vendor’s benchmark and not your cleanest samples.
3. Test diarization on your actual recording format — mono or stereo — not a demo file.
4. Get an all-in price with every feature you’ll use enabled, at your real call volume.
5. Confirm data residency and certification for every region you operate in, in writing.
6. Check integration effort honestly. Does it plug into your existing CRM, QA, and WFM stack, or does it require rebuilding around it?
7. Pilot on one team or one site before a full rollout, and measure the KPI you actually care about — AHT, QA coverage, FCR — not just transcription accuracy in isolation.
Common Mistakes When Adopting STT/TTS
● Buying real-time infrastructure for a post-call use case. Expensive, and unnecessary if nobody is using the transcript live.
● Evaluating on the vendor’s benchmark instead of your own audio. The single most common and most costly mistake.
● Ignoring mono-channel diarization until after go-live. By then it’s a production problem, not a procurement decision.
● Treating TTS as a solved commodity. A voice that sounds fine in a demo can sound stilted or mispronounce product names and customer names constantly in production — test on your actual vocabulary.
● No plan for the escalation path. STT and TTS improve automation; they don’t remove the need for a human on complex or sensitive calls. If your rollout has no clean handoff to a human, automation just relocates the frustration instead of removing it.
● Underestimating latency for real-time use cases. A live agent-assist or voice bot that lags behind the conversation is worse than no automation at all — see voice agent latency for what actually causes the delay.
Conclusion
Speech-to-text and text-to-speech are mature technology now — the failures that still happen aren’t usually about whether the AI can transcribe or speak. They’re about procurement shortcuts: trusting a benchmark number instead of your own audio, buying real-time infrastructure for a batch use case, missing a mono-channel diarization gap until it’s in production, or signing a rate card that doesn’t include the features you’ll actually turn on. Test on your hardest calls, price the fully-loaded cost, confirm data handling in writing, and pilot before you scale. Get those five things right and the technology does what the pitch deck promised. Skip them and you’ll spend the first two quarters after go-live finding out why the demo didn’t match your call floor.
FAQs
What is speech to text for a call center?
It's the use of automatic speech recognition (ASR) to convert live or recorded customer calls into text, so conversations can power live agent assist, QA scoring, analytics, and CRM updates instead of staying trapped in audio files.
Why doesn't a vendor's accuracy benchmark predict my results?
Vendor benchmarks are typically measured on clean, high-fidelity audio with native accents and single speakers. Real call-center audio is compressed telephony, often around 8kHz, with background noise, accents, and code-switching — all of which degrade accuracy in ways a clean benchmark never tests. Evaluate on your own recordings, not the published number.
What is the difference between real-time and post-call speech-to-text?
Post-call transcription becomes available after the call ends and is used for QA, analytics, and compliance review. Real-time transcription streams as the call happens and can power live agent assist and voice bots. They require different infrastructure, so decide which one your use case actually needs before buying.
Why does mono-channel audio make diarization harder?
Diarization identifying who said what is far more reliable when the agent and customer are recorded on separate channels. Many legacy call-recording systems record both speakers mixed into a single mono track, which is a harder separation problem and typically produces lower accuracy. Test diarization on your actual recording format before buying.
Is call recording data covered by privacy regulations?
Generally yes — a recorded voice is treated as personal data under regimes like the EU's GDPR and equivalent regional laws elsewhere. That means a transcription vendor's certifications and data-handling practices become part of your own compliance exposure, so confirm data residency, retention, and certification before signing.




