8kHz Telephony Benchmark 2026: What Narrowband PSTN Audio Actually Does to Voice AI Accuracy in India

A hospital group in Hyderabad ran a four-week pilot and killed it. The vendor had demoed an appointment-reminder agent that transcribed Telugu near-perfectly and read back appointment times in a voice the procurement committee described as indistinguishable from a person. In production the same system misheard one appointment date in six and produced a voice that patients over sixty repeatedly asked to repeat itself.
Nothing about the model changed between the demo and the pilot. What changed was that the demo ran over a WebRTC connection at 48kHz and the pilot ran over a PSTN trunk at 8kHz through a G.711 codec, and then through a mobile leg that re-encoded to AMR-NB.
This is the most predictable and least measured failure in Indian voice AI procurement. Vendors demo on wideband audio because that is what a browser gives them. Production runs on narrowband because that is what the Indian telephone network gives you. The gap between those two conditions is large, quantifiable, and almost never disclosed.
This post measures it. Word error rate degradation by codec and language, digit and alphanumeric accuracy, text-to-speech quality loss, and what to do about each.
What narrowband actually removes
A 48kHz WebRTC stream carries frequency content up to roughly 20kHz. A PSTN call carries 300Hz to 3,400Hz. That is not a small trim; it removes most of the acoustic evidence that distinguishes several classes of sound.
Fricatives lose their identity. The difference between /s/, /f/ and /th/ lives largely above 4kHz. Below that ceiling they converge. This is why "fifteen" and "sixteen" collapse on phone calls, and why it happens in every language.
Retroflex consonants blur. Indian languages make heavy use of retroflex stops, and the cues that separate them from their dental counterparts sit in high-frequency transitions that narrowband discards. Hindi ट versus त, Tamil ட versus த. This is an India-specific degradation that English-centric benchmarks never surface.
Aspiration cues weaken. The aspirated/unaspirated distinction that separates क from ख, प from फ carries in a burst of high-frequency energy. Narrowband attenuates it.
Nasal place of articulation flattens. Distinguishing म, न and ण relies on spectral detail that survives poorly.
Add codec compression on top of the bandwidth limit and the picture worsens. G.711 is bandwidth-limited but not heavily compressed. G.729 compresses to 8kbps with a vocoder that models speech as an excitation plus filter, which is fine for intelligibility and destructive for the fine spectral detail recognisers use. AMR-NB, ubiquitous on Indian mobile legs, adapts its bitrate to network conditions and can drop to 4.75kbps on a congested cell, at which point the audio is barely more than a sketch of the original.
Methodology
We recorded a fixed 900-utterance test set spoken by 45 speakers across Hindi, Tamil, Telugu, Marathi and Bengali, balanced across genders and across Tier-1, Tier-2 and Tier-3 speaker origins. The set covers four utterance types: conversational sentences, spoken digit strings of 10 to 14 digits, alphanumeric strings such as vehicle registrations and PNRs, and Indian proper names.
Each recording was captured once at 48kHz studio quality, then passed through a codec chain simulating each production path:
- Wideband reference: 48kHz uncompressed, the demo condition.
- Opus wideband: 16kHz, VoIP path.
- G.711 a-law: 8kHz, standard PSTN.
- G.729: 8kHz compressed, common on cost-optimised trunks.
- AMR-NB 12.2k: 8kHz mobile, good conditions.
- AMR-NB 4.75k: 8kHz mobile, congested cell.
- Tandem: G.711 to AMR-NB, simulating a landline-originated call terminating on a mobile, which is a very common Indian production path and the one nobody tests.
Identical recogniser, identical settings, one variable. Word error rate for sentences, string error rate for digits and alphanumerics, name error rate for proper nouns.
Word error rate by codec
Conversational sentences, WER percentage, averaged across five languages:
| Codec path | WER | Degradation vs wideband |
|---|---|---|
| Wideband 48kHz reference | 6.1% | baseline |
| Opus 16kHz | 7.4% | +1.3pp |
| G.711 a-law 8kHz | 11.8% | +5.7pp |
| G.729 8kHz | 15.2% | +9.1pp |
| AMR-NB 12.2k | 14.6% | +8.5pp |
| AMR-NB 4.75k | 23.9% | +17.8pp |
| Tandem G.711 to AMR-NB | 19.7% | +13.6pp |
The headline: a system demoed at 6.1% WER runs at 11.8% on a plain PSTN call and at 19.7% on the tandem path. That is a tripling of errors, and the tandem row is the one that matches a large share of real Indian outbound traffic.
By language, G.711 8kHz versus wideband:
| Language | Wideband WER | G.711 WER | Degradation |
|---|---|---|---|
| Hindi | 5.4% | 10.2% | +4.8pp |
| Marathi | 6.3% | 11.9% | +5.6pp |
| Bengali | 6.8% | 12.7% | +5.9pp |
| Telugu | 6.1% | 12.4% | +6.3pp |
| Tamil | 6.0% | 13.1% | +7.1pp |
Tamil and Telugu degrade more than Hindi. The retroflex and aspiration density in Dravidian phonology means more of the discriminating information sits in the frequency band that narrowband removes. Any vendor quoting a single "Indian languages" accuracy figure is averaging across a 2.3-point spread that matters if your customers are in Chennai rather than Delhi.
Speaker origin compounds this. Tier-3 Hindi speakers in our set degraded 1.4× more than Tier-1 speakers on the same codec path, because regional phonology moves the utterance further from the model's training distribution and narrowband removes the evidence needed to recover. This is a different mechanism from the code-switching problem covered in the Patna versus Delhi Hindi post, and the two stack.
Digits and alphanumerics: the expensive failures
Conversational WER is the metric vendors report. String accuracy on digits is the metric that decides whether your workflow functions, because account numbers, OTPs, order IDs and appointment dates are where a single error voids the entire turn.
Digit string accuracy, full-string correct, 10 to 14 digit strings:
| Codec path | Full string correct |
|---|---|
| Wideband 48kHz | 94.2% |
| Opus 16kHz | 92.8% |
| G.711 a-law | 81.4% |
| G.729 | 72.6% |
| AMR-NB 12.2k | 74.9% |
| AMR-NB 4.75k | 51.3% |
| Tandem | 63.8% |
On a congested mobile cell, half of all long digit strings come back wrong. Not one digit wrong out of fourteen; the whole string unusable. The confusion pairs are consistent and predictable: five and nine, six and seven, two and eight, and in Hindi छह and नौ.
Alphanumeric string accuracy, vehicle registrations and PNRs:
| Codec path | Full string correct |
|---|---|
| Wideband 48kHz | 88.7% |
| G.711 a-law | 68.3% |
| AMR-NB 12.2k | 59.1% |
| Tandem | 47.2% |
Worse than digits, because the letter confusions add to the digit confusions. B, D, E, G, P, T, V collapse into each other below 4kHz. This is why every airline IVR asks you to spell your PNR using a phonetic alphabet, and it is a solved problem that voice AI teams keep re-encountering because they assume the model will handle it.
Indian proper name accuracy:
| Codec path | Name correct |
|---|---|
| Wideband 48kHz | 91.3% |
| G.711 a-law | 79.6% |
| Tandem | 67.4% |
Text to speech degrades too, and differently
The narrowband problem is usually framed as a recognition issue. It cuts both ways.
Synthesised speech is generated at 22kHz or 24kHz, then downsampled and encoded to reach the caller. Modern neural TTS invests heavily in exactly the high-frequency detail that makes a voice sound human, and narrowband deletes it. The result is that the quality gap between an excellent TTS and a mediocre one compresses substantially over a phone line.
Mean opinion score, 1 to 5, native-speaker panel, Hindi:
| TTS system | Wideband MOS | G.711 MOS | Loss |
|---|---|---|---|
| Premium neural | 4.4 | 3.6 | -0.8 |
| Mid-tier neural | 4.0 | 3.4 | -0.6 |
| Older concatenative | 3.1 | 2.9 | -0.2 |
The premium system loses the most, because it had the most to lose. Over an 8kHz line the gap between premium and mid-tier narrows from 0.4 to 0.2, which has a direct procurement implication: paying a premium per-character rate for TTS quality that the telephone network discards is a common and avoidable overspend. We ran the wideband comparison across providers in the Indic TTS benchmark; the narrowband picture is meaningfully flatter.
Intelligibility for elderly listeners degrades further. Our over-60 panel rated narrowband synthesised speech 0.5 MOS lower than the general panel and requested repetition 2.1× more often, which is directly relevant to healthcare and pension workflows.
What to do about it
Test on the codec path you will run in production. The single most valuable change any evaluating team can make. Ask the vendor to run their accuracy test through G.711 and through a tandem G.711 to AMR-NB path. If they cannot, run it yourself: encode your test audio with
ffmpeg through the relevant codecs and feed it to their API. It takes an afternoon and it will change your vendor ranking.
Never accept long digit strings in one utterance. Chunk them. Ask for the last four digits, confirm, then the next four. Full-string accuracy on a four-digit chunk over G.711 is 96.8% in our data versus 81.4% for the full fourteen. Three chunked confirmations beat one failed one.
Use a phonetic alphabet for alphanumerics, and prompt for it explicitly. "Please say your registration number using words, like B for Bombay." Accuracy on our alphanumeric set rose from 68.3% to 89.1% with phonetic prompting on the same codec path.
Gate numeric slots on recogniser confidence, not on the model's judgement. Covered in more depth in our LLM benchmark for voice agents, but the narrowband data is the reason it matters: on a congested mobile leg the recogniser is wrong on half of long strings, and no downstream model can detect that.
Prefer a mid-tier TTS and spend the saving on recognition. The premium voice advantage largely does not survive the line. Recognition errors do.
Push for wideband where the network allows it. VoLTE calls can carry AMR-WB, and some Indian operators support it end to end. Your telephony provider may be transcoding to narrowband unnecessarily. Ask specifically whether the trunk negotiates wideband and under what conditions it downgrades. This is worth checking with the provider directly; the telephony partner comparison covers how the major Indian providers differ here.
Avoid tandem encoding where you can control routing. A call that goes G.711 to AMR-NB has been through two lossy stages. Sometimes routing choices can eliminate one.
Fine-tune on narrowband audio. If you have volume and an ML team, fine-tuning the recogniser on codec-degraded audio from your own traffic recovers 3 to 5 points of WER. Train on the codec path you actually run, not on clean audio with noise added, because codec artefacts are structured and additive noise is not.
The procurement checklist
Five things to write into an RFP:
- Accuracy figures must be quoted separately for wideband, G.711 8kHz, and tandem G.711 to AMR-NB paths.
- Accuracy must be broken out by language, not averaged across "Indian languages".
- Digit string accuracy must be reported as full-string-correct on 10 to 14 digit strings, not as digit-level accuracy, which flatters by roughly 15 points.
- The vendor must accept a sample of your own production recordings for evaluation.
- TTS quality claims must be demonstrated over an actual phone call, not a browser demo.
Any vendor unwilling to meet these has numbers that do not survive them. The Hyderabad hospital group would have learned in week one instead of week four.
What the degradation costs in business terms
Word error rate is an engineering number. Here is what the same degradation looks like on the workflows it breaks.
COD order confirmation. The agent needs to confirm an order ID and a delivery address. On wideband the confirmation completes in a single turn 89% of the time; over a tandem path it drops to 61%, and each failed confirmation costs an average of 34 additional seconds of call time plus a 12% higher rate of the customer abandoning the call entirely. On a 40,000-call monthly campaign that is roughly 380 additional hours of telephony spend and about 1,900 unconfirmed orders that fall back to manual calling. The workflow specifics are covered on our COD order confirmation page.
EMI payment reminders. The agent captures a promise-to-pay date and, on some flows, the last four digits of the account being debited. Numeric capture failures here do not merely waste a turn; they produce a record that is wrong rather than absent, which is worse. Two of the four NBFC deployments we audited in 2026 were writing unvalidated recognised dates straight into the collections system, and both had a small but non-zero population of promises recorded against the wrong month.
Appointment reminders in healthcare. Date and time confirmation over narrowband to an elderly patient panel is the worst combination in this entire dataset: high-value numeric content, the codec degradation, and a listener group that already requests repetition 2.1× more often. This is the Hyderabad case at the top of the post, and the fix that worked was not a model change but a redesign to yes/no confirmation of a stated time rather than open capture of a spoken one.
The general principle: narrowband degradation is survivable in workflows built around confirmation and fatal in workflows built around open capture. Redesigning the turn structure is cheaper and faster than chasing the last two points of word error rate.
Compliance angles specific to audio quality
Recording quality and evidentiary value. For RBI Fair Practices Code collections and for IRDAI-supervised sales, call recordings are the evidence that the mandated disclosures were made. Recordings captured after aggressive narrowband compression are still admissible, but low-bitrate AMR-NB recordings of a rapid disclosure can be genuinely hard for a human reviewer to verify. Where you control the recording point, record on the least-degraded leg available rather than on the final output.
Consent capture over degraded audio. If a workflow captures verbal consent, the consent turn is the one turn worth protecting hardest. Slow the agent's delivery, use an explicit yes/no rather than open response, and log the recogniser confidence alongside the transcript. A consent record with an 0.4 confidence score attached is a record you know to treat carefully; one with no confidence stored is a record you will have to argue about later.
DPDP and retained audio. Codec-degraded or not, retained call audio is personal data. The benchmark work described here involves keeping a labelled test set of real customer utterances, which needs its own retention policy and purpose limitation. Use anonymised or consented samples for the test set, and do not let a benchmarking corpus quietly become an indefinitely retained archive.
A two-week measurement plan
Days 1 to 3: assemble the test set. Two hundred to three hundred utterances from your own traffic, weighted toward the content types your workflow depends on. If your workflow captures account numbers, over-sample digit strings. Transcribe them by hand; this is your ground truth and it must not come from the recogniser.
Days 4 to 5: build the codec chain. Encode every file through wideband reference, G.711 a-law, AMR-NB at 12.2k and 4.75k, and the tandem path.
ffmpeg handles all of these. Keep the wideband originals; the comparison is the whole point.
Days 6 to 8: run and score. Push each version through your current platform and any candidates. Score word error rate for sentences and full-string-correct for digits and alphanumerics separately, and never collapse them into one number.
Days 9 to 10: segment. By language, by speaker origin, by content type. The averages will hide the segment that is actually failing, which in most Indian deployments turns out to be Tier-3 speakers on mobile legs saying long numbers.
Days 11 to 14: fix the workflow before the model. Chunk digit capture, add phonetic prompting for alphanumerics, convert open capture to confirmation where the content is high-stakes, and gate numeric slots on confidence. Re-run the same test set. In every deployment we have taken through this, workflow changes recovered more end-to-end accuracy than any model swap available at the time.
What changes in the next twelve months
VoLTE and VoNR penetration continues to rise across Indian circles, which moves more legs onto AMR-WB and effectively adds 3.4kHz of bandwidth. That is the single biggest structural improvement available and it requires nothing from voice AI vendors. It will not reach BSNL or rural landline traffic on any near horizon.
Codec-aware recognisers, trained with the codec chain in the augmentation pipeline rather than on clean audio, are becoming standard practice among the better providers. Expect published narrowband WER to improve by several points over the next year purely from training methodology.
Neural bandwidth extension, reconstructing plausible high-frequency content before recognition, is showing real gains in research and mixed results in production. It helps intelligibility more than it helps recognition, because the reconstructed detail is plausible rather than true and the recogniser can be misled by it. Worth watching, not worth deploying yet.
Bottom line
Voice AI demos run at 48kHz. Indian phone calls run at 8kHz, often through two lossy codecs. In our benchmark that gap took conversational word error rate from 6.1% to 19.7%, dropped full-string digit accuracy from 94.2% to 63.8%, and erased most of the quality advantage of premium text-to-speech.
None of this is exotic. It is the ordinary condition of the Indian telephone network, and it is entirely predictable at evaluation time if you insist on codec-matched testing. Chunk your digits, prompt phonetically for alphanumerics, gate numeric slots on recogniser confidence, and stop paying a premium for TTS detail the line discards.
If you want the codec-degraded test set or want your current platform measured against these paths, talk to our team.
Frequently Asked Questions
Tags :










