All Blogs

    Voice AI Concurrency Benchmark India 2026: What Breaks at 500 Simultaneous Calls

    16 Mins ReadSep 29, 2026
    Voice AI Concurrency Benchmark India 2026: What Breaks at 500 Simultaneous Calls

    The procurement call that prompted this post lasted eleven minutes. A VP of Collections at a mid-size NBFC in Mumbai had run a voice AI pilot for six weeks, liked the numbers, and wanted to scale from 4,000 calls a day to 60,000 before the next quarter's delinquency cycle. His question to the vendor was simple: can your platform handle it. The vendor said yes. Both of them meant something different by the word, and neither of them said a number.

    He was asking about calls per day. The vendor was answering about calls per day. The thing that would actually break in week two was concurrency, which neither of them had mentioned, and the calls-per-second ceiling on his SIP trunk, which nobody had looked up.

    That gap is where most Indian voice AI rollouts stall. Not at the pilot, which works. At the ramp.

    What this post argues

    Concurrency, not volume, is the binding constraint on outbound voice AI in India, and the ceiling is almost never set by the voice AI platform itself. It is set by whichever of five stacked layers runs out of headroom first: the SIP trunk's calls-per-second cap, the ASR provider's streaming session quota, the LLM's requests-per-minute limit, the TTS concurrent stream allowance, or your own orchestration layer's connection pool. This post gives measured degradation curves from 25 to 750 simultaneous call legs, the arithmetic to convert a daily target into a concurrency number, the rate-limit ceilings you should ask every vendor to put in writing, and a four-week load-test plan you can hand to your platform team.

    Why capacity planning got harder in 2026

    Two things changed. The first is that Indian outbound volumes climbed sharply once per-outcome pricing made high-volume dialing affordable. A collections team that ran 8,000 calls a day in 2024 because each one cost real money now runs 60,000 because the marginal call is cheap. The arithmetic of the business changed faster than the arithmetic of the infrastructure.

    The second is architectural. Voice AI in 2024 was mostly a single model doing speech to text, a rules engine, and a text to speech voice. In 2026 a single call leg holds open a streaming ASR session, one or more LLM inference requests per turn, a streaming TTS connection, and often a function call out to a CRM or payments API mid-conversation. One call is no longer one resource. It is four to six concurrent resources, each metered separately, each with its own rate limit, each capable of becoming the bottleneck on its own.

    The result is that capacity failures in 2026 are rarely clean. The system does not refuse the 501st call. It accepts it, queues it somewhere invisible, and degrades the quality of all 500 calls already in flight.

    The arithmetic nobody does before signing

    Start with the conversion that matters. Concurrency is not volume. It is the number of call legs alive at the same instant.

    The formula:

    Concurrent channels = (dial attempts per hour × average leg duration in seconds) / 3600
    

    The trap is "dial attempts", not "connected calls". In Indian outbound, connect rates run 18 to 32 percent depending on the hour, the list age and whether the number is Tier 1 or Tier 3. Every unconnected attempt still occupies a channel for the ring duration.

    Work a real target. An NBFC wants 50,000 connected conversations a day, spread across the two productive windows Indian outbound actually has, roughly 11am to 1pm and 5pm to 8pm, so call it five productive hours.

    InputValue
    Connected calls needed per hour10,000
    Connect rate on a 30-day-old NBFC list25%
    Dial attempts required per hour40,000
    Unanswered attempts per hour30,000
    Average ring time before abandon20 seconds
    Average answered call duration48 seconds
    Channel-seconds from unanswered legs600,000
    Channel-seconds from answered legs480,000
    Total channel-seconds per hour1,080,000
    Concurrent channels required300

    Three hundred simultaneous channels to deliver fifty thousand conversations. Most buyers, asked to guess, say thirty.

    Now the second number, which is the one that actually bites. Calls per second, or CPS, is the rate at which your trunk will let you originate new calls. Indian carriers commonly provision 5 to 30 CPS per trunk unless you negotiate otherwise, and the default on a new account is usually at the low end. At 10 CPS, ramping from zero to 300 concurrent channels takes a minimum of 30 seconds even if everything downstream is instant. At 5 CPS it takes a minute. Dialers that try to fill the pipe faster get SIP 503 responses, and badly written retry logic turns those into a thundering herd that makes the problem worse.

    CPS is the single most commonly missed number in Indian voice AI procurement. Ask for it in writing. See our telephony integration guide for how this interacts with carrier selection.

    The five ceilings, and which one you hit first

    Each layer has an independent quota. Your effective concurrency is the minimum across all of them, not the number the voice AI vendor quotes.

    LayerTypical ceilingHow it failsWho sets it
    SIP trunk channels100 to 2,000 per trunkSIP 503, no circuit availableCarrier (Plivo, Exotel, Ozonetel, Knowlarity, Tata Tele, Twilio)
    SIP calls per second5 to 30 CPSOrigination refused, ramp stallsCarrier, negotiable
    Streaming ASR sessions100 to 500 concurrentSession refused or transcript lagASR provider
    LLM requests and tokens per minuteVaries widely by tierHTTP 429, then mid-call silenceModel provider
    TTS concurrent streams50 to 500Audio stutter, gaps, truncationTTS provider
    Orchestration connection poolSelf-imposedQueueing, unbounded latencyYou

    In the deployments we have load-tested, the binding constraint is the LLM rate limit about half the time, SIP CPS about a third of the time, and TTS streams most of the rest. The voice AI platform's own advertised concurrency is almost never the real ceiling, which is why it is a nearly useless number to shop on.

    One consequence worth stating plainly: a vendor who answers "we support unlimited concurrency" is telling you they have not load-tested with your model provider on your account tier. Unlimited is not a capacity claim. It is the absence of one.

    Measured degradation from 25 to 750 channels

    The numbers below come from staged load tests on Indian mobile numbers over a mixed Jio, Airtel and Vi footprint, using a Hindi-plus-English agent on a collections script with a 48-second target handle time. Time to first byte is measured from end of caller speech to first audible agent audio. Barge-in miss rate is the percentage of caller interruptions the agent failed to yield to within 300ms.

    Concurrent legsTTFB p50TTFB p95Barge-in miss rateCall drop rateUsable
    25620ms940ms1.2%0.3%Yes
    100680ms1,180ms2.1%0.5%Yes
    250790ms1,640ms4.8%1.1%Yes, with headroom warnings
    400930ms2,210ms7.9%2.0%Marginal
    5001,120ms2,890ms11.3%3.4%No
    7501,880ms5,200ms22.6%9.7%No

    The shape matters more than any single row. Degradation is not linear. Between 25 and 250 channels the p50 moves 170ms, which a caller will not consciously notice. Between 400 and 750 it moves nearly a second, and p95 more than doubles. The knee sits between 250 and 400 on this configuration.

    What a caller experiences at each stage is worth spelling out, because the metrics understate it. At p95 TTFB of 1.6 seconds the agent sounds thoughtful. At 2.9 seconds it sounds broken, and Indian callers in particular start talking over it, which drives the barge-in miss rate up, which produces the overlapping-speech failure that makes people hang up. The drop rate and the barge-in rate are not independent variables. They feed each other past the knee.

    For the underlying latency budget that produces these numbers, see our sub-500ms latency architecture breakdown and the voice AI latency benchmarks for India.

    Why the knee moves

    The same platform will knee at a different point on your traffic. Four variables move it materially:

    Codec and bandwidth. Narrowband 8kHz G.711 traffic, which is most Indian PSTN, costs less CPU per leg than wideband but produces worse ASR, so some stacks compensate with a heavier model that costs more. Our 8kHz narrowband ASR and TTS benchmark covers the accuracy side of this tradeoff.

    Language mix. Code-switched Hindi and English turns trigger longer ASR hypotheses and more LLM tokens per turn than clean English. A Hindi-heavy collections list will knee 15 to 25 percent earlier than an English-heavy support queue on identical infrastructure.

    Function calls mid-conversation. Every CRM read, UPI link generation or payment status check inserts an external API round trip into the turn. Under load these queue behind each other. A script with two function calls per conversation knees noticeably earlier than one with none.

    Retry storms. The most common self-inflicted wound. A failed dial that retries immediately, at scale, converts a small carrier hiccup into sustained CPS exhaustion.

    What actually goes wrong past the ceiling

    Seven failure modes, in rough order of how often we see them.

    Silent queueing. The orchestration layer accepts more calls than downstream can serve and queues them rather than rejecting. Nothing errors. Nothing alerts. TTFB climbs across every call in flight, including the ones that were fine. This is the worst failure because your dashboards stay green while call quality collapses. Fix: bound your queues explicitly and shed load at the edge, loudly.

    LLM 429 mid-call. The model provider rate-limits you between turn three and turn four. The agent goes silent. The caller says hello twice and hangs up. Fix: provision a fallback model on a separate quota and fail over within the turn, not after it.

    TTS stream starvation. Concurrent synthesis streams exceed quota and the audio arrives in fragments. Callers describe this as the agent stuttering or cutting out. Fix: pre-synthesise the fixed segments of the script, which in a collections flow is often 40 percent of the words spoken.

    Carrier throttling. SIP 503 or 486 at a rate that rises with your CPS. Often the carrier is protecting the destination network, not you. Fix: rate-limit origination on your side to just under the negotiated CPS, with jitter.

    DLT scrubbing bottleneck. TRAI DLT consent scrubbing has to happen at dial time, not at queue-build time, because consent can be withdrawn between the two. At scale the scrubbing service becomes a serialisation point. Fix: batch-check with a short TTL cache, and confirm with your provider what their scrub throughput actually is.

    Recording write backpressure. Every call writes audio. At 500 concurrent legs that is a sustained write load, and if your storage layer applies backpressure it propagates up into the media path. Fix: write to a local buffer and ship asynchronously.

    Barge-in degradation. Voice activity detection gets less responsive as CPU contention rises. The agent talks over people. This is measurable before it is audible, which makes it a good early-warning metric.

    The numbers to hold vendors to

    When you are comparing platforms, these are the questions that separate a real capacity answer from a sales one. Ask for each in writing, with the measurement conditions attached.

    QuestionWhat a good answer looks like
    Sustained concurrent legs at p95 TTFB under 1.5sA number, plus the codec, language mix and script length it was measured on
    Calls per second on originationA number, plus whether it is negotiable and at what cost
    ASR concurrent session quotaA number, plus behaviour on exhaustion (refuse vs degrade)
    LLM rate limit and fallback behaviourNamed fallback model, in-turn failover, measured failover latency
    TTS concurrent stream quotaA number, plus whether pre-synthesis is supported
    Behaviour at ceilingExplicit load shedding with an error, not silent queueing
    Time to ramp zero to full concurrencySeconds, derived from CPS

    A vendor who can answer all seven has load-tested. A vendor who answers the first one with a large round number and cannot answer the rest has not. This is a cheaper filter than a pilot.

    For how the major India telephony providers differ on the origination side, see our comparison of telephony partners for voice AI in India, and for the commercial side, voice AI pricing in India. If Twilio is on your shortlist specifically, the Caller Digital versus Twilio comparison covers the India-specific origination differences.

    Compliance interacts with concurrency

    Two regulatory constraints have throughput characteristics, and both are routinely discovered late.

    TRAI DLT scrubbing must occur at dial time. Consent registered against a template and a sender can be withdrawn at any point, so a scrub performed when the campaign list was built is not a defence if the call goes out six hours later. At low volume this is invisible. At 40,000 dial attempts an hour the scrub service is in the hot path of every single call, and its own rate limit becomes yours.

    DPDP Act 2023 requires purpose-bound consent and an audit trail per call. The audit write is another per-call resource. Under load, teams sometimes make audit logging asynchronous and lossy to protect latency, which is an understandable engineering instinct and a compliance problem. The audit trail is the thing you will be asked to produce. Make it durable and bound your concurrency to what you can durably log.

    RBI Fair Practices Code constrains collections calling windows, which compresses your productive hours and therefore raises the concurrency you need for a given daily volume. A team that could spread 50,000 calls across twelve hours would need half the concurrency of one restricted to five. Compliance narrows the window, and the window sets the peak. For sector specifics see our BFSI industry page and the EMI and payment reminder use case.

    Inbound concurrency is a different problem

    Everything above assumes outbound, where you control the dial rate and can shape traffic to fit your ceiling. Inbound inverts the problem. The caller decides when to call, so your peak concurrency is set by your customers rather than by your dialer, and the one lever you have on the outbound side, slowing down, does not exist.

    Indian inbound traffic is spikier than most capacity models assume. A missed-call-back campaign, an SMS blast, a payment-due date and a delivery exception notification all produce inbound bursts that arrive within minutes of the trigger. A single SMS to 200,000 borrowers on EMI due date will return a call spike that is an order of magnitude above the day's baseline, concentrated in the first fifteen minutes. Teams that sized inbound capacity on average concurrency discover this once.

    Three practical consequences. First, size inbound on the 99th percentile minute, not the busy hour, because an hourly average hides a burst entirely. Second, coordinate outbound campaigns with inbound capacity, since the two share every layer below the application, and an aggressive outbound push during an inbound spike will starve both. Third, build the queue experience deliberately: an inbound caller who hits capacity should get an acknowledgement and a callback commitment, not a ring-out, because an abandoned inbound call from a customer who was trying to pay you is considerably more expensive than a failed outbound attempt.

    The asymmetry is worth holding onto. Outbound capacity failures cost you throughput. Inbound capacity failures cost you the customers who were already trying to reach you.

    A four-week load test you can actually run

    Most teams either skip load testing or attempt it once, a week before go-live, at full target volume, which tells you only that something broke. Staged testing tells you what and when.

    Week 1: instrument, do not load. Before adding traffic, confirm you can measure TTFB per turn, barge-in yield time, per-layer error rates and queue depth at every stage. If you cannot see queue depth, you cannot detect silent queueing, which is the failure you are most likely to hit. Run at current production volume and record the baseline.

    Week 2: find the knee. Ramp in steps: 25, 50, 100, 150, 250, 400. Hold each step for at least twenty minutes, because some failures are cumulative rather than instantaneous, particularly storage backpressure and connection pool exhaustion. Record the full table above at each step. Stop when p95 TTFB crosses 1.5 seconds or barge-in miss rate crosses 5 percent, whichever comes first. That is your knee.

    Week 3: break it deliberately. Push 30 percent past the knee and watch how it fails. You are testing whether the system sheds load cleanly or degrades silently. Then test each layer's failure in isolation: revoke the LLM quota, throttle TTS, force SIP 503s. Confirm each one produces a visible error and a sensible caller experience, which usually means a graceful handoff or a callback promise rather than dead air. Also test the ramp specifically: how long from zero to knee-level concurrency, and does the dialer respect CPS.

    Week 4: run the real shape. Production traffic is not flat. Indian outbound has two sharp peaks and a long dead middle. Replay a real day's shape at target volume, including the 11am ramp, which is the steepest and the one most likely to hit CPS limits. Include a retry storm simulation. Sign off on the concurrency number you can hold through the 11am ramp, not the one you can hold at 3pm.

    Set your production ceiling at roughly 70 percent of measured knee. The headroom absorbs list-quality variance, carrier variability and the fact that your script will get longer over time, which it always does.

    For the connect-rate assumptions that feed the arithmetic at the top of this post, see our outbound connect rate benchmark for India.

    Bottom line

    Concurrency is the number that determines whether a voice AI deployment scales, and it is not the number most buyers ask about. Convert your daily conversation target into concurrent channels using dial attempts rather than connects, and the number will be three to ten times higher than your intuition. Then find the real ceiling, which lives in whichever of the telephony, ASR, LLM, TTS and orchestration layers has the least headroom on your account, not in the platform's marketing page. Test in stages, find the knee, and run production at seventy percent of it. A platform that degrades loudly at its limit is worth more than one that claims not to have one.

    Frequently Asked Questions

    Kanan Richhariya

    Kanan Richhariya

    Other Blogs

    1.png
    Voice AI & Voice Technology

    Cost Per Resolved Contact in India 2026: Chat vs Voice AI vs Human Agent Economics

    Kanan Richhariya

    Publish: Aug 5, 2026

    2.png
    Voice AI & Voice Technology

    Agentic Chatbots for Customer Care in India 2026: What Actually Resolves a Ticket

    Kanan Richhariya

    Publish: Aug 5, 2026

    2.png
    Voice AI & Voice Technology

    India Voice AI Market Size 2026: Conversational AI, STT, TTS and Speech Analytics Reconciled

    Kanan Richhariya

    Publish: Aug 4, 2026

    1.png
    Voice AI & Voice Technology

    CCaaS Pricing in India 2026: Per-Agent vs Pay-As-You-Go TCO for a 20-Seat Contact Centre

    Kanan Richhariya

    Publish: Aug 4, 2026

    Real-Time Voice AI Diagnostics_ From Customer Support to Health ScreeningI (2).png
    Voice Automation Strategies

    Welcome and Onboarding Calls with Voice AI in India 2026: The First-72-Hours Playbook That Cuts Early Churn

    Kanan Richhariya

    Publish: Jul 23, 2026

    Real-Time Voice AI Diagnostics_ From Customer Support to Health ScreeningI.png
    Voice Automation Strategies

    Bill Payment Reminder Calls in India 2026: The Voice AI Playbook for Utilities, Subscriptions, Broadband and Recharge

    Kanan Richhariya

    Publish: Jul 23, 2026

    AI Voice Deepfake Fraud in India 2026.png
    Voice AI & Voice Technology

    AI Voice Deepfake Fraud in India 2026: How Legitimate AI Calling Stays on the Right Side of Trust

    Kanan Richhariya

    Publish: Jul 20, 2026

    Voice AI for small business.png
    Voice Automation Strategies

    Voice AI for Small Businesses and MSMEs in India 2026: The Owner's Playbook for AI Calling Without an Enterprise Budget

    Kanan Richhariya

    Publish: Jul 20, 2026

    Open AI Realtime.png
    Voice AI & Voice Technology

    OpenAI Realtime API vs Voice AI Platforms India 2026: What It Actually Takes to Ship AI Calling Agents on a Raw Model API

    Kanan Richhariya

    Publish: Jul 20, 2026

    Voice AI vs BPO.png
    Voice Automation Strategies

    Voice AI vs Philippines BPO for Outbound Calling 2026: The Real Cost Per Minute Math

    Kanan Richhariya

    Publish: Jul 20, 2026

    Caller Digital

    © 2025 Caller Digital | All Rights Reserved