Sales

How to Scale Voice AI Agents to 1,000+ Concurrent Calls Without Breaking the Bank

Albert W.
October 10, 2026

A production voice AI agent costs between $0.012 and $0.147 per active call minute, depending on architecture. A fully managed stack (Vapi + Deepgram + a frontier LLM + ElevenLabs + Twilio) runs about $0.147/min before fixed fees. A self-hosted control plane with open-weight models runs about $0.012/min.

At 1 million billable minutes a month with capacity for 1,000 concurrent calls, that gap is $158,600 versus $13,590 a month: a 91.4% reduction. Even if self-hosted GPUs sit idle around the clock, the saving stays above 86%.

Three levers produce almost all of the savings:

  • Remove the wrapper markup. Managed platforms add $0.05–$0.055/min plus $8–$10 per concurrent line above the included tier.
  • Replace proprietary TTS and LLM endpoints. Together they are 51% of a managed stack's per-minute cost; open-weight models cut them by more than 98%.
  • Skip work you have already done. Semantic caching and dynamic model routing avoid paying for repeat turns and over-sized models.

The hidden costs of voice AI

A voice agent's bill is not one price. It is five metered services stacked in sequence, each with its own markup, plus fixed fees that only appear at scale.

The five-stage pipeline

Every spoken turn passes through the same chain:

  1. Speech-to-Text (STT): streaming transcription of the caller's audio.
  2. LLM reasoning: the model decides what to say.
  3. Text-to-Speech (TTS): the reply is synthesized into audio.
  4. Orchestration: a control plane manages turn-taking, interruptions and state.
  5. Telephony / transport: audio travels over SIP, PSTN or WebRTC.

Managed platforms add a platform layer on top of every model vendor. Vapi charges $0.05/min for hosting and Retell $0.055/min for voice infrastructure, with models billed separately; Bland bundles the LLM, STT and TTS into a flat $0.12–$0.14/min. Stacked with vendor costs, the all-in price lands between $0.11 and $0.31 per minute.

Typical managed-stack cost per minute

ComponentExample vendorCost per minuteShare of total
OrchestrationVapi hosting fee$0.050034%
Text-to-SpeechElevenLabs Flash~$0.040027%
LLMFrontier model (GPT-5.6 class, long context)~$0.035024%
TelephonyTwilio outbound$0.014010%
Speech-to-TextDeepgram Nova-3 (pay-as-you-go list)$0.00775%
Total~$0.1467100%

Orchestration and TTS alone are 61% of the bill. Neither is the part that makes an agent smart.

Five costs that don't show up in a pricing calculator

1. Latency extends billable time. Every millisecond the caller waits is a millisecond you pay all five vendors for. Illustration: a 5-minute call with 30 agent turns and 700 ms of extra delay per turn adds 21 seconds, or 7% more billable time. Slow agents also cause more "Hello? Are you there?" repeats, which add turns. (Dialers have the same problem — see our guide to the causes of dialer latency.)

2. Call time is billed, not speech time. Platforms bill the whole connection; Retell, for example, states that silence and hold time are billable. Pauses, hold time and network buffering keep the meter running; industry estimates put the inflation at 12% to 20% over actual spoken audio.

3. Concurrency penalties. Platforms include only 4 to 30 concurrent lines, depending on plan. Beyond that, Vapi charges $10 and Retell $8 per line per month. On Vapi's Core package (10 lines included), running 1,000 concurrent lines means $9,900 a month in line fees before a single call is answered.

4. Compliance add-ons. Vapi's HIPAA add-on costs $2,000/month. Zero Data Retention is now included in Vapi's paid packages, but Retell and Bland reserve BAAs for enterprise contracts.

5. TTS character density. TTS is billed per character, and one spoken minute averages 750 to 900 characters. At ElevenLabs Flash's $40 per million characters, that is $0.030 to $0.036 per minute; ElevenLabs' own estimate is about $0.04/min. Verbose prompts, spelled-out numbers or stray formatting push character counts, and costs, above what linear models predict.

The voice AI cost-per-minute formula

The true cost of one call minute is the sum of five variable costs plus your fixed monthly fees spread across monthly volume:

Ccall = CSTT + CLLM + CTTS + COrch + CTel + CFixed ÷ MMonthly
TermWhat it measuresHow to calculate it
CSTTSpeech-to-text per minuteVendor rate per audio minute
CLLMLLM inference per minuteTokens per minute × price per token (or GPU-hour cost ÷ minutes served)
CTTSText-to-speech per minuteCharacters per minute × price per character
COrchOrchestration / control plane per minutePlatform fee, or your own compute ÷ minutes
CTelTelephony per minuteSIP / PSTN carrier rate
CFixedFixed monthly feesLine reservations, compliance add-ons, cluster costs
MMonthlyBillable minutes per monthYour volume

The fixed-fee term is why volume changes the answer. The same $11,900 in monthly fees adds $0.012/min at 1 million minutes but $0.119/min at 100,000.

Unit economics: managed APIs vs. an integrated control plane

Moving from a managed vendor stack to a self-hosted control plane cuts cost per million minutes from $158,629 to $13,590, while cutting P95 latency from over 850 ms to under 400 ms.

The comparison below assumes 1 million billable minutes per month with capacity for 1,000 concurrent calls.

MetricTier 1: Fully managed stackTier 2: Optimized commercial APIsTier 3: Self-hosted control plane
Orchestration$0.0500 (Vapi hosting)$0.0130 (LiveKit Cloud: $0.010 agent session + $0.003 SIP)$0.0020 (self-hosted LiveKit on Kubernetes, estimate)
Speech-to-Text$0.0077 (Deepgram Nova-3, pay-as-you-go list)$0.0065 (Deepgram Nova-3, Growth list)$0.0065 (Deepgram Nova-3, Growth list)
LLM~$0.0350 (frontier GPT-5.6-class model, long context)~$0.0010 (GPT-5.6 Luna)$0.00019 (Llama-3.1-8B FP8 on vLLM, NVIDIA L4)
Text-to-Speech$0.0400 (ElevenLabs Flash)$0.0090 (Inworld Realtime TTS 2.0 via LiveKit)$0.0005 (Kokoro-82M, self-hosted)
Telephony$0.0140 (Twilio outbound)$0.0050 (Telnyx SIP outbound)$0.0032 (Telnyx SIP inbound)
Variable cost per minute$0.1467$0.0345$0.0124
Fixed monthly overhead$11,929 (Vapi Core $29 + 990 extra lines + HIPAA)$500 (LiveKit Scale plan)$1,200 (Kubernetes control plane)
Total per 1M minutes$158,629$35,000$13,590
Saving vs. Tier 1—77.9%91.4%
P95 latency (audio-to-audio)850–1,300 ms450–650 ms220–380 ms
Concurrency ceiling4–30 lines included, then $10/lineLiveKit Scale up to 600 sessions; Deepgram Growth up to 225 streams; Enterprise beyondScales with cluster nodes

STT is now Tier 3's largest cost. Deepgram's limited-time streaming promotion ($0.0042/min on Growth) would bring Tier 3 to about $0.010/min.

Where the $0.00019 LLM figure comes from

Self-hosted LLM cost depends on how many calls one GPU can serve at once. Here is the math for Llama-3.1-8B in FP8 on vLLM, using NVIDIA L4 GPUs (24 GB VRAM).

  1. Streams per GPU: an L4 serves up to 64 concurrent streams without dropped requests, keeping P95 time-to-first-token under 300 ms.
  2. Throughput check: a voice conversation needs about 150 output tokens per minute. 64 streams × 150 = 9,600 tokens/min, or 160 tokens/sec, far below the L4's ~947 tokens/sec ceiling. VRAM for the KV cache is the real limit, not compute.
  3. GPUs needed: 1,000 streams ÷ 64 per GPU = 15.6, rounded up to 16 GPUs.
  4. Hourly cost: 16 × $0.70/hour = $11.20/hour.
  5. Cost per minute: 16 GPUs at full load serve 60,000 call minutes per hour. $11.20 ÷ 60,000 = $0.000187/min.
CLLM = (16 × $0.70) ÷ (1,000 × 60) = $0.000187 per minute

The utilization caveat. That figure assumes the GPUs are busy. If you keep all 16 GPUs running 24/7 ($8,064/month) but only serve 1 million minutes, the LLM cost is about $0.008/min. Tier 3's total then rises to roughly $21,460 per million minutes: still 86% below Tier 1. Autoscaling GPU nodes with traffic closes most of that gap.

Reference architecture for 1,000+ concurrent calls

A self-hosted control plane runs on Kubernetes (GKE or EKS) in three layers, so no third party sits between the caller and your models.

  1. Telephony and transport: calls enter through wholesale SIP trunking (for example Telnyx) into a LiveKit cluster.
  2. Orchestration: autoscaled agent workers coordinate each turn: streaming STT, turn-taking, barge-in handling and the semantic cache lookup.
  3. Inference: a GPU pool serves the open-weight LLM (Llama-3.1-8B on vLLM) and TTS (Kokoro-82M). Only cache misses reach it.

Resilience at 1,000 calls. Three safeguards keep conversations alive during spikes:

  • Circuit breakers with backoff: repeated 429 or 503 errors trip a breaker and reroute to a fallback (Deepgram to AssemblyAI, local Llama to a cloud endpoint).
  • Barge-in handling: each worker keeps a 200 ms audio ring buffer. When the caller interrupts, it cancels agent audio, halts TTS and flushes the LLM queue without dropping frames.
  • Built-in Zero Data Retention: transcripts and audio live only in RAM, never on disk, with SRTP/TLS end to end. This replaces paid compliance add-ons such as Vapi's $2,000/month HIPAA tier.

How to cut latency and billable tokens

Four techniques move a voice stack from Tier 1 to Tier 3 economics: self-hosting open-weight models, semantic caching, dynamic model routing and speculative generation.

1. Self-host open-weight TTS and LLMs

TTS and the LLM are the largest swappable costs, and open-weight models now undercut commercial APIs by 37× or more.

TTS: Kokoro-82M. This 82-million-parameter model (StyleTTS2 architecture) uses under 1 GB of VRAM. It reaches a real-time factor of 0.02 to 0.03, meaning 10 seconds of audio is synthesized in about 200 ms.

TTS optionDeploymentFirst-audio latencyCost per 1M characters
Kokoro-82M (RTX 5090)Self-hosted28 ms (reported)~$0.30 equivalent
Kokoro-82M (RTX 3090)Self-hosted40 ms (reported)~$0.50 equivalent
Kokoro-82M (DeepInfra)Managed endpoint170 ms (reported)$0.99
Cartesia Sonic 3.6Cloud API, Scale plan128 ms (reported)~$37
ElevenLabs FlashCloud API~75 ms (vendor-stated)$40

Even on a managed endpoint, Kokoro-82M is about 40× cheaper than ElevenLabs Flash and 37× cheaper than Cartesia Sonic. Its advantage is price, not speed: on GPU-hosted servers it is competitive, but ElevenLabs quotes about 75 ms for Flash.

LLM: 8B models on vLLM. FP8 quantization shrinks Llama-3.1-8B or Qwen3-8B to about 8 GB, leaving over 15 GB of a 24 GB GPU for KV cache. Qwen3-8B scales to 256 concurrent requests on one A10 or L4 without raising error rates. For Llama-3.1-8B, capping concurrency at 64 per GPU keeps P95 time-to-first-token under 300 ms.

2. Semantic audio and response caching

Up to 40% of turns in support, scheduling and logistics calls are repeats: business hours, account status, policy disclosures. A semantic cache answers them without touching the LLM or TTS.

  1. Embed: convert the streaming transcript into a vector.
  2. Look up: query a vector index (Redis Vector Search or Qdrant) of past prompt-response pairs.
  3. Compare: a cosine similarity of 0.92 or higher counts as a hit.
  4. Serve: on a hit, stream pre-synthesized audio straight from memory.
  5. Fall back: on a miss, run the normal LLM + TTS path and cache the result.

Cached turns respond in under 30 ms and cost $0.00 in LLM and TTS fees.

3. Dynamic model routing

Most turns don't need a frontier model. Route each turn by complexity:

RouteHandlesModel
Tier A: micro-intents"Yes", "Tomorrow at 3 PM", balance checks, form fillingFine-tuned 3B model (Llama-3.2-3B, Qwen-2.5-3B) or a deterministic state machine
Tier B: standard flowMulti-turn conversation and reasoningLlama-3.1-8B FP8 cluster
Tier C: edge casesAmbiguous requests, policy exceptionsFrontier model via API (e.g. Claude, GPT-5.6)

Routing 70% of turns to Tiers A and B can cut average LLM cost by more than 80% versus sending every turn to a frontier model.

4. Speculative decoding and streaming speculative solving

Time-to-first-token is what callers perceive as lag. Speculation attacks it at two levels.

Speculative decoding (model level). A small draft model (e.g. a 1B model) proposes several tokens; the 8B target model verifies them in one forward pass. This gives a 2× to 3× speedup with no change to output quality.

Streaming speculative solving (pipeline level). Traditional pipelines wait 400–600 ms of silence before acting, then run STT, LLM and TTS in sequence. Speculative solving starts generating before the caller finishes.

ApproachSilence waitProcessingFirst audioInterruption cost
Traditional sequential VAD400–600 msSequential: VAD → STT → LLM → TTS~1,200 msHigh: full context reset
Streaming speculative solving100–150 msParallel: partial transcripts → speculative draft220–300 msNear zero: stream cancelled instantly

It works in four steps:

  1. Continuous partials: streaming STT sends partial transcripts as the caller speaks.
  2. Semantic triggers: a lightweight classifier predicts when a phrase is nearly complete ("Can you move my appointment to tomorrow at…").
  3. Pre-emptive generation: the LLM drafts a reply before the caller finishes.
  4. Validate and commit: if the final transcript matches, audio plays immediately. If the caller changes course ("…actually, make it Friday"), the draft is discarded unheard.

The result is 200–300 ms voice-to-voice latency, matching natural human turn-taking.

A phased path from prototype to 1,000+ concurrent calls

Don't self-host on day one. Move down the cost curve as volume justifies the engineering work.

PhaseConcurrent callsStackGoalApprox. cost per minute
1. Prototype0–50Managed platform (Vapi or Retell) + Deepgram, a frontier LLM, ElevenLabsIterate fast on prompts, flows and voice design$0.11–$0.31
2. Decouple50–250LiveKit Cloud control plane + commercial APIs (Deepgram Growth, GPT-5.6 Luna, budget TTS)Remove wrapper fees; ~78% savings~$0.035
3. Self-host250–1,000+Self-hosted Kokoro-82M + Llama-3.1-8B FP8 on vLLM, semantic cache, speculative solvingLowest cost and latency at scale~$0.010–$0.012

The phase boundaries are rules of thumb. The real trigger is when monthly savings exceed the cost of the engineers maintaining the stack.

Want a voice AI agent without stitching together five vendors? See the Bearworks AI Agent, or book a demo to talk through your call volume.

Frequently asked questions

How much does it cost to build a voice AI agent at scale?

Running cost ranges from about $0.147 per minute on a fully managed stack to about $0.012 per minute on a self-hosted control plane. At 1 million minutes a month, that is roughly $158,600 versus $13,590, including fixed fees.

What is the cheapest STT and TTS stack for a voice agent?

For TTS, Kokoro-82M costs about $0.30–$0.99 per million characters, self-hosted or on DeepInfra, versus $37–$40 for Cartesia Sonic and ElevenLabs Flash. For STT, Deepgram Nova-3 lists at $0.0065 per streaming minute on its Growth plan ($0.0042 during its current promotion). Pair them with a self-hosted LiveKit control plane and low-cost SIP trunking such as Telnyx ($0.0032–$0.005/min).

Why are managed voice AI platforms so expensive at scale?

They add a platform fee ($0.05/min on Vapi, $0.055/min on Retell) on top of every model vendor, bill silence and hold time, and charge $8–$10 per concurrent line beyond the 4–30 included. At 1,000 lines on Vapi, line fees alone are about $9,900 a month.

How many GPUs do I need for 1,000 concurrent voice calls?

For the LLM, about 16 NVIDIA L4 GPUs running Llama-3.1-8B in FP8 on vLLM, at 64 concurrent streams each. That costs about $11.20 per hour at $0.70 per GPU-hour.

What latency should a production voice agent target?

Aim for under 500 ms voice-to-voice; human turn-taking sits around 200–300 ms. Managed stacks typically land at 850–1,300 ms P95. Self-hosted pipelines with speculative solving reach 220–380 ms.

Does semantic caching work for voice agents?

Yes, for repetitive use cases. Up to 40% of turns in support and scheduling calls are repeats. Cache hits respond in under 30 ms and skip LLM and TTS costs entirely.

Assumptions behind these numbers

  • Volume: 1 million billable minutes per month, provisioned for a 1,000-call concurrency peak.
  • LLM usage: ~150 output tokens per call minute.
  • TTS usage: 750–900 characters per spoken minute (~800 used for per-minute TTS figures).
  • GPU pricing: NVIDIA L4 at ~$0.70/hour (Google Cloud g2-standard-4, us-east4 on-demand).
  • Pricing date: vendor list prices checked October 10, 2026, excluding limited-time promotions. Prices change often; check each vendor before budgeting.
  • Excluded: engineering salaries, monitoring and observability tooling, and embedding-model cost for the semantic cache.

Sources

Bearworks Parallel Dialer Software

Turn Calls into Deals
‍
with Bearworks

Ready to get started? Book a demo:






Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.