← Studio library

Hosted voice agents, the state of the art

Reference notes. Source retained; historical claims may need rechecking.

Studio reference illustration

Hosted conversational voice agents — state of the art, August 2026

Researched 06-08-2026 for Related Studio referencethe Guide page (studio). Prime criteria: interruption (barge-in) and fluidity (voice-to-voice latency, turn-taking), for a RAG-grounded, cloned-French-voice guide embedded on a static government web page. Facts labeled: [docs] = official documentation, [vendor] = official marketing/blog claim, [3rd-party] = external measurement. Companion report: Related Studio referenceopen-stacks-06-08-2026.md.


1. ElevenLabs Agents (Conversational AI) — the incumbent

Interruption / barge-in

Turn-taking

Latency

Connection

New since June 2026 (beyond the above)

Pricing


2. OpenAI Realtime API


3. Google Gemini Live API


4. New/notable hosted entrants (2026)


5. xAI Grok Voice


Comparison table

Platform Interruption quality Voice-to-voice latency Cloned voice French Web embed, no server Price/min
ElevenLabs Agents Best-configured: dedicated turn model (turn_v3), eagerness, interruption_mode, ignore-terms, background-voice filter Components documented (75 ms TTS + 150 ms ASR + <100 ms orch.); "sub-second" e2e [vendor] Yes (incl. professional clones) Yes (32-lang Flash v2.5) Yes (public agent widget/ID) $0.08 overage + LLM (~$0.08–0.12 all-in)
OpenAI Realtime (gpt-realtime-2.1) Good: server VAD + semantic VAD; "more reliable" in 2.1 [vendor] Native s2s; p95 −25% in 2.1 [vendor]; no official ms No (fixed voices) Yes No (server mints token) Token-based; measured $0.05–0.46/min; mini ~$0.02
Gemini Live (3.1 flash live, preview) Good native barge-in, auto VAD; still preview Native audio, low; no official ms No (8 prebuilt voices) Yes (70 langs) No (ephemeral token via server) ~$0.005 in + $0.018 out /min — cheapest
xAI Grok Voice (think-fast-2.0) Basic: threshold server VAD only "Low-latency" [vendor]; no figures Yes (Custom Voices API) Yes (20+ langs) No (ephemeral token via server) $0.05 (1.0) / $0.08 (2.0) + tools
Cartesia Line Platform-managed barge-in Sub-90 ms TTS model [vendor]; no e2e figure Yes Yes (Sonic multilingual) Telephony-first; web unclear ~$0.06 [3rd-party]
Deepgram Voice Agent Native barge-in + turn prediction [vendor] Fast STT stack; no official e2e No (stock Aura-2) Weak (Aura-2 EN-focused) No $0.075 bundled
Retell Provider-dependent ~700 ms measured [3rd-party] Via ElevenLabs voices Yes No (server creates web call) from ~$0.07 + costs
Vapi Provider-dependent, configurable ~1,450 ms default measured [3rd-party] Via ElevenLabs/Cartesia Yes Yes (public-key web SDK) $0.05 + providers
Hume EVI 4-mini Good (empathic turn model) [vendor] No official figures Yes (clone + design) Yes (EVI 4-mini only) No (token via server) $0.04–0.07

Reading for the guide decision

Source previewDownload original source