# Hosted conversational voice agents — state of the art, August 2026 *Researched 06-08-2026 for [the Guide page](../index.html) (studio). Prime criteria: **interruption (barge-in)** and **fluidity (voice-to-voice latency, turn-taking)**, for a RAG-grounded, cloned-French-voice guide embedded on a static government web page. Facts labeled: **[docs]** = official documentation, **[vendor]** = official marketing/blog claim, **[3rd-party]** = external measurement. Companion report: [open-stacks-06-08-2026.md](open-stacks-06-08-2026.md).* --- ## 1. ElevenLabs Agents (Conversational AI) — the incumbent ### Interruption / barge-in - Detection is **server-side**, not client VAD: a proprietary turn-taking model + VAD decides barge-in; the client receives events and stops playback. The server streams a `vad_score` event (probability 0–1 the user is speaking) and, on interruption, an `agent_response_correction` event with the truncated transcript. **[docs]** [Client events](https://elevenlabs.io/docs/eleven-agents/customization/events/client-events) - Interruptions are enabled by selecting `interruption` as a client event in the agent's Advanced tab; disable for content that must be delivered in full. **[docs]** [Conversation flow](https://elevenlabs.io/docs/eleven-agents/customization/conversation-flow) - **New (6 July 2026):** `disable_interruptions` replaced by `interruption_mode` with three states: `allow` · `disable_during_tool` · `disable_during_tool_and_turn`, with per-tool overrides. **[docs]** [Changelog 6-07-2026](https://elevenlabs.io/docs/changelog/2026/7/6) - **New (22 June 2026):** `interruption_ignore_terms` (array of strings) — suppress turn detection on specific phrases (backchannels like "oui", "d'accord" won't cut the agent off). **[docs]** [Changelog 22-06-2026](https://elevenlabs.io/docs/changelog/2026/6/22) - **New (3 Aug 2026):** `vad` config gains `background_voice_detection` — filters background speakers so they don't trigger false barge-ins. **[docs]** [Changelog 3-08-2026](https://elevenlabs.io/docs/changelog/2026/8/3) ### Turn-taking - `turn_eagerness`: `eager` | `normal` | `patient`. `turn_timeout`: 1–30 s. `soft_timeout_config`: filler audio ("Hmm… yes") when the LLM is slow — timeout 0.5–8.0 s, optional LLM-generated filler, up to 7 extra messages, `randomize_fillers`, `max_soft_timeouts_per_generation` 1–8. **[docs]** [Conversation flow](https://elevenlabs.io/docs/eleven-agents/customization/conversation-flow), [Changelog 8-06](https://elevenlabs.io/docs/changelog/2026/6/8), [22-06](https://elevenlabs.io/docs/changelog/2026/6/22) - **New (8 June 2026):** selectable turn model — `turn_model`: `turn_v2` | `turn_v3`, **default `turn_v3`**, per agent or per workflow node. **[docs]** [Changelog 8-06-2026](https://elevenlabs.io/docs/changelog/2026/6/8) - The turn system combines VAD with prosody/semantic signals ("read the meaning of what's being said to intuit when a turn has finished"); "speculative turn-taking" is in production. **[vendor]** [Interaction models blog](https://elevenlabs.io/blog/interaction-models) ### Latency - Component figures: Flash v2.5 TTS **~75 ms model latency** (explicitly excludes app + network; 32 languages incl. French) **[docs]** [Models](https://elevenlabs.io/docs/overview/models); Scribe v2 Realtime ASR **~150 ms** (now the default ASR provider, replacing `elevenlabs`, since 8-06-2026) **[docs]** [Models](https://elevenlabs.io/docs/overview/models), [Changelog 8-06](https://elevenlabs.io/docs/changelog/2026/6/8); orchestration overhead "**<100 ms**" **[vendor]** [Orchestration engine blog](https://elevenlabs.io/blog/unpacking-elevenagents-orchestration-engine). - End-to-end: "sub-second pipeline" **[vendor, no measured figure published]** [Interaction models](https://elevenlabs.io/blog/interaction-models). Independent benchmark of the TTS alone: ~255 ms median TTFB with network. **[3rd-party]** [Deepgram analysis](https://deepgram.com/learn/is-elevenlabs-real-time-what-developers-need-to-know) ### Connection - JS SDK: voice conversations default to **WebRTC**, text-only to WebSocket; explicit `connectionType: 'webrtc' | 'websocket'`. WebRTC audio fixed at pcm_48000. Auth: public agents connect with just the agent ID (works from a static page); private agents need a server to mint a signed URL (WS) or conversation token (WebRTC). **[docs]** [JS SDK](https://elevenlabs.io/docs/eleven-agents/libraries/java-script), [WebRTC token](https://elevenlabs.io/docs/eleven-agents/api-reference/conversations/get-webrtc-token), [WebRTC blog](https://elevenlabs.io/blog/conversational-ai-webrtc) ### New since June 2026 (beyond the above) - **`eleven_realtime_v1_mini`** is now the default realtime agent model (new realtime model family, 3-08-2026 — not yet on the public models overview page). **[docs]** [Changelog 3-08-2026](https://elevenlabs.io/docs/changelog/2026/8/3) - Agent **reasoning**: `enable_reasoning_summary`, reasoning in transcripts, streamed via `onAgentReasoningResponsePart` (SDK). **[docs]** [7-06](https://elevenlabs.io/docs/changelog/2026/7/6), [3-08](https://elevenlabs.io/docs/changelog/2026/8/3) - New `knowledge_base` **system tool** with `enabled_strategies`; RAG `use_agent_defaults`. **[docs]** [3-08](https://elevenlabs.io/docs/changelog/2026/8/3) - `agent_response_complete` WebSocket event (reliable end-of-turn detection client-side); `pre_tool_speech` (`auto`/`force`/`off`). **[docs]** [Changelog April–June](https://elevenlabs.io/docs/changelog) - Per-conversation `cost_fiat` (USD) + platform charge breakdowns in the API. **[docs]** [7-06](https://elevenlabs.io/docs/changelog/2026/7/6) ### Pricing - All plans: overage **$0.080/min**; burst concurrency **$0.160/min**; **LLM billed separately** on top (deducted from credits). Included minutes: Free 15 · Starter 75 · Creator 275 · Pro 1,238 · Scale 3,738 · Business 12,375. **[docs]** [Agents pricing](https://elevenlabs.io/pricing/agents) - Third-party summaries put all-in cost at ~$0.08–0.12/min depending on LLM. **[3rd-party]** [Cekura](https://www.cekura.ai/blogs/elevenlabs-pricing) --- ## 2. OpenAI Realtime API - **Current model: `gpt-realtime-2.1`** + `gpt-realtime-2.1-mini` (released ~6 July 2026), plus `gpt-realtime-translate` and `gpt-live-transcribe`. True speech-to-speech (no STT→LLM→TTS cascade). **[docs]** [Realtime guide](https://developers.openai.com/api/docs/guides/realtime), [community announcement](https://community.openai.com/t/new-realtime-models-on-the-api-gpt-realtime-2-1-and-gpt-realtime-2-1-mini/1385896) - 2.1 claims: better alphanumeric handling, better silence/noise handling, "**more reliable interruption behavior** when a user speaks over the model", **p95 latency lowered ≥25%** via caching, configurable reasoning effort. **[vendor claims, no absolute ms figures published]** [DataNorth summary](https://datanorth.ai/news/openai-releases-gpt-realtime-2-1-voice-models) - Interruption: `turn_detection` = **server VAD** (threshold/silence) or **`semantic_vad`** — a classifier scoring the probability the user has finished speaking; less likely to cut users off mid-thought. **[docs]** [VAD guide](https://developers.openai.com/api/docs/guides/realtime-vad) - Voices: **built-in voices only — no custom or cloned voices**; nothing in the 2.1 release changes this. **[docs — absence of any cloning feature]** [Realtime docs](https://developers.openai.com/api/docs/guides/realtime). This remains the blocker for the guide's cloned voice. - Browser use: WebRTC for browsers, but a **server must mint ephemeral credentials** (`POST /v1/realtime/client_secrets`) — a static page alone cannot connect safely. **[docs]** [Realtime guide](https://developers.openai.com/api/docs/guides/realtime) - Pricing (official, token-based): 2.1 audio **$32/1M in · $64/1M out** (cached audio in $0.40); mini **$10/1M in · $20/1M out**. **[docs]** [Pricing](https://developers.openai.com/api/docs/pricing). Real-world: **$0.18–0.46/min uncached; $0.05–0.10/min with caching** (measured over 4,000 sessions); mini ≈ $0.016/min base math. **[3rd-party]** [HackerNoon](https://hackernoon.com/openai-realtime-api-pricing-in-2026-real-world-data-from-4000-measured-sessions), [eesel](https://www.eesel.ai/blog/gpt-realtime-mini-pricing) --- ## 3. Google Gemini Live API - **Status: still preview.** Models: `gemini-2.5-flash-native-audio-preview-12-2025` and the newer **`gemini-3.1-flash-live-preview`**; "may change before becoming stable, more restrictive rate limits". **[docs]** [Pricing page](https://ai.google.dev/gemini-api/docs/pricing), [Live overview](https://ai.google.dev/gemini-api/docs/live) - Voices: **prebuilt Google voices only** (Puck, Charon, Kore, Fenrir, Aoede, Leda, Orus, Zephyr) — **no voice cloning**; open feature requests confirm the gap. **[docs + 3rd-party]** [Voice config](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/live-api/configure-language-voice), [adk-docs issue #487](https://github.com/google/adk-docs/issues/487) - Barge-in: native — "users can interrupt the model at any time"; configurable automatic activity detection; affective dialog and proactive audio are its signature features. **[docs]** [Live API](https://ai.google.dev/gemini-api/docs/live), [capabilities](https://ai.google.dev/gemini-api/docs/live-api/capabilities) - 70 languages incl. French; connection is a **stateful WebSocket**; **ephemeral tokens** recommended for client-side, but they are still minted server-side. **[docs]** [Live API](https://ai.google.dev/gemini-api/docs/live) - Pricing: 2.5 native audio **$3.00/1M audio-in · $12.00/1M audio-out**; 3.1 flash live listed at **$0.005/min audio-in · $0.018/min audio-out** — by far the cheapest hosted option (~$0.02/min mixed). **[docs]** [Pricing](https://ai.google.dev/gemini-api/docs/pricing) --- ## 4. New/notable hosted entrants (2026) - **Cartesia Line**: code-first voice-agent platform on Cartesia's SSM models; **cloned voices yes**, sub-90 ms TTS model latency **[vendor]**, barge-in platform-managed; **$0.06/min** reported **[3rd-party]**. [Line launch](https://www.cartesia.ai/blog/introducing-line-for-voice-agents/), [Voximplant press](https://www.globenewswire.com/news-release/2026/02/12/3237440/0/en/Voximplant-Brings-Cartesia-Line-Voice-Agents-into-Real-Calls.html) - **Deepgram Voice Agent API** (GA): single API bundling STT (Nova-3/Flux) + LLM + TTS (Aura-2) with **native barge-in and turn-taking prediction**; bundled **$4.50/hr = $0.075/min [vendor]**; no self-serve arbitrary cloning; **French TTS coverage weak** (Aura-2 English-focused). [Voice Agent API](https://deepgram.com/product/voice-agent-api), [GA post](https://deepgram.com/learn/voice-agent-api-generally-available) - **Retell AI**: telephony-first agent platform (web SDK exists); cloned voices via integrated ElevenLabs/other TTS; **from ~$0.07/min [vendor]**; measured default voice loop ~700 ms **[3rd-party]**. [Pricing analysis](https://www.retellai.com/blog/ai-voice-agent-pricing-full-cost-breakdown-platform-comparison-roi-analysis), [Softcery calculator](https://softcery.com/ai-voice-agents-calculator) - **Vapi**: orchestrator — **$0.05/min platform fee + provider costs**; cloned voices via ElevenLabs/Cartesia plug-ins; web SDK with public key (embeds without own server); measured default latency ~1,450 ms (config-dependent) **[3rd-party]**. [Softcery](https://softcery.com/ai-voice-agents-calculator), [ainora comparison](https://ainora.lt/blog/ai-voice-agent-cost-per-minute-2026) - **Hume EVI**: **EVI 3 (EN/ES only)** and **EVI 4-mini (supports French)**; voice cloning from a sample and voice design from text **[docs]**; **$0.04–0.07/min**, 5 free min/mo **[docs]**. [EVI versions](https://dev.hume.ai/docs/speech-to-speech-evi/configuration/evi-version), [Voices](https://dev.hume.ai/docs/empathic-voice-interface-evi/configuration/voices), [Pricing](https://www.hume.ai/pricing) --- ## 5. xAI Grok Voice - **Public API: yes.** Grok Voice Agent API, **OpenAI Realtime API-compatible** (`wss://api.x.ai/v1/realtime?model=…` — OpenAI SDKs work by swapping the base URL). Models: `grok-voice-think-fast-1.0`, flagship **`grok-voice-think-fast-2.0`**; `grok-voice-latest` alias moves to 2.0 on **5 Aug 2026**. **[docs]** [Voice Agent API](https://docs.x.ai/developers/model-capabilities/audio/voice-agent), [LiteLLM provider page](https://docs.litellm.ai/docs/providers/xai_realtime) - **Cloned voices: yes** — "Clone any voice from a short reference clip with the Custom Voices API"; 80+ built-in voices. **[docs]** [Voice Agent API](https://docs.x.ai/developers/model-capabilities/audio/voice-agent) - **French: yes** — 20+ languages "with native-quality accents", automatic language detection and mid-conversation switching. **[docs]** [Voice overview](https://docs.x.ai/developers/model-capabilities/audio/voice) - Interruption: **server VAD only, threshold-based** — `turn_detection.threshold` 0.1–0.9 (default 0.85), `silence_duration_ms` 0–10000, or manual. No semantic/prosodic turn model documented — a configuration generation behind ElevenLabs turn_v3 and OpenAI semantic_vad. **[docs]** [Voice Agent API](https://docs.x.ai/developers/model-capabilities/audio/voice-agent) - Latency: "low-latency" **[vendor, no published figures]**. - RAG: yes — `file_search` over uploaded document collections (vector stores), plus web search and X search server-side tools. **[docs]** [Voice Agent API](https://docs.x.ai/developers/model-capabilities/audio/voice-agent) - Static page without server: **no** — ephemeral tokens recommended for browsers but minted server-side with the API key. **[docs]** [Voice Agent API](https://docs.x.ai/developers/model-capabilities/audio/voice-agent) - Pricing: **$0.05/min** (1.0) · **$0.08/min** (2.0), voices included; tools billed per call (web/X search $5/1k, document search $2.50/1k); 100 concurrent sessions, 30-min max session. **[3rd-party reporting of official rates]** [eesel](https://www.eesel.ai/blog/grok-voice-agent-builder-pricing), [aipricing.guru](https://www.aipricing.guru/xai-pricing/). Also: no-code **Voice Agent Builder** beta since 1 July 2026 (phone-oriented). **[3rd-party]** [eesel](https://www.eesel.ai/blog/grok-voice-agent-builder) --- ## Comparison table | Platform | Interruption quality | Voice-to-voice latency | Cloned voice | French | Web embed, no server | Price/min | |---|---|---|---|---|---|---| | **ElevenLabs Agents** | Best-configured: dedicated turn model (turn_v3), eagerness, `interruption_mode`, ignore-terms, background-voice filter | Components documented (75 ms TTS + 150 ms ASR + <100 ms orch.); "sub-second" e2e [vendor] | **Yes** (incl. professional clones) | Yes (32-lang Flash v2.5) | **Yes** (public agent widget/ID) | $0.08 overage + LLM (~$0.08–0.12 all-in) | | **OpenAI Realtime (gpt-realtime-2.1)** | Good: server VAD + semantic VAD; "more reliable" in 2.1 [vendor] | Native s2s; p95 −25% in 2.1 [vendor]; no official ms | **No** (fixed voices) | Yes | No (server mints token) | Token-based; measured $0.05–0.46/min; mini ~$0.02 | | **Gemini Live (3.1 flash live, preview)** | Good native barge-in, auto VAD; still **preview** | Native audio, low; no official ms | **No** (8 prebuilt voices) | Yes (70 langs) | No (ephemeral token via server) | ~$0.005 in + $0.018 out /min — cheapest | | **xAI Grok Voice (think-fast-2.0)** | Basic: threshold server VAD only | "Low-latency" [vendor]; no figures | **Yes** (Custom Voices API) | Yes (20+ langs) | No (ephemeral token via server) | $0.05 (1.0) / $0.08 (2.0) + tools | | **Cartesia Line** | Platform-managed barge-in | Sub-90 ms TTS model [vendor]; no e2e figure | Yes | Yes (Sonic multilingual) | Telephony-first; web unclear | ~$0.06 [3rd-party] | | **Deepgram Voice Agent** | Native barge-in + turn prediction [vendor] | Fast STT stack; no official e2e | No (stock Aura-2) | **Weak** (Aura-2 EN-focused) | No | $0.075 bundled | | **Retell** | Provider-dependent | ~700 ms measured [3rd-party] | Via ElevenLabs voices | Yes | No (server creates web call) | from ~$0.07 + costs | | **Vapi** | Provider-dependent, configurable | ~1,450 ms default measured [3rd-party] | Via ElevenLabs/Cartesia | Yes | Yes (public-key web SDK) | $0.05 + providers | | **Hume EVI 4-mini** | Good (empathic turn model) [vendor] | No official figures | **Yes** (clone + design) | Yes (EVI 4-mini only) | No (token via server) | $0.04–0.07 | ## Reading for the guide decision - **ElevenLabs remains the only platform combining all four requirements**: cloned voice + built-in RAG + French + serverless static-page embed. Its interruption stack is also the most configurable (turn_v3, interruption_mode, ignore-terms, background-voice filtering — all shipped June–Aug 2026). - **The only credible challenger on the two prime qualities** is OpenAI gpt-realtime-2.1 (native s2s fluidity, semantic VAD) — but it still has **no cloned voices** and needs a token server. - **xAI Grok Voice is the notable 2026 newcomer**: cloned voices + RAG + French + OpenAI-compatible API at $0.05–0.08/min, but its turn-taking is plain threshold VAD and it needs a server for browser tokens. - **Gemini Live is the price disruptor** (~10× cheaper) but is still in preview and has no cloned voices.