# Open / self-hosted real-time voice agents — state of the art, 06-08-2026 *Researched 06-08-2026 for [the Guide page](../index.html) (studio). Scope: interruption (barge-in / full-duplex) + fluidity (voice-to-voice latency), French, voice cloning, single rented GPU. Documented figures marked (doc); vendor claims marked (claim). Companion report: [hosted-platforms-06-08-2026.md](hosted-platforms-06-08-2026.md).* --- ## 1. NVIDIA Nemotron voice agent blueprint - **Current release: v2.0.0, 9 July 2026** (v1.0.0 was 3 March 2026). Re-architecture with multiple pipelines: generic-assistant, **multilingual-assistant**, omni-assistant (Nemotron 3 Omni replaces ASR+LLM), frontend-backend. Pipecat under the hood, upgraded to 1.3.0; **Smart Turn detection replaced VAD-silence** for turn-taking. [Releases](https://github.com/NVIDIA-AI-Blueprints/nemotron-voice-agent/releases) - **French: yes, via the multilingual pipeline** — Parakeet 1.1B RNNT multilingual ASR + Magpie TTS Multilingual, fixed language per session. [Repo](https://github.com/NVIDIA-AI-Blueprints/nemotron-voice-agent) - **Magpie TTS Multilingual 357M**: 12 languages incl. fr-FR; latest v2607 (21 July 2026); NVIDIA Open Model License, commercial use OK. **Zero-shot cloning was REMOVED from this open release "for security reasons"** (doc, model card). Cloning survives only in the separate **Magpie TTS Zeroshot NIM** container (3–10 s reference clip, on-prem NIM — NVIDIA's enterprise container channel, not plain open weights). [Model card](https://huggingface.co/nvidia/magpie_tts_multilingual_357m) · [NIM voice-cloning docs](https://docs.nvidia.com/nim/speech/latest/tts/voice-cloning.html) - **Latency**: "sub-second E2E" (claim, no measured figure in repo). - **GPU (doc)**: workstation = **1 GPU ≥ 72 GB or 2 GPUs ≥ 40 GB** (2×A40 48GB ≈ $0.88/h on RunPod would qualify; one 4090 does not); DGX Spark 128 GB unified; "cloud CPU-only" mode calls NVIDIA-hosted NIMs (not sovereign). License of the blueprint code: BSD-2-Clause. - **Nemotron 3 VoiceChat (12B full-duplex)**: **still early-access only** — apply-for-access form, demo on build.nvidia.com. Frank's seat: approved 27-07-2026 (see memory `reference-nemotron-voicechat.md`). [Artificial Analysis](https://artificialanalysis.ai/articles/nemotron-3-voicechat-leader-speech-pareto) - **NEW — NemotronLabs VoiceChat-11B, open weights released 3 Aug 2026**: end-to-end full-duplex S2S, ~450 ms turn-taking (doc, model card), yields instantly on barge-in. **English only. "Ready for research purposes only."** OpenMDW v1.1 license; validated GPUs A100/H100/H200/B100/B200/RTX-6000 (no consumer card listed). [Model card](https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B) ## 2. Kyutai (Moshi · Unmute · DSM) - **STT `stt-1b-en_fr`**: ~1B, **English+French, 0.5 s delay, built-in semantic VAD** — the turn-taking signal. CC-BY-4.0 weights. Rust server: H100 serves 400 real-time streams (doc); trivially fits a 4090 slice. [Repo](https://github.com/kyutai-labs/delayed-streams-modeling) · [kyutai.org/stt](https://kyutai.org/stt/) - **TTS 1.6B** (July 2025): en+fr, true streaming (starts before full text). **Cloning still restricted** — the voice-embedding model is not released; you pick from the curated/donated voice repository (Voice Donation Project ran to Feb 2026). [kyutai.org/tts](https://kyutai.org/tts/) - **Pocket TTS (13 Jan 2026) — the policy reversal**: 100M params, **CPU-only, ~200 ms to first audio (doc), 6 languages incl. French, voice cloning from any WAV locally** (`export-voice`), MIT code + open weights. [Repo](https://github.com/kyutai-labs/pocket-tts) - **Unmute** (their production cascade, powers unmute.sh): STT-1B + any OpenAI-compatible LLM + TTS 1.6B, **~450 ms response in production (doc)**, self-host from **one 16 GB GPU** (STT 2.5 + TTS 5.3 + LLM 6.1 GB), Docker Compose to Swarm. [Unmute docs](https://kyutai-labs-unmute.mintlify.app/introduction) - **Moshi**: full-duplex, 7B, **English only**, CC-BY-4.0 — now mostly a base for forks (PersonaPlex). Production-usable in Aug 2026: the STT/TTS/Unmute pieces (en+fr); Moshi itself stays research. ## 3. Pipecat - **Latest: v1.7.0, 1 Aug 2026** (v1.4.0 was 17 June 2026). [PyPI](https://pypi.org/project/pipecat-ai/) · [Releases](https://github.com/pipecat-ai/pipecat/releases) - **Smart Turn v3 → v3.1**: semantic end-of-turn on raw audio, 8 MB, **23 languages incl. French, 12 ms CPU / 3–7 ms GPU (doc)**, BSD-2 with open training data; **v3.1 (drop-in) improves accuracy**. It is now the **default turn-stop strategy, weights bundled with pipecat-ai** — no config needed. [Daily blog v3](https://www.daily.co/blog/announcing-smart-turn-v3-with-cpu-inference-in-just-12ms/) · [v3.1](https://www.daily.co/blog/improved-accuracy-in-smart-turn-v3-1/) · [HF](https://huggingface.co/pipecat-ai/smart-turn-v3) - **v1.7.0 gem: `PocketTTSService`** — Kyutai pocket-tts local, CPU-only, **French + voice cloning**, native in the framework the June plan already chose. - Other since 1.4: TTFA (time-to-first-audio) metrics (1.5), Media-over-QUIC transport + DeepgramFlux (1.6), STT usage metrics (1.7). - **Full-duplex**: no local full-duplex model plugin; `realtime_service_mode` (1.4) optimizes hosted speech-to-speech services (OpenAI Realtime, Gemini Live). Barge-in remains the classic VAD + Smart Turn interrupt — mature, not model-native. ## 4. LiveKit Agents - **Turn detection moved to a unified AUDIO end-of-turn model** (`livekit.agents.inference.TurnDetector`), replacing the text-based English/Multilingual models (deprecated, removed in 2.0). **v1 = highest accuracy but served on LiveKit Cloud inference only; v1-mini = runs locally on CPU, free** — the self-hosted option. 14 languages incl. **French**. [Docs](https://docs.livekit.io/agents/build/turns/turn-detector) · [Turn detector plugin](https://github.com/livekit/agents/tree/main/livekit-plugins/livekit-plugins-turn-detector) - Legacy multilingual model: ONNX on CPU, <500 MB RAM, fine-tuned from Qwen2.5-0.5B (doc). [HF](https://huggingface.co/livekit/turn-detector) · [Blog](https://livekit.com/blog/improved-end-of-turn-model-cuts-voice-ai-interruptions-39) - Barge-in: Silero VAD triggers interruption; semantic model commits the turn — same architecture class as Pipecat. Fully self-hosted possible with v1-mini, but the **best detector is cloud-tied**, a sovereignty minus vs Pipecat's bundled Smart Turn. ## 5. Open full-duplex / speech-to-speech models (beyond Moshi & Nemotron) - **PersonaPlex-7B (NVIDIA, Feb 2026)** — Moshi-based full-duplex with role prompt + voice conditioning; **best open conversational dynamics (91.0% per Artificial Analysis)**; English only; code MIT, weights NVIDIA Open Model License (gated); ~1×24 GB feasible. [Repo](https://github.com/NVIDIA/personaplex) · [HF](https://huggingface.co/nvidia/personaplex-7b-v1) · [Paper](https://arxiv.org/pdf/2602.06053) - **FLM-Audio 7B (CofeAI)** — native full-duplex ("natural monologue" dual training), **English+Chinese**, Apache-2.0, open weights, ~1×24 GB. [Repo](https://github.com/cofe-ai/flm-audio) · [Paper](https://arxiv.org/abs/2509.02521) - **Freeze-Omni (Tencent)** — streaming S2S over a frozen LLM, zh/en, open code+weights, strongest speech reasoning of the open duplex set (33.9%). [Repo](https://github.com/Tencent/Freeze-Omni) - **Qwen3-Omni-30B-A3B (Alibaba)** — Apache-2.0, streaming S2S with natural turn-taking (not literal listen-while-speaking); **speech OUT in 10 languages incl. French**; 3 preset voices, no cloning; **~79 GB BF16 (doc)** → 80 GB-class GPU. Closest thing to "French speech-to-speech open model". [Repo](https://github.com/QwenLM/Qwen3-Omni) - **DuplexOmni** (June 2026) — true full-duplex built on Qwen3-Omni; **paper stage, weights not confirmed**. [arXiv](https://arxiv.org/html/2606.09186v1) - Meta: nothing open full-duplex. Mistral: Voxtral = audio understanding/STT only. Fish Audio: TTS only. - **The headline: no open full-duplex model speaks French today.** Full-duplex in French remains proprietary (or wait for DuplexOmni-class releases). ## 6. Streaming TTS — French + zero-shot cloning - **Qwen3-TTS (22 Jan 2026) — the new XTTS-v2 successor.** Apache-2.0, 0.6B/1.7B, **10 languages incl. French, 3 s cloning (needs transcript of the reference), true streaming, e2e latency "as low as 97 ms" (claim)**; vLLM day-0 (offline mode so far). Easily 1×4090. [Repo](https://github.com/QwenLM/Qwen3-TTS) · [Simon Willison](https://simonwillison.net/2026/Jan/22/qwen3-tts/) - **Chatterbox Multilingual v3 (Resemble)** — MIT, 21–25 languages incl. French, zero-shot cloning, PerTh watermark baked in; Turbo ~75 ms (claim, vendor); 65.3%-preferred-over-ElevenLabs study is **vendor-run**. 1×4090. [HF](https://huggingface.co/ResembleAI/chatterbox) · [Resemble](https://www.resemble.ai/learn/models/chatterbox-multilingual) - **CosyVoice 3 / Fun-CosyVoice3-0.5B (Dec 2025)** — Apache-2.0, 9 languages incl. French, zero-shot cloning, chunk-aware streaming (Qwen2.5-0.5B backbone). 1×4090. [HF](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512) - **Kyutai TTS 1.6B** — en/fr, best-in-class streaming, but cloning restricted (voice repo only). **Pocket TTS** — fr + free local cloning, CPU, ~200 ms, lower fidelity (100M). - **XTTS-v2**: frozen upstream (Coqui defunct, idiap fork = maintenance only) — superseded on every axis by Qwen3-TTS/CosyVoice3. **IndexTTS-2**: zh/en only, no French. --- ## Decision table | Stack | Interruption / full-duplex | Voice-to-voice latency | French | Cloning | GPU | Maturity | |---|---|---|---|---|---|---| | **Pipecat 1.7 + Kyutai STT-1B + vLLM + Qwen3-TTS** | Barge-in via bundled Smart Turn v3.1 (12 ms, 23 langs); no full-duplex | Sub-second achievable (STT 0.5 s delay + LLM + TTS 97 ms first packet) | ✔ | ✔ 3 s zero-shot | **1×4090** ($0.34/h) | Production — all parts shipped, permissive licenses | | **Kyutai Unmute (as-is)** | Barge-in via semantic VAD | **~450 ms (doc, their prod)** | ✔ | ✖ voice repo only (swap TTS to fix) | 1×16 GB (A40 comfortable) | Production — powers unmute.sh | | **LiveKit Agents self-hosted** | Barge-in; best turn model is cloud-tied, local v1-mini is weaker | Comparable to Pipecat | ✔ | Depends on TTS plugged | 1×4090 | Production | | **Nemotron blueprint v2** | Barge-in + Smart Turn; full-duplex only via early-access VoiceChat | "Sub-second" (claim) | ✔ (Magpie multilingual) | ✖ removed from open weights; NIM container only | **≥72 GB or 2×40 GB** | Production but heavy, NVIDIA-channel dependent | | **NemotronLabs VoiceChat-11B** | **TRUE full-duplex, ~450 ms** | ~450 ms (doc) | ✖ English only | ✖ | A100/H100-class | Research only (license says so) | | **Moshi / PersonaPlex-7B** | TRUE full-duplex, best open dynamics | ~200 ms frame-level | ✖ English only | Voice conditioning (PersonaPlex) | 1×24 GB | Research/demo | | **Qwen3-Omni-30B** | Streaming S2S, natural turn-taking, not full-duplex | Low (no doc figure) | ✔ speech out | ✖ 3 preset voices | ~80 GB | Usable, heavy | **Bottom line.** Full-duplex exists open (Moshi, PersonaPlex, VoiceChat-11B) but **none speaks French** and the best is research-licensed. For French + cloning today, the June plan's shape survives with two upgrades: **Qwen3-TTS replaces XTTS-v2** (Apache-2.0, streaming, 3 s cloning, French) and **Smart Turn v3.1 comes free inside Pipecat 1.7** for barge-in; Kyutai STT-1B gives French streaming ASR with semantic VAD. All of it on one RTX 4090; the RAG LLM size is the only thing that could push to the A40/L40S. Kyutai's Unmute proves the 450 ms cascade in production; NVIDIA's blueprint is the polished-but-heavy alternative (72 GB, cloning stripped from open weights). Watch for a French full-duplex release (DuplexOmni-class, or Nemotron 3 VoiceChat leaving early access) — that is the next genuine jump.