← Studio library

Open voice agents, the state of the art

Reference notes. Source retained; historical claims may need rechecking.

Studio reference illustration

Open / self-hosted real-time voice agents — state of the art, 06-08-2026

Researched 06-08-2026 for Related Studio referencethe Guide page (studio). Scope: interruption (barge-in / full-duplex) + fluidity (voice-to-voice latency), French, voice cloning, single rented GPU. Documented figures marked (doc); vendor claims marked (claim). Companion report: Related Studio referencehosted-platforms-06-08-2026.md.


1. NVIDIA Nemotron voice agent blueprint

2. Kyutai (Moshi · Unmute · DSM)

3. Pipecat

4. LiveKit Agents

5. Open full-duplex / speech-to-speech models (beyond Moshi & Nemotron)

6. Streaming TTS — French + zero-shot cloning


Decision table

Stack Interruption / full-duplex Voice-to-voice latency French Cloning GPU Maturity
Pipecat 1.7 + Kyutai STT-1B + vLLM + Qwen3-TTS Barge-in via bundled Smart Turn v3.1 (12 ms, 23 langs); no full-duplex Sub-second achievable (STT 0.5 s delay + LLM + TTS 97 ms first packet) ✔ ✔ 3 s zero-shot 1×4090 ($0.34/h) Production — all parts shipped, permissive licenses
Kyutai Unmute (as-is) Barge-in via semantic VAD ~450 ms (doc, their prod) ✔ ✖ voice repo only (swap TTS to fix) 1×16 GB (A40 comfortable) Production — powers unmute.sh
LiveKit Agents self-hosted Barge-in; best turn model is cloud-tied, local v1-mini is weaker Comparable to Pipecat ✔ Depends on TTS plugged 1×4090 Production
Nemotron blueprint v2 Barge-in + Smart Turn; full-duplex only via early-access VoiceChat "Sub-second" (claim) ✔ (Magpie multilingual) ✖ removed from open weights; NIM container only ≥72 GB or 2×40 GB Production but heavy, NVIDIA-channel dependent
NemotronLabs VoiceChat-11B TRUE full-duplex, ~450 ms ~450 ms (doc) ✖ English only ✖ A100/H100-class Research only (license says so)
Moshi / PersonaPlex-7B TRUE full-duplex, best open dynamics ~200 ms frame-level ✖ English only Voice conditioning (PersonaPlex) 1×24 GB Research/demo
Qwen3-Omni-30B Streaming S2S, natural turn-taking, not full-duplex Low (no doc figure) ✔ speech out ✖ 3 preset voices ~80 GB Usable, heavy

Bottom line. Full-duplex exists open (Moshi, PersonaPlex, VoiceChat-11B) but none speaks French and the best is research-licensed. For French + cloning today, the June plan's shape survives with two upgrades: Qwen3-TTS replaces XTTS-v2 (Apache-2.0, streaming, 3 s cloning, French) and Smart Turn v3.1 comes free inside Pipecat 1.7 for barge-in; Kyutai STT-1B gives French streaming ASR with semantic VAD. All of it on one RTX 4090; the RAG LLM size is the only thing that could push to the A40/L40S. Kyutai's Unmute proves the 450 ms cascade in production; NVIDIA's blueprint is the polished-but-heavy alternative (72 GB, cloning stripped from open weights). Watch for a French full-duplex release (DuplexOmni-class, or Nemotron 3 VoiceChat leaving early access) — that is the next genuine jump.

Source previewDownload original source