| Pipecat 1.7 + Kyutai STT-1B + vLLM + Qwen3-TTS |
Barge-in via bundled Smart Turn v3.1 (12 ms, 23 langs); no full-duplex |
Sub-second achievable (STT 0.5 s delay + LLM + TTS 97 ms first packet) |
✔ |
✔ 3 s zero-shot |
1×4090 ($0.34/h) |
Production — all parts shipped, permissive licenses |
| Kyutai Unmute (as-is) |
Barge-in via semantic VAD |
~450 ms (doc, their prod) |
✔ |
✖ voice repo only (swap TTS to fix) |
1×16 GB (A40 comfortable) |
Production — powers unmute.sh |
| LiveKit Agents self-hosted |
Barge-in; best turn model is cloud-tied, local v1-mini is weaker |
Comparable to Pipecat |
✔ |
Depends on TTS plugged |
1×4090 |
Production |
| Nemotron blueprint v2 |
Barge-in + Smart Turn; full-duplex only via early-access VoiceChat |
"Sub-second" (claim) |
✔ (Magpie multilingual) |
✖ removed from open weights; NIM container only |
≥72 GB or 2×40 GB |
Production but heavy, NVIDIA-channel dependent |
| NemotronLabs VoiceChat-11B |
TRUE full-duplex, ~450 ms |
~450 ms (doc) |
✖ English only |
✖ |
A100/H100-class |
Research only (license says so) |
| Moshi / PersonaPlex-7B |
TRUE full-duplex, best open dynamics |
~200 ms frame-level |
✖ English only |
Voice conditioning (PersonaPlex) |
1×24 GB |
Research/demo |
| Qwen3-Omni-30B |
Streaming S2S, natural turn-taking, not full-duplex |
Low (no doc figure) |
✔ speech out |
✖ 3 preset voices |
~80 GB |
Usable, heavy |