A person in the page you can talk with. You ask, typed or spoken; she answers in her own voice, only from facts we wrote. Named by Frank, 06-08-2026.
Two different people, often the same face. The speaker performs a prepared script: press play, she presents the page, you listen. The guide holds a conversation: you ask anything in her domain, she finds the answer in her fact files and speaks it back. Zalia is both on the same site: she narrates the Comores leaflets as a speaker, and answers questions on the SARL conditions page as a guide.
A recorded presentation, perfect diction, zero risk: she can only say what we recorded. Her page: Speakers.
A real dialogue. She listens, answers, can be cut off. Riskier and harder, which is why she answers only from her fact files, under a charter, watched by a guard.
The brain is ours; the mouth is a choice. Bricks 1 to 4 are the brain: they are engine-independent, proven on Zalia, and reused whatever voice technology carries the conversation. Bricks 5 and 6 are the mouth and ears: the part the tools verdict decides.
The reference. Text and voice, in French, on the SARL page. 5 fact files + charter, tested 9 for 9, trick questions included. Her editor is live at zalia-admin.
Agent, voice, pronunciation and 13 fact files all built; her talk buttons were withdrawn 02-08-2026 because the conversation was never proven live. The diagnosis.
Narration only, by verdict of 19-07-2026. Her legally verified leaflets wait as ready knowledge; wake her when a public assistant is wanted. The topic.
The bar, set by Frank 06-08-2026. Two qualities decide whether talking with a guide feels like talking with a person:
Mid-sentence, by speaking, the way you would with a person. She stops at once and listens. The trade calls this barge-in; the strongest form is full duplex: she listens even while she speaks.
She answers fast enough that silence never becomes awkward, and she knows when you have finished speaking without you pressing anything. Measured as voice-to-voice delay: a person answers in about 0.2 seconds.
Where we stand, measured. Zalia today, typed question to first spoken syllable: 1.8 to 3.3 seconds (the four laws, measured 26-07-2026). Interrupting by typing a new question works and cuts her cleanly; interrupting by VOICE, mid-answer, has never been verified on any of our guides. That verification, and the choice of the fastest engine, is what the tools verdict below decides.
The sweep. Two researchers went through vendor documentation and repositories on 06-08-2026, one on hosted platforms, one on open self-hosted stacks. Full evidence, every claim sourced: hosted report · open report.
Still the only platform with all four of our needs: cloned voice, grounding, French, and talking from a static page with no server. And since June it shipped exactly what Frank asks for: a better turn model (v3, now default), interruption modes, ignore-words so "oui" and "d'accord" do not cut her off, a filter against background voices, and filler sounds while she thinks. Our agents use none of these yet.
Real and serious: cloned voices, French, grounding, about $0.05 to 0.08 a minute. But its ear is a simple volume threshold, a generation behind ElevenLabs and OpenAI at knowing when you finished speaking, and it needs a server of ours for the page to connect. OpenAI is the fluidity reference but still refuses cloned voices; Gemini is 10 times cheaper but preview, no clone.
Kyutai runs a proven half-second conversation loop in French on one small GPU, and our June plan survives with two upgrades: Qwen3-TTS (French, 3-second cloning, streaming) and turn detection now bundled free in Pipecat. All of it fits one rented card at ~$0.34 an hour. The models that truly listen while speaking answer in half a second and can be cut off like a person, but none speaks French yet; our NVIDIA seat watches that door.
| Way | Can be cut off | Flows | Her cloned voice | French | From a static page | Price |
|---|---|---|---|---|---|---|
| ElevenLabs (today) | Yes, best controls of the market | Sub-second claimed; 1.8–3.3 s measured on Zalia before tuning | Yes | Yes | Yes | ~$0.08–0.12/min |
| OpenAI Realtime 2.1 | Yes, semantic ear | Native speech-to-speech, fast | No | Yes | No, needs our server | ~$0.05–0.46/min |
| Grok Voice 2.0 | Basic, volume threshold only | No published figures | Yes | Yes | No, needs our server | $0.05–0.08/min |
| Gemini Live | Yes, native | Fast, native audio | No | Yes | No, needs our server | ~$0.02/min, cheapest |
| Our machine: Kyutai loop | Yes, semantic ear | ~0.5 s proven in production | Not her exact clone yet (their voices, or swap the mouth) | Yes | No, needs our server | ~$0.34/h flat |
| Our machine: Pipecat cascade (June plan, refreshed) | Yes, turn detection bundled | Sub-second achievable, to measure | Yes, 3 s cloning (Qwen3-TTS) | Yes | No, needs our server | ~$0.34/h flat |
| Full-duplex models (listen while speaking) | Like a person | ~0.45 s documented | No | No, English only | No | Research licenses |
Frank's direction, 06-08-2026. The table above is the paper ranking; the ear decides. First audition the systems side by side in French and hear the quality: how each one takes a cut-off, how the conversation flows. Then imitate the winning quality at the lowest possible cost (the per-minute engines against the rented machine at a flat ~$0.34 an hour). The bench is track 1 below; the day a French full-duplex model lands, it joins the bench the same week.
The tracks, ranked. Adapted from the interaction diagnosis of 02-08-2026, reordered for the two qualities above.
| Track | Next task | Effort | It unlocks | |
|---|---|---|---|---|
| 1 | The audition: hear the quality side by side | Talk in French with each candidate on its own demo: Zalia with her new controls set (turn model v3, interruption mode, ignore-words, background-voice filter), OpenAI Realtime, Gemini Live, Grok Voice, Kyutai's public loop. Cut each one off mid-answer; score by ear | Hours | The quality bar, heard, not read |
| 2 | The grounded bench, then the cheapest imitation | For the audition's finalists: same fact files, same ten questions, live test pages and clips on the Desktop; measure delay and cost per minute; then reproduce the winning quality at the lowest cost (tuning, or the rented-machine cascade: Pipecat + Kyutai ears + Qwen3-TTS mouth) | 1–2 days | The winner, and the cheapest way to its quality |
| 3 | Guard every guide | Write Valentina's witness questions; give the existing guard a schedule that runs | Hours | Answers that stay right as facts change |
| 4 | One-click knowledge refresh | Generalize Zalia's editor (edit a fact, press Publish) to every guide | 1–2 days | Colleagues keep the knowledge current |
| 5 | The face that answers | Real-time mouth movement against live audio, on a rented GPU | Days, exploratory | The guide herself visibly answering |
| Way | What it is | Cost |
|---|---|---|
| ElevenLabs, hosted | What Zalia uses today: real-time loop, cloned voice, no server of ours. | Per minute of conversation; ~148 EUR at 2,000 min/month |
| Our own machine, rented | The whole loop on a rented graphics computer (RunPod, billed by the second), per the sovereign plan. | ~0.34 to 1 USD per hour of machine, flat whatever the traffic |
The standing rule. Before renewing any per-minute service, the open path on the rented machine is checked and both costs are said.
| What | File |
|---|---|
| The whole method (the playbook) | voice-agent skill |
| Zalia's architecture (fact files, audiences, compiler) | zalia-rag-architecture.md |
| The guard | rag-gardien skill (retired) |
| The self-hosted plan, with costs | agent-vocal-souverain-plan.md |
| Where conversation stood on 02-08-2026 | the interaction diagnosis |
| The NVIDIA seat (full duplex, early access) | nemotron-voicechat.md |
| The Comores stream (Zalia's home) | comores topic |