Studio › The guide · 06-08-2026 · method: voice-agent skill · sibling pages: Voice · Speakers

The guide

A person in the page you can talk with. You ask, typed or spoken; she answers in her own voice, only from facts we wrote. Named by Frank, 06-08-2026.

The guide answers; the speaker presents

Two different people, often the same face. The speaker performs a prepared script: press play, she presents the page, you listen. The guide holds a conversation: you ask anything in her domain, she finds the answer in her fact files and speaks it back. Zalia is both on the same site: she narrates the Comores leaflets as a speaker, and answers questions on the SARL conditions page as a guide.

The speaker

One-way, prepared

A recorded presentation, perfect diction, zero risk: she can only say what we recorded. Her page: Speakers.

The guide

Two-way, live

A real dialogue. She listens, answers, can be cut off. Riskier and harder, which is why she answers only from her fact files, under a charter, watched by a guard.

Six bricks make a guide

The brain is ours; the mouth is a choice. Bricks 1 to 4 are the brain: they are engine-independent, proven on Zalia, and reused whatever voice technology carries the conversation. Bricks 5 and 6 are the mouth and ears: the part the tools verdict decides.

1
Fact filesShort files, one subject each, every number traced to its source, labeled with WHO may see it. What is not in her files cannot be said. Zalia's architecture.
2
The charterHer conduct, compiled into her instructions: what she may say, what she never says, when she sends the visitor to a human. She informs, she never decides a personal case.
3
The compilerSends only the right audience's files to the agent. Internal notes can never leak: the choice is made at compile time, not trusted to the agent.
4
The guardA robot that re-asks witness questions and trick questions on a schedule, and flags any wrong, leaked or drifted answer. rag-gardien skill (retired).
5
The voiceHer cloned voice with a pronunciation dictionary, so names are always said right. The clone belongs to an engine; the recording sample is the portable asset.
6
The ears and the floorThe real-time loop: hearing the visitor, knowing when they finished, being interruptible. Today rented from ElevenLabs; the one brick we do not own.
Three guides exist; one is live
Zalia · Comores · live

The reference. Text and voice, in French, on the SARL page. 5 fact files + charter, tested 9 for 9, trick questions included. Her editor is live at zalia-admin.

Valentina · AZFA · wired, closed

Agent, voice, pronunciation and 13 fact files all built; her talk buttons were withdrawn 02-08-2026 because the conversation was never proven live. The diagnosis.

Grace · Lesotho · parked

Narration only, by verdict of 19-07-2026. Her legally verified leaflets wait as ready knowledge; wake her when a public assistant is wanted. The topic.

A real conversation can be interrupted, and it flows

The bar, set by Frank 06-08-2026. Two qualities decide whether talking with a guide feels like talking with a person:

Quality 1

You can cut her off

Mid-sentence, by speaking, the way you would with a person. She stops at once and listens. The trade calls this barge-in; the strongest form is full duplex: she listens even while she speaks.

Quality 2

The conversation flows

She answers fast enough that silence never becomes awkward, and she knows when you have finished speaking without you pressing anything. Measured as voice-to-voice delay: a person answers in about 0.2 seconds.

Where we stand, measured. Zalia today, typed question to first spoken syllable: 1.8 to 3.3 seconds (the four laws, measured 26-07-2026). Interrupting by typing a new question works and cuts her cleanly; interrupting by VOICE, mid-answer, has never been verified on any of our guides. That verification, and the choice of the fastest engine, is what the tools verdict below decides.

The tools, judged on 06-08-2026: keep ElevenLabs, tune her new interruption controls

The sweep. Two researchers went through vendor documentation and repositories on 06-08-2026, one on hosted platforms, one on open self-hosted stacks. Full evidence, every claim sourced: hosted report · open report.

Hosted: ElevenLabs keeps the crown

Still the only platform with all four of our needs: cloned voice, grounding, French, and talking from a static page with no server. And since June it shipped exactly what Frank asks for: a better turn model (v3, now default), interruption modes, ignore-words so "oui" and "d'accord" do not cut her off, a filter against background voices, and filler sounds while she thinks. Our agents use none of these yet.

The newcomer: Grok Voice

Real and serious: cloned voices, French, grounding, about $0.05 to 0.08 a minute. But its ear is a simple volume threshold, a generation behind ElevenLabs and OpenAI at knowing when you finished speaking, and it needs a server of ours for the page to connect. OpenAI is the fluidity reference but still refuses cloned voices; Gemini is 10 times cheaper but preview, no clone.

Our own machine: real, close, not yet French full-duplex

Kyutai runs a proven half-second conversation loop in French on one small GPU, and our June plan survives with two upgrades: Qwen3-TTS (French, 3-second cloning, streaming) and turn detection now bundled free in Pipecat. All of it fits one rented card at ~$0.34 an hour. The models that truly listen while speaking answer in half a second and can be cut off like a person, but none speaks French yet; our NVIDIA seat watches that door.

WayCan be cut offFlowsHer cloned voiceFrenchFrom a static pagePrice
ElevenLabs (today)Yes, best controls of the marketSub-second claimed; 1.8–3.3 s measured on Zalia before tuningYesYesYes~$0.08–0.12/min
OpenAI Realtime 2.1Yes, semantic earNative speech-to-speech, fastNoYesNo, needs our server~$0.05–0.46/min
Grok Voice 2.0Basic, volume threshold onlyNo published figuresYesYesNo, needs our server$0.05–0.08/min
Gemini LiveYes, nativeFast, native audioNoYesNo, needs our server~$0.02/min, cheapest
Our machine: Kyutai loopYes, semantic ear~0.5 s proven in productionNot her exact clone yet (their voices, or swap the mouth)YesNo, needs our server~$0.34/h flat
Our machine: Pipecat cascade (June plan, refreshed)Yes, turn detection bundledSub-second achievable, to measureYes, 3 s cloning (Qwen3-TTS)YesNo, needs our server~$0.34/h flat
Full-duplex models (listen while speaking)Like a person~0.45 s documentedNoNo, English onlyNoResearch licenses

Frank's direction, 06-08-2026. The table above is the paper ranking; the ear decides. First audition the systems side by side in French and hear the quality: how each one takes a cut-off, how the conversation flows. Then imitate the winning quality at the lowest possible cost (the per-minute engines against the rented machine at a flat ~$0.34 an hour). The bench is track 1 below; the day a French full-duplex model lands, it joins the bench the same week.

What we improve next

The tracks, ranked. Adapted from the interaction diagnosis of 02-08-2026, reordered for the two qualities above.

TrackNext taskEffortIt unlocks
1The audition: hear the quality side by sideTalk in French with each candidate on its own demo: Zalia with her new controls set (turn model v3, interruption mode, ignore-words, background-voice filter), OpenAI Realtime, Gemini Live, Grok Voice, Kyutai's public loop. Cut each one off mid-answer; score by earHoursThe quality bar, heard, not read
2The grounded bench, then the cheapest imitationFor the audition's finalists: same fact files, same ten questions, live test pages and clips on the Desktop; measure delay and cost per minute; then reproduce the winning quality at the lowest cost (tuning, or the rented-machine cascade: Pipecat + Kyutai ears + Qwen3-TTS mouth)1–2 daysThe winner, and the cheapest way to its quality
3Guard every guideWrite Valentina's witness questions; give the existing guard a schedule that runsHoursAnswers that stay right as facts change
4One-click knowledge refreshGeneralize Zalia's editor (edit a fact, press Publish) to every guide1–2 daysColleagues keep the knowledge current
5The face that answersReal-time mouth movement against live audio, on a rented GPUDays, exploratoryThe guide herself visibly answering
What it costs
WayWhat it isCost
ElevenLabs, hostedWhat Zalia uses today: real-time loop, cloned voice, no server of ours.Per minute of conversation; ~148 EUR at 2,000 min/month
Our own machine, rentedThe whole loop on a rented graphics computer (RunPod, billed by the second), per the sovereign plan.~0.34 to 1 USD per hour of machine, flat whatever the traffic

The standing rule. Before renewing any per-minute service, the open path on the rented machine is checked and both costs are said.

Where the material lives
WhatFile
The whole method (the playbook)voice-agent skill
Zalia's architecture (fact files, audiences, compiler)zalia-rag-architecture.md
The guardrag-gardien skill (retired)
The self-hosted plan, with costsagent-vocal-souverain-plan.md
Where conversation stood on 02-08-2026the interaction diagnosis
The NVIDIA seat (full duplex, early access)nemotron-voicechat.md
The Comores stream (Zalia's home)comores topic