Studio › People · 02-08-2026 · method: site-speaker skill

The speaker

A person in the page who presents it in a real voice.

How to get one

Ask Claude: "Put a speaker on this page."

Then say four things:

  • still or moving
  • the size — face · half bust · bust · full body
  • the person — an existing speaker (Valentina, Zalia, Grace) · a photo you give · or describe the person you want and the image is generated
  • the voice — an existing voice · one picked from the voice library (say the language, the tone, man or woman) · or a clone of a real person, with their consent, from 30 seconds of clean speech

Claude builds it with the site-speaker recipe.

Valentina, the speaker, on the free-zones page
Valentina, live
The two speakers

Still — she speaks but does not move. What Zalia does on the Comores leaflets and Grace in Lesotho: a photo and a recorded voice. Quick and almost free — only the voice costs.

Moving — the new capability: she speaks and her face moves with the words. Valentina. One master clip (1.20 EUR, once) plus centimes per part.

A still speaker: the round photo that speaks without moving Still speaker Speaks, does not move. Zalia · Grace · the AZFA circle.
The moving speaker: Valentina, her face moves as she speaks Moving speaker Speaks and moves. Valentina, live.
The sizes (for the moving speaker)

How much of the person the page shows. Every size speaks.

Face To the bottom of the neck.
Half bust Head and shoulders only.
Bust To the chest — Valentina today. Live
Valentina full body, head to feet Full body To the feet. The image is ready; the speaking version not yet.
Buttons and features
Presenting — works today
Her greeting card: Hola, soy Valentina, Sí escuchar After the click: two icons, pause and stop

Her card offers one button. After the click she presents the page part by part, and two icons remain: pause, stop.

Talking — not ready yet

A real conversation with the visitor. We do not know how to do it well yet; its button stays off the page until we do. Where we stand, exactly: Interaction.

Interaction — talking with the speaker

How it works. The visitor asks, typed or spoken; an agent finds the answer in a small set of fact files we wrote, and speaks it back in the speaker's own voice. She answers only from those files — nothing invented, internal notes can never leak. This grounding is called RAG; our method is the voice-agent skill.

Zalia — live

The reference. Text and voice, in French, on the Comores SARL page. 5 fact files + a conduct charter; tested 9 for 9, trick questions included.

Valentina — ready, closed

Her agent, voice, pronunciation and 13 fact files are all wired; the buttons were withdrawn 02-08-2026. No visitor can talk with her yet.

Grace — parked

Narration only, by verdict of 19-07-2026. Her legally verified leaflets wait as the knowledge; wake her when a public assistant is wanted.

What is missing, for Valentina. A full voice test on the live page (a page opened from disk cannot use the microphone) · cutting her off mid-answer, never verified · the greeting is long in voice mode · the face does not answer — during a conversation the bust steps aside and only the voice replies, which is the real meaning of "not ready" · switching text to voice loses the dialogue · cost per minute unmeasured and the agent has no domain lock. Full detail: the diagnosis.

TrackNext taskEffortIt unlocks
1Prove Valentina's conversationRun the full voice round trip on the live page, fix what fails, restore the two buttons for Frank's verdict~1 dayThe first speaker a visitor can talk with
2Guard every speaker's answersWrite Valentina's witness questions; give the existing guard a schedule that runsHoursAnswers that stay right as facts change
3One-click knowledge refreshBuild the fact editor (edit, press Publish); reuse for every speaker1–2 daysColleagues keep the knowledge current
4The face that answersTest real-time mouth movement on a rented GPU against live audioDays, exploratoryThe speaker herself visibly answering
5Wake GraceWhen wanted: run the method on her leaflet corpus~1 dayLesotho citizens ask instead of read
6The sovereign betA bet, not a task: NVIDIA's self-hosted voice model — natural interruption, under a third of a second — English-only today; we asked about SpanishWatchConversation on our own machines, at centimes
How one is made
  1. Portrait on green — one photo, head and shoulders, so she can be cut out.
  2. Voice — reuse or clone; acronyms spelled as they must sound, locked in a dictionary.
  3. Script in parts — the page's argument, 6 to 10 pieces of 15 to 30 seconds.
  4. One paid master clip — about 1.20 EUR, once per person; the living eyes.
  5. The factory — repaints only the mouth per recording, under a centime a part. Drawn, step by step.
  6. Cut-out in pairs — one file for Chrome, one for Safari, always both.
  7. Page wiring — idle breath, card, icons, the slow glide; then the verification walk.
What it costs
PieceWhenCost
Master clipOnce per personAbout 1.20 EUR
Each spoken partPer segmentUnder a centime (rented GPU)
Voice recordingPer scriptElevenLabs minutes, or free with Voicebox
Where the material lives
WhatFile
The whole recipe (the playbook)site-speaker skill
Picture → talking video → cutoutphoto-avatar-video skill
The talking feature, when readyvoice-agent skill · Voice
Frank himself explaining a documentfrank-avatar skill
The recipes and the rulesMake pages move · criteria
The factorythe factory, drawn · musetalk.md
Valentina's page, clips, recordingsvalentina-video.html · busto/ · narracion/
Hola, soy Valentina
Shall I present this page?