Studio › Voice · 31-07-2026 · method: voice-agent skill

Give a page a voice

So a person can listen instead of read, and ask instead of search. Three shapes, all working today — and the recipe is installable.

1

A page that reads itself

Recorded segments. The reader presses play, the voice explains, the page scrolls and opens the right part by itself.

Live: Valentina on the free zones page · the Comores leaflets narrated by Zalia · the Lesotho governance explainer.
2

A guide that answers

The visitor asks a question, by voice or by typing, and a grounded voice answers from a small knowledge base, never from the open internet.

Live: Valentina (free zones), Zalia (Comores), Frank's own guide on the Lesotho capital and shares change.
3

A voice that moves a face

The same recorded voice drives a photo, so a person appears on the page and speaks. This crosses into People.

Live: Valentina's talking bust, eight segments, no frame around her.

Make a voice

Cloning. Thirty seconds of clean speech is enough. Hosted clones live at ElevenLabs and keep their quality across languages; the free local alternative is Voicebox on the Mac, which costs nothing and never leaves the machine.

Saying names right. Every voice needs a small pronunciation list. Acronyms are the trap: written plainly, a voice reads AZFA as a word. Spell it in the script the way it must sound ("a zeta efe a"). This is a rule, not a detail: Frank asked for that correction three times before it stuck.

The sound. A fixed treatment gives every recording the same body and presence, so segments recorded weeks apart still sound like one person.

Four traps, already paid for

  1. Typing is the door, not the microphone. A page opened as a file cannot capture a microphone at all. Typed question, spoken answer works everywhere; the microphone needs the page to be served.
  2. A clone is lost on the wrong model. Clones are trained per voice model. On the wrong one the person sounds like a stranger. Check before delivering.
  3. A public guide needs no key in the page. It connects without secrets, so the page can be handed to anyone, but anyone can also use its minutes: reserve it for documents given to a few people, or lock it to one site.
  4. Never let it speak first. A voice that starts by itself when a page opens ambushes the reader. The reader presses play.

What it costs

WayWhat it isCost
ElevenLabs, hostedCloned voice, recorded narration and the live guide. What Valentina and Zalia use today.Subscription plus minutes
Voicebox, on the MacCloning and narration, offline, nothing sent anywhere.Free
Our own machineThe whole live guide hosted by us on a rented graphics computer (RunPod, ~0.34 USD/h, billed by the second). NVIDIA's Nemotron early access belongs here: it answers in under a third of a second and can be interrupted like a person — English only today, so a bet, not a tool.About 0.34 USD an hour of machine

Take the recipe

In the pack, installable now (npx skills add gfrankgva/studio-skills): voice-agent — the whole method: facts labeled by audience, a conduct charter, a guard that replays witness questions. Also packed: photo-avatar-video (a photo that speaks) and site-speaker (the full speaker). All six: Install our skills.

Where conversation stands today, speaker by speaker (Zalia live and tested · Valentina wired, closed · Grace parked), with the ranked improvement tracks: the interaction diagnosis. The speaker herself — sizes, buttons, factory: People.