So a person can listen instead of read, and ask instead of search. Three shapes, all working today — and the recipe is installable.
Recorded segments. The reader presses play, the voice explains, the page scrolls and opens the right part by itself.
The visitor asks a question, by voice or by typing, and a grounded voice answers from a small knowledge base, never from the open internet.
The same recorded voice drives a photo, so a person appears on the page and speaks. This crosses into People.
Cloning. Thirty seconds of clean speech is enough. Hosted clones live at ElevenLabs and keep their quality across languages; the free local alternative is Voicebox on the Mac, which costs nothing and never leaves the machine.
Saying names right. Every voice needs a small pronunciation list. Acronyms are the trap: written plainly, a voice reads AZFA as a word. Spell it in the script the way it must sound ("a zeta efe a"). This is a rule, not a detail: Frank asked for that correction three times before it stuck.
The sound. A fixed treatment gives every recording the same body and presence, so segments recorded weeks apart still sound like one person.
| Way | What it is | Cost |
|---|---|---|
| ElevenLabs, hosted | Cloned voice, recorded narration and the live guide. What Valentina and Zalia use today. | Subscription plus minutes |
| Voicebox, on the Mac | Cloning and narration, offline, nothing sent anywhere. | Free |
| Our own machine | The whole live guide hosted by us on a rented graphics computer (RunPod, ~0.34 USD/h, billed by the second). NVIDIA's Nemotron early access belongs here: it answers in under a third of a second and can be interrupted like a person — English only today, so a bet, not a tool. | About 0.34 USD an hour of machine |
In the pack, installable now (npx skills add gfrankgva/studio-skills): voice-agent — the whole method: facts labeled by audience, a conduct charter, a guard that replays witness questions. Also packed: photo-avatar-video (a photo that speaks) and site-speaker (the full speaker). All six: Install our skills.
Where conversation stands today, speaker by speaker (Zalia live and tested · Valentina wired, closed · Grace parked), with the ranked improvement tracks: the interaction diagnosis. The speaker herself — sizes, buttons, factory: People.