Reference previewStudio › Speakers › how one is made · 01-09-2026

How a speaker is made, A to Z

What this is. The one sequence every session follows to put a speaking person on a page. Each step says what it is, the tool that does it, the check that proves it, and what it costs. Written 01-09-2026 on Frank's instruction, after the same knowledge sat scattered across four skills, a criteria list and three topics.

The two rules above all the steps. The words and the voice are approved BEFORE any picture exists. And a section is born whole: one section, one clip, from the locked portrait plus its approved audio. Ordering separate takes and stitching them is what produced jumping transitions, drifting size and more than $120 in one month.

0. Open the studio first

Reference previewThe Speakers page fixes the vocabulary and what is possible today: the two kinds (still · moving), the four sizes (face · half bust · bust · full body). Sizes are always asked in those words, never in pixels.

1. Who speaks

Three roads: take one that exists (Valentina, Gulnura, Zalia, Grace) · invent a new one · a real person. For a real person you FILM their movement; you never ask an engine to invent it (four generated masters of Frank were rejected, 05-08-2026). Wizard for a colleague: Reference previewmake-your-speaker.

2. Lock the face

ONE green portrait, approved by Frank, never re-invented between batches. A wrong face invalidates a whole batch (it did, on the first ABC round). Keep it with the person's other born-once assets, never in a temporary folder.

3. Choose the voice, and approve it

Before any video. Either a library or cloned voice (ElevenLabs, cents a line) or an engine's own native voice. When there is a choice, Frank picks blind: he chose the native voice over three clones. Acronyms go in a pronunciation dictionary so they can never regress.

4. Freeze the words

The script comes from the page's own text, cut one section per page block. Breath groups of about 25 words, pace 1.9 to 2.1 words a second. Frank approves the text, then it is frozen. A word changed after this point is a re-render, not a purchase.

5. Render the audio, and listen

The whole script as sound, in the approved voice, for cents. This is where rhythm and pronunciation are judged, by ear, before a single picture is paid for.

6. Birth the pictures

One clip per section, from the locked portrait plus that audio: Reference previewnacer.py. No fifteen-second cap, no seam inside a section, the same face throughout. Before renting: python3 nacer.py --sweep, and let the tool check the balance and the ceiling. Its two-step proof clip runs first, so a fault costs minutes.

The section's length chooses the route, and this is not a preference. The free Grok session births at most 15 seconds, so any section longer than that comes back as several takes to be joined, and joined takes show their seams: measured 01-09-2026 on the live ABC home page, one 35-second file carried seven visible discontinuities, the worst jumping four and a half times the clip's own movement. So: a section under 15 seconds can be born whole on the free route; a section longer than 15 seconds can only be born whole on the machine, from the portrait plus its approved audio. Promising "one clip, no seams" on the free route for a long section is a promise that cannot be kept.

For a speaker whose approved voice IS an engine's native voice (Valentina), the takes come from the free Grok web session (Reference previewgrok-speaker). Her voice is not lost on the machine route: the audio of her native takes is what drives it, which is exactly how the clips Frank accepted on 28-08 were made.

7. Measure every clip

la-vara.py (in 3 - Projects/abc-colombia/valentina/): face size, headroom, shoulder gap, alpha, mouth opening, rhythm, posture at the joins. Match the loudness: the machine returns the voice about 14 dB down. A clip that fails is re-born, never patched.

A take must END AT REST, and that is judged on the last frames, not on the last second of sound. The engine holds the mouth open through trailing silence, so a silence-to-rest measure on the audio passes a take that finishes mid-vowel. Seen on the ABC note, 01-09-2026: the first part ends with the mouth wide open and the brows raised, the next opens on a closed smile, and the cut between them is what a visitor calls a bad transition. Check the closing frames for a closed, calm mouth; when several sections are joined, that check is what keeps the joins invisible.

8. Cut her out

Chromakey sampling the green the engine rendered, not the portrait's green. Two twins: webm (VP9 alpha) and mov (HEVC alpha, or Safari shows nothing). Never ProRes, it is forty times the size. Prove the alpha by forced decode: ffprobe lies about it.

9. Wire her into the page

Reference previewsite-speaker: two stacked video elements swapping only once the next clip has painted (no white flash), an idle clip that does not silently mouth words, one card, pause and stop, the page travelling with her.

10. Prove it where it plays

la-vara-pagina.py on the live bytes, at three widths, never a local copy. Drive a real Chrome, not a bundled Chromium: the bundled build carries no H.264 by licence, so it cannot decode our clips and its playback checks skip in silence while the page still reports PASA. That hole hid a broken pause on all three ABC pages for weeks (Valentina seat, 01-09-2026). A check that cannot run must fail, never pass. Then the phone: solid inside (alpha ≥ 250), the right file types served, Low Power Mode, the block fixed to the four sides rather than measured in screen units. A clip that passes every measure can still be broken by the page.

11. Frank decides

His eye and his ear are the acceptance test, and blind comparison is how a choice is settled. Publishing to a draft while testing is standing-authorised; publishing to the live page needs his word, every time.

The rule that keeps sessions from diverging

A number becomes a rule only when it is measured, and an unmeasured number is marked as a guess in the same sentence. On 01-09-2026 two seats spent days apart because this session wrote "A100 only, 250 GB of disk" into the shared tool from a single crash it had not diagnosed: the crash was container memory, not the card, and the disk was a guess that billed for nothing. The other seat measured a cheap card working and the three pages falling from about $200 to about $60.

So: every session touching a speaker reads Reference previewthe board before assuming, and writes its findings there the moment it has them. Each speaker folder carries a CLAUDE.md that says so, and it loads by itself.

What each step costs

StepCost
Portraitcents (image model)
Voice, per lineabout 2 cents (ElevenLabs)
Audio of a whole scriptcents
Pictures, free route (Grok web session)$0 marginal, inside the $30 monthly subscription
Pictures, machine route$33 to $37 a published minute as measured, which Frank re-opened on 31-08: see the brief
Pictures, buying takes from an engineabout $5 a minute, three times more with the retakes
Conversation with visitors$0.08 a conversation minute (ElevenLabs convai)

Where the rest is written

Door: 5 - Handovers/topics/speakers.md · the 45 measured verdicts: Reference previewcriteria.md · the tool and its seven fixes: Reference previewnacer · the always-loaded rule: ~/.claude/CLAUDE.md §Visual output · memory feedback-speaker-sections-born-whole.