What this is. The one sequence every session follows to put a speaking person on a page. Each step says what it is, the tool that does it, the check that proves it, and what it costs. Written 01-09-2026 on Frank's instruction, after the same knowledge sat scattered across four skills, a criteria list and three topics.
The two rules above all the steps. The words and the voice are approved BEFORE any picture exists. And a section is born whole: one section, one clip, from the locked portrait plus its approved audio. Ordering separate takes and stitching them is what produced jumping transitions, drifting size and more than $120 in one month.
The Speakers page fixes the vocabulary and what is possible today: the two kinds (still · moving), the four sizes (face · half bust · bust · full body). Sizes are always asked in those words, never in pixels.
Three roads: take one that exists (Valentina, Gulnura, Zalia, Grace) · invent a new one · a real person. For a real person you FILM their movement; you never ask an engine to invent it (four generated masters of Frank were rejected, 05-08-2026). Wizard for a colleague:
make-your-speaker.
ONE green portrait, approved by Frank, never re-invented between batches. A wrong face invalidates a whole batch (it did, on the first ABC round). Keep it with the person's other born-once assets, never in a temporary folder.
Before any video. Either a library or cloned voice (ElevenLabs, cents a line) or an engine's own native voice. When there is a choice, Frank picks blind: he chose the native voice over three clones. Acronyms go in a pronunciation dictionary so they can never regress.
The script comes from the page's own text, cut one section per page block. Breath groups of about 25 words, pace 1.9 to 2.1 words a second. Frank approves the text, then it is frozen. A word changed after this point is a re-render, not a purchase.
The whole script as sound, in the approved voice, for cents. This is where rhythm and pronunciation are judged, by ear, before a single picture is paid for.
One clip per section, from the locked portrait plus that audio: ![]()
nacer.py. No fifteen-second cap, no seam inside a section, the same face throughout. Before renting: python3 nacer.py --sweep, and let the tool check the balance and the ceiling. Its two-step proof clip runs first, so a fault costs minutes.
The section's length chooses the route, and this is not a preference. The free Grok session births at most 15 seconds, so any section longer than that comes back as several takes to be joined, and joined takes show their seams: measured 01-09-2026 on the live ABC home page, one 35-second file carried seven visible discontinuities, the worst jumping four and a half times the clip's own movement. So: a section under 15 seconds can be born whole on the free route; a section longer than 15 seconds can only be born whole on the machine, from the portrait plus its approved audio. Promising "one clip, no seams" on the free route for a long section is a promise that cannot be kept.
For a speaker whose approved voice IS an engine's native voice (Valentina), the takes come from the free Grok web session (
grok-speaker). Her voice is not lost on the machine route: the audio of her native takes is what drives it, which is exactly how the clips Frank accepted on 28-08 were made.
la-vara.py (in 3 - Projects/abc-colombia/valentina/): face size, headroom, shoulder gap, alpha, mouth opening, rhythm, posture at the joins. Match the loudness: the machine returns the voice about 14 dB down. A clip that fails is re-born, never patched.
A take must END AT REST, and that is judged on the last frames, not on the last second of sound. The engine holds the mouth open through trailing silence, so a silence-to-rest measure on the audio passes a take that finishes mid-vowel. Seen on the ABC note, 01-09-2026: the first part ends with the mouth wide open and the brows raised, the next opens on a closed smile, and the cut between them is what a visitor calls a bad transition. Check the closing frames for a closed, calm mouth; when several sections are joined, that check is what keeps the joins invisible.
Chromakey sampling the green the engine rendered, not the portrait's green. Two twins: webm (VP9 alpha) and mov (HEVC alpha, or Safari shows nothing). Never ProRes, it is forty times the size. Prove the alpha by forced decode: ffprobe lies about it.
site-speaker: two stacked video elements swapping only once the next clip has painted (no white flash), an idle clip that does not silently mouth words, one card, pause and stop, the page travelling with her.
la-vara-pagina.py on the live bytes, at three widths, never a local copy. Drive a real Chrome, not a bundled Chromium: the bundled build carries no H.264 by licence, so it cannot decode our clips and its playback checks skip in silence while the page still reports PASA. That hole hid a broken pause on all three ABC pages for weeks (Valentina seat, 01-09-2026). A check that cannot run must fail, never pass. Then the phone: solid inside (alpha ≥ 250), the right file types served, Low Power Mode, the block fixed to the four sides rather than measured in screen units. A clip that passes every measure can still be broken by the page.
His eye and his ear are the acceptance test, and blind comparison is how a choice is settled. Publishing to a draft while testing is standing-authorised; publishing to the live page needs his word, every time.
A number becomes a rule only when it is measured, and an unmeasured number is marked as a guess in the same sentence. On 01-09-2026 two seats spent days apart because this session wrote "A100 only, 250 GB of disk" into the shared tool from a single crash it had not diagnosed: the crash was container memory, not the card, and the disk was a guess that billed for nothing. The other seat measured a cheap card working and the three pages falling from about $200 to about $60.
So: every session touching a speaker reads
the board before assuming, and writes its findings there the moment it has them. Each speaker folder carries a CLAUDE.md that says so, and it loads by itself.
| Step | Cost |
|---|---|
| Portrait | cents (image model) |
| Voice, per line | about 2 cents (ElevenLabs) |
| Audio of a whole script | cents |
| Pictures, free route (Grok web session) | $0 marginal, inside the $30 monthly subscription |
| Pictures, machine route | $33 to $37 a published minute as measured, which Frank re-opened on 31-08: see the brief |
| Pictures, buying takes from an engine | about $5 a minute, three times more with the retakes |
| Conversation with visitors | $0.08 a conversation minute (ElevenLabs convai) |
Door: 5 - Handovers/topics/speakers.md · the 45 measured verdicts:
criteria.md · the tool and its seven fixes:
nacer · the always-loaded rule: ~/.claude/CLAUDE.md §Visual output · memory feedback-speaker-sections-born-whole.