Reference previewStudio › Speakers · 28-08-2026 · method: site-speaker skill

The speaker

Choose the production route explicitly. The current portrait-and-approved-audio route and the older master-clip repaint route have different inputs. Master-clip rules below apply only to legacy repainting, not to every new speaker. A speaking video is not a conversational agent.

Speaking presenter exampleCompare presenter, animation and lip-sync tools

A person in the page who presents it in a real voice. She is standing on this page right now — press her card to hear her.

Making one? The whole sequence is on one page: Reference previewhow a speaker is made, A to Z, eleven steps, each with its tool, its check and its cost.

How a section is born, decided 28-08-2026

One section, one clip. A page section is born whole from a locked portrait plus its approved audio, on a rented machine. The words and the voice are settled before any picture exists, so there is no fifteen-second cap, no seam between takes, and a changed word is a re-render, not a purchase.

How Frank decided. Three clips of the same welcome line, judged blind: her current take against two born on the machine. His verdict, « the 3 are good », « good voice and good image ». He also heard, before any measurement, that the machine clips were quieter: the engine returns the voice about 14 dB down.

The old way, and why it stopped. Ordering separate takes and stitching them is what produced jumping transitions, drifting size and more than $120 in one month. A session that reaches for it should know it was measured and left behind.

The tool. Reference previewnacer.py, portrait plus one recording per section: it rents the machine, births the clips, levels the voice and gives the machine back. Two numbers to carry: the mouth reads 18 to 19 % of the face on the machine against 13.9 % native, both accepted; and the house mouth band was read on native clips only, so it reads high here.

What it costs. About $33 to $37 for one published minute of video: the machine spends roughly 23 minutes of its own time per second of speech, and it rents for $1.39 to $1.59 an hour (Reference previewrunpod.io/pricing, checked 31-08-2026). Slower and dearer than the old per-take method — no seams costs more than seams. Full sum: What it costs.

What it does not fix. The machine paints the face onto the voice it is given. A recording that runs slow, runs fast, or says a word wrong stays wrong however well the face renders. Approve the voice first.

To put a speaker in your page: give your Claude the recipe, then ask

First, once per machine. Your Claude does not know this recipe until you give it to it. The recipe is a small file called site-speaker. Install it with one command in your Claude Code:

npx skills add gfrankgva/studio-skills

Or take the file yourself: Reference previewdownload site-speaker and put it in ~/.claude/skills/site-speaker/SKILL.md. More ways, and what else is in the pack: further down this page.

Then just ask, in your own words: « Put a speaker on this page. » Say two more things — still or moving, and which size — and it builds it with the site-speaker recipe.

Before you make her speak, legacy master-clip rules.

1. Legacy route only: repaint the master clip, never a photograph. The mouth engine takes --reuse-base <master.mp4>. Given --photo instead, it invents the whole face and the result is flat, soft-mouthed and out of sync — refused on sight, every time. Everything but the mouth must be copied frame for frame from the clip you paid for.

2. A mouth painted on a still has never passed. MuseTalk was refused 06-08-2026 (« the mouth is closed, she doesn't look natural »), LatentSync on a still refused 10-08-2026 (« pas synchro avec les paroles »). What passes is the native clip, where the voice and the mouth are born together, or a repaint on that clip.

Which speaker do you want?

This is the first decision, and the only one that changes the work. Everything below — still or moving, the sizes, what she can do, what it costs — is the same whichever you pick.

Once you have chosen, tell Claude Code “put a speaker on this page” and say two more things: still or moving, and which size. It builds it with the site-speaker recipe.

The month of takes, and the six rules it left

Where it comes from. Valentina and Gulnura between 31-07 and 28-08-2026: thirty handovers, more than $120 of engine time, and one night that put a speaker on a live site with the wrong words in her mouth. Everything below was paid for once. The night, written up · the money, counted.

One cause under three faults. Jumping transitions, a rhythm that changes, a face that drifts in size: three seams, one reason. Every take is born as its own scene, capped at fifteen seconds, then stitched to its neighbours. On Gulnura's evening, 42 clips bought became 17 takes and 7 published parts: 520 seconds paid for 170 seconds on the page.

The route, settled. Reverse the order. Approve the voice first, cheaply, then birth each page section as ONE clip from the locked portrait plus that audio. No fifteen second cap, no seam inside a section, the same face throughout. Frank judged Reference previewthree versions of one real segment blind on 28-08-2026, a native take against two born on a rented machine from the same approved audio, and kept the machine route for new speakers: « the 3 are good », « good voice and good image ». Valentina stays on her native voice.

The number that comes with it. The machine hands the voice back about 14 decibels under where a page expects it, its own attenuation and not the recording: measured at −8.5 against −22.9. Frank heard it before anything was measured. Anything driving that engine levels the sound afterwards, which Reference previewnacer.py does in one pass. The door for all of it is the speakers topic.

Rule 1. A band, never a ceiling. A take is accepted only between 1.85 and 2.10 words a second. The old rule rejected anything above 2.3, so the slow takes were redone and the brisk ones survived, and one visit mixed 1.75 with 2.17. The ear hears that seam.

Rule 2. When the rule changes, every take is reborn. Never only the failing ones. Sameness across a page beats saving a few dollars on the takes that scraped through.

Rule 3. Four layers, on the served page. What she says, transcribed from the clips the live site serves, equals the caption, equals the approved pack, and the pack matches the page sections. One command checks all four before and after every publish, so a colleague editing the page flags the affected parts within minutes. Her words passed every script check that night and still trailed the page, because no check tied the script to the screen. Working example: walk-served.py.

Rule 4. Freeze the text before recording. Recording starts only on a page whose text its owner has declared frozen. A later edit is not forbidden, it simply re-opens the affected segments and they rejoin the queue.

Rule 5. A ruler that invents a defect costs money. The transcriber is what decides whether a take worth a dollar is kept or bought again, and ours misreads: the same clip returned « Azfa », « AZFA » and « AFFA » on three passes, and heard relaciona as relación a. Calibrate every ruler against work already approved before it judges anything new, and give it the list of names it must not guess.

Rule 6. Read what renders, never the source. Any claim about what a visitor sees comes from the living page. A dead text field in the code produced a false alarm about the captions, and cost a day.

The free road, and when to leave it. Grok's own subscription session births takes at no cost per second: 32 takes for nothing on 21-08. The paid interface is $0.081 a second, and it is what $120 went to when the free road was down all night. Check the road before launching a batch. Recipe: grok-speaker.

The exception on voices, measured 24-08. The house rule says take a library voice rather than clone an invented person. It stops being true when one engine makes the face and the voice together. Same sentence, same image, a clone against five professional voices: Frank chose the engine's own voice three times out of three. It costs $0.65 a line against about two centimes, and it is paid because the ear asks for it, not because there is nothing else.

The two speakers

Still, she speaks but does not move. What Zalia does on the Reference previewComores leaflets and Grace in Reference previewLesotho: a photo and a recorded voice. Quick and almost free, only the voice costs.

Moving: she speaks and her face moves with the words. Valentina. Since 28-08-2026 the standard needs only a locked portrait plus her approved voice, born whole on a rented machine — no master clip: how, and what it costs. A speaker built before that date (a master clip, then a cheap repaint per part) keeps working the old way: Reference previewthat route, step 2.

A still speaker: the round photo that speaks without moving Still speaker Speaks, does not move. Zalia · Grace · the AZFA circle.
The moving speaker: Valentina, her face moves as she speaks Moving speaker Speaks and moves. Valentina, Reference previewlive.
The sizes (for the moving speaker)

How much of the person the page shows. Every size speaks.

Face To the bottom of the neck.
Half bust Head and shoulders only.
Bust To the chest, Valentina today. Reference previewLive
Valentina full body, head to feet Full body To the feet. The image is ready; the speaking version not yet.
How sharp she is, and what that costs in time

The engines that make a face return a small picture, about 640 pixels. That size, not the mouth, is what makes a speaker look soft. Enlarging her four times and putting a film grain back fixes it - but it is slow, and it is the only slow step in the whole build. So it is a choice, and here is the same face under each one. Look first, read after.

Valentina as she comes out of the engine, not enlarged Leave her as she is No waiting. What the AZFA page has shown since July.
Valentina enlarged with the quick network, with grain The quick finish About 35 minutes of computer time per minute of speaker.
Valentina enlarged with the slow network, with grain The slow finish About 13 hours per minute of speaker. Same enlargement, a heavier network inside.

What that means for a real page. Valentina speaks for two minutes and forty-five seconds on the AZFA page. That is under two hours of unattended computer time on the quick finish, and about forty hours on the slow one. Both are free, it is your machine, not a service.

Ask for it by name. Tell Claude Code “use the quick finish” or “use the slow finish”. If you say nothing, it uses the quick one and shows you the two side by side before spending a night on the other.

Measured 05-08-2026 on a Mac, on Valentina's own frames at 640 pixels: 1.5 seconds a frame against 29 to 35. Enlarging is the same ×4 in both; only the network differs. Three things that look like they would help and do not: cutting the frame into tiles (worse), half precision (no change), and a rented GPU (22 seconds a frame, and it cannot hand the file back).

Buttons and features
Presenting, works today
Her greeting card: Hola, soy Valentina, Sí escuchar After the click: two icons, pause and stop

Her card offers one button. After the click she presents the page part by part, and two icons remain: pause, stop.

Talking, not ready yet

A real conversation with the visitor. We do not know how to do it well yet; its button stays off the page until we do. Where we stand, exactly: Interaction.

Interaction, talking with the speaker

How it works. The visitor asks, typed or spoken; an agent finds the answer in a small set of fact files we wrote, and speaks it back in the speaker's own voice. She answers only from those files, nothing invented, internal notes can never leak. This grounding is called RAG; our method is the voice-agent skill. A speaker who converses is the guide, and she has her own page: Reference previewThe guide (06-08-2026).

Zalia, live

The reference. Text and voice, in French, on the Reference previewComores SARL page. 5 fact files + a conduct charter; tested 9 for 9, trick questions included.

Valentina, ready, closed

Her agent, voice, pronunciation and 13 fact files are all wired; the buttons were withdrawn 02-08-2026. No visitor can talk with her yet.

Grace, parked

Narration only, by verdict of 19-07-2026. Her legally verified leaflets wait as the knowledge; wake her when a public assistant is wanted.

What is missing, for Valentina. A full voice test on the live page (a page opened from disk cannot use the microphone) · cutting her off mid-answer, never verified · the greeting is long in voice mode · the face does not answer, during a conversation the bust steps aside and only the voice replies, which is the real meaning of "not ready" · switching text to voice loses the dialogue · cost per minute unmeasured and the agent has no domain lock. Full detail: Reference previewthe diagnosis.

TrackNext taskEffortIt unlocks
1Prove Valentina's conversationRun the full voice round trip on the live page, fix what fails, restore the two buttons for Frank's verdict~1 dayThe first speaker a visitor can talk with
2Guard every speaker's answersWrite Valentina's witness questions; give the existing guard a schedule that runsHoursAnswers that stay right as facts change
3One-click knowledge refreshBuild the fact editor (edit, press Publish); reuse for every speaker1–2 daysColleagues keep the knowledge current
4The face that answersTest real-time mouth movement on a rented GPU against live audioDays, exploratoryThe speaker herself visibly answering
5Wake GraceWhen wanted: run the method on her leaflet corpus~1 dayLesotho citizens ask instead of read
6The sovereign betA bet, not a task: NVIDIA's self-hosted voice model, natural interruption, under a third of a second, English-only today; we asked about SpanishWatchConversation on our own machines, at centimes
How one is made
  1. Portrait on green, one photo, head and shoulders, so she can be cut out.
  2. Voice, reuse or clone; acronyms spelled as they must sound, locked in a dictionary.
  3. Script in parts, the page's argument, 6 to 10 pieces of 15 to 30 seconds.
  4. One master clip: once per person, the living eyes. Two routes now, Veo 3.1 Fast on our own key ($0.80) or Magic Hour (~1.20 EUR); never skip it, a free animate-a-still route gives flat, dead eyes. Reference previewBoth routes, drawn.
  5. The factory, repaints only the mouth per recording, under a centime a part. Reference previewDrawn, step by step.
  6. The brief for a new speaker: the sheet you copy for each person, with what never changes (the flat green, the rhythm, one take per part). Reference previewThe template to copy.
  7. Cut-out in pairs, one file for Chrome, one for Safari, always both.
  8. Page wiring, idle breath, card, icons, the slow glide; then the verification walk.

Since 07-08-2026, the best-looking parts are born, not repainted. Frank rejected every processed version of a part and chose the raw clip three times running, so a part is now bought born: the engine speaks that exact line, at its own size, and the only step before the page is removing the green. About $1 a part, once. The factory stays for long narration and for people we filmed. Reference previewThe verdict, on the factory page · Reference previewcriteria 20 and 21.

Since 28-08-2026, step 4 is skippable. A locked portrait plus the approved voice is enough; the machine (Reference previewnacer.py) births the whole section, no master clip, no repaint. Steps 1, 2, 3, 6, 7 and 8 still apply. Detail and cost: above.

Putting her online, four things that only break on a phone

All four were invisible on the desktop preview and all four hit Frank on his phone within minutes of the ABC note going live. Check them before saying a speaker is online.

  1. The still picture must be a keyed PNG, never the portrait JPG. The page falls back to it whenever video cannot play, and a JPG still carries the green background: she appears inside a rectangle. Cut it out with the same recipe as the clips, and add a version mark to the file name or phones keep the old one.
  2. The voice-only fallback must hand over to the video, not lock it out. The small sound file always wins the race on a phone; our player then refused the video that arrived a second later, so she stayed a still picture with a voice, forever. The video takes over at the same word. Test on a slow connection, never on the machine that made it.
  3. Low Power Mode blocks even a silent video on iPhone, so she was simply absent and the card floated over nothing. If nothing has painted 1.5 seconds after the page loads, show the cut-out still; any video that starts hides it again.
  4. A button's own label can swallow its tap. The stop control's word sits over the button once visible; on a phone the tap landed on the word. Let the visible label take clicks too, and make stop unconditional so no stale state can make it do nothing.

And "online" means the live address serves it. A push is not a deploy: rapid pushes queue behind each other and Frank read old text three times in a row. Wrap the publish so it only reports success after the live page proves the change.

What it costs

Corrected 31-08-2026. This table used to price the old way: one master clip, then a mouth repainted on each part. Since the route was settled on 28-08 a section is born whole, so the money is counted per published minute, not per take. The door for all of it is the speakers topic.

Paid once, per person. A locked portrait, and her voice approved before any picture exists: cents a line at ElevenLabs, or nothing at all with Voicebox on the Mac. Getting the voice right first is the whole saving, because then a changed word is a re-render and not a purchase.

Then, per published minute, in one look:

OptionForCost, per published minuteCatch
Still speakerA simple page, smallest budgetVoice only: cents a line, or free with VoiceboxShe does not move
Machine, the locked portrait plus our approved audio (rented computer). The route, since 28-08-2026The finished page, chosen on the eye, not on the price$33 to $37, measuredSlow, hours of machine time; it does not fix a bad recording
Grok, through his own subscription (skill grok-speaker)A quick free draft, testing the wordsNothing beyond the subscriptionA fresh scene each take, capped at 15 seconds — seams, jumping posture
Grok, bought by the secondRarely worth it now the free session existsAbout $5, and $15 or more once three takes are bought for one keptSame lottery as the free session, but billed
HeyGen Avatar IIIWorth pricing before a long narrationAbout $1 to $4, their price list, never measured by usBlocked on credit today; a different account, a different face
A conversation instead of a narration (ElevenLabs convai)A visitor who asks, not one who watches$0.08 the conversation minuteVoice only, the bust idles

Where the $33 comes from, corrected 31-08-2026. This table said $1 to $3 on 28-08, then $35 to $47 earlier today — both from a guessed hourly rate. The measurement, from the engine test's own record: two clips of 6.4 seconds, about 2.5 hours of A100 each; round 6 of that same test proved the cheaper A40 cannot even load the model. A published minute is 9.4 clips, so about 23 hours of A100, and Reference previewrunpod.io/pricing, checked 31-08-2026, prices that machine at $1.39 (PCIe) to $1.59 (SXM4) the hour: $33 to $37. The whole test, failed rounds included, cost $17 for 12.8 seconds of video.

So the machine route is the most expensive one we have, not the cheapest. It was chosen on 28-08 because Frank's eye picked it blind, and that reason stands on its own. But anything that plans a long narration on it should price it at $33 to $37 the minute, look again at HeyGen at $1 to $4, and remember that Grok through the subscription still costs nothing beyond the subscription.

One rule before any per-minute service: check what the open model on the rented computer does first, and say both prices. That is how the month that cost more than $120 ends.

Give it to a colleague

This whole page is a skill, a small file of instructions their own assistant reads and follows: a person in the page who presents it in her own voice. They install it once; after that they simply say what they want, in their own words, and their assistant does it, on their own machine, with their own accounts, nothing of Frank's.

One command, in their Claude Code:

npx skills add gfrankgva/studio-skills

That installs this skill together with the studio's six others.

Or take the file itself: Reference previewdownload the skill (one file, put it in their ~/.claude/skills/site-speaker/) · Reference previewread it on GitHub · Reference previewthe whole pack.

Without the skill it still works, the method on this page still works by hand. The skill only means their assistant already knows it.

Where the material lives
WhatFile
The whole recipe (the playbook)site-speaker skill
Picture → talking video → cutoutphoto-avatar-video skill
The talking feature, when readyvoice-agent skill · Reference previewVoice
Frank himself explaining a documentfrank-avatar skill
The recipes and the rulesReference previewMake pages move · Reference previewcriteria
The factoryReference previewthe factory, drawn · musetalk.md
Hélène's master clip, made with Veo 3.1, both traps found and the rejected free-route evidencehow-it-was-made.md · Reference previewfactory.html, step 2
Valentina's page, clips, recordingsvalentina-video.html · busto/ · narracion/
Hola, soy Valentina
Shall I present this page?