--- name: site-speaker description: "Put a speaking person on any web page: a transparent talking bust who greets the visitor, guides them through the page section by section with her own voice, and can be paused or stopped. The full reproducible recipe (portrait, voice, segments, video, transparency, page wiring, UX), proven end to end on Valentina for AZFA, 31-07-2026. TRIGGER when you want 'a Valentina on this site', 'someone who explains this page', 'a speaker/presenter on the page', 'the same video guide for X'. NOT for a live question-answer agent alone (that is /voice-agent), NOT for a single talking clip with no page (that is /photo-avatar-video)." --- *From Frank's Studio — studio.smartrules.ai · install: npx skills add gfrankgva/studio-skills* # A speaking person on a site **What this builds.** A person stands in the corner of a page, breathing quietly. A small translucent card beside her says who she is and offers one action. The visitor presses it: she speaks, the page travels with her section by section, and two discreet controls let them pause or stop. No frame, no rectangle, no widget: she is part of the page. **Reference build (copy from it, it works):** Valentina on [azfa.eregistrations.dev/valentina-video](https://azfa.eregistrations.dev/valentina-video). View the page source — the whole wiring block sits at the end of the HTML. --- ## Phase 0 — decide, in one block of questions 1. **Who speaks?** An existing persona with a voice already (reuse it) or a new person (a photo + a voice to pick or clone). 2. **What does she say?** The page's own argument, cut into 6 to 10 parts, one per section. 3. **Which page?** Work on a COPY of it, never the live page, until the owner approves. 4. **Language and register** (usted/tú, formal/warm). Consent: a real person's face and voice need that person's agreement. Ask before generating. --- ## Phase 1 — the portrait on green The video engines need her on a flat chroma background so she can be cut out. - Take the persona's existing photo. Reframe it as a **head-and-shoulders bust on flat chroma green** with one Gemini image edit (model `gemini-3-pro-image`, key from your keys file). - **Free alternatives for the image work:** the same edit runs at zero cost in [Google AI Studio](https://aistudio.google.com) (Nano Banana 2, free daily quota, ~10 requests/min) — good for manual retries and variants. A full-body version of the persona is made the same way — always demand feet visible, matching clothes, flat green, "keep her EXACTLY as she is". - The prompt must say twice: replace the ENTIRE background with uniform green **and** keep the person EXACTLY as she is (same face, expression, hair, clothes, lighting). Otherwise the model repaints her. - Get the portrait approved before anything is generated from it. ## Phase 2 — the voice - **Reuse before cloning.** Check the ElevenLabs shared library first: `GET /v1/shared-voices?search=`. Valentina's "Pilar Durán" (`x6LHvMgpXmty838MUqHh`) was already there, so no clone was needed and any account can use it. - **A pronunciation dictionary is mandatory**, created once and attached to everything she says: `POST /v1/pronunciation-dictionaries/add-from-rules` with `type: "alias"` rules, then attach to the agent via `conversation_config.tts.pronunciation_dictionary_locators`. Acronyms are the trap: written plainly, a voice reads AZFA as a word. Spell it as it must sound. The shape that won, after three rounds: **`azeta-efe-a`** (hyphens, no commas: commas make it drag) and **`unktád`** (accent for the stress). - **Never set `speed`.** A 0.85 speed made one segment 33.7 s where the others were 27; the ear hears it instantly as "boring, slow". Generate with `model_id` only, matching the other segments. - **Model choice:** `eleven_multilingual_v2` for recorded narration; **`eleven_flash_v2_5` for a live agent** (one second faster per answer, half the credits, same voice IF the clone shows `fine_tuned` for it in `GET /v1/voices/{id}` → `fine_tuning.state`). ### Match the clone to the master — the "muffled" fix A clone never comes back sounding quite like the voice it copies, and the owner's ear hears it as "muffled". That is not a taste question, it is **measurable**, and it is repairable in one ffmpeg pass. Do this for every new character, right after the clone. **Measure first, in two numbers.** Compare the reference audio and the clone's test line on (a) median pitch (autocorrelation F0) and (b) energy in six bands. A typical bad clone: +11 Hz too high and −7.6 dB in 2–4 kHz — the presence band. Boomy plus no presence IS the muffled sound. **Then correct, one filter chain** — pitch by resampling, presence by a wide bell — and re-measure until every band sits within about 1 dB: ``` ffmpeg -i clone.mp3 -af "asetrate=44100*R,aresample=44100,atempo=1/R,\ equalizer=f=3000:width_type=h:width=2200:g=+7.6,\ loudnorm=I=-16:TP=-1.5:LRA=11" -ar 44100 -ac 1 out.wav ``` `R = master_pitch / clone_pitch`. The gain figures are **not constants** — they are your measured gaps. **Three ways the reference itself dulls the clone — check all three before blaming the voice service.** 1. **A resampling accident.** An ffmpeg `concat` of mixed-rate clips silently produced a 192000 Hz reference once. Force `-ar 44100` on every input. 2. **`remove_background_noise=true`.** The denoiser eats presence. Send `false`. 3. **One voice only in the reference.** A generator's native speech drifts between takes (one character measured 197–228 Hz across four generations). Cloning a mix averages them into a blur — keep only takes whose pitch matches the approved one. **And say it plainly to the owner: a clone is a near-match, not a copy.** Judged side by side, the original wins. Good enough for volume (any line, any language, under a centime); a hero line may be worth generating at the master engine. ## Phase 3 — the script, cut into segments One segment per page block, 15 to 30 seconds each. Keep the mapping in a `guion.json`: `{block, tab, audio, text}` where `block` is the element id to travel to and `tab` an optional panel to open. Generate one mp3 per segment (`seg1.mp3` … `segN.mp3`). ## Phase 4 — the moving picture ### The standing method: buy ONE master base, then repaint it forever The rule (31-07-2026): **every person gets one good Magic Hour clip, kept as their master base.** Everything after that is made by the factory repainting that clip, for about a tenth of the paid price. This is the only way to get paid-quality eyes and expression at factory prices, because the factory alone invents the face and it looks flat. - **How long should the master be? 25 to 30 seconds.** Not "a few". In speech mode the worker extends a short base with a plain repeat (`-stream_loop`), so every restart is a visible jump and her blinks repeat on a fixed cycle. At 25 to 30 s, a normal segment needs no loop at all and a long one loops at most once. Cost of the master: about **1.20 EUR once** (48 credits per second). - **Shoot the master right:** on the flat green, calm settings (`stable`, intensity 0.15), mouth active (any words, they get repainted), and ideally beginning and ending in the same neutral pose so a repeat is invisible. - **Also keep a 4-second silent idle** (mouth closed, one blink) for her resting state. - **Store both** beside the person's assets and never delete them: they are the origin of every future video for that person. - Improvement worth making when it matters: give the talk base the same ping-pong treatment the idle already gets, so any length loops seamlessly. **The upgrade, built and measured 31-07-2026.** MuseTalk's second mode, "avatar/realtime", prepares a base clip ONCE (face boxes, masks, latents cached) and then runs only the two fast models per audio. The naive path re-detects the face on every frame of every job; preparing once per person and reusing the cache turned **138 s of GPU per clip into 11.5 to 18.7 s** (67 fps on a 4090, about 2.6x faster than real time), **under half a centime a clip**. Preparation costs 80 to 218 s once per bust, and the cache travels with `--cache .tar.gz` so it survives the worker dying. Numbers and cache design: the worker repo, [github.com/gfrankgva/musetalk-worker](https://github.com/gfrankgva/musetalk-worker). ### The three routes, cheapest first: - **The factory (under 1 centime per minute of finished video once the bust is prepared).** A MuseTalk worker on a rented GPU (RunPod serverless, RTX 4090 ≈ $0.34/h; set `` in your environment). Prepare once: `python3 tools/musetalk.py --prepare --reuse-base --cache .cache.tar.gz`. Then every segment: `python3 tools/musetalk.py --reuse-base --cache .cache.tar.gz --audio --out `. From a photo instead: `--photo `. Worker + tooling: [github.com/gfrankgva/musetalk-worker](https://github.com/gfrankgva/musetalk-worker). Cold start 130 to 240 s, then about 12 s per 30-second segment. - **Keep the master base short for preparation.** Preparation is per frame of base video and the base is only ever ping-pong looped, so a 4-second base renders identically to a 28-second one, prepares in ~12 s instead of ~200 s, and caches in 2.5 MB instead of 16 MB. The cost is that the head motion repeats every 8 s. - **THE WORKHORSE: repaint the mouth on the master base.** `--reuse-base --audio `. Everything except the mouth (eyes, blinks, head, expression) is copied frame for frame from the base, so a paid clip's quality is preserved while the words change. This is how Valentina's first sentence was corrected without paying again. - **Magic Hour (about 2.4 EUR/minute)** to create the master base, and for anything where the face fills the frame. Submit via the API and poll every 20 s — Magic Hour has no job-list endpoint, so print the job id first or a lost poll loses the job. Settings: `generation_mode: "stable"`, `intensity: 0.15`, never `prompted`. - Also make one **4-second silent idle clip** (mouth closed, gentle breathing, one blink) for her resting state. ### Two mouth engines — pick by how big the mouth is on screen The factory has a second engine, and they trade sharpness against price. | | **MuseTalk** | **LatentSync** | |---|---|---| | The mouth, at 6× | flat waxy gradient, a visible texture break at the neck | real lip border, volume, a highlight, skin texture on the chin | | Price | **$0.008 per minute of video** | **$0.147** — 18× dearer, 19× slower; a 10 s line is about 2 centimes | The sensible split: MuseTalk for long page narration, LatentSync for **face**-sized speakers and hero clips, where the mouth is large on screen and the softness shows. One implementation trap: LatentSync's upstream `restore_img` inverse-warps the whole face crop, costing the eyes it never repaints ~27% of their detail — composite on the model's own mouth mask instead (the public worker image `ghcr.io/gfrankgva/latentsync-worker-pub` does this by default). ### The standard finish — every video ends this way, in this order The generators return small frames (Grok 544 square, others near it), and that resolution — not the mouth engine — is what makes a face look soft. Whatever engine made the clip: 1. **Repaint the mouth at native size.** Never upscale first: the repainted mouth would end up softer than the face around it. 2. **Upscale ×4** with Real-ESRGAN, free on your own GPU: `pip install spandrel`, download `RealESRGAN_x4plus.pth` (github.com/xinntao/Real-ESRGAN releases), extract frames with ffmpeg, run the model per frame, reassemble at the same fps. About 9 s a frame on an Apple-silicon Mac; 544 becomes 2176. 3. **Add a medium grain**, a new pattern every frame, luma only so colour and the green stay clean (numpy: `rng.normal(0, 4.0, frame.shape[:2])` added equally to RGB, then clip). **Never skip it** — without it the skin is too even and reads as retouched. 4. **Then** the chromakey of Phase 5 — keying last, on sharp frames, gives better edges. **The grain must never be visible to the key.** Grain noise drifts dark, low-saturation pixels (hair) toward the key colour and the chromakey eats them (solid hair pixels measured 27.7% -> 2.6%; the clip reads bright and washed out on a light page). So compute the transparency from the CLEAN upscaled frames and take the colour from the GRAINED frames, merged: `[1:v]chromakey=:0.14:0.08,format=yuva420p,alphaextract,erosion,erosion[al];[0:v]despill=type=green[rgb];[rgb][al]alphamerge[out]` (input 0 = grained, 1 = clean). And sample the key green from CLEAN frames - grain clipping falsifies the sampled value. **Deliver two sizes.** Grain compresses badly: the same clip is 24 MB at 2176 and 0.7 MB at 1088. The big one to keep, the small one for the page. **Two explanations that look right and are wrong — they were measured, do not re-argue them.** The upscaler does NOT shift skin colour (a cheek moved R−0.6 G+0.4 B−0.7 out of 255) and does NOT remove micro-texture (it measures higher). It makes skin EVEN, and evenness is what reads as retouched; the grain works by breaking the evenness. ## Phase 5 — cut her out **Keying is the LAST step of the standard finish above** — key the upscaled, grained clip, never the raw render. The working recipe, per clip: chromakey at the sampled green + despill + **two alpha erosions** to shave the rim, into a WebM (Chrome/Firefox/Edge) and a HEVC .mov (Safari; the HEVC-alpha encode needs a Mac). **Two free quality wins in the keying, verified by execution (31-07-2026):** convert to `format=yuva444p` BEFORE the chromakey (`yuva420p` makes the filter read half-resolution colour, which is what blocks up hair edges), and feather the matte with `gblur=sigma=1.5:planes=8` (alpha only) rather than shrinking it with erosion. A matte gamma curve via `geq` on the alpha reproduces the softer falloff OBS uses for hair. **Why still a green screen (checked 31-07-2026).** The modern cut-out models that would remove the need for it (MatAnyone 1 and 2, SAM2Matting, VideoMaMa, MaGGIe) are all **non-commercial licences**, so they are closed for commercial and client work. The only commercially usable ones are RobustVideoMatting (GPL-3.0, weaker on hair than chromakey) and, worth adding when hair edges bother you, **AGED despill** (MIT, no machine learning, touches only the semi-transparent edge pixels, 4K at 60 fps): github.com/AdolphGong/aged-despill. Two traps, both cost an hour: - **Sample the green from the GENERATED video**, not the portrait (Valentina's came out `#00a547`). - **Never raise the similarity above ~0.14 and never set an explicit despill mix**: dark hair goes semi-transparent and reads grey on light pages. The rim is fixed by erosion, never by stronger keying. ## Phase 6 — the page Copy the script block at the end of the reference page (view source at [azfa.eregistrations.dev/valentina-video](https://azfa.eregistrations.dev/valentina-video)). What matters: **The idle loop, measured fix (31-07-2026).** A hard-cut loop shows a visible jump (seam 13x the normal frame-to-frame change). Ping-pong removes the seam perfectly but reverses her motion. The answer is a short crossfade, which keeps motion forward, preserves the alpha channel and costs 0.2 s of clip: ```bash ffmpeg -y -c:v libvpx-vp9 -i idle.webm -filter_complex \ "[0:v]format=yuva420p,split[a][b];[a]trim=start=0.25,setpts=PTS-STARTPTS[body];\ [b]trim=start=0:end=0.25,setpts=PTS-STARTPTS[head];\ [body][head]xfade=transition=fade:duration=0.25:offset=3.5[v]" -map "[v]" \ -c:v libvpx-vp9 -pix_fmt yuva420p -b:v 0 -crf 30 -auto-alt-ref 0 -an idle-loop.webm ``` `offset` = clip duration minus twice the crossfade. **Sweep the duration and measure**: the residual is not smooth (0.25 s scored 0.49x, 0.28 s scored 2.85x on one clip). On TALK clips, place the window over a closed-mouth pause or the blend ghosts two mouth shapes. **Her.** A borderless `