---
name: photo-avatar-video
description: "Turn ONE picture of a person into an animated talking video in that person's voice — and optionally into a transparent 'bust' that melts into a web page with no frame. TRIGGER when you have a picture and want it to speak ('make her talk', 'animated video with the voice of the person in the picture', 'a video avatar of X', 'a talking bust of X on the page'). NOT for a live interactive Q&A agent (use /voice-agent — this skill produces recorded video)."
---
*From Frank's Studio — studio.smartrules.ai · install: npx skills add gfrankgva/studio-skills*
# Photo → talking avatar video (any person)
**What this makes.** From one photo: a lip-synced talking video, voiced by a clone of the person (or a chosen stock voice). Optionally: a **transparent talking bust** ("busto" = bust; webm+mov alpha pair) that sits on a page with no rectangle — the person becomes part of the design. Proven end to end 31-07-2026 on Valentina ([azfa.eregistrations.dev/valentina-video](https://azfa.eregistrations.dev/valentina-video)).
## The five sizes (ask which one)
1. **Face only** — corner chip; lip-sync shows most; any photo works.
2. **Face + neck** — the portrait; any photo works.
3. **Small half body (bust)** — head/shoulders/chest; the default for a page corner; any decent photo.
4. **Big half body (waist up)** — hands may appear (engines invent hands: keep motion stable); needs a wider photo OR a Gemini zoom-out extension (extend the canvas while keeping the original face pixels).
5. **Full body** — same pipeline with a full-body source; mouth small, gestures dominant; entrance scenes, not narration.
**Where it came from.** Settled after a measured engine comparison (2026-07-27: Magic Hour won) plus the alpha-video research in the Studio motion notes ([studio.smartrules.ai](https://smartrules.ai/studio)).
## Phase 0 — wizard (ask before building)
Ask only what is missing, in one numbered block:
1. **Script** — what must the person say? (text, or a document to narrate)
2. **Voice** — a voice sample of the person (≥30 s clean speech, or a video to extract it from)? An existing ElevenLabs voice_id? Or pick a stock voice?
3. **Shape** — standalone video (mp4, framed) or transparent bust for a web page?
4. **Length** — per clip. >60 s means stitching; propose cutting the script into segments instead.
Consent check: cloning a real person's voice and face needs that person's OK. Confirm the person agreed before generating anything.
## Phase 1 — voice
**Keys:** your keys file (`ELEVENLABS_API_KEY`, `MAGICHOUR_API_KEY`, `GEMINI_PAID_KEY`). Pass keys to curl via a temp 0600 config file, never on the command line.
- **Clone from sample:** ElevenLabs IVC — `POST /v1/voices/add` with 1–3 clean samples. From a video: `ffmpeg -i in.mp4 -vn -ac 1 -ar 44100 sample.wav` first. Sovereign/free alternative: a self-hosted voice-clone TTS server (e.g. Voicebox; takes a WAV reference ≤30 s).
- **Existing voice:** reuse the voice_id (Valentina = "Pilar Durán" `x6LHvMgpXmty838MUqHh`, in the ElevenLabs shared library).
- **Generate speech:** model `eleven_multilingual_v2` (the family that keeps clones trained — check `fine_tuning.state` before delivery). A small curl loop, one mp3 per script block, is all the tooling needed.
- **If audio is already provided** (e.g. existing narration segments): skip straight to Phase 2 — the video engine takes an audio file as-is.
## Phase 2 — video (photo + audio → talking clip)
**Engine: the factory is DEFAULT once proven on a job; Magic Hour is the fallback.** The standing rule: before paying any per-minute service, check whether an open model on a rented GPU does the job, and say the cost of both. Magic Hour (settled 2026-07-27 after measured comparison; ~48 credits/s ≈ 4 centimes/s, up to 60 s per generation, no watermark, and the ONLY platform whose website credits fund the API — the wallet trap) stays the fallback for when the factory endpoint isn't live yet or loses on quality/turnaround.
- Driver: a small script that submits the Magic Hour job, prints the job id first, then polls every 20 s (Magic Hour has no job-list endpoint — a lost poll loses the job).
- Settings that won the sweep: `generation_mode: "stable"`, `intensity: 0.15` (intensity is inert; **never** `prompted` — measured motion 29.89 vs 5.86, the whole scene swims).
- Photo prep: a bust-framed portrait (head + shoulders, eyes to camera, mouth closed). If the source photo has a busy background and a bust is wanted, first regenerate/edit it onto a **flat chroma green** background (a Gemini image edit, or a Higgsfield/Krea image model). Keep the face pixels honest: if the model repaints the face, composite the original face pixels back and take only clothing/background from the edit.
- **Verify motion** before accepting: mean abs frame diff between consecutive frames; 5–7 is right, >15 means the background is swimming — regenerate.
- **Cost ladder (verified 31-07-2026, EUR per minute of output).** **The factory floor ≈ 0.02–0.05: LivePortrait (idle base) + MuseTalk 1.5 on a RunPod 4090 ($0.34/h)** — mouth-only editing, the most chroma-key-stable option; setup + one-line command in the worker repo, [github.com/gfrankgva/musetalk-worker](https://github.com/gfrankgva/musetalk-worker). Magic Hour ≈ 2.4 (fallback, subscription-capped). Cheaper than Magic Hour but not the factory: WaveSpeed InfiniteTalk/Wan-S2V 480p ≈ 1.6 (plain REST, pay-per-use); Kling avatar v2 via [fal.ai](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/standard) ≈ 3.1 (no minimum — the official Kling API needs ~$700 resource packs, and Kling WEBSITE credits never fund the API, re-confirmed 31-07-2026 error 1102; website credits are usable only by hand in their app). Avoid: fal's InfiniteTalk listing (~11 EUR/min for the same model), OmniHuman for calm busts (expensive, too much motion).
## Phase 3 — transparent bust (optional, for "melted into the site")
Only when the destination is a web page. Full doctrine + embed grammar: the Studio motion notes ([studio.smartrules.ai](https://smartrules.ai/studio)).
1. Generate the clip on the flat green background (Phase 2).
2. Key + convert to the two-format alpha pair (Mac only for step b):
```bash
# a) green mp4 -> WebM VP9 with alpha (Chrome/Firefox/Edge)
ffmpeg -i in.mp4 \
-vf "chromakey=0x00d000:0.12:0.08,despill=type=green,format=yuva420p" \
-c:v libvpx-vp9 -pix_fmt yuva420p -b:v 0 -crf 30 -auto-alt-ref 0 bust.webm
# b) -> HEVC with alpha for Safari (hevc_videotoolbox, hardware, Mac only)
ffmpeg -c:v libvpx-vp9 -i bust.webm \
-c:v hevc_videotoolbox -allow_sw 1 -alpha_quality 0.75 -vtag hvc1 bust.mov
```
Tune the key color to the actual green (`ffprobe`/eyedropper a frame); check edges on hair.
3. Embed borderless — Safari picks the first playable source; the bottom fade is what "melts" the torso into the page:
```html
```
4. File weight: alpha video is heavy. Keep busts ≤20–30 s per file, 480–720 px tall, and lazy-load. Smallest-file alternative when weight bites: the `stacked-alpha-video` npm web component (double-height video, alpha as luma, WebGL recombine).
5. An **idle loop** (2–4 s of subtle breathing, generated from the same photo with no speech audio) crossfaded with the speaking clips makes the person feel present between narration segments.
## Verification (never skip)
- Play the result — judge by ear and eye, on the real file, before delivering.
- For a web bust: open the page in Safari AND Chrome (the two alpha codecs differ); check the hair edge and the bottom fade.
- Cost sanity: announce estimated credits BEFORE generating (seconds × 48 Magic Hour credits). Spend one cheap test call on any NEW platform before promising anything (the wallet trap).
## Proven values (Valentina, 31-07-2026 — start from these)
- Green portrait via Gemini `gemini-3-pro-image`: ask for "perfectly uniform, flat, evenly lit chroma-key green screen (solid green #00B140)" + "keep the person EXACTLY as she is". The model returns a slightly different green — **sample a corner pixel of the GENERATED VIDEO** before keying (Valentina's came out `#00a547`).
- Chromakey: `chromakey=0x00a547:0.14:0.08,despill=type=green` (similarity 0.14, blend 0.08) **+ two alpha erosions** to shave the green rim: `...,format=yuva420p,split[k1][k2];[k2]alphaextract,erosion,erosion[al];[k1][al]alphamerge`. Two traps proven 31-07-2026: raising similarity above ~0.14 makes DARK HAIR semi-transparent (reads grey on light pages), and any explicit `despill mix` greys dark colors — the rim is fixed by erosion, never by stronger keying.
- Embed trap: a **muted, hidden video refuses to play** ("video-only background media paused to save power") — segments must play unmuted (they carry the voice); only the idle loop is muted, and it must stay visible.
- Safari file: `-c:v hevc_videotoolbox -allow_sw 1 -alpha_quality 0.75 -q:v 55 -vtag hvc1`, audio mapped from the original green mp4 (`-map 0:v -map 1:a`). ffprobe reports `yuv420p` even when alpha is present — verify by eye in Safari, not by probe.
- Poll loops: wrap every status call in try/except and keep polling — a transient 502 on a poll killed a healthy render once.
- ~48 credits/second. **Check the balance BEFORE submitting a batch**: a mid-batch 402 leaves half the jobs unsubmitted (submit cheapest-first, or compare total seconds × 48 against the balance shown in any 402 message).
## Known traps
- **The wallet trap** — website subscription credits ≠ API credits everywhere except Magic Hour. Test with one cheap call first.
- ElevenLabs `speed` unreliable at small steps; choose by ear.
- An approved render is a specific FILE — never re-roll approved segments.
- Long audio: cut the script, generate per segment; never one 3-minute take.
- Chroma green clothing/eyes will key out — check the portrait before generating 8 segments.
## The standard finish — every video ends this way, in this order
The generators return small frames (Grok 544 square, others near it), and that resolution — not the mouth engine — is what makes a face look soft. Whatever engine made the clip:
1. **Repaint the mouth at native size.** Never upscale first: the repainted mouth would end up softer than the face around it.
2. **Upscale ×4** with Real-ESRGAN, free on your own GPU: `pip install spandrel`, download `RealESRGAN_x4plus.pth` (github.com/xinntao/Real-ESRGAN releases), extract frames with ffmpeg, run the model per frame, reassemble at the same fps. About 9 s a frame on an Apple-silicon Mac; 544 becomes 2176.
3. **Add a medium grain**, a new pattern every frame, luma only so colour and the green stay clean (numpy: `rng.normal(0, 4.0, frame.shape[:2])` added equally to RGB, then clip). **Never skip it** — without it the skin is too even and reads as retouched. This level was chosen from four versions rendered side by side; render options, never reason about them.
4. **Then** chromakey. Keying last, on sharp frames, gives better edges. **The grain must never be visible to the key.** Grain noise drifts dark, low-saturation pixels (hair) toward the key colour and the chromakey eats them (solid hair pixels measured 27.7% -> 2.6%; the clip reads bright and washed out on a light page). So compute the transparency from the CLEAN upscaled frames and take the colour from the GRAINED frames, merged: `[1:v]chromakey=:0.14:0.08,format=yuva420p,alphaextract,erosion,erosion[al];[0:v]despill=type=green[rgb];[rgb][al]alphamerge[out]` (input 0 = grained, 1 = clean). And sample the key green from CLEAN frames - grain clipping falsifies the sampled value.
**Deliver two sizes.** Grain compresses badly: the same clip is 24 MB at 2176 and 0.7 MB at 1088. The big one to keep, the small one for the page.
**Two explanations that look right and are wrong — they were measured, do not re-argue them.** The upscaler does NOT shift skin colour (a cheek moved R−0.6 G+0.4 B−0.7 out of 255) and does NOT remove micro-texture (it measures higher). It makes skin EVEN, and evenness is what reads as retouched; the grain works by breaking the evenness.
## The birth of a character — do this ONCE, at the best quality available
**A character is born once, at the best quality money can buy that day — face AND voice. Everything afterwards is reproduced by the factory for centimes.** The master assets are permanent: the factory copies them forever and can never exceed them.
1. **Survey before buying.** Video engines change every few weeks; check what is best TODAY and its price per second before spending. Judge a master clip on the FACE alone (eyes alive, natural blinks, still camera, chroma background intact) — the mouth gets repainted anyway.
2. **Buy the face once** — one master clip from the green portrait.
3. **Harvest the voice at the same time.** Engines that generate video with native speech give the character a voice that matches the face. Generate ~30-40 s of continuous speech, extract the audio, strip silences, normalise loudness, and **clone it once** into a voice service. From then on the voice is yours: any script, any language, under a centime a line, identical forever — instead of paying per second for a different voice every time. (Clone endpoints commonly cap uploads around 11 MB — send mp3, not raw WAV.)
4. **Then the factory reproduces** — mouth repainted per sentence on a rented GPU, the cloned voice speaking every new line. The character never changes face or voice again.