--- name: photo-avatar-video description: "Turn ONE picture of a person into an animated talking video in that person's voice — and optionally into a transparent 'bust' that melts into a web page with no frame. TRIGGER when you have a picture and want it to speak ('make her talk', 'animated video with the voice of the person in the picture', 'a video avatar of X', 'a talking bust of X on the page'). NOT for a live interactive Q&A agent (use /voice-agent — this skill produces recorded video)." --- *From Frank's Studio — studio.smartrules.ai · install: npx skills add gfrankgva/studio-skills* # Photo → talking avatar video (any person) **What this makes.** From one photo: a lip-synced talking video, voiced by a clone of the person (or a chosen stock voice). Optionally: a **transparent talking bust** ("busto" = bust; webm+mov alpha pair) that sits on a page with no rectangle — the person becomes part of the design. Proven end to end 31-07-2026 on Valentina ([azfa.eregistrations.dev/valentina-video](https://azfa.eregistrations.dev/valentina-video)). ## The five sizes (ask which one) 1. **Face only** — corner chip; lip-sync shows most; any photo works. 2. **Face + neck** — the portrait; any photo works. 3. **Small half body (bust)** — head/shoulders/chest; the default for a page corner; any decent photo. 4. **Big half body (waist up)** — hands may appear (engines invent hands: keep motion stable); needs a wider photo OR a Gemini zoom-out extension (extend the canvas while keeping the original face pixels). 5. **Full body** — same pipeline with a full-body source; mouth small, gestures dominant; entrance scenes, not narration. **Where it came from.** Settled after a measured engine comparison (2026-07-27: Magic Hour won) plus the alpha-video research in the Studio motion notes ([studio.smartrules.ai](https://smartrules.ai/studio)). ## Phase 0 — wizard (ask before building) Ask only what is missing, in one numbered block: 1. **Script** — what must the person say? (text, or a document to narrate) 2. **Voice** — a voice sample of the person (≥30 s clean speech, or a video to extract it from)? An existing ElevenLabs voice_id? Or pick a stock voice? 3. **Shape** — standalone video (mp4, framed) or transparent bust for a web page? 4. **Length** — per clip. >60 s means stitching; propose cutting the script into segments instead. Consent check: cloning a real person's voice and face needs that person's OK. Confirm the person agreed before generating anything. ## Phase 1 — voice **Keys:** your keys file (`ELEVENLABS_API_KEY`, `MAGICHOUR_API_KEY`, `GEMINI_PAID_KEY`). Pass keys to curl via a temp 0600 config file, never on the command line. - **Clone from sample:** ElevenLabs IVC — `POST /v1/voices/add` with 1–3 clean samples. From a video: `ffmpeg -i in.mp4 -vn -ac 1 -ar 44100 sample.wav` first. Sovereign/free alternative: a self-hosted voice-clone TTS server (e.g. Voicebox; takes a WAV reference ≤30 s). - **Existing voice:** reuse the voice_id (Valentina = "Pilar Durán" `x6LHvMgpXmty838MUqHh`, in the ElevenLabs shared library). - **Generate speech:** model `eleven_multilingual_v2` (the family that keeps clones trained — check `fine_tuning.state` before delivery). A small curl loop, one mp3 per script block, is all the tooling needed. - **If audio is already provided** (e.g. existing narration segments): skip straight to Phase 2 — the video engine takes an audio file as-is. ## Phase 2 — video (photo + audio → talking clip) **Engine: the factory is DEFAULT once proven on a job; Magic Hour is the fallback.** The standing rule: before paying any per-minute service, check whether an open model on a rented GPU does the job, and say the cost of both. Magic Hour (settled 2026-07-27 after measured comparison; ~48 credits/s ≈ 4 centimes/s, up to 60 s per generation, no watermark, and the ONLY platform whose website credits fund the API — the wallet trap) stays the fallback for when the factory endpoint isn't live yet or loses on quality/turnaround. - Driver: a small script that submits the Magic Hour job, prints the job id first, then polls every 20 s (Magic Hour has no job-list endpoint — a lost poll loses the job). - Settings that won the sweep: `generation_mode: "stable"`, `intensity: 0.15` (intensity is inert; **never** `prompted` — measured motion 29.89 vs 5.86, the whole scene swims). - Photo prep: a bust-framed portrait (head + shoulders, eyes to camera, mouth closed). If the source photo has a busy background and a bust is wanted, first regenerate/edit it onto a **flat chroma green** background (a Gemini image edit, or a Higgsfield/Krea image model). Keep the face pixels honest: if the model repaints the face, composite the original face pixels back and take only clothing/background from the edit. - **Verify motion** before accepting: mean abs frame diff between consecutive frames; 5–7 is right, >15 means the background is swimming — regenerate. - **Cost ladder (verified 31-07-2026, EUR per minute of output).** **The factory floor ≈ 0.02–0.05: LivePortrait (idle base) + MuseTalk 1.5 on a RunPod 4090 ($0.34/h)** — mouth-only editing, the most chroma-key-stable option; setup + one-line command in the worker repo, [github.com/gfrankgva/musetalk-worker](https://github.com/gfrankgva/musetalk-worker). Magic Hour ≈ 2.4 (fallback, subscription-capped). Cheaper than Magic Hour but not the factory: WaveSpeed InfiniteTalk/Wan-S2V 480p ≈ 1.6 (plain REST, pay-per-use); Kling avatar v2 via [fal.ai](https://fal.ai/models/fal-ai/kling-video/ai-avatar/v2/standard) ≈ 3.1 (no minimum — the official Kling API needs ~$700 resource packs, and Kling WEBSITE credits never fund the API, re-confirmed 31-07-2026 error 1102; website credits are usable only by hand in their app). Avoid: fal's InfiniteTalk listing (~11 EUR/min for the same model), OmniHuman for calm busts (expensive, too much motion). ## Phase 3 — transparent bust (optional, for "melted into the site") Only when the destination is a web page. Full doctrine + embed grammar: the Studio motion notes ([studio.smartrules.ai](https://smartrules.ai/studio)). 1. Generate the clip on the flat green background (Phase 2). 2. Key + convert to the two-format alpha pair (Mac only for step b): ```bash # a) green mp4 -> WebM VP9 with alpha (Chrome/Firefox/Edge) ffmpeg -i in.mp4 \ -vf "chromakey=0x00d000:0.12:0.08,despill=type=green,format=yuva420p" \ -c:v libvpx-vp9 -pix_fmt yuva420p -b:v 0 -crf 30 -auto-alt-ref 0 bust.webm # b) -> HEVC with alpha for Safari (hevc_videotoolbox, hardware, Mac only) ffmpeg -c:v libvpx-vp9 -i bust.webm \ -c:v hevc_videotoolbox -allow_sw 1 -alpha_quality 0.75 -vtag hvc1 bust.mov ``` Tune the key color to the actual green (`ffprobe`/eyedropper a frame); check edges on hair. 3. Embed borderless — Safari picks the first playable source; the bottom fade is what "melts" the torso into the page: ```html ``` 4. File weight: alpha video is heavy. Keep busts ≤20–30 s per file, 480–720 px tall, and lazy-load. Smallest-file alternative when weight bites: the `stacked-alpha-video` npm web component (double-height video, alpha as luma, WebGL recombine). 5. An **idle loop** (2–4 s of subtle breathing, generated from the same photo with no speech audio) crossfaded with the speaking clips makes the person feel present between narration segments. ## Verification (never skip) - Play the result — judge by ear and eye, on the real file, before delivering. - For a web bust: open the page in Safari AND Chrome (the two alpha codecs differ); check the hair edge and the bottom fade. - Cost sanity: announce estimated credits BEFORE generating (seconds × 48 Magic Hour credits). Spend one cheap test call on any NEW platform before promising anything (the wallet trap). ## Proven values (Valentina, 31-07-2026 — start from these) - Green portrait via Gemini `gemini-3-pro-image`: ask for "perfectly uniform, flat, evenly lit chroma-key green screen (solid green #00B140)" + "keep the person EXACTLY as she is". The model returns a slightly different green — **sample a corner pixel of the GENERATED VIDEO** before keying (Valentina's came out `#00a547`). - Chromakey: `chromakey=0x00a547:0.14:0.08,despill=type=green` (similarity 0.14, blend 0.08) **+ two alpha erosions** to shave the green rim: `...,format=yuva420p,split[k1][k2];[k2]alphaextract,erosion,erosion[al];[k1][al]alphamerge`. Two traps proven 31-07-2026: raising similarity above ~0.14 makes DARK HAIR semi-transparent (reads grey on light pages), and any explicit `despill mix` greys dark colors — the rim is fixed by erosion, never by stronger keying. - Embed trap: a **muted, hidden video refuses to play** ("video-only background media paused to save power") — segments must play unmuted (they carry the voice); only the idle loop is muted, and it must stay visible. - Safari file: `-c:v hevc_videotoolbox -allow_sw 1 -alpha_quality 0.75 -q:v 55 -vtag hvc1`, audio mapped from the original green mp4 (`-map 0:v -map 1:a`). ffprobe reports `yuv420p` even when alpha is present — verify by eye in Safari, not by probe. - Poll loops: wrap every status call in try/except and keep polling — a transient 502 on a poll killed a healthy render once. - ~48 credits/second. **Check the balance BEFORE submitting a batch**: a mid-batch 402 leaves half the jobs unsubmitted (submit cheapest-first, or compare total seconds × 48 against the balance shown in any 402 message). ## Known traps - **The wallet trap** — website subscription credits ≠ API credits everywhere except Magic Hour. Test with one cheap call first. - ElevenLabs `speed` unreliable at small steps; choose by ear. - An approved render is a specific FILE — never re-roll approved segments. - Long audio: cut the script, generate per segment; never one 3-minute take. - Chroma green clothing/eyes will key out — check the portrait before generating 8 segments. ## The standard finish — every video ends this way, in this order The generators return small frames (Grok 544 square, others near it), and that resolution — not the mouth engine — is what makes a face look soft. Whatever engine made the clip: 1. **Repaint the mouth at native size.** Never upscale first: the repainted mouth would end up softer than the face around it. 2. **Upscale ×4** with Real-ESRGAN, free on your own GPU: `pip install spandrel`, download `RealESRGAN_x4plus.pth` (github.com/xinntao/Real-ESRGAN releases), extract frames with ffmpeg, run the model per frame, reassemble at the same fps. About 9 s a frame on an Apple-silicon Mac; 544 becomes 2176. 3. **Add a medium grain**, a new pattern every frame, luma only so colour and the green stay clean (numpy: `rng.normal(0, 4.0, frame.shape[:2])` added equally to RGB, then clip). **Never skip it** — without it the skin is too even and reads as retouched. This level was chosen from four versions rendered side by side; render options, never reason about them. 4. **Then** chromakey. Keying last, on sharp frames, gives better edges. **The grain must never be visible to the key.** Grain noise drifts dark, low-saturation pixels (hair) toward the key colour and the chromakey eats them (solid hair pixels measured 27.7% -> 2.6%; the clip reads bright and washed out on a light page). So compute the transparency from the CLEAN upscaled frames and take the colour from the GRAINED frames, merged: `[1:v]chromakey=:0.14:0.08,format=yuva420p,alphaextract,erosion,erosion[al];[0:v]despill=type=green[rgb];[rgb][al]alphamerge[out]` (input 0 = grained, 1 = clean). And sample the key green from CLEAN frames - grain clipping falsifies the sampled value. **Deliver two sizes.** Grain compresses badly: the same clip is 24 MB at 2176 and 0.7 MB at 1088. The big one to keep, the small one for the page. **Two explanations that look right and are wrong — they were measured, do not re-argue them.** The upscaler does NOT shift skin colour (a cheek moved R−0.6 G+0.4 B−0.7 out of 255) and does NOT remove micro-texture (it measures higher). It makes skin EVEN, and evenness is what reads as retouched; the grain works by breaking the evenness. ## The birth of a character — do this ONCE, at the best quality available **A character is born once, at the best quality money can buy that day — face AND voice. Everything afterwards is reproduced by the factory for centimes.** The master assets are permanent: the factory copies them forever and can never exceed them. 1. **Survey before buying.** Video engines change every few weeks; check what is best TODAY and its price per second before spending. Judge a master clip on the FACE alone (eyes alive, natural blinks, still camera, chroma background intact) — the mouth gets repainted anyway. 2. **Buy the face once** — one master clip from the green portrait. 3. **Harvest the voice at the same time.** Engines that generate video with native speech give the character a voice that matches the face. Generate ~30-40 s of continuous speech, extract the audio, strip silences, normalise loudness, and **clone it once** into a voice service. From then on the voice is yours: any script, any language, under a centime a line, identical forever — instead of paying per second for a different voice every time. (Clone endpoints commonly cap uploads around 11 MB — send mp3, not raw WAV.) 4. **Then the factory reproduces** — mouth repainted per sentence on a rented GPU, the cloned voice speaking every new line. The character never changes face or voice again.