--- name: build-your-own-factory description: "Set up your OWN talking-video factory: open models (MuseTalk 1.5 + LivePortrait) on a rented serverless GPU, one paid master clip per person, then every new sentence for centimes. TRIGGER when the user says 'set up my own talking-video factory', 'make avatar video cheap', 'run MuseTalk on a rented GPU', 'stop paying per minute for talking avatars'. NOT for producing a single clip (that is /photo-avatar-video) and NOT for putting the speaker on a web page (that is /site-speaker) — this skill builds the INFRASTRUCTURE both of those can then use." --- *From Frank's Studio — https://smartrules.ai/studio · install: npx skills add gfrankgva/studio-skills* # Build your own talking-video factory **What you build.** A serverless render farm for talking-bust video, owned by you. One RunPod endpoint runs a ready-made public worker image; you pay for GPU seconds only while a job runs. Per person you buy ONE good paid clip (the "master", ~1.20 EUR, its living eyes and expression are the whole point), and from then on every new sentence is made by repainting only the mouth of that master to match new audio: about 10 seconds of GPU and well under a centime per sentence, at the paid clip's visual quality. The chain, drawn: [the factory](https://smartrules.ai/studio/people/factory.html). ## What you need - **A RunPod account + API key** ([runpod.io](https://www.runpod.io), a few dollars of credit). A rented RTX 4090 starts around 0.34 USD/h; the serverless tier used here bills ~0.00031 USD per second, only while rendering. - **One master clip per person** (Magic Hour or similar, ~1.20 EUR for 25–30 s): the person on a flat green background, calm, mouth moving (the words do not matter — they get repainted). Also keep a **4-second silent idle** (mouth closed, one blink) for the resting state. - **ffmpeg** locally, for cutting out the green. Nothing to build or train: the worker image is public and the model weights download themselves at first cold start. ## Set it up once The public image is **`ghcr.io/gfrankgva/musetalk-worker-pub:latest`** (source: [github.com/gfrankgva/musetalk-worker](https://github.com/gfrankgva/musetalk-worker), built by GitHub Actions). One worker, four modes: `idle` (photo → calm idle loop), `prepare` (base video → cached face boxes/latents/masks), `talk` (base + audio → green mp4), `diag` (health check). Via the RunPod REST API (or do the same in the console: Serverless → New Endpoint → Docker Image): ```bash # 1. Template curl -s -X POST https://rest.runpod.io/v1/templates \ -H "Authorization: Bearer " -H "Content-Type: application/json" \ -d '{"name":"musetalk-worker","imageName":"ghcr.io/gfrankgva/musetalk-worker-pub:latest","isServerless":true,"containerDiskInGb":40}' # 2. Endpoint (templateId from the answer above) curl -s -X POST https://rest.runpod.io/v1/endpoints \ -H "Authorization: Bearer " -H "Content-Type: application/json" \ -d '{"name":"musetalk","templateId":"","gpuTypeIds":["NVIDIA GeForce RTX 4090"],"workersMin":0,"workersMax":1,"idleTimeout":180,"scalerType":"QUEUE_DELAY","scalerValue":4}' ``` The settings that matter, whichever route you take: - **GPU: RTX 4090 24 GB** (fallback: A40 48 GB) · container disk **40 GB** · no network volume needed. - **Workers min 0, max 1** — a batch must not be split across workers with cold caches. - **Idle timeout 180 s** — one warm worker then serves a whole batch of sentences from its local cache. That is up to 3 minutes of idle billing (~$0.06) against re-preparing (~$0.03–0.07) on every cold clip; it wins as soon as you render more than one sentence. Note the endpoint id from the answer — `` below. Then verify for pennies: ```bash curl -s -X POST https://api.runpod.ai/v2//runsync \ -H "Authorization: Bearer " -H "Content-Type: application/json" \ -d '{"input":{"mode":"diag"}}' ``` `diag` returns the build sha, the CUDA state of both venvs and a real model import — seconds of GPU, no render. The first call also triggers the one-time cold start (see traps). ## Render a sentence Every call is `POST https://api.runpod.ai/v2//runsync` with the same auth header (use `/run` + poll `/status/` for long jobs). Inputs go as `*_url` or `*_base64`; use URLs for anything big (RunPod caps payloads around 10–20 MB). Output comes back as `video_base64` (or a URL if you pass an `r2` config). **The workhorse call — repaint the master's mouth with new audio:** ```json {"input": {"mode": "talk", "base_video_url": "https://.../master-green.mp4", "audio_url": "https://.../sentence.mp3", "bbox_shift": 0}} ``` Out: a green-background mp4, same length as the audio. Eyes, blinks, head and hair are copied frame for frame from the master; only the mouth is new. `bbox_shift` is the mouth-openness knob, per portrait: lips barely move → try 5; too much teeth → negative. **The cache pattern — prepare once, ~10 s per sentence after.** The first `talk` against a new base video prepares it inline (~2 min of GPU) and caches the result on the worker. Sentences rendered back to back then take ~10 s each on the warm worker. To keep the cache beyond the worker's life, prepare explicitly and hold the tarball yourself: ```json {"input": {"mode": "prepare", "base_video_url": "https://.../master-green.mp4", "return_cache": true}} ``` Keep the returned `cache_base64` (~16 MB for a 28 s base) and send it back on later `talk` calls as `cache_base64` — a cold worker then skips preparation entirely. The cache key is a hash of the base video + parameters, so a re-encoded base is a new cache. **From a photo only (no master yet):** `{"input":{"mode":"idle","image_url":"https://.../portrait-green.jpg","idle_seconds":10}}` returns a calm animated loop you can use as a base — good for a quick test, but see the last trap before shipping it. ## The standard finish — every video ends this way, in this order The generators return small frames (Grok 544 square, others near it), and that resolution — not the mouth engine — is what makes a face look soft. Whatever engine made the clip, it ends with these steps IN THIS ORDER: 1. **Repaint the mouth at native size** (the render step above). Never upscale first, or the repainted mouth ends up softer than the face around it. 2. **Upscale ×4** with Real-ESRGAN, free on your own GPU: `pip install spandrel`, download `RealESRGAN_x4plus.pth` (github.com/xinntao/Real-ESRGAN releases), then per frame: extract with ffmpeg, run the model, reassemble at the same fps. About 9 s a frame on an Apple-silicon Mac (`mps`); 544 becomes 2176. 3. **Add a medium grain**, a new pattern every frame, luma only so colour and the green stay clean (numpy: `rng.normal(0, 4.0, frame.shape[:2])` added equally to RGB, then clip). **Never skip it** — without it the skin is too even and reads as retouched. 4. **Only then cut out the green** (next section). Keying last, on sharp frames, gives better edges. **The grain must never be visible to the key.** Grain noise drifts dark, low-saturation pixels (hair) toward the key colour and the chromakey eats them (solid hair pixels measured 27.7% -> 2.6%; the clip reads bright and washed out on a light page). So compute the transparency from the CLEAN upscaled frames and take the colour from the GRAINED frames, merged: `[1:v]chromakey=:0.14:0.08,format=yuva420p,alphaextract,erosion,erosion[al];[0:v]despill=type=green[rgb];[rgb][al]alphamerge[out]` (input 0 = grained, 1 = clean). And sample the key green from CLEAN frames - grain clipping falsifies the sampled value. **Deliver two sizes.** Grain compresses badly: the same clip is 24 MB at 2176 and 0.7 MB at 1088. The big one to keep, the small one for pages. **Two explanations that look right and are wrong — they were measured, do not re-argue them.** The upscaler does NOT shift skin colour (a cheek moved R−0.6 G+0.4 B−0.7 out of 255) and does NOT remove micro-texture (it measures higher). It makes skin EVEN, and evenness is what reads as retouched; the grain works by breaking the evenness. ## Cut out the green This is the LAST step of the standard finish above — key the upscaled, grained clip, never the raw render. The factory returns the bust on its original green; keying happens on your machine. Two verified quality wins are baked into this recipe: convert to **`format=yuva444p` BEFORE the chromakey** (half-resolution colour otherwise blocks up hair edges) and **feather the matte with `gblur` on the alpha plane** (`planes=8`) instead of eroding it. ```bash # Sample the ACTUAL green from the rendered clip first (it shifts from the source photo; e.g. 0x00a547) KEY="chromakey=0x00a547:0.12:0.05,despill=type=green,gblur=sigma=1.5:planes=8" # WebM with alpha (Chrome/Firefox/Edge) ffmpeg -y -i sentence.mp4 -vf "format=yuva444p,$KEY" \ -c:v libvpx-vp9 -pix_fmt yuva420p -b:v 0 -crf 30 -auto-alt-ref 0 -c:a libopus sentence.webm # HEVC .mov with alpha (Safari; encoder available on macOS) ffmpeg -y -i sentence.mp4 -vf "format=yuva444p,$KEY" \ -c:v hevc_videotoolbox -allow_sw 1 -alpha_quality 0.75 -vtag hvc1 -c:a aac sentence.mov ``` **Always produce the pair.** A page needs both: `.webm` for Chrome/Firefox/Edge, `.mov` for Safari. A missing twin fails silently, on one browser only. ## The measured numbers Measured on a live 4090 endpoint (serverless flex, $0.00031/s), 640×640 base, ~27 s sentences: | What | Time | Cost | |---|---|---| | Master clip (paid), once per person | — | ~1.20 EUR | | Prepare, once per master | ~2 min GPU (80–218 s, CPU-host dependent) | $0.03–0.07 | | **One sentence, warm worker** | **~10 s GPU (11.5–18.7 s)** | **$0.004–0.006** | | First sentence on a cold worker, cache sent back | ~36 s | $0.011 | | Cold start (image known to the machine) | 130–240 s | unbilled | | A whole 8-sentence narration | prepare + 8 clips | **$0.06–0.10** | | Per minute of finished video, bust prepared | ~25 s GPU | ~$0.008 | ## The traps - **Cold start is 130–240 s** (the weights, ~5 GB, download on each cold start; unbilled). The very first pull of the image on a fresh machine is much longer. A "no GPU available" timeout right after setup is usually the cold start, not a missing GPU: warm the endpoint with `{"input":{"mode":"diag"}}` and call again. Attaching a small network volume makes the weight download one-time — at the price of pinning the endpoint to one datacenter. - **Keep the master base short: 25–30 s.** Preparation cost and cache size are per FRAME of base video, and the base is only ever looped — a longer master buys nothing but a longer prepare. Under ~25 s the loop repeat starts to show in long sentences. - **Never raise chromakey similarity above ~0.14** and never set an explicit despill mix: dark hair goes semi-transparent and reads grey on light pages. A green rim is fixed by the alpha feather, never by keying harder. - **The factory alone invents flat eyes.** From a photo, `idle` mode produces a competent but lifeless base — the eyes do not live. Always repaint a paid master clip: the factory's whole economy is paid quality once, factory prices forever. - **Pin, then prove.** Point the template at an immutable image tag (`:`, the Actions build pushes both), never trust `:latest`; every job answer carries `build_sha`. A warm worker keeps serving the old image after you repoint — set workers max to 0, let the pool drain, set it back. ## The master clip, since 03-08-2026: Veo 3.1 on your own Google key The step that used to be paid is now centimes on the same Gemini key that makes the images. Image-to-video from the person's green-screen portrait: 8 s, 720p, living eyes, natural blinks, camera static, ~40 s of waiting. - Model `veo-3.1-generate-preview` (also `-fast-`, `-lite-`), method `predictLongRunning` on `https://generativelanguage.googleapis.com/v1beta`. Check `GET /v1beta/models?key=…` before assuming a video capability is missing. - Body: `instances[0] = {prompt, image:{bytesBase64Encoded, mimeType}}`, `parameters = {aspectRatio:"16:9", resolution:"720p", durationSeconds:8, personGeneration:"allow_adult"}`. - Prompt shape that worked: speaks warmly to camera, lips move continuously, eyes alive and blinking, head barely moves, camera completely static, flat green background unchanged, photorealistic. - Two traps: it adds black bars around the portrait (crop with ffmpeg `cropdetect`); and the green-portrait prompt must NAME the garment explicitly or the image model silently repaints the clothes. - What she says in the master does not matter — the mouth is repainted per sentence afterwards.