# Birth a section: one clip, one portrait, one voice **What this is.** The route Frank approved on 28-08-2026, after the blind test on Valentina's portada: a page section is born WHOLE from a locked portrait plus its approved audio, instead of being ordered as separate takes and stitched. No fifteen-second cap, no seams between takes, the same face from start to end, retakes near zero because the words and the voice are decided before any picture exists. **Run it.** ``` python3 nacer.py --portrait retrato-verde.png --audio parte-01.wav parte-02.wav ``` Optional: `--scale 3.0` (mouth follows the sound more strongly; 2.5 is softer, both accepted by Frank), `--steps`, `--match` (LUFS), `--out`. **What it does.** Rents an A100 80GB, installs the engine with every fix listed below, downloads its weights, births one clip per audio file, levels the voice, brings everything home and returns the machine. It terminates the machine on success, on failure and on Ctrl-C. **Measured (27/28-08-2026, Valentina's portada).** Face identity and framing hold. Mouth opening 18.0 % of the face at scale 2.5 and 19.2 % at 3.0, against her native 13.9 %; Frank accepted both, so the house band was calibrated on native clips only and reads high here. The engine returns the voice about 14 decibels down: the tool puts it back. **Cost.** Corrected 31-08-2026: the machine is an A100 80GB, $1.39 (PCIe) to $1.59 (SXM4) an hour on [runpod.io/pricing](https://www.runpod.io/pricing), not $2.5. At roughly 2.5 hours per 6.4-second clip un-tuned, a published minute (9.4 clips) costs **$33 to $37**, not "a few dollars": that line was arithmetic that never got checked against the real per-hour rate. Tuning (TeaCache, fewer steps, serverless) is untested; do not promise "under a dollar" until it is measured. **The seven faults that cost a night, all fixed inside `remote-base.sh`.** Do not "simplify" any of these lines. | Fault | The line that fixes it | |---|---| | Invented image tag | `runpod/pytorch:1.1.0-cu1281-torch260-ubuntu2204`, read off Docker Hub | | `huggingface-cli` is a deprecated no-op | downloads through python `snapshot_download` | | One bad pin kills a batch install | per-line requirement install, then an import preflight BEFORE the 80 GB download | | New transformers refuses the legacy wav2vec line | guarded patch plus an audio-encoder smoke test | | transformers 4.49 pulls hub 1.x, which breaks diffusers | the era trio pinned together | | A40 container memory kills the 14B load | A100 80GB only | | CLIP hard-asserts flash-attention | prebuilt wheel, cxx11abiTRUE, `--force-reinstall` | ## Renting without burning money (01-09-2026, after Frank: "trop d'argent dépensé en échecs") **Where the money went**, measured on the round that proved the route: $7.50 of real computing (two clips at 40 steps), $1.40 across six failed setups (they died in minutes, the preflight did its work), and **$2.25 on one machine nobody was watching**, orphaned by a hung connection until another session found it. So the waste was not the crashes; it was an unattended machine, and paying full price to discover a fault. **Five guards, now inside the tool.** 1. **Nothing is rented before three answers**: no machine of ours is already renting, the balance covers the ceiling, and the ceiling is stated. A run that dies for lack of credit has paid for nothing. 2. **A ceiling in dollars** (`--max-usd`, $12 by default): the machine is returned when the run has cost that much, whatever it is doing. 3. **A two-step clip on one second of sound, before the real batch.** Every fault we ever hit shows up there, for a few minutes of machine instead of hours. 4. **The machine is returned on every path**: success, failure, Ctrl-C, and a hung connection can no longer escape the loop (that is exactly what orphaned the $2.25 machine). 5. **A sweeper**: `python3 nacer.py --sweep`, on its own, kills every machine of ours. Run it when you start and when you finish, and any time you are unsure. It is free and it takes a second. **What would remove most of the remaining cost**, not yet built: bake the environment and the weights into an image or a network volume, so no round ever pays again for the same 80 GB download and the same twenty minutes of installing. **Corrected 01-09-2026 by measurement, not by argument.** The 14B model **loads on a 48 GB card**, unquantised (audit round, pod `kfa3q47kh4o8rf`, 91 minutes, about $0.67): the card served is an **L40S billing $0.99/h**, read off the running machine, so a published minute is **about $23**, against $33 to $37 on an A100. Two corrections come with it: the August "dies on container memory" was **InfiniteTalk**, a different engine, and the round that died on `assert FLASH_ATTN_2_AVAILABLE` was running an older script that **never installed flash-attention** (MultiTalk sets that flag False only on `ModuleNotFoundError`, so the package was absent, not incompatible with the card). Which of the five cards was assigned is not recorded, so the rule is **"a 48 GB card from the list"**, never a named one. What is still unmeasured: a full generation completing on that card, one round of about $0.70. **Two corrections from the Valentina seat, 01-09-2026** (hers, measured on her own round, mine to fold in properly when her figures land): - **The disk is billed, and 250 GB was my guess, not a measurement.** The weights need about 110 GB, so the default is now `--disk 140`. She reports the oversized disk eating a large share of a short run's cost. - **A cheaper card may be enough.** The A40 failure of 27-08 was container memory, not the GPU. She reports the model loading on a cheap card, with one piece of software to recompile for it. `--gpu` now takes any card list, cheapest first. Her three pages fall from about $200 to about $60 on that path. **Then.** Measure the mouth (`3 - Projects/abc-colombia/valentina/medir-boca.py`), key to transparent twins (webm VP9 alpha + mov HEVC alpha, never ProRes), wire with [site-speaker](/Users/unctad/.claude/skills/site-speaker/SKILL.md), and check the live bytes, never the local copy. **Doctrine**: [Speakers](../../index.html) · topic `5 - Handovers/topics/speakers.md` · assessment `8 - Plans/speaker-videos-cheap-and-consistent.html`.