--- name: make-your-speaker description: "Take someone from one photo to their own speaker: a face that moves, their own voice, in four sizes (face, half bust, bust, full body), ready for any web page or video. A guided wizard for a person who has never done this — it asks, it shows options, it says what each step costs before spending. TRIGGER when someone says 'I want my own avatar', 'make me a speaker', 'I want my face talking on our site', 'help my colleague make an avatar'. NOT for putting an EXISTING speaker on a page (that is /site-speaker), NOT for a single talking clip from a picture (that is /photo-avatar-video)." --- *From Frank's Studio — https://smartrules.ai/studio · install: npx skills add gfrankgva/studio-skills · the guided page: https://studio.smartrules.ai/people/make-your-own.html* # Make your own speaker **What this produces.** A person hands you one photo and about five minutes of their own recording. They leave with their face moving and speaking in their own voice, cut out of its background, in whichever of the four sizes they chose, ready to drop onto a page or into a video. **About three dollars, once. Every line they say afterwards costs under a centime.** **You are a wizard here, not a tool.** The person has never done this. Ask in one block, not a drip. **Show options whenever there is something to look at** — never describe a look in words when you can render it. Say what a step costs BEFORE spending. --- ## Phase 0 — which of the three roads, then consent **A speaker is reached by three roads, and the person must name theirs before anything else:** 1. **Take one that exists** (Valentina, Zalia, Grace…) — free, nothing to supply. That is NOT this skill: stop here and use `/site-speaker` to put her on the page. 2. **Make a NEW speaker who is not you** — an invented person, described by nationality, age, gender, dress; no photo, no recording (~$1.50). 3. **Make an avatar of YOURSELF** (or a colleague) — one photo and five minutes of their recording (~$3). If they did not say which, ask this first, as the very first question. Roads 2 and 3 continue below. **Consent next.** A real person's face and voice need that person's agreement before anything is generated. If the photo is not of the person in front of you, stop and get it in writing. **The road determines everything in the middle of this skill** — road 3 is a real person, road 2 an invented character: | | **A real person** (you, a colleague, a minister) | **An invented character** (a guide, a mascot, a persona) | |---|---|---| | Where the movement comes from | **Film them.** Their own recording, or 60 s on a phone. | **Generate it.** Grok Imagine from the green portrait. | | Why | A person recognises their own stillness before they recognise their own face, and no prompt describes it. Four paid attempts to prompt one real man's manner all failed. | The character has no manner yet. The model inventing one is exactly what you want. | Everything else — portrait, voice, finish, delivery — is the same for both. **Then the short block:** their name and how they want to be addressed · the language(s) they will speak · formal or warm · where the speaker will appear (a page, a video, both). --- ## Phase 1 — the photo Ask for **one** good photo: face lit from the front, eyes open, no sunglasses, shoulders inside the frame, as high resolution as they have. A photo from a phone beats a crop from a video every time — video frames are small and soft, and everything downstream inherits that. --- ## Phase 2 — the four sizes, on green Generate their portrait on flat chroma green in **all four sizes**, from the same photo, with an image model (Gemini `gemini-3-pro-image`, or the same edit free in Google AI Studio): **face** (to the bottom of the neck) · **half bust** (head and shoulders) · **bust** (to the chest) · **full body** (to the feet) These are the only names for the sizes. Never pixels, never invented words. **The prompt must say two things, twice:** 1. Replace the **ENTIRE** background with a uniform, flat, evenly lit chroma green `#00B140` — no gradient, no shadow, no texture. 2. Keep the person **EXACTLY** as they are — same face, same expression, same hair strand for strand, same skin tone, same clothes, same head angle, same lighting. Do not restyle, beautify, slim or rejuvenate. Say it once and the model repaints them. A terracotta shirt came back as a mustard knit; name the garment explicitly. **Then show all four side by side, at the size they will really appear on the page, and let them choose.** Write the choice down. --- ## Phase 3 — the skin Nearly every portrait comes back a little too red. **Render three levels of red removed** — 12%, 15%, 22% — on their chosen size, and let them pick. Our first build chose 15%. Separate the green from the person first, and touch **only** the skin: the chroma must stay perfect for the cut-out. ```python R, G, B = a[:,:,0], a[:,:,1], a[:,:,2] green = (G > R*1.25) & (G > B*1.25) skin = (~green) & (R > G) & (G >= B) a[:,:,0] = np.where(skin, R - (R - G) * k, R) # k = 0.12 / 0.15 / 0.22 ``` --- ## Phase 4 — the master clip The master is bought or filmed **ONCE**. Everything afterwards copies it and can never exceed it, so this is the only moment where paying for the best is right. ### If it is a real person — capture the movement 1. **Do they already have a video of themselves?** Any decent recording of them talking to a camera is a master. Its resolution does not matter: Phase 6 fixes that. Use it. 2. **Otherwise, ask for 60 seconds on their phone**: plain wall behind them, daylight from the front, camera at eye height, and they simply talk and listen the way they normally do. A modern phone shoots 4K, which is several times anything a generator returns. 3. Trim to the calmest 15–30 seconds where they are still, facing the camera, not gesturing. **Do not buy a generated master for a real person before trying this.** Four were bought for one man and all four were rejected: too theatrical, then too frozen, then "not natural at all", then "the smile is not nice". ### If it is an invented character — generate it **Grok Imagine**, 15 s from the green portrait, **$1.21 a call** (it does not cap at 8 s). Survey the engines first — they change every few weeks. ``` POST https://api.x.ai/v1/videos/generations {"model":"grok-imagine-video-1.5","duration":15, "image":{"url":""}, "prompt":"..."} → {"request_id": ...}, then poll GET /v1/videos/{id} until status "done" ``` **Prompt for RESTRAINT, not expression.** Almost still · no broad smile · eyebrows fixed · head moving a few degrees · slow natural blinks · the camera does not move · the green stays flat. Asking for warmth, or for someone "speaking about something they care about", produces raised brows and wide eyes. **The mouth is repainted later anyway, so the master only needs living eyes.** **When the right energy is uncertain, generate TWO at different levels and let them choose.** Two calls at once is cheaper than four one at a time. **Strip the invented voice.** Grok gives the character a voice; for a real person it is never used. --- ## Phase 5 — their voice **Ask for about three minutes, recorded on their phone, in a quiet room, speaking WITH ENERGY** — the way they talk when they are convincing someone. **The energy matters more than the microphone.** A clone built from calm samples produces a calm narrator forever. This is the single most common way a speaker ends up lifeless. Clone it once (ElevenLabs `POST /v1/voices/add`, multipart): - **Raw WAV, 44.1 kHz, mono.** Force `-ar 44100` on every input: an ffmpeg `concat` of mixed-rate clips silently upsampled one reference to 192000 Hz and dulled the result. - **`remove_background_noise=false`.** The denoiser eats presence. - Keep it under the **11 MB** upload cap — 20 to 40 seconds is enough for an instant clone, so raw WAV fits. - **One voice only in the reference.** Never mix takes that sound different; the clone averages them into a blur. **Say this plainly to the person: a clone is a near-match, not a copy.** Judged side by side against a real recording, the real one wins. It is good enough for volume — every line, any language, under a centime — and a hero line may be worth recording for real. **If it sounds muffled**, it is measurable, not taste: compare median pitch and six bands of energy against the real voice, then correct in one ffmpeg pass (a wide bell around 3 kHz for presence, `asetrate`+`atempo` for pitch). Detail: `site-speaker` §"Match the clone to the master". **A pronunciation dictionary is mandatory** for any acronym or name. A voice guesses, and its guess drifts between takes. --- ## Phase 6 — the finish. Every video ends this way, in this order 1. **Repaint the mouth at native size.** MuseTalk (cheap, $0.008 a minute) or LatentSync (sharper lips and chin, $0.147 a minute — worth it for a **face**-sized speaker). Never upscale first, or the new mouth ends up softer than the face around it. 2. **Upscale ×4** with Real-ESRGAN — runs on the person's own Mac GPU, about 9 s a frame, free. `spandrel` + `RealESRGAN_x4plus.pth`. 3. **Add a medium grain**, a new pattern on every frame, luma only so colour and the green stay clean: `rng.normal(0, 4.0, frame.shape[:2])` added to RGB. **Never skip it.** Without it the skin is too even and reads as retouched. 4. **Only then remove the green.** Keying last, on sharp frames, gives better edges; keep similarity ≤0.14 and fix any rim by feathering the alpha, never by keying harder. **The grain must never be visible to the key.** Grain noise drifts dark, low-saturation pixels (hair) toward the key colour and the chromakey eats them (solid hair pixels measured 27.7% -> 2.6%; the clip reads bright and washed out on a light page). So compute the transparency from the CLEAN upscaled frames and take the colour from the GRAINED frames, merged: `[1:v]chromakey=:0.14:0.08,format=yuva420p,alphaextract,erosion,erosion[al];[0:v]despill=type=green[rgb];[rgb][al]alphamerge[out]` (input 0 = grained, 1 = clean). And sample the key green from CLEAN frames - grain clipping falsifies the sampled value. **Two explanations that look right and are wrong — do not re-argue them, they were measured.** The upscaler does **not** shift skin colour (a cheek moved R−0.6 G+0.4 B−0.7 out of 255) and does **not** remove micro-texture (it measures higher). It makes skin **even**, and evenness is what reads as retouched. The grain works by breaking the evenness. --- ## Phase 7 — deliver - **Transparent WebM** (VP9, `yuva420p`) for Chrome, Firefox, Edge **and a HEVC .mov** (`hvc1`, alpha) for Safari. Both, always. - **Two sizes: the big one to keep, ~1088 px for pages.** Grain compresses badly — the same clip is 24 MB at 2176 and 0.7 MB at 1088. - **A one-page demo where they see themselves speaking**, on a light background, and open it in their browser. - **Store the masters in a permanent folder with a README saying what each one cost.** They are the origin of every future video of that person. Never leave them in a temporary folder: two sessions once had to rescue a set that was one wipe from being gone. **Verify the transparency before showing it:** decode one frame with `-pix_fmt rgba` and check corners at alpha 0, face at 255, and tens of thousands of semi-transparent edge pixels. Without `-pix_fmt rgba` the alpha reads 255 everywhere and will fool you. --- ## The traps, in one list - A **generated master for a real person** — capture, don't generate. - **Prompting a master for expression** instead of restraint. - **Upscaling before the mouth repaint**, or **keying before the upscale**. - **Skipping the grain.** - **Cloning a voice from calm samples**, or from a mix of takes, or through a denoiser, or via MP3. - **A base64 image to Grok**, or a public URL whose host serves the wrong content-type. - **Masters left in a temporary folder.** - **Describing a look in words** when you could have rendered two versions and asked.