A person who stands on your page and speaks: face moving, real voice, no frame around them. Start by choosing one of three, because the rest of the work depends on it.
First, once per machine. Your Claude does not know this recipe until you give it to it. The recipe is a small file called make-your-speaker. Install it with one command in your Claude Code:
npx skills add gfrankgva/studio-skills
Or take the file yourself:
download make-your-speaker and put it in ~/.claude/skills/make-your-speaker/. More ways, and what else is in the pack: further down this page.
Then just ask, in your own words: « Make me my own speaker. » It will ask which of the three you want, then take you through it.
Three speakers are built and stored, with their voices: Valentina (Spanish, moving), Zalia (French, still), Grace (English, still). They are shown, at every size, on
the Speakers page.
Tell Claude Code:
put Valentina on this page
Nothing to generate, nothing to pay. If the language or the register does not fit, do not force it — take choice 02.
You are not choosing from a catalogue, you are writing a description. "A Salvadoran woman of about forty, warm, in a light blouse, plain background." Ask for three, look at them side by side, keep one. Nobody real is involved, so there is nothing to clear and nobody to ask.
| Where to make the face | Cost | Good for |
|---|---|---|
| Google AI Studio — aistudio.google.com | free daily quota | The default. Try variants by hand, keep the one you like. |
| The Gemini image API | fractions of a centime | The same, from a session, in batches of three. |
| Krea, Ideogram, Flux | free tiers | A different look when every Gemini face feels alike. |
| generated.photos | paid, licensed | A catalogue you filter by age, gender and ethnicity, when you would rather pick than describe. |
Do not use a stock photo of a real person. Unsplash, Pexels, Freepik, Adobe Stock and Getty licences generally forbid making a model appear to speak, endorse or represent something — which is exactly what a speaker does. A face found through an image search is a lead, never a source.
A historical figure is the exception. A portrait in the public domain can be used directly: search Wikimedia Commons through its API, and search the painter's name as well as the sitter's — that is what surfaces the museum-grade copies. Proven on Rousseau.
Then the master clip is generated from that face on a green background — Grok Imagine, 15 seconds, $1.21, bought once. For an invented person, generating the movement is correct: they have no manner yet, so a model inventing one is exactly what you want.
And the voice? They have none, so you pick one from a library, or clone a voice you have permission to use. The recording step of choice 03 does not apply.
Why filmed and not generated. Four generated masters were bought for one real man, from the same approved portrait, and he rejected all four: too theatrical, then too frozen, then "not natural at all", then "the smile is not nice". A person recognises their own stillness before they recognise their own face, and no written instruction describes it. Rewording is not the lever. Film is.
What only you can do, and it is five minutes. Sixty seconds of video — plain wall, daylight from the front, camera at eye height, just talking and listening the way you normally do. Then three minutes of audio, speaking the way you speak when you are convincing someone. The energy matters more than the microphone: a voice cloned from calm samples sounds like a bored narrator forever, and no setting fixes it afterwards.
Consent, if the face is a colleague's and not yours. A real person's face and voice need that person's agreement in writing before anything is generated.