{{SPEAKER}}, the standing brief

What she is, what every take must respect, and what we need back. Written for the assistant that generates her clips. Last revised 21-08-2026.

Read this before generating anything. It replaces re-explaining the job in every message. If a rule here contradicts a message in the chat, the message wins and this page gets corrected.

The chat is disposable, this page is not. When a conversation is closed by mistake or its context expires, open a new one and paste a single line: "Read {{THIS_PAGE_URL}} and follow it. Then rebuild Valentina's identity lock and wait for my first take." Everything you need is here, because you cannot read our disk.

Who she is

{{SPEAKER}} is the speaker of {{PROJECT}}. {{PRONOUN}} stands cut out in the corner of a page, over the text, and reads it part by part in {{LANGUAGE}}, addressing the reader as {{ADDRESS}}.

{{SPEAKER}} was born in Grok Imagine on {{BIRTHDAY}} from one still portrait on flat green. That native clip is the reference every later take is measured against. Live pages: {{PAGES}}. The resting clip we key the identity from is {{IDLE_URL}}.

If this is a new chat, start here

Chats get closed. Nothing is lost, because her identity does not live in the conversation, it lives in a file on this site. Rebuild it before generating anything:

  1. Fetch {{IDLE_URL}}, take its first frame, and flatten it onto solid #00FF00. That JPEG is the identity lock.
  2. Generate every take image to video from that file. Never text-to-image a new woman: the first batch that did it invented a different person and all of it was thrown away.
  3. Show the first take before making the rest, so we can confirm it is her.

If the lock is cleared from your sandbox mid-session, rebuild it from the same URL and carry on; the face stays the same.

Why the green must stay flat

We remove it to make her transparent on the page. A gradient, a shadow cast on the backdrop or a moving background leaves a halo around her. Flat, uniform, edge to edge, and nothing else in frame: no captions, no titles, no logo, no watermark, no music, no second voice.

The house band, measured

WhatWhere it must landWhy that number
The wordsExactly as written, nothing added, no improvising, no translationThe page shows the same sentence; a changed word is a wrong page
Pace1.9 to 2.1 words a second, hard ceiling 2.3The clips the owner approved sit at 2.1; faster reads as a busy mouth
Mouth opening14 to 18 percent of her own face width, peaks at or under 22The approved reference measures {{MOUTH_MEDIAN}} median, {{MOUTH_PEAK}} peak. A mouth a third wider was rejected on sight, and so was one that barely opened
StartAbout one second of silence, mouth closed, then {{PRONOUN_LOWER}} speaksThe assembly needs that beat to join takes without a jump
If she finishes earlySilence, soft smile, mouth closedGiven a longer clip than the sentence needs, the engine invents sounds. It has happened three times
CameraCompletely static, no zoom, no pan, framing heldThe face must keep the same width across every clip on a page
Letters{{LETTERS_RULE}}{{WORD_TRAPS}}

How to order and deliver

What we do with them, so you know what breaks

Each take is trimmed to the end of its speech, keyed against the green, and warped so her face is the same width and the same place in every clip of the page. Then the takes of a page are joined into a single clip with a 0.7 second breath between sentences: we never switch files mid-speech, because the cut is audible. A take that fails any line of the table above is not published, it is asked again.

What we would like back, beyond the files

Say plainly what you can and cannot do, with numbers: longest single take, whether the identity holds across sessions, resolution and frame rate, watermark or not, and what the weekly pool has left. When a take fails our table we will tell you which line and by how much; a short answer on what to change is worth more than a new take generated blind.