# Motion stream: Frank's criteria **Source file.** Frank's view: [criteria.html](criteria.html) ← Motion home: [index.html](index.html) The verdict loop for this stream. Read BEFORE producing any motion page or document. Add a numbered criterion, with the why, every time Frank corrects or approves. 1. **Show, don't label (31-07-2026, motion front door).** A concept card with an abstract label ("A person, melted in" · "Calm UI, glass type") conveys nothing: every concept shown on a page carries a REAL example image (screenshot, thumbnail) so Frank can judge "is this what I want?" on sight. A named site gets its thumbnail + one plain line. **Why:** the purpose of these pages is to understand very quickly and be operational; words that do not serve that purpose are eliminated. 2. **No unexplained jargon, ever (31-07-2026, same verdict).** GSAP, Three.js, WebGL and the like: either the term is worth a one-line plain explanation, or it does not appear. A mention Frank cannot picture is space taken from him. **Why:** "I have no idea what GSAP is, so either you explain if it is worth it, or you don't mention." 3. **Every element earns its place or goes (31-07-2026, same verdict).** A card whose purpose is not clear on sight ("AI video, Higgsfield Hailuo": what is it FOR?) is deleted, not decorated. **Why:** "whatever doesn't convey any clear meaning should be eliminated." 4. **Everything Frank reads in this library is an HTML page (31-07-2026).** Markdown stays only as the machine source behind it; every index and back-link points at the HTML view. **Why:** "I see that you still write documents in md for the UI/UX document. They all should be in HTML." 5. **Sibling controls are twins (31-07-2026, Valentina pills).** Neighboring controls share one component, one container, one baseline; screenshot every UI state and check each pair of neighbors before delivering. Prefer discreet twin icon buttons over text pills. **Why:** "Two different sizes of buttons, not aligned: you should never produce something like this." Global rule: `feedback_html_output_rules.md` §26. 6. **A person in a page introduces herself with a card (31-07-2026, Valentina).** Beside the person, a white card: "Hola, soy [name]", one short question, ONE primary action in the accent color ("Sí, escuchar") and one quiet text alternative ("Conversar"); the card hides while she speaks and returns when she stops. **Why:** "This is the card I like", chosen over icon pairs: with only two actions, plain words beat symbols that need decoding. Live: [azfa.eregistrations.dev/valentina-video](https://azfa.eregistrations.dev/valentina-video). 7. **A card leads to an introduction of its topic, never to a raw source (31-07-2026, the Voice card).** Every card on the Studio front page opens a page that introduces its subject plainly: what we do with it, the real examples, the tools, the method. If that page does not exist, BUILD IT before shipping the card; do not point the card at one example, one leaflet or one section and hope the reader infers the topic. **Why:** the Voice card showed a French leaflet about trader status and opened it. "One expects an introduction to the topic, very simply, what do we do with voice." 8. **The word is speaker; sizes settled (02-08-2026, the Speakers page).** The person on a page is a **speaker**, never "talking bust", never a genre term. Its forms are **sizes**, never "rungs": **face** (to the bottom of the neck) · **half bust** (head and shoulders) · **bust** (to the chest, Valentina today) · **full body** (to the feet). Every size SPEAKS; a still picture is not a size but the separate option before the speaker ("just an image, not speaking"). There are **two kinds of speaker**, always clearly distinguished: the **still speaker** (speaks, does not move, Zalia in the Comores, Grace in Lesotho; a photo + a recorded voice, almost free) and the **moving speaker** (speaks and her face moves, Valentina; the new capability, master clip + centimes per part). The still speaker stays on offer BECAUSE it is cheaper and quicker; the sizes above apply to the moving speaker. **Why:** "Is it a talking bust or a speaker? If it is a speaker, just say speaker." · "I don't know what rungs is." · "A speaker, so far, in the Comoros and Lesotho leaflets was a still speaker. What is new now is that it is a moving, animated speaker. Let's show the distinction clearly." 9. **An illustration shows one thing only (02-08-2026, same verdict).** An image that illustrates the speaker contains ONLY the speaker, no unrelated buttons, controls or badges in the frame, and never a control that no longer exists on the live page. **Why:** "You make people think that you have two buttons… this is not the case. Don't mix the problems." 10. **Studio pages: concise, visual, collapsible (02-08-2026, saved as standing rules at Frank's request).** (a) As concise as possible, cut every sentence that a picture or a line can replace. (b) Show images wherever possible instead of text. (c) Every section is collapsible (`
`). **Why:** "Less text and more visual presentation" · "Make sections collapsible… Same for the other sections." 11. **Features are told apart as presenting vs talking (02-08-2026, the Speakers page).** The speaker has two features: **presenting** (she reads the page part by part, works today) and **talking** (a real conversation, not ready yet; its button stays off until it is). A page about the speaker shows the real button, and the real icons that appear after the click. **Why:** "Stay for the time being on the presenter. Talking, we don't know how to do it yet." 12. **A Studio page is a cookbook (02-08-2026, the motion page).** Titles name dishes ("Make pages move", "Use video, not code"), sections are recipes (what you get · what you need · the steps · the cost · where it is proven), and no words about ourselves: no "doctrine", no "manifesto" register, no essay sentences. **Why:** "This is a practical manual, not a report nor a manifest… our model is a cook book, should be all very concrete." 13. **A cloned voice is not a copy of the master; the master voice stays the master (05-08-2026, Valentina's voice).** Frank compared Grok's own native speech against a cloned-and-corrected version of it, twice, and chose the original both times: *"the original is clearer, all the others are a bit étouffé"*, then *"voz A is much better."* The correction was real and measured, the clone had an 11 Hz pitch offset and a 7.6 dB hole in the 2–4 kHz presence band, closed to 1.6 dB across the whole spectrum, and it still lost. **So: measure and correct every clone (it is cheap and it helps), but do not claim the factory reproduces the master's voice until Frank says it does.** The open question for the next session is whether a professional-grade clone (about 30 minutes of ONE stable voice) closes the gap, or whether hero lines must simply be generated at the master engine. **Why:** the ear decides, and the ear has now decided twice against the clone. 14. **Upscale a person, then put the grain back, Frank's recipe (05-08-2026, the upscale test).** Real-ESRGAN at 4× makes edges real again (lip border, hair strands, collar) but leaves the skin too even, which reads as retouched: *"it looks less natural because it removes the imperfections."* **Adding a medium film grain on top is what he chose**, grain new on every frame, luma only so the colour and the green stay clean, sigma about 4 on 0–255. Order in the factory, fixed: repaint the mouth at native size → upscale the finished clip → grain → chromakey last. **Two things measured, because both are natural explanations that turn out to be wrong.** (a) **The colour does not shift.** He read the upscaled skin as a different colour; the cheek measures R−0.6 G+0.4 B−0.7 out of 255, which is nothing. The eye reads *smoothness* as colour. (b) **The upscaler does not remove micro-texture**, it measures higher, not lower (2.38 against 1.28). It REGULARISES: irregular blotches become an even surface, and evenness is what reads as retouched. Grain works not by adding texture back but by breaking that evenness. **The cost to know:** grain is expensive to compress, the same clip is 24 MB at 2176 and 0.7 MB at 1088, so deliver two sizes, big to keep, small for pages. **Why:** I argued from the measurements that grain could not help, and his eye said otherwise on the rendered clip. Render the options, never reason about them. 15. **For an invented character you generate the movement; for a real person you capture it (05-08-2026, four rejected masters).** Four Grok masters were bought from the same approved portrait of Frank, $4.84 in all, and he rejected every one: *"totally exaggerated, this is not me at all"* · too frozen, *"I'm not moving, so make it natural"* · *"not natural at all"* · *"the smile is not nice"*. The failures are not a wording problem. Prompting swings the model between theatrical and dead and never lands on **that particular person's** way of being still, because the model is inventing a personality for someone who already has one. **So: a generated master is right for Valentina, who does not exist. For a real person, start from film of them, their own existing recording, or sixty seconds on a phone against a plain wall, and let the new pattern fix the resolution.** The mouth is repainted afterwards either way. **Why:** a person recognises their own stillness before they recognise their own face, and no prompt describes it. 16. **A capability anyone can use needs a door on the front page (05-08-2026, the Studio).** The wizard, the page explaining it and its top-menu links all existed, and Frank still could not see it: *"I don't see prominently in the studio index page either 'make your own speaker'. It should be linked to the homepage. People should see it."* **A top-menu link is not visibility. A card on the front door is.** And the page that lists a capability must first answer the reader's real question, *take one that exists, or make my own?*, before it explains anything else. **Why:** the Studio is read by people who did not build it. What is one click away for us is invisible to them. 17. **Three ways to a speaker, and they are the page's spine (05-08-2026, Make your own speaker).** Anyone arriving must choose one of three BEFORE reading anything else, and the page must be built around that choice, not mention it: **take one that exists** (free, nothing to supply) · **make a NEW speaker who is not you**, a person who does not exist, described by nationality, age, gender, dress; no photo, no recording, nobody to ask (~$1.50) · **make an avatar of yourself**, one photo and five minutes of your own recording (~$3). Frank had to say it twice before the page was restructured: naming the routes inside a section is not the same as making them the spine. **A stock photo of a real person is not one of the three.** **Say plainly what cannot be done:** Unsplash, Pexels, Freepik, Adobe Stock and Getty licences generally forbid making a real model appear to speak or endorse, which is exactly what a speaker does; a face found by image search is a lead, never a source. A historical figure in the public domain is the exception, search Wikimedia Commons by API, and search the painter's name too. **Why:** *"What is not clear is that you can choose somebody else's avatar, and then you need to be given a list of places where you can find pictures according to the nationality and the gender."* People will otherwise take the first stock photo they find. 18. **A section is a door, not a container (05-08-2026, the Speakers page).** When a capability has several routes, the section's front page IS the choice: a question, then one card per route, each opening its own page, and everything common to all routes stays behind on the front page, once. Frank's own words: *"We should start from the current speaker page, and then you have the three choices. This brings you to other pages, and we leave on the speaker page what is common to all types of speakers: the size, the cost, and all."* Two corollaries he had to state separately: **the choice must be visually big** (a table of rows reads as reference, three large cards read as a decision), and **a screenshot standing in for a live thing is deleted once the live thing is on the page**. **Why:** a reader who has not chosen yet cannot use anything else on the page. 19. **The grain must never be visible to the key (05-08-2026, Valentina's rebuilt part).** Frank rejected the first part rebuilt with the standard finish: *"it's kind of bright, missing real colors, so today is much better."* Measured cause (ibero session): the grain, not the network, her cheek was pixel-identical, but grain noise on dark, low-saturation pixels (hair, jacket) drifts them toward the key colour, so `chromakey` reads hair as background: mean alpha 236→198, fully solid hair pixels 27.7%→2.6%. Against a light page that reads as bright and washed out. **The fix keeps criterion 14's order and splits the inputs: alpha from the CLEAN upscaled frames, colour from the GRAINED frames, merged** (`[1:v]chromakey…,alphaextract,erosion,erosion[al];[0:v]despill[rgb];[rgb][al]alphamerge`). Corollary: **sample the key green from CLEAN frames**, grain clipping at zero lifts near-zero channels and reads `0x00a547` as `0x09a34b`, wrongly. **Why:** his eye caught in seconds what the recipe's wording would have shipped on every part. 20. **The master ships as born; only the key may touch it (06-08-2026, the network pair, then the bar).** Two verdicts the same evening. First, the same 8 s clip through the identical finish twice, quick and strong enlarging network, side by side, both rejected: *"Both are bad. They're not of good quality, so we need to think of something else."* Then his eye ranked the Desktop clips, in two passes that end in one place: the processed native part first read as *"good finally"* next to the pair, but against the raw clips he corrected it, *"this Saludo Nativo is not good… the good ones"* are the untouched 544 px Grok originals (the birthday clip, the master): *"It's so wonderful to see."* He himself noted the best one is *"much smaller than the previous one."* Even light processing (×2 enlarge + grain + a careless key) had already killed it; only the raw clip holds the wonder. The ranking is exactly inverse to the processing: measured, the native master carries MORE micro-texture (HF 8.2) than any processed output, the "real colors" are born in the clip, and each stage after birth (enlarger smoothing, synthetic grain replacing organic texture) subtracts what his eye misses. **So: a speaker's part is BORN native, the engine speaks the exact line, at the master's own size, and the only step between birth and the page is the key (split inputs per criterion 19, at native size). No enlarger, no added grain on a native clip. Display the native size; never inflate it, the page shows her smaller than 544 anyway.** The repaint was then convicted by its isolating test the same night: born-1080 base, mouth repainted, no enlarger, *"number 4 is bad, mouth makes strange moves"* (the base was a SPEAKING Veo master, so the base's own jaw motion fought the repainted mouth; a silent-base variant is untested, but the native road won and closes the question). Criterion 14 still governs the one case where an enlarger is unavoidable. Confirmed the same night on the keyed twin: transparency alone preserves the wonder, the raw ABC intro and its key-only version were both approved on sight (*"yes 3 is good and 4 also"*). **Why:** the pair was built to pick a network, and his eye disqualified the chain instead. 21. **A part is born in takes, at a speaking pace, never crammed into one clip (07-08-2026, the ABC narration).** Frank on the first native set: *"Valentina talked too fast, in particular in the saludo."* The cause was arithmetic, not delivery: the engine caps a clip at 15 seconds (it refuses 20 outright), so 49 words became 3.3 words per second, and asking the model to "speak slowly" inside a full clip changes nothing, it still fits the text in. **The fix is to cut the part at its own natural pauses and buy one take per clause, each with room to breathe; measured after, 1.9 to 2.1 words per second, approved on the ear.** Three rules make the takes hold together, each earned from a rejection: the prompt must give a **runway** (*"silent for a brief moment, takes a small calm breath"*) or the first word is clipped, *"3 starts with missing word"*; the raw takes must be **trimmed at both ends** (0.3 s before, 0.4 s after) or the join carries dead air, *"there's a silence, and then it starts with an incomplete word"*; and every take must be **verified before it counts**, transcribed and diffed against the script (one take said *"la convierta"* for *"la convierte"*), onset ≥ 0.25 s, pace ≤ 2.3 words per second. **Why:** the words were right and the clip was still wrong; pace is a measurable property of a clip, so it is checked like any other, not hoped for. 22. **Solid inside, soft at the edge, and served as video (11-08-2026, Valentina on the ABC note, on a phone).** Frank, on his phone: *"Valentina is shown on a kind of background."* There was no background: her keyed clips were only **86% opaque inside the silhouette** (median alpha 219 of 255), so the note's own lines of text read straight through her jacket and hair. Her AZFA clips measure 95% (median 243) and never drew the complaint. **On a desktop the flaw hides**, she stands over an empty white margin, and nothing shows through; on a phone the text runs right behind her. **So: measure the interior alpha of every keyed part before it ships** (decode one frame, take the median alpha of the pixels above 128; it must be ≥ 250), and when it is low, lift it, anything under 24 to zero, 200 and above to 255, straight ramp between, which keeps the anti-aliased edge (recipe: `3 - Projects/abc-colombia/valentina/opacar-clips.sh`). **And check how the server types the file:** the ABC portal answered `application/octet-stream` for `.mov` and `audio/webm` for `.webm`; Safari refuses a `