← Studio library

Frank's criteria for motion

Reference notes. Source retained; historical claims may need rechecking.

Studio reference illustration

Motion stream: Frank's criteria

Source file. Frank's view: Related Studio referencecriteria.html

← Motion home: Related Studio referenceindex.html

The verdict loop for this stream. Read BEFORE producing any motion page or document. Add a numbered criterion, with the why, every time Frank corrects or approves.

  1. Show, don't label (31-07-2026, motion front door). A concept card with an abstract label ("A person, melted in" · "Calm UI, glass type") conveys nothing: every concept shown on a page carries a REAL example image (screenshot, thumbnail) so Frank can judge "is this what I want?" on sight. A named site gets its thumbnail + one plain line. Why: the purpose of these pages is to understand very quickly and be operational; words that do not serve that purpose are eliminated.
  2. No unexplained jargon, ever (31-07-2026, same verdict). GSAP, Three.js, WebGL and the like: either the term is worth a one-line plain explanation, or it does not appear. A mention Frank cannot picture is space taken from him. Why: "I have no idea what GSAP is, so either you explain if it is worth it, or you don't mention."
  3. Every element earns its place or goes (31-07-2026, same verdict). A card whose purpose is not clear on sight ("AI video, Higgsfield Hailuo": what is it FOR?) is deleted, not decorated. Why: "whatever doesn't convey any clear meaning should be eliminated."
  4. Everything Frank reads in this library is an HTML page (31-07-2026). Markdown stays only as the machine source behind it; every index and back-link points at the HTML view. Why: "I see that you still write documents in md for the UI/UX document. They all should be in HTML."
  5. Sibling controls are twins (31-07-2026, Valentina pills). Neighboring controls share one component, one container, one baseline; screenshot every UI state and check each pair of neighbors before delivering. Prefer discreet twin icon buttons over text pills. Why: "Two different sizes of buttons, not aligned: you should never produce something like this." Global rule: feedback_html_output_rules.md §26.
  6. A person in a page introduces herself with a card (31-07-2026, Valentina). Beside the person, a white card: "Hola, soy [name]", one short question, ONE primary action in the accent color ("Sí, escuchar") and one quiet text alternative ("Conversar"); the card hides while she speaks and returns when she stops. Why: "This is the card I like", chosen over icon pairs: with only two actions, plain words beat symbols that need decoding. Live: Related Studio referenceazfa.eregistrations.dev/valentina-video.
  7. A card leads to an introduction of its topic, never to a raw source (31-07-2026, the Voice card). Every card on the Studio front page opens a page that introduces its subject plainly: what we do with it, the real examples, the tools, the method. If that page does not exist, BUILD IT before shipping the card; do not point the card at one example, one leaflet or one section and hope the reader infers the topic. Why: the Voice card showed a French leaflet about trader status and opened it. "One expects an introduction to the topic, very simply, what do we do with voice."
  8. The word is speaker; sizes settled (02-08-2026, the Speakers page). The person on a page is a speaker, never "talking bust", never a genre term. Its forms are sizes, never "rungs": face (to the bottom of the neck) · half bust (head and shoulders) · bust (to the chest, Valentina today) · full body (to the feet). Every size SPEAKS; a still picture is not a size but the separate option before the speaker ("just an image, not speaking"). There are two kinds of speaker, always clearly distinguished: the still speaker (speaks, does not move, Zalia in the Comores, Grace in Lesotho; a photo + a recorded voice, almost free) and the moving speaker (speaks and her face moves, Valentina; the new capability, master clip + centimes per part). The still speaker stays on offer BECAUSE it is cheaper and quicker; the sizes above apply to the moving speaker. Why: "Is it a talking bust or a speaker? If it is a speaker, just say speaker." · "I don't know what rungs is." · "A speaker, so far, in the Comoros and Lesotho leaflets was a still speaker. What is new now is that it is a moving, animated speaker. Let's show the distinction clearly."
  9. An illustration shows one thing only (02-08-2026, same verdict). An image that illustrates the speaker contains ONLY the speaker, no unrelated buttons, controls or badges in the frame, and never a control that no longer exists on the live page. Why: "You make people think that you have two buttons… this is not the case. Don't mix the problems."
  10. Studio pages: concise, visual, collapsible (02-08-2026, saved as standing rules at Frank's request). (a) As concise as possible, cut every sentence that a picture or a line can replace. (b) Show images wherever possible instead of text. (c) Every section is collapsible (<details open>). Why: "Less text and more visual presentation" · "Make sections collapsible… Same for the other sections."
  11. Features are told apart as presenting vs talking (02-08-2026, the Speakers page). The speaker has two features: presenting (she reads the page part by part, works today) and talking (a real conversation, not ready yet; its button stays off until it is). A page about the speaker shows the real button, and the real icons that appear after the click. Why: "Stay for the time being on the presenter. Talking, we don't know how to do it yet."
  12. A Studio page is a cookbook (02-08-2026, the motion page). Titles name dishes ("Make pages move", "Use video, not code"), sections are recipes (what you get · what you need · the steps · the cost · where it is proven), and no words about ourselves: no "doctrine", no "manifesto" register, no essay sentences. Why: "This is a practical manual, not a report nor a manifest… our model is a cook book, should be all very concrete."
  13. A cloned voice is not a copy of the master; the master voice stays the master (05-08-2026, Valentina's voice). Frank compared Grok's own native speech against a cloned-and-corrected version of it, twice, and chose the original both times: "the original is clearer, all the others are a bit étouffé", then "voz A is much better." The correction was real and measured, the clone had an 11 Hz pitch offset and a 7.6 dB hole in the 2–4 kHz presence band, closed to 1.6 dB across the whole spectrum, and it still lost. So: measure and correct every clone (it is cheap and it helps), but do not claim the factory reproduces the master's voice until Frank says it does. The open question for the next session is whether a professional-grade clone (about 30 minutes of ONE stable voice) closes the gap, or whether hero lines must simply be generated at the master engine. Why: the ear decides, and the ear has now decided twice against the clone.
  14. Upscale a person, then put the grain back, Frank's recipe (05-08-2026, the upscale test). Real-ESRGAN at 4× makes edges real again (lip border, hair strands, collar) but leaves the skin too even, which reads as retouched: "it looks less natural because it removes the imperfections." Adding a medium film grain on top is what he chose, grain new on every frame, luma only so the colour and the green stay clean, sigma about 4 on 0–255. Order in the factory, fixed: repaint the mouth at native size → upscale the finished clip → grain → chromakey last. Two things measured, because both are natural explanations that turn out to be wrong. (a) The colour does not shift. He read the upscaled skin as a different colour; the cheek measures R−0.6 G+0.4 B−0.7 out of 255, which is nothing. The eye reads smoothness as colour. (b) The upscaler does not remove micro-texture, it measures higher, not lower (2.38 against 1.28). It REGULARISES: irregular blotches become an even surface, and evenness is what reads as retouched. Grain works not by adding texture back but by breaking that evenness. The cost to know: grain is expensive to compress, the same clip is 24 MB at 2176 and 0.7 MB at 1088, so deliver two sizes, big to keep, small for pages. Why: I argued from the measurements that grain could not help, and his eye said otherwise on the rendered clip. Render the options, never reason about them.
  15. For an invented character you generate the movement; for a real person you capture it (05-08-2026, four rejected masters). Four Grok masters were bought from the same approved portrait of Frank, $4.84 in all, and he rejected every one: "totally exaggerated, this is not me at all" · too frozen, "I'm not moving, so make it natural" · "not natural at all" · "the smile is not nice". The failures are not a wording problem. Prompting swings the model between theatrical and dead and never lands on that particular person's way of being still, because the model is inventing a personality for someone who already has one. So: a generated master is right for Valentina, who does not exist. For a real person, start from film of them, their own existing recording, or sixty seconds on a phone against a plain wall, and let the new pattern fix the resolution. The mouth is repainted afterwards either way. Why: a person recognises their own stillness before they recognise their own face, and no prompt describes it.
  16. A capability anyone can use needs a door on the front page (05-08-2026, the Studio). The wizard, the page explaining it and its top-menu links all existed, and Frank still could not see it: "I don't see prominently in the studio index page either 'make your own speaker'. It should be linked to the homepage. People should see it." A top-menu link is not visibility. A card on the front door is. And the page that lists a capability must first answer the reader's real question, take one that exists, or make my own?, before it explains anything else. Why: the Studio is read by people who did not build it. What is one click away for us is invisible to them.
  17. Three ways to a speaker, and they are the page's spine (05-08-2026, Make your own speaker). Anyone arriving must choose one of three BEFORE reading anything else, and the page must be built around that choice, not mention it: take one that exists (free, nothing to supply) · make a NEW speaker who is not you, a person who does not exist, described by nationality, age, gender, dress; no photo, no recording, nobody to ask (~$1.50) · make an avatar of yourself, one photo and five minutes of your own recording (~$3). Frank had to say it twice before the page was restructured: naming the routes inside a section is not the same as making them the spine. A stock photo of a real person is not one of the three. Say plainly what cannot be done: Unsplash, Pexels, Freepik, Adobe Stock and Getty licences generally forbid making a real model appear to speak or endorse, which is exactly what a speaker does; a face found by image search is a lead, never a source. A historical figure in the public domain is the exception, search Wikimedia Commons by API, and search the painter's name too. Why: "What is not clear is that you can choose somebody else's avatar, and then you need to be given a list of places where you can find pictures according to the nationality and the gender." People will otherwise take the first stock photo they find.
  18. A section is a door, not a container (05-08-2026, the Speakers page). When a capability has several routes, the section's front page IS the choice: a question, then one card per route, each opening its own page, and everything common to all routes stays behind on the front page, once. Frank's own words: "We should start from the current speaker page, and then you have the three choices. This brings you to other pages, and we leave on the speaker page what is common to all types of speakers: the size, the cost, and all." Two corollaries he had to state separately: the choice must be visually big (a table of rows reads as reference, three large cards read as a decision), and a screenshot standing in for a live thing is deleted once the live thing is on the page. Why: a reader who has not chosen yet cannot use anything else on the page.
  19. The grain must never be visible to the key (05-08-2026, Valentina's rebuilt part). Frank rejected the first part rebuilt with the standard finish: "it's kind of bright, missing real colors, so today is much better." Measured cause (ibero session): the grain, not the network, her cheek was pixel-identical, but grain noise on dark, low-saturation pixels (hair, jacket) drifts them toward the key colour, so chromakey reads hair as background: mean alpha 236→198, fully solid hair pixels 27.7%→2.6%. Against a light page that reads as bright and washed out. The fix keeps criterion 14's order and splits the inputs: alpha from the CLEAN upscaled frames, colour from the GRAINED frames, merged ([1:v]chromakey…,alphaextract,erosion,erosion[al];[0:v]despill[rgb];[rgb][al]alphamerge). Corollary: sample the key green from CLEAN frames, grain clipping at zero lifts near-zero channels and reads 0x00a547 as 0x09a34b, wrongly. Why: his eye caught in seconds what the recipe's wording would have shipped on every part.
  20. The master ships as born; only the key may touch it (06-08-2026, the network pair, then the bar). Two verdicts the same evening. First, the same 8 s clip through the identical finish twice, quick and strong enlarging network, side by side, both rejected: "Both are bad. They're not of good quality, so we need to think of something else." Then his eye ranked the Desktop clips, in two passes that end in one place: the processed native part first read as "good finally" next to the pair, but against the raw clips he corrected it, "this Saludo Nativo is not good… the good ones" are the untouched 544 px Grok originals (the birthday clip, the master): "It's so wonderful to see." He himself noted the best one is "much smaller than the previous one." Even light processing (×2 enlarge + grain + a careless key) had already killed it; only the raw clip holds the wonder. The ranking is exactly inverse to the processing: measured, the native master carries MORE micro-texture (HF 8.2) than any processed output, the "real colors" are born in the clip, and each stage after birth (enlarger smoothing, synthetic grain replacing organic texture) subtracts what his eye misses. So: a speaker's part is BORN native, the engine speaks the exact line, at the master's own size, and the only step between birth and the page is the key (split inputs per criterion 19, at native size). No enlarger, no added grain on a native clip. Display the native size; never inflate it, the page shows her smaller than 544 anyway. The repaint was then convicted by its isolating test the same night: born-1080 base, mouth repainted, no enlarger, "number 4 is bad, mouth makes strange moves" (the base was a SPEAKING Veo master, so the base's own jaw motion fought the repainted mouth; a silent-base variant is untested, but the native road won and closes the question). Criterion 14 still governs the one case where an enlarger is unavoidable. Confirmed the same night on the keyed twin: transparency alone preserves the wonder, the raw ABC intro and its key-only version were both approved on sight ("yes 3 is good and 4 also"). Why: the pair was built to pick a network, and his eye disqualified the chain instead.
  21. A part is born in takes, at a speaking pace, never crammed into one clip (07-08-2026, the ABC narration). Frank on the first native set: "Valentina talked too fast, in particular in the saludo." The cause was arithmetic, not delivery: the engine caps a clip at 15 seconds (it refuses 20 outright), so 49 words became 3.3 words per second, and asking the model to "speak slowly" inside a full clip changes nothing, it still fits the text in. The fix is to cut the part at its own natural pauses and buy one take per clause, each with room to breathe; measured after, 1.9 to 2.1 words per second, approved on the ear. Three rules make the takes hold together, each earned from a rejection: the prompt must give a runway ("silent for a brief moment, takes a small calm breath") or the first word is clipped, "3 starts with missing word"; the raw takes must be trimmed at both ends (0.3 s before, 0.4 s after) or the join carries dead air, "there's a silence, and then it starts with an incomplete word"; and every take must be verified before it counts, transcribed and diffed against the script (one take said "la convierta" for "la convierte"), onset ≥ 0.25 s, pace ≤ 2.3 words per second. Why: the words were right and the clip was still wrong; pace is a measurable property of a clip, so it is checked like any other, not hoped for.
  22. Solid inside, soft at the edge, and served as video (11-08-2026, Valentina on the ABC note, on a phone). Frank, on his phone: "Valentina is shown on a kind of background." There was no background: her keyed clips were only 86% opaque inside the silhouette (median alpha 219 of 255), so the note's own lines of text read straight through her jacket and hair. Her AZFA clips measure 95% (median 243) and never drew the complaint. On a desktop the flaw hides, she stands over an empty white margin, and nothing shows through; on a phone the text runs right behind her. So: measure the interior alpha of every keyed part before it ships (decode one frame, take the median alpha of the pixels above 128; it must be ≥ 250), and when it is low, lift it, anything under 24 to zero, 200 and above to 255, straight ramp between, which keeps the anti-aliased edge (recipe: 3 - Projects/abc-colombia/valentina/opacar-clips.sh). And check how the server types the file: the ABC portal answered application/octet-stream for .mov and audio/webm for .webm; Safari refuses a <video> that is not typed as video, so the transparent twin can never play on an iPhone while Chrome, which sniffs, shows it fine. Why: two defects that a desktop preview cannot see, on the two things a speaker is made of, her alpha and her delivery.
  23. A speaker who is waiting must LOOK silent, not merely be muted (11-08-2026, the ABC note). The waiting clip on the note was the Grok master with its audio stripped, so before anyone pressed Sí, escuchar she stood in the corner mouthing her English birth line, "Hey there, it's so good to see you." Frank read it off her lips within seconds of opening the page: "Valentina is speaking in English. She should not. She should wait." Muting removes the sound and leaves the performance. So the waiting loop is its own clip, mouth closed, small smile, breathing, one blink, never a speaking take with the audio removed. And when the waiting clip comes from a different engine or a different frame size than the speaking parts, align it by measurement, fitting by HEIGHT and never by width, her width breathes as she gestures (357-419 px across the ABC parts), her height does not. Fitting the ABC idle by width made it 4 px taller than the parts and cut her crown; Frank saw both faces of it immediately ("you cut her head at the top", "size changes when she is moving to a different clip"). Fit top and bottom to the parts', centre horizontally, and treat a bounding box whose top is 0 as a cut, never as a fit. Worked example, both within 1 px: 3 - Projects/abc-colombia/valentina/rehacer-idle.sh. Why: the illusion is a person waiting to be asked; a person who is already talking to herself is a different, worse illusion.
  24. Judge a cut-out person over a DARK image, never over the page's own light background (12-08-2026, Vox Populi). Frank, looking at the presentation and planning a photo behind the speaker: "there's a kind of white, do you know why?" Two causes, both invisible on the light page they were built on. One: a mask fade at her bottom edge (linear-gradient(#000 94%, transparent)), added so her body would not "cut like a knife", but she is anchored to the bottom of the screen, so it is the SCREEN that cuts her and a clean edge there is invisible; all the fade did was turn the last centimetre of her body into fog. Two: her resting still was 82% opaque inside while her speaking clips were 97%, so the page showed through the state a visitor sees FIRST. Same defect family as criterion 22, found one day later on another speaker built by another route, which is the point: it is not a clip problem, it is a checking problem. So: measure the still and the clips alike (median interior alpha ≥ 250), drop bottom fades, and look at her composited over a dark photograph before shipping, light on light hides everything. Why: every one of these defects was invisible in the environment where it was built, and obvious in the one where it will be used.
  25. A speaker made of several clips must END each one where she BEGINS it (12-08-2026, la nota ABC). Clips generated one at a time all start from the same portrait and end wherever the engine left the character. Measured on the eight ABC parts: overlap between two starts 0.975-0.993, overlap between one end and the next start only 0.892-0.966, so every join snapped her back from mid-gesture to neutral, and Frank named the seams one by one ("entre el primero y el segundo hay un salto, y lo mismo entre el segundo y el tercero, y así") even after a 240 ms dissolve was added. A dissolve blends two images; it cannot make two different poses one. The fix is in the clip, not the page: append a settle, 0.45 s in which the last frame cross-fades back to the clip's own first frame, with the audio padded by the same silence so no word is touched (all seven seams went to 0.970-0.994). Two rules follow: measure a seam, never judge it by eye (silhouette overlap between the last frame of one clip and the first of the next), and check for trailing silence before promising a trim, ours had none, which is why the answer had to be adding rather than cutting. Recipe: 3 - Projects/abc-colombia/valentina/asentar-partes.py.
  26. Measure the speaker's FACE in every clip: the engine does not hold its framing (12-08-2026, la nota ABC). Frank, after two fixes aimed at the seam: "Valentina still very different from a clip to another." Face detection frame by frame gave the answer the silhouette had hidden: the eight parts all START identical (face 149-154 px wide, centre 282,128 in a 544 frame) but INSIDE each one the engine pushes in and pulls back, the face travels between 137 and 164 px, up to 18 % of its width, and her head wanders 20 px. Every part therefore ends on its own framing and the next resets to the portrait's. Silhouette measurements miss this entirely, because she is cut at the frame's bottom edge: her bounding box stays the same while she grows. So the acceptance measure for a multi-clip speaker is her face, width and centre, at the start, the middle and the end of every clip. Fix it by warping: measure per frame, smooth over ~1.2 s (the drift is slow, her real head motion is not), and map every frame onto ONE reference shared by all the clips and the waiting loop. Two anchors were tried and rejected, both because they moved the problem instead of removing it, anchoring her bottom edge made her rise and fall by 70 px, and filling the gap under her shoulders left a grey band at the screen edge; the gap (52 px at worst) belongs below the screen, pushed there by the page. Recipes: 3 - Projects/abc-colombia/valentina/estabilizar-partes.py and rehacer-idle.py. Why: three rounds were spent on the seam, dissolve, then settle tails, while the cause was inside the clips. A defect the user reports twice is a defect measured wrongly, not a defect fixed weakly.
  27. ffmpeg -i without an explicit decoder silently drops alpha, on BOTH WebM/VP9 and HEVC/.mov (19-08-2026, la nota ABC). Frank saw a solid blue rectangle behind Valentina on four of eight parts and called it "un marco". The repaint (LatentSync) and the chromakey (new_pattern_batch.py) were both innocent, measured with the correct decode, out/abc-seg1-2176.webm (the direct finish output) already had a clean silhouette, alpha median 250+. The break was one step later, in the small downscale helper copy-out-to-originales.sh: it read the 2176 alpha webm with a plain ffmpeg -i "$src2176" …, and ffmpeg picked its default VP9 decoder, which has no alpha support, and silently filled every pixel opaque instead of erroring. The corrupted, alpha-less 544 file then fed estabilizar-partes.py, whose only real transparency came from its own affine-warp padding at the edges, the "frame" Frank saw is exactly that padding. The fix is a one-line decoder flag on the READ side: ffmpeg -c:v libvpx-vp9 -i src.webm … (confirmed A/B on the same bytes: without it, alpha mean 255/100% opaque; with it, alpha mean ≈123, correct silhouette). The same trap exists on the HEVC/.mov side, ffmpeg -i clip.mov -f rawvideo -pix_fmt rgba - can return a fully-opaque image if the file's alpha ever went missing upstream, and there is no separate flag to force it (Apple's auxiliary alpha track is read by ffmpeg's native hevc decoder when it is genuinely present, but nothing warns you when it silently is not). So: never trust a naive ffmpeg -i decode to prove alpha is correct, on either container. For WebM/VP9, always force -c:v libvpx-vp9 on the input before judging a clip broken or fixed. For HEVC/.mov, verify via its VP9/WebM sibling instead of the .mov directly whenever one exists, that sidesteps the ambiguity entirely. Any script in this pipeline that pipes ffmpeg -i *.webm into rawvideo/rgba without the explicit decoder is one silent ffmpeg-version or codec-negotiation change away from repeating this bug. Why: ffmpeg failed closed instead of failing loud, a wrong decoder produced a valid-looking opaque image instead of an error, so three pipeline stages (encode, downscale, stabilise) all "succeeded" while carrying no real alpha.
  28. The mouth is measured against the audio: coupling, silences, and lag per third (20-08-2026, la nota ABC). Frank on the new note: "voz no sincronizada, boca con movimientos raros… no se abre lo suficiente." Two hypotheses, both settled by measuring the final bytes. (a) Progressive temporal drift from the stabiliser's -shortest/apad, REFUTED: 24 fps and exact frame counts through the whole chain (499 frames in = 499 out per part; audio duration = video duration within 0.06 s). The desync is not in the containers. (b) The repainted mouth, CONFIRMED, and now it has numbers. Three measurements separate a native part from a repainted one on the same speaker, same ruler (3 - Projects/abc-colombia/valentina/verificar-sincronia.py): audio↔mouth correlation (native r 0.42–0.58 · repainted 0.08–0.35, the repainted mouth barely tracks the sound), silence→rest ratio (native 0.32–0.45 · repainted 0.52–0.74, the mouth keeps moving through the audio's silences), and best-lag stability per third (native 0±2 frames · repainted swinging −10 to +12 inside one clip; one part ran ~0.3 s early throughout). And the flat mouth had a cause upstream of every container: the base. The 19-08 run repainted over the RESTING base (grok-base-silent-idle, mouth closed) instead of the speaking ping-pong base, LatentSync inherits the jaw it is given: median tooth-pixel height in speech frames 11 px against 54 px on the native part (diagnosis session, mouth crop anchored to the stabilised face). So the shipping gate for any speaking part now includes the sync test, silence→rest ≤ 0.55, |lag| ≤ 2 frames in every third with r ≥ 0.25, AND the opening test: median mouth opening in speech ≥ 60 % of an approved reference clip. The speaking base is declared ONCE, in the recipe, with its why (3 - Projects/abc-colombia/valentina/rehacer-partes.sh); the gate is wired INSIDE the path (verificar-clips.py), no PASS, no copy. The controlled experiment (20-08) then closed the door the diagnosis could not: the same four clone audios repainted over the SPEAKING base, the gate passed alpha, face, settle and duration, and still failed every sync leg: the mouth now opens, but on the BASE's rhythm, not the audio's, four frames inside one 0.5 s audio silence show her wide open and articulating, and the lag still swings ±11 frames. So the repaint route fails on either base: resting base → mouth stays closed; speaking base → two rhythms fight (criterion 20's conviction, now photographed). A repainted part cannot pass what a native part passes: the 12-08 verdict, todo repintado pierde contra el Grok crudo, now has its numbers. Why: the eye said "raro" three times while every container measurement was perfect; the coupling between mouth and sound, and the mouth's opening, had never been measured, so the defect kept shipping as "fixed".
  29. A sync meter must read the clip's OWN frame rate, and the gate must be calibrated against work already approved (20-08-2026, la nota ABC). The first open-engine candidate was reported to Frank as drifting late: mouth lag growing +2 → +5 → +10 frames across ten seconds. It was not the engine. The clip is 25 fps and the meter binned its audio at 24 fps, the rate hard-coded for our own clips: 245 bins at 1/24 s consume 10.21 s of audio against 9.80 s of video, which fabricates exactly 0.41 s, exactly 10 frames, of perfectly linear late drift. Re-measured at its true rate, the same bytes give lag 0/0/0 and the best silence→rest ratio of any clip we hold, better than the native master. A fabricated defect is worse than a missed one: it convicts an innocent engine and sends the next session down a road that does not exist. So a meter that will be pointed at foreign clips reads width, height and frame rate from each file, and locates the mouth by detecting the face rather than by a fixed window; a hard-coded constant that is true for our pipeline becomes a lie the moment the ruler travels. And the thresholds themselves are calibrated against clips Frank has already approved: ours measure lag 0/−3/−4 on the live page he called good, so a gate demanding |lag| ≤ 2 would have failed his own approved work; it stands at |lag| ≤ 4 with drift ≤ 5 for that reason. A gate that fails what the customer already accepted is not strict, it is broken. Why: the same session that convicted a real defect with numbers then invented one with numbers. Measurement earns trust only when the ruler is checked against known-good and known-bad before it judges anything new.
  30. Her mouth has a size, and it is a BAND, not a floor (20-08-2026, la nota ABC). Frank on the first open-engine clip: "solo el clip nativo es bueno, en los dos otros la boca se abre demasiado." Measured as vertical mouth opening in percent of her own face width, so it compares across engines and frame sizes: native master 27.8 % median (p95 35.9) · open engine 36.9 % (p95 45.6) · the repaint on the live page 33.6 %. So his eye rejected a mouth a third wider than the one he approved. This is the mirror image of the 19-08 defect, where a repaint over a resting base opened too little and he said "no se abre lo suficiente". One property, two ways to fail, so the check is a band around the approved clip, never a floor: 24 % to 32 % of face width at the median. A "≥ 60 % of the reference" rule passes an over-articulated mouth cleanly, which is exactly how this one reached him. The knob that moves it is the audio guidance strength (sample_audio_guide_scale in MultiTalk, default 4.0 in a recommended 3-5 band; audio_guidance_scale in EchoMimic, best 1.8-2): higher pushes the mouth harder at the audio, lower calms it. Ours ran at the default, at the top of its band, untested lower. Why: the session had a number for silence and for lag and none for size, so the one property his eye judged was the one nothing measured.
  31. The house style for a speaking take: calm, restrained, and the smile survives the widest moment (Frank, 21-08-2026, the ABC note, chosen blind in an A/B). He said the new takes opened the mouth too much. The ruler said otherwise: median opening 16.7 percent of her face width against 16.6 on his own reference clip. Two things were really happening. The pace: the free engine speaks at 2.4 to 2.7 words a second where his approved clips sit at 2.1, and a fast mouth reads as a mouth that moves too much. The peak: his references top out at 20 to 21 percent, several new takes reached 22.7, and the eye judges the widest instant, not the median. Two calmer versions of the same sentence were then generated and shown side by side at their widest frame; he picked B without knowing which the session preferred, and the session had picked B too. B is now the order for every take: about 1.9 words a second, small restrained lip movements, and at the widest moment the mouth still keeps her smile, never a dropped jaw. Measured band: median 14 to 18 percent, peak at or under 21. Ruler: 3 - Projects/abc-colombia/valentina/medir-boca.py. Why: the first answer to "the mouth opens too much" was a number that said he was wrong. He was not wrong; the meter was measuring the wrong thing. When his eye and the ruler disagree, look for the property the ruler is not measuring, and show him the comparison rather than the argument.
  32. Lo que se publica es la PÁGINA, no la toma, y la vara se pasa entera antes de publicar, no después de que él lo note (Frank, 21-08-2026, las páginas ABC). Tres defectos llegaron a su oído el mismo día, y los tres nacieron del mismo error de granularidad. El empalme: cada frase era un archivo y la página saltaba de uno a otro, así que ella cortaba en seco y arrancaba de golpe; medí cada toma por separado y ninguna medida miraba la costura, que es donde vivía el defecto. Cura: una sola toma por página o por sección, unida con 0,7 s de respiración (ensamblar-tomas.py), y si el dibujo tiene que avanzar, se mueve por tiempo dentro del clip único, nunca cambiando de archivo. El barrido tardío: gateé toma a toma mientras construía y solo pasé la auditoría completa cuando él se quejó; ocho de dieciséis fallaban. Cura: auditar-pagina.py sobre TODAS las piezas, y su tabla limpia es condición de publicar. Y la vara que se queda corta: él ve un desfase de nueve cuadros que mi umbral dejaba pasar; cada vez que su ojo encuentra algo que el metro aprobó, el metro gana una columna y la vara se aprieta contra el clip que él acaba de rechazar. Es el mismo linaje que los criterios 29, 30 y 31: tamaño de boca, luego picos, luego ritmo, ahora costura. Why: medir la pieza no es medir la experiencia. La ilusión se rompe entre las piezas, que es justo donde nadie estaba mirando.
  33. Su GEOMETRÍA se mide como se mide su voz, en todas las piezas y en los tres estados (Frank, 21-08-2026, la nota conceptual). Él lo dijo entero en dos frases: "no deja de aumentar y bajar su tamaño… a veces se pega al bottom y a veces no" y "está pegada al bottom, pero su cabeza está ligeramente cortada". Medido: siete de las nueve piezas de la nota fallaban. Cuatro la dejaban flotar sobre el borde de abajo, hasta 53 px en la primera parte; la de espera tenía la coronilla contra el techo del cuadro (aire 0) y su cara 17 px más arriba que las habladas, así que ella saltaba cada vez que él la paraba. Tres causas, cada una con su lección. Una: la vara medía la TOMA, nunca la PIEZA EN LA PÁGINA. Todas nuestras reglas miran palabras, ritmo, boca y desfase; la geometría se midió una sola vez, el 12-08, y después se dio por resuelta. Dos: el estabilizador prometía en su docstring lo que su código no hacía: "la última fila opaca de cada columna se prolonga hasta el borde", y sólo medía el hueco y lo imprimía en una línea amable que se leía como un informe de trabajo hecho. Un comentario no es una prueba: lo que no se ejecuta, no existe. Tres: el ancla vivía copiada en tres archivos (estabilizar-partes.py en 128, acabar.py sustituyendo el texto a 146 sobre una copia, rehacer-idle.py con la suya), así que las partes habladas se mudaron y la espera se quedó atrás, y ningún archivo podía delatarlo porque los dos números nunca se encontraban. Y el defecto vivía en los estados que nunca visitamos: parada, pausa y reanudación, que es de donde salieron sus dos capturas. Así que: el ancla en UN solo archivo (valentina/encuadre.json), leído por el estabilizador, el rehacedor de la espera y la vara · auditar-encuadre.py barre TODAS las piezas de la página, la de espera incluida, con cuatro columnas : cara, aire sobre la coronilla (cero es una cabeza cortada), hueco bajo los hombros (más de cero es flotar) y alfa : y su tabla limpia es condición de publicar · pegar-al-borde.py hace lo que el docstring prometía, y prolonga la última fila SÓLIDA, nunca la del borde semitransparente, que pinta una banda gris · y la prueba final se hace en la página, en los tres estados, con capturas. Why: cada vez que apareció un defecto escribimos la vara de ese defecto; nunca escribimos la vara que mira lo que él mira, que es ella, entera, en la página, en todos sus estados.
  34. La herramienta que repara el defecto de hoy es la que fabrica el de mañana, y su propia regla nunca lo verá (Frank, 21-08-2026, la nota conceptual, después de varias horas). El 12-08 dijo "cambia de tamaño de un clip a otro" y le construimos un estabilizador anclado en la caja del detector de caras. Arregló el desajuste ENTRE clips y metió uno nuevo DENTRO de cada uno, porque esa caja crece cuando ella abre la boca: cada cuadro se escalaba al ritmo de su habla. Nueve días después dijo "no deja de avanzar y recular", y también "flota", que era el mismo mecanismo por abajo (al encoger el cuadro, sus hombros se despegaban del borde, hasta 53 px). Nadie midió el clip DESPUÉS de la herramienta con una regla distinta de la que la herramienta usaba para corregir. Una herramienta que corrige con una medida es perfecta según esa medida, por construcción: la regla del detector de caras decía que todo estaba en su sitio mientras el ojo de Frank veía lo contrario. Con la regla honesta (su PELO, que no se mueve al hablar, valentina/medir-cabeza.py) los números salieron por fin comparables: toma cruda del motor 2,8 % de vaivén y 0 px de sube y baja · la página que él aprobó 4,2 % y 11 px · lo que publicamos 6,9 a 11,4 % y de 18 a 33 px. Tres reglas distintas dieron tres respuestas distintas el mismo día, y su ojo acertó las tres veces. De ahí las reglas que quedan: toda corrección se verifica con una medida INDEPENDIENTE de la que corrige · esa medida se calibra contra trabajo que él ya aprobó Y contra trabajo que él ya rechazó, o no distingue nada · una medida que mide una parte del cuerpo que se mueve al hablar no sirve para medir el encuadre · y cuando su ojo y la regla no coinciden, la regla está mirando otra propiedad, que es el criterio 31 otra vez, ahora con una segunda factura pagada. Why: pasamos nueve días mejorando la medida equivocada. El defecto no era el motor, ni el keying, ni la página: era nuestra propia reparación anterior, y ninguna de nuestras varas podía verlo porque todas descendían de ella.
  35. Una cura se aplica a TODOS los niveles donde vive el defecto, y el empalmador es uno solo (Frank, 23-08-2026, la parte 1 de la nota). Por la mañana curamos las costuras ENTRE las partes de una página: un solo archivo, fundido de 0,35 s, respiración de 0,7 s. Por la tarde él escuchó la parte 1 rehecha y dijo "son como tres clips pegados juntos, hay como dos interrupciones". El mismo defecto, un nivel más adentro: dentro de una parte, tres tomas se seguían pegando con la herramienta vieja, que sostiene el último cuadro y CORTA. Y la interrupción no estaba en la imagen, estaba en el sonido: 1,46 y 1,40 segundos de silencio en mitad de una frase que continúa, contra 0,68 s de pausa natural dentro de una toma nativa suya. Salía de una suma que nadie hizo: 0,45 s de cola que yo dejaba en cada toma, más 0,7 s de respiración que añadía el empalmador, más el segundo de silencio con que nace cada toma. Dos varas mías dijeron que las dos versiones eran iguales (solape de siluetas 0,979 contra 0,985; diferencia de imagen entre cuadros 4,5 contra 3,7, cuando una toma sin costura marca 4,3): las dos miraban la imagen, y la propiedad estaba en el audio. Así que: un solo empalmador en todo el sistema (unir-partes.py, con --respiro corto dentro de una parte y largo entre partes; ensamblar-tomas.py se queda solo con quitar el verde de UNA toma) · las tomas se recortan a su habla antes de unirlas, de modo que la pausa sea una decisión y no la suma de tres colas · y la compuerta se mide sobre el archivo terminado, no sobre la intención: ningún silencio de empalme por encima de 0,8 s, umbral puesto entre lo que él acepta (0,68 natural) y lo que rechazó (1,46). Why: dos herramientas que hacen lo mismo de dos maneras garantizan que una de ellas siga estando mal, y arreglar un nivel sin bajar al siguiente sólo mueve el defecto de sitio.
  36. LA VARA: la calidad mínima de cualquier clip, en un solo comando (Frank, 24-08-2026, "fija eso como la regla absoluta para todos los clips"). Ningún clip de un presentador se publica sin pasar valentina/la-vara.py, y ninguna página sin pasarla sobre TODAS sus piezas, la de espera incluida. Siete columnas, y cada una nació de una queja suya: cara 143 a 165 px, la misma mediana en todas las piezas · aire sobre la coronilla ≥ 4 px, cero es una cabeza cortada · hueco bajo los hombros = 0, más es flotar · alfa interior ≥ 250, menos es que la página se transparenta a través de ella · pausa de empalme ≤ 0,8 s, más se oye como una interrupción · ritmo 1,9 a 2,1, tope 2,3 · boca 14 a 18 % de su cara, pico ≤ 22. Y dos comprobaciones que no son columnas: las palabras siguen todas después de unir, comparadas por sonido, y la prueba final se hace en la página, en sus tres estados, hablando, en pausa y parada. Por qué no era así antes, que es lo que él preguntó: cada regla nació como respuesta a una queja y se quedó viviendo donde cayó, una en cada script y dos sólo escritas en la skill, así que el listón nunca existió como un objeto y cada sesión pasaba la mitad que recordaba. Why: una regla que hay que recordar no es una regla, es una intención; sólo cuenta la que se ejecuta y devuelve PASA o NO PASA.
  37. One speaker, ONE implementation. Three copies is why a fix never lands (25-08-2026, el sitio ABC). Frank, after an afternoon of chained defects: "¿por qué teníamos tamaños distintos en las distintas páginas, por qué cada vez que tocas algo se descompone?" The answer was structural, not careless. Valentina existed three times: the front page used the shared valentina/guia.js; componentes.html carried its own inline copy of the same CSS; nota-conceptual.html a third with #nar-video. Each was written by a different session on a different day, and each kept its own numbers. So she measured 190 px on the front page, 118 in componentes, 150 in the note, nobody decided that, and every repair had to be made three times or it silently missed two pages. The project had already learned this exact lesson for ONE number, and written it into encuadre.json: "un número que vive en tres archivos se desincroniza; este es el único". It was never generalised from the number to the whole component. So: a speaker, a card, a control bar, a highlight, are ONE piece with one home. A page declares its content and its cue times; it never re-implements the mechanism. Before repairing anything a speaker does, count how many copies of it exist, and if the answer is more than one, say so before touching a line. Why: every hour of that afternoon bought one page's worth of fix while two pages stayed broken, and Frank kept finding the same defect somewhere else.

  38. The bar measured the CLIP and nothing measured the PAGE (25-08-2026, el sitio ABC). la-vara.py is a good instrument and every clip passed it. And yet the page was broken in ways no clip measurement can see: on a phone the invitation card was set to display:none, so she stood mute in the corner with no way to ask her to speak; her card sat on top of the UN credit; the player stopped her after ten seconds; the highlight lit the wrong door. A perfect clip inside a broken page is a broken page. Nine defects in one afternoon reached Frank's eye, and each cost a full round trip: he sees it, reports it, it is diagnosed, repaired, published, he looks again. So the page has its own bar, run before he is shown anything (3 - Projects/abc-colombia/valentina/la-vara-pagina.py): at phone, wide-phone and desktop, on the LIVE bytes, it checks that the card is there with her name and her question, that she measures the same on every page, that nothing of hers covers text, that the page does not overflow sideways, that the title is not glued to the top, that she plays past the opening seconds, and that the highlight follows the video. Its first run caught a defect that would otherwise have shipped. Why: the customer's eye is the most expensive check in the house, and it was being used as the first one.

  39. A fix is not done until it is PUBLISHED, and saying "fixed" about a local change is a lie by omission (25-08-2026). Twice in one afternoon the repairs were made, measured and reported as done while Frank was looking at the live site, where nothing had changed: "los dos problemas que reporté no se arreglaron". The measurements were true and worthless, because they described a copy nobody could see. So: every claim about a defect names where it was verified, and the only place that counts for a defect the customer reported is the live bytes he will look at. Measure locally to work; measure live to speak.

  40. Every fix is a new defect until its neighbours are measured (25-08-2026). Three of the afternoon's defects were created by that same afternoon's repairs. Moving the card next to her, which is the house rule, put a backdrop-filter beside a transparent video: iOS pushes such a video into its own compositing layer and loses the alpha, so she appeared inside a blue box on Frank's iPhone, and the clips were innocent, measured at 51.5 % transparent on both twins. Making her bigger, which he asked for, made pre-existing seam jumps visible for the first time. Adding the highlight with setTimeout lit the wrong door whenever the video took a moment to load, because the clock ran while the video did not: the cue must be read from currentTime, which never lies, not from a timer. So a repair is measured at several widths, at both ends of the scroll, and on the states around it, before it is called a repair. Why: a fix aimed at exactly one screenshot fits exactly one screenshot.

  41. The joins are measured in TWO planes, and the ruler that was missing is the eye's (26-08-2026, la portada ABC). Frank named six transitions on the front page, one after another: "la transición hacia el recorrido en pantallas no es buena, las 2 siguientes tampoco." Two different defects were hiding under one complaint, and the bar caught neither. The ear's plane: three joins held silences of 1.04, 1.09 and 1.06 s where the house limit is 0.80; past that it stops sounding like a breath and starts sounding like an interruption. The bar did measure pauses, but against the loose page threshold of 2.5 s, so they passed. The eye's plane: the silhouette overlap across the joins measured 0.850 and 0.890 where the house asks 0.97, and nothing measured it at all. So she changed posture while crossing the pause, which no dissolve repairs, because a dissolve blends two images and cannot make two poses one (criterion 25). Both rulers now live in la-vara.py (salto_de_postura), and a clip failing either cannot be published; apretar-costuras.py shortens a long join by removing frames from the CENTRE of the silence, where she is still, taking exactly the same from the audio at the same place so the sync never reopens. And one measured negative worth keeping: giving each take more of its calm run-up makes the join WORSE, not better, because the run-up is half a second of a widening smile before the first word. Why: a complaint about "the transitions" is not one defect, and the plane nobody measures is the one the customer sees.

  42. The audio is squared to the video per part, in FRAMES, before anything is joined (25/26-08-2026). The mouth fell behind the voice, worse towards the end, exactly as Frank described it. The joiner wrote video by counting frames and padded audio by counting seconds, assuming the two agree. They do not: when the breath is shorter than the fade, the hold collapses to zero but the fade's frames are still written, so the video gains two frames at every join. Five joins, ten frames, 0.417 s, and the file itself showed audio 0.397 s shorter than video. Every joiner squares each part's audio to that part's own video duration, in frames, before adding the breath; and the finished file is checked with video frames over fps against real audio samples, which must agree to the millisecond. Why: two units for one timeline is a drift generator, and it hides because each single join is imperceptible.

  43. Un vídeo PARADO no se enseña nunca: en reposo manda una foto con transparencia (26-08-2026, la portada ABC). Frank, por tercera vez en tres semanas y ya sin paciencia: "todavía V tiene un fondo, fix una vez por todas, es un problema recurrente." Las veces anteriores se buscó la causa en el archivo, y el archivo estaba bien: medido otra vez el 26-08, alfa mediana 255 por dentro, la mitad del cuadro a cero, esquinas limpias, y la propia miniatura de Apple (QuickLook, el mismo motor que Safari) lo pinta transparente. La causa estaba en el reproductor. Cuando el navegador no deja arrancar el clip, y iPhone en ahorro de energía no lo deja, el reproductor hacía visible el vídeo igual, parado; Safari pinta ese primer cuadro sin transparencia y sale ella dentro de un rectángulo con el fondo del plato. La cura tiene tres partes, y las tres van juntas: (1) mientras ningún clip suene de verdad manda una foto suya con transparencia (reposo.webp, 22 KB), que ningún navegador puede pintar con fondo; (2) lo que se ve se decide por el ESTADO REAL de reproducción (currentTime > 0 && !paused && !ended), no por la promesa de play(), que resuelve en sitios donde el vídeo no ha empezado; (3) Safari se lleva siempre el .mov, detectado por su nombre de navegador y no solo por canPlayType, porque su webm no sabe de transparencia y enseñaría el fondo entero. Se prueba así: cargar la página sin tocar nada y comprobar que manda la foto; pulsar y comprobar que manda el vídeo; parar y comprobar que vuelve la foto.

  44. El recorte viaja DENTRO del vídeo, no en la transparencia del archivo (26-08-2026, la portada ABC, misma noche que el 43). La foto en reposo del criterio 43 tapó el defecto en el arranque, y el recuadro volvió en cuanto ella empezó a hablar: "vuelve después de un rato". Medido por geometría en su captura: el rectángulo mide lo que mide su cuadro de vídeo, no la tarjeta. iPhone no aplica la transparencia del archivo, ni la del .mov con capa alfa de Apple (comprobada: la capa 1 existe en el flujo, y la miniatura de Apple en el Mac la pinta transparente) ni la del webm. La cura que no depende de nadie: cada clip viaja en UN archivo opaco, H.264, con el color arriba y su recorte en gris abajo, y la página los junta cuadro a cuadro en un <canvas>, pasando el gris del recorte a la transparencia del color. Todos los navegadores entienden un vídeo opaco. Se hace así: [0:v]format=rgba,split=2[c][m];[c]format=yuv420p[cv];[m]alphaextract,format=gray,format=yuv420p[mv];[cv][mv]vstack=inputs=2[v]. Pesa menos (la portada baja de 3,9 a 2,5 MB) y desaparece la segunda copia por clip. Dos avisos: el lienzo y el vídeo tienen que ser del mismo sitio, o el navegador prohíbe leer el lienzo; y si esa lectura falla, manda la foto en reposo antes que enseñar un recuadro.

  45. En el teléfono, la altura de la pantalla se mide en dvh, nunca en svh (27-08-2026, la portada ABC). Frank vio tres noches el mismo defecto: "Valentina sigue empujando los textos para arriba, ya no se ve el título". No era ella. svh es la pantalla con la barra del navegador puesta; cuando Safari la esconde, la pantalla crece, el bloque de 100svh se queda corto, por abajo asoma el fondo, y ese hueco deja mover la página hasta que el título se va detrás de la hora. Lo que se movía en la pantalla era la presentadora, así que parecía suya la culpa. La regla: cualquier bloque que deba llenar la pantalla del teléfono se mide en dvh, con svh antes como respaldo para navegadores viejos; y si debajo no hay nada, se bloquea la rodada (html, body { height: 100dvh; overflow: hidden }) dentro de la consulta de teléfono. Se prueba así, y sin esta prueba no está hecho: a 393 por 640, la altura de la página tiene que ser igual a la pantalla, y un scrollTo(0,300) tiene que dejarla en cero.

  46. Y ni dvh basta: el bloque que llena la pantalla del teléfono va FIJO a los cuatro lados (27-08-2026, la portada ABC, la cuarta vuelta sobre el mismo defecto). Frank lo describió con dos capturas separadas por dos segundos: al entrar, perfecta; después, una banda oscura abajo, todo sube y el título desaparece. Y añadió el dato que descarta la causa fácil: "antes no había carrusel y teníamos el mismo problema". La causa medida: Safari cambia el alto visible DOS veces al abrir, primero con su barra de traducción, que aparece y se va sola, y después escondiendo la de direcciones. Un bloque medido en svh o en dvh no sigue esos saltos al vuelo: se queda con el alto de antes, deja ver el fondo por abajo, y el hueco mueve la página. La cura: position: fixed; inset: 0 en la consulta de teléfono, que mide siempre la pantalla del instante pase lo que pase con las barras, más html, body { height: 100%; overflow: hidden } cuando debajo del bloque no hay nada. Se prueba cambiando el alto de la ventana SIN recargar: de 640 a 720, el bloque tiene que pasar a 720 en el mismo momento, con el título entero y la foto cubriendo. Esa prueba es la que faltaba en las tres vueltas anteriores.

Source previewDownload original source