Source
LongCat-Video repository
Read the primary source.
Reference illustration, not a screenshot of this source.
Open audio-driven avatar model · Video · Speakers
Turns one portrait and an approved audio track into a speaking video, face and body born together. Open weights from Meituan.

A portrait speaking our own approved audio, when InfiniteTalk misses lip shapes.
Not runnable on the Mac (NVIDIA cards only), and not tested by its authors in French or Spanish.
Portrait image; Approved audio; Optional prompt
Speaking video, 480p or 720p, up to 2 minutes a job on WaveSpeed
Hosted API (WaveSpeed $0.04 a second at 480p, fal $0.15) or open weights on a rented card; not run by us yet
Provider capabilities and historical Studio notes are distinguished below. Current account access, credits and runtime readiness were not tested.
WaveSpeed $0.04 a second at 480p, $0.08 at 720p, 3-second minimum; fal $0.15 and $0.30 (read 15-09-2026). Open weights: rented card time.
Provider-hosted on WaveSpeed or fal; self-hosted on a rented card.
Authors evaluated English and Chinese; French untested.
The strongest open claims on lip sync (Whisper-Large audio encoder, 8-step distillation, MIT weights), but not yet run on our material. The rating moves after the French bilabial test on « tableau de bord ».
Decision: Test it on the lip shapes InfiniteTalk misses before using it for a whole speaker.
Same family as InfiniteTalk: the whole person is generated from portrait and audio, not a mouth repainted on an existing clip.
Official model card and provider prices read 15-09-2026; no output test on our speakers
Catalogue review: 2026-09-15. The evidence note identifies what was verified. This date does not imply a fresh tool test or verification of every linked source.
Source
Read the primary source.
Reference illustration, not a screenshot of this source.
Source
Read the primary source.
Reference illustration, not a screenshot of this source.
Source
Read the primary source.
Reference illustration, not a screenshot of this source.
Source
Read the primary source.
Reference illustration, not a screenshot of this source.