
Source
AuK-Flash official model card
Read the primary source.
Reference illustration, not a screenshot of this source.
Speech-generation engine · Voice
Tencent's AuK (announced as "Nano Banana for audio"), here in its distilled four-step version: speech from a reference voice or a written voice description, plus editing, cleaning and separation of existing speech through written instructions.

Repairing an existing voice track by instruction (swap one mispronounced word, change emotion, remove breaths or noise) and making extra takes of a speaker's voice from a short reference, in non-commercial work.
Not for commercial use while it depends on the Qwen research encoder, not a streaming voice, and French is still untested.
Written instruction with the text to speak; Reference audio to clone, or source audio to edit
Speech audio at 24 kHz; Edited, cleaned or separated recording
Self-hosted Python runtime on a CUDA GPU, or the MLX branch on Apple Silicon; Gradio, ComfyUI and SGLang-Omni integrations
Provider capabilities and historical Studio notes are distinguished below. Current account access, credits and runtime readiness were not tested.
MIT weights, self-hosted GPU or Mac compute. Commercial use also needs a separate licence for the Qwen encoder.
Self-hosted if run locally; the public demo runs on Hugging Face.
Model card lists Chinese and English; instruction templates and evaluations are English and Chinese only. Spanish worked in one Studio test (27 September 2026). French untested.
Editorial task fit. Spanish now works in one Studio test, but its required encoder is non-commercial and it needs a large GPU or the Apple Silicon branch, so it stays below IndexTTS 2 / 2.5 (three stars) and Qwen3-TTS (four stars) as a general engine. Its instruction editing of existing speech goes further than any other engine in this task.
Decision: Use the public demo to repair or extend an existing voice track (one word, one emotion, extra takes) and judge it by ear. Do not build a commercial product on it while the Qwen encoder licence stays research-only.
The editing range is what sets it apart: replace spoken words, change emotion or timbre, reduce an accent, add or remove breaths and laughs, convert whisper, denoise and keep one speaker, each through a written instruction. Studio test, 27 September 2026, AuK Base on the public demo: a 15-second Spanish reference from an existing Studio speaker and a 20-word Spanish sentence. Speech recognition returned every word exactly; median pitch 204 Hz against 198 Hz for the reference (+6 Hz, half the 12 Hz drift measured between takes of the engine that invented that voice); loudness -11.5 against -12.1 LUFS; spectral centroid 1,075 Hz against 941 Hz, so slightly brighter. Timbre match still needs a listening check. The authors' report compares it with Qwen3-TTS and VoxCPM2 on English and Chinese test sets only; Studio has not verified those results.
Official model card, repository, licence files and technical report reviewed; one Spanish zero-shot take measured by Studio on the public demo (27 September 2026), not yet judged by ear
Catalogue review: 2026-09-27. The evidence note identifies what was verified. This date does not imply a fresh tool test or verification of every linked source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.

Source
Read the primary source.
Reference illustration, not a screenshot of this source.