Use with AI
Back to voice comparisons

AuK-Flash

Speech-generation engine · Voice

Tencent's AuK (announced as "Nano Banana for audio"), here in its distilled four-step version: speech from a reference voice or a written voice description, plus editing, cleaning and separation of existing speech through written instructions.

AuK-Flash capability illustration
Official Hugging Face social-thumbnail image for Tencent's AuK-Flash model card, retrieved 14 September 2026.

Use it for

Repairing an existing voice track by instruction (swap one mispronounced word, change emotion, remove breaths or noise) and making extra takes of a speaker's voice from a short reference, in non-commercial work.

Choose something else for

Not for commercial use while it depends on the Qwen research encoder, not a streaming voice, and French is still untested.

Give it

Written instruction with the text to speak; Reference audio to clone, or source audio to edit

Get back

Speech audio at 24 kHz; Edited, cleaned or separated recording

How it is operated

Self-hosted Python runtime on a CUDA GPU, or the MLX branch on Apple Silicon; Gradio, ComfyUI and SGLang-Omni integrations

Availability

Provider capabilities and historical Studio notes are distinguished below. Current account access, credits and runtime readiness were not tested.

How charging works

MIT weights, self-hosted GPU or Mac compute. Commercial use also needs a separate licence for the Qwen encoder.

Where processing happens

Self-hosted if run locally; the public demo runs on Hugging Face.

Language support

Model card lists Chinese and English; instruction templates and evaluations are English and Chinese only. Spanish worked in one Studio test (27 September 2026). French untested.

Choose an underlying speech engine

★★☆☆☆ 2/5 for this task

Editorial task fit. Spanish now works in one Studio test, but its required encoder is non-commercial and it needs a large GPU or the Apple Silicon branch, so it stays below IndexTTS 2 / 2.5 (three stars) and Qwen3-TTS (four stars) as a general engine. Its instruction editing of existing speech goes further than any other engine in this task.

Decision: Use the public demo to repair or extend an existing voice track (one word, one emotion, extra takes) and judge it by ear. Do not build a commercial product on it while the Qwen encoder licence stays research-only.

Limits and things to check

The editing range is what sets it apart: replace spoken words, change emotion or timbre, reduce an accent, add or remove breaths and laughs, convert whisper, denoise and keep one speaker, each through a written instruction. Studio test, 27 September 2026, AuK Base on the public demo: a 15-second Spanish reference from an existing Studio speaker and a 20-word Spanish sentence. Speech recognition returned every word exactly; median pitch 204 Hz against 198 Hz for the reference (+6 Hz, half the 12 Hz drift measured between takes of the engine that invented that voice); loudness -11.5 against -12.1 LUFS; spectral centroid 1,075 Hz against 941 Hz, so slightly brighter. Timbre match still needs a listening check. The authors' report compares it with Qwen3-TTS and VoxCPM2 on English and Chinese test sets only; Studio has not verified those results.

Evidence

Official model card, repository, licence files and technical report reviewed; one Spanish zero-shot take measured by Studio on the public demo (27 September 2026), not yet judged by ear

Catalogue review: 2026-09-27. The evidence note identifies what was verified. This date does not imply a fresh tool test or verification of every linked source.

Put it to work