Use with AI

Choose voice tools by the task

Compare alternatives doing the same job.

Build your own conversation stack

Start with Pipecat. Match its inputs and operating requirements to your job.

Pipecat capability illustration

Voice-agent framework

Pipecat

★★★★★ 5/5 for this task

Best fit within this set for composing a provider-selectable pipeline; engineering and deployment are still required.

Give it
Audio transport; ASR/LLM/TTS components
Get back
Working voice-agent pipeline
How it runs
Application code with chosen providers
Processing location
Depends on the selected application, backend and hosting route.
Charging model
Check application, model and compute charges separately.
Choose it when
Use when you need control over model providers and the conversation pipeline.
Watch for
Each provider adds its own cost, privacy, latency and availability constraints.

Source preview

Full evidence and recipe
Hugging Face speech-to-speech capability illustration

Reference application stack

Hugging Face speech-to-speech

★★★☆☆ 3/5 for this task

Useful for learning and prototyping; adaptation and deployment checks remain.

Give it
Audio; Models and runtime
Get back
Speech-to-speech pipeline
How it runs
Reference code and selected models
Processing location
Depends on the selected application, backend and hosting route.
Charging model
Check application, model and compute charges separately.
Choose it when
Use as a reference or starting implementation for an open-model pipeline.
Watch for
Hardware needs depend on the selected components.

Source preview

Full evidence and recipe

Choose an underlying speech engine

Start with Kokoro or CosyVoice. Match its inputs and operating requirements to your job. Other equally rated options appear below.

Kokoro capability illustration

Speech engine or client

Kokoro

★★★★☆ 4/5 for this task

Editorial task fit based on the documented role and integration requirements. Not the choice when reproducing a particular reference speaker is mandatory; check language and voice support.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Engine/client, often exposed through voice-pro
Processing location
Local when self-hosted.
Charging model
Local inference compute; check distribution terms.
Choose it when
Choose for lightweight speech synthesis using available voices.
Watch for
Not the choice when reproducing a particular reference speaker is mandatory; check language and voice support.

Source preview

Full evidence and recipe
Edge-TTS capability illustration

Speech engine or client

Edge-TTS

★★★☆☆ 3/5 for this task

Editorial task fit based on the documented role and integration requirements. This is an online client, not a local model. Do not infer a production service guarantee from a working client.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Engine/client, often exposed through voice-pro
Processing location
Online service.
Charging model
Client uses an online service; check applicable terms.
Choose it when
Use for a simple online speech-generation route when that service dependency is acceptable.
Watch for
This is an online client, not a local model. Do not infer a production service guarantee from a working client.

Source preview

Full evidence and recipe
F5-TTS capability illustration

Speech engine or client

F5-TTS

★★★☆☆ 3/5 for this task

Editorial task fit based on the documented role and integration requirements. Verify model weights, licence and reference preparation. Application code and pretrained weights may have different terms.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Engine/client, often exposed through voice-pro
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Evaluate as a self-hosted reference-voice engine inside a supported application.
Watch for
The project distinguishes MIT code from non-commercial pretrained weights. Verify the intended use against the selected model terms.

Source preview

Full evidence and recipe
CosyVoice capability illustration

Speech engine or client

CosyVoice

★★★★☆ 4/5 for this task

Editorial task fit based on the documented role and integration requirements. Pin the version and test the intended language, reference voice and streaming mode.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Engine/client, often exposed through voice-pro
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Choose when you need a controllable multilingual model route in an application you can operate.
Watch for
Pin the version and test the intended language, reference voice and streaming mode.

Source preview

Full evidence and recipe
Qwen3-TTS capability illustration

Speech-generation engine

Qwen3-TTS

★★★★☆ 4/5 for this task

Editorial task fit based on the documented role and integration requirements. VoiceDesign, CustomVoice and Base are different models. Pick the correct one; advertised cloning input length is not an output-quality guarantee.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Model runtime or compatible application
Processing location
Local/self-hosted, or selected cloud provider.
Charging model
Model compute; deployment or hosted-service charges.
Choose it when
Choose when you need instruction-led voice design or a supported cloning route.
Watch for
VoiceDesign, CustomVoice and Base are different models. Pick the correct one; advertised cloning input length is not an output-quality guarantee.

Source preview

Full evidence and recipe
Chatterbox capability illustration

Speech-generation engine

Chatterbox

★★★★☆ 4/5 for this task

Editorial task fit based on the documented role and integration requirements. Versions and runtimes vary; compare pronunciation and delivery by ear before production.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Model runtime or compatible application
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Choose for reference-voice speech generation when its supported version and language fit.
Watch for
Versions and runtimes vary; compare pronunciation and delivery by ear before production.

Source preview

Full evidence and recipe
IndexTTS 2 / 2.5 capability illustration

Speech-generation engine

IndexTTS 2 / 2.5

★★★☆☆ 3/5 for this task

Editorial task fit based on the documented role and integration requirements. Check the released model and supported controls. A documented feature is not a verified local integration.

Give it
Text; Reference audio where supported
Get back
Speech audio
How it runs
Model runtime or compatible application
Processing location
Self-hosted/runtime dependent.
Charging model
Model compute; check the selected release and weights.
Choose it when
Evaluate when precise control over the generated speech is needed.
Watch for
Do not transfer capabilities between versions: the 2.0 release notes say precise duration control was not enabled; 2.5 documents a speaking-speed duration factor. Check supported languages.

Source preview

Full evidence and recipe
AuK-Flash capability illustration

Speech-generation engine

AuK-Flash

★★☆☆☆ 2/5 for this task

Editorial task fit. Spanish now works in one Studio test, but its required encoder is non-commercial and it needs a large GPU or the Apple Silicon branch, so it stays below IndexTTS 2 / 2.5 (three stars) and Qwen3-TTS (four stars) as a general engine. Its instruction editing of existing speech goes further than any other engine in this task.

Give it
Written instruction with the text to speak; Reference audio to clone, or source audio to edit
Get back
Speech audio at 24 kHz; Edited, cleaned or separated recording
How it runs
Self-hosted Python runtime on a CUDA GPU, or the MLX branch on Apple Silicon; Gradio, ComfyUI and SGLang-Omni integrations
Processing location
Self-hosted if run locally; the public demo runs on Hugging Face.
Charging model
MIT weights, self-hosted GPU or Mac compute. Commercial use also needs a separate licence for the Qwen encoder.
Choose it when
Use the public demo to repair or extend an existing voice track (one word, one emotion, extra takes) and judge it by ear. Do not build a commercial product on it while the Qwen encoder licence stays research-only.
Watch for
The model card, instruction templates and published evaluations cover English and Chinese only. One Studio test on 27 September 2026 produced correct Spanish (see comparison notes); French is untested.; It needs the Qwen2.5-Omni-3B encoder, whose Qwen Research License allows non-commercial use only. The AuK weights themselves are MIT.; Output length is fixed in advance: pass a target duration, or let the optional prompt enhancer estimate it through a separate language model. No streaming. The demo caps a take at 30 seconds.; Launch-day peak GPU memory was about 25 GB; a 16 September update cut the encoder by about 7.5 GB and added CPU offload. Apple Silicon runs from a separate MLX branch; audio.cpp added offline GGUF runs on 20 September.; The free demo allows about one 30-second take a day without a Hugging Face token (measured 27 September 2026).

Source preview

Full evidence and recipe

Conversation services

Start with ElevenLabs Agents or OpenAI Realtime. Match its inputs and operating requirements to your job. Other equally rated options appear below.

ElevenLabs Agents capability illustration

Hosted agent platform

ElevenLabs Agents

★★★★☆ 4/5 for this task

A managed platform fits guide deployment, but the actual page still needs microphone, interruption and recovery tests.

Give it
Live speech; Knowledge and tools
Get back
Spoken responses; Tool actions
How it runs
Hosted agent configuration and application integration
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Use for a hosted guide with configured knowledge and tools.
Watch for
Studio guide instances are not all proven live; test the deployed instance separately.

Source preview

Full evidence and recipe
OpenAI Realtime capability illustration

Realtime model API

OpenAI Realtime

★★★★☆ 4/5 for this task

Strong fit for custom integration; requires more application work than a hosted agent builder.

Give it
Live audio; Instructions; Tool definitions
Get back
Audio/text responses; Tool calls
How it runs
Realtime API plus application transport
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Choose for a custom conversational application where you control the integration.
Watch for
Session behaviour, voices, access and supported models must be checked for the selected API release.

Source preview

Full evidence and recipe
Gemini Live capability illustration

Realtime multimodal API

Gemini Live

★★★★☆ 4/5 for this task

A differentiated route for multimodal interaction; not an acoustic-quality ranking against other voice APIs.

Give it
Streaming audio; Optional visual context
Get back
Audio/text responses
How it runs
Live API and application integration
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Choose when live voice must interact with visual context in a Google-based integration.
Watch for
Verify model availability, session limits and interruption in the actual application.

Source preview

Full evidence and recipe
Grok Voice Agent API capability illustration

Realtime model API

Grok Voice Agent API

★★★☆☆ 3/5 for this task

A relevant alternative; no current Studio guide integration was tested here.

Give it
Live audio; Tools
Get back
Spoken responses; Tool calls
How it runs
Voice Agent API
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Evaluate when an xAI-based voice integration is specifically useful.
Watch for
Do not infer API access from a consumer subscription.

Source preview

Full evidence and recipe

Dictate into an app

Start with FreeFlow. Match its inputs and operating requirements to your job.

FreeFlow capability illustration

Dictation application

FreeFlow

★★★★★ 5/5 for this task

Existing recorded daily workflow plus configurable providers. This is a workflow recommendation, not a claim of best recognition accuracy.

Give it
Microphone speech
Get back
Text in the active app
How it runs
Mac app with Groq or OpenAI-compatible providers
Processing location
Configured provider: local/self-hosted or cloud.
Charging model
Free application; provider or local-compute costs depend on configuration.
Choose it when
Default for the existing dictation workflow; choose the backend deliberately.
Watch for
Cloud use depends on backend; local providers are also supported.

Source preview

Full evidence and recipe
FluidVoice capability illustration

Dictation application

FluidVoice

★★★★☆ 4/5 for this task

A strong local-first alternative. Engine and enhancement settings, rather than the app name, determine privacy and results.

Give it
Microphone speech; Recorded audio
Get back
Dictated text; File transcripts
How it runs
Mac app; local models or optional cloud enhancement
Processing location
Local-first; optional cloud enhancement must be checked.
Charging model
Application/local-compute costs; optional providers have their own charges.
Choose it when
Choose when local processing is the priority; compare your own vocabulary against FreeFlow.
Watch for
Confirm language/model support and speaker-label output in the installed version.

Source preview

Full evidence and recipe
Wispr Flow capability illustration

Dictation application

Wispr Flow

★☆☆☆☆ 1/5 for this task

Personal suitability only: Frank previously rejected its recognition. No claim about other users or the current release.

Give it
Speech
Get back
Dictated text
How it runs
Desktop/mobile product; plan and platform vary
Processing location
Check the service and privacy settings before sensitive dictation.
Charging model
Service plan; check the current tier.
Choose it when
Reconsider only if a new trial fixes the recognition problems previously reported.
Watch for
Current version has not been retested; this rating is based on recorded user preference.

Source preview

Full evidence and recipe

Generate narration and cloned speech

Start with ElevenLabs. Match its inputs and operating requirements to your job.

ElevenLabs capability illustration

Hosted voice platform

ElevenLabs

★★★★★ 5/5 for this task

Fits the existing Studio narration workflow and exposes a documented API. Five stars means this workflow default, not the best voice in a blind test.

Give it
Text; Optional authorised reference audio
Get back
Speech audio
How it runs
Web application and API
Processing location
Hosted voice processing.
Charging model
Hosted plan and usage allowances.
Choose it when
Default for the established hosted narration route; use voice approval before producing video.
Watch for
Choose the appropriate product: narration, dubbing, transcription and agents are different tasks. Review the output by ear.

Source preview

Full evidence and recipe
Voicebox capability illustration

Local voice application

Voicebox

★★★★☆ 4/5 for this task

A local application is easier to adopt than assembling engines directly; exact features depend on the installed release.

Give it
Text; Reference audio
Get back
Generated speech; Edited voice stories
How it runs
Desktop application and integrations
Processing location
Check local model selection and any connected services.
Charging model
Application plus selected model/compute costs.
Choose it when
Start here for a local voice-production interface already recorded in the kit.
Watch for
The upstream app has expanded beyond TTS. Check configured backends before promising fully local operation.

Source preview

Full evidence and recipe
VoiceStudio capability illustration

Multi-engine voice application

VoiceStudio

★★★★☆ 4/5 for this task

Broader workflow coverage than a single TTS engine, with more configuration and model choices.

Give it
Text; Reference audio; Video or audio for dubbing
Get back
Speech; Dubs; Transcripts
How it runs
Application with separately selected engines
Processing location
Local engines, unless connected to a remote backend.
Charging model
Local compute and storage; engine/weight terms vary.
Choose it when
Choose when you need to compare or route several local engines in one workspace.
Watch for
Capabilities and licensing depend on the selected engine and weights; language catalogue size is not equal tested quality.

Source preview

Full evidence and recipe

Isolate voice from mixed audio

Start with Demucs. Match its inputs and operating requirements to your job.

voice-pro capability illustration

Audio workflow application

voice-pro

★★★★☆ 4/5 for this task

Choose the integrated interface when you want vocal separation plus surrounding audio tools. Demucs is an underlying separation engine, not a separate competing app.

Give it
Audio or video
Get back
Separated vocals; Subtitles; Speech
How it runs
Gradio interface around several tools
Processing location
Depends on the selected application, backend and hosting route.
Charging model
Check application, model and compute charges separately.
Choose it when
Use the interface for an assisted isolation workflow; use Demucs directly for a scripted separation step.
Watch for
Edge-TTS is an online client; installing the bundle does not make every feature local.

Source preview

Full evidence and recipe
Demucs capability illustration

Source-separation model

Demucs

★★★★★ 5/5 for this task

The right specialised layer for vocal separation; compare stem quality on the actual recording.

Give it
Mixed audio
Get back
Separated audio stems
How it runs
Local model or an application such as voice-pro
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Use before transcription or mixing when music masks speech.
Watch for
Separation can introduce artefacts; it cannot recover every obscured word.

Source preview

Full evidence and recipe

Open dialogue models

Start with Moshi. Match its inputs and operating requirements to your job.

Moshi capability illustration

Dialogue model and framework

Moshi

★★★☆☆ 3/5 for this task

Distinct from a modular ASR-LLM-TTS stack; useful to evaluate, with substantial deployment work.

Give it
Live audio
Get back
Spoken dialogue
How it runs
Self-hosted model/runtime
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Evaluate when research-level control over full-duplex dialogue is the objective.
Watch for
Validate languages, hardware and behaviour on the intended conversation.

Source preview

Full evidence and recipe
Nemotron VoiceChat capability illustration

Dialogue model

Nemotron VoiceChat

Not rated

Official model identified, but there is no local integration or comparative test in this review.

Give it
Audio stream
Get back
Conversational audio
How it runs
Model/runtime deployment
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Keep as a candidate for a self-hosted dialogue evaluation.
Watch for
Verify the exact model card, hardware and supported deployment.

Source preview

Full evidence and recipe

Speaker-labelled transcription

Start with VibeVoice-ASR. Match its inputs and operating requirements to your job.

VibeVoice-ASR capability illustration

Speech-recognition model

VibeVoice-ASR

★★★☆☆ 3/5 for this task

Good capability fit on paper, but the required output needs verification in the chosen runtime before becoming the default.

Give it
Long audio recording
Get back
Structured transcript; Speaker labels; Timestamps
How it runs
Model/runtime deployment
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Evaluate when who-spoke-when output matters.
Watch for
Do not advertise structured output as proven locally. Streaming ASR is a separate release.

Source preview

Full evidence and recipe
VibeVoice-ASR-Streaming capability illustration

Streaming recognition model

VibeVoice-ASR-Streaming

Not rated

New release identified; no local suitability evidence yet.

Give it
Incoming audio stream
Get back
Incremental transcript
How it runs
Streaming ASR integration
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Evaluate for live transcription where streaming is necessary.
Watch for
Language coverage and hotword behaviour differ by release; untested here.

Source preview

Full evidence and recipe

Transcribe a recording

Start with Whisper or Groq speech-to-text. Match its inputs and operating requirements to your job. Other equally rated options appear below.

Whisper capability illustration

Speech-recognition model

Whisper

★★★★☆ 4/5 for this task

A clear, reusable transcription baseline. Language and hardware affect results; no universal accuracy ranking is implied.

Give it
Audio file
Get back
Transcript; Timed segments
How it runs
Local runtime or a hosted service running Whisper
Processing location
Local when run locally; cloud when used through a hosted endpoint.
Charging model
Local compute, or hosted-service charges.
Choose it when
Use as a general multilingual transcription engine or baseline.
Watch for
Check names and technical terms; distinguish the open model from hosted transcription products.

Source preview

Full evidence and recipe
Groq speech-to-text capability illustration

Hosted inference service

Groq speech-to-text

★★★★☆ 4/5 for this task

Practical fit for the recorded automated ingestion route; no current speed or cost claim.

Give it
Audio file
Get back
Transcript
How it runs
Hosted API
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Use when cloud processing is acceptable and an existing API workflow is useful.
Watch for
Audio leaves the local device; service quotas and supported models must be checked.

Source preview

Full evidence and recipe
Parakeet TDT capability illustration

Speech-recognition model

Parakeet TDT

★★★★☆ 4/5 for this task

A focused ASR alternative; compare the same recording and language rather than vendor headline benchmarks.

Give it
Audio
Get back
Transcript
How it runs
Model runtime or host application
Processing location
Self-hosted if deployed locally, otherwise provider-dependent.
Charging model
Compute and model/weight terms; hosted wrappers charge separately.
Choose it when
Use through FluidVoice or a supported runtime when the selected model covers the language.
Watch for
Select the version explicitly; timestamps and formatting depend on the integration.

Source preview

Full evidence and recipe
VibeASR.cpp / BitNet capability illustration

CPU inference runtime

VibeASR.cpp / BitNet

★★★★☆ 4/5 for this task

The existing local record demonstrates a useful route. Output expectations must follow that runtime, not the broader family.

Give it
Recorded audio
Get back
Transcript
How it runs
Local CPU runtime
Processing location
Local execution.
Charging model
Local CPU compute and storage.
Choose it when
Use for the recorded local transcription route when a plain transcript is sufficient.
Watch for
The private measured record remains authoritative for actual output fields and performance.

Source preview

Full evidence and recipe

Translate and dub a recording

Start with ElevenLabs Dubbing. Match its inputs and operating requirements to your job.

ElevenLabs Dubbing capability illustration

Hosted dubbing workflow

ElevenLabs Dubbing

★★★★☆ 4/5 for this task

A complete managed dubbing route; compare translation, timing and pronunciation before publication.

Give it
Audio or video; Optional approved transcript
Get back
Translated dub
How it runs
Hosted application/API
Processing location
Provider-hosted processing; check the selected route.
Charging model
Hosted plan or usage charges; current price not quoted.
Choose it when
Choose managed dubbing when hosted processing and manual review are acceptable.
Watch for
Documented dubbing is not realtime. Language support and limits can change.

Source preview

Full evidence and recipe

Unresolved identities

No evidence-based starting preference is assigned yet. Compare the documented inputs, outputs and operating requirements, then verify the option before using it.

Meta Voice Transcribe capability illustration

Unresolved product label

Meta Voice Transcribe

Not rated

An ambiguous name cannot support a fair comparison.

Give it
Unverified
Get back
Unverified
How it runs
Unverified
Processing location
Depends on the selected application, backend and hosting route.
Charging model
Check application, model and compute charges separately.
Choose it when
Keep for follow-up identification.
Watch for
No current primary source identified.

Capability diagram

Full evidence and recipe