Qwen custom voice design model

Qwen 3 TTS Voice Design turns casting notes into speech

Use it when the voice should be described, not picked: age, accent, texture, pace, emotion, and language all belong in the brief.

text to audio From 1 credit per 100 characters.

Voice Generation

Describe a custom voice in natural language, then synthesize speech in ten supported languages.

From 1 credit per 100 characters.

Your generated audio will appear here.

Enter a prompt and click 'Generate' to start.

Examples

Qwen 3 TTS voice-design example

Listen to the generated result, then review the script, direction, and settings used to make it.

Qwen 3 TTS Voice Design

The lighthouse keeper

An elderly lighthouse keeper voice designed entirely from a natural-language description.

Listen for
The gravel texture, coastal lilt, pauses on the ellipses, and warmth underneath the age.
Settings
Language · English
View script and direction

Voice: An elderly lighthouse keeper, male, in his seventies. A gravelly, weathered voice with a gentle coastal English lilt, slow and deliberate pacing, thoughtful pauses between phrases, warm underneath the roughness.

Prompt: Forty years I kept that light... and I still remember the night the storm took the Eleanor. The sea was black as ink — and then, out of nowhere, this tiny green flare, bobbing in the waves. A survivor. One small light... answering mine. That's the thing about lighthouses, lad. They don't save ships. They just remind people... where the land is.

Model context

What is voice design?

Voice design means describing the speaker you want instead of selecting only from a preset list. Qwen 3 TTS Voice Design turns that description into a generated voice for the supplied script.

It is useful for characters, narrators, localized reads, and any project where the voice's identity is part of the creative brief.

Model strengths

Where voice design beats presets

Use it when the voice should be described, not picked: age, accent, texture, pace, emotion, and language all belong in the brief.

  • Character voices with age, accent, texture, pacing, and emotional notes.
  • Multilingual speech where the language selection and delivery style need to match the script.
  • Narration tests where casting language is faster than browsing many preset voices.

Prompt guide

A Qwen 3 TTS casting brief

Write the voice description first

Describe age, gender presentation when relevant, accent, vocal texture, energy, pace, and emotional temperature.

Match language to script

Choose the language deliberately when the script is not obvious or when pronunciation matters.

Use punctuation for acting

Commas, ellipses, short sentences, and paragraph breaks help shape breath and emphasis.

Change one trait at a time

If the voice is close, adjust one direction such as age, pace, or warmth instead of rewriting everything.

Frequently asked questions

How is Qwen 3 TTS different from Eleven V3?

Qwen 3 TTS designs a voice from a description. Eleven V3 starts from preset voices with stability and similarity controls.

What should a voice description include?

Include age range, accent, tone, texture, speed, emotion, and the recording context you want.

Can it support character voices?

Yes. Character voices are a strong use case because the model responds to casting-style descriptions.

When should I leave language on auto?

Auto is useful for simple scripts. Pick a language manually when pronunciation or localization quality matters.