Skip to main content
POST

Authorizations

Authorization
string
header
required

Modellix API Key. Format: Bearer <your_api_key>

Body

application/json

Shared Qwen-Audio 3.0 TTS request fields. Submit as an async task; poll GET /api/v1/tasks/{task_id} for the audio URL in result.resources. Qwen-Audio TTS does not expose hot_fix or Markdown filtering.

text
string
required

Text to synthesize. Required. Max 20,000 Unicode characters. Supports plain text and SSML (set enable_ssml to true) per Alibaba documentation.

Required string length: 1 - 20000
Example:

"There is a large garden behind my house."

voice
string
required

Voice ID for qwen-audio-3.0-tts-plus only. Use a Plus system voice (e.g. longanlingxin, longanlufeng) or a Plus base/cloned voice ID from the official Qwen-Audio-TTS voice list: https://help.aliyun.com/zh/model-studio/qwen-audio-tts-voice-list . Do not mix Flash system voices with Plus.

Required string length: 1 - 256
Example:

"longanlingxin"

format
enum<string>
default:mp3

Audio encoding format. Default mp3.

Available options:
mp3,
pcm,
wav,
opus
Example:

"mp3"

sample_rate
enum<integer>
default:22050

Audio sample rate in Hz. Default 22050.

Available options:
8000,
16000,
22050,
24000,
44100,
48000
Example:

24000

volume
integer
default:50

Output volume. Default 50. Range 0 (silent) to 100 (maximum).

Required range: 0 <= x <= 100
Example:

50

rate
number
default:1

Speech rate multiplier. Default 1.0. Range 0.5 (slow) to 2.0 (fast).

Required range: 0.5 <= x <= 2
Example:

1

pitch
number
default:1

Pitch multiplier. Default 1.0. Range 0.5 (lower) to 2.0 (higher).

Required range: 0.5 <= x <= 2
Example:

1

bit_rate
integer

Audio bit rate in kbps. Optional. Range 6 to 510. Only supported when format is opus; do not use for mp3, pcm, or wav.

Required range: 6 <= x <= 510
Example:

32

instruction
string

Optional speaking-style instruction. Enforced weighted length ≤100 (CJK ideographs / Han count as 2; other characters count as 1).

Example:

"Speak in a friendly customer-service tone."

language_hint
enum<string>

Target language hint for pronunciation (numbers, symbols, minor languages).

Available options:
zh,
en,
fr,
de,
ja,
ko,
ru,
pt,
th,
id,
vi,
es,
it,
ms,
fil,
ar
Example:

"en"

seed
integer
default:0

Random seed for reproducible synthesis when text, voice, and other parameters are identical. Default 0. Range 0 to 65535.

Required range: 0 <= x <= 65535
Example:

0

enable_ssml
boolean
default:false

Whether to parse text as SSML. Default false. When true, text must follow Alibaba SpeechSynthesizer SSML rules.

Example:

false

enable_aigc_tag
boolean
default:false

Embed AIGC invisible watermark into wav/mp3/opus output. Supported on Qwen-Audio 3.0 TTS Plus and Flash.

aigc_propagator
string

AIGC ContentPropagator. Only effective when enable_aigc_tag is true.

aigc_propagate_id
string

AIGC PropagateID. Only effective when enable_aigc_tag is true.

Response

Task submitted successfully. Poll GET /api/v1/tasks/{task_id} until the task completes; synthesized audio is returned on the task result.

Response object for asynchronous task submission.

code
integer
required

Response code, 0 indicates success

Example:

0

message
string
required

Response message

Example:

"success"

data
object
required

Task submission details. Poll GET /api/v1/tasks/{task_id} until status is success; audio appears in result.resources.