Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
Input | string | :heavy_check_mark: | Text to synthesize | Hello world |
InputReferences | []components.SpeechInputReference | :heavy_minus_sign: | Reference content for stateless voice cloning: one input_audio part carrying the voice sample, optionally accompanied by one text part with its transcript. Only routed to endpoints that support voice cloning. | [ { βinput_audioβ: { βdataβ: βdata:audio/wav;base64,UklGRuQXDABXQVZFβ¦β }, βtypeβ: βinput_audioβ }, { βtextβ: βI used to rule the world.β, βtypeβ: βtextβ } ] |
Model | string | :heavy_check_mark: | TTS model identifier | mistralai/voxtral-mini-tts-2603 |
Provider | *components.SpeechRequestProvider | :heavy_minus_sign: | Provider-specific passthrough configuration | |
ResponseFormat | *components.SpeechRequestResponseFormat | :heavy_minus_sign: | Audio output format | pcm |
Speed | *float64 | :heavy_minus_sign: | Playback speed multiplier. Only used by models that support it (e.g. OpenAI TTS). Ignored by other providers. | 1 |
Voice | *string | :heavy_minus_sign: | Voice identifier (provider-specific). | en_paul_neutral |