How consistent the voice sounds across the generation. Lower = more emotional variation and expressiveness (but can sound erratic). Higher = steady, predictable tone.
similarity_boost
0.0 - 1.0
0.75
How closely to match the original voice sample. Higher sounds more like the source voice but may amplify audio artifacts or background noise from the original recording.
style
0.0 - 1.0
0.0
Exaggerates the unique characteristics of the voice’s speaking style (v2+ and v3 models only). Higher values make the voice more “characterful” but can reduce stability.
speed
0.25 - 4.0
1.0
Speech speed multiplier. 1.0 = normal speed. Range is 0.25-4.0 for the REST API; the Agents Platform restricts to 0.7-1.2.
use_speaker_boost
boolean
true
Post-processing that enhances voice clarity and similarity to the original. Generally leave this on unless you’re experiencing artifacts.