StepAudio 3Gen

TURN IMAGINATION INTO SOUND

模型介绍

Sounds human. Goes beyond speech.

全要素生成

写下整个场景,
直接听见成片。

人物、动作、空间与音乐不再分开制作。模型理解一段完整描述,并在同一条时间线上完成组织。

评测表现

在盲测两两对比评估中,StepAudio 3 Gen 在 TTS 与 Voice Design 两项任务上,分别以 1755.3 和 1668.5 的 Elo 分数位列图示参评模型第一,图示对阵的总体胜率分别为 82.0% 和 75.5%,展现出在语音真人感与音色设计方面的领先表现。

Human-like TTS

Chinese Human-Likeness Arena: StepAudio 3 Gen Elo 1755.3, ranked first among the six models shown, 211.1 points ahead of the next shown model. Head-to-head results across the five opponents shown: 410 wins, 39 ties and 51 losses over 500 trials, for an 82.0% win rate including ties in the denominator. StepAudio 3 Gen win rates: Qwen 73%, StepAudio 2.5 TTS 78%, Doubao 79%, Inworld 90%, MiniMax 90%.

Voice Design

Voice Design Arena: StepAudio 3 Gen Elo 1668.5, ranked first among the five models shown; lead over the next shown model 97.0 points. Head-to-head results across four shown opponents: 148 wins, 9 ties and 39 losses in 196 trials, with 49 trials per opponent. Overall win rate 75.5%, counting ties in the denominator. StepAudio 3 Gen win rates: Qwen3-TTS-VD 65.3%, MOSS-VoiceGenerator 71.4%, Ming-omni-tts-0.5B 79.6%, VoiceSculptor 85.7%.

准备好开始创作了吗?

即将开放