← All model drops
Live drop
StepAudio 3 Realtime
StepFunDropped Sep 15, 2026
Full-duplex real-time voice model with think-while-speaking and native Voice Agent tool calling.
About this release
Live drop
StepFunDropped Sep 15, 2026
Full-duplex real-time voice model with think-while-speaking and native Voice Agent tool calling.
About this release
StepAudio 3 Realtime is StepFun's flagship real-time conversational audio model. It unifies audio understanding, turn management, adaptive reasoning, and tool calling in a single bidirectional stream, supporting natural interruptions, backchannels, emotion/paralinguistic cues, and background tasks that continue while the conversation proceeds. Part of the broader StepAudio 3 family that also includes ASR, TTS, Gen, and Music variants.
Read the official StepFun launch pageFirst builds incoming — this lane fills in automatically as people start building with StepAudio 3 Realtime. Know one? Submit the URL.