StepFun Releases StepAudio 2.5 Realtime: An End-to-End Voice Model with Roleplay-Specific RLHF and Paralinguistic Comprehension

TL;DR AI
2 min readKey summary
StepFun launched StepAudio 2.5 Realtime, a real-time end-to-end voice LLM for Chinese and English.
The model uses a single audio-in/audio-out pipeline, supports customizable personas, and is available via WebSocket API.
StepFun says it was trained with million-scale persona augmentation, roleplay-focused RLHF, and fused speech understanding and generation.
The company reports strong benchmark results, including paralinguistic comprehension, pointing to richer interactive voice assistants and roleplay use cases.
