Switch language한국어
Back to the list

StepFun Releases StepAudio 2.5 Realtime: An End-to-End Voice Model with Roleplay-Specific RLHF and Paralinguistic Comprehension

TL;DR AI

Key summary

2 min read
  1. StepFun launched StepAudio 2.5 Realtime, a real-time end-to-end voice LLM for Chinese and English.

  2. The model uses a single audio-in/audio-out pipeline, supports customizable personas, and is available via WebSocket API.

  3. StepFun says it was trained with million-scale persona augmentation, roleplay-focused RLHF, and fused speech understanding and generation.

  4. The company reports strong benchmark results, including paralinguistic comprehension, pointing to richer interactive voice assistants and roleplay use cases.

Read the original