Thinking Machines shows off preview of near-realtime AI voice and video conversation with new 'interaction models'

TL;DR AI
2 min readKey summary
Thinking Machines unveiled a research preview of multimodal “interaction models” built for near-real-time voice and video conversation.
The system processes audio and visual input in 200ms chunks and is designed to listen, speak, and see at the same time.
It combines an always-on interaction model with a separate background reasoning model for slower thinking.
A limited preview will open later, with broader release planned this year.
The approach could push AI beyond turn-based chat toward more continuous, natural live interaction.
