Thinking Machines Lab ships its first model and argues interactivity is what OpenAI gets wrong about voice

TL;DR AI
2 min readKey summary
Thinking Machines Lab previewed TML-Interaction-Small, its first model, built for more natural voice AI interactions.
The model processes audio, video, and text in 200-millisecond chunks, replacing rigid turn-taking with direct stream processing.
A separate background model handles longer reasoning, letting the interaction model stay fast and responsive in live conversations.
TML says the system beats OpenAI’s GPT-Realtime-2 and Google’s Gemini Live on interaction benchmarks like FD-bench v1.5.
The approach could support interruptions, overlapping speech, and lower-latency, context-aware responses in full-duplex voice assistants.

