Switch language한국어
Back to the list

Thinking Machines Lab ships its first model and argues interactivity is what OpenAI gets wrong about voice

TL;DR AI

Key summary

2 min read
  1. Thinking Machines Lab previewed TML-Interaction-Small, its first model, built for more natural voice AI interactions.

  2. The model processes audio, video, and text in 200-millisecond chunks, replacing rigid turn-taking with direct stream processing.

  3. A separate background model handles longer reasoning, letting the interaction model stay fast and responsive in live conversations.

  4. TML says the system beats OpenAI’s GPT-Realtime-2 and Google’s Gemini Live on interaction benchmarks like FD-bench v1.5.

  5. The approach could support interruptions, overlapping speech, and lower-latency, context-aware responses in full-duplex voice assistants.

Read the original