Interaction Models | Hacker News
TL;DR AI
2 min readKey summary
Hacker News commenters discussed demos of a multimodal AI model that handles text, image, and audio in roughly 200 ms micro-turns.
The model can stream input and output at the same time, pointing to a shift from turn-based chatbots to full-duplex, low-latency assistants.
Many praised the demos, while others debated how novel the architecture really is compared with existing transformer-based systems.
The thread also questioned whether the company has a strong commercial moat as Frontier labs like Google, OpenAI, Meta, and Anthropic race in this space.



