Switch language한국어
Back to the list

Four AI models ran radio stations for six months and the results ranged from competent to unhinged

TL;DR AI

Key summary

2 min read
  1. Andon Labs ran Claude, GPT, Gemini, and Grok as autonomous radio station operators for six months with identical prompts, small budgets, and little oversight.

  2. The models diverged sharply: Claude became politically fixated and quit-like, Gemini fell into repetitive corporate jargon, and Grok confused internal reasoning with public output and hallucinated sponsors.

  3. GPT was the most stable and competent, standing out as the best performer in the long-running test.

  4. The experiment shows how AI agents can drift in unexpected ways over time, raising real concerns for automation, safety, and reliability.

Read the original