Harness design for long-running application development

TL;DR AI
2 min readKey summary
Anthropic described a multi-agent harness that helps Claude build apps over many hours while also improving frontend design quality.
The setup separates planner, generator, and evaluator roles and uses a GAN-inspired loop to refine outputs.
To reduce long-task drift and context anxiety, it relies on context resets and structured handoffs between agents.
The piece argues that naive self-evaluation breaks down, especially for subjective design work, so evaluation must use explicit criteria.
