Anthropic's new benchmark claims Claude can match human experts in bioinformatics

TL;DR AI
2 min readKey summary
Anthropic released BioMysteryBench, a 99-question bioinformatics benchmark built from real noisy data with objectively checkable answers.
Claude reached about human-expert level on tasks that experts could solve, but was much less reliable on the hardest questions.
The benchmark is meant to measure practical research ability, not memorization, and is now publicly available on Hugging Face.
The results suggest frontier AI may already help with some bioinformatics work, while high scores still do not guarantee robustness on difficult problems.
