CryptanalysisBench Introduces a Framework to Measure LLM Cryptanalysis Capabilities

TL;DR AI
2 min readKey summary
Researchers introduced CryptanalysisBench, a three-tier benchmark for measuring large language models’ cryptanalysis skills.
The benchmark spans known practical breaks, production-strength cryptographic primitives, and frontier-level challenges.
Results on models including Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM 5.2 showed much better performance on easier or scaled-down tasks than on full-strength problems.
The framework is meant to track cryptographic weakness-finding ability and improve both defensive research and risk assessment.
