How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

TL;DR AI
2 min readKey summary
Researchers ran 400 autonomous penetration-testing trials against the same honeypot and vulnerable services across four LLMs.
The models did not develop persistent refusals after re-prompting, but exploit success varied sharply by model.
Gemini 2.5 Flash-Lite and GPT-4o-mini achieved many successful exploits, while qwen2.5-coder:14b performed worse overall.
Claude Sonnet 4 was constrained by upstream API overloads, and the paper found distinct failure patterns and limited credential reuse across conversation histories.
