I benchmarked Claude Code's caveman plugin against "be brief." | Hacker News
TL;DR AI
2 min readKey summary
A Hacker News discussion examined a benchmark of Claude Code response-compression methods, comparing the Caveman plugin with a simple “be brief” prompt.
Across 24 prompts and five arms, both approaches used nearly identical tokens and delivered similar quality, with full key-point coverage and no unsafe misses.
Commenters said the results are hard to trust because each prompt-arm pair appears to have been run only once, leaving run-to-run variance unmeasured.
The thread raises a broader question about whether prompt hacks and plugins meaningfully improve coding agents, or whether benchmark design matters more.



