Benchmark test “mcpbench” measures the performance of GPT-5.6 and Claude Opus from the perspective of “Can they build an MCP server?”

TL;DR AI
2 min readKey summary
Cloudflare engineer Matt Carey released mcpbench, a new benchmark for measuring how well AI models can implement MCP clients and servers.
The test evaluates success rates across models such as GPT-5.6, Claude Opus, Claude Sonnet, and Kimi, using both old and new MCP specs.
Results compare performance with and without documentation, showing that newer MCP changes are harder to handle from prior knowledge alone.
The benchmark aims to gauge how accurately models can adapt to external tool integration in real-world use.
