Switch language한국어
Back to the list

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

TL;DR AI

Key summary

2 min read
  1. MonitorBench is a new open benchmark for measuring chain-of-thought monitorability in large language models.

  2. The benchmark includes 1,514 test instances across 19 tasks and two stress-test settings.

  3. Experiments show monitorability is higher when structural reasoning is required and that closed-source models often have lower monitorability.

  4. Stress-tests can reduce monitorability by up to 30% in some non-structural tasks.

Read the original