Switch language한국어
Back to the list

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced PlanningBench, a taxonomy-guided framework that turns real planning workflows into scalable, diverse, and verifiable tasks for LLMs.

  2. The benchmark uses constraint-driven synthesis and automatic verification to create controllable planning data with adaptive difficulty.

  3. Frontier open-source and closed-source models were evaluated on the benchmark, revealing room for improvement in complex planning and instruction following.

  4. Training with reinforcement learning on verified examples further improved model performance on planning tasks.

Read the original