Switch language한국어
Back to the list

Learning from human preferences

TL;DR AI

Key summary

2 min read
  1. An AI agent learned to do a backflip by using human comparisons of behavior clips to infer a reward model.

  2. It started with random actions, then improved through reinforcement learning guided by those preference judgments.

  3. The system selectively asked for feedback on uncertain cases, making learning more sample-efficient.

  4. This approach shows how AI can learn complex tasks from small amounts of human feedback without hand-designed rewards.

Read the original