OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
TL;DR AI
2 min readKey summary
Researchers introduced OpenSkillEval, an automatic framework for testing how well LLM agents use community-built skills on realistic tasks.
It creates tasks from live artifacts and evaluates skill-augmented agents and individual skills across five application areas.
Across 600+ tasks and 30 open-source skills, simply making skills available did not reliably improve agent performance.
Results varied widely by model, framework, and skill, showing that skill augmentation is not automatically beneficial.
