OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

TL;DR AI
2 min readKey summary
Researchers introduced OpenSkillEval, an automatic benchmark for testing open-source skills on realistic downstream tasks for LLM agents.
The framework generates tasks from real-world artifacts across five application areas and evaluates 30 skills on 600+ instances.
Results show that skill usage and performance gains vary widely by model and agent framework.
The study suggests popular skills are not universally helpful and that skill choice should be evaluated in context.
