PAPER·17 hours agoSecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident ResponseHugging Face Papers
PAPER·20 hours agoSkillRise: Agentic Reinforcement Learning for Cross-Task Skill EvolutionHugging Face Papers
PAPER·yesterdayOmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingarXiv
PAPER·yesterdaySecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident ResponsearXiv
PAPER·2 days agoOmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingHugging Face Papers
PAPER·2 days agoRSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-ImprovementarXiv
PAPER·3 days agoSimulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM AgentsarXiv