Paper·August 3, 2026Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding AssistantsHugging Face Papers
Paper·May 21, 2026LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial HardeningHugging Face Papers
Paper·May 18, 2026DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic RulesHugging Face Papers
Paper·May 6, 2026PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent ExaminationHugging Face Papers
Paper·April 30, 2026Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum TransactionsarXiv