PAPER·August 3, 2026Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding AssistantsHugging Face Papers
TECH·July 31, 2026DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis | Hacker NewsHacker News
CODING·May 21, 2026AI-generated accessibility, an update — frontier models still fail, but skills change the gameDev.to
PAPER·May 21, 2026LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial HardeningHugging Face Papers
PAPER·May 18, 2026DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic RulesHugging Face Papers
TECH·May 15, 2026Show HN: Find the best local LLM for your hardware, ranked by benchmarks | Hacker NewsHacker News
TECH·May 15, 2026Poetiq’s Meta-System Automatically Builds a Model-Agnostic Harness That Improved Every LLM Tested on LiveCodeBench Pro Without Fine-TuningMarkTechPost
PAPER·May 6, 2026PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent ExaminationHugging Face Papers
PAPER·April 30, 2026Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum TransactionsarXiv
TECH·April 26, 2026Top 7 Benchmarks That Actually Matter for Agentic Reasoning in Large Language ModelsMarkTechPost