Beyond Retrieval: A Multitask Benchmark and Model for Code Search
TL;DR AI
2 min readKey summary
Researchers released CoREB, a new contamination-limited benchmark for code search.
It evaluates text-to-code, code-to-text, and code-to-code retrieval across five programming languages with graded judgments.
The benchmark goes beyond first-stage retrieval to better reflect real developer search behavior and avoid data contamination.
Results show short keyword-style queries remain difficult for current models, while the fine-tuned CoREB-Reranker beats existing baselines more consistently than off-the-shelf rerankers.
