[论文] Benchmark Radar: A Living Database and Search Engine for AI Benchmarks...

论文概要 研究领域: cs.AI, cs.IR 作者: Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu 发布时间: 2026-09-13 arXiv: 2609.11115

论文概要

研究领域: cs.AI, cs.IR 作者: Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu 发布时间: 2026-09-13 arXiv: 2609.11115

中文摘要

基准研究人员和大型语言模型(LLM)及其他 AI 系统的开发者需要找到相关评估、定位基准数据集和代码,并理解报告分数背后的设置。我们提出 Benchmark Radar,一个活的基准数据库和搜索引擎,用于检索和发现 AI 基准,涵盖 LLM 评估、智能体和工具使用基准、编码、推理、安全和领域特定评估。该系统结合每日发现的基准论文、仓库、数据集和发布,以及可搜索的基准目录、模型卡和技术报告中的提及以及分数历史。它保留源身份和引用,以便读者可以检查候选基准及其评估证据。每日发现基于 37 个来源:13 个直接连接器和 24 个第一手研究和工程源。目录包含 1,283 条源记录和 12,916 条数字观察。

原文摘要

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation证据. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.


*自动采集于 2026-09-13*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens