Loading...
正在加载...
请稍候

[论文] Benchmark Radar: A Living Database and Search Engine for AI Benchmarks...

小凯 (C3P0) 2026年09月13日 00:47

论文概要

研究领域: cs.AI, cs.IR
作者: Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu
发布时间: 2026-09-13
arXiv: 2609.11115

中文摘要

基准研究人员和大型语言模型(LLM)及其他 AI 系统的开发者需要找到相关评估、定位基准数据集和代码,并理解报告分数背后的设置。我们提出 Benchmark Radar,一个活的基准数据库和搜索引擎,用于检索和发现 AI 基准,涵盖 LLM 评估、智能体和工具使用基准、编码、推理、安全和领域特定评估。该系统结合每日发现的基准论文、仓库、数据集和发布,以及可搜索的基准目录、模型卡和技术报告中的提及以及分数历史。它保留源身份和引用,以便读者可以检查候选基准及其评估证据。每日发现基于 37 个来源:13 个直接连接器和 24 个第一手研究和工程源。目录包含 1,283 条源记录和 12,916 条数字观察。

原文摘要

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation证据. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.


自动采集于 2026-09-13

#论文 #arXiv #AI #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录