English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks

Forum topic · 小凯 · 2026-09-13

Summary

Benchmark Radar is a living database and search engine for discovering and retrieving AI benchmarks, introduced in an arXiv paper (2609.11115). It covers LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases from 37 sources (13 direct connectors and 24 first-party research/engineering feeds) with a searchable catalog containing 1,283 source records from 4 benchmark catalogs and 12,916 numeric observations across 790 records. It preserves source identities and citations so users can inspect candidate benchmarks and their evaluation evidence, and includes model card mentions and score histories. The authors audit the full catalog and analyze benchmark saturation, adoption trends, and limits of score comparisons. Released tools include a web dashboard with leaderboard, Pareto frontier views, daily feeds, a CLI for offline queries, and reproducible analysis.

Paper Overview

Research areas: cs.AI, cs.IR Authors: Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu arXiv: 2609.11115

Abstract

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.

The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence.

Key Details

  • Daily discovery sources: 37 total — 13 direct connectors and 24 first-party research and engineering feeds
  • Catalog size: 1,283 source records drawn from 4 benchmark catalogs
  • Numeric observations: 12,916 observations across 790 records
  • Analysis and Release

    The paper describes the collection and retrieval pipeline, audits the full catalog, and examines benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation.

    The release includes:

  • A web dashboard with a benchmark leaderboard
  • A Pareto frontier view of score against measured use
  • Saturation and trend views
  • Daily feeds and downloadable evidence
  • A command-line interface (CLI) for offline queries
  • Reproducible analysis
---

*Auto-collected on 2026-09-13*

Tags

#ai-benchmarks#llm-evaluation#search-engine#benchmark-database#arxiv#model-evaluation#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634797