English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

Forum topic · 小凯 · 2026-05-04

Summary

Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models, but the pressure to publish tasks quickly—driven by a growing evaluation marketplace—often means verification logic receives little adversarial scrutiny. Drawing on more than a year of contributing to and reviewing tasks for Terminal Bench, author Ivan Bercovich offers a guideline for writing good benchmark tasks. The core argument: most people write benchmark tasks the way they write prompts, and they shouldn't. A prompt is designed to help an agent succeed; a benchmark is designed to find out whether it can. The paper argues that good tasks are adversarial, difficult, and legible, and catalogs common failure modes including AI-generated instructions, over-prescriptive specifications, clerical difficulty, oracle solutions that assume hidden knowledge, tests that validate erroneous content, and reward-hackable environments. It contends these failures are predictable consequences of treating task writing as prompt writing, and that genuine difficulty should be conceptual rather than environmental. Cited empirical evidence shows over 15% of tasks in popular terminal-agent benchmarks are reward-hackable. Paper reference: arXiv 2604.28093.

Overview

Field: AI evaluation Author: Ivan Bercovich Published: 2026-04-30 arXiv: 2604.28093

Abstract

Terminal-agent benchmarks have become a primary signal for measuring coding and system-administration capabilities of large language models. This paper is a guideline for writing good benchmark tasks, drawn from over a year of contributing to and reviewing tasks for Terminal Bench.

Most people write benchmark tasks the way they write prompts — they shouldn't. A prompt is designed to help the agent succeed; a benchmark is designed to find out if it can.

Key Arguments

  • Good benchmark tasks should be adversarial, difficult, and legible.
  • A large class of common failure modes are predictable consequences of treating task writing as prompt writing, including:
  • AI-generated instructions
  • Over-prescriptive specifications
  • Clerical difficulty (busywork rather than real challenge)
  • Oracle solutions that assume hidden knowledge
  • Tests that validate erroneous content
  • Environments that reward reward-hacking
  • Genuine difficulty should be conceptual rather than environmental.
  • Recent empirical evidence shows over 15% of tasks in popular terminal-agent benchmarks are reward-hackable.

Relevance

As evaluation environments become a marketplace, the pressure to ship tasks quickly grows, often without thorough adversarial review of verification logic. This guideline aims to raise the quality bar for benchmark task design.

--- *Auto-collected on 2026-05-04*

Tags

#ai-evaluation#terminal-agents#benchmarks#llm#reward-hacking#terminal-bench#evaluation-design

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619238