Summary
WildClawBench (arXiv:2505.07235) is a benchmark for evaluating LLM and vision-language agents that operate through command-line interface (CLI) harnesses. Unlike existing agent benchmarks that rely on synthetic sandboxes, short-horizon tasks, mock-service APIs, and final-answer checking, WildClawBench is a native-runtime benchmark featuring 60 human-authored, bilingual, multimodal tasks. It is designed to test whether agents can complete realistic, long-horizon work in the same runtimes where they are actually deployed, rather than in idealized test environments. Authored by Shuangrui Ding, Xuanlang Dai, and Long Xing in the NLP field, the benchmark addresses a key gap in agent evaluation: measuring performance on extended, real-world workflows with genuine tools and services. This makes it a valuable resource for researchers developing autonomous agents and CLI-based assistants who need more realistic evaluation methodology.
Overview
- Research area: NLP
- Authors: Shuangrui Ding, Xuanlang Dai, Long Xing
- Published: 2025-05-09
- arXiv: 2505.07235
Abstract
Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However, most agent benchmarks still rely on synthetic sandboxes, short-horizon tasks, mock-service APIs, and final-answer checks, leaving open whether agents can complete realistic long-horizon work in the runtimes where they are deployed.
This work presents WildClawBench, a native-runtime benchmark of 60 human-authored, bilingual, multimodal tasks spanning real-world, long-horizon agent workflows.
Key Contributions
- Native-runtime evaluation: Tasks run in the actual deployment runtime rather than isolated sandboxes.
- Human-authored tasks: 60 bilingual, multimodal tasks designed by humans to reflect realistic work.
- Long-horizon focus: Emphasis on extended, multi-step workflows instead of short, single-answer tasks.
Links
- Paper: https://arxiv.org/abs/2505.07235
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177619922