English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Forum topic · 小凯 · 2026-08-15

Summary

AutoDesign is a framework that treats multimodal content transformation (such as academic paper-to-poster generation) as a long-horizon agentic process centered on a model-harness system. Unlike static existing paradigms, AutoDesign uses a meta-harness optimizer that guides a code agent to recursively improve its harness based on rollout feedback, aligning with human design priors and accumulating reusable experience. The authors introduce PosterBench, a benchmark with a 100-paper Main Track across five disciplines and a 10-paper controlled subset. On the Main Track, AutoDesign scores 78.32, beating Claude Design by 7.45 points. Across seven code-agent-model configurations, the learned DesignHarness lifts average PosterBench scores from 54.99 to 67.39 (+12.4%). In fully autonomous mode, it runs 253 tool calls and 11 editing turns in 40 minutes for under $3, achieving average conference-poster quality, and tops a system-blind human preference study. Paper: arXiv 2608.13560.

Overview

Field: NLP Authors: Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li Published: 2026-08-13 arXiv: 2608.13560

Abstract

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability.

In this paper, the authors present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve the harness based on rollout feedback.

Benchmark and Results

To instantiate and evaluate the framework, the authors focus on the academic paper-to-poster generation task and introduce PosterBench, comprising:

  • A 100-paper Main Track spanning five disciplines
  • PosterBench-mini, a shared 10-paper subset for controlled evaluation
  • Key results:

  • On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points.
  • Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%).
  • In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under $3, reaching average conference-poster quality in human evaluation.
  • A system-blind human study demonstrates that AutoDesign achieves the highest human preference among evaluated systems.
---

*Source: zhichai.net forum, auto-collected 2026-08-15.*

Tags

#autodesign#agentic-ai#poster-generation#benchmark#nlp#paper-to-poster#meta-optimization#posterbench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633493