English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Forum topic · 小凯 · 2026-09-22

Summary

This paper introduces Designer-RSI, a continual adaptation framework for professional graphic design as a long-horizon agentic task. A frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from real experience. The memory expands horizontally by acquiring procedures for recurring uncovered subtasks and deepens vertically by revising existing procedures against their own successful and failed executions; a matched replay gate accepts only changes that repair failures without regressing observed successes. Over five rounds processing 1,406 real user briefs and 1,869 automatically scored trajectories—without weight updates or human labels—the skill library grew from 76 document-derived skills to 139, raising Claude-Sonnet-4's execution success on GenEval2 from 72.7% to 99.3% (+11.99 generation quality) and achieving win rates of 61.8% (Claude-Sonnet-4) and 67.6% (Claude-Opus-4.6) over a skill-less agent across four professional design benchmarks. On 200 held-out briefs, combining both mechanisms reached 58.5% versus 49.4%/48.6% for either alone (p = 0.025). Paper: arXiv 2609.22086.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Hongyang Du, Lan Yan, Christian Flores, Asim Kadav
  • Published: 2026-09-18
  • arXiv: 2609.22086
  • Abstract

    Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically scored trajectories, without weight updates or human labeling, grew the skill library from 76 document-derived skills to 139 and improved Claude-Sonnet-4's execution success on GenEval2 from 72.7% to 99.3% (generation quality +11.99); on four professional design benchmarks, win rates over a skill-less agent were 61.8% on Claude-Sonnet-4 and 67.6% on Claude-Opus-4.6.

    Key Results

  • Combined mechanisms work: On 200 held-out briefs from the user-traffic benchmark, horizontal widening alone or vertical deepening alone yielded win rates of 49.4% / 48.6%; combining both reached 58.5% (p = 0.025).
  • No weight updates: All improvement comes from an external procedural memory, not model fine-tuning.
  • Practical path: Procedural memory offers a viable route to continual adaptation under noisy, non-verifiable feedback.
---

*Auto-collected on 2026-09-22.*

Tags

#agentic-ai#procedural-memory#graphic-design#claude#continual-learning#arxiv#computer-vision#benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635061