English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SciDiagramEdit: Learning to Edit Scientific Diagrams from Natural Paper Revisions

Forum topic · 小凯 · 2026-07-18

Summary

SciDiagramEdit is a benchmark and skill-evolution framework for instruction-driven editing of scientific figures, introduced in an arXiv paper (2607.15272) by Yasheng Sun, Zezi Zeng, Jürgen Schmidhuber, and colleagues. The work addresses the routine but time-consuming task of editing figures during manuscript revision—relabeling components, rearranging panels, and restyling visuals. Automating this via natural-language instructions is difficult because scientific figures are dense infographics combining schematics, plots, photos, captions, and arrows under a tight visual grammar. SciDiagramEdit operates on editable vector sources, letting users inspect and co-edit individual primitives alongside an agent. The benchmark is built from before/after figure pairs extracted from arXiv version histories, each anchored to the authors' actual revision intent. The framework uses skill-evolution-based agent learning, where an agent proposer iteratively refines the agent's skill specifications from multi-turn execution trajectories. Resulting skills progressively improve editing accuracy on a validation set, demonstrating that natural paper revisions are an effective training signal for instruction-driven figure editing.

Overview

Field: NLP Authors: Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber Published: 2026-07-16 arXiv: 2607.15272

Abstract

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight visual grammar to advance a specific argument.

To address this, the authors present SciDiagramEdit, a benchmark and skill-evolution framework that learns from natural paper revisions and operates on the figure's editable vector source, where users can inspect and co-edit individual primitives alongside the agent.

Key Points

  • Benchmark from real revisions: The benchmark extracts before/after figure pairs from arXiv version histories, with each pair anchored to the original authors' own revision intent.
  • Editable vector source: Unlike raster-based approaches, the framework works directly on editable vector graphics, enabling fine-grained inspection and co-editing of individual primitives (shapes, labels, arrows) together with the agent.
  • Skill-evolution agent learning: To handle the diversity of editing instructions, an agent proposer continuously refines the agent's skill specifications based on multi-turn execution trajectories.
  • Results: The evolved skills progressively improve editing accuracy on a validation set, showing that natural paper revisions serve as an effective training signal for instruction-driven figure editing.
  • Links

  • arXiv paper: https://arxiv.org/abs/2607.15272

Tags

#nlp#scientific-figures#diagram-editing#benchmark#agent-learning#arxiv#document-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433589