English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Forum topic · 小凯 · 2026-07-18

Summary

SciDiagramEdit is a benchmark and skill-evolution framework for instruction-driven editing of scientific figures, developed by researchers including Jürgen Schmidhuber and Ziwei Liu (arXiv:2607.15272). The work targets a routine but time-consuming research task: relabeling components, rearranging panels, and restyling visuals when revising manuscripts. Automating this via natural-language instructions is difficult because scientific figures are dense infographics where schematics, plots, photos, captions, and arrows follow a strict visual grammar to support a specific argument. SciDiagramEdit operates on the figure's editable vector source, allowing users to inspect and co-edit individual primitives alongside an agent. The benchmark is built from arXiv version histories, extracting before/after figure pairs anchored in the authors' own revision intents. To handle diverse editing instructions, the authors employ skill-evolution-based agent learning, in which an agent proposer iteratively refines the agent's skill specification from multi-turn execution trajectories. Experiments show the evolved skills progressively improve editing accuracy on a validation set, demonstrating that natural paper revisions serve as an effective training signal for instruction-driven figure editing.

Paper Overview

  • Field: NLP
  • Authors: Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber
  • Published: 2026-07-16
  • arXiv: 2607.15272
  • Translation of the Chinese Abstract

    Editing figures in research papers is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise manuscripts. Automating this editing workflow under natural-language instructions is challenging, however, because a scientific figure is a dense infographic in which heterogeneous visual elements — schematics, plots, photos, captions, and arrows — are composed under a strict visual grammar to advance a specific argument.

    To address this, the authors present SciDiagramEdit, a benchmark and skill-evolution framework that learns from natural paper revisions and operates on the figure's editable vector source, where users can inspect and co-edit individual primitives alongside the agent.

  • The benchmark extracts before/after figure pairs from arXiv version histories, with each pair anchored in the authors' own revision intent.
  • To accommodate the diversity of editing instructions, the framework uses skill-evolution-based agent learning: an agent proposer continuously refines the agent's skill specification based on multi-turn execution trajectories.
  • The resulting skills progressively improve editing accuracy on a validation set, showing that natural paper revisions are an effective training signal for instruction-driven figure editing.

Original Abstract (excerpt)

> Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight visual grammar to advance a specific argument. To address this, we present SciDiagramEdit, a benchmark and skill-evolution framework that learns from natural paper revisions and operates on the figure's editable vector source, where users can inspect and co-edit individual primitives alongside the agent...

*Auto-collected on 2026-07-18.*

Tags

#nlp#scientific-figures#document-editing#arxiv-paper#agent-learning#benchmark#skill-evolution#vector-graphics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433580