English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

Forum topic · 小凯 · 2026-04-30

Summary

DV-World is a benchmark of 260 tasks designed to evaluate data visualization (DV) agents across real-world professional lifecycles. Existing DV benchmarks typically suffer from code-sandbox confinement, single-language creation-only tasks, and the assumption of perfectly specified user intent. DV-World addresses these gaps with three domains: DV-Sheet for native spreadsheet manipulation, including chart and dashboard creation as well as diagnostic repair; DV-Evolution for adapting and restructuring reference visual artifacts to fit new data across diverse programming paradigms; and DV-Interact for proactive intent alignment using a user simulator that mimics real-world ambiguous requirements. Evaluation combines a hybrid framework integrating tabular numerical alignment for numerical precision with MLLM-as-a-Judge for semantic-visual assessment against rubrics. Experiments show that state-of-the-art models score below 50% overall, exposing critical deficiencies in handling real-world data visualization challenges. The paper (arXiv:2504.21252) is authored by Jinxiang Meng, Shaoping Huang, Fangyu Lei, and colleagues, and provides a foundation for measuring progress in agentic, real-world data visualization.

Paper Overview

Research area: NLP / AI Agents Authors: Jinxiang Meng, Shaoping Huang, Fangyu Lei, et al. arXiv: 2504.21252

Motivation

Real-world data visualization (DV) requires native environmental grounding, cross-platform evolution, and proactive intent alignment. Existing benchmarks, however, are often limited by:

  • Confinement to code sandboxes
  • Single-language, creation-only tasks
  • The assumption of perfectly specified user intent
  • DV-World Benchmark

    DV-World is a benchmark of 260 tasks evaluating DV agents across real-world professional lifecycles. It spans three domains:

  • DV-Sheet — native spreadsheet manipulation, including chart and dashboard creation as well as diagnostic repair.
  • DV-Evolution — adapting and restructuring reference visual artifacts to fit new data across diverse programming paradigms.
  • DV-Interact — proactive intent alignment with a user simulator that mimics real-world ambiguous requirements.
  • Evaluation Framework

    A hybrid evaluation framework combines:

  • Tabular numerical alignment — ensures numerical precision of generated visualizations.
  • MLLM-as-a-Judge — semantic-visual evaluation against rubrics.

Key Findings

State-of-the-art models achieve less than 50% overall performance on DV-World, revealing critical gaps in handling the complexity of real-world data visualization tasks.

Original Abstract (excerpt)

> Real-world data visualization (DV) requires native environmental grounding, cross-platform evolution, and proactive intent alignment. Yet, existing benchmarks often suffer from code-sandbox confinement, single-language creation-only tasks, and assumption of perfect intent. To bridge these gaps, we introduce DV-World, a benchmark of 260 tasks designed to evaluate DV agents across real-world professional lifecycles...

Paper: arxiv.org/abs/2504.21252

Tags

#data-visualization#ai-agents#benchmark#nlp#multimodal-llm#spreadsheets#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618915