VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic Labels

Overview

VisAnatomy is a large-scale corpus of real-world SVG charts designed to support chart understanding beyond coarse chart-type labels. Rather than treating a visualization as a single category such as “bar chart” or “line chart,” the dataset decomposes each chart into semantically meaningful components, including marks, axes, legends, labels, grouping structures, and visual encodings. This enables a more faithful representation of the chart’s underlying structure and of the mappings between data values and visual channels such as position, color, and size. The corpus is intended to reflect the diversity and complexity of charts encountered in the wild rather than only synthetic examples.

The dataset contains 942 charts generated by more than 50 tools across 40 chart types, with annotations for over 383,000 graphical elements. Compared with existing chart corpora, which often provide only high-level labels, VisAnatomy offers finer-grained semantic information that is better suited for downstream tasks such as semantic role inference, chart decomposition, chart type classification, and accessibility-oriented navigation. By capturing both structural and stylistic variation in real-world visualizations, VisAnatomy provides a more realistic foundation for developing and evaluating chart understanding systems.

Project Website

https://visanatomy.github.io/

Publications

VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic Labels, VIS 2025.
Chen Chen, Hannah K. Bako, Peihong Yu, John Hooker, Jeffrey Joyal, Simon C. Wang, Samuel Kim, Jessica Wu, Aoxue Ding, Lara Sandeep, Alex Chen, Chayanika Sinha, Zhicheng Liu