Paper-Plot-Skills: Let AI Draw Your Paper Figures — From Matplotlib Tuning Hell to One-Line Plotting
> GitHub: https://github.com/Trae1ounG/paper-plot-skills > Author: Trae1ounG (CUHK-Shenzhen) > Positioning: An AI Skill toolbox — paper figure reproduction and generation
---
1. A Breakdown Moment Every Researcher Knows
Your experiment works great. Now you need a bar chart comparing baselines and your method for the paper.
You open matplotlib, write 50 lines of code, and get:
- Default blue, zero differentiation
- System default sans-serif font that clashes with the paper
- Wrong bar spacing — too wide looks empty, too narrow looks cramped
- No error bars, no significance markers, no legend
- PNG saved with white margins, insufficient dpi, blurry when zoomed
- Bar charts: comparison experiments (A vs B), ablations (5 variants)
- Line charts: training curves, performance vs. hyperparameters, scaling laws
- Scatter plots: visualizations (t-SNE/UMAP), distribution comparisons
- Radar charts: multi-dimensional method comparison (accuracy, speed, memory, generalization)
- serif (Times New Roman style): traditional, formal, academic — favored by classic CV/ML venues (early ICCV, NeurIPS)
- sans-serif (Arial/Helvetica style): modern, clear, screen-friendly — favored by newer venues and systems papers
- LaTeX Computer Modern: matching figure fonts to LaTeX body text is the most professional approach
- Print-friendly: many papers are printed in grayscale; palettes must remain distinguishable
- Colorblind-friendly: red-green colorblindness affects ~8% of men; don't rely on red vs. green alone
- Differentiation: 3–5 series per figure, each instantly recognizable
- Subtle: figures support the data; they shouldn't distract
- Most journals require 300 dpi (print quality); some require 600 dpi for line art; web display needs only 72–150 dpi
- paper-plot-skills defaults to 300 dpi, satisfying most journal requirements out of the box
- A user uploaded a paper screenshot (issue #1)
- The AI analyzed the table layout, two-row results, and strong/weak highlight shading
- Generated script:
plot-from-image/scripts/classwise_iou_table.py - The reproduced figure is nearly identical to the original
- See a figure you like? Reproduce its style without manual tuning
- Reviewer says your figure looks too similar to another paper's? Show the reproduction script as evidence of independent implementation
- Advisor says "follow top-conference style"? Just feed a screenshot to the AI
- plot-from-data: parameter injection + template rendering — output is standard matplotlib code, not a black box
- plot-from-image: visual analysis + parameter inference — generating reproducible scripts
- Trae1ounG (2026). Paper-Plot-Skills. https://github.com/Trae1ounG/paper-plot-skills
- Preset styles sourced from: MemEvolve, SPICE, Self-Distillation, DAPO, SiameseNorm, MemGen, Meta-Harness, DoRA
After 2 hours of tuning and 10 StackOverflow threads, it looks okay — but still a notch below the figures in the neighboring group's paper.
The problem? Not your data, not your experiments — you simply don't know how to tune matplotlib. Paper figures aren't just "plotting data"; they are visual storytelling. Fonts, colors, spacing, annotations, and layout all shape the reviewer's first impression.
Trae1ounG's paper-plot-skills solves exactly this. It doesn't teach you matplotlib — it packages "top-conference figure aesthetics" as an AI Skill you invoke with one sentence.
---
2. Core Design: Extracting "Style Parameters" from 9 Real Paper Figures
The key insight: top papers' figure styles are highly consistent — and can be systematically extracted.
The author selected 9 figures with valuable plotting styles from papers he'd read and did two things:
1. Style decomposition: each figure is a "parameter system"
Example — a bar chart from the MemEvolve paper:
| Style dimension | Parameters | |---|---| | Font | serif (Times New Roman style), a paper classic | | Layout | paired bars — baseline and method side by side | | Annotations | gain arrows + percentages, showing improvements directly | | Y-axis | independent scales per subplot, not forced uniform | | Palette | low-saturation contrast, avoiding gaudy colors |
Example — a training curve from the DAPO paper:
| Style dimension | Parameters | |---|---| | Font | sans-serif, modern and clean | | Reference lines | horizontal dashed lines for key thresholds | | Break lines | vertical dashed lines marking key training events | | Spines | four-sided frame, outward ticks, professional look | | Legend | placed separately, not covering curves |
These parameters are written as .md docs (bar_paired_delta.md, line_training_curve.md, etc.), each paired with a matplotlib script template.
2. Two usage modes
Mode 1: plot-from-data — "draw my data in this style"
You say: "Plot my data using bar_grouped_hatch style." The AI:
1. Reads the bar_grouped_hatch.md style parameters
2. Reads your data (CSV/JSON/plain text)
3. Fills the template script, generating a dpi=300 publication-ready figure
Mode 2: plot-from-image — "reproduce this figure for me"
Upload a paper screenshot, and the AI: 1. Analyzes aspect ratio, fonts, colors, and layout structure 2. Infers matplotlib parameters automatically 3. Generates a reproducible Python script
---
3. Nine Preset Styles Covering the Most Common Paper Figure Types
| Style | Type | Source paper | Key features |
|---|---|---|---|
| bar_paired_delta | Bar | MemEvolve | Paired bars + gain arrows, serif font |
| bar_grouped_hatch | Bar | SPICE | Grouped bars + hatch fill for main method, value labels |
| line_confidence_band | Line | Self-Distillation | EMA smoothing + confidence band, LaTeX font |
| line_training_curve | Line | DAPO | Vertical break lines + horizontal references, sans-serif |
| line_loss_with_inset | Line | SiameseNorm | L-shaped spines + axis-end arrows + right-side zoom inset |
| scatter_tsne_cluster | Scatter | MemGen | t-SNE clusters + rounded colored annotation boxes, dotted grid |
| scatter_broken_axis | Scatter | Meta-Harness | Broken X-axis dual panel, multiple marker types |
| radar_dual_series | Radar | DoRA | Octagonal dashed concentric grid, two-method comparison |
These choices cover roughly 80% of ML/AI paper figure scenarios:
---
4. Technical Details: Why These "Style Parameters" Matter
4.1 Fonts: serif vs. sans-serif is a field convention, not just aesthetics
Each style specifies its font deliberately, matching the source paper.
4.2 Spine design: L-shaped vs. full frame is an information-density choice
| Spine style | Use case | Impression | |---|---|---| | Full frame | Line charts needing precise reading | Rigorous, engineering-oriented | | L-shaped (left+bottom) | Trend-emphasizing curves | Clean, modern | | Frameless (axes only) | Minimalist visualization | Premium, designed | | Open frame | Grouped comparisons | Open, comparative |
4.3 Color: not "pretty" but "distinguishable + printable + colorblind-friendly"
4.4 dpi=300 is not arbitrary
---
5. plot-from-image: From Screenshot to Script
1. Input: a paper screenshot (PDF capture, even a phone photo of a monitor)
2. Analysis: the AI identifies elements — number of bars/curves, colors, fonts, layout
3. Inference: visual features are translated into matplotlib parameters
4. Output: a .py script that reproduces the original figure
A real case: classwise_iou
Why this matters:
---
6. Feynman View: What Is This Project Really?
Q1: How is this different from matplotlib templates?
Templates (e.g., seaborn's paper context) only solve basics — font sizes, line widths, color cycles. But paper-figure aesthetics go further: paired bars with gain arrows, L-shaped spines with insets, broken axes with mixed markers. paper-plot-skills packages "advanced matplotlib techniques" as one-sentence Skills — it presets aesthetic judgment, not just parameters.
Q2: What role does AI play?
In both modes, AI is an assistant, not a replacement. You still choose the style, understand the data, and judge the output.
Q3: Where's the ceiling?
1. Style count: 9 styles cover ~80% of cases; specialized charts (manifold visualizations, network graphs, heatmaps) are not yet covered 2. Field limitations: styles come from ML/AI papers; biomedical, physics, and chemistry conventions differ (error bars, log-log axes) 3. Interactive figures: static matplotlib only; modern publishing increasingly accepts interactive HTML figures 4. Ecosystem lock-in: matplotlib only — no ggplot2 (R), Plotly, or TikZ
Q4: The most valuable lesson is the "aesthetic decomposition" methodology
The 9 scripts are less valuable than the method of dissecting figures — systematically extracting font, spine style, palette, annotation, and layout from any paper figure. This methodology transfers to any visualization tool, and even without using the toolbox, it teaches you to review your own figures more professionally.
---
7. Use Cases & Recommendations
| Scenario | Recommended mode | Why |
|---|---|---|
| Have experiment data, need comparison chart | plot-from-data + bar_paired_delta | Most common, immediate results |
| Love a paper figure, want its style | plot-from-image | Upload screenshot, automatic analysis |
| Need t-SNE visualization | plot-from-data + scatter_tsne_cluster | Professional cluster annotation |
| Ablation with 5 methods | plot-from-data + bar_grouped_hatch | Grouped bars + hatch, clearly separated |
| Training curves showing convergence | plot-from-data + line_confidence_band | EMA smoothing + confidence band |
| Multi-dimensional comparison | plot-from-data + radar_dual_series | Radar, 8 dimensions at a glance |
---
8. One-Sentence Summary
> Paper-Plot-Skills' core insight: paper figure quality is determined by aesthetic parameters, not just data. Nine styles distilled from real top-conference papers turn matplotlib's "tuning hell" into "one-line plotting." plot-from-image compresses the see-a-good-figure → reproduce-style cycle from hours to minutes. Researchers should spend time on experiments, not on color palettes.
---
References