PPTBench: Can Coding Agents Accurately Reconstruct Flowcharts into Editable Slides?

Einsia AI's PPTBench benchmark evaluates coding agents on reconstructing 500 scientific flowcharts into editable PowerPoint slides, revealing that even top models like GPT-6 Astra fail 20% of semantic and rendering tests, with visual self-inspection correlating strongly (r=0.88) with success.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
PPTBench: Can Coding Agents Accurately Reconstruct Flowcharts into Editable Slides?

Einsia AI's Navers Lab has released PPTBench , a benchmark designed to test whether coding agents can understand visual targets — such as flowcharts and architecture diagrams from scientific papers — and reconstruct them as fully editable, native PowerPoint (PPTX) files. The benchmark comprises 500 tasks sourced from real scientific publications across 13 domains (systems & software, AI, computer vision, language & speech, quantum & fundamental physics, astronomy, robotics, biomedicine, etc.), covering diverse layout types including multi-panel figures, hierarchical structures, repeated grids, complex connectors, and text-dense diagrams.

Task Definition

Given a flowchart image, the agent must generate code that produces a single-slide PPTX using native PowerPoint objects — text boxes, shapes, connectors, groups — not simply embed the original image. The agent must simultaneously solve three sub-problems: perceive the visual structure (nodes, text, connections), map it accurately onto a 2D canvas (positions, sizes, spacing), and deliver a valid, editable PPTX file.

Agentic Judge: Three-Stage Evaluation

Because the same visual element can be expressed in many different object trees (e.g., a 2×2 grid can be four rectangles or one rectangle plus two lines), PPTBench introduces an Agentic Judge that evaluates the final rendered artifact rather than enforcing a fixed object representation.

Artifact Gate : Checks that the PPTX parses and renders, is a single slide, contains sufficient native editable objects, and does not cheat by pasting the input image. Failures score 0.

Stage 1 – Semantic Correctness : Determines whether the reconstructed flowchart expresses the same process as the reference. A reversed arrow or missing node changes the meaning, so visual similarity alone is insufficient.

Stage 2 – Readability : Filters out artifacts with severe rendering or layout issues that make the slide practically unreadable despite semantic correctness.

Stage 3 – Visual Fidelity : For artifacts passing the first two stages, compares overall layout and local visual differences, recording structured defect categories (text/layout, local graphics, global layout).

Experimental Results

The team tested 10 models across 36 configurations , each running all 500 tasks, with three evaluation rounds per result — totalling 54,000 judgment records .

GPT-6 Astra High leads with 77.34 points, producing valid artifacts on all 500 tasks and passing both semantic and readability gates on 80.8% of them.

Kimi K3 High scores 67.80; GPT-5.6 Sol Max 49.28; Qwen 3.8 Max XHigh 47.69.

Even the best configuration fails to pass both gates on nearly 20% of tasks .

Failure Analysis

Out of 18,000 total results (36 configs × 500 tasks):

394 (2.19%) failed the Artifact Gate (invalid PPTX, multi-slide, image pasting, etc.).

11,526 (64.03%) were rejected at Stage 1 due to semantic errors — the dominant failure mode.

Of the remaining 6,080, 365 failed Stage 2 (rendering/text gating), leaving 5,715 (31.75%) for detailed visual scoring.

In the detailed scoring, text and layout issues contributed 53.0% of deductions, local graphics 34.3%, and global layout 12.8%. The single largest defect was unexpected text wrapping .

These defects are invisible in the generation code itself; they only appear upon rendering. Hence, agents that can visually inspect their own output have a decisive advantage.

Reasoning Effort vs. Performance

For GPT-6 Astra, increasing reasoning effort (Low → Medium → High → XHigh → Max) shows a non-monotonic pattern:

Average score: 59.42 → 68.16 → 77.34 → 72.60 → 68.76 .

Gate pass rate (semantic + readability): 63.0% → 71.0% → 80.8% → 75.0% → 71.0% .

Conditional detail score (given gate pass): 94.32 → 95.00 → 95.72 → 96.50 → 96.85 .

Higher effort mainly improves the likelihood of getting the process structure right (gate pass rate), while detail precision improves only marginally. The peak at High effort suggests that maxing out reasoning budget is not a reliable default.

Visual Self-Inspection Drives Improvement

Analyzing agent call trajectories across 27 configurations, the paper finds a strong correlation between the number of visual review rounds (where the agent renders and inspects its own output) and final score: Pearson r = 0.881 . In contrast, the number of code modification rounds correlates weakly (r = 0.255). Visual inspection gives the agent direct feedback — e.g., whether nodes overlap, arrows point correctly, text overflows — enabling targeted fixes.

Implications

PPTBench demonstrates that frontier agents can perform complex visual reconstruction, but a clear gap remains between "can reconstruct" and "reliably reconstruct correctly." Semantic understanding is the primary bottleneck ; once semantics are largely correct, text handling, layout, and local graphic fidelity become the next hurdles. The benchmark's core question is not merely "which model is strongest" but: When an agent sees a visual target, can it truly comprehend the underlying structure and reliably encode that understanding into code? This is the central challenge for the next phase of Visual Coding.

Resources: Paper arXiv:2609.29718 (https://arxiv.org/abs/2609.29718), Project page https://lab.einsia.ai/pptbench, GitHub https://github.com/Einsia/PPTBench, AgentGit sessions https://agent-git.com/@einsia/ppt-bench.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI AgentsBenchmarkingMultimodal EvaluationVisual CodingGPT-6 AstraAgentic JudgeFlowchart UnderstandingPPTBench
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.