PPTBench Benchmark Exposes Visual Coding Limits: GPT-6 Astra Leads but 20% Fail Semantic Checks

Einsia AI's PPTBench evaluates coding agents on reconstructing 500 scientific flowcharts into editable PPTX slides using a three-stage Agentic Judge; GPT-6 Astra High scores 77.34 with 80.8% passing semantic and rendering checks, yet semantic understanding remains the primary bottleneck across all models.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
PPTBench Benchmark Exposes Visual Coding Limits: GPT-6 Astra Leads but 20% Fail Semantic Checks

PPTBench: Benchmark for Visual Coding Agents

Einsia AI's Navers Lab released PPTBench , a benchmark testing whether coding agents can understand visual targets and translate them into structured, editable code. The benchmark comprises 500 tasks derived from real scientific paper diagrams — flowcharts and architecture figures — spanning 13 domains including systems, AI, computer vision, quantum physics, and biomedicine. Each task requires the agent to reconstruct a given diagram as a native PPTX file using PowerPoint objects (text boxes, shapes, connectors, groups) rather than embedding the image.

Agentic Judge Evaluation Framework

The evaluation employs an Agentic Judge with three sequential stages:

Artifact Gate : Checks whether the generated PPTX is parsable, renders as a single slide, contains sufficient native editable objects, and is not merely an image paste. Invalid artifacts score zero.

Stage 1 – Semantic Correctness : Determines if the reconstructed diagram preserves the original process logic. A single reversed arrow or missing node changes the meaning; visual similarity alone is insufficient.

Stage 2 – Readability : Filters out severe rendering or layout errors that make the slide unusable despite semantic correctness.

Stage 3 – Visual Fidelity : Compares layout and local visual details, tolerating different object representations (e.g., a checkmark drawn as two lines vs. a polygon) as long as the rendered result matches closely.

Experimental Results

The team tested 10 models across 36 configurations , each running all 500 tasks, with three evaluation rounds per result — totaling 54,000 judgment records . GPT-6 Astra High achieved the highest score of 77.34 , producing valid artifacts for every task and passing both semantic and rendering gates on 80.8% of tasks. Kimi K3 High followed at 67.80, while GPT-5.6 Sol Max and Qwen 3.8 Max XHigh scored 49.28 and 47.69 respectively.

Error Analysis

Error analysis across all 18,000 results reveals the dominant failure mode: 394 artifacts (2.19%) failed the Artifact Gate, but 11,526 (64.03%) were rejected for semantic errors — misunderstanding node relationships, arrow directions, or hierarchical structure. Only 5,715 results (31.75%) reached the visual fidelity stage, where text and layout issues accounted for 53.0% of detail penalties (unexpected line breaks being the top issue), local graphics 34.3%, and layout 12.8%.

Reasoning Effort Impact

Increasing reasoning effort (Low → Medium → High → XHigh → Max) on GPT-6 Astra raised the average score from 59.42 to a peak of 77.34 at High, with gate passage rising from 63.0% to 80.8%. However, further increasing effort to XHigh and Max reduced overall scores (72.60 and 68.76) because gate passage dropped to 71.0% despite slightly higher conditional detail scores (96.85). This indicates deeper reasoning mainly improves structural correctness, not fine-grained visual precision, and that maximum effort is not a reliable default.

Trajectory Analysis: Visual Self-Review

Trajectory analysis of 27 configurations showed a strong Pearson correlation ( 0.881 ) between the number of visual self-review rounds and final score, versus only 0.255 for code modification rounds. Agents that repeatedly rendered and inspected their own output — checking node spacing, arrow direction, text overflow — performed significantly better. The authors argue this "look-back" capability is the core of visual coding: autonomously deciding when to observe, interpreting what is seen, and making targeted corrections.

Resources

PPTBench's goal is to measure a foundational capability: when an agent sees a visual target, can it reliably understand the structure and encode that understanding into executable, editable code? The benchmark, evaluation framework, and data are publicly available at:

Project page: https://lab.einsia.ai/pptbench

GitHub: https://github.com/Einsia/PPTBench

arXiv: https://arxiv.org/abs/2609.29718

AgentGit (partial evaluation sessions): https://agent-git.com/@einsia/ppt-bench

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsBenchmarkMultimodal EvaluationVisual CodingGPT-6 AstraAgentic JudgePPTBenchFlowchart Reconstruction
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.