Machine Learning Algorithms & Natural Language Processing
Oct 2, 2026 · Artificial Intelligence
PPTBench Benchmark Exposes Visual Coding Limits: GPT-6 Astra Leads but 20% Fail Semantic Checks
Einsia AI's PPTBench evaluates coding agents on reconstructing 500 scientific flowcharts into editable PPTX slides using a three-stage Agentic Judge; GPT-6 Astra High scores 77.34 with 80.8% passing semantic and rendering checks, yet semantic understanding remains the primary bottleneck across all models.
AI AgentsAgentic JudgeBenchmark
0 likes · 13 min read
