Let Generative Models Draw Spatial Answers, Not Text Coordinates – Zhejiang’s Agentic Evaluation Framework
The article critiques coordinate‑based spatial benchmarks, introduces the ProVisE framework that lets image‑generation models answer by drawing, describes the automated Agentic Builder for protocol creation, presents the 14‑task SpatialGen‑Bench, and reports that generative models show direct spatial intuition while complementing text‑output VLMs, with detailed failure analysis.
