From One‑Shot Answers to Action Chains: S‑Agent Advances Spatial Intelligence
S‑Agent redefines spatial intelligence by replacing single‑shot answers with a multi‑step action chain that combines a vision‑language model for task planning, specialized depth and pose models for 3D alignment, and a spatial expert that converts geometry into usable evidence, achieving state‑of‑the‑art zero‑shot scores on MMSI‑Bench and ViewSpatial‑Bench and further improvements after distilling 29.2 k trajectories into an 8‑billion‑parameter model.
