Social Intelligence: The Missing Third Pillar of AGI Beyond Large Models and Robots
The article argues that while symbolic AI (e.g., GPT‑5.5, DeepSeek) and embodied robotics represent two mature AI domains, true experience‑based intelligence arises from social cognition, and Zhijing's SoMBench, Zing models, and Actio framework demonstrate a concrete technical path toward this third AGI pillar.
AI Landscape and the Need for Social Intelligence
Current AI research has formed two mature tracks: symbolic‑space intelligence (e.g., GPT‑5.5, DeepSeek) excels at language understanding, logical reasoning, and code generation; physical‑space intelligence (embodied robots) excels at perception, action, and environment control. Richard Sutton announced at WAIC 2026 that AI is entering an “experience era,” where agents must generate their own experience through continual interaction with the real world.
Experience is rooted in social relations—recursive beliefs such as “I know that you know that I know,” eye‑cue reading, and inference of unspoken intentions. This social cognition is presented as the missing third pole of AGI.
Zhijing Technical Stack
The Zhijing team (Institute of Computing Technology, Chinese Academy of Sciences) released a technical stack comprising:
SoMBench evaluation suite
Zing multi‑parameter social‑mind models
Actio deployment framework
SoMBench: Social‑Mind “Physical Exam”
SoMBench builds a capability map that decomposes “understanding people” into locatable cognitive gaps. It covers 17 secondary dimensions, 71 fine‑grained tasks, and 3,481 instances reviewed by 15 experts (87 % retained). The benchmark provides scores and error‑pattern diagnostics through cross‑question validation, option perturbation, and shortcut detection.
In tests on 20 top models, the best model achieved a correct‑rate of 72.08 % and none of the 17 secondary dimensions reached the 90 % ceiling, indicating a substantial gap in current models.
Zing Models: Training Targets Evolve with Capability
Zing replaces a fixed dataset and loss with a self‑evolving system called FLARE. When the diagnostic set reveals a shortfall, FLARE generates new training samples that target the same capability but vary semantics, scenes, and reasoning paths. The diagnostic set is strictly isolated from the frozen evaluation set.
Training proceeds in two stages:
General social‑mind foundation built with a Mixed‑Reward GRPO reinforcement‑learning framework that dynamically combines process rewards and verifiable rewards.
Specialized reinforcement for emotion understanding, intent inference, norm internalization, and perspective taking.
OPD online distillation adds token‑level teacher constraints during reinforcement learning to keep the model on validated reasoning trajectories.
Performance results:
Zing‑27B (multimodal) achieves a composite mean of 79.80 on five international social‑mind benchmarks, surpassing GPT‑5.5’s 78.44. On the HiToM recursive‑belief task, Zing‑27B scores 82.00 versus GPT‑5.5’s 77.92, becoming the first open‑source hundred‑billion‑parameter model to exceed GPT‑5.5 on this benchmark.
Zing‑32B (text) matches DeepSeek‑V4‑Pro with 76.14 versus 75.90.
Zing‑8B/14B (lightweight) outperform much larger baselines on HiToM: 3‑step 67.50 % (baseline 45.83 %) and 4‑step 65.83 % (baseline 43.75 %). The gains demonstrate that structured training, not sheer scale, drives social‑intelligence improvements.
Actio: Deployment Layer with Selective, Inspectable Execution
Actio’s Harness orchestrator activates four support modules based on task‑mind variables:
PRISM : 76 layered psychological and social skills.
Starling Memory : explicit “who‑believes‑what‑when” records.
SAGE : cross‑task reusable reasoning experience.
Gated RAG : on‑demand retrieval of cultural knowledge.
This design enables selective, inspectable execution of social‑mind capabilities.
Future Extensions
Three real‑world impact directions are envisioned:
Empowering robots to move from mere manipulation to “mindful” interaction, e.g., inferring whether a user’s frown signals hot tea or bad mood.
Transforming complex decision‑making from intuition to rehearsed scenario analysis, allowing stakeholders to preview cascading social effects before acting.
Synchronizing multiple agents—humans, robots, AI agents—through a shared mind‑state protocol that resolves the “you think I understand, I think you understand” risk.
Technical Resources
GitHub repository: https://github.com/Zhijing-AI/Zing
arXiv technical report: https://arxiv.org/abs/2607.23740
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
