Social Intelligence: The Missing Third Pillar of AGI Beyond Large Models and Robots

The article argues that while symbolic AI (e.g., GPT‑5.5, DeepSeek) and embodied robotics represent two mature AI domains, true experience‑based intelligence arises from social cognition, and Zhijing's SoMBench, Zing models, and Actio framework demonstrate a concrete technical path toward this third AGI pillar.

Machine Heart
Machine Heart
Machine Heart
Social Intelligence: The Missing Third Pillar of AGI Beyond Large Models and Robots

AI Landscape and the Need for Social Intelligence

Current AI research has formed two mature tracks: symbolic‑space intelligence (e.g., GPT‑5.5, DeepSeek) excels at language understanding, logical reasoning, and code generation; physical‑space intelligence (embodied robots) excels at perception, action, and environment control. Richard Sutton announced at WAIC 2026 that AI is entering an “experience era,” where agents must generate their own experience through continual interaction with the real world.

Experience is rooted in social relations—recursive beliefs such as “I know that you know that I know,” eye‑cue reading, and inference of unspoken intentions. This social cognition is presented as the missing third pole of AGI.

Zhijing Technical Stack

The Zhijing team (Institute of Computing Technology, Chinese Academy of Sciences) released a technical stack comprising:

SoMBench evaluation suite

Zing multi‑parameter social‑mind models

Actio deployment framework

SoMBench: Social‑Mind “Physical Exam”

SoMBench builds a capability map that decomposes “understanding people” into locatable cognitive gaps. It covers 17 secondary dimensions, 71 fine‑grained tasks, and 3,481 instances reviewed by 15 experts (87 % retained). The benchmark provides scores and error‑pattern diagnostics through cross‑question validation, option perturbation, and shortcut detection.

In tests on 20 top models, the best model achieved a correct‑rate of 72.08 % and none of the 17 secondary dimensions reached the 90 % ceiling, indicating a substantial gap in current models.

Zing Models: Training Targets Evolve with Capability

Zing replaces a fixed dataset and loss with a self‑evolving system called FLARE. When the diagnostic set reveals a shortfall, FLARE generates new training samples that target the same capability but vary semantics, scenes, and reasoning paths. The diagnostic set is strictly isolated from the frozen evaluation set.

Training proceeds in two stages:

General social‑mind foundation built with a Mixed‑Reward GRPO reinforcement‑learning framework that dynamically combines process rewards and verifiable rewards.

Specialized reinforcement for emotion understanding, intent inference, norm internalization, and perspective taking.

OPD online distillation adds token‑level teacher constraints during reinforcement learning to keep the model on validated reasoning trajectories.

Performance results:

Zing‑27B (multimodal) achieves a composite mean of 79.80 on five international social‑mind benchmarks, surpassing GPT‑5.5’s 78.44. On the HiToM recursive‑belief task, Zing‑27B scores 82.00 versus GPT‑5.5’s 77.92, becoming the first open‑source hundred‑billion‑parameter model to exceed GPT‑5.5 on this benchmark.

Zing‑32B (text) matches DeepSeek‑V4‑Pro with 76.14 versus 75.90.

Zing‑8B/14B (lightweight) outperform much larger baselines on HiToM: 3‑step 67.50 % (baseline 45.83 %) and 4‑step 65.83 % (baseline 43.75 %). The gains demonstrate that structured training, not sheer scale, drives social‑intelligence improvements.

Actio: Deployment Layer with Selective, Inspectable Execution

Actio’s Harness orchestrator activates four support modules based on task‑mind variables:

PRISM : 76 layered psychological and social skills.

Starling Memory : explicit “who‑believes‑what‑when” records.

SAGE : cross‑task reusable reasoning experience.

Gated RAG : on‑demand retrieval of cultural knowledge.

This design enables selective, inspectable execution of social‑mind capabilities.

Future Extensions

Three real‑world impact directions are envisioned:

Empowering robots to move from mere manipulation to “mindful” interaction, e.g., inferring whether a user’s frown signals hot tea or bad mood.

Transforming complex decision‑making from intuition to rehearsed scenario analysis, allowing stakeholders to preview cascading social effects before acting.

Synchronizing multiple agents—humans, robots, AI agents—through a shared mind‑state protocol that resolves the “you think I understand, I think you understand” risk.

Technical Resources

GitHub repository: https://github.com/Zhijing-AI/Zing

arXiv technical report: https://arxiv.org/abs/2607.23740

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelsbenchmarkAGIsocial intelligenceActioZing
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.