Tagged articles

spatial reasoning

11 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 9, 2026 · Artificial Intelligence

S-Space: Inside Multimodal Models' Internal 3D Map and Its Reasoning Failures

Researchers discover S-Space, a stable low-dimensional spatial workspace in multimodal models that encodes object positions, but models confuse coordinate systems and fail to reliably rotate spatial representations, revealing a gap between having spatial representations and using them for reasoning.

GPT-6 AstraQwenS-Space
0 likes · 15 min read
S-Space: Inside Multimodal Models' Internal 3D Map and Its Reasoning Failures
HyperAI Super Neural
HyperAI Super Neural
Sep 3, 2026 · Artificial Intelligence

UrbanGround: Testing MLLM Agents in Real-Scale Hong Kong City Navigation

Researchers from Shanghai Jiao Tong University, NUS, Meituan, and others built UrbanGround, a real-scale interactive Hong Kong environment using 3D geographic data, to evaluate MLLM agents on local perception, navigation persistence, and dynamic adaptation, revealing models excel at local tasks but fail at long-range navigation and dynamic recovery.

Hong KongMLLMUrbanGround
0 likes · 15 min read
UrbanGround: Testing MLLM Agents in Real-Scale Hong Kong City Navigation
Sohu Tech Products
Sohu Tech Products
Aug 26, 2026 · Artificial Intelligence

DeepSeek-V4-Flash-Vision-Exp Multimodal Evaluation: Strong Recognition but Over-Inference on Real-World Context

The author evaluates DeepSeek's new multimodal model across four visual reasoning challenges, finding excellent recognition and structured reasoning capabilities but a consistent tendency to hallucinate real-world details not present in images, a common limitation in vision-language models.

DeepSeekHallucinationOCR
0 likes · 9 min read
DeepSeek-V4-Flash-Vision-Exp Multimodal Evaluation: Strong Recognition but Over-Inference on Real-World Context
PaperAgent
PaperAgent
Aug 22, 2026 · Artificial Intelligence

DeepSeek’s Hidden Multimodal Model: Technical Deep‑Dive and Unexpected Bugs

The article reviews DeepSeek‑V4‑Flash‑Vision‑Exp, exposing a misidentification bug, detailing its visual‑primitive approach, impressive spatial‑reasoning benchmarks, and a highly compressed KV‑cache architecture that balances performance with efficiency.

DeepSeekKV cache compressionMoE
0 likes · 4 min read
DeepSeek’s Hidden Multimodal Model: Technical Deep‑Dive and Unexpected Bugs
Machine Heart
Machine Heart
Jun 24, 2026 · Artificial Intelligence

From Pixels to Words: A Native Vision-Language Model Unifies Images and Video

The paper introduces NEO‑ov, a native vision‑language model that discards external visual encoders, feeding raw pixels directly into a unified transformer, and demonstrates competitive performance on image, multi‑image, and video tasks—including fine‑grained perception and spatial reasoning—while outlining its three‑stage training pipeline and current limitations.

Qwenbenchmarkmultimodal
0 likes · 13 min read
From Pixels to Words: A Native Vision-Language Model Unifies Images and Video
Machine Heart
Machine Heart
Jun 24, 2026 · Artificial Intelligence

How APEIRIA Breaks the Black‑Box Barrier of 3D MLLMs (ICML 2026)

The paper introduces APEIRIA, a three‑stage curriculum that distills neuro‑symbolic program traces into 3D multi‑modal LLMs, enabling transparent spatial reasoning while preserving open‑vocabulary understanding, and demonstrates strong benchmark gains, modular upgrades, and zero‑shot generalization.

3D MLLMModular AINeuro-Symbolic Reasoning
0 likes · 11 min read
How APEIRIA Breaks the Black‑Box Barrier of 3D MLLMs (ICML 2026)
AI Architecture Hub
AI Architecture Hub
Jun 23, 2026 · Artificial Intelligence

Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More

This article reviews seven recent AI research papers, detailing how SpatialClaw enables code‑based spatial reasoning for vision‑language models, SkillWeaver introduces compositional skill routing, PreAct compiles agent actions into reusable state‑machines, and other works advance world‑model inference, self‑designing RL environments, collective skill‑tree search, and process‑aligned reinforcement learning for diffusion LLMs.

Large Language Modelsagent reasoningdiffusion models
0 likes · 15 min read
Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More
HyperAI Super Neural
HyperAI Super Neural
Mar 27, 2026 · Artificial Intelligence

Open-Source Reasoning Datasets: NVIDIA, OpenAI, Labs – Math, Spatial, Wiki QA

HyperAI has compiled a collection of high‑quality open‑source reasoning datasets—including Open‑RL, CHIMERA, Nemotron‑Math‑v2, OmniSpatial, FrontierScience, HotpotQA, VCR, and CIRR—covering math, multi‑step STEM problems, spatial reasoning, scientific tasks, wiki QA, and visual commonsense, all available for download or online use.

NVIDIAOpenAImultimodal
0 likes · 9 min read
Open-Source Reasoning Datasets: NVIDIA, OpenAI, Labs – Math, Spatial, Wiki QA
AI Algorithm Path
AI Algorithm Path
Feb 16, 2026 · Artificial Intelligence

Why Visual Tokenizers Bridge the Gap Between Pixels and Meaning

Vision‑language models turn continuous images into discrete tokens through patch extraction, encoding, and projection, enabling Transformers to reason jointly over vision and text, but this compression introduces limits in spatial reasoning, counting, and resolution sensitivity that users must understand.

Multimodal FusionSelf-AttentionVision-Language Models
0 likes · 22 min read
Why Visual Tokenizers Bridge the Gap Between Pixels and Meaning
AntTech
AntTech
Jul 17, 2025 · Artificial Intelligence

How M2-Reasoning-7B Achieves State‑of‑the‑Art Spatial Reasoning in Multimodal AI

M2-Reasoning-7B, an open‑source 7B multimodal model from Ant Group, combines a high‑quality data pipeline with dynamic multi‑task training and a novel reward function to deliver state‑of‑the‑art performance on both general and spatial reasoning benchmarks, surpassing many larger competitors.

Large Language ModelM2-Reasoningbenchmark
0 likes · 9 min read
How M2-Reasoning-7B Achieves State‑of‑the‑Art Spatial Reasoning in Multimodal AI