Tagged articles

LPRoPE

3 articles · Page 1 of 1
Data Party THU
Data Party THU
Jun 10, 2026 · Artificial Intelligence

How Visual Para-Thinker Tackles Visual Hallucination with a Clever Parallel Reasoning Design

The article introduces Visual Para-Thinker, a parallel reasoning framework for large vision‑language models that mitigates attention drift and visual hallucination by employing path‑aware attention, learnable parallel rotary position embeddings, and hybrid block‑and‑scan visual token partitions, and validates the approach with extensive multimodal benchmarks.

LPRoPEMultimodal BenchmarksParallel Attention
0 likes · 10 min read
How Visual Para-Thinker Tackles Visual Hallucination with a Clever Parallel Reasoning Design
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
May 24, 2026 · Artificial Intelligence

The First Visual‑Language Parallel Thinking Framework: Unpacking Its Core Mechanisms

The paper introduces Visual Para-Thinker, a parallel‑thinking framework for large‑scale visual‑language models that uses visual‑centered block and scan path partitions, Path‑aware Attention and Learnable Parallel Rotary Position Embedding, and demonstrates consistent gains across counting, visual search, hallucination and grounding benchmarks.

LPRoPEPa-Attentionbenchmark evaluation
0 likes · 11 min read
The First Visual‑Language Parallel Thinking Framework: Unpacking Its Core Mechanisms
Machine Heart
Machine Heart
May 24, 2026 · Artificial Intelligence

Inside the First Vision-Centric Parallel Thinking Framework for Vision-Language Models

The article introduces Visual Para-Thinker, the first parallel reasoning framework tailored for large‑scale vision‑language models, explains its block and scan visual path divisions, details the Path‑aware Attention and Learnable Parallel Rotary Position Embedding mechanisms, and presents experimental results showing significant gains on visual perception benchmarks.

LPRoPEPath-aware AttentionVision-Language Models
0 likes · 9 min read
Inside the First Vision-Centric Parallel Thinking Framework for Vision-Language Models