Tagged articles

PUMA

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 16, 2026 · Artificial Intelligence

Brake Overthinking in Long‑Reasoning Models by Detecting Semantic Redundancy

Long‑thinking LLMs often waste 41‑52% of tokens after the final answer; the PUMA framework detects when reasoning stops producing new semantic information, enabling early exit that cuts average token usage by 26.2% while keeping accuracy stable and even improving speed across multiple benchmarks.

PUMAearly exitllm-inference
0 likes · 9 min read
Brake Overthinking in Long‑Reasoning Models by Detecting Semantic Redundancy