SAGE: Topological Guidance Mitigates Long-Horizon Reasoning Biases, 8x Lean Pass Rate

Researchers from Virginia Tech, UW-Madison, and Dartmouth introduce SAGE, a post-training method that uses symbolic closure analysis to diagnose exploration and accumulation biases in long-horizon reasoning, applying algebraic sparsification and hyperbolic structure guidance to improve sampling and provide early feedback, achieving near 8x Lean verification pass rate on Andrews-Curtis tasks and outperforming baselines across 12 benchmarks.

Machine Heart
Machine Heart
Machine Heart
SAGE: Topological Guidance Mitigates Long-Horizon Reasoning Biases, 8x Lean Pass Rate

Problem: Long-Horizon Reasoning Biases

Large language models excel at short reasoning tasks but often fail on long chains where each step appears locally valid yet the final answer is wrong. Post-training with only terminal rewards cannot identify which early step caused the deviation. The paper uses the Andrews–Curtis (AC) group transformation task as a stress test: 1,190 constructed instances where each step's legality is checkable, but legal moves do not guarantee reaching the trivial representation. Success is only confirmed at the end of the sequence.

Symbolic Closure Analysis (SCA): Diagnosing Two Biases

Feasible Prefix Set

SCA defines the locally feasible prefix set F as all sequences where every step obeys the task's local rules. F is prefix-closed: once a step is illegal, no continuation can make the prefix legal. However, F contains many paths that never reach the goal; the true successful trajectories are a strict subset unknown during training.

Exploration Bias

Let Bₜ be the total branching factor at depth t and Bₜᶠ the locally feasible branches. The ratio of feasible prefixes at depth T is roughly the product of Bₜᶠ/Bₜ across layers. When many layers have few feasible branches, this ratio shrinks exponentially with depth, making it extremely hard for sparse terminal rewards to randomly sample useful trajectories. The paper decomposes policy gradient variance into three parts: variance within feasible trajectories, variance within infeasible trajectories, and variance due to different mean signals between the two groups. Concentrating sampling on F eliminates the latter two terms, improving learning signal quality.

Geometric intuition of exploration bias: feasible prefix ratio constrained by product of feasible branch ratios across layers.
Geometric intuition of exploration bias: feasible prefix ratio constrained by product of feasible branch ratios across layers.

Accumulation Bias

With only terminal rewards and KL regularization, early decisions receive little corrective feedback. The paper derives an upper bound on the total variation distance between the optimal policy and the reference policy: it depends on the reference policy's success probability p , maximum terminal reward Rmax , and KL penalty strength λ . When successful trajectories are rare and rewards are weak relative to the KL constraint, the bound is small, meaning training cannot substantially move away from the reference model's early choices. Small initial deviations then persist along the long chain.

Upper bound on policy deviation showing dependence on success probability, reward magnitude, and KL penalty.
Upper bound on policy deviation showing dependence on success probability, reward magnitude, and KL penalty.

SAGE: Structural Admissibility-Guided Exploration

SAGE translates the two SCA design goals into complementary structural potential functions used during training only. At inference, the learned policy runs without extra computation.

Algebraic Sparsification (Targets Exploration Bias)

The system represents the unsolved structure as a residual . For each candidate operation, it measures how much of the residual is explained by the algebraic subspace associated with that operation. Candidates explaining more residual receive higher soft compatibility scores, increasing their sampling probability without discarding other locally legal moves.

Hyperbolic Structure Guidance (Targets Accumulation Bias)

The current state and the goal structure are embedded into a hyperbolic space suited for hierarchical relationships. The hyperbolic distance between them provides a step-wise feedback signal: prefixes closer to the goal get higher scores, giving the optimizer a gradient before the terminal reward appears. This distance is a training-time structural signal, not a proof of final correctness.

SAGE architecture: old policy proposes candidates, algebraic sparsification and hyperbolic guidance reweight sampling, terminal reward and average structural score compute group relative advantage for policy update.
SAGE architecture: old policy proposes candidates, algebraic sparsification and hyperbolic guidance reweight sampling, terminal reward and average structural score compute group relative advantage for policy update.

Training Procedure

The old policy generates a set of candidate reasoning steps.

SAGE applies the two potential functions to softly re-rank candidates for training-time sampling.

After the trajectory ends, the terminal reward and the average structural score are combined into a group relative advantage to update the policy.

The paper provides a conditional concentration result: if the potential functions separate all feasible from infeasible trajectories by a positive margin Δ , the reweighted probability ratio between the two groups is multiplied by at most exp(−λΔ) . Whether this condition holds depends on the task-specific structural priors.

Experimental Results

Evaluation covers 12 benchmarks (7 closed-form math, 4 free-text reasoning, 1 AC long-horizon symbolic task) across 7 model families. Baselines include SFT, GRPO, EMPO, and GRPO-PRM (process reward model). All methods use comparable rollout budgets, generation lengths, and decoding constraints.

Closed-Form Mathematical Reasoning

On Qwen3.5 models, SAGE average accuracies: 2B → 42.11%, 9B → 47.95%, 35B → 64.86%. Improvements over the strongest same-size post-training baseline: +1.83 pp, +3.02 pp, +2.49 pp respectively. Not every single benchmark is won; e.g., on AIME the 35B model scores 62.04% vs. EMPO's 62.31%.

Closed-form math results across model sizes and baselines.
Closed-form math results across model sizes and baselines.

Free-Text Reasoning

At 9B: MMLU-Pro rises from 37.91% (EMPO) to 39.99% (SAGE); BBH-H from 44.04% to 45.31%; ARC-C from 39.73% to 42.04%; GPQA slightly drops from 21.11% to 20.86%. At 35B, SAGE leads all four metrics: BBH-H 69.07%, ARC-C 66.41%.

Free-text reasoning results at 9B and 35B scales.
Free-text reasoning results at 9B and 35B scales.

AC Long-Horizon Symbolic Task

Three metrics: AC transformation validity (local step legality), AC path solving (full path reaches goal), Lean verification pass rate (formal proof check). Across six backbone models, SAGE improves transformation validity by 19.2–26.0 pp and Lean pass rate by 13.2–20.8 pp. On Qwen3, SAGE's Lean pass rate approaches 8× the base model.

Direct comparison on Qwen3.5-35B:

GRPO: 54.28% validity / 23.05% path solving / 14.64% Lean pass

GRPO-PRM: Lean pass 17.36%

SAGE: 59.83% validity / 31.76% path solving / 23.69% Lean pass

AC task results comparing base models, GRPO, GRPO-PRM, and SAGE across multiple metrics.
AC task results comparing base models, GRPO, GRPO-PRM, and SAGE across multiple metrics.

Conclusion and Future Work

SAGE demonstrates that structural guidance during training—algebraic sparsification for exploration bias and hyperbolic structure guidance for accumulation bias—substantially improves long-horizon reasoning without requiring step-level gold annotations or inference-time overhead. The method shifts learning from pure terminal rewards to leveraging the reasoning process's intrinsic structure.

Future work must address constructing reliable structural signals for tasks where rules are less explicit than in AC. Nevertheless, this work shows a promising direction: teaching models to choose the next step more effectively by understanding the structural landscape of the problem.

References

Paper: https://arxiv.org/abs/2609.30192

Code: https://github.com/Susan571/SAGE-NeurIPS2026

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelsPost-TrainingNeurIPS 2026Andrews-CurtisLean VerificationLong-Horizon ReasoningStructural GuidanceSymbolic Closure Analysis
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.