SAGE: Topological Guidance Mitigates Long-Horizon Reasoning Biases, 8x Lean Pass Rate
Researchers from Virginia Tech, UW-Madison, and Dartmouth introduce SAGE, a post-training method that uses symbolic closure analysis to diagnose exploration and accumulation biases in long-horizon reasoning, applying algebraic sparsification and hyperbolic structure guidance to improve sampling and provide early feedback, achieving near 8x Lean verification pass rate on Andrews-Curtis tasks and outperforming baselines across 12 benchmarks.
Problem: Long-Horizon Reasoning Biases
Large language models excel at short reasoning tasks but often fail on long chains where each step appears locally valid yet the final answer is wrong. Post-training with only terminal rewards cannot identify which early step caused the deviation. The paper uses the Andrews–Curtis (AC) group transformation task as a stress test: 1,190 constructed instances where each step's legality is checkable, but legal moves do not guarantee reaching the trivial representation. Success is only confirmed at the end of the sequence.
Symbolic Closure Analysis (SCA): Diagnosing Two Biases
Feasible Prefix Set
SCA defines the locally feasible prefix set F as all sequences where every step obeys the task's local rules. F is prefix-closed: once a step is illegal, no continuation can make the prefix legal. However, F contains many paths that never reach the goal; the true successful trajectories are a strict subset unknown during training.
Exploration Bias
Let Bₜ be the total branching factor at depth t and Bₜᶠ the locally feasible branches. The ratio of feasible prefixes at depth T is roughly the product of Bₜᶠ/Bₜ across layers. When many layers have few feasible branches, this ratio shrinks exponentially with depth, making it extremely hard for sparse terminal rewards to randomly sample useful trajectories. The paper decomposes policy gradient variance into three parts: variance within feasible trajectories, variance within infeasible trajectories, and variance due to different mean signals between the two groups. Concentrating sampling on F eliminates the latter two terms, improving learning signal quality.
Accumulation Bias
With only terminal rewards and KL regularization, early decisions receive little corrective feedback. The paper derives an upper bound on the total variation distance between the optimal policy and the reference policy: it depends on the reference policy's success probability p , maximum terminal reward Rmax , and KL penalty strength λ . When successful trajectories are rare and rewards are weak relative to the KL constraint, the bound is small, meaning training cannot substantially move away from the reference model's early choices. Small initial deviations then persist along the long chain.
SAGE: Structural Admissibility-Guided Exploration
SAGE translates the two SCA design goals into complementary structural potential functions used during training only. At inference, the learned policy runs without extra computation.
Algebraic Sparsification (Targets Exploration Bias)
The system represents the unsolved structure as a residual . For each candidate operation, it measures how much of the residual is explained by the algebraic subspace associated with that operation. Candidates explaining more residual receive higher soft compatibility scores, increasing their sampling probability without discarding other locally legal moves.
Hyperbolic Structure Guidance (Targets Accumulation Bias)
The current state and the goal structure are embedded into a hyperbolic space suited for hierarchical relationships. The hyperbolic distance between them provides a step-wise feedback signal: prefixes closer to the goal get higher scores, giving the optimizer a gradient before the terminal reward appears. This distance is a training-time structural signal, not a proof of final correctness.
Training Procedure
The old policy generates a set of candidate reasoning steps.
SAGE applies the two potential functions to softly re-rank candidates for training-time sampling.
After the trajectory ends, the terminal reward and the average structural score are combined into a group relative advantage to update the policy.
The paper provides a conditional concentration result: if the potential functions separate all feasible from infeasible trajectories by a positive margin Δ , the reweighted probability ratio between the two groups is multiplied by at most exp(−λΔ) . Whether this condition holds depends on the task-specific structural priors.
Experimental Results
Evaluation covers 12 benchmarks (7 closed-form math, 4 free-text reasoning, 1 AC long-horizon symbolic task) across 7 model families. Baselines include SFT, GRPO, EMPO, and GRPO-PRM (process reward model). All methods use comparable rollout budgets, generation lengths, and decoding constraints.
Closed-Form Mathematical Reasoning
On Qwen3.5 models, SAGE average accuracies: 2B → 42.11%, 9B → 47.95%, 35B → 64.86%. Improvements over the strongest same-size post-training baseline: +1.83 pp, +3.02 pp, +2.49 pp respectively. Not every single benchmark is won; e.g., on AIME the 35B model scores 62.04% vs. EMPO's 62.31%.
Free-Text Reasoning
At 9B: MMLU-Pro rises from 37.91% (EMPO) to 39.99% (SAGE); BBH-H from 44.04% to 45.31%; ARC-C from 39.73% to 42.04%; GPQA slightly drops from 21.11% to 20.86%. At 35B, SAGE leads all four metrics: BBH-H 69.07%, ARC-C 66.41%.
AC Long-Horizon Symbolic Task
Three metrics: AC transformation validity (local step legality), AC path solving (full path reaches goal), Lean verification pass rate (formal proof check). Across six backbone models, SAGE improves transformation validity by 19.2–26.0 pp and Lean pass rate by 13.2–20.8 pp. On Qwen3, SAGE's Lean pass rate approaches 8× the base model.
Direct comparison on Qwen3.5-35B:
GRPO: 54.28% validity / 23.05% path solving / 14.64% Lean pass
GRPO-PRM: Lean pass 17.36%
SAGE: 59.83% validity / 31.76% path solving / 23.69% Lean pass
Conclusion and Future Work
SAGE demonstrates that structural guidance during training—algebraic sparsification for exploration bias and hyperbolic structure guidance for accumulation bias—substantially improves long-horizon reasoning without requiring step-level gold annotations or inference-time overhead. The method shifts learning from pure terminal rewards to leveraging the reasoning process's intrinsic structure.
Future work must address constructing reliable structural signals for tasks where rules are less explicit than in AC. Nevertheless, this work shows a promising direction: teaching models to choose the next step more effectively by understanding the structural landscape of the problem.
References
Paper: https://arxiv.org/abs/2609.30192
Code: https://github.com/Susan571/SAGE-NeurIPS2026
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
