SAGE Corrects Long-Horizon Reasoning Drift via Structural Guidance, 8x AC Verification
NeurIPS 2026 paper SAGE introduces structural guidance via Symbolic Closure Analysis to fix long-horizon reasoning biases, using algebraic sparsification and hyperbolic distance feedback during training, outperforming baselines on 12 benchmarks and achieving nearly 8x Lean verification pass rate on Andrews-Curtis tasks.
Introduction
Large language models excel at short reasoning tasks but often drift off course in long-horizon reasoning, where each step appears locally valid yet the final answer fails. A NeurIPS 2026 paper from Virginia Tech, University of Wisconsin–Madison, and Dartmouth College introduces SAGE (Structural Admissibility-Guided Exploration), a post-training method that uses structural signals to correct two types of bias: exploration bias and accumulation bias.
Problem: Sparse Rewards in Long Reasoning
Post-training with outcome-only rewards has improved mathematical reasoning, but long tasks suffer from sparse, delayed rewards. A solution path may involve dozens or hundreds of transformations, yet training only observes success or failure at the end, making it hard to identify which early step caused the drift. The paper uses the Andrews–Curtis (AC) task as a stress test: given a group presentation, the model must apply legal AC transformations to reach the trivial presentation. Each step's legality is checkable, but legal moves do not guarantee progress toward the goal; the true success signal appears only at the sequence end. The authors evaluate on 1,190 constructed AC instances, amplifying the "locally legal, globally hard" dilemma.
Symbolic Closure Analysis (SCA): Diagnosing Two Biases
SCA defines the locally feasible set F as all prefixes where every step obeys the task's local operation rules. F is prefix-closed: once a step becomes illegal, no continuation can make the prefix legal again. However, F contains many paths that never reach the solution; the true successful trajectories are a smaller, unknown subset.
Exploration Bias
At depth t , let Bₜ be the total branching factor and Bₜᶠ the locally feasible branches. The fraction of feasible prefixes at depth T roughly scales with the product of Bₜᶠ/Bₜ across layers. When many layers have few feasible branches, this fraction shrinks exponentially, making it harder for random sampling under sparse rewards to hit useful trajectories. The paper decomposes policy-gradient variance into three parts: variance within feasible trajectories, variance within infeasible trajectories, and variance due to differing mean signals between the two groups. Concentrating sampling on F eliminates the latter two terms, explaining why branch selection affects learning signal quality.
Accumulation Bias
With only terminal rewards, early decisions lack direct corrective feedback. Under KL-regularized optimization, if the reference policy finds a successful trajectory with probability p , the terminal reward is bounded by Rmax , and KL penalty strength is λ , the total variation distance between the optimal and reference policies is bounded. When successful trajectories are extremely rare and rewards are weak relative to the KL constraint, this bound is small, so training struggles to deviate significantly from the reference model's early choices, allowing small deviations to persist along long chains.
SAGE: Turning Diagnosis into Training Guidance
SAGE implements two complementary structural potential functions that score candidate reasoning steps during training without providing ground-truth solutions.
Algebraic Sparsification (Targets Exploration Bias)
The system represents the unsolved structure as a "residual" and compares how much of this residual each candidate operation's algebraic subspace can explain. Operations explaining more residual receive higher soft compatibility scores, increasing their sampling probability without discarding other locally legal moves.
Hyperbolic Structural Guidance (Targets Accumulation Bias)
The current state and goal structure are mapped into a hyperbolic space suited for hierarchical relationships. The distance between them yields a step-wise feedback signal: prefixes closer to the goal receive higher scores, providing guidance before the terminal reward appears. This "closeness" is a training-time structural signal, not a proof of final correctness.
Training Procedure
An old policy proposes a set of candidate steps. SAGE reweights these candidates using the two potential functions. After the trajectory ends, the terminal reward and average structural score are combined into a group-relative advantage for policy update. The paper proves a conditional concentration result: if the potential functions separate all locally feasible from infeasible trajectories by a positive margin Δ , the reweighted probability ratio between the two groups is at most multiplied by exp(−λΔ) . Whether this condition holds depends on the task's structural priors.
These structural modules add candidate scoring and distance computation during training. At inference, the learned policy is used directly without running the scoring modules, so there is no extra inference overhead.
Experimental Results
Evaluation covers 7 closed-form math benchmarks, 4 free-form reasoning benchmarks, and the AC long-horizon symbolic task — 12 benchmarks total across 7 model families. Baselines include SFT, GRPO, EMPO, and GRPO-PRM (process reward model). All methods use comparable rollout budgets, generation lengths, and decoding constraints.
Closed-Form Mathematical Reasoning
On Qwen3.5 models, SAGE average accuracy: 2B → 42.11%, 9B → 47.95%, 35B → 64.86%. These exceed the strongest same-size post-training baselines by 1.83, 3.02, and 2.49 percentage points respectively. Not every single benchmark is won; e.g., on AIME the 35B model scores 62.04% vs. EMPO's 62.31%.
Free-Form Reasoning
At 9B, MMLU-Pro rises from 37.91% (EMPO) to 39.99% (SAGE); BBH-H from 44.04% to 45.31%; ARC-C from 39.73% to 42.04%. GPQA slightly drops from 21.11% to 20.86%. At 35B, SAGE leads on all four: BBH-H 69.07%, ARC-C 66.41%.
AC Long-Horizon Task
Three metrics: AC transform validity (local step legality), ACPathSolving (full path reaches goal), and Lean verification pass rate (formal proof check). Across six backbone models, SAGE improves transform validity by 19.2–26.0 percentage points and Lean pass rate by 13.2–20.8 percentage points. On Qwen3, SAGE's Lean pass rate approaches 8× the base model . Direct comparison on Qwen3.5-35B: GRPO achieves 54.28%/23.05%/14.64% (validity/path-solving/Lean); SAGE reaches 59.83%/31.76%/23.69%. GRPO-PRM scores 17.36% on Lean, still below SAGE's 23.69%.
Conclusion
As LLMs generate longer reasoning chains, step count does not guarantee staying on a solution path. Terminal rewards only indicate final correctness, not where the drift began. SCA shows how feasible branches become exponentially scarce with depth and why sparse terminal rewards cannot easily correct early choices. SAGE translates this analysis into a structural post-training method: algebraic sparsification steers sampling toward promising feasible branches, hyperbolic structural guidance gives prefix-level feedback via goal-distance in hyperbolic space, and both signals are used for policy updates. Training requires no gold-standard step-by-step annotations, and inference incurs no extra structural computation. Results across math, free-form, and AC tasks demonstrate substantial accuracy gains. This work opens a new training direction: teaching models to choose the next step by learning from the reasoning process's structure, not just from final outcomes. Future work must verify whether reliable structural signals can be built for tasks with less explicit rules than AC.
Paper: https://arxiv.org/abs/2609.30192
Code:
https://github.com/Susan571/SAGE-NeurIPS2026Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
