Two ACL 2026 Papers Solve Overthinking and Skill Forgetting in LLMs
This article summarizes two high-impact ACL 2026 papers: DRP reduces reasoning token usage by 64% while improving accuracy via skill-aware step pruning, and SAGE doubles skill reuse success rates in agents through sequential rollout and skill-integrated rewards.
Paper 1: DRP — Distilled Reasoning Pruning with Skill-aware Step Decomposition
Background: Reasoning models like o1 and DeepSeek-R1 rely on long chain-of-thought (CoT), but smaller models tend to overthink — producing redundant backtracking and self-correction that wastes tokens and can mislead the model. Existing solutions are imperfect: inference-time pruning risks cutting off reasoning prematurely, while distillation suffers from a style mismatch where the teacher emits concise Short-CoT but the student generates verbose Long-CoT, making the student unable to learn from the teacher.
Method: Student Generates, Teacher Prunes, Student Learns
DRP reverses the usual distillation flow:
Skill-aware Step Decomposition: The teacher segments the student's generated CoT into atomic steps, labeling each with the skill used (e.g., "read quantity", "algebraic representation"). This yields finer granularity (average 12.6 steps vs. 8.3 for sentence-level splitting) and more stable semantic boundaries.
Step-level Pruning: For each step the teacher chooses one of four actions — keep, delete, rewrite, or merge — then rewrites the step in the student's original tone so the pruned trace still "sounds like the student".
Supervised Fine-tuning: The pruned, concise trajectories are used for SFT, teaching the student a "short but accurate" reasoning habit.
DRP framework overview: student generates long CoT, teacher prunes, then distilled back to student.
Skill decomposition → four-action pruning → distillation back to student.
Results: Tokens Cut 60%, Accuracy Rises
GSM8K: Average tokens drop from 917 to 328 ( -64% ), accuracy improves from 91.7% to 94.1% .
AIME24: Token reduction 43% with no accuracy loss.
1.5B model: Largest beneficiary — GSM8K accuracy jumps 12.7% , showing concise supervision compensates for limited capacity.
Token distribution shifts left, long tail disappears.
Core insight: Good training CoT must not only be short but also structurally resemble the student's own writing — this is the prerequisite for knowledge transfer.
Paper 2: SAGE — Reinforcement Learning for Self-Improving Agent with Skill Library
Background: LLM agents trained with RL on fixed scenarios perform well but fail to generalize to new environments — they memorize answers rather than learning how to learn. A skill library (reusable functions from successful trajectories) is a natural solution, but current approaches rely on prompt-driven skill invocation, making skill quality and reuse rates unreliable.
Method: Embed Skill Acquisition into the Reward Function
SAGE builds a CodeAct-style skill-library agent where the agent emits Python functions as skills, supporting use, generate, update, and save operations. Two key designs:
Sequential Rollout: Instead of rolling out on a single task, the agent executes a chain of tasks within the same scenario . Skills generated in earlier tasks are automatically added to the library and available for later tasks. This allows the reward signal from "later task succeeded because earlier task generated a good skill" to propagate back.
Skill-integrated Reward: Beyond task outcome reward, when a skill generated in task q₁ is invoked by task q₂ and q₂ also succeeds, both trajectories receive an extra bonus — explicitly encouraging "generate well" and "use well".
Training follows two stages: SFT warm-start (Claude 3.5 expert trajectories) → SAGE RL fine-tuning.
Top: skill-library agent interaction loop. Bottom: Sequential Rollout and Skill-integrated Reward mechanism.
Results: Tokens -59%, Skill Success Rate Doubles
vs. GRPO baseline without skill library on AppWorld:
SGC: 51.8% → 60.7% (+8.9%)
TGC: reaches 72.0%
Efficiency gains: interaction steps -26% , generated tokens -59% — skill reuse compresses complex operation sequences into single calls.
Skill use success rate rises from base 0.31 to 0.72 (more than double).
Shared Principle
Both papers — one looking inward at reasoning traces, the other outward at environment interactions — share a core belief: "Skills" are decomposable, depositable, reusable structural units. DRP extracts and compresses skill steps from reasoning trajectories; SAGE stores and reuses skill functions from environment interactions. Both elevate "leveraging past experience" from prompt engineering to a principled training signal.
Paper title: DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition
Paper link: https://arxiv.org/abs/2505.13975
Paper title: Reinforcement Learning for Self-Improving Agent with Skill Library
Paper link: https://arxiv.org/abs/2512.17102Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
