Two ACL 2026 Papers Solve Overthinking and Skill Forgetting in LLMs

This article summarizes two high-impact ACL 2026 papers: DRP reduces reasoning token usage by 64% while improving accuracy via skill-aware step pruning, and SAGE doubles skill reuse success rates in agents through sequential rollout and skill-integrated rewards.

PaperAgent
PaperAgent
PaperAgent
Two ACL 2026 Papers Solve Overthinking and Skill Forgetting in LLMs

Paper 1: DRP — Distilled Reasoning Pruning with Skill-aware Step Decomposition

Background: Reasoning models like o1 and DeepSeek-R1 rely on long chain-of-thought (CoT), but smaller models tend to overthink — producing redundant backtracking and self-correction that wastes tokens and can mislead the model. Existing solutions are imperfect: inference-time pruning risks cutting off reasoning prematurely, while distillation suffers from a style mismatch where the teacher emits concise Short-CoT but the student generates verbose Long-CoT, making the student unable to learn from the teacher.

Method: Student Generates, Teacher Prunes, Student Learns

DRP reverses the usual distillation flow:

Skill-aware Step Decomposition: The teacher segments the student's generated CoT into atomic steps, labeling each with the skill used (e.g., "read quantity", "algebraic representation"). This yields finer granularity (average 12.6 steps vs. 8.3 for sentence-level splitting) and more stable semantic boundaries.

Step-level Pruning: For each step the teacher chooses one of four actions — keep, delete, rewrite, or merge — then rewrites the step in the student's original tone so the pruned trace still "sounds like the student".

Supervised Fine-tuning: The pruned, concise trajectories are used for SFT, teaching the student a "short but accurate" reasoning habit.

DRP framework overview
DRP framework overview

DRP framework overview: student generates long CoT, teacher prunes, then distilled back to student.

DRP three-step flow
DRP three-step flow

Skill decomposition → four-action pruning → distillation back to student.

Results: Tokens Cut 60%, Accuracy Rises

GSM8K: Average tokens drop from 917 to 328 ( -64% ), accuracy improves from 91.7% to 94.1% .

AIME24: Token reduction 43% with no accuracy loss.

1.5B model: Largest beneficiary — GSM8K accuracy jumps 12.7% , showing concise supervision compensates for limited capacity.

Token distribution shifts left, long tail disappears.

Token distribution shift
Token distribution shift

Core insight: Good training CoT must not only be short but also structurally resemble the student's own writing — this is the prerequisite for knowledge transfer.

Paper 2: SAGE — Reinforcement Learning for Self-Improving Agent with Skill Library

Background: LLM agents trained with RL on fixed scenarios perform well but fail to generalize to new environments — they memorize answers rather than learning how to learn. A skill library (reusable functions from successful trajectories) is a natural solution, but current approaches rely on prompt-driven skill invocation, making skill quality and reuse rates unreliable.

Method: Embed Skill Acquisition into the Reward Function

SAGE builds a CodeAct-style skill-library agent where the agent emits Python functions as skills, supporting use, generate, update, and save operations. Two key designs:

Sequential Rollout: Instead of rolling out on a single task, the agent executes a chain of tasks within the same scenario . Skills generated in earlier tasks are automatically added to the library and available for later tasks. This allows the reward signal from "later task succeeded because earlier task generated a good skill" to propagate back.

Skill-integrated Reward: Beyond task outcome reward, when a skill generated in task q₁ is invoked by task q₂ and q₂ also succeeds, both trajectories receive an extra bonus — explicitly encouraging "generate well" and "use well".

Training follows two stages: SFT warm-start (Claude 3.5 expert trajectories) → SAGE RL fine-tuning.

SAGE framework: skill-library agent and sequential rollout
SAGE framework: skill-library agent and sequential rollout

Top: skill-library agent interaction loop. Bottom: Sequential Rollout and Skill-integrated Reward mechanism.

Results: Tokens -59%, Skill Success Rate Doubles

vs. GRPO baseline without skill library on AppWorld:

SGC: 51.8% → 60.7% (+8.9%)

TGC: reaches 72.0%

Efficiency gains: interaction steps -26% , generated tokens -59% — skill reuse compresses complex operation sequences into single calls.

Skill use success rate rises from base 0.31 to 0.72 (more than double).

SAGE results on AppWorld
SAGE results on AppWorld

Shared Principle

Both papers — one looking inward at reasoning traces, the other outward at environment interactions — share a core belief: "Skills" are decomposable, depositable, reusable structural units. DRP extracts and compresses skill steps from reasoning trajectories; SAGE stores and reuses skill functions from environment interactions. Both elevate "leveraging past experience" from prompt engineering to a principled training signal.

Paper title: DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition
Paper link: https://arxiv.org/abs/2505.13975

Paper title: Reinforcement Learning for Self-Improving Agent with Skill Library
Paper link: https://arxiv.org/abs/2512.17102
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

reinforcement learningLLM reasoningskill libraryACL 2026DRPSAGEagent skill reusechain-of-thought pruning
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.