From Experience to Ability: How Agentic Skills Form, Internalize, and Self‑Evolve

In this MLNLP Academic Talk, Tsinghua PhD candidate Wu Jinyang presents his research on Agentic Skill formation, internalization, and continual self‑evolution, detailing three projects—ThoughtICR, TemplateRL, and SEED—that connect contextual reasoning, reinforcement learning, and autonomous skill growth.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
From Experience to Ability: How Agentic Skills Form, Internalize, and Self‑Evolve

Wu Jinyang, a fourth‑year PhD student in the Automation department at Tsinghua University, researches large‑model inference, agentic reinforcement learning, and self‑evolution, with a focus on the formation, internalization, and continual update of Agentic Skills.

The talk titled “From Experience to Ability: The Formation, Internalization and Self‑Evolution of Agentic Skill” outlines three recent research efforts.

ThoughtICR shifts context learning from mimicking concrete examples to reusing abstract reasoning patterns, automatically extracting inference skills that can transfer across tasks.

TemplateRL incorporates these extracted reasoning skills into reinforcement learning, guiding high‑quality trajectory exploration so that explicit policy experience gradually internalizes as model capability.

SEED extracts skills from the trajectories generated by the current policy and converts the resulting behavior improvements into fine‑grained training signals, achieving a synergistic co‑evolution of skill extraction and policy learning.

Collectively, these works illustrate how Agentic Skills can bridge contextual reasoning, reinforcement learning, and autonomous self‑evolution, offering a pathway for agents to move from merely completing tasks to continuously learning from experience.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelsReinforcement LearningSEEDself‑evolutionAgentic SkillTemplateRLThoughtICR
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.