Tagged articles

ALFWorld

2 articles · Page 1 of 1
Data Party THU
Data Party THU
Sep 10, 2026 · Artificial Intelligence

Harness Continual Learning: How Agents Evolve Without Model Fine-Tuning

Researchers from Nanjing University and University of Wollongong propose Harness Continual Learning (HCL), a framework where AI agents continuously adapt by evolving their system-level harness—prompts, memory, skills, tools, and routing—while keeping the base model frozen, demonstrating improved performance and controlled forgetting across diverse benchmarks.

ALFWorldAgent Continual LearningDeepSeek-V4
0 likes · 16 min read
Harness Continual Learning: How Agents Evolve Without Model Fine-Tuning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 5, 2026 · Artificial Intelligence

StepOPSD: Precise Step‑Level Error Detection for Multi‑Turn Agent RL

StepOPSD adds a post‑hoc, step‑aware distillation stage to multi‑turn agent reinforcement learning, splitting rollouts into controllable steps, using successful trajectories as hindsight teachers to compute token‑level advantage adjustments, and demonstrating significant gains on ALFWorld and Search‑QA tasks where reward misalignment is most severe.

ALFWorldAdvantage WeightingSearch QA
0 likes · 13 min read
StepOPSD: Precise Step‑Level Error Detection for Multi‑Turn Agent RL