Tagged articles

policy improvement

2 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL

The paper by researchers from Beihang, Peking University and Meituan proposes PIRL, a new RL‑post‑training perspective that treats policy improvement as the optimization objective, and PIPO, a plug‑and‑play framework that adds a verification loop to amplify beneficial updates and suppress harmful ones, demonstrating consistent gains across math reasoning, code and tool‑use tasks.

Importance SamplingPIPOPIRL
0 likes · 9 min read
Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL
AI Algorithm Path
AI Algorithm Path
May 19, 2025 · Artificial Intelligence

Understanding Policy Evaluation and Improvement in Reinforcement Learning

This article explains how to solve Bellman equations, use iterative policy‑evaluation methods, apply the policy‑improvement theorem, and combine both steps in policy iteration, value iteration, and asynchronous variants, illustrated with a 5‑state example and a 4×4 gridworld.

Bellman equationGridWorldgeneralized policy iteration
0 likes · 15 min read
Understanding Policy Evaluation and Improvement in Reinforcement Learning