How Libra Allocates Resources for Agentic RL Post‑Training and Boosts Throughput Up to 3×
The paper presents Libra, a resource‑management system for Agentic RL post‑training that jointly optimizes training and rollout GPU allocation using a global planner, heterogeneous inference clusters, a causality‑driven multi‑level feedback queue, and an elastic hybrid pool, achieving up to three‑fold throughput gains and up to 2.5× faster reward convergence.
