Tagged articles

multi-agent-rl

5 articles · Page 1 of 1
DaTaobao Tech
DaTaobao Tech
Jul 13, 2026 · Artificial Intelligence

Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning

The article details how Taobao Live upgraded its static workflow to a low‑latency Agentic architecture, applied AgentTuning distillation and RLVR to curb hallucinations, and introduced a Multi‑Agent RL framework that separates tool‑calling and reply generation, achieving significant gains in factual correctness, helpfulness, and overall performance.

Agentic RLLLMdigital human
0 likes · 23 min read
Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning
DeepHub IMBA
DeepHub IMBA
May 19, 2026 · Artificial Intelligence

A 2026 Survey of LLM‑Focused RL: From PPO to DPO, GRPO, and Multi‑Agent RL

The article reviews five years of LLM‑centric reinforcement learning, tracing the evolution from early Q‑learning to PPO, then to Direct Preference Optimization, Group Relative Policy Optimization, and finally multi‑agent RL, detailing each method’s mechanics, strengths, failure modes, practical considerations, and emerging open‑source toolchains.

DPOGRPOLLM alignment
0 likes · 33 min read
A 2026 Survey of LLM‑Focused RL: From PPO to DPO, GRPO, and Multi‑Agent RL
Bilibili Tech
Bilibili Tech
Aug 30, 2022 · Artificial Intelligence

Neural MMO Massive AI Team Survival Challenge: Advances in Multi‑Agent Decision AI

The IJCAI‑2022 Neural MMO Massive AI Team Survival Challenge demonstrated that deep reinforcement‑learning agents can achieve sophisticated cooperation and competition among 128 agents in a large‑scale MMO‑style world, highlighting the growing focus on decision‑AI, the effectiveness of self‑play and CTDE, and the platform’s potential for future research into population‑level behavior, economics, and complex real‑world decision making.

AI competitionDecision AIMassive AI
0 likes · 11 min read
Neural MMO Massive AI Team Survival Challenge: Advances in Multi‑Agent Decision AI