Tagged articles

agentic reinforcement learning

4 articles · Page 1 of 1
Tencent Advertising Technology
Tencent Advertising Technology
Jul 17, 2026 · Artificial Intelligence

AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)

AdPilot, the first end‑to‑end autonomous advertising agent, reformulates ad delivery as a Markov decision process and combines structured memory, LLM‑enhanced reasoning, and a GRPO‑based reinforcement‑learning engine, while the newly released AdBench benchmark evaluates its superior performance across 38 scenarios and 7,600 instances, outperforming strong baselines by up to 11.76%.

AdBenchAdPilotKDD 2026
0 likes · 16 min read
AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)
Data Party THU
Data Party THU
Oct 20, 2025 · Artificial Intelligence

How Agentic RL Enables a 14B LLM to Outperform Giant Models – Inside rStar2‑Agent

This article analyzes the rStar2‑Agent paper, revealing how Agentic Reinforcement Learning, the GRPO‑RoC algorithm, a high‑throughput code‑execution service, and a three‑stage training recipe let a modest 14‑billion‑parameter model surpass much larger LLMs on challenging math benchmarks.

AI researchArtificial IntelligenceLLM
0 likes · 18 min read
How Agentic RL Enables a 14B LLM to Outperform Giant Models – Inside rStar2‑Agent
DataFunTalk
DataFunTalk
Sep 18, 2025 · Artificial Intelligence

How Tongyi DeepResearch Turns Chatty AI into a Research Powerhouse

Tongyi DeepResearch, an open‑source AI model and framework, achieves SOTA on multiple Deep Research benchmarks by combining fully open‑source models, frameworks, and data pipelines, and introduces novel agentic pre‑training, fine‑tuning, and reinforcement‑learning methods to enable complex multi‑step reasoning and real‑world applications.

AI researchOpen Sourceagentic reinforcement learning
0 likes · 14 min read
How Tongyi DeepResearch Turns Chatty AI into a Research Powerhouse