Tagged articles

agentic reinforcement learning

4 articles · Page 1 of 1
Tencent Advertising Technology
Tencent Advertising Technology
Jul 17, 2026 · Artificial Intelligence

AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)

AdPilot, the first end‑to‑end autonomous advertising agent, reformulates ad delivery as a Markov decision process and combines structured memory, LLM‑enhanced reasoning, and a GRPO‑based reinforcement‑learning engine, while the newly released AdBench benchmark evaluates its superior performance across 38 scenarios and 7,600 instances, outperforming strong baselines by up to 11.76%.

AdBenchAdPilotKDD 2026
0 likes · 16 min read
AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)
AI2ML AI to Machine Learning
AI2ML AI to Machine Learning
Feb 24, 2026 · Artificial Intelligence

Optimizing Structured Processes in the Large‑Model Era: From Reasoning to Agentic RL

The article analyzes how large‑model development has moved from reasoning to the agentic stage, compares open‑source and closed‑source capabilities, details Reasoning RL versus Agentic RL designs, and proposes skill‑centric data and verification mechanisms to close the performance gap.

DeepSeekGLM-5Large Language Models
0 likes · 10 min read
Optimizing Structured Processes in the Large‑Model Era: From Reasoning to Agentic RL
DataFunTalk
DataFunTalk
Sep 18, 2025 · Artificial Intelligence

How Tongyi DeepResearch Turns Chatty AI into a Research Powerhouse

Tongyi DeepResearch, an open‑source AI model and framework, achieves SOTA on multiple Deep Research benchmarks by combining fully open‑source models, frameworks, and data pipelines, and introduces novel agentic pre‑training, fine‑tuning, and reinforcement‑learning methods to enable complex multi‑step reasoning and real‑world applications.

AI researchagentic reinforcement learningopen-source
0 likes · 14 min read
How Tongyi DeepResearch Turns Chatty AI into a Research Powerhouse