Tagged articles

online RL

3 articles · Page 1 of 1
DataFunSummit
DataFunSummit
Jul 7, 2026 · Artificial Intelligence

Ant Group OpAgent: Online RL‑Powered Open‑Domain Browser Automation Agent

The article details Ant Group's OpAgent, an open‑domain browser automation agent that overcomes perception, timeliness, and implicit interaction challenges through a three‑stage pipeline of multi‑task supervised fine‑tuning, online reinforcement learning, and a four‑module Planner‑Grounder‑Reflector‑Summarizer architecture, achieving a 71.6% Pass@1 score on WebArena and releasing all code and models publicly.

OpAgentWebArenaagent architecture
0 likes · 15 min read
Ant Group OpAgent: Online RL‑Powered Open‑Domain Browser Automation Agent
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning

The paper introduces EMCES, a method that injects episodic memory into controllable diffusion models and uses a hash‑based state representation to generate high‑value synthetic samples, dramatically improving sample efficiency and downstream reinforcement‑learning performance while cutting storage and time costs.

Episodic MemoryHashingOffline RL
0 likes · 14 min read
EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning
Kuaishou Tech
Kuaishou Tech
Nov 25, 2025 · Artificial Intelligence

How Flow‑GRPO Boosts Image Generation Accuracy to 95% with Online Reinforcement Learning

Flow‑GRPO introduces online reinforcement learning into flow‑matching models by converting deterministic ODE sampling to stochastic SDE sampling and reducing denoising steps, raising SD‑3.5‑Medium's GenEval accuracy from 63% to 95%—surpassing GPT‑4o—and demonstrating strong gains in complex composition, text rendering, and human‑preference alignment across multiple generative tasks.

AI researchFlow Matchingdeep learning
0 likes · 8 min read
How Flow‑GRPO Boosts Image Generation Accuracy to 95% with Online Reinforcement Learning