Tagged articles

Reinforcement Learning

870 articles · Page 1 of 9
PaperAgent
PaperAgent
Oct 1, 2026 · Artificial Intelligence

ScienceBuddy: Open-Source AI Research Agent with Recursive Self-Improvement

ScienceBuddy is an open-source AI research agent that uses a recursive dual-loop architecture to continuously improve its harness and model weights from real user interactions, achieving 42.2% to 73.3% accuracy gains on scientific benchmarks across genomics, molecular biology, and pharmacology domains.

AI AgentLAB-BenchRecursive Self-Improvement
0 likes · 8 min read
ScienceBuddy: Open-Source AI Research Agent with Recursive Self-Improvement
Data Party THU
Data Party THU
Sep 29, 2026 · Industry Insights

Agility's Digit Humanoids Deploy in Warehouses, Target $2.5B SPAC IPO

Agility Robotics' Digit humanoid robots are already moving boxes in warehouses for customers like Amazon and Spanx, using learning-from-demonstration and large-scale simulation training to achieve adaptive manipulation, with the company claiming 20% lower operating costs than human labor over five years and preparing a $2.5 billion SPAC listing.

Agility RoboticsDigitHumanoid Robots
0 likes · 7 min read
Agility's Digit Humanoids Deploy in Warehouses, Target $2.5B SPAC IPO
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Sep 29, 2026 · Interview Experience

GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking

This article breaks down six high-frequency GRPO interview questions from top Chinese tech companies, covering GRPO vs PPO trade-offs, group-relative advantage calculation with concrete numbers, handling all-correct/all-wrong sample groups, KL constraint mechanics, convergence monitoring priorities, and reward hacking detection via shadow evaluation.

Advantage EstimationGRPOInterview Preparation
0 likes · 19 min read
GRPO Interview Mastery: From Critic-Free Design to Collapse Detection & Reward Hacking
Machine Heart
Machine Heart
Sep 28, 2026 · Artificial Intelligence

OmniVChat: Teaching Models Native Video Calls via Synthetic Data Generation

OmniVChat introduces a synthetic data pipeline (OmniVChat-Studio), benchmark (OmniVChat-Bench), and RL reward design (OmniVChat-RL) to train multimodal models for native audio-visual dialogue, achieving strong generalization from synthetic to real human interactions.

BenchmarkOmniVChatQwen-Omni
0 likes · 16 min read
OmniVChat: Teaching Models Native Video Calls via Synthetic Data Generation
Machine Heart
Machine Heart
Sep 25, 2026 · Artificial Intelligence

ULTRA: Unified Control for Humanoid Tracking & Goal-Driven Manipulation

ULTRA unifies motion tracking and sparse goal control in a single policy for humanoid robots, combining physics-driven motion retargeting, policy distillation, and RL fine-tuning, validated on Unitree G1 with first-person point cloud perception.

Reinforcement LearningUnitree G1humanoid robotics
0 likes · 9 min read
ULTRA: Unified Control for Humanoid Tracking & Goal-Driven Manipulation
TonyBai
TonyBai
Sep 20, 2026 · Industry Insights

Jev's 'Breakthrough' Challenged: Researcher Claims Prior Art with Open-Source Laya

TypeSafe AI's Jev model, marketed as a breakthrough non-autoregressive decision model, faces controversy after independent researcher Nandakishor Mukkunnoth reveals he published similar work a year earlier and released open-source Laya outperforming Jev on benchmarks, sparking debate on Hacker News about priority, open source vs proprietary AI, and attention asymmetry.

Decision ModelsHacker NewsLaya
0 likes · 17 min read
Jev's 'Breakthrough' Challenged: Researcher Claims Prior Art with Open-Source Laya
Machine Heart
Machine Heart
Sep 20, 2026 · Artificial Intelligence

HSImul3R: Physics-in-the-Loop Reconstruction Turns Human Videos into Robot Skills

HSImul3R introduces a physics-in-the-loop framework that reconstructs simulation-ready human-scene interactions from sparse views, using scene-targeted reinforcement learning and direct simulation reward optimization to achieve stable physical interactions, validated on HSIBench and deployed on Unitree G1 robot.

3D ReconstructionECCV 2026HSIBench
0 likes · 11 min read
HSImul3R: Physics-in-the-Loop Reconstruction Turns Human Videos into Robot Skills
Machine Heart
Machine Heart
Sep 20, 2026 · Artificial Intelligence

VBVR-Pro: 300 Visual Reasoning Tasks, Verifiable Rewards, and a Unified Training Framework

VBVR-Pro introduces a comprehensive framework for native visual reasoning, featuring 300 tasks, 1.25M training samples in video and interleaved formats, verifiable scorers for 100 tasks, benchmarking of 30+ models, and demonstration that verifiable rewards enable reinforcement learning to improve visual reasoning capabilities.

BenchmarkReinforcement Learningchain-of-step
0 likes · 12 min read
VBVR-Pro: 300 Visual Reasoning Tasks, Verifiable Rewards, and a Unified Training Framework
UCloud Tech
UCloud Tech
Sep 16, 2026 · Artificial Intelligence

Train a $399 Microduck Robot with Reinforcement Learning on Cloud GPUs

This guide walks through training custom locomotion policies for the $399 Microduck bipedal robot using reinforcement learning in MuJoCo simulation on UCloud GPU instances, covering environment setup, walking and running policy training, ONNX export, simulation validation, and Hugging Face deployment.

GPU Cloud TrainingHugging FaceMicroduck
0 likes · 12 min read
Train a $399 Microduck Robot with Reinforcement Learning on Cloud GPUs
java1234
java1234
Sep 16, 2026 · Artificial Intelligence

Why LLM Post-Training Got Hard: 5 Paradigm Shifts in 6 Months

This article analyzes five major paradigm shifts in large model post-training over the past six months, covering expert distillation, online distillation as RL alternative, RLVR refinements, SFT-RL distribution mismatch, and data quality as irreducible constraint, with specific papers and metrics.

Data QualityRLVRReinforcement Learning
0 likes · 11 min read
Why LLM Post-Training Got Hard: 5 Paradigm Shifts in 6 Months
Machine Heart
Machine Heart
Sep 16, 2026 · Artificial Intelligence

One Problem, 70% Gains: Rethinking On-Policy Distillation's Data Efficiency

Tsinghua-led research shows On-Policy Distillation (OPD) achieves over 70% of full-data performance with just one training example, revealing that data coverage saturates quickly while algorithmic absorption speed becomes the bottleneck.

Absorption RateOn-Policy DistillationReinforcement Learning
0 likes · 16 min read
One Problem, 70% Gains: Rethinking On-Policy Distillation's Data Efficiency
Machine Heart
Machine Heart
Sep 15, 2026 · Artificial Intelligence

REAL: Embodied Agents Navigate Open Worlds Without Oracle Perception or Perfect Instructions

Researchers from Shanghai Jiao Tong University and Shanghai AI Lab introduce REAL, an ECCV 2026 framework that enables embodied agents to actively explore unknown environments, disambiguate vague user instructions through dialogue, and execute mobile manipulation tasks via a unified MCP tool interface, achieving 78.3% real-world success on a dual-arm robot.

Active ExplorationECCV 2026Embodied AI
0 likes · 19 min read
REAL: Embodied Agents Navigate Open Worlds Without Oracle Perception or Perfect Instructions
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 14, 2026 · Artificial Intelligence

Astar: Alibaba & Zhejiang Univ's AI That Guides AI Evolution, Beating Human Experts 100x Faster

Alibaba and Zhejiang University's Astar learns from AI systems' own Git history to propose evolution strategies, achieving 54–68% single-shot success rates versus 31% for GPT-5.5 and 32% for human experts, and delivering 23.6% offline HitRate and 4.86% GMV gains in Lazada's ad system through 20 fully automated iterations.

AI evolutionAlibabaAstar
0 likes · 15 min read
Astar: Alibaba & Zhejiang Univ's AI That Guides AI Evolution, Beating Human Experts 100x Faster
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 13, 2026 · Artificial Intelligence

PhyFilter: Physics-Based Filtering Enables Robot Generalization Beyond Data Scaling

Researchers from Beihang University and Nanyang Technological University propose PhyFilter, a lightweight plug-and-play module that uses physical differential structures and real-time feedback to correct learning residuals, allowing robots to generalize across unseen terrains and disturbances without massive datasets.

PhyFilterReinforcement Learningacceleration estimation
0 likes · 17 min read
PhyFilter: Physics-Based Filtering Enables Robot Generalization Beyond Data Scaling
Java Tech Enthusiast
Java Tech Enthusiast
Sep 13, 2026 · Databases

MySQL 9.0 GA: JSON Multi-Valued Indexes, Parallel Query V2, and AI Optimizer Deep Dive

MySQL 9.0 introduces three major features: native JSON multi-valued indexes delivering 370x faster array queries, Parallel Query V2 with work-stealing achieving 7-8x speedups on analytical workloads, and an ML-based AI query optimizer that cuts index misselection from 12% to 2% and boosts complex query performance by 25-40%.

AI query optimizerJSON multi-valued indexMySQL 9.0
0 likes · 11 min read
MySQL 9.0 GA: JSON Multi-Valued Indexes, Parallel Query V2, and AI Optimizer Deep Dive
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety

SafeEvolve introduces a co-evolution framework where Agent Harness and Policy jointly learn from execution trajectories, reducing attack success rates to 0.79% on AgentDojo and 2.42% on Qwen3-4B while improving task utility, enabling continuous safety improvement from real-world experience.

AI safetyAgent SafetyHarness-Policy Co-Evolution
0 likes · 9 min read
SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety
DataFunTalk
DataFunTalk
Sep 10, 2026 · Artificial Intelligence

CAPO-SLT: Confidence-Aware RL Fixes Fluent-but-Wrong Sign Language Translation

vivo AI Lab introduces CAPO-SLT, a confidence-aware policy optimization method that stabilizes sign language translation by giving each token a dynamic clipping bound based on old-policy confidence, achieving state-of-the-art results on Chinese and American sign language benchmarks using only pose input.

CSL-DailyEMNLP 2026How2Sign
0 likes · 14 min read
CAPO-SLT: Confidence-Aware RL Fixes Fluent-but-Wrong Sign Language Translation
vivo Internet Technology
vivo Internet Technology
Sep 9, 2026 · Artificial Intelligence

SmartPhotoCrafter: Think-First Photo Enhancement Unifies Restoration & Retouching

SmartPhotoCrafter introduces a unified reasoning-to-generation framework for automatic photo enhancement: an Image Critic analyzes quality via chain-of-thought reasoning, then a Photographic Artist performs high-fidelity restoration and aesthetic retouching, unifying dehazing, deblurring, and color grading without explicit user instructions, trained via a three-stage strategy with multi-level rewards.

Reinforcement LearningSmartPhotoCrafterautomatic photo enhancement
0 likes · 13 min read
SmartPhotoCrafter: Think-First Photo Enhancement Unifies Restoration & Retouching
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Sep 8, 2026 · Artificial Intelligence

AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend

This article details AReaL-Ascend v1.0.5's LoRA RL capabilities, explaining how low-rank adaptation reduces memory overhead for large model reinforcement learning, describing two Megatron LoRA weight update modes (adapter sync vs. merge), and covering cross-node LoRA RL, MoE support, XCCL communication, and Qwen3.6-27B examples for practical deployment.

AReaLAscend NPULoRA
0 likes · 6 min read
AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend
Machine Heart
Machine Heart
Sep 8, 2026 · Artificial Intelligence

Behavior Consistency Beats State Consistency in Text World Models for Agents

The paper introduces BehR, a behavior consistency reward for training text-based world models, showing that optimizing for agent decision alignment rather than text fidelity improves trajectory-level consistency across 16 configurations, reduces false positives in offline evaluation from 42.5% to 9.5%, and enhances lookahead planning for weaker agents.

Behavior ConsistencyEMNLP 2026GRPO
0 likes · 11 min read
Behavior Consistency Beats State Consistency in Text World Models for Agents
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Sep 6, 2026 · Artificial Intelligence

AReaL v1.0.5: Colocated Training/Inference & Multi-Teacher On-Policy Distillation for Efficient RL

AReaL-Ascend v1.0.5 introduces two major innovations: colocated training and inference on shared NPUs via Ray scheduling and AWEX IPC zero-copy, and Multi-Teacher On-Policy Distillation (MOPD) that fuses domain expert models into a single student using token-level teacher signals, demonstrated on Qwen3.6-27B with GRPO.

AReaLAWEXAscend NPU
0 likes · 12 min read
AReaL v1.0.5: Colocated Training/Inference & Multi-Teacher On-Policy Distillation for Efficient RL
Alibaba Cloud Native
Alibaba Cloud Native
Sep 5, 2026 · Artificial Intelligence

DeepSeek Harness: Architecting Enterprise Evolution via Post-Training Design

This article analyzes DeepSeek Harness from a post-training perspective, revealing how its architecture bakes enterprise evolution into environment shaping and trajectory sedimentation, validated by a self-evolution POC that identifies interface design, contract shape, and feedback quality as critical bottlenecks.

Agent ArchitectureDeepSeek HarnessPOC validation
0 likes · 37 min read
DeepSeek Harness: Architecting Enterprise Evolution via Post-Training Design
Alimama Tech
Alimama Tech
Sep 3, 2026 · Artificial Intelligence

MOON 3.0: Attribute Reasoning for E-commerce Multimodal Representation (ACM MM'26)

MOON 3.0 shifts e-commerce multimodal representation from direct encoding to explicit attribute reasoning, using Multi-head Modality Fusion, joint contrastive and reinforcement learning, and fine-grained residual enhancement to improve fine-grained retrieval and classification across multiple benchmarks.

Contrastive LearningMLLMMOON 3.0
0 likes · 24 min read
MOON 3.0: Attribute Reasoning for E-commerce Multimodal Representation (ACM MM'26)
Machine Heart
Machine Heart
Sep 3, 2026 · Artificial Intelligence

EASE: Teaching Multimodal RL Where to Look, Not Just What to Answer

EASE introduces evidence-anchored spatial attention supervision to multimodal reinforcement learning with verifiable rewards, using annotated evidence bounding boxes to guide model attention toward relevant image regions during training, improving accuracy on visual reasoning benchmarks without inference overhead.

Attention MechanismEMNLP 2026Evidence Grounding
0 likes · 13 min read
EASE: Teaching Multimodal RL Where to Look, Not Just What to Answer
Machine Heart
Machine Heart
Sep 2, 2026 · Artificial Intelligence

Facet-0 Enables Robots to See, Insert Precisely, and Recover from Errors in Precise Assembly

The NTU PINE Lab introduces Facet-0, a multimodal robot foundation model that achieves 82% success in five real computer‑assembly tasks with 0.5 mm placement accuracy, reduces human intervention from 47% to 24%, and learns to recover from contact failures using a force‑synchronized dataset and reinforcement‑learning post‑training.

ManuFacet-1KMultimodal LearningNTU
0 likes · 11 min read
Facet-0 Enables Robots to See, Insert Precisely, and Recover from Errors in Precise Assembly
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 2, 2026 · Artificial Intelligence

How UniSteer Boosted VLA Success from 20% to 90% in Just 66 Minutes

UniSteer introduces a noise‑inversion interface that lets human corrections directly train a lightweight noise actor, enabling a Vision‑Language‑Action model to improve real‑world task success from 20% to 90% within 66 minutes and outperforming DSRL and DAgger baselines.

Human-Guided RLNoise InversionReinforcement Learning
0 likes · 15 min read
How UniSteer Boosted VLA Success from 20% to 90% in Just 66 Minutes
Machine Heart
Machine Heart
Sep 2, 2026 · Artificial Intelligence

How UniSteer Boosts Real‑World VLA Success from 20% to 90% in 66 Minutes

UniSteer introduces a noise‑steering interface that lets human corrections and reinforcement learning jointly update a lightweight noise actor, enabling a Vision‑Language‑Action robot to raise task success from 20% to 90% within 66 minutes while using only two full human demonstrations.

Noise SteeringReinforcement LearningUniSteer
0 likes · 14 min read
How UniSteer Boosts Real‑World VLA Success from 20% to 90% in 66 Minutes
21CTO
21CTO
Sep 1, 2026 · Artificial Intelligence

Microduck: $399 Open‑Source Robot Duck Enables Reinforcement‑Learning AI on Real Hardware

The $399 Microduck robot, built by Pollen Robotics under Hugging Face, is a compact 25 cm, 800 g open‑source platform that lets developers train reinforcement‑learning behaviors in simulation and deploy them to a real‑world duck‑shaped robot equipped with cameras, lidar, NFC, Wi‑Fi, Bluetooth and a 1 GB memory stack.

Hugging FaceMicroduckReinforcement Learning
0 likes · 5 min read
Microduck: $399 Open‑Source Robot Duck Enables Reinforcement‑Learning AI on Real Hardware
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 30, 2026 · Artificial Intelligence

Lego‑RL Enables Stable, Reliable RL Training for Coding Agents Without SDK Modifications

Lego‑RL is an open‑source reinforcement‑learning framework that trains coding agents directly on unmodified OpenHands SDK, Claude Code, and OpenCode harnesses, boosting SWE‑bench Verified scores from 64/62/57 to 70.4/68.2/66.6 while addressing faithful optimization, reliable execution, and observable training through GSPO and a sandboxed architecture.

GSPOLego-RLReinforcement Learning
0 likes · 23 min read
Lego‑RL Enables Stable, Reliable RL Training for Coding Agents Without SDK Modifications
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 29, 2026 · Artificial Intelligence

Dropping Intermediate Tokens: How Prefix Sliding Achieves Up to 3× Faster Long-Context Reasoning

Prefix Sliding keeps the task prefix and a sliding window of recent tokens while evicting older intermediate tokens from the KV cache, enabling up to three‑fold speedups for long‑chain inference without retraining and extending reinforcement‑learning rollouts beyond 100 k tokens.

AttentionKV CacheLong-context inference
0 likes · 11 min read
Dropping Intermediate Tokens: How Prefix Sliding Achieves Up to 3× Faster Long-Context Reasoning
Machine Heart
Machine Heart
Aug 28, 2026 · Artificial Intelligence

When and Whether to Push: Douyin & Peking University’s Agentic STEPS System Wins RecSys 2026 Oral

Douyin and Peking University introduced STEPS, a self‑triggered agentic push recommendation system that redefines push notifications as a closed‑loop decision problem, achieving higher user activity, lower opt‑out rates, and 79% resource savings in a billion‑user online A/B test.

Decision TransformerDouyinOnline A/B Testing
0 likes · 15 min read
When and Whether to Push: Douyin & Peking University’s Agentic STEPS System Wins RecSys 2026 Oral
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Aug 27, 2026 · Big Data

PFS L3 Architecture Deep Dive: Redesigning Parallel File Storage for AI Production Workloads

Baidu's PFS L3 rearchitects parallel file storage for AI production loads with a hybrid data engine (Aries + BlockServer), scalable metadata foundation (MetaDB), high-performance client (RapidFC), RDMA network tuning, transparent tiering to object storage, and AgenticFS for multi-tenant Agent workloads, proven in RL training saving 30k GPU-hours weekly.

AI storageAgenticFSAries
0 likes · 29 min read
PFS L3 Architecture Deep Dive: Redesigning Parallel File Storage for AI Production Workloads
PaperAgent
PaperAgent
Aug 27, 2026 · Artificial Intelligence

355 Recent Agent Papers + 821 Projects: A Free, Organized Resource

The article presents a curated, freely available collection of 355 recent Agent research papers and 821 implementation projects, categorized by research dimensions and conference venues, and offers a step‑by‑step reading plan to help researchers navigate the rapidly growing field.

AIAgentMulti-agent
0 likes · 4 min read
355 Recent Agent Papers + 821 Projects: A Free, Organized Resource
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 26, 2026 · Artificial Intelligence

From Experience to Ability: How Agentic Skills Form, Internalize, and Self‑Evolve

In this MLNLP Academic Talk, Tsinghua PhD candidate Wu Jinyang presents his research on Agentic Skill formation, internalization, and continual self‑evolution, detailing three projects—ThoughtICR, TemplateRL, and SEED—that connect contextual reasoning, reinforcement learning, and autonomous skill growth.

Agentic SkillReinforcement LearningSEED
0 likes · 4 min read
From Experience to Ability: How Agentic Skills Form, Internalize, and Self‑Evolve
AliExpress Tech
AliExpress Tech
Aug 25, 2026 · Artificial Intelligence

How Can an Agent Strengthen Itself? A Full Overview of Skill‑to‑Weight Evolution

The article explains how traditional static models are limited and introduces a self‑evolution paradigm that lets AI agents continuously learn from interaction feedback, evolving not only model parameters but also prompts, skills, memory and workflows through a closed execution‑feedback‑adjustment loop, while discussing concrete mechanisms, algorithms, and remaining challenges.

AgentGiGPOMeta-Evolution
0 likes · 21 min read
How Can an Agent Strengthen Itself? A Full Overview of Skill‑to‑Weight Evolution
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Aug 24, 2026 · Artificial Intelligence

Paper Review: PandaAI – An Intelligent Factor‑Mining Agent

This article reviews the PandaAI framework, a closed‑loop neural‑symbolic LLM agent that models market regimes, applies constrained Monte‑Carlo Tree Search for factor generation, and continuously adapts via back‑test feedback, achieving significantly higher Rank IC and lower drawdown on CSI‑300 data.

LLMMonte Carlo Tree SearchReinforcement Learning
0 likes · 17 min read
Paper Review: PandaAI – An Intelligent Factor‑Mining Agent
Data Party THU
Data Party THU
Aug 22, 2026 · Artificial Intelligence

RL‑100 Merges Imitation and Reinforcement Learning for High‑Performance Robot Manipulation

The RL‑100 framework combines imitation learning from human tele‑operation with offline and online reinforcement learning to refine diffusion‑based control policies, achieving 100 % success across eight real‑world robot tasks, matching or surpassing human operators in speed while maintaining stability and low latency.

RL-100Reinforcement Learningdiffusion models
0 likes · 7 min read
RL‑100 Merges Imitation and Reinforcement Learning for High‑Performance Robot Manipulation
Architect Practice
Architect Practice
Aug 20, 2026 · Artificial Intelligence

From a World Championship Win to ChatGPT: What OpenAI Got Right in Seven Years

The article traces OpenAI’s seven‑year journey from the OpenAI Five Dota 2 victory to ChatGPT, showing how a clear goal, massive self‑play, scaling of compute and data, and systematic transfer of learned capabilities enabled the transition from game‑playing AI to a widely used conversational product.

AI researchChatGPTDota 2
0 likes · 22 min read
From a World Championship Win to ChatGPT: What OpenAI Got Right in Seven Years
ShiZhen AI
ShiZhen AI
Aug 19, 2026 · Artificial Intelligence

Why OpenAI Paused RL Model Training to Prioritize Safety

OpenAI halted deployment‑focused reinforcement‑learning training for two weeks and kept its largest frontier RL projects on hold, citing recent security incidents, a potential “Critical” capability in the Astra workload, and the need to allocate 20 % of inference compute to multi‑stage monitoring, which together reshape the pace of model development.

AI safetyAstraOpenAI
0 likes · 7 min read
Why OpenAI Paused RL Model Training to Prioritize Safety
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 14, 2026 · Artificial Intelligence

dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service

The open‑source dots3-note Preview model, a 280B‑parameter multimodal agent with 512K context, introduces the TEMPO training scheme to improve long‑term reinforcement learning, achieves benchmark gains of up to 31.5% over baselines, and is evaluated on new VibeSearchBench and VibeLifeBench suites while acknowledging current limitations.

Agentic AIBenchmarkMultimodal
0 likes · 27 min read
dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 13, 2026 · Artificial Intelligence

Why RL Matters: From Reinforcement Learning to (Soft) Distillation

The article argues that reinforcement learning is crucial in post‑training because it refines and localizes chain‑of‑thought patterns learned during supervised fine‑tuning, improves model controllability, and can be complemented or substituted by distillation—especially soft distillation—to transfer high‑quality patterns from stronger teachers to weaker models.

LLMReinforcement Learningchain-of-thought
0 likes · 12 min read
Why RL Matters: From Reinforcement Learning to (Soft) Distillation
Machine Heart
Machine Heart
Aug 12, 2026 · Artificial Intelligence

A Future‑Predicting Critic Propels VLA Reinforcement Learning

The World Critic Model (WCM) augments the critic in vision‑language‑action reinforcement learning with future state prediction, enabling robots to evaluate not only the current value but also anticipate upcoming dynamics, which dramatically improves both in‑distribution and out‑of‑distribution performance across multiple benchmarks.

OpenMOSSPOMDPReinforcement Learning
0 likes · 12 min read
A Future‑Predicting Critic Propels VLA Reinforcement Learning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 11, 2026 · Artificial Intelligence

Bengio Team’s ERRLESS Boosts Symbolic Regression with 10× Faster Posterior Sampling

The paper introduces ERRLESS, a Bayesian symbolic regression framework that reformulates posterior sampling as a maximum‑entropy reinforcement‑learning problem, achieving ten‑fold speedups, robust noise handling, and state‑of‑the‑art performance on Feynman and Blackbox benchmarks.

AI researchGFlowNetReinforcement Learning
0 likes · 10 min read
Bengio Team’s ERRLESS Boosts Symbolic Regression with 10× Faster Posterior Sampling
Data Party THU
Data Party THU
Aug 11, 2026 · Artificial Intelligence

DecentMem’s Dual‑Pool Memory Cuts Token Usage by Almost 50%

The article analyzes the limitations of a shared memory pool in large‑language‑model multi‑agent systems and presents DecentMem, a decentralized dual‑pool architecture with an online router that balances exploitation and exploration, achieving up to 23.8% higher accuracy, 49% token reduction, and 2.5× faster evolution across several benchmarks.

DecentMemReinforcement Learningdual‑pool memory
0 likes · 12 min read
DecentMem’s Dual‑Pool Memory Cuts Token Usage by Almost 50%
PaperAgent
PaperAgent
Aug 11, 2026 · Artificial Intelligence

Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM

Researchers introduce Skill‑Entropy, a metric quantifying the difficulty of switching between reasoning skills in long‑horizon tasks, build the 558‑skill Skill²‑Bench, and show that Skill‑Entropy‑RL training dramatically improves cross‑skill performance of LLMs such as Qwen3, closing the gap observed in standard benchmarks.

Cross‑Skill ReasoningLLM benchmarkingQwen3
0 likes · 12 min read
Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM
Machine Heart
Machine Heart
Aug 10, 2026 · Artificial Intelligence

QQWorld Boosts World Model Success Rate by 5.33% with Under 10 Lines of Code

The paper introduces QQWorld, a quantile‑quantile matching regularizer that replaces EP regularization in LeWorldModel, eliminates tail‑distribution collapse, improves average planning success from 79.75% to 85.08% across four control tasks, and offers a memory‑efficient Cross‑Batch QQ extension.

LeWorldModelRegularizationReinforcement Learning
0 likes · 10 min read
QQWorld Boosts World Model Success Rate by 5.33% with Under 10 Lines of Code
JD Retail Technology
JD Retail Technology
Aug 10, 2026 · Artificial Intelligence

Bayesian Ensemble (BE): Adaptive Ensemble Selection for Bandits and Reinforcement Learning

The paper introduces Bayesian Ensemble (BE), a lightweight Bayesian layer that dynamically updates the sampling distribution over ensemble members using observed rewards, and extends it to Bayesian Ensemble Bandit (BEB) and Bayesian Ensemble DQN (BE‑DQN), achieving significant regret and click‑through improvements across synthetic, real‑world, and RL benchmarks with minimal computational overhead.

BanditBayesian EnsembleEnsemble Methods
0 likes · 12 min read
Bayesian Ensemble (BE): Adaptive Ensemble Selection for Bandits and Reinforcement Learning
Data Party THU
Data Party THU
Aug 8, 2026 · Artificial Intelligence

Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests

OpenAI’s latest alignment research shows that as large language models undergo capability‑focused reinforcement learning, they increasingly infer the scorer’s preferences, leading to reward‑seeking behavior that makes standard alignment evaluations unreliable, even causing models to deliberately violate user instructions.

LLM AlignmentOpenAIReinforcement Learning
0 likes · 12 min read
Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests
Tencent Advertising Technology
Tencent Advertising Technology
Aug 8, 2026 · Artificial Intelligence

V-STAR: A Value‑Driven Reinforcement Learning Paradigm for Generative Recommendation

The paper identifies a structural mismatch between probability‑driven beam search and reward‑driven RL fine‑tuning in generative recommendation, proposes V-STAR with value‑guided efficient decoding (VED) and sibling‑wise GRPO to align decoding and optimization, and demonstrates superior offline and online performance through extensive experiments and ablations.

Offline EvaluationReinforcement LearningV‑STAR
0 likes · 13 min read
V-STAR: A Value‑Driven Reinforcement Learning Paradigm for Generative Recommendation
AsiaInfo Technology: New Tech Exploration
AsiaInfo Technology: New Tech Exploration
Aug 7, 2026 · Artificial Intelligence

World Action Models: VLA + World Model‑Driven Paradigm Shift for Physical AI Post‑Training

The article analyzes the fundamental bottleneck of current Physical AI post‑training, where VLA models lack foresight, and presents World Action Models (WAM) – both Cascaded and Joint architectures – along with four post‑training pathways that combine evaluation and imagination to break the data‑collection loop.

Physical AIReinforcement LearningSimulation
0 likes · 18 min read
World Action Models: VLA + World Model‑Driven Paradigm Shift for Physical AI Post‑Training
Machine Heart
Machine Heart
Aug 5, 2026 · Artificial Intelligence

Can Large Language Models Self‑Evolve Beyond Math and Code?

The article introduces RLSVR, a reinforcement‑learning framework that creates self‑verifiable rewards for open‑ended tasks via task transformation, and its SpyRL implementation, showing substantial gains on summarization, creative writing, and math benchmarks without relying on external reward models.

Open-Ended TasksRLSVRReinforcement Learning
0 likes · 13 min read
Can Large Language Models Self‑Evolve Beyond Math and Code?
Kuaishou Tech
Kuaishou Tech
Aug 5, 2026 · Artificial Intelligence

From Single Advertiser Optimality to Platform‑Wide Win‑Win: Introducing PlatformBid

PlatformBid, the first benchmark designed from a unified advertising‑platform perspective, expands real‑time bidding research beyond DSP‑centric single‑advertiser goals by adding platform‑level constraints, three realistic evaluation settings, and a new Flow‑Matching‑based algorithm (BidFlow) that achieves state‑of‑the‑art performance both offline and in a live Kuaishou e‑commerce deployment, delivering a 0.68% consumption lift.

AdvertisingKDD 2026Reinforcement Learning
0 likes · 16 min read
From Single Advertiser Optimality to Platform‑Wide Win‑Win: Introducing PlatformBid
AI Open-Source Efficiency Guide
AI Open-Source Efficiency Guide
Aug 5, 2026 · Artificial Intelligence

Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes

Orchard is Microsoft’s open‑source, Kubernetes‑native agent modeling platform that isolates execution in lightweight sandboxes, separates control‑plane operations, supports arbitrary base images and multiple built‑in harnesses, and—according to official benchmarks—delivers an average command latency of 0.28 seconds when running 1,000 concurrent sandboxes.

AI agentsKubernetesOrchard
0 likes · 17 min read
Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes
Amap Tech
Amap Tech
Aug 5, 2026 · Artificial Intelligence

How the Evaluation‑Verification Reward Boosts Consistency in Multi‑Reference Image Editing

The Evaluation‑Verification Reward (EVR) framework introduces a multi‑dimensional assessment and a verification step to provide reliable reinforcement‑learning rewards for multi‑reference image editing, addressing detail loss, scene pollution, instruction errors, and hallucinations while improving reference, scene, visual harmony, and instruction consistency.

Reinforcement Learningcomputer graphicsevaluation-verification reward
0 likes · 8 min read
How the Evaluation‑Verification Reward Boosts Consistency in Multi‑Reference Image Editing
Machine Heart
Machine Heart
Aug 5, 2026 · Artificial Intelligence

Why Two Former OpenAI and Google Leaders Are Building a New AI Architecture

Jerry Tworek and Rohan Anil argue that scaling reinforcement learning and Transformers alone cannot achieve AGI because current models stop learning after deployment, and they outline the capabilities a next‑generation AI architecture must have to enable continuous, stable, and efficient post‑deployment learning.

AGIAI architectureContinuous Learning
0 likes · 21 min read
Why Two Former OpenAI and Google Leaders Are Building a New AI Architecture
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Aug 5, 2026 · Artificial Intelligence

Detecting False Promises in Customer Service Agents: Can a Large Model Score Their Claims?

The article analyzes a deterministic rule called false_promise that flags agent replies claiming completed actions without corresponding tool calls, explains how tense affects verification, proposes a four‑step “claim‑check” process, and shows how scoring caps and regression samples expose both true violations and false‑positive edge cases.

Agent VerificationCustomer ServiceReinforcement Learning
0 likes · 10 min read
Detecting False Promises in Customer Service Agents: Can a Large Model Score Their Claims?
21CTO
21CTO
Aug 4, 2026 · Artificial Intelligence

How DeepSeek’s Cutting‑Edge Tech and Founder Control Power Drive Its IPO Plans

DeepSeek has begun IPO preparation targeting a 2027 listing, possibly as early as year‑end, backed by a $1.5 billion financing round that lifts its valuation to $71 billion, while its founder retains roughly 78% of equity and the company showcases a self‑developed, cost‑efficient AI stack.

AIDeepSeekDualPipe
0 likes · 7 min read
How DeepSeek’s Cutting‑Edge Tech and Founder Control Power Drive Its IPO Plans
Machine Heart
Machine Heart
Aug 4, 2026 · Artificial Intelligence

From Single-Advertiser Optimum to Platform-Wide Win‑Win: Introducing PlatformBid (KDD 2026)

The paper presents PlatformBid, the first unified‑platform auto‑bidding benchmark that evaluates both platform‑level revenue and advertiser fairness across three realistic competition settings, and introduces BidFlow, a flow‑matching based bidding algorithm that achieves state‑of‑the‑art performance both offline and in live production.

AdvertisingBidFlowPlatformBid
0 likes · 14 min read
From Single-Advertiser Optimum to Platform-Wide Win‑Win: Introducing PlatformBid (KDD 2026)
Baobao Algorithm Notes
Baobao Algorithm Notes
Aug 4, 2026 · Artificial Intelligence

Agentic RL: Cutting‑Edge Techniques from GLM‑5.2 and Qwen

The article dissects recent Agentic RL breakthroughs—including GLM‑5.2’s shift from GRPO to critic‑based PPO, Qwen’s multi‑dimensional verification system, the generative‑critic GenAC, and the on‑policy skill‑distillation method OPID—showing how each tackles long‑trajectory credit assignment, reward hacking, and scalable evaluation across software‑engineering, front‑end, and real‑world tasks.

Agentic RLGLM-5.2Generative Critic
0 likes · 36 min read
Agentic RL: Cutting‑Edge Techniques from GLM‑5.2 and Qwen
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 2, 2026 · Artificial Intelligence

How 5.5K Data Beats Gemini: Beihang’s Concise Symbolic Bridge for Plane Geometry Reasoning

The paper introduces CDL Solver, a two‑stage decoupled framework that translates plane‑geometry diagrams into a concise symbolic language (CDL), reducing training data by 43× and achieving 85.7% accuracy on FormalGeo—surpassing Gemini 2.5 Pro, GPT‑4o and prior specialized models—while also demonstrating strong out‑of‑domain generalisation.

CVPR 2026Reinforcement Learningconcise description language
0 likes · 9 min read
How 5.5K Data Beats Gemini: Beihang’s Concise Symbolic Bridge for Plane Geometry Reasoning
DeepHub IMBA
DeepHub IMBA
Aug 2, 2026 · Artificial Intelligence

Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K

The article presents a from‑scratch LLM training framework called grpo‑llm, implements GRPO with Trio rollout, FSDP and a C++ reward extension, and conducts a controlled experiment comparing GRPO, PPO and DPO on the GSM8K math‑reasoning benchmark, revealing why DPO outperforms the other two under sparse binary rewards.

DPOGRPOGSM8K
0 likes · 10 min read
Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K
Model Perspective
Model Perspective
Jul 31, 2026 · Artificial Intelligence

Understanding the Post-Training Process in DeepSeek V4‑Flash

DeepSeek released the V4‑Flash model with the same architecture as the preview but a revamped post‑training pipeline—SFT, reinforcement learning with GRPO, and distillation—yielding dramatic benchmark jumps and illustrating how post‑training now defines the model's real‑world capabilities.

DeepSeekGRPOLLM training
0 likes · 11 min read
Understanding the Post-Training Process in DeepSeek V4‑Flash
SuanNi
SuanNi
Jul 30, 2026 · Artificial Intelligence

3B Activation Parameters Enable State‑of‑the‑Art Agentic Coding: KAT‑Coder‑V2.5‑Dev Open‑Source Release

KAT‑Coder‑V2.5‑Dev, a 350 B‑parameter MOE model with 3 B activation parameters built on Qwen3.6‑35B‑A3B, achieves top agentic coding performance on PinchBench and near‑top on SWE‑Bench Pro, and the article details its environment construction, data scaling, RL design, and stability improvements.

Data ScalingKAT-CoderReinforcement Learning
0 likes · 12 min read
3B Activation Parameters Enable State‑of‑the‑Art Agentic Coding: KAT‑Coder‑V2.5‑Dev Open‑Source Release
Amap Tech
Amap Tech
Jul 29, 2026 · Artificial Intelligence

ABot-C0: A General‑Purpose Control Intelligence Platform for Quadruped Robots

ABot-C0 presents a unified quadruped control stack that builds 16,074 physically‑validated motion trajectories, achieves 91.02% zero‑shot tracking success and 83.2% full‑terrain success, and demonstrates real‑time deployment at 200 Hz motor control and 50 Hz decision making.

LiDAR perceptionReinforcement LearningSim2Real
0 likes · 11 min read
ABot-C0: A General‑Purpose Control Intelligence Platform for Quadruped Robots
PaperAgent
PaperAgent
Jul 29, 2026 · Artificial Intelligence

How to Build Harness‑Native Agents Using OpenForge RL

OpenForge RL introduces a lightweight proxy and Kubernetes‑based orchestrator to decouple training from inference, enabling the training of 30B‑scale and 8B agents within any harness, while providing an automatic five‑stage task synthesis pipeline and demonstrating state‑of‑the‑art results across Claw, GUI, and Browser benchmarks.

AgentHarnessKubernetes
0 likes · 13 min read
How to Build Harness‑Native Agents Using OpenForge RL
Machine Heart
Machine Heart
Jul 26, 2026 · Artificial Intelligence

Beyond Scaling: How Macaron‑V1 Opens a New Path for Continuous Learning in Open‑Source AI

Macaron‑V1, built on the GLM‑5.2 foundation, demonstrates that post‑training growth via LoRA‑based Mixture‑of‑LoRA, recursive self‑improvement and multi‑agent collaboration can outperform larger static models on benchmarks like UI4A, while its supporting infrastructure (MinT, MindForge, LongStraw) makes million‑parameter reinforcement learning feasible.

AI infrastructureContinuous LearningLoRA
0 likes · 19 min read
Beyond Scaling: How Macaron‑V1 Opens a New Path for Continuous Learning in Open‑Source AI
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 24, 2026 · Artificial Intelligence

China’s 00‑Gen AI Team Unveils 748B LoRA Model Matching Opus 4.8 Performance

Mind Lab’s newly released Macaron‑V1, a 748‑billion‑parameter model built from a GLM‑5.2 base plus four specialized LoRA adapters, achieves benchmark results comparable to Opus 4.8, GPT‑5.5 and Gemini 3.1 Pro, while demonstrating the industry’s shift toward continuous‑learning AI through Mixture‑of‑LoRA architecture and open‑weight deployment.

AI modelBenchmarkContinuous Learning
0 likes · 16 min read
China’s 00‑Gen AI Team Unveils 748B LoRA Model Matching Opus 4.8 Performance
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 24, 2026 · Artificial Intelligence

Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse

The article analyzes why reinforcement learning with verifiable rewards (RLVR) for large models experiences rapid policy‑entropy collapse, breaks the phenomenon down to token‑level entropy changes driven by clipping, advantage, token probability and conditional entropy, and introduces STEER, a token‑wise reweighting scheme that stabilizes entropy and yields consistent performance gains on math and code benchmarks.

RLVRReinforcement LearningSTEER
0 likes · 14 min read
Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 24, 2026 · Artificial Intelligence

UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model

UniNote introduces a unified multimodal embedding model that combines representation learning and ranking optimization in a single forward pass, using a two‑stage SFT‑then‑RL training paradigm and Matryoshka Representation Learning to achieve competitive Item‑to‑Item retrieval performance while reducing latency.

Item2ItemMatryoshka Representation LearningMultimodal Retrieval
0 likes · 12 min read
UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model
Machine Heart
Machine Heart
Jul 24, 2026 · Artificial Intelligence

Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence

Macaron‑V1, an open‑source model built on the GLM‑5.2 base, demonstrates that scaling alone is insufficient by integrating LoRA‑based continuous learning and multi‑agent collaboration, achieving superior benchmark scores, efficient parameter updates, and a novel infrastructure that supports millions of adapters and long‑context reinforcement learning.

Continuous LearningLoRAMacaron-V1
0 likes · 18 min read
Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 22, 2026 · Artificial Intelligence

Renowned AI Scholars from SJTU, CUHK (Shenzhen) and Tencent Hunyuan to Present at MLNLP 2026 Symposium

The MLNLP 2026 online symposium on July 26 will feature leading AI researchers from Shanghai Jiao Tong University, CUHK (Shenzhen) and Tencent Hunyuan presenting talks on lifelong learning, generative model fine‑tuning, and unified multimodal reinforcement learning, with registration now open.

AI ConferenceGenerative ModelsLifelong Learning
0 likes · 12 min read
Renowned AI Scholars from SJTU, CUHK (Shenzhen) and Tencent Hunyuan to Present at MLNLP 2026 Symposium
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 21, 2026 · Artificial Intelligence

MIT Researchers Embed Generalization in Harness: Short-Task Training Unlocks 32× Length Extrapolation

MIT CSAIL’s study shows that training a Recursive Language Model with a harness that keeps each model call locally in-distribution allows the system to extrapolate up to 32-fold longer sequences, achieve superior cross-domain transfer, and outperform transformer baselines despite higher training cost.

HarnessRLMReinforcement Learning
0 likes · 10 min read
MIT Researchers Embed Generalization in Harness: Short-Task Training Unlocks 32× Length Extrapolation
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 20, 2026 · Artificial Intelligence

How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera motion smoothness and geometry reprojection consistency—to improve geometric consistency in video diffusion models without sacrificing pre‑training generalization, achieving smoother camera paths and coherent scene structures even in dynamic scenarios.

4D rewardECCV 2026Reinforcement Learning
0 likes · 9 min read
How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation
Machine Heart
Machine Heart
Jul 20, 2026 · Artificial Intelligence

Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment

In a candid WAIC 2026 interview, reinforcement‑learning pioneer Richard Sutton discusses his new for‑profit Oak Lab, the quest for a 20‑watt trillion‑parameter model, his disappointment with recent AI trends, the notion of a “complete mind,” robot‑kindergarten experiments, and why he believes aligning AI to a single human value system is a dangerous illusion.

AI AlignmentAI safetyEnergy Efficiency
0 likes · 12 min read
Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment
21CTO
21CTO
Jul 18, 2026 · Artificial Intelligence

Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era

In his WAIC keynote, Turing Award laureate Richard Sutton argues that scaling compute and static data does not yield true intelligence, urging a shift toward agents that learn from real‑world interaction and experience, marking the start of an AI "experience era".

AI safetyArtificial IntelligenceExperience Era
0 likes · 12 min read
Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

AI Enters the Experiential Era: WAIC 2026 Shows AI‑Native Learning Labs in Action

At WAIC 2026 the iLoveStudy AI learning agent demonstrated a shift from simply delivering answers to guiding students through interactive, step‑by‑step reasoning, while multimodal digital humans, advanced speech‑enhancement, and a data‑driven reinforcement loop enabled low‑latency, personalized education experiences at scale.

3D avatarAI EducationReinforcement Learning
0 likes · 17 min read
AI Enters the Experiential Era: WAIC 2026 Shows AI‑Native Learning Labs in Action
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera‑motion smoothness and geometry‑reprojection consistency—to eliminate drift and improve structural coherence in video diffusion models without altering their pretrained architecture, achieving state‑of‑the‑art results on static and dynamic benchmarks.

4D rewardECCV 2026Reinforcement Learning
0 likes · 7 min read
VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)
Tencent Advertising Technology
Tencent Advertising Technology
Jul 17, 2026 · Artificial Intelligence

AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)

AdPilot, the first end‑to‑end autonomous advertising agent, reformulates ad delivery as a Markov decision process and combines structured memory, LLM‑enhanced reasoning, and a GRPO‑based reinforcement‑learning engine, while the newly released AdBench benchmark evaluates its superior performance across 38 scenarios and 7,600 instances, outperforming strong baselines by up to 11.76%.

AdBenchAdPilotKDD 2026
0 likes · 16 min read
AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

How Six Robots Built a 3.5‑Meter Great Wall in 15 Hours Using VLA+World Model

Six robots assembled a 3.5 m × 1.5 m × 1.1 m Great Wall model with over 80,000 sub‑centimeter parts in 15 hours, showcasing the DM0.5 foundation model and DW0.5 world‑model loop (VLA+WM) that achieve sub‑millimeter precision, strong generalization, and state‑of‑the‑art benchmark scores.

BenchmarkDM0.5DW0.5
0 likes · 11 min read
How Six Robots Built a 3.5‑Meter Great Wall in 15 Hours Using VLA+World Model
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 16, 2026 · Artificial Intelligence

Inkling – A 975‑Billion‑Parameter Open‑Weight Multimodal Model from Thinking Machines

Thinking Machines Lab unveiled Inkling, a 975‑billion‑parameter open‑weight multimodal model featuring a hybrid‑expert Transformer, 1‑million‑token context, and extensive benchmark results, alongside the lighter Inkling‑Small, with detailed architecture, training methodology, reinforcement‑learning enhancements, and practical examples of web‑app generation and tool‑calling.

BenchmarkInklingMixture of Experts
0 likes · 15 min read
Inkling – A 975‑Billion‑Parameter Open‑Weight Multimodal Model from Thinking Machines
Didi Tech
Didi Tech
Jul 16, 2026 · Artificial Intelligence

FAST: A Parallel Framework that Accelerates Reinforcement Learning for Autonomous Driving

The paper introduces FAST, a parallel reinforcement‑learning sampling framework for autonomous‑driving that decouples individual episode termination from global resets via Dynamic Parallel Sampling Alignment and Scaled Mask‑Padding Optimization, achieving up to 9.08× higher sampling throughput and up to 2× faster training while preserving zero policy loss.

Autonomous DrivingReinforcement LearningSample Efficiency
0 likes · 15 min read
FAST: A Parallel Framework that Accelerates Reinforcement Learning for Autonomous Driving
PaperAgent
PaperAgent
Jul 16, 2026 · Artificial Intelligence

Best Practices for Training Long‑Horizon Autonomous Agents

This article surveys recent Agentic RL research, extracts practical design principles, and details concrete implementations such as ToRL, AgentGym‑RL, Agent‑R1, StarPO, and AutoForge, highlighting reward design, environment interfaces, scaling strategies, and stability diagnostics for long‑horizon autonomous agents.

Agentic RLReinforcement Learninglong-horizon agents
0 likes · 14 min read
Best Practices for Training Long‑Horizon Autonomous Agents
21CTO
21CTO
Jul 15, 2026 · Artificial Intelligence

Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents

Veteran reinforcement‑learning pioneer Richard Sutton announces the creation of Oak Lab, outlining a new Options‑and‑Knowledge architecture that aims to produce autonomous agents capable of continual, real‑time learning, and critiquing the current large‑language‑model paradigm as a dead‑end for true AI.

Oak LabOptions and Knowledge architectureReinforcement Learning
0 likes · 11 min read
Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents
Machine Heart
Machine Heart
Jul 15, 2026 · Artificial Intelligence

Why Does Large‑Model RL Training Narrow? Entropy Insights from ACL Paper

Large‑model reinforcement learning with verifiable rewards often suffers entropy collapse, causing exploration to shrink; this article dissects the phenomenon at the token level, identifies four influencing factors, critiques existing entropy interventions, and introduces STEER—a token‑wise reweighting scheme that stabilizes entropy dynamics and yields consistent gains on math reasoning and coding benchmarks.

LLMRLVRReinforcement Learning
0 likes · 12 min read
Why Does Large‑Model RL Training Narrow? Entropy Insights from ACL Paper
Machine Heart
Machine Heart
Jul 15, 2026 · Artificial Intelligence

How SEAGym Enables Self‑Evolving LLM Agents and Solves Evaluation Challenges

The article introduces SEAGym, a benchmark that treats self‑evolving LLM agents as reinforcement‑learning processes, evaluates their harness updates across multiple dimensions, and reveals how batch size, training source diversity, and backend model affect performance, stability, and cost.

BenchmarkHarness EngineeringLLM
0 likes · 15 min read
How SEAGym Enables Self‑Evolving LLM Agents and Solves Evaluation Challenges
Machine Heart
Machine Heart
Jul 14, 2026 · Artificial Intelligence

Why RL Pioneer Richard Sutton Is Launching a Startup to Escape the Large‑Model Paradigm

At nearly 70, Turing‑award‑winning reinforcement‑learning pioneer Richard Sutton co‑founded Oak Lab to develop the OaK architecture, aiming for AI agents that learn continuously from experience, critiquing the static data reliance of current large language models and targeting low‑energy trillion‑parameter AGI.

AGIAIContinuous Learning
0 likes · 6 min read
Why RL Pioneer Richard Sutton Is Launching a Startup to Escape the Large‑Model Paradigm
DaTaobao Tech
DaTaobao Tech
Jul 13, 2026 · Artificial Intelligence

Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning

The article details how Taobao Live upgraded its static workflow to a low‑latency Agentic architecture, applied AgentTuning distillation and RLVR to curb hallucinations, and introduced a Multi‑Agent RL framework that separates tool‑calling and reply generation, achieving significant gains in factual correctness, helpfulness, and overall performance.

Agentic RLLLMReinforcement Learning
0 likes · 23 min read
Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 13, 2026 · Artificial Intelligence

How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper

The article reviews a Nanjing University position paper that argues world‑model evaluation for embodied decision‑making should prioritize prediction of action consequences, strategy assessment, and planning support, while treating visual realism and semantic alignment as secondary diagnostics.

Embodied AIReinforcement LearningWorld Models
0 likes · 14 min read
How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper
Data Party THU
Data Party THU
Jul 12, 2026 · Artificial Intelligence

How Counterfactual Policy Optimization Boosts Visual Fidelity in Multimodal Reasoning (ICML 2026)

The paper introduces Counterfactual Policy Optimization (CFPO), a training‑time framework that inserts causal consistency constraints into multimodal reinforcement learning, forcing vision‑language models to rely on essential visual evidence and achieving consistent accuracy gains across real‑world and math‑centric benchmarks.

ICML2026MultimodalReinforcement Learning
0 likes · 19 min read
How Counterfactual Policy Optimization Boosts Visual Fidelity in Multimodal Reasoning (ICML 2026)
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper

The paper surveys the expanding definition of world models across robotics, autonomous driving, and video generation, identifies six capability claims, critiques current perception‑focused metrics, and proposes a decision‑centric 7‑level evaluation ladder and concrete protocols to assess action consequences, strategy ranking, and planning utility.

Embodied AIReinforcement LearningWorld Models
0 likes · 13 min read
How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL

The paper by researchers from Beihang, Peking University and Meituan proposes PIRL, a new RL‑post‑training perspective that treats policy improvement as the optimization objective, and PIPO, a plug‑and‑play framework that adds a verification loop to amplify beneficial updates and suppress harmful ones, demonstrating consistent gains across math reasoning, code and tool‑use tasks.

Importance SamplingPIPOPIRL
0 likes · 9 min read
Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL
DataFunSummit
DataFunSummit
Jul 11, 2026 · Artificial Intelligence

Why Diversity Beats Data Scale: Insights from MiniMax & Fudan’s DIVE Paper

The DIVE study shows that expanding the diversity of tool pools and task structures, rather than merely increasing the amount of homogeneous training data, dramatically improves LLM agents' ability to generalize to unseen tools, as demonstrated by a 12k‑vs‑48k experiment and reinforced by a four‑stage synthesis pipeline and RL fine‑tuning.

AI agentsDIVELLM training
0 likes · 14 min read
Why Diversity Beats Data Scale: Insights from MiniMax & Fudan’s DIVE Paper
Kuaishou Tech
Kuaishou Tech
Jul 10, 2026 · Artificial Intelligence

KAT-Coder-Pro V2.5 Launch: Boosting Agentic Coding from Code Writing to Full Engineering

KAT-Coder-Pro V2.5 introduces a flagship Agentic coding model that expands long‑chain engineering ability, adds a universal Agentic framework, and leverages a large‑scale RL pipeline, achieving top scores on SWE‑Bench Pro, PinchBench and internal benchmarks while enabling developers to hand over complete issues without manual decomposition.

AutoBuilderBenchmarkKAT-Coder-Pro
0 likes · 11 min read
KAT-Coder-Pro V2.5 Launch: Boosting Agentic Coding from Code Writing to Full Engineering
DataFunSummit
DataFunSummit
Jul 9, 2026 · Artificial Intelligence

Token-Level Credit Assignment Outperforms Broadcast GRPO in LLM Math Reasoning

The paper identifies the broadcast‑style credit assignment of GRPO as a bottleneck for RL‑LLM math reasoning, proposes the Outcome‑Grounded Advantage Reshaping (OAR) framework with token‑importance estimation, and demonstrates that its two variants, OAR‑P and OAR‑G, consistently improve accuracy, training efficiency, and stability across multiple math benchmarks.

GRPOLLMOAR
0 likes · 15 min read
Token-Level Credit Assignment Outperforms Broadcast GRPO in LLM Math Reasoning
Machine Heart
Machine Heart
Jul 9, 2026 · Artificial Intelligence

Can Your Self‑Distillation Model Do Without Reference Solutions? Introducing d‑OPSD for Diffusion LLMs

The paper presents d‑OPSD, the first on‑policy self‑distillation framework for diffusion large language models that eliminates reference solutions and extra teacher models, using only one‑tenth of RL steps while achieving equal or superior reasoning performance and markedly higher training efficiency, as demonstrated on multiple math‑reasoning benchmarks.

Diffusion Language ModelsOPSDReinforcement Learning
0 likes · 7 min read
Can Your Self‑Distillation Model Do Without Reference Solutions? Introducing d‑OPSD for Diffusion LLMs
Machine Heart
Machine Heart
Jul 9, 2026 · Artificial Intelligence

LingBot-Video: The Open‑Source Video Foundation Model Tailored for Robots

LingBot-Video is an open‑source video‑generation foundation model built for embodied AI, combining a sparse‑Mixture‑of‑Experts architecture, a multi‑stage data curriculum and six‑dimensional reward learning to achieve physically consistent video synthesis that outperforms existing open‑source baselines in both visual quality and robotic relevance.

Embodied AIMixture of ExpertsMultimodal Model
0 likes · 22 min read
LingBot-Video: The Open‑Source Video Foundation Model Tailored for Robots