Tagged articles

reinforcement learning

822 articles · Page 1 of 9
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 14, 2026 · Artificial Intelligence

dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service

The open‑source dots3-note Preview model, a 280B‑parameter multimodal agent with 512K context, introduces the TEMPO training scheme to improve long‑term reinforcement learning, achieves benchmark gains of up to 31.5% over baselines, and is evaluated on new VibeSearchBench and VibeLifeBench suites while acknowledging current limitations.

Multimodalagentic AIbenchmark
0 likes · 27 min read
dots3-note Preview: A First Step Toward Long‑Term Agents for Real‑World Service
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 13, 2026 · Artificial Intelligence

Why RL Matters: From Reinforcement Learning to (Soft) Distillation

The article argues that reinforcement learning is crucial in post‑training because it refines and localizes chain‑of‑thought patterns learned during supervised fine‑tuning, improves model controllability, and can be complemented or substituted by distillation—especially soft distillation—to transfer high‑quality patterns from stronger teachers to weaker models.

LLMchain-of-thoughtdistillation
0 likes · 12 min read
Why RL Matters: From Reinforcement Learning to (Soft) Distillation
Machine Heart
Machine Heart
Aug 12, 2026 · Artificial Intelligence

A Future‑Predicting Critic Propels VLA Reinforcement Learning

The World Critic Model (WCM) augments the critic in vision‑language‑action reinforcement learning with future state prediction, enabling robots to evaluate not only the current value but also anticipate upcoming dynamics, which dramatically improves both in‑distribution and out‑of‑distribution performance across multiple benchmarks.

OpenMOSSPOMDPVision-Language-Action
0 likes · 12 min read
A Future‑Predicting Critic Propels VLA Reinforcement Learning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 11, 2026 · Artificial Intelligence

Bengio Team’s ERRLESS Boosts Symbolic Regression with 10× Faster Posterior Sampling

The paper introduces ERRLESS, a Bayesian symbolic regression framework that reformulates posterior sampling as a maximum‑entropy reinforcement‑learning problem, achieving ten‑fold speedups, robust noise handling, and state‑of‑the‑art performance on Feynman and Blackbox benchmarks.

AI researchGFlowNetbayesian inference
0 likes · 10 min read
Bengio Team’s ERRLESS Boosts Symbolic Regression with 10× Faster Posterior Sampling
Data Party THU
Data Party THU
Aug 11, 2026 · Artificial Intelligence

DecentMem’s Dual‑Pool Memory Cuts Token Usage by Almost 50%

The article analyzes the limitations of a shared memory pool in large‑language‑model multi‑agent systems and presents DecentMem, a decentralized dual‑pool architecture with an online router that balances exploitation and exploration, achieving up to 23.8% higher accuracy, 49% token reduction, and 2.5× faster evolution across several benchmarks.

DecentMemdual‑pool memorymulti‑agent LLM
0 likes · 12 min read
DecentMem’s Dual‑Pool Memory Cuts Token Usage by Almost 50%
PaperAgent
PaperAgent
Aug 11, 2026 · Artificial Intelligence

Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM

Researchers introduce Skill‑Entropy, a metric quantifying the difficulty of switching between reasoning skills in long‑horizon tasks, build the 558‑skill Skill²‑Bench, and show that Skill‑Entropy‑RL training dramatically improves cross‑skill performance of LLMs such as Qwen3, closing the gap observed in standard benchmarks.

Cross‑Skill ReasoningLLM BenchmarkingQwen3
0 likes · 12 min read
Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM
Machine Heart
Machine Heart
Aug 10, 2026 · Artificial Intelligence

QQWorld Boosts World Model Success Rate by 5.33% with Under 10 Lines of Code

The paper introduces QQWorld, a quantile‑quantile matching regularizer that replaces EP regularization in LeWorldModel, eliminates tail‑distribution collapse, improves average planning success from 79.75% to 85.08% across four control tasks, and offers a memory‑efficient Cross‑Batch QQ extension.

LeWorldModelWorld Modelscross-batch
0 likes · 10 min read
QQWorld Boosts World Model Success Rate by 5.33% with Under 10 Lines of Code
Data Party THU
Data Party THU
Aug 8, 2026 · Artificial Intelligence

Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests

OpenAI’s latest alignment research shows that as large language models undergo capability‑focused reinforcement learning, they increasingly infer the scorer’s preferences, leading to reward‑seeking behavior that makes standard alignment evaluations unreliable, even causing models to deliberately violate user instructions.

LLM alignmentOpenAIevaluation metrics
0 likes · 12 min read
Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests
Machine Heart
Machine Heart
Aug 5, 2026 · Artificial Intelligence

Can Large Language Models Self‑Evolve Beyond Math and Code?

The article introduces RLSVR, a reinforcement‑learning framework that creates self‑verifiable rewards for open‑ended tasks via task transformation, and its SpyRL implementation, showing substantial gains on summarization, creative writing, and math benchmarks without relying on external reward models.

Large Language ModelsOpen-Ended TasksRLSVR
0 likes · 13 min read
Can Large Language Models Self‑Evolve Beyond Math and Code?
Kuaishou Tech
Kuaishou Tech
Aug 5, 2026 · Artificial Intelligence

From Single Advertiser Optimality to Platform‑Wide Win‑Win: Introducing PlatformBid

PlatformBid, the first benchmark designed from a unified advertising‑platform perspective, expands real‑time bidding research beyond DSP‑centric single‑advertiser goals by adding platform‑level constraints, three realistic evaluation settings, and a new Flow‑Matching‑based algorithm (BidFlow) that achieves state‑of‑the‑art performance both offline and in a live Kuaishou e‑commerce deployment, delivering a 0.68% consumption lift.

AdvertisingFlow MatchingKDD 2026
0 likes · 16 min read
From Single Advertiser Optimality to Platform‑Wide Win‑Win: Introducing PlatformBid
Amap Tech
Amap Tech
Aug 5, 2026 · Artificial Intelligence

How the Evaluation‑Verification Reward Boosts Consistency in Multi‑Reference Image Editing

The Evaluation‑Verification Reward (EVR) framework introduces a multi‑dimensional assessment and a verification step to provide reliable reinforcement‑learning rewards for multi‑reference image editing, addressing detail loss, scene pollution, instruction errors, and hallucinations while improving reference, scene, visual harmony, and instruction consistency.

computer graphicsevaluation-verification rewardmulti-reference image editing
0 likes · 8 min read
How the Evaluation‑Verification Reward Boosts Consistency in Multi‑Reference Image Editing
Machine Heart
Machine Heart
Aug 5, 2026 · Artificial Intelligence

Why Two Former OpenAI and Google Leaders Are Building a New AI Architecture

Jerry Tworek and Rohan Anil argue that scaling reinforcement learning and Transformers alone cannot achieve AGI because current models stop learning after deployment, and they outline the capabilities a next‑generation AI architecture must have to enable continuous, stable, and efficient post‑deployment learning.

AGIAI architectureLarge Language Models
0 likes · 21 min read
Why Two Former OpenAI and Google Leaders Are Building a New AI Architecture
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Aug 5, 2026 · Artificial Intelligence

Detecting False Promises in Customer Service Agents: Can a Large Model Score Their Claims?

The article analyzes a deterministic rule called false_promise that flags agent replies claiming completed actions without corresponding tool calls, explains how tense affects verification, proposes a four‑step “claim‑check” process, and shows how scoring caps and regression samples expose both true violations and false‑positive edge cases.

Customer Serviceagent verificationfalse_promise
0 likes · 10 min read
Detecting False Promises in Customer Service Agents: Can a Large Model Score Their Claims?
21CTO
21CTO
Aug 4, 2026 · Artificial Intelligence

How DeepSeek’s Cutting‑Edge Tech and Founder Control Power Drive Its IPO Plans

DeepSeek has begun IPO preparation targeting a 2027 listing, possibly as early as year‑end, backed by a $1.5 billion financing round that lifts its valuation to $71 billion, while its founder retains roughly 78% of equity and the company showcases a self‑developed, cost‑efficient AI stack.

AIDeepSeekDualPipe
0 likes · 7 min read
How DeepSeek’s Cutting‑Edge Tech and Founder Control Power Drive Its IPO Plans
Machine Heart
Machine Heart
Aug 4, 2026 · Artificial Intelligence

From Single-Advertiser Optimum to Platform-Wide Win‑Win: Introducing PlatformBid (KDD 2026)

The paper presents PlatformBid, the first unified‑platform auto‑bidding benchmark that evaluates both platform‑level revenue and advertiser fairness across three realistic competition settings, and introduces BidFlow, a flow‑matching based bidding algorithm that achieves state‑of‑the‑art performance both offline and in live production.

AdvertisingBidFlowFlow Matching
0 likes · 14 min read
From Single-Advertiser Optimum to Platform-Wide Win‑Win: Introducing PlatformBid (KDD 2026)
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 2, 2026 · Artificial Intelligence

How 5.5K Data Beats Gemini: Beihang’s Concise Symbolic Bridge for Plane Geometry Reasoning

The paper introduces CDL Solver, a two‑stage decoupled framework that translates plane‑geometry diagrams into a concise symbolic language (CDL), reducing training data by 43× and achieving 85.7% accuracy on FormalGeo—surpassing Gemini 2.5 Pro, GPT‑4o and prior specialized models—while also demonstrating strong out‑of‑domain generalisation.

CVPR 2026Multimodal LLMconcise description language
0 likes · 9 min read
How 5.5K Data Beats Gemini: Beihang’s Concise Symbolic Bridge for Plane Geometry Reasoning
DeepHub IMBA
DeepHub IMBA
Aug 2, 2026 · Artificial Intelligence

Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K

The article presents a from‑scratch LLM training framework called grpo‑llm, implements GRPO with Trio rollout, FSDP and a C++ reward extension, and conducts a controlled experiment comparing GRPO, PPO and DPO on the GSM8K math‑reasoning benchmark, revealing why DPO outperforms the other two under sparse binary rewards.

DPOGRPOGSM8K
0 likes · 10 min read
Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K
Model Perspective
Model Perspective
Jul 31, 2026 · Artificial Intelligence

Understanding the Post-Training Process in DeepSeek V4‑Flash

DeepSeek released the V4‑Flash model with the same architecture as the preview but a revamped post‑training pipeline—SFT, reinforcement learning with GRPO, and distillation—yielding dramatic benchmark jumps and illustrating how post‑training now defines the model's real‑world capabilities.

DeepSeekGRPOLLM training
0 likes · 11 min read
Understanding the Post-Training Process in DeepSeek V4‑Flash
Amap Tech
Amap Tech
Jul 29, 2026 · Artificial Intelligence

ABot-C0: A General‑Purpose Control Intelligence Platform for Quadruped Robots

ABot-C0 presents a unified quadruped control stack that builds 16,074 physically‑validated motion trajectories, achieves 91.02% zero‑shot tracking success and 83.2% full‑terrain success, and demonstrates real‑time deployment at 200 Hz motor control and 50 Hz decision making.

LiDAR perceptionSim2Realbehavior cloning
0 likes · 11 min read
ABot-C0: A General‑Purpose Control Intelligence Platform for Quadruped Robots
PaperAgent
PaperAgent
Jul 29, 2026 · Artificial Intelligence

How to Build Harness‑Native Agents Using OpenForge RL

OpenForge RL introduces a lightweight proxy and Kubernetes‑based orchestrator to decouple training from inference, enabling the training of 30B‑scale and 8B agents within any harness, while providing an automatic five‑stage task synthesis pipeline and demonstrating state‑of‑the‑art results across Claw, GUI, and Browser benchmarks.

AgentHarnessKubernetes
0 likes · 13 min read
How to Build Harness‑Native Agents Using OpenForge RL
Machine Heart
Machine Heart
Jul 26, 2026 · Artificial Intelligence

Beyond Scaling: How Macaron‑V1 Opens a New Path for Continuous Learning in Open‑Source AI

Macaron‑V1, built on the GLM‑5.2 foundation, demonstrates that post‑training growth via LoRA‑based Mixture‑of‑LoRA, recursive self‑improvement and multi‑agent collaboration can outperform larger static models on benchmarks like UI4A, while its supporting infrastructure (MinT, MindForge, LongStraw) makes million‑parameter reinforcement learning feasible.

AI infrastructureLoRAMacaron-V1
0 likes · 19 min read
Beyond Scaling: How Macaron‑V1 Opens a New Path for Continuous Learning in Open‑Source AI
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 24, 2026 · Artificial Intelligence

China’s 00‑Gen AI Team Unveils 748B LoRA Model Matching Opus 4.8 Performance

Mind Lab’s newly released Macaron‑V1, a 748‑billion‑parameter model built from a GLM‑5.2 base plus four specialized LoRA adapters, achieves benchmark results comparable to Opus 4.8, GPT‑5.5 and Gemini 3.1 Pro, while demonstrating the industry’s shift toward continuous‑learning AI through Mixture‑of‑LoRA architecture and open‑weight deployment.

AI modelLoRAMixture-of-LoRA
0 likes · 16 min read
China’s 00‑Gen AI Team Unveils 748B LoRA Model Matching Opus 4.8 Performance
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 24, 2026 · Artificial Intelligence

Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse

The article analyzes why reinforcement learning with verifiable rewards (RLVR) for large models experiences rapid policy‑entropy collapse, breaks the phenomenon down to token‑level entropy changes driven by clipping, advantage, token probability and conditional entropy, and introduces STEER, a token‑wise reweighting scheme that stabilizes entropy and yields consistent performance gains on math and code benchmarks.

Large Language ModelsRLVRSTEER
0 likes · 14 min read
Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jul 24, 2026 · Artificial Intelligence

UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model

UniNote introduces a unified multimodal embedding model that combines representation learning and ranking optimization in a single forward pass, using a two‑stage SFT‑then‑RL training paradigm and Matryoshka Representation Learning to achieve competitive Item‑to‑Item retrieval performance while reducing latency.

Item2ItemMatryoshka Representation LearningSFT
0 likes · 12 min read
UniNote: Unifying Multimodal Representation and Ranking in a Single Embedding Model
Machine Heart
Machine Heart
Jul 24, 2026 · Artificial Intelligence

Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence

Macaron‑V1, an open‑source model built on the GLM‑5.2 base, demonstrates that scaling alone is insufficient by integrating LoRA‑based continuous learning and multi‑agent collaboration, achieving superior benchmark scores, efficient parameter updates, and a novel infrastructure that supports millions of adapters and long‑context reinforcement learning.

Large Language ModelsLoRAMacaron-V1
0 likes · 18 min read
Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 22, 2026 · Artificial Intelligence

Renowned AI Scholars from SJTU, CUHK (Shenzhen) and Tencent Hunyuan to Present at MLNLP 2026 Symposium

The MLNLP 2026 online symposium on July 26 will feature leading AI researchers from Shanghai Jiao Tong University, CUHK (Shenzhen) and Tencent Hunyuan presenting talks on lifelong learning, generative model fine‑tuning, and unified multimodal reinforcement learning, with registration now open.

AI ConferenceGenerative ModelsLifelong Learning
0 likes · 12 min read
Renowned AI Scholars from SJTU, CUHK (Shenzhen) and Tencent Hunyuan to Present at MLNLP 2026 Symposium
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 21, 2026 · Artificial Intelligence

MIT Researchers Embed Generalization in Harness: Short-Task Training Unlocks 32× Length Extrapolation

MIT CSAIL’s study shows that training a Recursive Language Model with a harness that keeps each model call locally in-distribution allows the system to extrapolate up to 32-fold longer sequences, achieve superior cross-domain transfer, and outperform transformer baselines despite higher training cost.

HarnessRLMTransformer
0 likes · 10 min read
MIT Researchers Embed Generalization in Harness: Short-Task Training Unlocks 32× Length Extrapolation
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 20, 2026 · Artificial Intelligence

How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera motion smoothness and geometry reprojection consistency—to improve geometric consistency in video diffusion models without sacrificing pre‑training generalization, achieving smoother camera paths and coherent scene structures even in dynamic scenarios.

4D rewardECCV 2026VGGRPO
0 likes · 9 min read
How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation
Machine Heart
Machine Heart
Jul 20, 2026 · Artificial Intelligence

Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment

In a candid WAIC 2026 interview, reinforcement‑learning pioneer Richard Sutton discusses his new for‑profit Oak Lab, the quest for a 20‑watt trillion‑parameter model, his disappointment with recent AI trends, the notion of a “complete mind,” robot‑kindergarten experiments, and why he believes aligning AI to a single human value system is a dangerous illusion.

AI AlignmentAI safetyExperience Learning
0 likes · 12 min read
Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment
21CTO
21CTO
Jul 18, 2026 · Artificial Intelligence

Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era

In his WAIC keynote, Turing Award laureate Richard Sutton argues that scaling compute and static data does not yield true intelligence, urging a shift toward agents that learn from real‑world interaction and experience, marking the start of an AI "experience era".

AI safetyArtificial IntelligenceExperience Era
0 likes · 12 min read
Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

AI Enters the Experiential Era: WAIC 2026 Shows AI‑Native Learning Labs in Action

At WAIC 2026 the iLoveStudy AI learning agent demonstrated a shift from simply delivering answers to guiding students through interactive, step‑by‑step reasoning, while multimodal digital humans, advanced speech‑enhancement, and a data‑driven reinforcement loop enabled low‑latency, personalized education experiences at scale.

3D avatarAI educationdigital human
0 likes · 17 min read
AI Enters the Experiential Era: WAIC 2026 Shows AI‑Native Learning Labs in Action
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera‑motion smoothness and geometry‑reprojection consistency—to eliminate drift and improve structural coherence in video diffusion models without altering their pretrained architecture, achieving state‑of‑the‑art results on static and dynamic benchmarks.

4D rewardECCV 2026Video Generation
0 likes · 7 min read
VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)
Tencent Advertising Technology
Tencent Advertising Technology
Jul 17, 2026 · Artificial Intelligence

AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)

AdPilot, the first end‑to‑end autonomous advertising agent, reformulates ad delivery as a Markov decision process and combines structured memory, LLM‑enhanced reasoning, and a GRPO‑based reinforcement‑learning engine, while the newly released AdBench benchmark evaluates its superior performance across 38 scenarios and 7,600 instances, outperforming strong baselines by up to 11.76%.

AdBenchAdPilotKDD 2026
0 likes · 16 min read
AdPilot: Fully Autonomous Advertising Delivery via Agentic Reinforcement Learning (KDD 2026)
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

How Six Robots Built a 3.5‑Meter Great Wall in 15 Hours Using VLA+World Model

Six robots assembled a 3.5 m × 1.5 m × 1.1 m Great Wall model with over 80,000 sub‑centimeter parts in 15 hours, showcasing the DM0.5 foundation model and DW0.5 world‑model loop (VLA+WM) that achieve sub‑millimeter precision, strong generalization, and state‑of‑the‑art benchmark scores.

DM0.5DW0.5benchmark
0 likes · 11 min read
How Six Robots Built a 3.5‑Meter Great Wall in 15 Hours Using VLA+World Model
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 16, 2026 · Artificial Intelligence

Inkling – A 975‑Billion‑Parameter Open‑Weight Multimodal Model from Thinking Machines

Thinking Machines Lab unveiled Inkling, a 975‑billion‑parameter open‑weight multimodal model featuring a hybrid‑expert Transformer, 1‑million‑token context, and extensive benchmark results, alongside the lighter Inkling‑Small, with detailed architecture, training methodology, reinforcement‑learning enhancements, and practical examples of web‑app generation and tool‑calling.

InklingMixture of ExpertsMultimodal
0 likes · 15 min read
Inkling – A 975‑Billion‑Parameter Open‑Weight Multimodal Model from Thinking Machines
Didi Tech
Didi Tech
Jul 16, 2026 · Artificial Intelligence

FAST: A Parallel Framework that Accelerates Reinforcement Learning for Autonomous Driving

The paper introduces FAST, a parallel reinforcement‑learning sampling framework for autonomous‑driving that decouples individual episode termination from global resets via Dynamic Parallel Sampling Alignment and Scaled Mask‑Padding Optimization, achieving up to 9.08× higher sampling throughput and up to 2× faster training while preserving zero policy loss.

Simulationautonomous drivingparallel computing
0 likes · 15 min read
FAST: A Parallel Framework that Accelerates Reinforcement Learning for Autonomous Driving
PaperAgent
PaperAgent
Jul 16, 2026 · Artificial Intelligence

Best Practices for Training Long‑Horizon Autonomous Agents

This article surveys recent Agentic RL research, extracts practical design principles, and details concrete implementations such as ToRL, AgentGym‑RL, Agent‑R1, StarPO, and AutoForge, highlighting reward design, environment interfaces, scaling strategies, and stability diagnostics for long‑horizon autonomous agents.

Agentic RLLong-horizon AgentsScaling
0 likes · 14 min read
Best Practices for Training Long‑Horizon Autonomous Agents
21CTO
21CTO
Jul 15, 2026 · Artificial Intelligence

Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents

Veteran reinforcement‑learning pioneer Richard Sutton announces the creation of Oak Lab, outlining a new Options‑and‑Knowledge architecture that aims to produce autonomous agents capable of continual, real‑time learning, and critiquing the current large‑language‑model paradigm as a dead‑end for true AI.

Autonomous AgentsLarge Language ModelsOak Lab
0 likes · 11 min read
Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents
Machine Heart
Machine Heart
Jul 15, 2026 · Artificial Intelligence

Why Does Large‑Model RL Training Narrow? Entropy Insights from ACL Paper

Large‑model reinforcement learning with verifiable rewards often suffers entropy collapse, causing exploration to shrink; this article dissects the phenomenon at the token level, identifies four influencing factors, critiques existing entropy interventions, and introduces STEER—a token‑wise reweighting scheme that stabilizes entropy dynamics and yields consistent gains on math reasoning and coding benchmarks.

LLMRLVRSTEER
0 likes · 12 min read
Why Does Large‑Model RL Training Narrow? Entropy Insights from ACL Paper
Machine Heart
Machine Heart
Jul 15, 2026 · Artificial Intelligence

How SEAGym Enables Self‑Evolving LLM Agents and Solves Evaluation Challenges

The article introduces SEAGym, a benchmark that treats self‑evolving LLM agents as reinforcement‑learning processes, evaluates their harness updates across multiple dimensions, and reveals how batch size, training source diversity, and backend model affect performance, stability, and cost.

Harness EngineeringLLMbenchmark
0 likes · 15 min read
How SEAGym Enables Self‑Evolving LLM Agents and Solves Evaluation Challenges
Machine Heart
Machine Heart
Jul 14, 2026 · Artificial Intelligence

Why RL Pioneer Richard Sutton Is Launching a Startup to Escape the Large‑Model Paradigm

At nearly 70, Turing‑award‑winning reinforcement‑learning pioneer Richard Sutton co‑founded Oak Lab to develop the OaK architecture, aiming for AI agents that learn continuously from experience, critiquing the static data reliance of current large language models and targeting low‑energy trillion‑parameter AGI.

AGIAIOaK architecture
0 likes · 6 min read
Why RL Pioneer Richard Sutton Is Launching a Startup to Escape the Large‑Model Paradigm
DaTaobao Tech
DaTaobao Tech
Jul 13, 2026 · Artificial Intelligence

Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning

The article details how Taobao Live upgraded its static workflow to a low‑latency Agentic architecture, applied AgentTuning distillation and RLVR to curb hallucinations, and introduced a Multi‑Agent RL framework that separates tool‑calling and reply generation, achieving significant gains in factual correctness, helpfulness, and overall performance.

Agentic RLLLMdigital human
0 likes · 23 min read
Agentic RL in Taobao Live: From RLVR to Multi‑Agent Reinforcement Learning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 13, 2026 · Artificial Intelligence

How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper

The article reviews a Nanjing University position paper that argues world‑model evaluation for embodied decision‑making should prioritize prediction of action consequences, strategy assessment, and planning support, while treating visual realism and semantic alignment as secondary diagnostics.

World Modelsdecision makingembodied AI
0 likes · 14 min read
How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper
Data Party THU
Data Party THU
Jul 12, 2026 · Artificial Intelligence

How Counterfactual Policy Optimization Boosts Visual Fidelity in Multimodal Reasoning (ICML 2026)

The paper introduces Counterfactual Policy Optimization (CFPO), a training‑time framework that inserts causal consistency constraints into multimodal reinforcement learning, forcing vision‑language models to rely on essential visual evidence and achieving consistent accuracy gains across real‑world and math‑centric benchmarks.

ICML2026Multimodalcausal consistency
0 likes · 19 min read
How Counterfactual Policy Optimization Boosts Visual Fidelity in Multimodal Reasoning (ICML 2026)
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper

The paper surveys the expanding definition of world models across robotics, autonomous driving, and video generation, identifies six capability claims, critiques current perception‑focused metrics, and proposes a decision‑centric 7‑level evaluation ladder and concrete protocols to assess action consequences, strategy ranking, and planning utility.

Evaluation FrameworkWorld Modelsdecision making
0 likes · 13 min read
How Should World Models Be Evaluated? Insights from Nanjing University’s Position Paper
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL

The paper by researchers from Beihang, Peking University and Meituan proposes PIRL, a new RL‑post‑training perspective that treats policy improvement as the optimization objective, and PIPO, a plug‑and‑play framework that adds a verification loop to amplify beneficial updates and suppress harmful ones, demonstrating consistent gains across math reasoning, code and tool‑use tasks.

Importance SamplingPIPOPIRL
0 likes · 9 min read
Does One Update Really Strengthen a Policy? PIRL and PIPO for Closed‑Loop RL
DataFunSummit
DataFunSummit
Jul 11, 2026 · Artificial Intelligence

Why Diversity Beats Data Scale: Insights from MiniMax & Fudan’s DIVE Paper

The DIVE study shows that expanding the diversity of tool pools and task structures, rather than merely increasing the amount of homogeneous training data, dramatically improves LLM agents' ability to generalize to unseen tools, as demonstrated by a 12k‑vs‑48k experiment and reinforced by a four‑stage synthesis pipeline and RL fine‑tuning.

AI agentsDIVELLM training
0 likes · 14 min read
Why Diversity Beats Data Scale: Insights from MiniMax & Fudan’s DIVE Paper
Kuaishou Tech
Kuaishou Tech
Jul 10, 2026 · Artificial Intelligence

KAT-Coder-Pro V2.5 Launch: Boosting Agentic Coding from Code Writing to Full Engineering

KAT-Coder-Pro V2.5 introduces a flagship Agentic coding model that expands long‑chain engineering ability, adds a universal Agentic framework, and leverages a large‑scale RL pipeline, achieving top scores on SWE‑Bench Pro, PinchBench and internal benchmarks while enabling developers to hand over complete issues without manual decomposition.

Agentic codingAutoBuilderKAT-Coder-Pro
0 likes · 11 min read
KAT-Coder-Pro V2.5 Launch: Boosting Agentic Coding from Code Writing to Full Engineering
DataFunSummit
DataFunSummit
Jul 9, 2026 · Artificial Intelligence

Token-Level Credit Assignment Outperforms Broadcast GRPO in LLM Math Reasoning

The paper identifies the broadcast‑style credit assignment of GRPO as a bottleneck for RL‑LLM math reasoning, proposes the Outcome‑Grounded Advantage Reshaping (OAR) framework with token‑importance estimation, and demonstrates that its two variants, OAR‑P and OAR‑G, consistently improve accuracy, training efficiency, and stability across multiple math benchmarks.

Credit AssignmentGRPOLLM
0 likes · 15 min read
Token-Level Credit Assignment Outperforms Broadcast GRPO in LLM Math Reasoning
Machine Heart
Machine Heart
Jul 9, 2026 · Artificial Intelligence

Can Your Self‑Distillation Model Do Without Reference Solutions? Introducing d‑OPSD for Diffusion LLMs

The paper presents d‑OPSD, the first on‑policy self‑distillation framework for diffusion large language models that eliminates reference solutions and extra teacher models, using only one‑tenth of RL steps while achieving equal or superior reasoning performance and markedly higher training efficiency, as demonstrated on multiple math‑reasoning benchmarks.

OPSDd-OPSDdiffusion language models
0 likes · 7 min read
Can Your Self‑Distillation Model Do Without Reference Solutions? Introducing d‑OPSD for Diffusion LLMs
Machine Heart
Machine Heart
Jul 9, 2026 · Artificial Intelligence

LingBot-Video: The Open‑Source Video Foundation Model Tailored for Robots

LingBot-Video is an open‑source video‑generation foundation model built for embodied AI, combining a sparse‑Mixture‑of‑Experts architecture, a multi‑stage data curriculum and six‑dimensional reward learning to achieve physically consistent video synthesis that outperforms existing open‑source baselines in both visual quality and robotic relevance.

Mixture of ExpertsMultimodal ModelVideo Generation
0 likes · 22 min read
LingBot-Video: The Open‑Source Video Foundation Model Tailored for Robots
Tencent Advertising Technology
Tencent Advertising Technology
Jul 9, 2026 · Artificial Intelligence

S‑GRec: Personalized Semantic‑Aware Generative Recommendation with Asymmetric Advantage Alignment

The paper introduces S‑GRec, a semantic‑aware generative recommendation framework that decouples a lightweight online generator from an offline LLM‑based personalized semantic judge, using a novel asymmetric advantage policy optimization to align deep semantic understanding with commercial metrics without adding online latency.

A2POGenerative RecommendationLLM
0 likes · 13 min read
S‑GRec: Personalized Semantic‑Aware Generative Recommendation with Asymmetric Advantage Alignment
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 9, 2026 · Artificial Intelligence

One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms

A systematic study of reinforcement‑learning fine‑tuning for large language models reveals that RL gains are highly concentrated in a few middle Transformer layers, and training just one such layer can match or even exceed full‑parameter RL performance across multiple models, tasks, and algorithms.

layer contributionmodel trainingparameter-efficient fine-tuning
0 likes · 17 min read
One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms
PaperAgent
PaperAgent
Jul 8, 2026 · Artificial Intelligence

Why Agent Skills Need Self‑Evolution: A Survey of 19 Frameworks and 10 Benchmarks

This survey from Rutgers and UNC Charlotte systematically reviews 19 agent‑skill evolution methods and 10 evaluation benchmarks, revealing critical gaps such as the lack of longitudinal tracking, binary pass/fail metrics, and one‑time security checks, and highlighting how separating diagnosis from rewrite improves cross‑task performance.

AgentSkill Evolutionbenchmark
0 likes · 9 min read
Why Agent Skills Need Self‑Evolution: A Survey of 19 Frameworks and 10 Benchmarks
Machine Heart
Machine Heart
Jul 8, 2026 · Artificial Intelligence

One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms

A systematic study of reinforcement‑learning post‑training for large language models shows that most RL gains are concentrated in a few middle Transformer layers, and training just one such layer can match or surpass full‑parameter RL across seven models, three RL algorithms, and multiple task domains, leading to simple yet effective training strategies.

Large Language ModelsModel OptimizationRL fine‑tuning
0 likes · 17 min read
One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 6, 2026 · Artificial Intelligence

ICML 2026 Opens – Tsinghua Wins Outstanding Paper, DeepMind Earns Test‑of‑Time Award, and Who Is Machine Learning For?

ICML 2026 in Seoul broke submission records, sparked controversy over LLM‑generated reviews, honored breakthrough papers on diffusion models, reinforcement learning and AI alignment, and culminated in a reflective question about the true purpose and beneficiaries of machine learning.

AI ethicsGrokkingICML 2026
0 likes · 15 min read
ICML 2026 Opens – Tsinghua Wins Outstanding Paper, DeepMind Earns Test‑of‑Time Award, and Who Is Machine Learning For?
JD Retail Technology
JD Retail Technology
Jul 6, 2026 · Artificial Intelligence

GROLE: Instance-Level Expert Routing for Incremental Learning in Oxygen AIIC Models

The paper introduces GROLE, a two‑stage incremental‑learning framework that builds a frozen pool of task‑specific LoRA experts and trains a lightweight instance‑level selector via reinforcement‑learning‑based gradient‑free optimization with Dirichlet sampling, achieving superior stability‑plasticity trade‑offs and state‑of‑the‑art results on multiple CL benchmarks.

GROLELarge Language ModelsLoRA
0 likes · 14 min read
GROLE: Instance-Level Expert Routing for Incremental Learning in Oxygen AIIC Models
Machine Heart
Machine Heart
Jul 6, 2026 · Artificial Intelligence

ICML 2026 Awards Unveiled: Breakthroughs in Diffusion Models, AI Alignment, and Reinforcement Learning

ICML 2026 announced ten award‑winning papers, highlighting novel insights such as the flexibility trap in diffusion language models, high‑accuracy sampling for diffusion, risks of AI alignment tools, a random‑matrix view of diffusion consistency, grokking in ridge regression, and an asynchronous deep‑RL framework, each accompanied by concise abstracts and links.

AI AlignmentGrokkingICML 2026
0 likes · 16 min read
ICML 2026 Awards Unveiled: Breakthroughs in Diffusion Models, AI Alignment, and Reinforcement Learning
Machine Heart
Machine Heart
Jul 5, 2026 · Artificial Intelligence

Why Larger Blocks Hurt Diffusion Language Model Inference and How T* Solves It

The article analyzes the trade‑off in masked diffusion language models where larger generation blocks increase parallelism but degrade reasoning, and shows how the T* progressive block‑scaling method using trajectory‑aware reinforcement learning stabilizes training and boosts accuracy across block sizes, with up to 15 % gains on MATH500.

Block ScalingMATH500T*
0 likes · 8 min read
Why Larger Blocks Hurt Diffusion Language Model Inference and How T* Solves It
IT Services Circle
IT Services Circle
Jul 3, 2026 · Artificial Intelligence

Ornith-1.0: The New Open‑Source Agentic Coding King with MIT License

Ornith-1.0, an open‑source model family released under the MIT license, tops multiple Agentic Coding benchmarks (SWE‑Bench Verified 82.4, Terminal‑Bench 77.5, etc.), spans from 9B to 397B parameters, and introduces joint reinforcement‑learning optimization of scaffold and solution to reshape AI‑assisted programming.

AI coding agentsAgentic codingOrnith-1.0
0 likes · 13 min read
Ornith-1.0: The New Open‑Source Agentic Coding King with MIT License
Machine Heart
Machine Heart
Jul 3, 2026 · Artificial Intelligence

ICML 2026: Enabling Multimodal Large Models to Reason Over Time with the Open‑Source TaRO Framework

The paper introduces the Temporal‑Aware Reasoning Optimization (TaRO) framework, which equips multimodal video large models with time‑aware reasoning via template‑based exploration, a temporal‑sensitivity reward, and progressive curriculum learning, achieving state‑of‑the‑art zero‑shot performance on several video temporal grounding benchmarks, including long‑video datasets.

TaROTemporal ReasoningVideo Temporal Grounding
0 likes · 9 min read
ICML 2026: Enabling Multimodal Large Models to Reason Over Time with the Open‑Source TaRO Framework
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 2, 2026 · Artificial Intelligence

Perfect Scores, Hidden Flaws: Qwen & Fudan Reveal Coding Agent Reward Issues

The article analyses how coding agents exploit unit‑test rewards by rewriting tests, explains why reward signals are only proxies for underspecified human intent, and argues that trustworthy AI requires a co‑evolving verification system rather than a single perfect validator.

AI safetyReward Designcoding agents
0 likes · 19 min read
Perfect Scores, Hidden Flaws: Qwen & Fudan Reveal Coding Agent Reward Issues
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

Perfect Scores, Hidden Flaws: Qwen and Fudan Expose Reward Design Dilemmas in Coding Agents

The article analyzes how coding agents can game test‑based rewards by altering verification signals, argues that reward signals are merely proxies for human intent, and proposes a co‑evolving verification system—combining scalable, faithful, and robust components—to reliably guide reinforcement‑learning agents.

AI safetyReward Designcoding agents
0 likes · 20 min read
Perfect Scores, Hidden Flaws: Qwen and Fudan Expose Reward Design Dilemmas in Coding Agents
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning

The paper introduces EMCES, a method that injects episodic memory into controllable diffusion models and uses a hash‑based state representation to generate high‑value synthetic samples, dramatically improving sample efficiency and downstream reinforcement‑learning performance while cutting storage and time costs.

Episodic MemoryHashingOffline RL
0 likes · 14 min read
EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning
Bilibili Tech
Bilibili Tech
Jul 1, 2026 · Artificial Intelligence

FATE Series (SABER & CASTER) Debuts at ACL 2026: Advanced LLM Reasoning

At ACL 2026 in San Diego, Bilibili’s tech team introduced the FATE series—SABER, which reduces overthinking in LLMs with a token‑budgeted switchable training, and CASTER, a community‑aware evaluation system built on Social‑CoT and the MEDEA framework that outperforms GPT‑5.2 and Claude‑4.5‑Opus on the new CASTER‑Bench, while also promoting the B‑UP talent recruitment program.

LLMMultimodal EvaluationSocial CoT
0 likes · 10 min read
FATE Series (SABER & CASTER) Debuts at ACL 2026: Advanced LLM Reasoning
Amap Tech
Amap Tech
Jun 30, 2026 · Artificial Intelligence

Six ECCV 2026 Papers – Vision, Video Generation, Visual‑Language Navigation

ECCV 2026 received 10,473 submissions and accepted 2,883 (27.5%); Gaode contributed six papers spanning computer vision, generative video, and visual‑language navigation, each presenting novel reinforcement‑learning or multimodal frameworks, new datasets, and benchmark results that outperform prior state‑of‑the‑art methods.

ECCV 2026Video Generationcomputer vision
0 likes · 13 min read
Six ECCV 2026 Papers – Vision, Video Generation, Visual‑Language Navigation
AntTech
AntTech
Jun 30, 2026 · Artificial Intelligence

FinixDoc Tackles Hard Financial Document Parsing with a 4B Model and Open‑Source Benchmark

FinixDoc is an end‑to‑end financial document parsing system built on a 4‑billion‑parameter Qwen3‑VL model that outperforms open‑source baselines on the newly released FinixDocBench, handling low‑quality, complex and ultra‑large documents through a specialized training pipeline and evaluation matrix.

FinixDocBenchQwen3-VL-4Bcontrastive learning
0 likes · 11 min read
FinixDoc Tackles Hard Financial Document Parsing with a 4B Model and Open‑Source Benchmark
Tencent Cloud Developer
Tencent Cloud Developer
Jun 30, 2026 · Artificial Intelligence

Why Claude Leads in Code Generation: A Deep Dive into Its Systemic Advantage

The article analyses why Claude’s code‑writing ability outperforms rivals, tracing its edge to a combination of verifiable‑reward reinforcement learning, Constitutional AI safety guards, a product‑driven data flywheel, multi‑level reward shaping, and continuous human‑in‑the‑loop evaluation on benchmarks such as SWE‑bench.

AI safetyAnthropicClaude
0 likes · 34 min read
Why Claude Leads in Code Generation: A Deep Dive into Its Systemic Advantage
AI Architecture Hub
AI Architecture Hub
Jun 30, 2026 · Artificial Intelligence

How to Fine‑Tune LLMs in 2026: Overcome the 30‑40% Error Wall with GRPO and RULER

Teams building LLM‑powered products often hit a wall where 30‑40% of responses are wrong and the model never learns from mistakes; the article explains how modern fine‑tuning using GRPO‑based reinforcement learning and the open‑source ART framework, together with the RULER reward‑free evaluator, lets small open‑source models surpass larger ones in cost, latency, and accuracy.

ART frameworkAgent TrainingGRPO
0 likes · 9 min read
How to Fine‑Tune LLMs in 2026: Overcome the 30‑40% Error Wall with GRPO and RULER
Machine Heart
Machine Heart
Jun 29, 2026 · Artificial Intelligence

How MWA™'s Long‑Sequence Bidirectional Physical Causal Chain Sets a New Record in Embodied AI

The article presents MWA™, the first long‑sequence bidirectional physical causal chain hidden‑space world model, details its bidirectional dynamics, latent‑action pre‑training, three‑gradient constraints and AnyPhys negative‑sample system, and shows it achieved a 75.2% success rate on the RoboCasa GR1 TableTop benchmark, surpassing leading competitors.

AnyPhysRoboCasa benchmarkbidirectional dynamics
0 likes · 14 min read
How MWA™'s Long‑Sequence Bidirectional Physical Causal Chain Sets a New Record in Embodied AI
Alipay Experience Technology
Alipay Experience Technology
Jun 29, 2026 · Artificial Intelligence

How a 4B‑Parameter UI‑UX Model Outperforms 235B Models in Detecting App Experience Flaws

The open‑source 4B UI‑UX multimodal model achieves a 0.7963 SOTA score on the UXBench benchmark, surpassing much larger models such as Claude‑4.5‑Sonnet and Qwen3‑VL‑Thinking, thanks to reward routing, asymmetric rewards, and data‑alchemy techniques, and it can be installed and deployed with a few pip commands.

AI for UXModel EvaluationMultimodal LLM
0 likes · 11 min read
How a 4B‑Parameter UI‑UX Model Outperforms 235B Models in Detecting App Experience Flaws
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 28, 2026 · Artificial Intelligence

Why the Log‑Ratio Reward in OPD Is Fundamentally Flawed and Should Be Replaced

The paper reveals that the unbounded log‑ratio reward used in vanilla On‑Policy Distillation causes extreme gradient variance, early‑stage instability, and poor final performance, and demonstrates that replacing the log with a bounded Box‑Cox power transform (PowerOPD) resolves these issues while improving accuracy, efficiency, and memory usage.

Box-CoxLarge Language ModelsOPD
0 likes · 16 min read
Why the Log‑Ratio Reward in OPD Is Fundamentally Flawed and Should Be Replaced
Machine Heart
Machine Heart
Jun 28, 2026 · Artificial Intelligence

Can AI Learn on the Job? RLVR, OPSD, and Dreaming for the Next‑Gen Training Paradigm

The article examines Dwarkesh Patel’s view that future AI must move beyond one‑off pre‑training to continual, on‑the‑job learning, discussing Reinforcement Learning with Verifiable Rewards (RLVR), the need for "grindable" tasks, and emerging approaches like on‑policy self‑distillation (OPSD) and "dreaming" to write real‑world experience back into model weights.

AI Training ParadigmsOn‑policy Self‑DistillationRLVR
0 likes · 12 min read
Can AI Learn on the Job? RLVR, OPSD, and Dreaming for the Next‑Gen Training Paradigm
Machine Heart
Machine Heart
Jun 28, 2026 · Artificial Intelligence

Why Robot AI Is Harder Than Large‑Scale Models: A First‑Principles Analysis

The article breaks down robot AI to a simple function mapping observations to actions, explains why latency, data diversity, and the need for split architectures make it far more challenging than training large language models, and surveys current solutions from edge‑cloud trade‑offs to action‑chunking and self‑learning.

AICloud ComputingData Collection
0 likes · 17 min read
Why Robot AI Is Harder Than Large‑Scale Models: A First‑Principles Analysis
Data Party THU
Data Party THU
Jun 27, 2026 · Artificial Intelligence

AI and Chemists Co-Develop TYR Inhibitors via Dual-Track Optimization

The study presents a dual-track strategy that combines deep reinforcement‑learning‑driven de novo molecular generation with expert‑guided medicinal chemistry to discover and optimize TYR inhibitors, demonstrating how AI expands chemical space while chemists ensure synthetic feasibility, leading to potent candidates such as AI10‑m15 with strong anti‑melanogenesis activity.

AI-driven drug discoveryTYR inhibitorchemical space exploration
0 likes · 8 min read
AI and Chemists Co-Develop TYR Inhibitors via Dual-Track Optimization
Amap Tech
Amap Tech
Jun 25, 2026 · Artificial Intelligence

ReaGeo: The First End‑to‑End LLM Geocoding Framework Linking Precise Mapping and Spatial Correlation

ReaGeo, a novel end‑to‑end geocoding system built on the Qwen2.5‑3B large language model, converts address text directly into Geohash sequences using chain‑of‑thought reasoning and GRPO reinforcement learning, achieving an average error of 119.6 m and 97.2 % accuracy within 500 m on Beijing data, surpassing commercial APIs and academic baselines while also modeling broader spatial correlation for line‑ and area‑type queries.

GeocodingLLMMap Search
0 likes · 15 min read
ReaGeo: The First End‑to‑End LLM Geocoding Framework Linking Precise Mapping and Spatial Correlation
Black & White Path
Black & White Path
Jun 25, 2026 · Artificial Intelligence

Can DeepSeek‑V4‑Fable’s AI Make Red Teams Redundant?

DeepSeek‑V4‑Fable, an autonomous AI agent built on a Chinese large‑model foundation and refined with SFT and GRPO, achieves a 58.7% overall solve rate on 300 held‑out CTF challenges, prompting a debate on its impact on red‑team workflows and security governance.

AICTFDeepSeek-V4-Fable
0 likes · 9 min read
Can DeepSeek‑V4‑Fable’s AI Make Red Teams Redundant?
Machine Heart
Machine Heart
Jun 24, 2026 · Industry Insights

Are Humanoid Robots Being Designed for Simulators? A Veteran’s Warning

The article warns that humanoid robot designers are sacrificing mechanical advantages—such as parallel joints and tendon‑driven hands—to make hardware easier for simulation, turning robust engineering principles into a simulation‑driven shortcut that risks limiting real‑world performance.

HardwareSim2RealSimulation
0 likes · 9 min read
Are Humanoid Robots Being Designed for Simulators? A Veteran’s Warning
Machine Heart
Machine Heart
Jun 24, 2026 · Artificial Intelligence

How APEIRIA Breaks the Black‑Box Barrier of 3D MLLMs (ICML 2026)

The paper introduces APEIRIA, a three‑stage curriculum that distills neuro‑symbolic program traces into 3D multi‑modal LLMs, enabling transparent spatial reasoning while preserving open‑vocabulary understanding, and demonstrates strong benchmark gains, modular upgrades, and zero‑shot generalization.

3D MLLMNeuro-Symbolic Reasoningchain-of-thought
0 likes · 11 min read
How APEIRIA Breaks the Black‑Box Barrier of 3D MLLMs (ICML 2026)
Ops Development & AI Practice
Ops Development & AI Practice
Jun 23, 2026 · Artificial Intelligence

Sovereign‑Free Routing: How Sakana AI’s Fugu Beats Claude Fable 5 Amid Geopolitical Constraints

Sakana AI’s newly released Fugu system uses a tiny 7B “commander” model to dynamically orchestrate a pool of global and local AI models, achieving a 73.7 % SWE‑bench Pro score that outperforms GPT‑5.5 and the heavily sanctioned Claude Fable 5, while illustrating a sovereign‑free routing strategy born from geopolitical and compute limitations.

AI geopoliticsBenchmarkingEvolutionary Algorithms
0 likes · 8 min read
Sovereign‑Free Routing: How Sakana AI’s Fugu Beats Claude Fable 5 Amid Geopolitical Constraints
AI Architecture Hub
AI Architecture Hub
Jun 23, 2026 · Artificial Intelligence

Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More

This article reviews seven recent AI research papers, detailing how SpatialClaw enables code‑based spatial reasoning for vision‑language models, SkillWeaver introduces compositional skill routing, PreAct compiles agent actions into reusable state‑machines, and other works advance world‑model inference, self‑designing RL environments, collective skill‑tree search, and process‑aligned reinforcement learning for diffusion LLMs.

Large Language Modelsagent reasoningdiffusion models
0 likes · 15 min read
Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 21, 2026 · Artificial Intelligence

xOPD Evolution: Mapping Recent OPD Improvements – Rephrased Same Problems vs. New Modules

This article surveys the latest on‑policy distillation (OPD) research, categorizing each work as either a reinterpretation of an existing problem or a modification of a different module, and highlights the experimental findings, design choices, and trade‑offs reported across the papers.

LLMOPDOn-Policy Distillation
0 likes · 31 min read
xOPD Evolution: Mapping Recent OPD Improvements – Rephrased Same Problems vs. New Modules
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 21, 2026 · Artificial Intelligence

Rank‑Only Rewards Accelerate One‑Step Text‑to‑Image Preference Optimization 3.5×

DrPO introduces a drifting‑field based, rank‑only reward mechanism for one‑step text‑to‑image models, enabling reinforcement‑learning‑after‑training without back‑propagating reward gradients; it speeds up training 3.51× versus DRaFT, works with non‑differentiable rewards, and improves generation quality on SD‑Turbo and SDXL‑Turbo.

DrPODrifting ModelHPSv3
0 likes · 11 min read
Rank‑Only Rewards Accelerate One‑Step Text‑to‑Image Preference Optimization 3.5×
Machine Heart
Machine Heart
Jun 21, 2026 · Artificial Intelligence

Why the Once‑Rejected PPO Algorithm Became a Pillar of Modern LLM Training

The article recounts how Proximal Policy Optimization, initially dismissed by NeurIPS 2017 for limited novelty, later became a cornerstone of RLHF and large‑language‑model training, illustrating how academic evaluation can miss long‑term impact, with parallels to other once‑rejected breakthroughs such as LSTM, SIFT and Dropout.

Algorithm RejectionLarge Language ModelsNeurIPS
0 likes · 5 min read
Why the Once‑Rejected PPO Algorithm Became a Pillar of Modern LLM Training
PaperAgent
PaperAgent
Jun 21, 2026 · Artificial Intelligence

What Drives AI Model Evolution? OpenAI’s New Findings on Beneficial Traits

OpenAI’s latest study shows that injecting just 5% of beneficial‑trait data into reinforcement‑learning training yields over 80% improvement across more than 50 alignment evaluations, revealing that a few underlying personality traits drive cross‑domain alignment and persist under adversarial pressure.

AI AlignmentLarge Language Modelsadversarial robustness
0 likes · 12 min read
What Drives AI Model Evolution? OpenAI’s New Findings on Beneficial Traits
Fighter's World
Fighter's World
Jun 21, 2026 · Artificial Intelligence

How Post‑Training Turns General AI into Enterprise‑Specific Intelligence

The article explains how post‑training shifts AI investment from massive pre‑training compute to targeted, low‑cost fine‑tuning that lets enterprises embed proprietary business logic, outlines the ROI formula based on data exclusivity, eval formalizability, and task frequency, and presents real‑world case studies.

AI economicsEnterprise AIevaluation
0 likes · 24 min read
How Post‑Training Turns General AI into Enterprise‑Specific Intelligence
Machine Heart
Machine Heart
Jun 21, 2026 · Artificial Intelligence

Is GRPO Obsolete? Why GLM‑5.2 Dropped It and What It Means for RL

GLM‑5.2 replaces the Group Relative Policy Optimization (GRPO) algorithm with a critic‑based PPO approach for long‑horizon tasks, arguing that GRPO’s group comparison breaks down on variable‑length trajectories, a shift that has sparked vigorous debate across the reinforcement‑learning community.

DeepSeekGLM-5.2GRPO
0 likes · 10 min read
Is GRPO Obsolete? Why GLM‑5.2 Dropped It and What It Means for RL
Machine Heart
Machine Heart
Jun 20, 2026 · Artificial Intelligence

DrPO: Ranking‑Only Rewards Boost One‑Step Text‑to‑Image Preference Optimization by 3.51×

DrPO introduces a ranking‑only reward that builds a drift field from on‑policy image samples to fine‑tune one‑step text‑to‑image models, achieving up to 3.51× faster training on large multimodal rewards, supporting non‑differentiable signals, and demonstrating superior quality across multiple benchmarks.

Drifting Preference Optimizationdrift fieldnon-differentiable reward
0 likes · 14 min read
DrPO: Ranking‑Only Rewards Boost One‑Step Text‑to‑Image Preference Optimization by 3.51×
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 18, 2026 · Artificial Intelligence

From Imitation to Optimization: Recent Advances in On-Policy Distillation

This article surveys the latest research on On-Policy Distillation for large language models, covering methods that improve training stability, self‑distillation frameworks, and detailed analyses of when and why OPD succeeds or fails, with concrete experimental results and practical insights.

Entropy-AwareLarge Language ModelsOn-Policy Distillation
0 likes · 19 min read
From Imitation to Optimization: Recent Advances in On-Policy Distillation
Kuaishou Tech
Kuaishou Tech
Jun 18, 2026 · Artificial Intelligence

Kuaishou Tech Team Highlights Multiple ICML 2026 Papers Across AI Domains

The Kuaishou technology team reports that several of its papers were accepted at the prestigious ICML 2026 conference—including a spotlight paper on metaphor video understanding, works on causal discovery for irregular time series, image super‑resolution, large‑scale notification dispatch, full‑order ranking, phase‑aware MoE for RL, end‑to‑end e‑commerce search, spatial‑reasoning rewards, a unified SWE benchmark, video temporal grounding, and interpretable transformers—while also inviting attendees to visit their booth B101 in Seoul.

ICML 2026KuaishouLarge Language Models
0 likes · 18 min read
Kuaishou Tech Team Highlights Multiple ICML 2026 Papers Across AI Domains
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 18, 2026 · Artificial Intelligence

Can a 3B Model Rival Claude Opus 4.5? Benchmark Gaps or Aggressive Post‑Training?

VibeThinker‑3B, a 3‑billion‑parameter language model built on Qwen2.5‑Coder‑3B, achieves scores within the range of 671 B‑parameter models on benchmarks such as LiveCodeBench, AIME26, IMO‑AnswerBench and GPQA, thanks to a two‑stage SFT, multi‑domain reinforcement learning, offline self‑distillation and a claim‑reliability (CLR) evaluator that together push its reasoning ability to the frontier.

Large Language ModelsParameter EfficiencyVibeThinker-3B
0 likes · 9 min read
Can a 3B Model Rival Claude Opus 4.5? Benchmark Gaps or Aggressive Post‑Training?
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 18, 2026 · Artificial Intelligence

TNT: Dynamic Token Limits Slash Reward Hacking in Mixed Inference Models Below 10%

The paper introduces Thinking‑Based Non‑Thinking (TNT), a reinforcement‑learning approach that sets a per‑question dynamic token ceiling for non‑thinking mode using the answer length from thinking mode, cutting reward‑hacking incidence to under 10% while boosting accuracy and cutting token usage by nearly half across several math benchmarks.

ACL 2026NLPTNT
0 likes · 10 min read
TNT: Dynamic Token Limits Slash Reward Hacking in Mixed Inference Models Below 10%
vivo Internet Technology
vivo Internet Technology
Jun 17, 2026 · Artificial Intelligence

BeautyGRPO: A New Reinforcement Learning Framework that Recreates Realistic Portraits

The CVPR 2026 paper introduces BeautyGRPO, a reinforcement‑learning framework that leverages the fine‑grained FRPref‑10K portrait‑retouching preference dataset and a novel Dynamic Path Guidance algorithm to simultaneously enhance skin texture, preserve identity features, and achieve superior aesthetic alignment, outperforming existing retouching models on objective metrics and user preference tests.

BeautyGRPOCVPR 2026FRPref-10K
0 likes · 9 min read
BeautyGRPO: A New Reinforcement Learning Framework that Recreates Realistic Portraits
Machine Heart
Machine Heart
Jun 17, 2026 · Artificial Intelligence

Why Massive GPU Farms Still Fail to Deliver Enterprise‑Ready AI—and How Jiuzhang’s AI Factory Solves It

Despite a surge to over 140 trillion daily token calls in China, enterprises find general large models can answer but cannot execute business workflows, a gap Jiuzhang Yunji addresses with its AI Factory that combines reinforcement‑learning‑driven professional model production, a five‑capability training platform, and an Inference OS to industrialize AI at scale.

AI infrastructureToken economyindustrial AI
0 likes · 22 min read
Why Massive GPU Farms Still Fail to Deliver Enterprise‑Ready AI—and How Jiuzhang’s AI Factory Solves It
Machine Heart
Machine Heart
Jun 17, 2026 · Artificial Intelligence

Why RL‑Trained Agents Still Fail to Reason Actively: The Information Self‑Locking Problem

The paper reveals that outcome‑based reinforcement learning often traps LLM agents in an information self‑locking regime where weak action selection and belief tracking prevent proper credit assignment, and introduces AREW, a lightweight advantage‑reweighting method that restores active reasoning across multiple tasks and models.

AREWAgentic RLLLM Agents
0 likes · 24 min read
Why RL‑Trained Agents Still Fail to Reason Actively: The Information Self‑Locking Problem
Machine Heart
Machine Heart
Jun 15, 2026 · Artificial Intelligence

HyVLA-0.5: Sub‑millimeter UMI Data and Real‑Robot Reinforcement Eliminate Heavy Tele‑operation

HyVLA-0.5, an open‑source embodied VLA model from Tencent Robotics X, leverages over 10,000 hours of sub‑millimeter UMI demonstration data and a novel FlowPRO reinforcement pipeline to achieve more than 90% success on simulated and real‑world tasks, while supporting cross‑embodiment transfer and asynchronous deployment.

FlowPROHyVLA-0.5UMI data
0 likes · 16 min read
HyVLA-0.5: Sub‑millimeter UMI Data and Real‑Robot Reinforcement Eliminate Heavy Tele‑operation
Top Architect
Top Architect
Jun 13, 2026 · Artificial Intelligence

What Is an Inference Large Language Model? A Visual Guide

The article explains inference‑type large language models, how they differ from traditional models by breaking questions into reasoning steps, the shift from training‑time to test‑time compute, scaling‑law insights, validation techniques, proposal‑distribution tricks, and the detailed training pipeline of DeepSeek‑R1, while also discussing failed experiments and future directions.

DeepSeek-R1Large Language Modelsinference models
0 likes · 20 min read
What Is an Inference Large Language Model? A Visual Guide
Bilibili Tech
Bilibili Tech
Jun 12, 2026 · Artificial Intelligence

A New UGC Video Evaluation Paradigm Built on 17 Billion Real User Interactions

The paper introduces CASTER, a multimodal AI system that uses Social‑CoT reasoning and the MEDEA framework to simulate diverse audience reactions, benchmarked on the large‑scale CASTER‑Bench dataset, and demonstrates superior performance over GPT‑5.2, Claude‑4.5‑Opus, and traditional VQA methods while already being deployed on Bilibili.

Community resonanceSocial CoTUGC video evaluation
0 likes · 9 min read
A New UGC Video Evaluation Paradigm Built on 17 Billion Real User Interactions