Tagged articles

GUI agents

6 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 23, 2026 · Artificial Intelligence

Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents

Workflow Gym introduces a realistic, long‑horizon benchmark covering 56 professional applications and 338 real‑world workflows, revealing that top GUI agents like Gemini 3.1 Pro achieve only about 30% single‑run success and exposing key failure modes such as consistency breaks and lack of domain knowledge.

AI performanceGUI agentsWorkflow Gym
0 likes · 14 min read
Workflow Gym: Moving Beyond Simulated Tests to Bridge the Gap for GUI Agents
DataFunSummit
DataFunSummit
Jul 10, 2026 · Artificial Intelligence

UI-MOPD Enables Cross‑Platform GUI Agents to Retain Skills Without Forgetting

The article analyzes why GUI agents trained on both desktop (mouse‑click) and mobile (touch) interactions suffer from behavior collapse and catastrophic forgetting, introduces the UI‑MOPD framework that assigns platform‑specific teachers through on‑policy distillation, and shows an 8B model surpassing a 235B baseline on OSWorld and MobileWorld benchmarks while preserving general GUI understanding.

GUI agentsOn-Policy DistillationUI‑MOPD
0 likes · 8 min read
UI-MOPD Enables Cross‑Platform GUI Agents to Retain Skills Without Forgetting
Machine Heart
Machine Heart
Apr 7, 2026 · Artificial Intelligence

How LaSM Pulls GUI Agents’ Attention Back from Deceptive Pop‑up Attacks

The paper introduces LaSM (Layer‑wise Scaling Mechanism), a training‑free weight‑scaling patch that restores GUI agents’ focus by selectively amplifying attention and MLP weights in critical middle layers, effectively defending against pop‑up‑style environment injection attacks without retraining or altering model architecture.

GUI agentsLaSMattention analysis
0 likes · 14 min read
How LaSM Pulls GUI Agents’ Attention Back from Deceptive Pop‑up Attacks
Baobao Algorithm Notes
Baobao Algorithm Notes
Jan 26, 2026 · Artificial Intelligence

From Search Ads to Foundation Models: My Journey Building the EvoCUA GUI Agent

The author explains why he transitioned from search advertising algorithms to foundation model research, outlines the four typical activities of base‑model teams, and shares detailed technical insights, experimental practices, and scaling strategies that led the EvoCUA GUI Agent to achieve open‑source SOTA on OSWorld.

AI researchGUI agentsexperiment methodology
0 likes · 17 min read
From Search Ads to Foundation Models: My Journey Building the EvoCUA GUI Agent
AI Frontier Lectures
AI Frontier Lectures
Mar 24, 2025 · Artificial Intelligence

What Can AI Agents Learn from the Latest AIR 2025 Research?

The article compiles insights from the AIR 2025 conference and related talks, covering the evolution of agents from reinforcement‑learning to LLM‑driven systems, novel agent architectures like AIDE, GUI agents, natural‑language reinforcement learning, and scaling advances in large language models such as Qwen, while highlighting key algorithms, benchmarks, and open research questions.

AI AgentsGUI agentsLarge Language Models
0 likes · 27 min read
What Can AI Agents Learn from the Latest AIR 2025 Research?