Tagged articles

long-horizon tasks

11 articles · Page 1 of 1
Old Zhang's AI Learning
Old Zhang's AI Learning
Sep 1, 2026 · Artificial Intelligence

Doubao-Seed-Evolving Tested: Coding, Multimodal, and Long-Horizon Agent Skills

The author evaluates Doubao-Seed-Evolving's latest upgrades across coding, agent, multimodal, and long-horizon skill execution using real-world tasks like PPT-to-website conversion, personal project management, SVG generation, 3D Rubik's cube animation, and a complex 30-minute web scraping skill, finding strong autonomous task decomposition, testing, and error recovery.

AI model testingDoubao-Seed-EvolvingLLM Evaluation
0 likes · 10 min read
Doubao-Seed-Evolving Tested: Coding, Multimodal, and Long-Horizon Agent Skills
Machine Heart
Machine Heart
Aug 29, 2026 · Artificial Intelligence

Recuris: A New Memory Paradigm That Boosts Performance from 3B Models to Claude Opus 5

Recuris introduces a compact task‑state‑driven memory architecture and gated recursive self‑improvement, enabling agents to use and evolve memory more reliably and delivering large, consistent gains from 3B open‑source models up to frontier models such as Claude Opus 5 across multiple long‑horizon benchmarks.

LLM AgentsRecurisRecursive Self-Improvement
0 likes · 11 min read
Recuris: A New Memory Paradigm That Boosts Performance from 3B Models to Claude Opus 5
PaperAgent
PaperAgent
Aug 15, 2026 · Artificial Intelligence

DeepSeek Harness Open‑Source and Alibaba’s LongHorizon‑Harness: MEA Loop Boosts Long‑Horizon AI Agents

The article introduces DeepSeek Harness and Alibaba’s LongHorizon‑Harness, explains their Manage‑Execute‑Audit (MEA) loop for explicit task‑state management, and shows benchmark improvements—WeaveBench up to 80.7%, OSWorld 3×, Terminal‑Bench 77.2%—while analyzing token costs, compute allocation, and case studies of failure recovery.

AI agentsDeepSeek HarnessLongHorizon-Harness
0 likes · 9 min read
DeepSeek Harness Open‑Source and Alibaba’s LongHorizon‑Harness: MEA Loop Boosts Long‑Horizon AI Agents
PaperAgent
PaperAgent
Aug 11, 2026 · Artificial Intelligence

Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM

Researchers introduce Skill‑Entropy, a metric quantifying the difficulty of switching between reasoning skills in long‑horizon tasks, build the 558‑skill Skill²‑Bench, and show that Skill‑Entropy‑RL training dramatically improves cross‑skill performance of LLMs such as Qwen3, closing the gap observed in standard benchmarks.

Cross‑Skill ReasoningLLM BenchmarkingQwen3
0 likes · 12 min read
Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM
Machine Heart
Machine Heart
Jun 21, 2026 · Artificial Intelligence

Is GRPO Obsolete? Why GLM‑5.2 Dropped It and What It Means for RL

GLM‑5.2 replaces the Group Relative Policy Optimization (GRPO) algorithm with a critic‑based PPO approach for long‑horizon tasks, arguing that GRPO’s group comparison breaks down on variable‑length trajectories, a shift that has sparked vigorous debate across the reinforcement‑learning community.

DeepSeekGLM-5.2GRPO
0 likes · 10 min read
Is GRPO Obsolete? Why GLM‑5.2 Dropped It and What It Means for RL
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 19, 2026 · Artificial Intelligence

AutoResearch SKILL Open‑Source: Framework for Long‑Horizon Autonomous Research

The Deli AutoResearch SKILL, now open‑sourced, presents a three‑layer framework that tackles cognitive loops, stalling, and runtime fragility in long‑horizon tasks by persisting state, detecting stalls, and using a heartbeat watchdog, and it includes a paper‑writing skill with self‑play experiments that achieve self‑rated scores up to 8.6.

RL experimentsSelf-Playautonomous agents
0 likes · 17 min read
AutoResearch SKILL Open‑Source: Framework for Long‑Horizon Autonomous Research
Machine Heart
Machine Heart
May 25, 2026 · Artificial Intelligence

Claude’s Pass Rate Under 4%: SaaS‑Bench Shatters the “Fully Automated Office” Dream

SaaS‑Bench evaluates AI agents on 23 real SaaS applications and 106 cross‑app, long‑horizon tasks, revealing that even the strongest model, Claude Opus 4.7, passes fewer than four percent of tasks and exposing four structural failure modes that separate benchmark scores from true office productivity.

AI agentsBenchmarkingClaude Opus
0 likes · 10 min read
Claude’s Pass Rate Under 4%: SaaS‑Bench Shatters the “Fully Automated Office” Dream
Machine Heart
Machine Heart
May 22, 2026 · Artificial Intelligence

HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits

HiF-VLA introduces a motion‑centric bidirectional spatiotemporal reasoning framework with a joint‑expert module that simultaneously predicts future visual motion and generates high‑precision action sequences, eliminating visual redundancy, cutting inference latency and memory usage, and achieving superior success rates on long‑horizon benchmarks such as CALVIN and LIBERO‑LONG.

HiF-VLAMotion RepresentationVision-Language-Action
0 likes · 9 min read
HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits
Machine Heart
Machine Heart
Apr 26, 2026 · Artificial Intelligence

Surpassing Claude Mythos and GPT‑5.5: Stanford’s New LLM‑as‑a‑Verifier Agent Framework

Stanford, Berkeley and Nvidia introduce LLM‑as‑a‑Verifier, a verification framework that scales verification compute, uses fine‑grained score tokens, repeated checks and criteria decomposition to boost agent performance, eliminate scoring ties and achieve SOTA results on Terminal‑Bench, surpassing Claude Mythos and GPT‑5.5 while improving safety in long‑horizon tasks.

Agent VerificationLLMLLM-as-a-Verifier
0 likes · 8 min read
Surpassing Claude Mythos and GPT‑5.5: Stanford’s New LLM‑as‑a‑Verifier Agent Framework
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Mar 12, 2026 · Artificial Intelligence

LongHorizonUI: A Unified Robust Framework for Long‑Horizon GUI Agent Automation

LongHorizonUI tackles the steep success‑rate drop of GUI agents on tasks longer than 10‑15 steps by introducing three tightly coupled modules—enhanced perception, deep reflective decision, and compensatory execution—and validates the approach on the new LongGUIBench benchmark with consistent performance gains across both app and game scenarios.

GUI automationICLR 2026benchmark
0 likes · 12 min read
LongHorizonUI: A Unified Robust Framework for Long‑Horizon GUI Agent Automation