Tagged articles

LLM agents

166 articles · Page 1 of 2
dbaplus Community
dbaplus Community
Sep 28, 2026 · Artificial Intelligence

Building the Agent Self-Evolution Flywheel: Evaluation → Memory → Implementation → Control

This article presents a comprehensive four-stage flywheel methodology for agent self-evolution—evaluation, memory, implementation, and human-in-the-loop control—detailing core challenges, engineering practices, and integration patterns to create a continuous improvement loop for AI agents.

AI EngineeringAgent Self-EvolutionEvaluation Systems
0 likes · 53 min read
Building the Agent Self-Evolution Flywheel: Evaluation → Memory → Implementation → Control
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 28, 2026 · Artificial Intelligence

Why Plan Mode in AI Coding Tools Is Already Dying

Ayman Nadeem, founder of YC-backed AI coding startup Nuanced, explains why the once-essential Plan Mode — where developers write detailed specs before code generation — is becoming obsolete as models grow capable of autonomous reasoning, iterative action, and self-correction without heavy upfront planning documents.

AI coding assistantsAI-generated documentationClaude Code
0 likes · 14 min read
Why Plan Mode in AI Coding Tools Is Already Dying
Architect
Architect
Sep 28, 2026 · Artificial Intelligence

Dream-RSI: How Agents 'Dream' on Past Searches to Optimize Exploration Budgets

Dream-RSI introduces a recursive self-improvement framework where AI agents improve exploration strategies by replaying historical search trees offline, enabling budget allocation decisions without retraining models, demonstrated across algorithm engineering, math optimization, and GPU kernel tasks with reduced search costs.

AI researchDream-RSIGoogle DeepMind
0 likes · 26 min read
Dream-RSI: How Agents 'Dream' on Past Searches to Optimize Exploration Budgets
AI Large Model Application Practice
AI Large Model Application Practice
Sep 28, 2026 · Artificial Intelligence

JEV in Agent Systems: 8 High-Speed Decision Patterns for LLM Agents

This article details eight practical scenarios where JEV (Judgment and Evaluation) models accelerate agent decision-making, including ReAct loop control, web automation, tool routing, safety guards, task evaluation, model routing, RAG filtering, and real-time robotics control, showing how lightweight classifiers reduce latency and token costs.

Agent SystemsDecision ModelsJev
0 likes · 20 min read
JEV in Agent Systems: 8 High-Speed Decision Patterns for LLM Agents
DataFunTalk
DataFunTalk
Sep 25, 2026 · Artificial Intelligence

MemoHarness: Agent Evolution Shifts from Model Weights to External Harness

MemoHarness introduces a six-dimensional external harness system that adapts agent behavior through case-based experience from execution trajectories, showing improvements on terminal, coding, and finance tasks while acknowledging limited experimental scale and selective cross-task transferability.

Agent HarnessFinanceAgentHarness Engineering
0 likes · 16 min read
MemoHarness: Agent Evolution Shifts from Model Weights to External Harness
DataFunSummit
DataFunSummit
Sep 23, 2026 · Artificial Intelligence

MemoHarness: Agent Evolution Shifts from Model Weights to External Harness Systems

MemoHarness introduces a six-dimensional harness framework where agents learn from execution trajectories during development, then adapt to new tasks via case-based retrieval—showing gains on terminal, coding, and finance tasks while acknowledging limited scale, selective transfer, and no online learning.

Agent HarnessFinanceAgentHarness Engineering
0 likes · 18 min read
MemoHarness: Agent Evolution Shifts from Model Weights to External Harness Systems
Amap Tech
Amap Tech
Sep 22, 2026 · Artificial Intelligence

DIANOIA: Multi-Agent Diagnosis & Repair for Navigation Tool Trajectories

Amap introduces DIANOIA, a multi-agent system that diagnoses and repairs low-confidence tool-call trajectories for navigation Live mode by generating diverse candidates, executing them in real tool environments, cross-reviewing failures, and synthesizing corrected sequences, boosting data quality for LLM training while reducing compute costs.

DIANOIAEMNLP 2026LLM agents
0 likes · 28 min read
DIANOIA: Multi-Agent Diagnosis & Repair for Navigation Tool Trajectories
Data Party THU
Data Party THU
Sep 22, 2026 · Artificial Intelligence

Why Your Multi-Agent System Is Costlier, Slower, and Worse Than a Single Agent

The article analyzes why multi-agent systems often underperform single agents, identifying context isolation as the key benefit only when tasks exceed a single context window, detailing six architectural patterns, cost multipliers up to 15x tokens, and a decision framework for choosing between multi-agent and single-agent approaches with proper engineering practices.

AI EngineeringCost OptimizationLLM agents
0 likes · 19 min read
Why Your Multi-Agent System Is Costlier, Slower, and Worse Than a Single Agent
Machine Heart
Machine Heart
Sep 17, 2026 · Artificial Intelligence

Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab

Shanghai AI Lab and university collaborators propose SHE and SafeEvolve, two frameworks that evolve agent safety by learning from execution trajectories: SHE updates a modular safety harness via trajectory-driven evolution, while SafeEvolve distills verified harness experience into the policy model through SFT and RL, reducing attack success rates on benchmarks.

Agent SafetyHarness EvolutionLLM agents
0 likes · 12 min read
Beyond Prompts: Harness-Policy Co-Evolution for Agent Safety by Shanghai AI Lab
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 16, 2026 · Artificial Intelligence

NetCanvas: Huawei GTS Gives AI Agents an Interactive Visual Topology for Network Troubleshooting

Huawei GTS's NetCanvas provides LLM agents with an interactive visual topology as external working memory, enabling them to 'watch' network topology during troubleshooting; on CTBench, pass rates jump from 30.3% to 54.5% (63.6% with priors), dual-firewall/ECMP tasks soar from ~10% to ~90%, with zero regressions and 26–45% token cost reduction.

CTBenchHuawei GTSLLM agents
0 likes · 11 min read
NetCanvas: Huawei GTS Gives AI Agents an Interactive Visual Topology for Network Troubleshooting
Machine Heart
Machine Heart
Sep 14, 2026 · Artificial Intelligence

T-Mem: Teaching AI Associative Recall Beyond Similarity Search

Tencent's T-Mem introduces associative recall for LLM agents by pre-storing trigger cues during memory writing, enabling retrieval via semantic associations rather than surface similarity, achieving SOTA on LoCoMo and LoCoMo-Plus benchmarks with minimal performance drop on associative tasks.

EMNLP 2026LLM agentsLOCOMO benchmark
0 likes · 13 min read
T-Mem: Teaching AI Associative Recall Beyond Similarity Search
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety

SafeEvolve introduces a co-evolution framework where Agent Harness and Policy jointly learn from execution trajectories, reducing attack success rates to 0.79% on AgentDojo and 2.42% on Qwen3-4B while improving task utility, enabling continuous safety improvement from real-world experience.

AI safetyAgent SafetyHarness-Policy Co-Evolution
0 likes · 9 min read
SafeEvolve: Co-Evolving Harness and Policy for Self-Improving Agent Safety
DataFunSummit
DataFunSummit
Sep 11, 2026 · Artificial Intelligence

Graph Engineering Restructures Agent Systems: From Harness to Ontology

This article reviews a 2026 paper on Graph Engineering for LLM agents, detailing the shift from individual agent intelligence to system intelligence via explicit task DAGs, runtime state management with checkpoints and replay, multi-agent coordination through capability modeling, and ontology engineering for shared semantics.

Agent CoordinationDAG SchedulingGraph Engineering
0 likes · 20 min read
Graph Engineering Restructures Agent Systems: From Harness to Ontology
PaperAgent
PaperAgent
Sep 11, 2026 · Artificial Intelligence

Google's Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Google introduces Procedural Graphs, a self-evolving graph structure that explicitly encodes procedural knowledge for LLM agents, enabling them to locate, extract, and generate contextual guidance from a (procedure, relation, procedure) triplet graph, achieving 85% survival from 0% in CFO simulation and winning 21 of 24 model-benchmark combinations across 7 benchmarks and 4 LLMs.

Agent ReliabilityGraph-based ReasoningLLM agents
0 likes · 9 min read
Google's Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
DataFunTalk
DataFunTalk
Sep 9, 2026 · Artificial Intelligence

Graph Engineering Restructures Agent Systems: From Harness to Ontology

A 2026 survey paper introduces Graph Engineering as the next phase for LLM agents, shifting focus from individual model capabilities to system-level organization via explicit task DAGs, runtime state management with provenance and recovery, capability-based agent coordination, and a graph-native control plane that treats tasks, agents, and state as first-class system objects.

Agent CoordinationDAG SchedulingGraph Engineering
0 likes · 22 min read
Graph Engineering Restructures Agent Systems: From Harness to Ontology
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 8, 2026 · Artificial Intelligence

Strong Model ≠ Strong Agent: PolyWorkBench Benchmarks Cross-Lingual Long-Horizon Workflows

PolyWorkBench introduces 67 cross-lingual long-horizon workflow tasks across 5 domains and 10 languages, revealing that top models like Claude Opus 4.8 show 22.5% performance variance across agent harnesses and significant drops on commerce tasks and low-resource languages due to language understanding and cross-lingual coordination errors.

Agent BenchmarksAgent HarnessClaude Opus
0 likes · 10 min read
Strong Model ≠ Strong Agent: PolyWorkBench Benchmarks Cross-Lingual Long-Horizon Workflows
Architecture Development Notes
Architecture Development Notes
Sep 8, 2026 · Artificial Intelligence

OpenTelemetry GenAI Semantic Conventions: Modeling Agent Runs as Span Trees

The article explains how OpenTelemetry GenAI semantic conventions address the gap in agent observability by defining a standardized span tree structure (invoke_agent, execute_tool, chat, etc.) and three critical fields (gen_ai.agent.name, gen_ai.conversation.id, gen_ai.operation.name) to capture tool calls, retries, and child-agent handoffs that auto-instrumentation misses.

GenAILLM agentsOpenTelemetry
0 likes · 10 min read
OpenTelemetry GenAI Semantic Conventions: Modeling Agent Runs as Span Trees
Machine Heart
Machine Heart
Sep 8, 2026 · Artificial Intelligence

Behavior Consistency Beats State Consistency in Text World Models for Agents

The paper introduces BehR, a behavior consistency reward for training text-based world models, showing that optimizing for agent decision alignment rather than text fidelity improves trajectory-level consistency across 16 configurations, reduces false positives in offline evaluation from 42.5% to 9.5%, and enhances lookahead planning for weaker agents.

Behavior ConsistencyEMNLP 2026GRPO
0 likes · 11 min read
Behavior Consistency Beats State Consistency in Text World Models for Agents
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Sep 8, 2026 · Artificial Intelligence

AI Agent Development: The Dual Challenge of Thinking Engineering & Distributed Systems

This article argues that AI agent development shifts from traditional coding to dual-system engineering: single agents require thinking logic design (prompt engineering, reasoning frameworks), while multi-agent systems demand distributed architecture skills (task graphs, state management, concurrency control), combining probabilistic reasoning with system reliability challenges.

AI AgentsLLM agentsLangGraph
0 likes · 14 min read
AI Agent Development: The Dual Challenge of Thinking Engineering & Distributed Systems
DataFunTalk
DataFunTalk
Sep 7, 2026 · Artificial Intelligence

Graph Engineering Rebuilds Agent Systems: From Harness to System Intelligence

A 2026 survey paper introduces Graph Engineering as the system layer that organizes LLM agents into reliable multi-agent workflows through explicit DAGs, runtime state management, fault tolerance, and a control plane, shifting focus from individual agent capabilities to system-level engineering.

Agent CoordinationDAG SchedulingGraph Engineering
0 likes · 21 min read
Graph Engineering Rebuilds Agent Systems: From Harness to System Intelligence
DataFunTalk
DataFunTalk
Sep 4, 2026 · Artificial Intelligence

MemoHarness: Agent Evolution Moves Outside the Model to External Harness Systems

This article analyzes MemoHarness, a framework that evolves AI agents by optimizing six external harness dimensions—context assembly, tool interaction, generation control, task orchestration, memory management, and output processing—instead of model weights, demonstrating gains on terminal, coding, and finance tasks while noting limited experimental scale and selective cross-task transfer.

Agent HarnessAgent MemoryExternal System Evolution
0 likes · 22 min read
MemoHarness: Agent Evolution Moves Outside the Model to External Harness Systems
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Sep 1, 2026 · Artificial Intelligence

XAlpha: AI Quant Researcher with Memory & Reflection for Alpha Discovery

XAlpha introduces a memory-driven AI quantitative researcher that automates the full hypothesis-to-code alpha discovery loop via a multi-brain architecture, integrating report knowledge, hypothesis planning, factor evolution, empirical validation, and feedback integration, achieving superior performance on CSI300.

AI quantitative researcherCSI300LLM agents
0 likes · 35 min read
XAlpha: AI Quant Researcher with Memory & Reflection for Alpha Discovery
AI Engineering
AI Engineering
Sep 1, 2026 · Artificial Intelligence

Why Long-Term LLM Agents Fail: The Same Design Choice Behind Two Deaths

Long‑term LLM agents suffer from ever‑slowing execution and context poisoning because they continuously append every observation, action, and reasoning step to the prompt, but the SKILL.state approach replaces this growing history with a compact mutable state, dramatically cutting token usage while boosting accuracy and robustness across diverse benchmarks.

BenchmarkGeminiGemma
0 likes · 11 min read
Why Long-Term LLM Agents Fail: The Same Design Choice Behind Two Deaths
Machine Heart
Machine Heart
Aug 29, 2026 · Artificial Intelligence

Recuris: A New Memory Paradigm That Boosts Performance from 3B Models to Claude Opus 5

Recuris introduces a compact task‑state‑driven memory architecture and gated recursive self‑improvement, enabling agents to use and evolve memory more reliably and delivering large, consistent gains from 3B open‑source models up to frontier models such as Claude Opus 5 across multiple long‑horizon benchmarks.

LLM agentsRecurisRecursive Self-Improvement
0 likes · 11 min read
Recuris: A New Memory Paradigm That Boosts Performance from 3B Models to Claude Opus 5
PaperAgent
PaperAgent
Aug 25, 2026 · Artificial Intelligence

A New Paradigm for Teaching Agents Tools: Insights from ACL 2026 ToolCPT

ToolCPT demonstrates that embedding real‑world tool knowledge during LLM pre‑training, rather than fine‑tuning, dramatically improves agent performance, using a mined corpus of 5.1 million proxy tools, detailed playbooks, and a 10 % tool‑data mix that yields up to 7.15‑point gains on benchmark tasks.

ACL 2026Agent BenchmarksLLM agents
0 likes · 8 min read
A New Paradigm for Teaching Agents Tools: Insights from ACL 2026 ToolCPT
Alibaba Cloud Native
Alibaba Cloud Native
Aug 19, 2026 · Artificial Intelligence

Reproducible Three‑Dimensional Evaluation of DeepSeek Harness on Alibaba Cloud AgentLoop

This article presents a reproducible, three‑dimensional deterministic evaluation framework (outcome, compliance, process) built on Alibaba Cloud AgentLoop, applies it to a 10‑task subset of terminal‑bench 2.1 to benchmark DeepSeek Harness against Codex, details the methodology, results, and future research directions.

AI agent assessmentAgentLoopDeepSeek Harness
0 likes · 26 min read
Reproducible Three‑Dimensional Evaluation of DeepSeek Harness on Alibaba Cloud AgentLoop
AI Engineer Programming
AI Engineer Programming
Aug 15, 2026 · Artificial Intelligence

Mastering Stateful AI Agent Orchestration with LangGraph

LangGraph is an open‑source framework that replaces linear LLM pipelines with graph‑based, stateful agents, offering loops, conditional branching, persistent checkpoints, human‑in‑the‑loop support, and built‑in monitoring, enabling complex multi‑step workflows that scale from simple chatbots to enterprise‑grade AI assistants.

AI workflowLLM agentsLangChain
0 likes · 20 min read
Mastering Stateful AI Agent Orchestration with LangGraph
DataFunTalk
DataFunTalk
Aug 9, 2026 · Artificial Intelligence

Beyond Model Scaling: How Agent Training Shifts from Bulk Environments to Designed Worlds

Recent ACL 2026 papers (EnvScaler, AgentScaler, Echoverse, and Beyond Simply Environment Scaling) reveal a transition from merely increasing the number of training environments to carefully designing environment distributions that improve agent performance, with empirical evidence showing both gains and diminishing returns.

Agent ScalingEnvironment ScalingExperience Distribution
0 likes · 17 min read
Beyond Model Scaling: How Agent Training Shifts from Bulk Environments to Designed Worlds
TonyBai
TonyBai
Aug 5, 2026 · Artificial Intelligence

Why Filesystem‑Based Memory Beats Expectations for LLM Agents – 5 Surprising Findings

A new multi‑institution study formalizes LLM‑agent memory as a three‑role filesystem store, evaluates six memory shapes across four dialogue and one skill benchmark, and reveals that structured organization halves retrieval cost but does not guarantee higher answer accuracy, with model personality and toolsets driving the shape of the memory store.

LLM agentsagent managementfilesystem memory
0 likes · 20 min read
Why Filesystem‑Based Memory Beats Expectations for LLM Agents – 5 Surprising Findings
Architect
Architect
Aug 4, 2026 · Artificial Intelligence

From Bug Fixes to Completed Work: Insights from Tencent’s WorkBuddy Bench

WorkBuddy Bench reveals why fixing a bug does not equal finishing a task, proposing a four‑layer completion model and a reproducible benchmark that evaluates agents across Code, Web, Office, and Security workspaces, showing how prompts, context, harnesses, loops and graphs must be verified to claim true completion.

AI evaluationAgent BenchmarkLLM agents
0 likes · 19 min read
From Bug Fixes to Completed Work: Insights from Tencent’s WorkBuddy Bench
Architect
Architect
Aug 3, 2026 · Artificial Intelligence

Separating Planning and Execution in LLM Agents: ArbiterOS Governance Kernel

As LLM agents gain the ability to read code, modify files, send emails and call APIs, ArbiterOS introduces a runtime governance layer that turns model intents into structured, traceable instructions, enabling policies to approve, block, or request confirmation before any high‑risk action is executed.

AI safetyArbiterOSLLM agents
0 likes · 23 min read
Separating Planning and Execution in LLM Agents: ArbiterOS Governance Kernel
DataFunSummit
DataFunSummit
Aug 1, 2026 · Artificial Intelligence

MemoHarness: How Agent Evolution Shifts to the External System

MemoHarness expands the notion of self‑evolving agents by keeping the language model frozen and continuously improving the surrounding control system—context assembly, tool interaction, generation settings, workflow orchestration, memory management, and output handling—demonstrating measurable gains on terminal, code‑generation, and finance tasks while highlighting limited experimental scale and selective cross‑task transfer.

Agent HarnessExperience LearningHarness Engineering
0 likes · 17 min read
MemoHarness: How Agent Evolution Shifts to the External System
Machine Heart
Machine Heart
Jul 23, 2026 · Artificial Intelligence

Teaching Agents to Evolve: The Hierarchical Skill Meta‑Evolving Framework HiSME

HiSME, a lightweight hierarchical skill meta‑evolution framework from Tsinghua and Huawei, enables LLM agents to accumulate execution experience without updating model parameters by evolving both task‑specific skills and the meta‑skills that generate and maintain them, improving performance on multi‑turn tool use and open‑world tasks.

HiSMELLM agentsMeta-Learning
0 likes · 11 min read
Teaching Agents to Evolve: The Hierarchical Skill Meta‑Evolving Framework HiSME
HyperAI Super Neural
HyperAI Super Neural
Jul 23, 2026 · Artificial Intelligence

ChemGraph: 13 Benchmarks Reveal LLM Agent’s Capabilities in Computational Chemistry

The Argonne National Laboratory team introduces ChemGraph, an LLM‑driven agent for computational chemistry, and evaluates it across 13 benchmark tasks, showing that small models excel on simple tasks while larger models and multi‑agent designs dramatically improve performance on complex molecular simulations.

AI automationBenchmarkChemGraph
0 likes · 11 min read
ChemGraph: 13 Benchmarks Reveal LLM Agent’s Capabilities in Computational Chemistry
PaperAgent
PaperAgent
Jul 18, 2026 · Artificial Intelligence

SkillOpt 2.0: A Leaner, Faster Self‑Evolving Agent

The article presents SkillOpt‑Lite, a stripped‑down self‑evolving agent pipeline that achieves lighter computation, faster convergence within the first few steps, and higher performance ceilings across multiple benchmarks, while exposing the underlying zero‑order optimization principles and validation requirements.

AgentBenchmarkLLM agents
0 likes · 10 min read
SkillOpt 2.0: A Leaner, Faster Self‑Evolving Agent
o-ai.tech
o-ai.tech
Jul 17, 2026 · Artificial Intelligence

When Does Trajectory Review Boost Agent Success? Five Key Factors Explained

Recent research shows that letting agents review their own execution trajectories can improve task success rates, but only under clear conditions such as reliable external feedback, concrete planning outputs, appropriate timing, sufficient model capability, and manageable cost‑benefit trade‑offs.

Agent EngineeringLLM agentsdynamic replanning
0 likes · 31 min read
When Does Trajectory Review Boost Agent Success? Five Key Factors Explained
Black & White Path
Black & White Path
Jul 17, 2026 · Information Security

A Complete AI Penetration Testing Landscape: 56 Open‑Source Agents, 73 Papers, and Key Models

The article surveys the emerging field of AI‑driven offensive security, cataloguing 56 open‑source penetration‑testing agents, 73 academic papers, six offensive models, benchmark suites, and DARPA AIxCC 2025 finalists, offering researchers and practitioners a consolidated view of tools, research trends, and evaluation frameworks.

AI modelsAI securityDARPA AIxCC
0 likes · 8 min read
A Complete AI Penetration Testing Landscape: 56 Open‑Source Agents, 73 Papers, and Key Models
Ray's Galactic Tech
Ray's Galactic Tech
Jul 16, 2026 · Artificial Intelligence

K8s, Kafka, Nacos Agent Platform to Prevent Token Bankruptcy and Skill Avalanches

The article details how a production‑grade Agent platform built on Kubernetes, Kafka, and Nacos addresses token budget overruns, uncontrolled skill execution, and RAG hallucinations by introducing a four‑layer runtime architecture, token pre‑allocation, explicit state management, dynamic governance policies, and robust skill specifications.

KafkaKubernetesLLM agents
0 likes · 32 min read
K8s, Kafka, Nacos Agent Platform to Prevent Token Bankruptcy and Skill Avalanches
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 15, 2026 · Artificial Intelligence

How SEAGym Tackles Evaluation Challenges of Self‑Evolving LLM Agents

The SEAGym benchmark reframes LLM agent evaluation from static success rates to dynamic harness evolution, offering multi‑view metrics, detailed snapshot diagnostics, and extensive experiments that reveal validation gains, OOD generalization gaps, batch‑size trade‑offs, and cross‑model transfer effects.

BenchmarkHarness EngineeringLLM agents
0 likes · 15 min read
How SEAGym Tackles Evaluation Challenges of Self‑Evolving LLM Agents
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 13, 2026 · Artificial Intelligence

LegalWorld: Building a Full‑Lifecycle Interactive Simulation World for Legal AI Agents

LegalWorld creates an interactive environment that simulates the entire Chinese civil litigation process—from consultation to second‑instance trial—using over 75,000 paired first‑ and second‑instance judgments, supports heterogeneous LLM agents, and provides the LongJud‑Bench benchmark to evaluate and improve legal AI performance.

LLM agentsLegal AILongJud-Bench
0 likes · 15 min read
LegalWorld: Building a Full‑Lifecycle Interactive Simulation World for Legal AI Agents
AI2ML AI to Machine Learning
AI2ML AI to Machine Learning
Jul 13, 2026 · Artificial Intelligence

From QA to Task‑Oriented Agents: Recent Trends in Large Language Models

The article surveys the latest advances in large language model agents, covering multi‑agent collaboration, long‑horizon planning, self‑evolution, trust and safety, test‑time scaling techniques, new foundation and multimodal models, open‑source and closed‑source breakthroughs, world‑model integration, and emerging vertical applications.

Foundation ModelsLLM agentsTest-Time Scaling
0 likes · 12 min read
From QA to Task‑Oriented Agents: Recent Trends in Large Language Models
Data Party THU
Data Party THU
Jul 13, 2026 · Artificial Intelligence

From QA to Task Completion: Survey of LLM Agent Systems and Harness Design

This survey argues that modern LLM agents should be viewed as a coupled system of a foundational model and an execution harness, analyzes the evolution from prompt engineering to harness engineering, defines six core harness responsibilities, examines task pressures, proposes richer evaluation metrics, and outlines future research directions.

Agent EngineeringExecution HarnessLLM agents
0 likes · 16 min read
From QA to Task Completion: Survey of LLM Agent Systems and Harness Design
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

Agentic Era: Shifting Recommendation from Platform-Centric to User-Governed

Recent research argues that the traditional platform‑centric recommendation paradigm is reaching its limits, proposing a user‑governed personalization model enabled by LLM agents that can aggregate cross‑platform data, with experimental evidence showing significant performance gains over platform‑only approaches.

Artificial IntelligenceLLM agentscross-platform data
0 likes · 20 min read
Agentic Era: Shifting Recommendation from Platform-Centric to User-Governed
AI Engineering
AI Engineering
Jul 11, 2026 · Artificial Intelligence

Why Logs Should Be the Agent Itself, Not Just a Byproduct

The article analyzes Yohei Nakajima’s "The Log is the Agent" paper, showing how ActiveGraph unifies goals, rules, tool calls, LLM responses, and artifacts into a single append‑only event log, enabling deterministic replay, cheap forking, and full provenance for LLM‑driven agents.

ActiveGraphAgent ArchitectureEvent Sourcing
0 likes · 13 min read
Why Logs Should Be the Agent Itself, Not Just a Byproduct
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

How AReaL 2.0 Accelerates Self‑Evolving Agents

AReaL 2.0 introduces an online reinforcement‑learning infrastructure that turns real‑world agent interactions into a learning loop, defining three pillars—trajectory data protocol, data proxy, and evolution control plane—to enable agents to not only execute tasks but continuously improve from their own experience.

AReaLAgentic RLLLM agents
0 likes · 16 min read
How AReaL 2.0 Accelerates Self‑Evolving Agents
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 25, 2026 · Artificial Intelligence

AutoResearch Advances: RUC & Microsoft Open‑Source Arbor Gives Agents Research Memory

Arbor, an open‑source autonomous research framework from RUC’s Gaoling AI Institute and Microsoft Research, structures the research loop with a growing hypothesis‑tree and insight back‑propagation, allowing agents to retain hypotheses, evidence, and failures, and achieves the best held‑out results on six real AO tasks, surpassing Codex and Claude Code.

AI research automationArbor frameworkLLM agents
0 likes · 18 min read
AutoResearch Advances: RUC & Microsoft Open‑Source Arbor Gives Agents Research Memory
Architect
Architect
Jun 25, 2026 · Artificial Intelligence

Why a Concise CLAUDE.md Entry File Is Critical for LLM Agents in Your Repo

The article explains how a short, well‑structured CLAUDE.md file injects the minimal yet essential context an LLM coding agent needs before it scans a repository, preventing common mis‑assumptions about tech stack, commands, boundaries, and completion criteria.

AGENTS.mdAI toolingCLAUDE.md
0 likes · 16 min read
Why a Concise CLAUDE.md Entry File Is Critical for LLM Agents in Your Repo
PaperAgent
PaperAgent
Jun 20, 2026 · Artificial Intelligence

Vertical Domain Agents Gain 88.5% Boost by Adapting the Runtime Interface, Not Retraining

The paper shows that many failures of deterministic LLM agents stem from mismatched model‑environment interfaces, and introduces LIFE‑HARNESS—a four‑layer runtime harness that extracts reusable failure patterns from training trajectories without updating model weights, delivering an average 88.5% relative performance gain across 126 model‑environment settings.

Deterministic AgentsLLM agentsLife-Harness
0 likes · 8 min read
Vertical Domain Agents Gain 88.5% Boost by Adapting the Runtime Interface, Not Retraining
Machine Heart
Machine Heart
Jun 19, 2026 · Artificial Intelligence

Which Multi‑Agent Communication Protocol Wins? UIUC Introduces ProtocolBench at ICML 2026

The UIUC team presents ProtocolBench, a systematic benchmark that compares four multi‑agent communication protocols across four realistic scenarios, revealing distinct trade‑offs in latency, reliability, and security, and proposes ProtocolRouter to automatically select the most suitable protocol per workload.

BenchmarkLLM agentsProtocolBench
0 likes · 14 min read
Which Multi‑Agent Communication Protocol Wins? UIUC Introduces ProtocolBench at ICML 2026
Code Mala Tang
Code Mala Tang
Jun 19, 2026 · Artificial Intelligence

Five Skeptical Questions About RTK’s Token Compression Claims

The article critically examines RTK’s token‑compression promises, exposing misleading savings metrics, silent‑failure bugs, missing task‑success benchmarks, its status as a fragile feature rather than a product, and the brittleness of its output parser, before offering concrete guidance on when to use it.

CLI output parsingLLM agentsRTK
0 likes · 8 min read
Five Skeptical Questions About RTK’s Token Compression Claims
HyperAI Super Neural
HyperAI Super Neural
Jun 18, 2026 · Artificial Intelligence

Weekly AI Paper Digest: D4RT 300× Faster 4D Reconstruction, SAI Theory Challenges AGI, and More

This week’s AI paper roundup covers DeepMind’s D4RT framework that accelerates dynamic 4D reconstruction by up to 300×, a Columbia‑NYU proposal of Superhuman Adaptable Intelligence that questions AGI, MIT‑UW findings on chatbot delusional spiraling, security risks of autonomous agents, a new ARA protocol for executable research artifacts, a vision of AI‑driven software engineering, and a memory‑caching approach that expands RNN capacity while reducing complexity.

AI safetyArtificial IntelligenceD4RT
0 likes · 11 min read
Weekly AI Paper Digest: D4RT 300× Faster 4D Reconstruction, SAI Theory Challenges AGI, and More
Machine Heart
Machine Heart
Jun 17, 2026 · Artificial Intelligence

Why RL‑Trained Agents Still Fail to Reason Actively: The Information Self‑Locking Problem

The paper reveals that outcome‑based reinforcement learning often traps LLM agents in an information self‑locking regime where weak action selection and belief tracking prevent proper credit assignment, and introduces AREW, a lightweight advantage‑reweighting method that restores active reasoning across multiple tasks and models.

AREWAgentic RLLLM agents
0 likes · 24 min read
Why RL‑Trained Agents Still Fail to Reason Actively: The Information Self‑Locking Problem
PaperAgent
PaperAgent
Jun 17, 2026 · Artificial Intelligence

Spatial-Agent: A New Concept‑Transformation Paradigm for Map Agents

The paper introduces Spatial‑Agent, which models geospatial question answering as a concept‑transformation process using a GeoFlow Graph intermediate representation, outlines a five‑step workflow, defines core concepts and functional roles, and demonstrates its effectiveness on MapEval‑API and MapQA benchmarks with detailed error and cost analyses.

BenchmarkGISGeoFlow Graph
0 likes · 13 min read
Spatial-Agent: A New Concept‑Transformation Paradigm for Map Agents
Machine Heart
Machine Heart
Jun 15, 2026 · Artificial Intelligence

Breaking the SWE‑bench Score‑Only Myth: Open‑Source Benchmark that Independently Measures Harnesses

The article critiques the reliance on raw SWE‑bench scores for programming agents, introduces the Claw‑SWE‑Bench benchmark and a dedicated adapter that isolates harness effects, and presents extensive experiments showing how model choice, harness design, and cost impact real-world coding performance across multiple languages.

BenchmarkHarnessLLM agents
0 likes · 14 min read
Breaking the SWE‑bench Score‑Only Myth: Open‑Source Benchmark that Independently Measures Harnesses
Java Tech Enthusiast
Java Tech Enthusiast
Jun 13, 2026 · Artificial Intelligence

Why Bigger 1M‑Token Windows Still Need Careful Context Engineering

Even though modern LLMs like DeepSeek‑V4, GPT‑5.5 and Claude Opus 4.7 support 1 million‑token windows, simply stuffing more data does not improve agent performance; effective Context Engineering—selecting, structuring, and managing the right information—remains essential for reliable results.

LLM agentsMemoryRAG
0 likes · 32 min read
Why Bigger 1M‑Token Windows Still Need Careful Context Engineering
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 12, 2026 · Artificial Intelligence

The Next Frontier for Large‑Scale LLM Agents: 17 Must‑Read Papers on Self‑Evolving Harnesses

This article surveys 17 recent core papers that explore how the system‑level harness surrounding large‑model agents can be automatically generated, evolved, and audited, covering topics such as system boundaries, failure‑driven improvement, memory and skill optimization, source‑level rewriting, scaling laws, aging, and safety.

Agent MemoryHarness EngineeringLLM agents
0 likes · 18 min read
The Next Frontier for Large‑Scale LLM Agents: 17 Must‑Read Papers on Self‑Evolving Harnesses
Data Party THU
Data Party THU
Jun 11, 2026 · Artificial Intelligence

Boost 18 LLM Agents Without Retraining Using LIFE‑HARNESS

The article introduces LIFE‑HARNESS, a runtime‑interface adaptation framework that keeps model weights unchanged, extracts reusable failure patterns from a single model's training trace, and achieves an average 88.5% relative performance gain across 18 LLM agents and 7 deterministic environments, with successful transfer to 17 other models.

Cross-Model TransferLLM agentsRuntime Harness
0 likes · 8 min read
Boost 18 LLM Agents Without Retraining Using LIFE‑HARNESS
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Jun 11, 2026 · Artificial Intelligence

Scaling Automated Formalization of Mathematics: Inside Meta’s AutoformBot and the ATLAS Lean 4 Library

Meta’s recent paper presents AutoformBot, a multi‑agent system that treats formalizing entire mathematics textbooks as a large‑scale software‑engineering project, generating the ATLAS Lean 4 library with over 45,000 declarations and demonstrating a 71 % success rate across 26 open‑access books.

AutoformBotFormal VerificationLLM agents
0 likes · 14 min read
Scaling Automated Formalization of Mathematics: Inside Meta’s AutoformBot and the ATLAS Lean 4 Library
Xike
Xike
Jun 9, 2026 · Artificial Intelligence

From Demo to Production: Key Practices for Engineering LLM Agents

The article explains how to transform a prototype LLM agent into a reliable production service by defining service contracts, externalizing configuration, handling session state, implementing multi‑level rate limiting, and integrating observability, deployment, and rollback mechanisms to avoid common engineering pitfalls.

ConfigurationLLM agentsdeployment
0 likes · 25 min read
From Demo to Production: Key Practices for Engineering LLM Agents
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 8, 2026 · Artificial Intelligence

Re‑evaluating the Token World of LLM Agents: A Dual‑View Economics Overview

The paper surveys the rapid growth of token consumption in LLM agents, proposes a dual‑view Token Economics framework that treats tokens as production factors, exchange media, and accounting units, and classifies optimization challenges from single‑agent efficiency to ecosystem‑level pricing, security, and future research directions.

AI Resource ManagementCost OptimizationLLM agents
0 likes · 10 min read
Re‑evaluating the Token World of LLM Agents: A Dual‑View Economics Overview
Alibaba Cloud Native
Alibaba Cloud Native
Jun 8, 2026 · Artificial Intelligence

Code Harness vs. Model-Driven Harness: Can Agent Control Be Expressed as Executable Natural Language?

The article reviews the "Natural-Language Agent Harnesses" paper, explains the distinction between code, middleware, and harness layers for LLM agents, introduces NLAH and IHR concepts, and details experimental evaluations that show natural‑language harnesses can match code‑based control while exposing new trade‑offs and risks.

Intelligent Harness RuntimeLLM agentsModule Ablation
0 likes · 13 min read
Code Harness vs. Model-Driven Harness: Can Agent Control Be Expressed as Executable Natural Language?
TechVision Expert Circle
TechVision Expert Circle
Jun 2, 2026 · Artificial Intelligence

How to Build a System with True Self‑Decision Capability

This article details a CTO‑driven initiative to give core business systems genuine self‑decision ability by defining the concept, presenting a four‑layer architecture, and sharing concrete design choices, safety mechanisms, governance practices, and real‑world lessons learned from an e‑commerce fulfillment use case.

KafkaLLM agentsMulti-Agent Architecture
0 likes · 14 min read
How to Build a System with True Self‑Decision Capability
Xike
Xike
Jun 2, 2026 · Artificial Intelligence

Why Agent Conversations Aren’t Just Chat Logs: Effective Context Management

The article explains that an Agent’s context is a structured snapshot built from role contracts, tool trajectories, and window budgeting, not a raw chat transcript, and details how proper context handling prevents forgetting, token bloat, and tool‑call mismatches in multi‑turn LLM workflows.

Context ManagementLLM agentsReAct
0 likes · 16 min read
Why Agent Conversations Aren’t Just Chat Logs: Effective Context Management
James' Growth Diary
James' Growth Diary
Jun 1, 2026 · Artificial Intelligence

How Hermes Implements Bounded Memory: Character Limits, Compression, and Snapshots to Prevent Overflow

The article details Hermes' bounded memory system, which uses character limits for persistent files, a three‑stage context compression pipeline, boundary alignment to protect tool calls, snapshot caching, triple redaction, and anti‑thrashing mechanisms, ensuring agents never overflow or lose critical information.

HermesLLM agentsbounded memory
0 likes · 16 min read
How Hermes Implements Bounded Memory: Character Limits, Compression, and Snapshots to Prevent Overflow
DataFunTalk
DataFunTalk
Jun 1, 2026 · Artificial Intelligence

Rethinking Agent Harness: Toward State‑Aware Runtime for Reliable LLM Agents

The article argues that improving large‑model agents requires more than bigger models or longer context windows; it calls for a stable, auditable, and recoverable runtime that manages state transitions, prevents error propagation, and enables trace‑native evaluation of long‑running agents.

Agent HarnessLLM agentsRuntime Engineering
0 likes · 13 min read
Rethinking Agent Harness: Toward State‑Aware Runtime for Reliable LLM Agents
ITPUB
ITPUB
May 30, 2026 · Artificial Intelligence

Is RAG Dead? How Grep Is Making a Comeback in LLM‑Powered Code Search

This article investigates the claim that Retrieval‑Augmented Generation (RAG) is obsolete by dissecting Claude Code’s grep‑driven search architecture, benchmarking its performance against traditional vector‑based retrieval, comparing it with Cursor and OpenAI Codex, and analyzing the trade‑offs of multi‑round agentic search.

Claude CodeCursorLLM agents
0 likes · 36 min read
Is RAG Dead? How Grep Is Making a Comeback in LLM‑Powered Code Search
Linyb Geek Road
Linyb Geek Road
May 29, 2026 · Artificial Intelligence

Agent Harness Architecture Deep Dive: From ReAct Loop to Production‑Grade AI System Design

The article argues that the real performance bottleneck of AI agents lies in the Agent Harness infrastructure rather than the model itself, and it systematically explains how prompt, context, and infrastructure layers, tool handling, memory, verification, error handling, and design trade‑offs shape production‑ready LLM agents.

AI infrastructureAgent HarnessContext Management
0 likes · 24 min read
Agent Harness Architecture Deep Dive: From ReAct Loop to Production‑Grade AI System Design
Code Mala Tang
Code Mala Tang
May 28, 2026 · Artificial Intelligence

When Claude Skills Need Determinism, Use Skillflows

The article analyzes Claude's natural‑language SKILL.md approach, highlights its flexibility and nondeterminism, and explains how adding a declarative skillflow.json graph enforces deterministic execution, auditability, lower token cost, and better consistency for high‑frequency, compliance‑critical tasks.

ClaudeCost OptimizationDeterminism
0 likes · 11 min read
When Claude Skills Need Determinism, Use Skillflows
ShiZhen AI
ShiZhen AI
May 27, 2026 · Artificial Intelligence

Turning Click‑Based Web Agents into Repeatable Scripts with Microsoft’s Open‑Source Webwright

Microsoft’s open‑source Webwright framework redefines browser agents by replacing step‑by‑step click actions with generated Playwright scripts, enabling repeatable, debuggable web tasks; the article details its architecture, workflow, benchmark results on Online‑Mind2Web and Odysseys, and discusses practical benefits and limitations.

BenchmarkGPT-5.4LLM agents
0 likes · 9 min read
Turning Click‑Based Web Agents into Repeatable Scripts with Microsoft’s Open‑Source Webwright
AI Engineering
AI Engineering
May 26, 2026 · Artificial Intelligence

Training Only the Skill Document While Keeping Model Weights Frozen (SkillOpt)

Microsoft Research introduces SkillOpt, a method that freezes large‑model weights and instead trains a natural‑language skill document as the sole learnable parameter, using a rollout‑reflect‑edit‑gate loop, achieving optimal results across 52 benchmark‑model‑environment combinations and demonstrating strong transferability.

LLM agentsSkillOptbenchmark evaluation
0 likes · 9 min read
Training Only the Skill Document While Keeping Model Weights Frozen (SkillOpt)
Code Mala Tang
Code Mala Tang
May 25, 2026 · Artificial Intelligence

Behind 95K Stars: browser-use’s LLM Browser Automation vs Playwright

browser-use, an open‑source MIT‑licensed LLM agent loop that compresses page DOM into an indexed list of interactive elements, lets large language models plan and execute web tasks, and is compared against Anthropic’s Computer Use, OpenAI’s Operator and traditional Playwright/Selenium, highlighting its flexibility, lower cost, but higher LLM usage and deployment trade‑offs.

Anthropic Computer UseLLM agentsMIT license
0 likes · 16 min read
Behind 95K Stars: browser-use’s LLM Browser Automation vs Playwright
HyperAI Super Neural
HyperAI Super Neural
May 25, 2026 · Artificial Intelligence

CVEvolve: Zero‑Code Autonomous Discovery of Scientific Image‑Processing Algorithms

CVEvolve, a no‑code autonomous agent framework from ANL, leverages large‑language‑model agents to discover, evaluate, and iterate scientific image‑processing algorithms without any programming, and demonstrates superior performance on X‑ray fluorescence registration, Bragg‑peak detection, and diffraction‑image segmentation compared with traditional baselines.

CVEvolveLLM agentsautonomous discovery
0 likes · 13 min read
CVEvolve: Zero‑Code Autonomous Discovery of Scientific Image‑Processing Algorithms
PaperAgent
PaperAgent
May 25, 2026 · Artificial Intelligence

DeepSeek’s Harness: How Agent Harness Engineering Is Shaping the Next LLM Agent Era

The article surveys DeepSeek’s Harness initiative, presenting the Binding‑Constraint Thesis, three‑stage evolution from prompt to harness engineering, the ETCLOVG seven‑layer architecture, and concrete benchmark evidence that harness‑only improvements far outweigh model upgrades, while detailing security, observability, and governance considerations for reliable LLM agents.

AI architectureAgent EvaluationAgent Harness Engineering
0 likes · 12 min read
DeepSeek’s Harness: How Agent Harness Engineering Is Shaping the Next LLM Agent Era
James' Growth Diary
James' Growth Diary
May 24, 2026 · Artificial Intelligence

Execution → Observation → Reflection → Improvement: How Hermes Closes the Skill Loop

The article dissects Hermes' background review mechanism, showing how a silent daemon thread performs post‑conversation reflection, writes valuable insights to a skill or memory store, shares prompt designs, fork‑agent isolation, priority update rules, and common pitfalls for building continuously learning LLM agents.

Background ReviewDaemon ThreadHermes
0 likes · 14 min read
Execution → Observation → Reflection → Improvement: How Hermes Closes the Skill Loop
Old Zhang's AI Learning
Old Zhang's AI Learning
May 21, 2026 · Artificial Intelligence

SkillOS: Enabling Agents to Self‑Manage Their Skills

SkillOS reframes skill management for LLM agents as a long‑horizon reinforcement‑learning problem, letting a trainable Skill Curator automatically insert, update, or delete markdown‑based skills, which the frozen Agent Executor then consumes, improving memory‑free performance and cross‑task transfer.

LLM agentsReinforcement LearningSelf-Evolving Agents
0 likes · 6 min read
SkillOS: Enabling Agents to Self‑Manage Their Skills
Data Party THU
Data Party THU
May 18, 2026 · Artificial Intelligence

How VIGIL’s Verify‑Before‑Execute Paradigm Defeats LLM Agent Tool Hijacking

VIGIL introduces a verify‑before‑commit framework that isolates tool‑stream injection attacks on LLM agents, using intent anchoring, perception sanitization, speculative reasoning, grounding verification, and validated trajectory memory, reducing attack success rates to 8‑12% while preserving task utility.

AI safetyLLM agentsSIREN benchmark
0 likes · 11 min read
How VIGIL’s Verify‑Before‑Execute Paradigm Defeats LLM Agent Tool Hijacking
LuTiao Programming
LuTiao Programming
May 17, 2026 · Artificial Intelligence

Why Your AI Keeps Going Off‑Track: The 4 Essential CLAUDE.md Directives

The article analyzes why AI coding assistants often stray from intended requirements, exposing a core judgment deficit, and shows how a concise four‑line CLAUDE.md file—detailing assumptions, minimal code, scoped changes, and verifiable success criteria—can dramatically improve AI behavior, reduce over‑design, and lower review costs.

AI codingCLAUDE.mdCode Review
0 likes · 11 min read
Why Your AI Keeps Going Off‑Track: The 4 Essential CLAUDE.md Directives
dbaplus Community
dbaplus Community
May 17, 2026 · Artificial Intelligence

Why Grep Is Replacing Vector Indexes: RAG Isn’t Dead, It’s Evolving

The article dissects Claude Code’s LLM‑driven Grep search, showing how multi‑round tool calls replace static vector‑based RAG, presents ripgrep performance benchmarks, compares Claude Code with Cursor and Codex, and argues that zero‑index search is optimal for local code bases while larger projects still need indexing.

Claude CodeLLM agentsRAG
0 likes · 36 min read
Why Grep Is Replacing Vector Indexes: RAG Isn’t Dead, It’s Evolving
PaperAgent
PaperAgent
May 11, 2026 · Artificial Intelligence

SkillOS: How Skill Governance Powers Self‑Evolving AI Agents

SkillOS addresses the one‑off nature of current LLM agents by introducing a closed‑loop system where a trainable Skill Curator continuously extracts, updates, and manages reusable skills from execution traces, leading to measurable gains in success rates, efficiency, and cross‑task generalization.

Grouped Task StreamsLLM agentsMeta-Strategy Skills
0 likes · 10 min read
SkillOS: How Skill Governance Powers Self‑Evolving AI Agents
Linyb Geek Road
Linyb Geek Road
May 10, 2026 · Artificial Intelligence

Designing Progressive Large‑Model Agents: Architecture, Frameworks, and Real‑World Practices

This article examines the evolution of large‑model agents, outlines four development stages, compares workflow, collaborative, and evolutionary frameworks, details core components such as perception, memory, planning, tools, and reflection, and explains how a progressive, loop‑based architecture can be applied across verticals like research, code generation, and complex workflow automation.

Agent ArchitectureAlphaEvolveLLM agents
0 likes · 9 min read
Designing Progressive Large‑Model Agents: Architecture, Frameworks, and Real‑World Practices
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
May 8, 2026 · Artificial Intelligence

T²PO: Uncertainty‑Guided Exploration Control for Stable Multi‑Turn Agent RL

The paper identifies inefficient exploration, termed "hesitation," as the root cause of instability in multi‑turn reinforcement learning for LLM agents and introduces T²PO, an uncertainty‑driven token‑ and turn‑level control framework that markedly improves training stability and performance on benchmarks such as WebShop, ALFWorld, and Search QA.

LLM agentsT2POexploration control
0 likes · 16 min read
T²PO: Uncertainty‑Guided Exploration Control for Stable Multi‑Turn Agent RL
PaperAgent
PaperAgent
May 4, 2026 · Artificial Intelligence

A Comprehensive Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution

This survey systematically reviews self‑evolving agents, explains why autonomous agents are needed, proposes a unified taxonomy of three evolution paradigms, analyzes model‑centric, environment‑centric, and co‑evolution approaches, and outlines future challenges in designing adaptive environments.

AI Agent TaxonomyCo-EvolutionEnvironment-Centric Evolution
0 likes · 14 min read
A Comprehensive Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution
AI Tech Publishing
AI Tech Publishing
May 1, 2026 · Artificial Intelligence

5 Counterintuitive Design Principles for Prompt Caching in Claude Code

The article details five counterintuitive design principles for Claude Code's prompt caching—optimizing prompt layout, using message‑based updates, never switching models or tools mid‑conversation, safely compressing context, and monitoring cache health—backed by concrete examples and up to 90% cost savings.

AI EngineeringClaude CodeLLM agents
0 likes · 10 min read
5 Counterintuitive Design Principles for Prompt Caching in Claude Code
AI Explorer
AI Explorer
Apr 30, 2026 · Industry Insights

AI Tech Daily: Key AI Industry Highlights for April 30 2026

The AI Tech Daily roundup highlights Microsoft's 123% AI revenue surge, groundbreaking GPT‑5.5 restrictions, DeepSeek's multimodal launch, Ant Group's zkDTVM benchmark record, a 23‑year‑old Linux kernel bug, Stripe's 288 AI‑focused features, and emerging trends in LLM agent orchestration and AI adoption metrics.

AI revenueDeepSeekGPT-5.5
0 likes · 4 min read
AI Tech Daily: Key AI Industry Highlights for April 30 2026
SuanNi
SuanNi
Apr 27, 2026 · Artificial Intelligence

How MIT’s RUBICON Cuts AI Agent Costs by 90% While Achieving 100% Accuracy

The paper shows that conventional LLM agents fail on real‑world enterprise data because of chaotic data sources, while the RUBICON architecture uses a minimal Agentic Query Language to let users direct data retrieval, achieving 100% accuracy with a much cheaper model and dramatically lower token and monetary costs.

Agentic Query LanguageBenchmarkData Integration
0 likes · 11 min read
How MIT’s RUBICON Cuts AI Agent Costs by 90% While Achieving 100% Accuracy
DeepNoMind
DeepNoMind
Apr 27, 2026 · Artificial Intelligence

Let AI Write Its Own “Employee Handbook”: The ACE Paradigm Battle

The ACE paper proposes a novel Agentic Context Engineering framework that lets LLM agents continuously improve by treating context as a self‑learning playbook, achieving large accuracy gains on the AppWorld benchmark while cutting latency and cost, and it critically compares ACE to fine‑tuning, RAG, and prompt‑optimization approaches.

ACEAgentic Context EngineeringAppWorld benchmark
0 likes · 38 min read
Let AI Write Its Own “Employee Handbook”: The ACE Paradigm Battle
AI Architecture Hub
AI Architecture Hub
Apr 23, 2026 · Artificial Intelligence

Why Prompt Caching Is Critical: Lessons from Building Claude Code

Prompt caching, a prefix‑matching technique that reuses prior LLM interactions, proved essential for Claude Code’s low latency and cost, and the article details counter‑intuitive practices such as arranging static prompts first, updating info via messages, avoiding mid‑session model or tool changes, and ensuring cache‑safe context forks.

AI EngineeringClaude CodeLLM agents
0 likes · 10 min read
Why Prompt Caching Is Critical: Lessons from Building Claude Code
AI Waka
AI Waka
Apr 22, 2026 · Artificial Intelligence

Hybrid MCP‑Skill Model: Keeping LLM Agent Skills Fresh

The article analyzes the trade‑offs between packaging new agent functionality as a static Skill versus a dynamic MCP server, proposes a hybrid thin‑CLI approach that combines the ease of Skills with the up‑to‑date guarantees of MCP, and illustrates the design with concrete code examples.

API VersioningCLI wrapperHybrid Architecture
0 likes · 7 min read
Hybrid MCP‑Skill Model: Keeping LLM Agent Skills Fresh
PaperAgent
PaperAgent
Apr 22, 2026 · Artificial Intelligence

How SkillClaw Enables Collective Evolution of Agent Skills in Real-World Use

SkillClaw introduces a centralized evolution framework that transforms user interactions into structured evidence, allowing LLM agents to refine, create, or skip skills based on aggregated success and failure patterns, with nightly validation ensuring only proven improvements are deployed, resulting in consistent performance gains across diverse tasks.

AI workflowBenchmarkLLM agents
0 likes · 13 min read
How SkillClaw Enables Collective Evolution of Agent Skills in Real-World Use
AntTech
AntTech
Apr 22, 2026 · Artificial Intelligence

How Multi‑Agent MCTS and Information‑Gain Rewards Are Transforming Mobile GUI and Search Agents

This article reviews two recent ICLR 2026 papers—M²‑Miner, a multi‑agent Monte‑Carlo Tree Search framework for low‑cost mobile GUI data mining, and IGPO, an information‑gain‑based reinforcement‑learning method that provides dense rewards for multi‑turn search agents—detailing their designs, experiments, and open‑source releases.

GUI Data MiningInformation GainLLM agents
0 likes · 8 min read
How Multi‑Agent MCTS and Information‑Gain Rewards Are Transforming Mobile GUI and Search Agents