Tagged articles

Self-Evolving Systems

10 articles · Page 1 of 1
PaperAgent
PaperAgent
Sep 11, 2026 · Artificial Intelligence

Google's Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Google introduces Procedural Graphs, a self-evolving graph structure that explicitly encodes procedural knowledge for LLM agents, enabling them to locate, extract, and generate contextual guidance from a (procedure, relation, procedure) triplet graph, achieving 85% survival from 0% in CFO simulation and winning 21 of 24 model-benchmark combinations across 7 benchmarks and 4 LLMs.

Agent ReliabilityGraph-based ReasoningLLM Agents
0 likes · 9 min read
Google's Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
TonyBai
TonyBai
Sep 9, 2026 · Artificial Intelligence

YC Debunks 'Model-Only' Myth: Harness, Not Model, Sets Agent Ceiling

YC Paper Club reveals how the same Claude Opus model scores 30% on ARC-AGI bare but reaches 95% with proper Harness, and NVIDIA's AVO hits 100%, proving agent runtime—not model weights—determines the performance ceiling.

ARC-AGIAgent RuntimeContinual Harness
0 likes · 20 min read
YC Debunks 'Model-Only' Myth: Harness, Not Model, Sets Agent Ceiling
Software Engineering 3.0 Era
Software Engineering 3.0 Era
Sep 7, 2026 · Artificial Intelligence

AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems

This comprehensive guide synthesizes 2026 industry practices from Xiaohongshu and Alipay to build production-ready AI Agent evaluation systems, covering metrics (Quality/Cost/Safety), three-tier evaluation granularities, Judge system design, OpenTelemetry-based observability, platform architecture with contract-driven test generation, dual flywheel offline/online loops, and self-evolving prompt optimization — moving evaluation from post-hoc verification to embedded engineering guardrails.

AI Agent EvaluationAgentOpsEvaluation Methodology
0 likes · 37 min read
AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems
Amap Tech
Amap Tech
Aug 14, 2026 · Artificial Intelligence

How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery

The article explains why a single successful AI‑generated code run is insufficient for enterprise software, and how AutoSDK uses built‑in observability, Loop Engineering, and a four‑stage "observe‑attribute‑intervene‑validate" loop—supported by concrete metrics, trace and log pillars—to achieve stable, continuously improving AI coding delivery.

AI codingLoop EngineeringObservability
0 likes · 17 min read
How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery
DataFunTalk
DataFunTalk
Jul 16, 2026 · Artificial Intelligence

Mastering Enterprise Agents: Protocols, Constraints, Self‑Evolution, and Cost

The live discussion reveals that stronger models hide subtle errors, shifting from chatbots to agents requires a cognitive upgrade, multi‑agent collaboration hinges on clear contracts, physical permissions trump prompts, and a three‑layer Rule‑Skill‑Hook framework plus careful handling of long context and self‑evolution are essential for reliable, cost‑effective enterprise AI deployment.

AI AgentsConstraint engineeringSelf-Evolving Systems
0 likes · 17 min read
Mastering Enterprise Agents: Protocols, Constraints, Self‑Evolution, and Cost
DevOps Cloud Academy
DevOps Cloud Academy
Jul 7, 2026 · Industry Insights

Agentic AI Enters Its Golden Era: How Intelligent Systems Are Reshaping Productivity

The article argues that the coming years will be a golden period for Agentic AI as intelligent agents evolve into an AI operating system that can decompose tasks, coordinate multiple agents, and fundamentally transform enterprise productivity, supported by emerging token economics, ontology‑driven infrastructure, and predictions from Gartner and industry leaders.

AI operating systemAgentic AIGartner
0 likes · 14 min read
Agentic AI Enters Its Golden Era: How Intelligent Systems Are Reshaping Productivity
ThinkingAgent
ThinkingAgent
Jun 18, 2026 · Artificial Intelligence

Evolving AI Skills Yield 116% Accuracy Boost – From Handwritten Prompts to Autonomous Creation

The 2026 Memento‑Skills system lets AI agents create and refine their own Skills, boosting General AI Assistants accuracy by 26.2% and Humanity's Last Exam accuracy by 116.2%, while outlining three generations of Skill evolution, self‑evolution frameworks, benchmark results, practical applications, and best‑practice guidelines for technical leaders.

AI AgentsBenchmarkingSelf-Evolving Systems
0 likes · 17 min read
Evolving AI Skills Yield 116% Accuracy Boost – From Handwritten Prompts to Autonomous Creation
Alibaba Cloud Developer
Alibaba Cloud Developer
May 22, 2026 · Artificial Intelligence

How Core Agent Concepts and Paradigms Have Evolved and the Rationale Behind Them

The article traces the evolution of AI agents from early ReAct‑style models through workflow‑based systems to autonomous and self‑evolving agents, analyzing six core dimensions—Prompt, Planning, Memory, Tools, Workflow, and Environment—and explains why each paradigm shift occurred, citing recent frameworks and research.

AI AgentsMemory ManagementPrompt Engineering
0 likes · 25 min read
How Core Agent Concepts and Paradigms Have Evolved and the Rationale Behind Them
Data Party THU
Data Party THU
Apr 28, 2026 · Artificial Intelligence

How MiniMax Drives Joint Evolution of Models and Harnesses

The article analyzes MiniMax’s strategy of co‑evolving large language models with a Harness framework, contrasting product philosophies, detailing a live MaxHermes demo that creates and refines reusable Skills, and explaining how this dual evolution reshapes the competitive focus from single‑turn Q&A to sustained, self‑improving agent workflows.

AI AgentsHermesMiniMax
0 likes · 14 min read
How MiniMax Drives Joint Evolution of Models and Harnesses
Tech Verticals & Horizontals
Tech Verticals & Horizontals
Jan 10, 2026 · Artificial Intelligence

Intelligent Agent System Levels 0‑4: From Core Reasoning to Self‑Evolving Agents

The article outlines a five‑tier taxonomy of intelligent agents—from a standalone language‑model reasoning engine lacking real‑time perception, through tool‑enabled problem solvers, context‑engineered planners, collaborative multi‑agent teams, up to self‑evolving systems that can create new tools or agents to fill capability gaps.

Multi-Agent CollaborationSelf-Evolving Systemsagent architecture
0 likes · 9 min read
Intelligent Agent System Levels 0‑4: From Core Reasoning to Self‑Evolving Agents