Tagged articles

dynamic planning

6 articles · Page 1 of 1
Architect
Architect
Sep 18, 2026 · Artificial Intelligence

What Fermat's Last Theorem Formalization Reveals About Multi-Agent Collaboration

Anthropic's 11-day project formalizing Fermat's Last Theorem in Lean with 30,000 machine-checked theorems exposes five critical patterns for multi-agent systems: verifiable artifacts, dynamic task graphs, evidence-based planning, verification-gated state changes, and recoverable execution state.

Fermat's Last TheoremFormal VerificationLean theorem prover
0 likes · 26 min read
What Fermat's Last Theorem Formalization Reveals About Multi-Agent Collaboration
Alibaba Middleware
Alibaba Middleware
Aug 12, 2026 · Operations

How STAROps Detects Unknown Anomalies with Intelligent Log Inspection

STAROps transforms raw logs into actionable insights by clustering log patterns, drilling down across dimensions with AI operators, and using an Agent that dynamically plans investigations, integrates UModel cross‑source mapping, and continuously refines findings to catch unknown anomalies before they become incidents.

AI operatorsObservabilityUModel
0 likes · 17 min read
How STAROps Detects Unknown Anomalies with Intelligent Log Inspection
JD Cloud Developers
JD Cloud Developers
Aug 4, 2026 · Artificial Intelligence

NaviAgent: Scalable Tool Orchestration for Oxygen Agents via Graph‑Driven Bilevel Planning

The paper introduces NaviAgent, a double‑layer architecture that separates LLM‑based planning from graph‑driven tool navigation, explicitly models API‑parameter dependencies, continuously updates the tool graph with execution feedback, and achieves up to 13.1 % higher task success rates on large‑scale API benchmarks.

AI AgentsGraph ModelingLLM
0 likes · 16 min read
NaviAgent: Scalable Tool Orchestration for Oxygen Agents via Graph‑Driven Bilevel Planning
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jun 23, 2026 · Artificial Intelligence

How AI Agents Can Slash Server‑Side End‑to‑End Test Costs

The article analyzes why server‑side end‑to‑end testing has become costly due to cross‑domain complexity, long chains, and combinatorial explosion, and demonstrates how an AI‑driven agent with dynamic planning, progressive knowledge‑base loading, and a debug‑first script generation loop can dramatically reduce automation effort from days to minutes.

AI AgentDebug-firstdynamic planning
0 likes · 17 min read
How AI Agents Can Slash Server‑Side End‑to‑End Test Costs
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Apr 3, 2026 · Artificial Intelligence

How Alibaba Cloud’s Ops‑Agentic‑Search Reached Human‑Level Performance on the GAIA Benchmark

Alibaba Cloud’s AI Search team introduces Ops‑Agentic‑Search, an enterprise‑grade AI agent framework that tackles core challenges of hallucination, task failure, and long‑term consistency, leverages the GAIA benchmark to demonstrate a 92.36% accuracy—matching human experts—and outlines its technical architecture, key mechanisms, use cases, and future open‑source contributions.

GAIA benchmarkOpenSearchdynamic planning
0 likes · 11 min read
How Alibaba Cloud’s Ops‑Agentic‑Search Reached Human‑Level Performance on the GAIA Benchmark
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Apr 2, 2026 · Artificial Intelligence

How Alibaba Cloud’s Ops‑Agentic‑Search Reached Human‑Level Performance on the GAIA Benchmark

The article explains the shift of AI agents from passive responders to proactive executors, outlines the challenges of hallucination, task failure, and consistency, introduces the GAIA benchmark, and details how Alibaba Cloud's Ops‑Agentic‑Search achieved a 92.36% accuracy—matching human experts—through global planning, reflection, dynamic context management, and a self‑evolving skills system.

AI AgentGAIA benchmarkOps-Agentic-Search
0 likes · 12 min read
How Alibaba Cloud’s Ops‑Agentic‑Search Reached Human‑Level Performance on the GAIA Benchmark