Alibaba Cloud Agentic Search: From Finding Answers to Completing Tasks

At the 2026 Yunqi Conference, Alibaba Cloud unveiled Agentic Search 2.0, a new AI search paradigm that evolves from answer generation to autonomous task execution via planning, tool use, and self-evolving memory, backed by a re-architected Elasticsearch engine delivering 60ms hot-query latency and 70% cost savings at hundred-billion-vector scale.

Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Agentic Search: From Finding Answers to Completing Tasks

Conference Context and Core Thesis

On September 24, 2026, at the Yunqi Conference AI Search sub-forum, Alibaba Cloud Intelligent Group Computing Platform Business Unit AI Search lead Xing Shaomin presented "New Paradigm × New Engine: Alibaba Cloud Agentic Search Reshapes AI Search." The talk systematically explained the industry trend of AI search evolving from answer generation to task execution, and shared technical practices of Agentic Search 2.0, Elasticsearch, FalconSeek 2.0, and the Knowledge Engine.

AI search is undergoing two simultaneous upgrades: the upper layer from a Q&A tool to a Search Agent capable of planning, execution, and self-evolution; the lower layer from a human-oriented retrieval system to a real-time context engine for Agents.

Alibaba Cloud proposes a "new paradigm" connecting enterprise knowledge, tools, execution environments, and feedback mechanisms, and a "new engine" carrying high-frequency, real-time, multi-tenant, and hybrid retrieval demands, making search a key entry point for enterprise Agents to complete complex tasks.

AI Search Enters Scale Phase: Value Shifts from Traffic to Task Outcomes

With generative AI adoption accelerating, users increasingly describe complex intent in natural language and expect systems to synthesize multi-source information directly. Public data shows Google AI Overviews, Google AI Mode, ChatGPT, and Perplexity user bases and usage frequencies growing continuously, indicating AI search has moved from proof-of-concept to scaled application.

AI search also changes user decision chains. Traditional search requires users to filter links, compare information, and judge; AI search completes more requirement clarification and solution screening before users reach business pages. A cited customer sample study found AI search sourced traffic conversion rates up to 9x traditional search (not a universal benchmark due to industry, sample, and attribution variations), signaling enterprises must shift from pure traffic volume to intent quality and final task results.

Three-Stage Evolution: From "I Know" to "Task Done"

Xing summarized AI search evolution in three stages:

Pre-2022: Keyword retrieval and link return.

2023–2024: Understanding questions, synthesizing information, generating answers.

2025–2026: Task execution — users provide goals, not keywords; systems perform step decomposition, tool selection, knowledge access, operation execution, and result delivery.

The endpoint of search shifts from 'I know' to 'task done.' Enterprises need not just a large model bolted onto a search box, but a system connecting knowledge, planning, execution, memory, and feedback.

Agentic Search 2.0: Enterprise-Grade Search Agent Architecture

Built on enterprise knowledge, Agentic Search targets scenarios like industry research reports, operations diagnosis, automated inspection, data insight, and e-commerce guidance. It composes an enterprise-grade Search Agent from:

Agent Loop

Sandbox

Tool Calling

Permission Audit

Memory

AutoSkill

PageIndex

AgenticWiki

The system ingests multimodal documents, online collaborative docs, enterprise knowledge bases, search engines, and databases, converting scattered data into citable, verifiable execution bases via hybrid retrieval and dynamic orchestration.

Memory Module: Balancing Effectiveness, Cost, and Context Length

In long tasks, memory is not "more is better"; the key is accurate selection, update, and forgetting. Internal September 2026 evaluations showed:

LoCoMo composite score: 96.69%

LongMemEval-V2 test version overall accuracy: 81.37%

Under the LoCoMo evaluation configuration, input tokens reduced by 68.18%

These results demonstrate the system's exploration of balancing effectiveness, cost, and context length.

Controlled Self-Evolution via Multi-Agent Collaboration

Self-evolution does not mean unconstrained automatic system changes. Tasks are completed by Planner, Researcher, Executor, and Reviewer agents collaborating. Execution trajectories, quality scores, anomaly corrections, user feedback, and citation evidence feed a unified experience stream, which separately drives Memory updates, Skill extraction and evaluation, and incremental knowledge base precipitation. Evolved capabilities are loaded on demand and remain observable, evaluable, and rollbackable, preventing accidental experiences from hardening into system behavior.

Elasticsearch Agent: Extending Search to Ops Diagnosis and Data Insight

The presentation demonstrated Alibaba Cloud Elasticsearch Agent for operations diagnosis and data insight:

Ops scenario: Aggregates cluster status, logs, metrics, and historical events, outputting structured diagnostic reports with abnormal nodes, key metrics, probable causes, suggested actions, and evidence chains.

Data insight scenario: Users pose business analysis tasks in natural language; Agent understands metric definitions, retrieves data, identifies issues, and forms improvement recommendations.

The product forms a "Detection Layer — Analysis Layer — Insight Layer" stack:

Bottom: Continuous collection of logs, metrics, events, monitoring, alerts.

Middle: Agentic Search provides knowledge, memory, Skills, sandbox, permissions.

Top: Agents, Tools, Skills form intelligent inspection, fault diagnosis, performance tuning, data insight.

Behind the simple natural language entry, permission isolation, tool governance, execution audit, and evidence traceability remain prerequisites.

Elasticsearch as Agent-Oriented Multi-Tenant Engine

When search directly serves Agents, results must be structured, machine-readable, citable real-time context — not just human-readable. Agent call frequency is higher, data updates faster, tenant counts larger, requiring full-text search, vector search, multi-tenant isolation, stable latency, and controllable cost. Xing emphasized: Search engines are becoming the cognitive and context infrastructure for Agents.

Architecture: Write/Query Decoupling, OSS Unified Storage, Slice-Based Tenancy

Alibaba Cloud Elasticsearch builds an Agent-oriented multi-tenant engine: write nodes and query nodes decoupled; OSS as unified persistent storage; tenant data organized by Slice; cold data no longer occupies compute and local disk long-term; loaded on-demand into memory or SSD cache at query time.

In a million-document test under the specific architecture and measurement scope:

Hot and cold query P99 latency: ~60ms and ~600ms respectively.

Overall cost reduction: ~70% (actual effect varies with data scale, hardware, access patterns).

Qoder Benchmark: Hundreds of Billions of Documents, Tens of Millions of Repos

Qoder, an AI Coding scenario, faces large code repos and docs, many tenants, continuous code growth, and need for real-time search after updates. Disclosed figures:

Supports hundreds of billions of documents, tens of millions of code repositories, and hundreds of billions of vectors .

Hot query P99: ~60ms .

Second-level code updates.

Under corresponding load and benchmark scope, comprehensive cost ~1/7 of comparable solutions .

FalconSeek 2.0: Cloud-Native Kernel for Agent Workloads

Core capabilities are powered by Alibaba Cloud Elasticsearch cloud-native kernel FalconSeek. It maintains Elasticsearch API and Query DSL compatibility, moves core query execution to a C++ kernel reading Lucene Segment files directly, and uses batch execution, memory lifecycle management, and query-specific optimizations to reduce CPU, memory, and GC overhead.

In vector search, FalconSeek first traverses HNSW graph for candidates, then reads raw float vectors for batch re-ranking, enabling full-text and vector capabilities to co-optimize within the same engine.

Knowledge Engine: Retrieval-Action-Memory-Knowledge Self-Evolution Loop

For long-running Agent-produced memory, skills, and knowledge, the Knowledge Engine uses Namespace as tenant and Agent identity boundary. Within a unified index system it fuses BM25, vector, filter, rerank, and personalized retrieval. Hot/Warm/Cold/Frozen lifecycle management handles data of different value and access frequency, forming a "Retrieval-Action-Memory-Knowledge" enterprise knowledge self-evolution closed loop.

Xing concluded: Memory makes context accumulable, Skills make experience reusable, Knowledge Engine lets enterprise assets flow continuously within security boundaries. Through collaboration of Agentic Search, Elasticsearch, FalconSeek, and Knowledge Engine, Alibaba Cloud aims to build complete infrastructure from search to action , providing real-time, reliable, verifiable, and cost-controllable context support.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Memory ManagementElasticsearchVector SearchAI SearchEnterprise SearchAgentic SearchSearch AgentsFalconSeek
Alibaba Cloud Big Data AI Platform
Written by

Alibaba Cloud Big Data AI Platform

The Alibaba Cloud Big Data AI Platform builds on Alibaba’s leading cloud infrastructure, big‑data and AI engineering capabilities, scenario algorithms, and extensive industry experience to offer enterprises and developers a one‑stop, cloud‑native big‑data and AI capability suite. It boosts AI development efficiency, enables large‑scale AI deployment across industries, and drives business value.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.