From One Alert Card to One‑Click RCA: How Tasting Scaled Intelligent Ops Across 10,000 Stores
Tasting built a cloud‑native, AI‑assisted alert platform that unifies multi‑source alarms, reduces noise, automates root‑cause analysis, and closes the feedback loop, turning each incident into a reusable operations asset for its ten‑thousand‑store restaurant chain.
Business challenge – A nationwide restaurant chain with over 10,000 stores generates alerts from dozens of sources (CloudMonitor, ARMS, SLS, custom alerts, certificates, etc.). The alerts are scattered across 4‑5 systems, each with its own format, contact configuration, and notification method, leading to high maintenance cost, duplicate "alert storms", and heavy reliance on on‑call engineers' personal experience for root‑cause diagnosis.
Key insight – After an alert fires, the real efficiency gain comes from how the incident is handled, not from detection alone. A unified, intelligent platform is needed to turn raw alerts into traceable events, automate analysis, and capture knowledge for future reuse.
Solution architecture
Unified alert entry – Alibaba Cloud Monitor 2.0 (CMS 2.0) aggregates alerts from ARMS, CMS 1.0, Prometheus, SLS, and custom sources via a webhook, normalizes fields, and performs convergence judgment.
Low‑code workflow orchestration – The platform parses the alert source, standardizes context (title, level, resource, status), then routes the event through convergence, group routing, and escalation policies (P1‑P3) to generate or update a Feishu alert card.
Feishu alert card + AI assistance – The card displays real‑time status, allows on‑call engineers to claim, view, and query the incident, and automatically injects context into STAROps AI for root‑cause diagnosis.
Three‑step feedback loop
AI provides a diagnostic suggestion based on the injected event context and historical root‑cause library.
Engineer manually fills the actual root cause in the card.
Engineer rates the AI suggestion as "accurate" or "inaccurate" and optionally adds a comment.
Knowledge accumulation – Confirmed root causes are stored in a historical root‑cause database. Subsequent similar alerts prioritize this history, improving AI accuracy and reducing diagnosis time.
Digital employee matrix (STAROps) –
Skill : Encapsulates a complete SOP for a specific failure scenario (trigger conditions, data collection steps, logic, output format).
Mission : Orchestrates multiple Skills in parallel, forming a long‑running inspection matrix that runs per business line and time window.
Agent : Executes the Mission, produces structured inspection reports, and feeds results back into the knowledge base.
Results – After deployment, on‑call engineers see far fewer duplicate alerts, resolve incidents entirely within Feishu, and experience a noticeable reduction in investigation time. The platform provides quantifiable metrics for alert noise, response speed, AI accuracy, and knowledge reuse, turning the notification tool into a continuously learning engineering system.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Alibaba Cloud Native
We publish cloud-native tech news, curate in-depth content, host regular events and live streams, and share Alibaba product and user case studies. Join us to explore and share the cloud-native insights you need.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
