From Visibility to Self‑Healing: ChangjieTong’s Observability and AI‑Powered Ops Journey
ChangjieTong transformed its SaaS‑based, multi‑tenant finance cloud platform by building a five‑layer observability stack on Alibaba Cloud CloudMonitor 2.0, integrating a UModel digital‑twin, and deploying AI‑driven inspection, self‑healing and capacity‑prediction loops, which lifted SLA from 99.9% to 99.995% and cut average fault‑resolution time from over 10 minutes to under 30 seconds.
Background: As a leading provider of cloud‑based finance services for millions of small enterprises, ChangjieTong faced three critical operational problems after rapid SaaS and cloud‑native migration—"invisible" user‑experience issues, overwhelming alarm noise, and manual, experience‑driven fault handling.
Five‑layer observability model: The team classified applications into three tiers and built a unified monitoring hierarchy covering (1) basic infrastructure (CPU, memory, disk, network), (2) middleware (Redis, databases, message queues), (3) application performance (GC frequency, thread state, pod latency), (4) business‑level signals (error logs, rate‑limit events), and (5) user‑experience metrics (HTTP status codes, response‑time spikes). Data from all layers are ingested into Alibaba Cloud CloudMonitor 2.0, which unifies logs (SLS), real‑time metrics (CMS), tracing, and events, providing petabyte‑scale storage with 50% lower cost than self‑built solutions.
Digital twin with UModel: CloudMonitor 2.0’s UModel capability creates a three‑dimensional topology of applications, resources, and tenants. Entities such as ECS, VPC, SLB, RDS, ACK clusters, pods, services, and custom CMDB items are modeled in a unified graph, enabling drill‑down from a single user‑experience alarm to the exact resource and tenant responsible.
AI‑enabled operational loops: Leveraging the unified data foundation, ChangjieTong deployed three AI scenarios—intelligent inspection (pre‑emptive risk scans for capacity trends, configuration compliance, slow SQL, etc.), self‑healing (automated actions for resource exhaustion, connection‑pool limits, node failures, traffic spikes), and capacity prediction (time‑series forecasting to anticipate resource shortages and guide FinOps cost‑control). Each loop follows a "detect → analyze → decide → act" cycle, with AI recommendations reviewed by humans before execution.
Evolution stages: The journey is divided into four phases. Phase 1 (SLA 99.9%) established lifecycle management, red‑line policies, and multi‑center gray‑release controls. Phase 2 (SLA 99.95%) added platform services, automated policy enforcement, and systematic debt remediation. Phase 3 (SLA 99.99%) codified the "0‑2‑5‑10" methodology—0 incidents, detection within 2 minutes (MTTI), root‑cause analysis within 5 minutes (MTTK), and remediation within 10 minutes (MTTF + MTTV). Phase 4 (SLA 99.995%) layered AI capabilities, delivering an intelligent cockpit, digital employees, and skill‑based AI modules that close the loop from perception to automated remediation.
Results: After the four‑stage rollout, SLA rose from 99.9% to 99.995%, reducing annual downtime from ~9 hours to <30 minutes. Average fault‑location time dropped from >10 minutes to <30 seconds, a >20× improvement in the MTTK segment. The system now prevents 99% of incidents before they affect users and limits manual intervention to high‑severity cases, freeing the on‑call team to focus on architectural improvements.
Organizational impact: The observability platform captured expert knowledge into a digital‑employee knowledge base, enabling rapid onboarding and consistent response quality. Insights from AI diagnostics feed back into development standards (e.g., slow‑SQL patterns become codified database guidelines), creating a virtuous cycle of continuous improvement.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Alibaba Cloud Native
We publish cloud-native tech news, curate in-depth content, host regular events and live streams, and share Alibaba product and user case studies. Join us to explore and share the cloud-native insights you need.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
