Tagged articles

observability

1324 articles · Page 1 of 14
DataFunTalk
DataFunTalk
Oct 5, 2026 · Artificial Intelligence

WeChat's Agentic OLAP Architecture: Observability, Memory & Access Control

WeChat's technical architecture team shares their exploration of adapting OLAP infrastructure for Agentic AI, covering distributed observability with Langfuse and ClickHouse, remote memory services using vector search and FUSE, and access protection mechanisms treating agents as authenticated identities.

Access ControlAgentic AIClickHouse
0 likes · 16 min read
WeChat's Agentic OLAP Architecture: Observability, Memory & Access Control
Architecture and Beyond
Architecture and Beyond
Oct 3, 2026 · R&D Management

Beyond Vibe Coding: What Engineers Must Do When Code Becomes a Black Box

The article analyzes risks of AI-generated 'vibe coding' where engineers lose internal code understanding, detailing consequences like hidden architectural decay, loss of tacit production knowledge, and inability to debug complex failures, then proposes practical strategies: defining white-box boundaries, independent verification, limiting change scope, retaining takeover capability, preserving dark knowledge, and reallocating engineering time toward constraint design and failure recovery.

AI-assisted developmentEngineering Managementblack-box systems
0 likes · 30 min read
Beyond Vibe Coding: What Engineers Must Do When Code Becomes a Black Box
Architects' Tech Alliance
Architects' Tech Alliance
Sep 30, 2026 · Artificial Intelligence

YuanNao Web Agent Review: 5 Enterprise-Ready Design Patterns from a Hands-On Test

A hands-on test of YuanNao Web Agent reveals five key design patterns — unified multi-agent portal, model-agent decoupling, transparent metering, enterprise usage dashboards, and layered security — that transform personal AI agents into a centrally managed, cost-trackable, and secure enterprise asset.

AI agentsAccess ControlEnterprise AI
0 likes · 8 min read
YuanNao Web Agent Review: 5 Enterprise-Ready Design Patterns from a Hands-On Test
ITPUB
ITPUB
Sep 30, 2026 · Artificial Intelligence

How AI Agents Handle High Concurrency: Admission, Backpressure, Async & Isolation

This article details a production-grade architecture for handling high concurrency in AI agent systems, covering real load estimation, admission control with bounded queues, async task processing, hierarchical concurrency budgets for models and tools, resource isolation via bulkheads, idempotency for state consistency, graded degradation strategies, and observability-driven capacity planning.

AI agentsCircuit BreakerConcurrency Budget
0 likes · 18 min read
How AI Agents Handle High Concurrency: Admission, Backpressure, Async & Isolation
dbaplus Community
dbaplus Community
Sep 29, 2026 · Industry Insights

From Monoliths to AI Agents: Architecture Evolution's Two Patterns and New Challenges

This article traces software architecture evolution from monoliths through primitive distributed systems, SOA, microservices, and cloud-native, highlighting two recurring patterns — finer decoupling and stronger fault isolation — and a shift from zero-failure goals to designing for resilience, while outlining four unprecedented challenges AI agents introduce: semantic hallucinations, stateful context, dynamic orchestration, and observability gaps.

AI agentsCloud NativeMicroservices
0 likes · 23 min read
From Monoliths to AI Agents: Architecture Evolution's Two Patterns and New Challenges
TonyBai
TonyBai
Sep 26, 2026 · Backend Development

Go Heap Profiles to Gain Complete Memory View via New memory_space Sample Type

A Go proposal (#79179) provisionally accepted to add a memory_space sample type to heap/alloc profiles, providing a full memory hierarchy from RSS down to live/dead heap, stacks, runtime, and non-Go memory, solving the long-standing gap where heap profiles show only about 50% of actual RSS.

Goheap profilememory profiling
0 likes · 12 min read
Go Heap Profiles to Gain Complete Memory View via New memory_space Sample Type
Data Bricklaying Diary
Data Bricklaying Diary
Sep 24, 2026 · Operations

Observability Isn't a Log Platform: Connecting Business Goals to Verifiable Runtime Facts

This article explains that observability is not centralized logging but a practice of defining business outcomes (SLIs/SLOs), correlating metrics, logs, and traces via stable business identifiers, designing actionable alerts, structuring dashboards along business flows, separating observability responsibilities, halting automation when evidence is unreliable, and using runtime data to continuously correct architectural assumptions.

AlertingSLI/SLOSRE
0 likes · 28 min read
Observability Isn't a Log Platform: Connecting Business Goals to Verifiable Runtime Facts
Tencent Technical Engineering
Tencent Technical Engineering
Sep 23, 2026 · Operations

How an AI Agent Cuts Error Code Triage from Hours to Minutes: Three Practices

This article details how an AI Agent on an orchestration platform automates error code root cause analysis by integrating knowledge bases, code graphs, observability data, and code hosting platforms, achieving 88% consistency with human annotations and zero hard conflicts across 50 test cases through a five-step investigation workflow, dual-path JSON parsing, and regression-tested prompt optimization.

AI AgentError Code TroubleshootingRoot Cause Analysis
0 likes · 26 min read
How an AI Agent Cuts Error Code Triage from Hours to Minutes: Three Practices
Random Bulletin
Random Bulletin
Sep 22, 2026 · Backend Development

Fault Prediction at 10M QPS: From Zero to Production-Ready System

This article details a practical roadmap for building production-grade fault prediction systems at massive scale, covering target selection, data governance, model evolution, time-window design, and safe action loops—emphasizing that reliable prediction requires engineering rigor beyond just model training.

Operationsfault predictionincident response
0 likes · 29 min read
Fault Prediction at 10M QPS: From Zero to Production-Ready System
Random Bulletin
Random Bulletin
Sep 20, 2026 · Backend Development

Automating Fault Localization at 10M QPS: Evidence Chains Over Manual Hunts

This article details how to build automated fault localization for 10M QPS systems by unifying entity identities, aligning timestamps, integrating change records, and applying a four-layer engine—anomaly normalization, temporal correlation, topological pruning, and causal scoring—to converge millions of anomalies into verifiable hypotheses while avoiding correlation-causation pitfalls through counterfactual evidence and phased rollout.

Causal InferenceRoot Cause Analysisautomated troubleshooting
0 likes · 35 min read
Automating Fault Localization at 10M QPS: Evidence Chains Over Manual Hunts
Random Bulletin
Random Bulletin
Sep 19, 2026 · Operations

Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation

This article explains how to reduce fault detection latency from minutes to seconds in 10M QPS systems by implementing layered signals, combined evidence detection, distributed judgment, event normalization, and safe automation guardrails, rather than simply increasing sampling frequency.

AlertingSLOdistributed systems
0 likes · 36 min read
Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation
Cloud Architecture
Cloud Architecture
Sep 19, 2026 · Databases

Beyond COUNT_STAR=0: Auditable MySQL Index Removal with performance_schema

This article presents a rigorous, production-ready framework for MySQL index governance that moves beyond simplistic zero-usage checks, using performance_schema observation windows, structural gatekeeping, invisible index canary testing, and quantified rollback metrics to safely identify and remove redundant indexes without risking query regressions.

MySQLSQL Tuningdatabase administration
0 likes · 29 min read
Beyond COUNT_STAR=0: Auditable MySQL Index Removal with performance_schema
Open Source Tech Hub
Open Source Tech Hub
Sep 18, 2026 · Industry Insights

AI Makes Code Cheap: The Rising Value of Facts, Guardrails & Self-Correcting Systems

As AI drives code generation costs toward zero, developers must shift focus from writing speed to guarding increasingly valuable assets: real-world facts, constraint-based guardrails, failure archives, self-correcting closed loops, and the ability to define standards in uncharted technical territories.

AI code generationJevons ParadoxTechnical Standards
0 likes · 10 min read
AI Makes Code Cheap: The Rising Value of Facts, Guardrails & Self-Correcting Systems
Ops Development & AI Practice
Ops Development & AI Practice
Sep 17, 2026 · Cloud Native

Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained

This article explains why OpenTelemetry Helm charts now recommend deploying Collector as three separate releases—otel-agent (DaemonSet for node metrics), otel-cluster (singleton Deployment for cluster metrics), and otel-gateway (scalable Deployment for trace ingestion)—detailing Presets simplification, lifecycle isolation, failure domains, and when to consolidate to two releases.

Cloud NativeCollectorDaemonSet
0 likes · 24 min read
Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained
Random Bulletin
Random Bulletin
Sep 17, 2026 · Operations

10M QPS Architecture #256: Shifting Mega-Promo Reliability from Reactive to Proactive

This article details a proactive framework for mega-promotion reliability at 10M QPS, covering business red lines, full-link capacity modeling, stress testing for system boundaries, actionable degradation and isolation, decision-centric observability, executable runbooks, structured war rooms, drills, and readiness gates to replace reactive firefighting.

capacity planningdegradation strategieshigh-concurrency architecture
0 likes · 41 min read
10M QPS Architecture #256: Shifting Mega-Promo Reliability from Reactive to Proactive
Alibaba Cloud Observability
Alibaba Cloud Observability
Sep 14, 2026 · Cloud Native

How Lemon Retail Achieved 70% Alert Convergence and Minute-Level MTTR with AI-Driven Cloud-Native Observability

Lemon, a food retail SaaS provider serving 20,000+ stores, unified logs, metrics, and traces into a full-chain digital twin using Alibaba Cloud CloudMonitor 2.0 and STAROps, deploying four intelligent operations layers that cut alert noise by 70% and reduced incident response and MTTR to minutes.

AIOpsAlert GovernanceCloud Native
0 likes · 16 min read
How Lemon Retail Achieved 70% Alert Convergence and Minute-Level MTTR with AI-Driven Cloud-Native Observability
Su San Talks Tech
Su San Talks Tech
Sep 14, 2026 · Backend Development

5 AI Patterns to Pinpoint Flaky Production Bugs

The article presents five practical patterns for using AI to debug intermittent production bugs: time-window log analysis, comparative field diffing, concurrent reproduction scripts, hypothesis validation with evidence tables, and integrating logging/database monitoring via MCP, emphasizing AI for retrieval/enumeration while humans retain fix responsibility.

AI debuggingMCPProduction Incidents
0 likes · 28 min read
5 AI Patterns to Pinpoint Flaky Production Bugs
Cloud Architecture
Cloud Architecture
Sep 13, 2026 · Backend Development

Production-Grade RAG with Spring AI: Verifiable, Rollbackable, Auditable Knowledge Base

This article details a production-ready customer service knowledge base built with Spring AI 2.0.1 and Milvus, covering immutable index versioning, tenant-isolated retrieval with parameterized filters, deterministic chunk IDs, idempotent ingestion pipelines, dual-index blue-green deployments, and comprehensive observability with automated rollback triggers.

KubernetesMilvusProduction Engineering
0 likes · 27 min read
Production-Grade RAG with Spring AI: Verifiable, Rollbackable, Auditable Knowledge Base
Random Bulletin
Random Bulletin
Sep 13, 2026 · Backend Development

Safe Configuration Rollout at Scale: Versioning, Canary, and Rollback Loops

The article analyzes the engineering challenges of moving configuration management from single-machine files to distributed cluster control planes, covering immutable versioning, schema validation layers, canary deployment by failure domains, two-phase commit for atomic switches, rollback with side-effect mitigation, observability across control and data planes, and maturity stages of config platforms.

Control Planecanary deploymentconfiguration management
0 likes · 47 min read
Safe Configuration Rollout at Scale: Versioning, Canary, and Rollback Loops
Golang Shines
Golang Shines
Sep 13, 2026 · Operations

AIOps: A Systematic Guide to Intelligent IT Operations

This article systematically explains AIOps (Artificial Intelligence for IT Operations), covering its definition as a capability combining big data, ML, and NLP; core value in reducing alert noise and cognitive load; four key components; complementary relationship with DevOps; domain-agnostic vs domain-centric implementation strategies; and the future of predictive operations with human-in-the-loop oversight.

AIOpsDevOpsIT Operations
0 likes · 8 min read
AIOps: A Systematic Guide to Intelligent IT Operations
LuTiao Programming
LuTiao Programming
Sep 11, 2026 · Backend Development

API Latency Spike to 8s: Not Slow SQL, But Long Transactions Holding DB Connections

An API's sudden latency spike from 200ms to 8 seconds was traced not to slow SQL but to database connection pool exhaustion caused by @Transactional methods holding connections during external HTTP calls to a payment service; splitting transactions and moving external calls outside the transaction boundary resolved the issue.

Database Connection PoolHikariCPMicroservices
0 likes · 15 min read
API Latency Spike to 8s: Not Slow SQL, But Long Transactions Holding DB Connections
Architecture Development Notes
Architecture Development Notes
Sep 11, 2026 · Artificial Intelligence

Control Inversion in Agent Tool Calling: Orchestration Loops Move Into Model-Generated Code

This article analyzes the shift from host-controlled to model-generated orchestration loops in AI agent tool calling, comparing Anthropic's API-level and LangChain's middleware approaches, examining cost control primitives, container lifecycle, debugging challenges, BFCL v4 evaluation data, and practical adoption criteria for programmatic tool calling.

AnthropicBFCL v4LangChain
0 likes · 17 min read
Control Inversion in Agent Tool Calling: Orchestration Loops Move Into Model-Generated Code
Woodpecker Software Testing
Woodpecker Software Testing
Sep 11, 2026 · Artificial Intelligence

AI Testing Performance Optimization: Compute, I/O, Scheduling & Observability Deep Dive

This article analyzes performance optimization for AI-driven testing tools across four dimensions—computation, I/O, scheduling, and observability—detailing practical architectural strategies like lightweight models, zero-copy data transfer, dynamic Kubernetes-based scheduling, and multi-layer observability, with real-world case studies showing significant latency and cost reductions.

AI testingKubernetes schedulingPerformance Optimization
0 likes · 9 min read
AI Testing Performance Optimization: Compute, I/O, Scheduling & Observability Deep Dive
Cloud Architecture
Cloud Architecture
Sep 10, 2026 · Operations

From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana

This comprehensive guide details building a production-grade database monitoring system using Prometheus and Grafana, covering SLI/SLO design, alerting strategies, architecture, metric selection, security, Prometheus configuration, Alertmanager routing, Grafana dashboards, scaling, incident response runbooks, anti-patterns, and operational processes to shift from reactive firefighting to proactive prevention.

AlertingAlertmanagerDatabase Monitoring
0 likes · 42 min read
From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana
Random Bulletin
Random Bulletin
Sep 10, 2026 · Operations

Canary Releases at 10M QPS: Making Gradual Rollouts Mandatory, Not Optional

This article explains why canary releases become essential at 10M QPS, detailing how to define canary units by identity, failure domain, and business scenario; set absolute traffic limits; use evidence-driven state machines with four metric layers; ensure version compatibility; choose rollback, feature-flag, or forward-fix strategies; build platform capabilities; avoid common pitfalls; and make canary the default deployment path.

10M QPSPlatform Engineeringcanary release
0 likes · 30 min read
Canary Releases at 10M QPS: Making Gradual Rollouts Mandatory, Not Optional
iQIYI Technical Product Team
iQIYI Technical Product Team
Sep 10, 2026 · Big Data

Inside iQIYI's Agent Team: Automating Cross-Service Big Data Diagnosis

iQIYI built a Big Data Assistant using an Agent Team architecture where a Coordinator delegates tasks to domain-specific Service Agents (Scheduling, Spark, Flink, ML, Data Lake, StarRocks) that collaborate via a shared blackboard to diagnose cross-service issues, evolving from rule-based workflows to LangGraph-based agents with Harness engineering for context management, tool control, and observability.

Agent TeamCross-Service DiagnosisHarness Engineering
0 likes · 21 min read
Inside iQIYI's Agent Team: Automating Cross-Service Big Data Diagnosis
Woodpecker Software Testing
Woodpecker Software Testing
Sep 10, 2026 · Frontend Development

Self-Healing UI Test Automation with SeleniumBase: A Practical Guide

This article details a three-step workflow for building self-healing UI test scripts using SeleniumBase, covering declarative locator fallbacks, DOM context inference, audit logging, and observability integration, reducing locator failure rates from 31% to 4.7% in a financial app regression suite.

CI/CDDOM inferenceSeleniumBase
0 likes · 8 min read
Self-Healing UI Test Automation with SeleniumBase: A Practical Guide
James' Growth Diary
James' Growth Diary
Sep 10, 2026 · Backend Development

The Five Acts of an AI Conversation Lifecycle: Why Architecture Matters

This article dissects the five-stage lifecycle of an AI agent conversation—from host submission through trusted identity, gateway routing, runtime turns, to audit logging—revealing why treating chat as a simple HTTP request leads to misdiagnosed failures, security gaps, and broken observability across multi-entry platforms.

AI agent platformAudit Loggingconversation lifecycle
0 likes · 23 min read
The Five Acts of an AI Conversation Lifecycle: Why Architecture Matters
Raymond Ops
Raymond Ops
Sep 9, 2026 · Operations

Linux Logging Deep Dive: Kernel, journald, rsyslog & 5 Real Fault Cases

This comprehensive guide dissects the Linux logging stack — kernel ring buffer, journald, rsyslog, logrotate, and service logs — with configuration details, command references, and five step-by-step troubleshooting cases covering SSH brute force, disk exhaustion, OOM kills, network packet loss, and systemd service failures.

LinuxOperationsTroubleshooting
0 likes · 47 min read
Linux Logging Deep Dive: Kernel, journald, rsyslog & 5 Real Fault Cases
Alibaba Cloud Native
Alibaba Cloud Native
Sep 8, 2026 · Cloud Native

Lemon's Intelligent Ops: 70% Alert Convergence, Minute-Level MTTR for 20K+ Retail Stores

Food retail digitalizer Lemon unified logs, metrics, and traces into a digital twin using Alibaba Cloud CloudMonitor 2.0 and STAROps with UModel, deploying unified alert governance, natural language observability, automated inspections, and AI-driven root cause analysis to achieve 70% alert convergence and minute-level MTTR across 20,000+ stores.

AIOpsAlert GovernanceAutomated Inspection
0 likes · 16 min read
Lemon's Intelligent Ops: 70% Alert Convergence, Minute-Level MTTR for 20K+ Retail Stores
Architecture Development Notes
Architecture Development Notes
Sep 8, 2026 · Artificial Intelligence

OpenTelemetry GenAI Semantic Conventions: Modeling Agent Runs as Span Trees

The article explains how OpenTelemetry GenAI semantic conventions address the gap in agent observability by defining a standardized span tree structure (invoke_agent, execute_tool, chat, etc.) and three critical fields (gen_ai.agent.name, gen_ai.conversation.id, gen_ai.operation.name) to capture tool calls, retries, and child-agent handoffs that auto-instrumentation misses.

GenAILLM AgentsOpenTelemetry
0 likes · 10 min read
OpenTelemetry GenAI Semantic Conventions: Modeling Agent Runs as Span Trees
Woodpecker Software Testing
Woodpecker Software Testing
Sep 8, 2026 · R&D Management

Shift-Left Testing 2026: AI Quality Gates, Contract-First Collaboration, and Observable Metrics

The article analyzes four 2026 shift-left testing trends: IDE-level contract testing, AI-driven risk-based quality gates, contract-first cross-team governance, and a three-dimensional observability model measuring defect detection time, cost ratio, and production rollback reduction, with real-world metrics from banking, e-commerce, and automotive sectors.

AI quality gatesContract TestingDevOps
0 likes · 9 min read
Shift-Left Testing 2026: AI Quality Gates, Contract-First Collaboration, and Observable Metrics
Software Engineering 3.0 Era
Software Engineering 3.0 Era
Sep 7, 2026 · Artificial Intelligence

AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems

This comprehensive guide synthesizes 2026 industry practices from Xiaohongshu and Alipay to build production-ready AI Agent evaluation systems, covering metrics (Quality/Cost/Safety), three-tier evaluation granularities, Judge system design, OpenTelemetry-based observability, platform architecture with contract-driven test generation, dual flywheel offline/online loops, and self-evolving prompt optimization — moving evaluation from post-hoc verification to embedded engineering guardrails.

AI Agent EvaluationAgentOpsEvaluation Methodology
0 likes · 37 min read
AI Agent Evaluation Guide: Building Observable, Evaluable, Self-Evolving Quality Systems
Cloud Architecture
Cloud Architecture
Sep 7, 2026 · Backend Development

Why a Single Timeout Spawned Two Risk Reviews: Production MCP Server Patterns

The article analyzes a timeout-induced duplicate risk review incident, then presents a comprehensive production-grade MCP server design covering stateless protocol alignment, schema validation, dual-key idempotency (request_key + business_key), UNKNOWN state machine, recovery workers, MCP Tasks integration, concurrency control, observability, and fault-injection testing to ensure exactly-once business effects.

MCPModel Context ProtocolProduction Engineering
0 likes · 27 min read
Why a Single Timeout Spawned Two Risk Reviews: Production MCP Server Patterns
Alibaba Cloud Native
Alibaba Cloud Native
Sep 6, 2026 · Artificial Intelligence

AI Agents Need a Semantic Layer, Not More Data: UnifiedModel Boosts Accuracy 10-20%

UnifiedModel provides an open-source semantic layer that organizes enterprise assets, data, and relationships into a queryable object graph, enabling AI agents to read metrics by object and trace root causes along relationships; experiments on DataAgentBench show 10-20% accuracy gains for four flagship models, with GLM-5.2 reaching 50.2% pass@1.

AI agentsDataAgentBenchRoot Cause Analysis
0 likes · 20 min read
AI Agents Need a Semantic Layer, Not More Data: UnifiedModel Boosts Accuracy 10-20%
Open Source Tech Hub
Open Source Tech Hub
Sep 6, 2026 · Backend Development

Designing Reliable Failure Boundaries in PHP: Errors, Exceptions & Result Types

This article explores how to design robust failure boundaries in PHP payment systems by distinguishing validation errors, expected business declines, transient provider failures, and programming defects, using result types for normal outcomes, translated exceptions for integration failures, idempotent retries, and observability without sensitive data.

PHPerror handlingexceptions
0 likes · 38 min read
Designing Reliable Failure Boundaries in PHP: Errors, Exceptions & Result Types
Architecture Development Notes
Architecture Development Notes
Sep 4, 2026 · Artificial Intelligence

75% Cheaper Cache Reads: Why Long-Running Agent Costs Now Depend on Prefix Stability

Anthropic's Fable 5.1 reduces cache read pricing from $1 to $0.25 per million tokens, shifting long-running agent cost bottlenecks from output to repeated stable prefix reads, making prefix stability, cache breakpoint placement, TTL tuning, and hit-rate observability critical architectural levers for cost control.

AI Agent ArchitectureAnthropicContext Management
0 likes · 15 min read
75% Cheaper Cache Reads: Why Long-Running Agent Costs Now Depend on Prefix Stability
TechVision Expert Circle
TechVision Expert Circle
Sep 4, 2026 · Artificial Intelligence

AI Bills Skyrocket Despite Cheaper Models: The Agent Cost Multiplier Effect

As model inference prices drop, AI costs surge because Agent architectures multiply model calls per user request; the article breaks down the four-layer cost structure and offers six practical governance tactics—model routing, prompt caching, call-chain slimming, token budgets, observability, and chargebacks—to build a sustainable AI FinOps practice.

AI cost managementAgent ArchitectureFinOps
0 likes · 15 min read
AI Bills Skyrocket Despite Cheaper Models: The Agent Cost Multiplier Effect
Cloud Architecture
Cloud Architecture
Sep 3, 2026 · Backend Development

TCC Is Not a Silver Bullet: Engineering Go Microservice Distributed Transactions with Performance Tuning

This article details the engineering implementation and performance optimization of TCC distributed transactions in Go microservices, using an order placement scenario with inventory locking and wallet freezing, covering data models, idempotent state machines, recovery workers, hotspot optimization, and observability.

Distributed TransactionsGoMicroservices
0 likes · 33 min read
TCC Is Not a Silver Bullet: Engineering Go Microservice Distributed Transactions with Performance Tuning
Tech Ocean
Tech Ocean
Sep 3, 2026 · Backend Development

AI Coding Is Nearly Free — Why Does Delivery Still Cost So Much?

The article argues that while AI tools like Claude Code accelerate initial code generation, the real engineering effort lies in defining failure states, untangling legacy systems, ensuring idempotency, rigorous testing, safe rollbacks, and post-launch observability — illustrated through an order resubmission feature that requires six non-negotiable steps before production.

AI codingBackend Engineeringidempotency
0 likes · 12 min read
AI Coding Is Nearly Free — Why Does Delivery Still Cost So Much?
Random Bulletin
Random Bulletin
Sep 2, 2026 · Operations

Automating Root‑Cause Analysis for Million‑QPS Systems: From Manual to AI‑Assisted

When a transaction‑success rate dropped at 02:13 AM and 186 alerts flooded the on‑call channel, engineers struggled to piece together fragmented evidence, highlighting why manual root‑cause analysis is slow at scale and how an evidence‑driven automated pipeline can narrow investigation space, rank candidates with confidence, and keep humans in the loop for safe remediation.

Root Cause Analysisautomationincident response
0 likes · 26 min read
Automating Root‑Cause Analysis for Million‑QPS Systems: From Manual to AI‑Assisted
Woodpecker Software Testing
Woodpecker Software Testing
Sep 2, 2026 · Operations

5 Common Pitfalls in Performance Regression Testing

In today’s fast‑paced agile and micro‑service environments, performance regression testing is often treated as optional, leading to severe TPS drops and hidden degradations; this article details five typical misconceptions, backs them with real‑world examples, and offers concrete practices to make performance regression a continuous, cross‑team responsibility.

CI/CDKubernetesLoad Testing
0 likes · 8 min read
5 Common Pitfalls in Performance Regression Testing
Code Mala Tang
Code Mala Tang
Sep 1, 2026 · Industry Insights

AI Agent Infrastructure Is Just 1996 Linux Sysadmin Practices Rebranded

The article maps modern AI Agent terminology — identity isolation, least privilege, sandbox, runtime, scheduler, observability, human-in-the-loop, secure execution, self-healing — to decades-old Linux concepts like user permissions, chmod, directories, sudo, systemd, cron, journalctl, SSH, Docker, and process supervision, arguing that Linux veterans already manage AI agents as just another untrusted user.

AI agentsDevOpsInfrastructure
0 likes · 3 min read
AI Agent Infrastructure Is Just 1996 Linux Sysadmin Practices Rebranded
Alibaba Cloud Observability
Alibaba Cloud Observability
Sep 1, 2026 · Cloud Native

Cross-Layer Root Cause Analysis: How STAROps & Yaochi Agent Trace Alerts to Missing Indexes & Blocking Lua Scripts

The article demonstrates how STAROps' full-stack correlation combined with Yaochi Agent's deep database diagnosis enables cross-layer root cause analysis across three real incidents: an RDS missing index, a Redis Lua script blocking the single thread, and an EVAL command causing CPU saturation, showing command-level diagnostic precision.

EVAL commandLua scriptRDS
0 likes · 14 min read
Cross-Layer Root Cause Analysis: How STAROps & Yaochi Agent Trace Alerts to Missing Indexes & Blocking Lua Scripts
Random Bulletin
Random Bulletin
Sep 1, 2026 · Operations

From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems

The article examines why manual alert handling stalls at massive scale, outlines the risks of naïve auto‑rollback, and presents a step‑by‑step framework—including event control planes, executable runbooks, safety guards, and staged automation—to reliably move from human‑only to fully automated incident response in high‑throughput environments.

OperationsRunbookalert automation
0 likes · 23 min read
From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems
Woodpecker Software Testing
Woodpecker Software Testing
Sep 1, 2026 · Cloud Native

Why 90% of Container Performance Issues Come From Poor Capacity Planning – An In‑Depth Look

The article explains how container performance testing must evolve from simple load simulation to chaos‑engineered, observability‑driven capacity planning, introduces a 4‑dimensional capacity model, and shows automated SLI‑based scaling using real‑world e‑commerce and finance case studies.

Control PlaneKubernetescapacity planning
0 likes · 7 min read
Why 90% of Container Performance Issues Come From Poor Capacity Planning – An In‑Depth Look
Alibaba Cloud Native
Alibaba Cloud Native
Aug 31, 2026 · Cloud Native

Kickstarting the Data Flywheel: Four Ways to Connect Agents to AgentLoop

This article explains how AgentLoop uses the OpenTelemetry protocol and probes to ingest high‑quality runtime data, offering four integration methods—one‑click generic agents, SDK framework integration, annotation‑based high‑code, and eBPF—demonstrated with a Claude Code customer‑service agent and end‑to‑end verification on the observation page.

AgentLoopCloud NativeData Ingestion
0 likes · 9 min read
Kickstarting the Data Flywheel: Four Ways to Connect Agents to AgentLoop
dbaplus Community
dbaplus Community
Aug 30, 2026 · Artificial Intelligence

Cut Alert Troubleshooting Time by 80% with LLM Agents: A Full Technical Walkthrough

The article details how an LLM‑driven Troubleshooter system automates data collection, root‑cause analysis, and recommendation generation for alerts, slashing median investigation time from about 20 minutes to 4.4 minutes across 11 services and over ten alert types, and presents architecture, tool design, observability, a real‑world case, performance metrics, and future roadmap.

Alert TroubleshootingLLMReAct agent
0 likes · 17 min read
Cut Alert Troubleshooting Time by 80% with LLM Agents: A Full Technical Walkthrough
Random Bulletin
Random Bulletin
Aug 29, 2026 · Operations

Alert Convergence at Scale: From Simple Deduplication to AI‑Driven Clustering

A 40‑second DB jitter triggered over 3,000 alerts, but by applying a five‑layer alert‑convergence strategy—deduplication, grouping, inhibition & silencing, dependency‑based aggregation, and AI‑powered clustering—teams can reduce noise by up to 90 %, turning a storm of notifications into a single actionable signal.

AIOpsAlertingdeduplication
0 likes · 19 min read
Alert Convergence at Scale: From Simple Deduplication to AI‑Driven Clustering
Data Bricklaying Diary
Data Bricklaying Diary
Aug 29, 2026 · Operations

From LLMOps to AgentOps: Operating Enterprise Agents Across Full Task Lifecycles

This article argues that enterprises need AgentOps, not just LLMOps, to manage AI agents that execute multi-step tasks with tools, state, and human oversight, detailing six key capabilities: task identity, state checkpoints, component versioning, end-to-end observability, task-level evaluation, and human-in-the-loop as a first-class operational state.

AI agentsAgentOpsLLMOps
0 likes · 15 min read
From LLMOps to AgentOps: Operating Enterprise Agents Across Full Task Lifecycles
Random Bulletin
Random Bulletin
Aug 28, 2026 · Operations

Log Correlation: From Isolated Entries to Linked Traces in High‑QPS Systems

The article explains how to turn millions of independent error logs into a coherent, request‑level waterfall and business‑level story by injecting trace_id, business keys, and cross‑signal foreign keys, while addressing async boundaries, sampling, naming consistency, and storage costs.

distributed tracinglog correlationmetrics
0 likes · 18 min read
Log Correlation: From Isolated Entries to Linked Traces in High‑QPS Systems
Data Bricklaying Diary
Data Bricklaying Diary
Aug 27, 2026 · R&D Management

System Architecture Isn't Tech Selection: The Critical Questions for Production Readiness

This article argues that system architecture starts from business goals and quality attributes, not technology choices, using a refund processing example to illustrate trade-offs across system boundaries, collaboration patterns, and six architectural perspectives, emphasizing that observability, fault tolerance, and security must be designed in and validated with independent evidence throughout the system's lifecycle.

Microservicesarchitectural trade-offsevidence-based validation
0 likes · 20 min read
System Architecture Isn't Tech Selection: The Critical Questions for Production Readiness
Woodpecker Software Testing
Woodpecker Software Testing
Aug 27, 2026 · Cloud Native

Adversarial Performance Testing: A Hands‑On Guide to Boost System Resilience

In today’s high‑concurrency, microservice‑driven cloud‑native world, traditional load testing often misses real‑world failure modes, so this guide introduces adversarial performance testing—injecting faults, latency, and malicious traffic—to expose hidden bottlenecks and build resilient systems.

Cloud NativeMicroservicesPerformance Optimization
0 likes · 8 min read
Adversarial Performance Testing: A Hands‑On Guide to Boost System Resilience
Woodpecker Software Testing
Woodpecker Software Testing
Aug 27, 2026 · Backend Development

Backend Performance Tuning: Emerging Trends and Strategies for the Next Three Years

The article examines how backend performance tuning is evolving from manual, experience‑driven cycles to data‑driven observability, closed‑loop AIOps, serverless/Wasm architectures, and energy‑aware optimization, outlining concrete examples, tools, and forecasts shaping the field through 2026.

AIOpsBackend PerformanceEnergy Efficiency
0 likes · 8 min read
Backend Performance Tuning: Emerging Trends and Strategies for the Next Three Years
Random Bulletin
Random Bulletin
Aug 26, 2026 · Operations

Log Standards at Million‑QPS Scale: From Free‑form to Strict Structured Logging

When a production outage forces a midnight investigation across five services, the lack of log standards turns a quick debug into an all‑night forensic hunt; the article explains how structured logs, a unified schema, level semantics, trace_id linking, and field‑level masking enforced by SDK, Lint and CI can eliminate these five walls and make logging scalable, searchable, and compliant.

CI enforcementStructured Logginghigh QPS
0 likes · 22 min read
Log Standards at Million‑QPS Scale: From Free‑form to Strict Structured Logging
Woodpecker Software Testing
Woodpecker Software Testing
Aug 26, 2026 · Operations

2026 Open‑Source Performance Testing Tools: From Load to Diagnosis

The article evaluates the evolution and practical capabilities of leading 2026 open‑source performance testing tools across twelve real‑world scenarios—ranging from financial API stress tests to IoT clusters and LLM latency—using a five‑dimensional model that assesses protocol coverage, native cloud‑native integration, intelligent diagnosis, generative collaboration, and compliance readiness.

Cloud NativeLoad Testingbenchmark
0 likes · 8 min read
2026 Open‑Source Performance Testing Tools: From Load to Diagnosis
Cloud Architecture
Cloud Architecture
Aug 25, 2026 · Backend Development

Designing Nginx for Million‑Scale WebSocket Connections: Architecture, Configuration, and Pitfalls

This article walks through the end‑to‑end design of a production‑grade Nginx‑based WebSocket gateway that can handle a million concurrent connections, covering the five essential requirements, detailed Nginx settings, backend gateway responsibilities, load‑balancing strategies, observability, common failure patterns, and step‑by‑step Go code examples.

GoKubernetesNginx
0 likes · 34 min read
Designing Nginx for Million‑Scale WebSocket Connections: Architecture, Configuration, and Pitfalls
Woodpecker Software Testing
Woodpecker Software Testing
Aug 25, 2026 · Industry Insights

End-to-End Load Testing: Deep Cost‑Benefit Analysis and Decision Framework

While end‑to‑end load testing is essential for high‑availability in e‑commerce, finance, and government systems, this article reveals hidden costs—environment, data, and coordination—quantifies benefits such as reduced errors, tighter capacity control, and faster fault detection, and offers a ROI matrix and practical decision guidelines.

Load TestingROIcapacity planning
0 likes · 8 min read
End-to-End Load Testing: Deep Cost‑Benefit Analysis and Decision Framework
Architect's Guide
Architect's Guide
Aug 25, 2026 · Cloud Native

API Gateway vs Load Balancer: How to Choose the Right Traffic Management Tool

This article compares load balancers and API gateways, explaining their layer‑4 vs layer‑7 focus, feature sets such as routing, authentication, observability, and extensibility, and outlines suitable scenarios like microservices, API publishing, and high‑throughput network entry to help readers select the appropriate component.

API GatewayMicroservicesauthentication
0 likes · 11 min read
API Gateway vs Load Balancer: How to Choose the Right Traffic Management Tool
Random Bulletin
Random Bulletin
Aug 24, 2026 · Operations

Scaling Log Volumes from GB to TB at Ten‑Million QPS: Cost‑Effective Strategies and Architecture

At ten‑million QPS, log data can explode from a few gigabytes to terabytes or even petabytes, triggering storage blow‑up, pipeline saturation, slow queries, runaway costs, and poor signal‑to‑noise, and the article breaks down ingest, index, and store costs while presenting edge sampling, label‑based indexing, tiered storage, and log‑to‑metric rollup as mitigation tactics.

ElasticsearchLokicost optimization
0 likes · 19 min read
Scaling Log Volumes from GB to TB at Ten‑Million QPS: Cost‑Effective Strategies and Architecture
DataFunSummit
DataFunSummit
Aug 24, 2026 · Artificial Intelligence

Why Powerful AI Agents Are Becoming More Like Traditional Software

Palantir's new Agent Stack shifts AI agents from short‑lived model‑prompt loops to a production‑grade architecture that adds state, events, effects, durable execution, observability and ontology, turning agents into reliable, governable software components for real‑world business tasks.

AI agentsDurable ExecutionPalantir
0 likes · 11 min read
Why Powerful AI Agents Are Becoming More Like Traditional Software
Woodpecker Software Testing
Woodpecker Software Testing
Aug 24, 2026 · Cloud Native

Microservice Performance Testing: Emerging Trends for the Next Three Years

The article argues that microservice performance testing must evolve from isolated load‑generation to topology‑aware, chaos‑integrated, AI‑driven practices, highlighting upcoming trends such as service‑mesh‑based traffic modeling, built‑in chaos‑as‑a‑test, and large‑language‑model‑assisted root‑cause analysis to prevent cascading failures.

AIOpsMicroserviceschaos engineering
0 likes · 7 min read
Microservice Performance Testing: Emerging Trends for the Next Three Years
Woodpecker Software Testing
Woodpecker Software Testing
Aug 23, 2026 · Operations

Open-Source Performance Testing Strategy: A Practical Guide

Performance bottlenecks cause over 63% of production failures, yet many teams rely on costly commercial tools; this article presents a lightweight, transparent, and evolvable open-source performance testing framework, detailing layered strategies, data-driven feedback loops, and common pitfalls to achieve sustainable quality assurance.

CI/CDchaos engineeringobservability
0 likes · 9 min read
Open-Source Performance Testing Strategy: A Practical Guide
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Aug 23, 2026 · Artificial Intelligence

Why AI Projects Look Great but Perform Poorly? A Practitioner’s Deep Retrospective

The article analyzes why AI projects that shine in proof‑of‑concepts often falter in production, highlighting four core challenges—probabilistic uncertainty, data quality, engineering complexity, and misleading accuracy metrics—and proposes four practical ways to break through these obstacles.

AI DeploymentMLOpsdata governance
0 likes · 13 min read
Why AI Projects Look Great but Perform Poorly? A Practitioner’s Deep Retrospective
LuTiao Programming
LuTiao Programming
Aug 22, 2026 · Backend Development

Why Spring Boot 4.1.1 Makes Controllers and JSON Unnecessary for Internal Microservice Calls

Spring Boot 4.1.1 now offers native gRPC support, letting Java microservices replace the usual REST controllers, DTOs, Feign clients and JSON payloads with a single .proto definition, generated code, and built‑in security, health, observability and testing features, while highlighting best practices and pitfalls.

JavaMicroservicesProtocol Buffers
0 likes · 12 min read
Why Spring Boot 4.1.1 Makes Controllers and JSON Unnecessary for Internal Microservice Calls
Random Bulletin
Random Bulletin
Aug 22, 2026 · Operations

From Manual Trace Queries to Intelligent Link Analysis at Million‑QPS Scale

The article walks through how link tracing evolves from manually searching individual traces to an automated, intelligent system that aggregates massive spans, derives RED metrics and service maps, performs critical‑path and differential analysis, auto‑detects anomalies, and ties together traces, metrics, and logs for rapid root‑cause identification.

RED metricscritical pathdifferential analysis
0 likes · 19 min read
From Manual Trace Queries to Intelligent Link Analysis at Million‑QPS Scale
MaGe Linux Operations
MaGe Linux Operations
Aug 22, 2026 · Operations

Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways

The article analyzes how the emergence of MCP, function calling, and agent toolchains transforms API gateway traffic, identifies blind spots in traditional monitoring, and proposes a three‑layer metric system—including request, inference, and tool‑call dimensions—along with concrete Prometheus metrics, alert rules, and implementation guidelines for reliable observability.

API GatewayMCPPrometheus
0 likes · 33 min read
Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways
Random Bulletin
Random Bulletin
Aug 21, 2026 · Operations

Scaling Link Tracing Storage: From Centralized Elasticsearch to Tiered Architecture

The article analyzes why storing massive tracing spans in a single Elasticsearch cluster fails at high QPS, outlines the four key challenges of trace data, and presents a three‑step engineering solution—sampling, hot‑warm‑cold tiered storage, and separating indexes from span payloads—while comparing major back‑ends such as Elasticsearch, Cassandra, ClickHouse, and Tempo.

ClickHouseElasticsearchTempo
0 likes · 19 min read
Scaling Link Tracing Storage: From Centralized Elasticsearch to Tiered Architecture
DataFunTalk
DataFunTalk
Aug 21, 2026 · Artificial Intelligence

Palantir’s Object Timeline: Measuring Enterprise Agent Work and Observability

Palantir’s new Object Timeline feature aggregates token usage, runtime, waiting time and Agentic Coverage for each business object, turning enterprise AI agents into observable work units and revealing how much work they actually perform, where bottlenecks occur, and what remains human‑driven.

AI AgentAgentic CoverageEnterprise AI
0 likes · 15 min read
Palantir’s Object Timeline: Measuring Enterprise Agent Work and Observability
Data Bricklaying Diary
Data Bricklaying Diary
Aug 21, 2026 · R&D Management

Feature Complete ≠ Production Ready: Why AI Coding Demands Engineering Discipline

This article argues that AI can rapidly generate functional code but cannot lower the engineering bar for production readiness, which requires risk-matched baselines, independent verification, and accountable gates — illustrated through a batch-import example showing the gap between happy-path code and real-world constraints like retries, idempotency, observability, and rollback.

AI-assisted developmentengineering gatesidempotency
0 likes · 21 min read
Feature Complete ≠ Production Ready: Why AI Coding Demands Engineering Discipline
Woodpecker Software Testing
Woodpecker Software Testing
Aug 21, 2026 · Operations

Intelligent, Adaptive, Observable Cache Strategy Testing in 2026

The article examines how cache testing has evolved in 2026 from simple hit‑rate checks to semantic contract verification, AI‑driven dynamic policies, and full‑stack observability, illustrating each shift with real‑world examples, metrics, and adversarial reinforcement testing techniques.

AI‑driven testingadversarial testingcache testing
0 likes · 6 min read
Intelligent, Adaptive, Observable Cache Strategy Testing in 2026
Woodpecker Software Testing
Woodpecker Software Testing
Aug 20, 2026 · Operations

Shift‑Left Performance Testing: Ensuring System Resilience from Early Development

The article explains how shifting performance testing left—embedding performance contracts, automated gates, and cultural practices throughout requirements, design, coding, and integration—prevents costly production incidents, improves defect interception rates, and transforms system resilience into a predictable, built‑in quality attribute.

CI/CDmicrobenchmarkobservability
0 likes · 9 min read
Shift‑Left Performance Testing: Ensuring System Resilience from Early Development
Random Bulletin
Random Bulletin
Aug 19, 2026 · Operations

Sampling at Ten‑Million QPS: From Full to Intelligent Adaptive Sampling

The article explains how sampling serves as the cost‑fidelity knob in distributed tracing, why full sampling collapses at ten‑million QPS, compares head‑based, tail‑based, and intelligent adaptive strategies, and shows how tools like OpenTelemetry, Jaeger, and AWS X‑Ray implement these approaches.

JaegerOpenTelemetryadaptive sampling
0 likes · 19 min read
Sampling at Ten‑Million QPS: From Full to Intelligent Adaptive Sampling
YiSu Grain
YiSu Grain
Aug 19, 2026 · Cloud Native

Day 60 Cloud‑Native Case Study: Service Governance, Reliable Messaging, and Observability

This Day 60 case study walks through a regional medical appointment platform that has been broken into micro‑services on a container cluster, asking you to select and justify service discovery, TCC/Saga, reliable messaging, Kubernetes, Service Mesh and observability measures, and to explain their benefits and trade‑offs.

Distributed TransactionsKubernetesReliable Messaging
0 likes · 36 min read
Day 60 Cloud‑Native Case Study: Service Governance, Reliable Messaging, and Observability
Tencent Cloud Middleware
Tencent Cloud Middleware
Aug 19, 2026 · Operations

How AI Gateway Makes Large-Model Calls Visible, Traceable, and Auditable

Enterprises deploying large-model APIs often struggle to see token usage, latency, and errors; the AI Gateway embeds metrics, structured logs, and distributed tracing at the gateway layer, providing token-level insights, request-level latency breakdowns, and full-chain auditability without code changes, as demonstrated in a real-world incident.

AI GatewayCloud NativeLLM
0 likes · 17 min read
How AI Gateway Makes Large-Model Calls Visible, Traceable, and Auditable
Architecture Development Notes
Architecture Development Notes
Aug 19, 2026 · Artificial Intelligence

Orchestrator-Worker Pattern: Engineering Dynamic Task Decomposition for AI Agents

This article explains the Orchestrator-Worker pattern for AI agents, where an orchestrator dynamically decomposes complex tasks into specialized workers, enabling parallel execution and reducing context interference, with practical engineering considerations for model selection, error handling, and observability.

AI agentsAgent ArchitectureContext Management
0 likes · 14 min read
Orchestrator-Worker Pattern: Engineering Dynamic Task Decomposition for AI Agents
DataFunTalk
DataFunTalk
Aug 18, 2026 · Artificial Intelligence

How Palantir’s Object Timeline Turns AI Agent Activity into Real‑World Production Metrics

Palantir’s new Object Timeline feature moves AI observability from model‑level traces to business‑object lifecycles, exposing token usage, runtime, waiting time and Agentic Coverage for each object, allowing enterprises to quantify how much work agents actually perform, where bottlenecks lie, and why simple automation percentages can be misleading.

AI AgentAgentic CoverageEnterprise AI
0 likes · 15 min read
How Palantir’s Object Timeline Turns AI Agent Activity into Real‑World Production Metrics
Cloud Architecture
Cloud Architecture
Aug 17, 2026 · Backend Development

Go Microservice Stability: Rate Limiting, Circuit Breaking, Degradation and K8s Production Architecture

The article walks through a real‑world traffic spike in an e‑commerce order service, explains why isolated techniques like rate limiting, circuit breaking or degradation are insufficient, and presents a complete, layered stability‑governance solution for Go microservices running on Kubernetes, complete with code, configuration, observability and testing guidance.

Circuit BreakerGoKubernetes
0 likes · 42 min read
Go Microservice Stability: Rate Limiting, Circuit Breaking, Degradation and K8s Production Architecture
Cloud Architecture
Cloud Architecture
Aug 17, 2026 · Backend Development

Comprehensive Guide to Building an Enterprise‑Grade Distributed ID System in Go

This article walks through the full design and production‑ready implementation of a Go‑based distributed ID service, comparing Snowflake and Leaf Segment algorithms, detailing a dual‑engine architecture, SDK caching, scaling on Kubernetes, observability, deployment, and performance testing for high‑throughput enterprise applications.

GoKubernetesLeaf Segment
0 likes · 36 min read
Comprehensive Guide to Building an Enterprise‑Grade Distributed ID System in Go
AI Large Model Application Practice
AI Large Model Application Practice
Aug 17, 2026 · Artificial Intelligence

From Sketching a Graph to Full‑Scale Graph Engineering: Key Practices

The article examines Graph Engineering as the disciplined process of turning multi‑agent collaboration diagrams into reliable, observable, and recoverable production systems, covering basic coordination patterns, state sharing, failure handling, observability, and a comparative look at leading frameworks such as LangGraph, Google ADK, OpenAI Agents SDK, and Claude Dynamic Workflows.

Failure RecoveryGraph Engineeringmulti-agent systems
0 likes · 17 min read
From Sketching a Graph to Full‑Scale Graph Engineering: Key Practices
Cloud Architecture
Cloud Architecture
Aug 14, 2026 · Cloud Native

Complete Guide to Go Microservice Logging and Tracing with OpenTelemetry (Industrial‑Grade Solution)

When an alarm rang at 2:17 AM, a Go order service’s P99 latency surged from 220 ms to 4.6 s and its error rate climbed to 1.8 %; the article explains why many teams still see limited value after adopting OpenTelemetry, identifies three missing pieces—stable trace IDs, end‑to‑end context propagation, and production‑ready pipelines—and delivers a step‑by‑step, code‑first blueprint for building an industrial‑grade observability stack that scales in Kubernetes.

Cloud NativeGoOpenTelemetry
0 likes · 40 min read
Complete Guide to Go Microservice Logging and Tracing with OpenTelemetry (Industrial‑Grade Solution)
Random Bulletin
Random Bulletin
Aug 14, 2026 · Operations

Why Metric Counts Explode from Thousands to Millions in High‑QPS Systems

A single high‑cardinality label can cause a monitoring system to crash as metric series jump from hundreds of thousands to millions, overwhelming storage, queries, collection, cost, and signal‑to‑noise; the article explains the root cause, impact dimensions, and practical mitigation strategies such as cardinality control, distributed TSDBs, downsampling, pre‑aggregation, and proper division of metrics, traces, and logs.

PrometheusTSDBhigh cardinality
0 likes · 17 min read
Why Metric Counts Explode from Thousands to Millions in High‑QPS Systems
Amap Tech
Amap Tech
Aug 14, 2026 · Artificial Intelligence

How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery

The article explains why a single successful AI‑generated code run is insufficient for enterprise software, and how AutoSDK uses built‑in observability, Loop Engineering, and a four‑stage "observe‑attribute‑intervene‑validate" loop—supported by concrete metrics, trace and log pillars—to achieve stable, continuously improving AI coding delivery.

AI codingLoop EngineeringPerformance Optimization
0 likes · 17 min read
How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery
phodal
phodal
Aug 14, 2026 · Artificial Intelligence

How Harness Inspector Makes AI Agent Deliveries Observable, Inspectable, and Traceable

Harness Inspector is a read‑only local workbench that unifies requirements, Agent sessions, file activity, and Git commits into a single interface, enabling developers to trace the full delivery chain from intent through process to output and assess which agent actions merit skill extraction.

AI AgentGitHarness Inspector
0 likes · 10 min read
How Harness Inspector Makes AI Agent Deliveries Observable, Inspectable, and Traceable
Big Data and Microservices
Big Data and Microservices
Aug 14, 2026 · Artificial Intelligence

Why 90% of AI Agent Deployments Fail: The Three Critical Pitfalls

A MIT report shows that 95% of AI Agent pilots flop, and this article breaks down the three common traps—treating agents as a cure‑all, ignoring human‑in‑the‑loop control, and lacking observability—while offering concrete case studies and practical mitigation steps.

AI AgentMIT reportdeployment pitfalls
0 likes · 13 min read
Why 90% of AI Agent Deployments Fail: The Three Critical Pitfalls
Cloud Architecture
Cloud Architecture
Aug 13, 2026 · Cloud Native

Kubernetes Node Maintenance: From Drain to True Zero‑Downtime Engineering

Many teams mistakenly believe that a simple `kubectl drain` guarantees safe node shutdown, but in production the risk spans the control plane, service discovery, long‑lived connections, load balancers and observability; this guide presents a repeatable, auditable, production‑grade process that turns node maintenance into a zero‑interruption engineering workflow.

Graceful ShutdownKubernetesOperator
0 likes · 41 min read
Kubernetes Node Maintenance: From Drain to True Zero‑Downtime Engineering
Woodpecker Software Testing
Woodpecker Software Testing
Aug 13, 2026 · Industry Insights

2026 Stress‑Testing ROI: When Is the Investment Worth It?

The article analyzes how AI‑assisted scenario generation, chaos‑as‑a‑service, and observability reshape stress‑testing costs in 2026, presenting ROI models, industry benchmarks, and critical thresholds that turn testing from a risk hedge into a growth lever.

AI-generated trafficCloud NativeROI
0 likes · 8 min read
2026 Stress‑Testing ROI: When Is the Investment Worth It?
Alibaba Middleware
Alibaba Middleware
Aug 12, 2026 · Operations

How STAROps Detects Unknown Anomalies with Intelligent Log Inspection

STAROps transforms raw logs into actionable insights by clustering log patterns, drilling down across dimensions with AI operators, and using an Agent that dynamically plans investigations, integrates UModel cross‑source mapping, and continuously refines findings to catch unknown anomalies before they become incidents.

AI operatorsCloud NativeUModel
0 likes · 17 min read
How STAROps Detects Unknown Anomalies with Intelligent Log Inspection
Alibaba Cloud Native
Alibaba Cloud Native
Aug 12, 2026 · Cloud Native

Alibaba Cloud and Datadog Release OpenTelemetry Go Compile‑Time Instrumentation v1 for Zero‑Code Observability

The OpenTelemetry Go Compile‑Time Instrumentation project, jointly launched by Alibaba Cloud and Datadog, fills the last observability gap for Go by injecting tracing and metrics code at build time, offering zero‑code instrumentation, no runtime overhead, and seamless CI/CD integration while comparing it with manual and eBPF approaches.

Cloud NativeCompile-Time InstrumentationGo
0 likes · 10 min read
Alibaba Cloud and Datadog Release OpenTelemetry Go Compile‑Time Instrumentation v1 for Zero‑Code Observability