Tagged articles

metrics

628 articles · Page 1 of 7
Data Bricklaying Diary
Data Bricklaying Diary
Sep 24, 2026 · Operations

Observability Isn't a Log Platform: Connecting Business Goals to Verifiable Runtime Facts

This article explains that observability is not centralized logging but a practice of defining business outcomes (SLIs/SLOs), correlating metrics, logs, and traces via stable business identifiers, designing actionable alerts, structuring dashboards along business flows, separating observability responsibilities, halting automation when evidence is unreliable, and using runtime data to continuously correct architectural assumptions.

ObservabilitySLI/SLOSRE
0 likes · 28 min read
Observability Isn't a Log Platform: Connecting Business Goals to Verifiable Runtime Facts
Ops Development & AI Practice
Ops Development & AI Practice
Sep 17, 2026 · Cloud Native

Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained

This article explains why OpenTelemetry Helm charts now recommend deploying Collector as three separate releases—otel-agent (DaemonSet for node metrics), otel-cluster (singleton Deployment for cluster metrics), and otel-gateway (scalable Deployment for trace ingestion)—detailing Presets simplification, lifecycle isolation, failure domains, and when to consolidate to two releases.

CollectorDaemonSetHelm
0 likes · 24 min read
Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained
FunTester
FunTester
Aug 29, 2026 · Operations

How to Shift Performance Testing Left in Development

The article explains why discovering performance problems just before release is risky and proposes moving performance testing earlier in the software lifecycle—defining measurable goals during requirements, validating components during development, checking service boundaries during integration, and using CI/CD for continuous feedback.

CI/CDLoad Testingearly testing
0 likes · 12 min read
How to Shift Performance Testing Left in Development
Random Bulletin
Random Bulletin
Aug 28, 2026 · Operations

Log Correlation: From Isolated Entries to Linked Traces in High‑QPS Systems

The article explains how to turn millions of independent error logs into a coherent, request‑level waterfall and business‑level story by injecting trace_id, business keys, and cross‑signal foreign keys, while addressing async boundaries, sampling, naming consistency, and storage costs.

Observabilitydistributed tracinglog correlation
0 likes · 18 min read
Log Correlation: From Isolated Entries to Linked Traces in High‑QPS Systems
MaGe Linux Operations
MaGe Linux Operations
Aug 22, 2026 · Operations

Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways

The article analyzes how the emergence of MCP, function calling, and agent toolchains transforms API gateway traffic, identifies blind spots in traditional monitoring, and proposes a three‑layer metric system—including request, inference, and tool‑call dimensions—along with concrete Prometheus metrics, alert rules, and implementation guidelines for reliable observability.

API GatewayMCPObservability
0 likes · 33 min read
Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways
DataFunTalk
DataFunTalk
Aug 21, 2026 · Artificial Intelligence

Palantir’s Object Timeline: Measuring Enterprise Agent Work and Observability

Palantir’s new Object Timeline feature aggregates token usage, runtime, waiting time and Agentic Coverage for each business object, turning enterprise AI agents into observable work units and revealing how much work they actually perform, where bottlenecks occur, and what remains human‑driven.

AI AgentAgentic CoverageEnterprise AI
0 likes · 15 min read
Palantir’s Object Timeline: Measuring Enterprise Agent Work and Observability
Smart Sea Tide
Smart Sea Tide
Aug 21, 2026 · Big Data

Designing Scalable Data Warehouses: Architecture, Modeling, Scheduling, and Metric Construction

The article explains how the surge of data in the DT era makes traditional storage insufficient, defines a data warehouse as a subject‑oriented, integrated, stable collection for decision support, outlines its development lifecycle—including integration, modeling, services, scheduling, metadata and quality management—and stresses that building a warehouse is an ongoing, iterative process driven by evolving business needs.

Data ModelingETLarchitecture
0 likes · 3 min read
Designing Scalable Data Warehouses: Architecture, Modeling, Scheduling, and Metric Construction
Tencent Cloud Middleware
Tencent Cloud Middleware
Aug 19, 2026 · Operations

How AI Gateway Makes Large-Model Calls Visible, Traceable, and Auditable

Enterprises deploying large-model APIs often struggle to see token usage, latency, and errors; the AI Gateway embeds metrics, structured logs, and distributed tracing at the gateway layer, providing token-level insights, request-level latency breakdowns, and full-chain auditability without code changes, as demonstrated in a real-world incident.

AI GatewayLLMObservability
0 likes · 17 min read
How AI Gateway Makes Large-Model Calls Visible, Traceable, and Auditable
Linyb Geek Road
Linyb Geek Road
Aug 15, 2026 · Operations

Key Metrics Every Ops Engineer Should Monitor

This article enumerates essential operational metrics—such as CPU, memory, disk and network I/O, response time, throughput, error rates, availability, MTBF/MTTR, security logs, and capacity‑planning indicators—explaining their meanings and recommended target values to help engineers comprehensively monitor system performance, stability, and efficiency.

OperationsPerformanceavailability
0 likes · 10 min read
Key Metrics Every Ops Engineer Should Monitor
Linyb Geek Road
Linyb Geek Road
Aug 15, 2026 · Operations

17 Essential IT Operations Metrics Everyone Should Know (AI Not Required)

The article outlines why monitoring key IT operations metrics is vital for performance, reliability, and cost control, then details 17 common metrics—including availability, failure rate, MTTR, MTBF, response time, throughput, error rate, capacity utilization, latency, data integrity, success rates, waiting time, backup success, recovery time, security patch time, server and network bandwidth utilization—providing definitions, calculation formulas, typical reference values, and applicable scenarios.

Capacity UtilizationIT OperationsMTTR
0 likes · 7 min read
17 Essential IT Operations Metrics Everyone Should Know (AI Not Required)
Amap Tech
Amap Tech
Aug 14, 2026 · Artificial Intelligence

How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery

The article explains why a single successful AI‑generated code run is insufficient for enterprise software, and how AutoSDK uses built‑in observability, Loop Engineering, and a four‑stage "observe‑attribute‑intervene‑validate" loop—supported by concrete metrics, trace and log pillars—to achieve stable, continuously improving AI coding delivery.

AI codingLoop EngineeringObservability
0 likes · 17 min read
How AutoSDK Builds a Self‑Evolving AI Coding Loop for Enterprise Delivery
samdeepthink
samdeepthink
Aug 12, 2026 · Operations

Why Observability Is More Than Monitoring: Finding the Root Cause Quickly

The article explains that observability goes beyond simple monitoring by combining metrics, logs, and traces to pinpoint where and why a system issue occurs, especially in microservice and cloud‑native environments, and stresses the importance of correlating data rather than merely collecting more.

AIOpsLogsObservability
0 likes · 3 min read
Why Observability Is More Than Monitoring: Finding the Root Cause Quickly
AgentGuide
AgentGuide
Aug 6, 2026 · Artificial Intelligence

How to Evaluate RAG Systems? Key Metrics and Frameworks Used in Projects

The article explains how to assess Retrieval‑Augmented Generation (RAG) projects using the open‑source Ragas framework, detailing four evaluation dimensions and breaking down specific retrieval and generation metrics such as precision, recall, answer correctness, relevance, and faithfulness.

AIRAGRAGAS
0 likes · 4 min read
How to Evaluate RAG Systems? Key Metrics and Frameworks Used in Projects
Linyb Geek Road
Linyb Geek Road
Aug 2, 2026 · Operations

How to Build a Systematic Enterprise Monitoring Architecture

This article outlines a comprehensive, step‑by‑step approach for constructing a systematic enterprise monitoring system, covering the four core technical modules (collection, data, operators, alerts), designing a layered metric framework, and establishing a health‑management lifecycle that includes proactive alert prevention, real‑time handling, and post‑incident review.

CMDBObservabilitySRE
0 likes · 21 min read
How to Build a Systematic Enterprise Monitoring Architecture
Random Bulletin
Random Bulletin
Jul 31, 2026 · Cloud Native

Scaling at Ten‑Million QPS: From Manual to Automatic Autoscaling

The article analyzes why manual capacity adjustments break down at ten‑million‑QPS scale, then walks through metric‑driven autoscaling, anti‑flapping algorithms, headroom planning, predictive scaling, stateful service challenges, and multi‑dimensional strategies to achieve a cost‑stable dynamic balance.

AutoscalingKubernetescapacity-planning
0 likes · 20 min read
Scaling at Ten‑Million QPS: From Manual to Automatic Autoscaling
AliExpress Tech
AliExpress Tech
Jul 21, 2026 · Artificial Intelligence

Fine-Grained Evaluation of AI Agents: Designing a Comprehensive Testing Framework

This article presents a comprehensive, fine-grained evaluation framework for AI agents that moves beyond traditional text-similarity metrics, defines architecture-aligned quality, cost, and performance indicators for each core module, describes dataset construction, LLM-as-Judge tasks, execution engine, and visual dashboards, and shares practical lessons and future directions.

AI AgentLLM-as-judgeTesting Framework
0 likes · 45 min read
Fine-Grained Evaluation of AI Agents: Designing a Comprehensive Testing Framework
Data Integration and Governance
Data Integration and Governance
Jul 15, 2026 · Big Data

How to Build a Data Warehouse: End-to-End Process from Source Data to Analytics

Many companies start data‑warehouse projects by merely extracting data, building a few tables and adding a BI report, only to face inconsistent metrics, unreadable tables, mismatched numbers and endless Excel rechecks; the article outlines a full end‑to‑end process—from source‑data inventory and stable ingestion to layered storage, modeling, metric unification and quality monitoring—to ensure trustworthy, reusable analytics.

AnalyticsData ModelingData Quality
0 likes · 17 min read
How to Build a Data Warehouse: End-to-End Process from Source Data to Analytics
Go Development Architecture Practice
Go Development Architecture Practice
Jul 14, 2026 · Operations

Embedded Monitoring Best Practice: Use go-commons for Built-in Service Health Reports

This article demonstrates how to quickly add lightweight, plug‑and‑play monitoring to a Go service using the open‑source go-commons library, showing installation, a minimal 50‑line example that exposes business QPS and system metrics via a single /metrics endpoint, and how to integrate it with Prometheus and Grafana.

GoPrometheusgo-commons
0 likes · 6 min read
Embedded Monitoring Best Practice: Use go-commons for Built-in Service Health Reports
ThinkingAgent
ThinkingAgent
Jul 14, 2026 · Operations

Why AI Agents Need Observability: Tracing, Monitoring, and SRE in the L8 Layer

A recent fintech chatbot failure exposed how missing tracing, cost attribution, and proper alerting can turn a three‑day incident into a three‑day investigation, prompting a detailed guide on the L8 observability layer that defines three pillars—Tracing, Metrics, Logs—and outlines best‑practice tooling, standards, and implementation steps for AI production systems.

AI ObservabilityCost attributionLangfuse
0 likes · 36 min read
Why AI Agents Need Observability: Tracing, Monitoring, and SRE in the L8 Layer
AI Engineer Programming
AI Engineer Programming
Jul 11, 2026 · Operations

Building an Observability Platform for LLM Agents with OpenTelemetry

This article explains why LLM agents need a dedicated observability platform, introduces OpenTelemetry’s core concepts and architecture, shows how to manually instrument Python code, enable automatic instrumentation, configure the Collector, handle common distributed‑system pitfalls, and extend OTel with agent‑specific semantics and evaluation loops.

CollectorLLM AgentObservability
0 likes · 20 min read
Building an Observability Platform for LLM Agents with OpenTelemetry
Big Data and Microservices
Big Data and Microservices
Jul 9, 2026 · Artificial Intelligence

How to Evaluate and Observe AI Agents: Optimizing Your Digital Employee

The article explains why traditional benchmark scores are insufficient for production AI agents and proposes a four‑dimensional evaluation framework—task success, step efficiency, cost, and safety—combined with an observability stack of metrics, structured logs, and full‑trace decision snapshots to continuously measure, debug, and improve digital employees.

AI agentsCost ManagementLLM
0 likes · 17 min read
How to Evaluate and Observe AI Agents: Optimizing Your Digital Employee
Xike
Xike
Jul 6, 2026 · Backend Development

Seeing the Full Request Journey: Completing Spring Boot Trace Integration

This guide shows how to extend an existing Prometheus‑Grafana‑Loki stack with SkyWalking to capture full request traces in Spring Boot, explaining trace fundamentals, automatic instrumentation, manual spans, log‑trace correlation, cross‑service topology, and production considerations.

LogsObservabilitySkyWalking
0 likes · 17 min read
Seeing the Full Request Journey: Completing Spring Boot Trace Integration
Raymond Ops
Raymond Ops
Jul 5, 2026 · Operations

Building a Basic Monitoring System from Zero: How to View CPU, Memory, Disk, and Network

This article walks you through setting up a complete monitoring stack with Prometheus, node_exporter, Grafana and Alertmanager, explains how to interpret the four core dimensions—CPU, memory, disk and network—using a structured troubleshooting workflow, and provides real‑world case studies, scripts and best‑practice recommendations.

AlertmanagerGrafanaLinux
0 likes · 35 min read
Building a Basic Monitoring System from Zero: How to View CPU, Memory, Disk, and Network
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jun 25, 2026 · Artificial Intelligence

AI Coding in Practice: Insights from ByteDance’s VP of Technology

ByteDance’s AI coding effort has grown over six‑fold in contribution rate, but the team highlights three real challenges—over‑reliance on simple metrics, turning fast Vibe Coding into stable deliverables, and coordinating diverse roles—offering data‑driven experiments and a systematic AI development roadmap.

AI codingHarnessSoftware Engineering
0 likes · 13 min read
AI Coding in Practice: Insights from ByteDance’s VP of Technology
Alibaba Cloud Observability
Alibaba Cloud Observability
Jun 22, 2026 · Cloud Native

Zero‑Code Full‑Stack Observability with OpenTelemetry eBPF: CloudMonitor 2.0’s In‑Kernel “Lens”

OpenTelemetry eBPF Instrumentation (OBI) injects a kernel‑level, zero‑code probe that automatically captures OpenTelemetry‑compatible traces, metrics, and logs for over 15 protocols—including HTTP, gRPC, MySQL, Redis, Kafka, and CUDA—while handling cross‑language context propagation, GPU tracing, and seamless integration with CloudMonitor 2.0.

ObservabilityOpenTelemetryZero-Code Monitoring
0 likes · 19 min read
Zero‑Code Full‑Stack Observability with OpenTelemetry eBPF: CloudMonitor 2.0’s In‑Kernel “Lens”
Code Mala Tang
Code Mala Tang
Jun 16, 2026 · Industry Insights

GitHub Star Inflation: Why 10,000 Stars No Longer Impress

The article analyzes how GitHub star counts have inflated—especially for AI tools—showing that a 20,000‑star threshold that once guaranteed a top‑10 spot now falls short, and explains why stars are becoming a noisy attention metric rather than a reliable quality indicator.

AIGitHubfake stars
0 likes · 13 min read
GitHub Star Inflation: Why 10,000 Stars No Longer Impress
Alibaba Cloud Observability
Alibaba Cloud Observability
Jun 15, 2026 · Cloud Native

Measuring AI Coding Impact from Individual to Organization with LoongSuite‑Pilot and SLS

This article details how LoongSuite‑Pilot captures heterogeneous AI coding agent events and leverages Alibaba Cloud Log Service (SLS) SQL dashboards to provide end‑to‑end, organization‑wide metrics—covering individual usage, team adoption, token consumption, skill and tool utilization—enabling R&D teams to quantify the real‑world effectiveness of AI coding assistants.

AI codingCloud LoggingDevOps
0 likes · 21 min read
Measuring AI Coding Impact from Individual to Organization with LoongSuite‑Pilot and SLS
Coder Trainee
Coder Trainee
Jun 13, 2026 · Artificial Intelligence

AI Agent Observability and Debugging: Building a Transparent Agent System

This article explains why AI agents behave like black boxes, introduces a three‑pillar observability framework (tracing, metrics, logging), demonstrates practical tracing with LangSmith and LangFuse, shows how to instrument agents with custom metrics, evaluate performance, and share best‑practice guidelines for production‑ready debugging.

AI AgentLangChainLangSmith
0 likes · 19 min read
AI Agent Observability and Debugging: Building a Transparent Agent System
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 12, 2026 · Operations

Why Open‑Source LoongSuite Pilot Is Needed as AI Coding Agents Become Core Infrastructure

The article analyzes how AI coding agents like Cursor, Claude Code, and Codex have become essential developer tools, yet suffer from almost zero observability, and explains how the open‑source LoongSuite Pilot provides a unified collection platform, semantic schema, security controls, dashboards, and ROI metrics to turn these agents into manageable infrastructure.

AI coding agentLoongSuite PilotObservability
0 likes · 27 min read
Why Open‑Source LoongSuite Pilot Is Needed as AI Coding Agents Become Core Infrastructure
TechVision Expert Circle
TechVision Expert Circle
Jun 9, 2026 · Artificial Intelligence

How CIOs Can Stop Being the Scapegoat in AI Projects

The article explains why many CIOs become blamed for AI project failures and provides a three‑layer governance framework, engineering‑focused architecture choices, a concrete observability and metrics system, and four actionable steps to turn the CIO into a responsible leader rather than a fall‑guy.

AI GovernanceAI architectureAgent Framework
0 likes · 14 min read
How CIOs Can Stop Being the Scapegoat in AI Projects
Alibaba Cloud Native
Alibaba Cloud Native
Jun 9, 2026 · Cloud Native

From Individual Productivity to Organizational Insight: Building AI Coding Metrics with LoongSuite‑Pilot and SLS

The article explains how to capture event‑level AI coding agent data using LoongSuite‑Pilot, align it with the LoongSuite GenAI semantic conventions, store it in Alibaba Cloud Log Service (SLS), and construct a multi‑layered SQL dashboard that turns personal usage signals into organization‑wide metrics for informed decision‑making.

AIDevOpsObservability
0 likes · 25 min read
From Individual Productivity to Organizational Insight: Building AI Coding Metrics with LoongSuite‑Pilot and SLS
Cloud Architecture
Cloud Architecture
May 30, 2026 · Operations

How to Build Production‑Grade Observability Metrics and Alerting for Batch Jobs

The article explains why batch processing tasks often slip out of control, defines a four‑layer observability model covering status, progress, quality and performance, proposes a unified task state machine and event flow, and provides concrete metric, logging, tracing and alerting designs—including Go and Java SDK examples—for reliable production‑level batch job monitoring.

Batch ProcessingGoObservability
0 likes · 34 min read
How to Build Production‑Grade Observability Metrics and Alerting for Batch Jobs
Tech Stroll Journey
Tech Stroll Journey
May 25, 2026 · Operations

How Linux Sends a Packet: From Process to NIC and the Key Metrics to Watch

The article walks through the Linux packet lifecycle—from the send() system call, through the transport and network layers, to the NIC driver—explaining each step, virtual‑network abstractions, and the essential bandwidth, latency, loss, conntrack, and socket buffer metrics to monitor when problems arise.

LinuxTCP/IPcontainer-networking
0 likes · 10 min read
How Linux Sends a Packet: From Process to NIC and the Key Metrics to Watch
Cloud Architecture
Cloud Architecture
May 23, 2026 · Cloud Native

Build a Production-Ready Observability Platform with OpenTelemetry

To solve fragmented monitoring in Java microservices, the article details how to construct a production‑grade observability platform using OpenTelemetry, covering unified data models, collector architecture, tracing, metrics, logging, sampling strategies, Kubernetes deployment, and practical guidelines for scaling, governance, and root‑cause analysis.

KubernetesObservabilityOpenTelemetry
0 likes · 37 min read
Build a Production-Ready Observability Platform with OpenTelemetry
Coder Trainee
Coder Trainee
May 21, 2026 · Cloud Native

Building Full Observability for Spring Cloud Microservices with Micrometer, Prometheus, and Grafana

After solving distributed transactions with Seata, this tutorial shows how to add complete observability to Spring Cloud microservices by integrating Micrometer, Prometheus, and Grafana, covering metrics pillars, configuration, custom business metrics, dashboard setup, alert rules, validation steps, and common pitfalls.

GrafanaMicrometerObservability
0 likes · 12 min read
Building Full Observability for Spring Cloud Microservices with Micrometer, Prometheus, and Grafana
TechVision Expert Circle
TechVision Expert Circle
May 19, 2026 · R&D Management

How to Make Your Tech Team Visible and Valuable to Management

Even with long hours, on‑time deliveries, and decreasing incidents, many tech teams struggle to show their business impact, leading to budget cuts; this article analyzes why the value translation breaks down and offers a concrete visualization framework to align engineering work with company metrics.

FinOpsR&D managementValue Visualization
0 likes · 10 min read
How to Make Your Tech Team Visible and Valuable to Management
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
May 19, 2026 · Backend Development

Why Logs Alone Fail in Spring Boot: Achieving True Observability

The article explains that relying solely on log statements in Spring Boot applications cannot reveal request identities, latency, async task health, failure details, or cross‑service flows, and demonstrates how to augment logs with MDC correlation IDs, Micrometer metrics, and Zipkin tracing for comprehensive observability.

MicrometerObservabilityZipkin
0 likes · 9 min read
Why Logs Alone Fail in Spring Boot: Achieving True Observability
DeepNoMind
DeepNoMind
May 12, 2026 · R&D Management

Managing Engineering Teams in the AI Era: Rethinking Processes, Structure, and Metrics

Fiona Fung explains how the Claude Code team shifted from code‑centric bottlenecks to verification, review, cross‑functional collaboration and safety, cutting outdated processes, flattening the org, leveraging Claude for PR automation, shifting left testing, and measuring impact with new onboarding, PR‑cycle, and AI‑assisted commit metrics.

AIClaude CodeFlat Organization
0 likes · 18 min read
Managing Engineering Teams in the AI Era: Rethinking Processes, Structure, and Metrics
AgentGuide
AgentGuide
May 3, 2026 · Artificial Intelligence

How to Evaluate an AI Agent Beyond Just Accuracy

Evaluating AI agents requires more than accuracy; you must measure task completion, execution trace, tool usage, latency, cost, error rates, and both explicit and implicit user feedback, using observability, offline smoke‑test and regression suites, and continuous online monitoring to create a closed‑loop improvement process.

AI AgentObservabilityOnline Testing
0 likes · 14 min read
How to Evaluate an AI Agent Beyond Just Accuracy
AI Engineer Programming
AI Engineer Programming
May 2, 2026 · Artificial Intelligence

From Demo to Production: How to Evaluate RAG Effectively

This guide outlines a comprehensive RAG evaluation framework covering failure modes, multi‑layer metrics, test‑set construction, open‑source tools, CI/CD quality gates, production monitoring, and special considerations for agentic RAG to ensure reliable, trustworthy retrieval‑augmented generation systems.

AILLMRAG
0 likes · 18 min read
From Demo to Production: How to Evaluate RAG Effectively
Alibaba Cloud Native
Alibaba Cloud Native
Apr 26, 2026 · Cloud Native

Seeing Inside Hermes: Full Visibility into Agent Execution with OpenTelemetry

The article introduces Alibaba Cloud's Hermes observability plugin built on OpenTelemetry, which transforms the previously opaque AI agent runtime into a fully traceable system by recording every reasoning step, tool invocation, token usage, latency, and security event, enabling precise cost attribution, performance analysis, and audit of high‑risk behaviors.

AI AgentHermesObservability
0 likes · 13 min read
Seeing Inside Hermes: Full Visibility into Agent Execution with OpenTelemetry
Smart Workplace Lab
Smart Workplace Lab
Apr 19, 2026 · Industry Insights

How to Turn AI-Boosted Productivity into Visible Performance Metrics

This article presents a practical framework for documenting AI‑enhanced work contributions, introducing a weekly performance‑evidence matrix that quantifies decision density, risk interception, and asset accumulation, along with communication scripts tailored to different manager types and step‑by‑step SOPs for archiving proof, helping professionals turn speed gains into measurable performance value.

AIEvidencePerformance
0 likes · 7 min read
How to Turn AI-Boosted Productivity into Visible Performance Metrics
PMTalk Product Manager Community
PMTalk Product Manager Community
Apr 10, 2026 · Artificial Intelligence

Why AI Product Evaluation Is Hard and How to Build a Scientific Assessment Framework

The article analyzes the unique challenges of evaluating AI products—output uncertainty, subjective criteria, over‑fitting risk, high cost, and vague metrics—compares traditional testing with AI testing, proposes a five‑step evaluation workflow, defines concrete metrics such as pass rate and efficiency gain, and illustrates the process with a real‑world sales‑script generation case study, concluding with five key success factors and future trends.

AI evaluationautomationcase study
0 likes · 13 min read
Why AI Product Evaluation Is Hard and How to Build a Scientific Assessment Framework
AI Step-by-Step
AI Step-by-Step
Apr 8, 2026 · Operations

How to Light Up the Black Box of LLM Agents with Full‑Stack Observability

The article explains why traditional logs are insufficient for LLM agents, outlines five observability dimensions—tracing, metrics, behavioral governance, state & memory, and evaluation—and provides concrete, open‑source‑based steps to instrument, monitor, and act on agent workloads in production.

Behavioral GovernanceLLM agentsObservability
0 likes · 11 min read
How to Light Up the Black Box of LLM Agents with Full‑Stack Observability
MaGe Linux Operations
MaGe Linux Operations
Apr 6, 2026 · Operations

Master Redis Monitoring: Essential Metrics, Scripts, and Alerting Strategies

This guide walks operations engineers through building a complete Redis monitoring system—covering why monitoring matters, which metrics to collect, how to gather them with Prometheus and Grafana, and practical Bash scripts for health checks, memory, persistence, replication, client connections, and alert thresholds.

GrafanaPrometheusRedis
0 likes · 31 min read
Master Redis Monitoring: Essential Metrics, Scripts, and Alerting Strategies
Alibaba Cloud Native
Alibaba Cloud Native
Apr 5, 2026 · Operations

How OpenClaw CMS Plugin v0.1.2 Turns Agent Tracing into Precise, Cost‑Effective Observability

The OpenClaw CMS observability plugin v0.1.2 solves the hidden‑trace problem by fully restoring multi‑round LLM execution, stabilizing concurrent chains, and introducing granular agent metrics, enabling developers, testers, and operators to debug faster, assess costs accurately, and improve cross‑team collaboration.

AgentObservabilityOpenClaw
0 likes · 8 min read
How OpenClaw CMS Plugin v0.1.2 Turns Agent Tracing into Precise, Cost‑Effective Observability
AgentGuide
AgentGuide
Apr 3, 2026 · Artificial Intelligence

How to Evaluate RAG Systems: Key Metrics and the Ragas Framework

The article explains how to assess Retrieval-Augmented Generation (RAG) projects using the Ragas automated evaluation framework, detailing four key dimensions—recall quality, answer faithfulness, answer relevance, and context utilization—and describes the underlying metrics for both retrieval and generation stages.

LLMRAGRAGAS
0 likes · 5 min read
How to Evaluate RAG Systems: Key Metrics and the Ragas Framework
DevOps Coach
DevOps Coach
Mar 26, 2026 · Industry Insights

Which DevOps Metrics Will Drive Business Success by 2026?

The article analyzes how traditional DevOps activity metrics are being replaced by outcome‑focused indicators that directly affect cost, delivery speed, reliability and overall business performance, citing New Relic and Flexera forecasts and outlining the metrics teams should adopt or discard by 2026.

DORADevOpsFinOps
0 likes · 13 min read
Which DevOps Metrics Will Drive Business Success by 2026?
Big Data Tech Team
Big Data Tech Team
Mar 18, 2026 · Big Data

From Zero to One: Building Enterprise Data Standards for Data Warehouses

This guide explains why data standards are essential for data warehouses, outlines the four categories of standards, and provides a step‑by‑step process—including research, framework design, template creation, review, implementation, and ongoing maintenance—to help practitioners and interviewees establish robust, business‑aligned data standards.

data standardizationdata warehousemetrics
0 likes · 10 min read
From Zero to One: Building Enterprise Data Standards for Data Warehouses
Woodpecker Software Testing
Woodpecker Software Testing
Mar 15, 2026 · R&D Management

Shift‑Left Testing: Transforming Teams from Reactive Bug‑Fixers to Proactive Quality Architects

The article explains how shift‑left testing evolves from a simple early‑testing tactic into a comprehensive team transformation that embeds quality into every stage of software delivery, detailing new roles, metrics, toolchains, and practical steps for test experts to become quality architects.

metricsquality engineeringshift-left testing
0 likes · 8 min read
Shift‑Left Testing: Transforming Teams from Reactive Bug‑Fixers to Proactive Quality Architects
PMTalk Product Manager Community
PMTalk Product Manager Community
Mar 15, 2026 · Product Management

7-Step Architecture Framework for AI Product Management: A Hands‑On Case Study

This article walks through a real‑world AI‑driven image generation system for cross‑border e‑commerce, detailing business pain points, stakeholder analysis, technical selection, MVP scope, architecture decisions, metric funnels, gray‑release strategy, and continuous evolution that cut per‑image cost to under ¥0.5 and delivery time to one minute.

AIarchitecturecase study
0 likes · 16 min read
7-Step Architecture Framework for AI Product Management: A Hands‑On Case Study
PMTalk Product Manager Community
PMTalk Product Manager Community
Mar 13, 2026 · Product Management

How AI Product Managers Should Rethink Funnel Analysis

In the AI era the classic funnel of exposure‑click‑register‑retain‑pay no longer reflects value creation, so product managers must shift the focus to effective task entry, first usable results, mid‑funnel adoption, retention of high‑impact tasks, and stable commercial metrics.

AIFunnel Analysisgrowth
0 likes · 24 min read
How AI Product Managers Should Rethink Funnel Analysis
Architect-Kip
Architect-Kip
Mar 4, 2026 · Operations

Essential SRE Monitoring and Alerting Standards: From Metrics to Incident Response

This guide outlines comprehensive SRE monitoring and alerting standards, covering core principles, log instrumentation, health‑check requirements, baseline resource and application metrics, alarm severity tiers, response SLAs, on‑call rotation, continuous optimization, and noise‑reduction mechanisms to ensure reliable service operation.

OperationsSREalerting
0 likes · 14 min read
Essential SRE Monitoring and Alerting Standards: From Metrics to Incident Response
Data Integration and Governance
Data Integration and Governance
Mar 4, 2026 · Big Data

9 Quantitative Metrics to Evaluate Your Data Warehouse—A Complete Guide

The article presents nine concrete, formula‑based metrics across completeness, reuse, and compliance dimensions—such as cross‑layer reference rate, summary query ratio, model reuse coefficient, lineage divergence, field description coverage, layering info coverage, domain ownership, naming compliance, and field‑consistency—to objectively assess data‑warehouse health and guide continuous improvement.

completenessdata governancedata warehouse
0 likes · 10 min read
9 Quantitative Metrics to Evaluate Your Data Warehouse—A Complete Guide
Alibaba Cloud Native
Alibaba Cloud Native
Mar 2, 2026 · Artificial Intelligence

How to Make AI Agents Auditable and Controlled with OpenClaw, SLS, and OTEL

This article explains how to combine OpenClaw session logs, application logs, and OpenTelemetry metrics in Alibaba Cloud SLS to answer who triggered an AI agent, what actions were taken, how much it cost, and whether the behavior is traceable, enabling a complete observability and security solution for AI agents.

AI AgentOTELObservability
0 likes · 26 min read
How to Make AI Agents Auditable and Controlled with OpenClaw, SLS, and OTEL
Woodpecker Software Testing
Woodpecker Software Testing
Mar 1, 2026 · Artificial Intelligence

Four Hidden Model Evaluation Pitfalls That Undermine AI Deployments

The article examines four common yet hidden model evaluation mistakes—confusing attractive metrics with business impact, using static test sets, ignoring statistical significance, and lacking fine‑grained attribution—illustrating each with real‑world cases and offering concrete practices to build a more robust, business‑aligned evaluation pipeline.

A/B testingAI Deploymentconcept drift
0 likes · 8 min read
Four Hidden Model Evaluation Pitfalls That Undermine AI Deployments
Yunqi AI+
Yunqi AI+
Feb 22, 2026 · R&D Management

Rethinking Product Development: How AI Reshapes the Value Stream, Not Just Code Speed

The article analyzes how AI has evolved from a code‑completion aid to a foundational operating system that forces product‑research teams to redesign the entire requirement‑to‑delivery value stream, outlining practical boundaries, pilot implementation, organizational role changes, metric shifts, and risk governance.

AIR&D managementSoftware Engineering
0 likes · 17 min read
Rethinking Product Development: How AI Reshapes the Value Stream, Not Just Code Speed
dbaplus Community
dbaplus Community
Feb 8, 2026 · Databases

Why Oracle AWR Is the Gold Standard for DB Performance and How Domestic Databases Compare

The article explains Oracle's Automatic Workload Repository (AWR) as a comprehensive performance‑diagnostic tool, breaks down its core functions, and then evaluates how several domestic databases such as Kingbase measure up in terms of report completeness, metric richness, SQL analysis, wait‑event handling, OS integration, and usability.

AWRDiagnosticsDomestic databases
0 likes · 21 min read
Why Oracle AWR Is the Gold Standard for DB Performance and How Domestic Databases Compare
Raymond Ops
Raymond Ops
Feb 2, 2026 · Operations

10 Essential PromQL Queries Every Ops Engineer Should Master

This article presents ten practical PromQL query examples covering CPU, memory, disk, network, database, Kubernetes, and business metrics, explains the underlying concepts, provides alert thresholds and best‑practice tips, and includes advanced optimization and alert‑rule design guidance for reliable monitoring.

ObservabilityPromQLPrometheus
0 likes · 22 min read
10 Essential PromQL Queries Every Ops Engineer Should Master
Ops Community
Ops Community
Jan 27, 2026 · Operations

Master Linux System Monitoring: Deep Dive into CPU, Memory, and I/O Metrics

This comprehensive guide explains how to collect and analyze Linux system metrics—including CPU usage, memory consumption, disk I/O, and load average—using native /proc and /sys interfaces, popular command‑line tools, and Prometheus Node Exporter, with practical scripts, configuration examples, and troubleshooting case studies for reliable performance monitoring and capacity planning.

LinuxPrometheusmetrics
0 likes · 39 min read
Master Linux System Monitoring: Deep Dive into CPU, Memory, and I/O Metrics
PMTalk Product Manager Community
PMTalk Product Manager Community
Jan 18, 2026 · Product Management

Cut Through the Fog: How Product Managers Can Re‑Anchor Value and Evolve

Amid slowing growth and noisy data, product managers face three crises—demand fog, value vacuum, and capability gaps; the article offers a step‑by‑step framework with real‑world cases to clarify user needs, align actions with business goals, strengthen technical and analytical skills, and make data‑driven decisions that turn feature work into measurable value.

decision-makinggrowth strategiesmetrics
0 likes · 14 min read
Cut Through the Fog: How Product Managers Can Re‑Anchor Value and Evolve
Woodpecker Software Testing
Woodpecker Software Testing
Jan 13, 2026 · User Experience Design

A Complete User Experience Testing Process: From Planning to Implementation

The article outlines a systematic, end‑to‑end UX testing workflow—defining goals, designing test plans, recruiting representative users, preparing materials, calibrating and managing test sessions, collecting quantitative and qualitative data, analyzing results with metrics like SUS and efficiency index, extracting actionable insights, and converting findings into concrete product improvements—highlighting how AI‑driven tools can boost test efficiency and business value.

AI Testing ToolsUX Researchmetrics
0 likes · 7 min read
A Complete User Experience Testing Process: From Planning to Implementation
Programmer DD
Programmer DD
Jan 12, 2026 · Artificial Intelligence

5 Counterintuitive Lessons for Evaluating AI Agents Effectively

This article shares five surprising, high‑impact lessons from Anthropic on building robust AI agent evaluation suites, covering early failure‑case collections, recognizing clever “failures,” focusing on outcomes over process, choosing the right success metrics, and the irreplaceable value of human review.

AI evaluationAnthropicagent testing
0 likes · 10 min read
5 Counterintuitive Lessons for Evaluating AI Agents Effectively
Huolala Tech
Huolala Tech
Jan 7, 2026 · Operations

How Exemplar Bridges the Last‑Mile Gap in Observability

Facing the “last mile” challenge of correlating metrics, logs, and traces, the article examines common heterogeneous storage architectures, critiques existing Exemplar implementations, and presents HuoLala’s end‑to‑end solution that treats Exemplar as an independent observable dimension, detailing its data model, SDK integration, collector, and interactive visualization.

ExemplarLogAggregationObservability
0 likes · 22 min read
How Exemplar Bridges the Last‑Mile Gap in Observability
Woodpecker Software Testing
Woodpecker Software Testing
Jan 5, 2026 · Backend Development

Five Core Dimensions of Maintainability Testing for Microservice Systems

This article presents a detailed, step‑by‑step guide to maintainability testing, defining five core dimensions—modularization, reusability, analysability, modifiability, and testability—along with their metrics, a relationship model, a comprehensive microservice e‑shop case study, concrete test scenarios, code examples, and best‑practice recommendations for improving software quality and delivery speed.

CI/CDDevOpsarchitecture
0 likes · 20 min read
Five Core Dimensions of Maintainability Testing for Microservice Systems
Woodpecker Software Testing
Woodpecker Software Testing
Jan 5, 2026 · Operations

Three Core Dimensions of Performance Testing: Time Behavior, Resource Utilization, and Capacity

This article breaks down performance testing into three essential dimensions—time behavior, resource utilization, and capacity—explains their key metrics, demonstrates a detailed e‑commerce flash‑sale case study, and shows how systematic testing and optimization can dramatically improve response times, throughput, and scalability.

JMeterLoad TestingPrometheus
0 likes · 12 min read
Three Core Dimensions of Performance Testing: Time Behavior, Resource Utilization, and Capacity
DevOps Coach
DevOps Coach
Dec 26, 2025 · Operations

10 Actionable Agile Metrics to Replace Velocity and Deliver Real Value

This article presents ten practical, measurable Agile metrics—each with a problem statement, improvement action, real‑world example, concise code snippet, and baseline—showing how teams can shift from velocity to telemetry that reveals flow, quality, and predictability.

Agilemetricstelemetry
0 likes · 20 min read
10 Actionable Agile Metrics to Replace Velocity and Deliver Real Value
DevOps Coach
DevOps Coach
Dec 22, 2025 · R&D Management

Why We Abandoned Scrum: Inside Our Developer‑Led Delivery Transformation

After discovering that traditional Agile rituals stifled high‑output engineering teams, we rebuilt our workflow around autonomous, domain‑owned squads using GitHub PRs, feature flags, and real‑time metrics, resulting in dramatically faster deployments, fewer incidents, and higher developer satisfaction.

Agile TransformationDeveloper-Led DeliveryFeature Flags
0 likes · 8 min read
Why We Abandoned Scrum: Inside Our Developer‑Led Delivery Transformation
Alibaba Cloud Observability
Alibaba Cloud Observability
Dec 15, 2025 · Cloud Native

How UModel PaaS API Simplifies Observability Queries with Unified Entity Search

This article explains how the UModel PaaS API abstracts complex observability concepts—such as EntitySet, DataSet, StorageLink, and Filter—into a unified, object‑oriented query interface, offering Table, Object, and metadata modes, code examples, UI and SDK usage, and AI‑agent integration for efficient, low‑maintenance monitoring.

AI AgentAPIObservability
0 likes · 16 min read
How UModel PaaS API Simplifies Observability Queries with Unified Entity Search
PMTalk Product Manager Community
PMTalk Product Manager Community
Dec 9, 2025 · Product Management

Real‑World AI Data Analysis Case for Product Managers: Iteration & Optimization

The article shows how product managers can avoid the disappointment of a feature that looks perfect but gets no users by building a complete data‑driven loop that combines user‑behavior and business metrics, walks through a real e‑commerce recommendation case, outlines data‑collection pitfalls, metric‑design methods, hypothesis‑driven analysis, testing procedures and concrete steps to turn insights into iterative product improvements.

AIData Analysiscase study
0 likes · 33 min read
Real‑World AI Data Analysis Case for Product Managers: Iteration & Optimization
DevOps Coach
DevOps Coach
Dec 8, 2025 · Operations

How to Quantify SRE ROI: Turning Reliability Metrics into Business Value

This article explains how SRE leaders can bridge the gap between technical reliability metrics and business outcomes by defining core SRE concepts, applying a step‑by‑step ROI formula, illustrating code‑level impact, avoiding common pitfalls, and looking ahead to AI‑driven reliability forecasting.

BusinessValueOperationsROI
0 likes · 10 min read
How to Quantify SRE ROI: Turning Reliability Metrics into Business Value
Ray's Galactic Tech
Ray's Galactic Tech
Nov 26, 2025 · Cloud Native

Mastering Kubernetes Performance Bottlenecks: The Ultimate Troubleshooting Guide

This comprehensive guide walks you through the seven key performance metrics, resource, application, and system component indicators, and provides step‑by‑step methods, advanced tips, and tool recommendations for diagnosing and resolving Kubernetes performance bottlenecks from cluster‑wide to pod‑level details.

KubernetesPerformancecloud-native
0 likes · 11 min read
Mastering Kubernetes Performance Bottlenecks: The Ultimate Troubleshooting Guide
IT Architects Alliance
IT Architects Alliance
Nov 25, 2025 · Operations

Making Architecture Decisions Observable with DevOps Monitoring

The article explains how to integrate architecture decision tracking into DevOps monitoring, detailing tagging, multi‑layer metric design, time‑window analysis, automated alerts, reporting, and continuous optimization to turn architectural choices into measurable, data‑driven outcomes.

DevOpsObservabilitycloud-native
0 likes · 9 min read
Making Architecture Decisions Observable with DevOps Monitoring
Architecture Digest
Architecture Digest
Nov 24, 2025 · Operations

Boost Java Service Performance with MyPerf4J: A High‑Speed, Low‑Impact Monitoring Tool

MyPerf4J is an open‑source, high‑performance Java monitoring and statistics tool that uses a JavaAgent for zero‑intrusion, records up to ten million method calls per second with nanosecond precision, and provides real‑time metrics such as QPS, latency percentiles, memory and GC stats, making it ideal for both development and production environments.

JavaAgentPerformance Monitoringjava
0 likes · 6 min read
Boost Java Service Performance with MyPerf4J: A High‑Speed, Low‑Impact Monitoring Tool
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Nov 20, 2025 · Artificial Intelligence

How to Build a Quantifiable Data Quality Framework for Dynamic Incremental RAG

This article explains why static RAG metrics don’t apply to dynamic pipelines, introduces five essential dimensions—Parseability, Deduplication, Relevance, Chunk Quality, and Freshness—and shows how to combine them into a weighted score that enables monitoring, alerts, and continuous improvement of dynamic RAG systems.

Data QualityDynamic RAGRetrieval-Augmented Generation
0 likes · 10 min read
How to Build a Quantifiable Data Quality Framework for Dynamic Incremental RAG
High Availability Architecture
High Availability Architecture
Nov 14, 2025 · Artificial Intelligence

Quantifying AI Programming Efficiency: A Traceable and Measurable System

This article outlines the challenges of tracking AI‑generated code and measuring AI contribution, reviews earlier ad‑hoc methods, and presents a comprehensive solution featuring a VSCode plugin for unified AI dialogue management and a cloud service that quantifies AI impact across projects, teams, and individual developers.

AIAnalyticsVSCode
0 likes · 9 min read
Quantifying AI Programming Efficiency: A Traceable and Measurable System
DevOps Coach
DevOps Coach
Nov 10, 2025 · Operations

How to Use SRE Metrics for Data‑Driven Reliability and Faster Releases

This guide explains the SRE framework—SLA, SLO, SLI hierarchy, golden signals, error budgets, and DORA metrics—showing how to instrument a Python app with OpenTelemetry, query Prometheus, avoid common pitfalls, and adopt a cultural and technical process that balances feature velocity with system stability.

DORAError BudgetGolden Signals
0 likes · 18 min read
How to Use SRE Metrics for Data‑Driven Reliability and Faster Releases
Architect
Architect
Nov 4, 2025 · Operations

How to Accurately Track API Calls per Minute: 5 Proven Monitoring Strategies

This article explores why precise per‑minute API call statistics are essential for performance bottleneck detection, capacity planning, security alerts, billing, and troubleshooting, and presents five practical implementations—including fixed‑window counters, sliding windows, AOP‑based interception, Redis time‑series storage, and Micrometer‑Prometheus integration—along with their trade‑offs and capacity‑planning guidelines.

Performance OptimizationPrometheusRedis
0 likes · 25 min read
How to Accurately Track API Calls per Minute: 5 Proven Monitoring Strategies
JakartaEE China Community
JakartaEE China Community
Nov 4, 2025 · Operations

How Logs, Traces, and Metrics Differ—and Why It Matters

Logs, tracing, and metrics each serve distinct monitoring goals—logs capture discrete events for debugging and audit, traces map request flows to pinpoint performance bottlenecks, and metrics provide time‑series health data; understanding their differences and integrating tools like ELK, OpenTelemetry, Prometheus, and Grafana enables robust observability.

ELKGrafanaLogs
0 likes · 7 min read
How Logs, Traces, and Metrics Differ—and Why It Matters
Programmer XiaoFu
Programmer XiaoFu
Oct 28, 2025 · Backend Development

6 Ways to Measure API Response Time in Java

This article examines six practical techniques for measuring the latency of online interfaces in Java, from simple System.currentTimeMillis() calls to advanced AOP, interceptors, filters, and production‑grade monitoring tools like Micrometer and APM, comparing their precision, intrusiveness, and suitable scenarios.

AOPPerformanceSpring
0 likes · 23 min read
6 Ways to Measure API Response Time in Java
Alibaba Cloud Developer
Alibaba Cloud Developer
Oct 27, 2025 · Artificial Intelligence

How to Build a Quantifiable AI Coding Efficiency Metric System

This article explains how, amid the rapid rise of AI‑assisted programming, a scientific and actionable R&D efficiency metric framework was designed, detailing core indicators such as AI code adoption rate, data collection methods, platform architecture, and practical insights from a large‑scale implementation.

AICodingMCP
0 likes · 18 min read
How to Build a Quantifiable AI Coding Efficiency Metric System
Raymond Ops
Raymond Ops
Oct 12, 2025 · Operations

Master PromQL: From Basics to Advanced Query Techniques

This comprehensive guide walks you through PromQL fundamentals, covering data types, gauge and counter metrics, time‑series concepts, query selectors, offsets, arithmetic and logical operators, vector matching, aggregation functions, and key Prometheus functions such as increase, rate, and histogram_quantile, with practical examples and visual illustrations.

PromQLPrometheusalerting
0 likes · 29 min read
Master PromQL: From Basics to Advanced Query Techniques