Observability Isn't a Log Platform: Connecting Business Goals to Verifiable Runtime Facts
This article explains that observability is not centralized logging but a practice of defining business outcomes (SLIs/SLOs), correlating metrics, logs, and traces via stable business identifiers, designing actionable alerts, structuring dashboards along business flows, separating observability responsibilities, halting automation when evidence is unreliable, and using runtime data to continuously correct architectural assumptions.
The previous article discussed how capacity, traffic, storage, and cost jointly constrain a business chain.
For these conclusions to hold continuously in production, we must know how much work the system is processing, how long users wait, which dependency is slowing down, and whether the business state is ultimately correct.
Many teams have built log platforms, monitoring dashboards, and alerting systems, yet during incidents they still search logs machine by machine. Tools can collect data but do not automatically tell us what to observe, how signals correlate, or who should act after seeing a problem.
Observability is not centralizing logs; it is making the business goals, state boundaries, dependency relationships, and capacity assumptions in the architecture become continuously observable, correlatable, and verifiable runtime facts.
Machine Health Does Not Equal Business Health
Continuing with the refund example from earlier.
A user submits a refund; the interface quickly returns "accepted", but twenty minutes later it still hasn't completed. At that time, application CPU is low, interface success rate is near normal, yet on-call staff cannot immediately answer:
Is this an occasional delay for one refund, or has a backlog already formed system-wide?
Is the task waiting for scheduling, or is the payment channel responding slowly?
Has the channel already succeeded, but the local state or business event hasn't updated?
Is retry recovering the problem or further amplifying channel pressure?
If monitoring only covers application liveness, CPU, memory, and HTTP status codes, the system may appear completely normal. But what users truly care about is not whether the interface returned 200, but whether the refund can complete correctly within the promised time.
Therefore, observability must first answer two types of questions:
1. What are users and the business experiencing? (symptoms)
2. Why is the system producing this result internally? (causes)Looking only at resource metrics may reveal a cause but not whether it affects users; looking only at business success rates may tell you the system is broken but not why.
Design the Questions You Need to Answer First
Observability is not about recording as much as possible; it is about deciding, before failures occur, which questions you need to answer from the runtime signals the system emits.
For the refund chain, I would start from the business outcome and drill down into internal dependencies layer by layer:
Question: Is the refund application accepted normally? Key Runtime Fact: Acceptance volume, rejection volume, interface error rate, acceptance latency
Question: Is the refund completed within the target time? Key Runtime Fact: On-time completion rate, end-to-end latency, count of timed-out unfinished tasks
Question: In which state does the task stay too long? Key Runtime Fact: Time spent in each state, age of earliest unfinished task, backlog count
Question: Which internal or external dependency has slowed down? Key Runtime Fact: Call volume, latency distribution, timeouts, rate limiting, unknown results
Question: Is retry recovering or amplifying the failure? Key Runtime Fact: Original request volume, actual call volume, retry count, recovery outcome
Question: Are channel and local states consistent? Key Runtime Fact: Channel query result, local state, reconciliation discrepancies, manual interventions
This list defines what needs to be seen first, then decides whether to express it via logs, metrics, or traces. Reversing the order leads teams to collect large amounts of framework-default data while lacking evidence for real business problems.
What Logs, Metrics, and Traces Each Answer
Logs, metrics, and distributed tracing are not three interchangeable tools. They record system behavior at different granularities and must be correlated through shared context.
Metrics — Better at answering: Whether a problem exists, its magnitude, and trend. Refund chain example: On-time completion rate, backlog, channel latency distribution. Main limitation: Aggregation loses individual requests and business details.
Logs — Better at answering: What event occurred at a moment, why state changed. Refund chain example: Task accepted, channel returned unknown, state entered reconciliation. Main limitation: Without structure and correlation IDs, only full-text search works.
Trace — Better at answering: Which nodes a technical call traversed, where time was spent. Refund chain example: Gateway, refund service, database, and channel call latency. Main limitation: May not cover the full lifecycle of long-running async business.
Metrics are suitable for first detecting "completion rate is dropping"; traces help locate "time is mainly spent in channel calls"; structured logs then provide concrete evidence for "why a specific task entered reconciliation".
These signals do not automatically form causal conclusions. Channel latency and refund timeout rising together only shows correlation; combining call chains, state changes, release records, and necessary reproduction or chaos experiments is required to judge whether the current explanation holds.
Technical Calls and Business Lifecycles Need Different Identifiers
A Trace ID usually suits correlating a single synchronous request or a segment of async execution, but should not be used as a business primary key spanning hours or even days.
A refund may go through interface acceptance, then async execution, channel query, reconciliation, and manual handling. These stages may correspond to multiple traces and must be connected by a stable business identifier.
Refund ID or business operation ID — Meaning: The same business fact. Suitable lifecycle: From acceptance to final state and subsequent reconciliation.
Request ID — Meaning: One ingress or retry attempt. Suitable lifecycle: One request.
Trace ID — Meaning: A cross-component technical execution path. Suitable lifecycle: One synchronous call or async execution segment.
Task ID or message ID — Meaning: A specific async work unit or delivery carrier. Suitable lifecycle: Scheduling, delivery, consumption, and retry.
Idempotency key — Meaning: Whether multiple attempts represent the same business intent. Suitable lifecycle: Deduplication and result reuse boundaries.
Channel request ID — Meaning: Operation recognized by the external system. Suitable lifecycle: Submission, query, receipt, and reconciliation.
Logs and traces should also record service, environment, instance, version, and release batch. Only then can you determine whether a problem started from a certain release, or appears only in a specific environment or instance.
However, do not use refund IDs, user identifiers, or task IDs directly as regular metric labels. Such high-cardinality fields generate massive time series and are better placed in controlled logs, traces, or business query systems. Sensitive data among them must be desensitized, access-controlled, and given retention periods per security and compliance requirements.
SLIs and SLOs Should Originate from User Outcomes
CPU utilization, memory usage, and queue backlog are important, but they are not the service outcomes users receive.
SLI (Service Level Indicator) represents the system's actual behavior. SLO (Service Level Objective) specifies the level this indicator should achieve over a time window. OpenTelemetry's observability guide also emphasizes that good SLIs should measure service from the user perspective.
SLI: Over the past 30 days, 99.93% of refunds completed within 10 minutes
SLO: Every 30 days, at least 99.9% of refunds should complete within 10 minutesFor example, the refund service can define two categories of indicators for "application acceptance" and "refund completion":
SLI: Acceptance result clarity rate. Definition of "Good Event": Eligible applications explicitly accepted or rejected within target time. Must Be Explicit at Design Time: Total event scope, whether legitimate business rejections count as failures.
SLI: On-time completion rate. Definition of "Good Event": Accepted refunds reaching correct final state within promised time. Must Be Explicit at Design Time: Start/end time, correct final state, business type, allowed exclusions.
SLI: State correctness rate. Definition of "Good Event": Local state, channel result, and subsequent events consistent within target window. Must Be Explicit at Design Time: Authoritative source, reconciliation window, discrepancy handling method.
Target numbers must be based on real business commitments and baseline data, not blindly filled with 99.99% to make the system appear more reliable. In low-traffic systems, a single failure can cause large short-window fluctuations; here you must combine event count, impact, and longer windows for judgment, not mechanically apply the same ratio alert.
If an SLO allows a few sub-standard events within the statistical window, that acceptable space is the error budget. It is not a quota for intentionally creating errors, but a gauge for current reliability risk and whether the team can continue releasing and experimenting.
Alerts Should Point to Impact and Action
The problem with many monitoring systems is not lack of alerts, but too many. CPU, memory, threads, connections, queues, dependencies, and every instance send separate messages; one failure may generate dozens of notifications, yet none clarifies what impact users suffered.
Alert design should prioritize user symptoms and SLO risk, then use dependency and resource signals to assist localization. Google SRE's four golden signals — latency, traffic, errors, saturation — serve as basic checks but must still map to concrete business outcomes.
Urgent notification — Trigger condition: Major user impact, ongoing, on-call must act immediately. Handling method: Notify current owner, enter incident response.
Ticket or backlog — Trigger condition: Risk accumulating, can be handled in next working period. Handling method: Create time-bound, owner-assigned remediation task.
Record and analyze — Trigger condition: Metric change not yet constituting actionable risk. Handling method: Retain in dashboards and events for trend analysis.
An actionable alert must include: which business goal is affected, impact scope and duration, current owner, key dashboards and query entry points, and suggested first step. If the recipient cannot take any action, it should not repeatedly disturb on-call as an urgent alert.
Dashboards Should Unfold Along the Business Chain
Dashboards are not about plastering all metrics on one big screen; they help on-call staff navigate from symptom to localization.
The refund chain can be organized in four layers:
1. Business outcomes: acceptance, on-time completion, state correctness
↓
2. Service level: traffic, latency, errors, saturation
↓
3. Dependencies & state: task backlog, database, messaging, payment channel
↓
4. Resources & observability infra: instances, CPU, memory, network, collection pipelineThe first layer answers "is the user being impacted?"; subsequent layers help answer "why?". Dashboards should allow drill-down via time, version, service, and correlation identifiers — not force on-call to switch systems and re-guess query conditions.
The observability system itself is a production chain that needs observing. When collection breaks, clocks drift, metrics lag, logs are lost, or queries become unavailable, "no alerts" cannot be interpreted as "no problems".
Observability Responsibility Cannot Rest Solely on the Platform Team
The platform team can build unified log, metric, trace, and alerting platforms, but they don't know when a refund is truly complete, nor can they independently decide which latencies require immediate human intervention.
Observability likewise requires three layers of responsibility:
Architecture design — Primary responsibility: Define key business outcomes, state boundaries, identifier propagation, SLI/SLO, failure models, and observability requirements.
Platform capability — Primary responsibility: Provide default collection, correlation context, storage/query, desensitization, retention, dashboards, and alert channels.
Service team — Primary responsibility: Add business instrumentation, maintain dashboards, alerts, runbooks, on-call duty, and correct implementation based on runtime data.
Business owners must also confirm completion timeliness, allowed delays, risk levels, and manual intervention priorities. The observability platform can only compute clearly defined indicators; it cannot replace business decisions.
Automated Mechanisms Must Stop When Evidence Is Unreliable
Observability signals often drive auto-scaling, rate limiting, circuit breaking, rollback, and failover. Once signals are missing or scopes are wrong, automation may continue executing based on illusion.
Automated mechanisms should not endlessly retry or amplify actions when:
Observation data is interrupted, delayed, or source cannot be confirmed for a long period;
Business state, external results, and technical metrics conflict with each other;
Auto-recovery executes but user impact does not converge or recurs;
Funds state modification, manual reconciliation results, or expanded blast radius are needed;
Existing runbooks do not cover the current failure combination.
At that point, the system should preserve current evidence and executed actions, hand off state to authorized personnel, rather than overwrite the scene with another round of automation.
Observation Data Should Reverse-Correct Architecture
Observability is not only for troubleshooting during failures; it should also verify previous architectural assumptions.
In the refund system, the team might have assumed: queues absorb short bursts after promotions, channel timeouts recover via retries, adding instances shortens task wait time. Runtime data may prove: actual backlog recovery exceeds business targets, retry success rate is low yet greatly increases channel calls, new instances are limited by database connections.
At this point, one should not merely adjust dashboard thresholds, but return to capacity budgets, retry boundaries, state models, and dependency contracts to correct the architectural design.
Architectural assumption
→ Observability signals
→ Runtime evidence
→ Gap analysis
→ Correct architecture, thresholds, or operating practicesRuntime data only represents phenomena observed in specific versions, traffic, environments, and time windows. It can falsify some assumptions, but cannot be extrapolated beyond its boundaries as permanently valid causal conclusions.
Observability Conclusions Also Need Evidence
Having collection agents installed and data on dashboards does not prove the system is observable.
Observability Architecture Judgment: Key business outcomes are measurable. Verification Target: Acceptance, completion, timeout, discrepancy match business facts. Evidence Source: Business reconciliation, metric validation, scenario testing.
Observability Architecture Judgment: Anomalous requests are traceable end-to-end. Verification Target: From business ID can query request, trace, task, and external call. Evidence Source: Fault injection, trace query, log correlation.
Observability Architecture Judgment: Alerts detect real user impact. Verification Target: Correctly notify when failure reaches threshold; no noisy alerts for non-impact fluctuations. Evidence Source: Alert drills, incident records, on-call feedback.
Observability Architecture Judgment: Observability pipeline itself is trustworthy. Verification Target: Collection, transport, storage, query interruptions are detectable. Evidence Source: Collection interruption drills, integrity checks, clock verification.
Observability Architecture Judgment: Runtime responsibility can truly be taken over. Verification Target: Non-original author can locate from alert and handle per runbook. Evidence Source: On-call drills, handover tests, handling records.
Every major change in state, dependency, retry, capacity, or release method should trigger a check whether existing signals and alerts remain valid. If a new final state is added but the on-time completion rate definition isn't updated, the resulting number may be precise but the conclusion wrong.
Agent Systems Likewise Cannot Just Save Conversations
Agent Runtime is also a stateful, dependency-bearing, externally-behaving running system. Merely saving user and model natural language conversations cannot explain which nodes the task passed through, which model and tool versions were used, why it entered waiting or failed, nor prove whether external actions actually completed.
Task ID, node execution, model calls, tool calls, policy decisions, external receipts, and human decisions need correlation via stable identifiers. A tool returning success or receiving an external receipt does not automatically equal business completion; high-risk actions still require verification against authoritative source state per contract. This is not to save more reasoning text, but to judge whether the Agent correctly completed the business task within controlled boundaries.
Small Systems Also Need a Minimum Observability Loop
Internal systems with short call chains, low traffic, and failures manageable manually during work hours do not need full distributed tracing, complex SLO platforms, or 24/7 on-call from day one.
The minimum loop should still include: the one or two outcome metrics users care most about, structured error logs, stable request identifiers, key dependency health status, one alert that reaches the owner, and a brief troubleshooting guide.
When cross-service localization repeatedly consumes time, issues cannot be reproduced, runtime responsibility begins to hand over, or the business has explicit SLO requirements, then continue adding traces, error budgets, alert tiers, and more complete on-call mechanisms.
Summary
Observability is not owning a log, metric, and trace platform; it is the team's ability, when new problems arise, to understand from trustworthy signals what users are experiencing, trace through which states and dependencies the problem passed, and find the next actionable step.
1. First define the business outcomes that need to be visible,
2. Then design signals, correlation identifiers, and the runtime loop.Machine liveness does not mean business health; data collected does not mean conclusions are trustworthy; alerts fired does not mean responsibility is closed. Only when runtime facts can reverse-verify and correct architecture does the observability loop truly close.
The next article will continue discussing fault tolerance and recovery: when these signals prove a failure has occurred, how should the system limit impact and restore to a correct state.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
