R&D Management 30 min read

Engineering Leadership: Shifting from Problem Solving to Problem Discovery

This article details how engineering managers can evolve from reactive problem-solving to proactive problem discovery by questioning task origins, validating early signals, designing limited experiments, balancing delegation with intervention, and verifying post-launch outcomes, using concrete examples like latency optimization trade-offs and release risk management.

Architecture and Beyond
Architecture and Beyond
Architecture and Beyond
Engineering Leadership: Shifting from Problem Solving to Problem Discovery

Introduction

The author describes a habit of reviewing their calendar to see where time goes. Much of it is spent solving problems — architecture, requirements, people — which feels busy and fulfilling. The article argues that problems, like testing and requirements, need to "shift left": if the same issues recur next month, we must ask what we actually solved today.

Where Do Problems Come From?

Engineers receive tasks with clear constraints (throughput, latency, compatibility, deadline). Clear constraints make technical discussion converge. But managers must also check those constraints themselves. For example, a business request to halve an API's response time could be met with indexing, caching, concurrency tuning, or resource reallocation — all valid engineering work. The manager first asks: what is the user actually waiting for? If the API is only a tiny part of a multi-step user journey, local optimization may not improve experience. If the wait is mostly in manual review, compressing server time is misdirected. If the response is already fast but users resubmit, the issue is feedback and flow. The point is not to dismiss performance work, but to verify that the technical metric truly represents the desired outcome. The author insists stakeholders describe the problem — who struggles, under what conditions, what loss occurs, what evidence exists — not the solution. Managers must own the source of the task. Solution review only checks the answer to a given question; it cannot correct the question itself.

Bian Que's Dilemma

The classic story of Bian Que and his three brothers illustrates an evaluation problem. The eldest brother treats illness before it appears, the middle treats it when minor, and Bian Que treats it when severe and visible. The visible cure gets the most recognition. In technology, when a problem becomes severe, the loss and the fix are both visible, so value is easily understood. Preventive work lacks that observable outcome: a team claims an architectural change avoided a major outage, but the outage might not have happened anyway due to low traffic or untouched code paths. "Prevention is important" does not justify any preventive investment. The author asks for concrete evidence: which functions would fail if a dependency goes down, are alternative paths available, what recovery requires, and how the team has verified this. Only then discuss the cost of redundancy — build, switchover, drills, ongoing maintenance. Unlike the legendary physician, technical managers face incomplete information, changing business, and judgments that new evidence can overturn. Therefore, we must record what we observed, what we inferred, how we plan to verify, and the cost of being wrong. Early discovery is only worth investing in when judgment quality and handling cost are both acceptable.

Write Symptoms Clearly

Often, responsibility is assigned the moment a problem is raised: delay → poor execution; production bug → weak quality awareness; repeated reviews → insufficient technical skill. These labels are short and map to familiar actions (push progress, add reviews, train), but they lack verifiable content. Taking delay as an example, the author breaks down the process: when did requirements become dev-ready, how long did coding take, how long was integration wait, where did rework happen, did acceptance criteria change? If a two-week task spent most time waiting for a dependency to confirm an interface, demanding faster coding won't address the main loss. Adding more status meetings only consumes coordination time. However, long wait time does not automatically mean "cross-team collaboration is broken." The wait could stem from unclear interface definitions, the other team lacking resources, or our own premature start before conditions were met. Each requires a different response. Adding a coordination role might just add another layer of message passing. The author proposes a logic chain: observed facts → explanations based on facts → candidate actions. For instance, several tasks stall at the same dependency (fact); the other team is resource-constrained (explanation to verify); adjust both schedules (candidate action). Writing them together prevents later discussion from silently accepting the unproven explanation. The closer the problem description stays to observable behavior, the easier it is for the team to discuss concrete changes. Spending energy explaining why a person is "not responsible enough" rarely yields an acceptably verifiable plan.

Find Early Signals

Post-incident analysis has concentrated evidence. Early problem detection deals with much fuzzier signals. The author prioritizes areas that currently show no loss but rely on sustained extra labor to stay afloat: frequent alerts requiring developer investigation, every release needing a specific person's manual confirmation, a data fix run daily, a module only a fixed few dare modify, projects that deliver on time only by cutting tests or compressing acceptance. These phenomena don't prove imminent loss of control, but they warrant inspection because the current results hide unrecorded manual effort. If we only track whether delivery finished or the service is up, we cannot know how much compensatory labor the team expends. The author asks the team to record at low cost: repetitive manual operations, time waiting for external decisions, unplanned work, and tasks that depend on specific individuals. The goal is to find recurring patterns, not to slice everyone's day into dozens of time slices. Overly fine granularity turns recording into a burden and distorts data around evaluation metrics. Technical observation requires similar trade-offs. The author does not demand full detailed logging on all paths just to avoid missing something. Sampling rate, retention, field count, and query needs must map to specific diagnostic problems, with a budget for collection overhead and storage cost. If a new dataset cannot answer any decision question, the preference is not to collect it. Alerts must also specify follow-up actions: when a metric crosses a threshold, who intervenes, what do they judge, how long can they wait? Alerts that cannot answer these go into an observation bucket first, avoiding direct interruption of on-call staff. Early detection needs sufficient information but cannot be achieved by unlimited data volume. The approach works backward from a few high-loss scenarios: before the result worsens, did we have a chance to see the change? Can existing information distinguish different causes? How can we add the missing piece at minimum cost?

How Deep to Question

After finding a recurring problem, it's easy to fall into the opposite extreme: keep asking "why" until reaching a grand explanation. A failed release → insufficient testing → tight schedule → business pressure → "organization needs stronger quality culture." The discussion ends, but the next release still has no concrete answer. The author recommends stopping the analysis at the level that is currently actionable and verifiable, while recording higher-level constraints. Suppose a failure involved a config change. After restoring service, continue to clarify: was the config validated, could pre-release checks have caught the anomaly, can the change be rolled back independently, why didn't existing checks cover this path? If validation is missing, discuss how to add it. If validation exists but was skipped due to time pressure, examine the bypass conditions and approval authority. Further up, if every delivery commitment implicitly compresses verification time, the scheduling method itself enters the improvement scope. These layers can coexist; there is no need to force a single "root cause." The author cares more about which change breaks a known failure path, which shortens recovery time, and which constraints must be accepted for now. Adding a release gate is a common choice, but used cautiously: it continuously consumes reviewer time and makes releases wait on another person's judgment. If a class of errors can be caught by deterministic checks, automated validation is evaluated first; manual review only adds value when business context judgment is required. Automation is a common strategy but must consider cost. A low-frequency exception that takes little time to handle may not justify a complex process. Conversely, high-frequency, stable operations that rely on manual checking need a stated reason. The depth of analysis should serve the choice of intervention. When further questioning no longer changes the action, the author stops expanding the discussion and turns to verification.

Problems Also Need Scheduling

As problem discovery improves, the backlog likely grows first. Code debt, process gaps, people dependencies, business assumption risks — if every discovery must be solved immediately, the team starts many improvement efforts in parallel and the original delivery plan loses credibility. Managers must own the decision to "not handle this for now." To decide whether a problem enters the schedule, the author looks at: what loss has already occurred, how future loss might expand, how reliable the judgment is, what the handling cost is, and whether delaying loses adjustment space. Simply put: what happens if we don't do it, and what does it cost to do it? (That's two questions.) Example: two improvements both reduce maintenance work. One can be done anytime; the other involves an interface contract that multiple teams will soon depend on. Once that contract spreads, migration requires coordinating many more participants. The author might prioritize the latter even if it saves less time right now. But these factors are not combined into a precise score for ranking by decimal points. Probability, impact, and effort contain large estimates; a formula cannot eliminate uncertainty. For problems with insufficient information, the author schedules separate verification work: a time-boxed window to check actual call patterns, run a capacity experiment, or replay a batch of historical tasks. After verification, decide whether to launch a full rebuild. Verification tasks need exit criteria; they cannot become open-ended "ongoing research" that consumes headcount indefinitely. The author also distinguishes two types of decisions: allowing a risk to exist temporarily, and committing to fix it by a certain date. The former requires recording who accepted the risk, the accepted scope, and the conditions for re-evaluation. Not all deferred items go into an endless technical debt list to be revisited only when the next incident occurs. Limited resources can only handle limited problems. A manager's job includes prioritization, deferral, and rejection — not just producing an ever-growing risk list.

Do Limited Validation First

Early-discovered problems often lack enough evidence for large-scale reconstruction. The priority is to find small experiments that can change the judgment. Suppose the team believes slow delivery stems from a hard-to-use common component library. Rebuilding the whole library is a large investment with a long validation cycle. Instead, pick one category of repetitive tasks, adjust one high-frequency interface, and compare integration steps, rework count, and total duration. The experiment does not need to prove the new approach is better in all scenarios. It only needs to answer: does the currently identified obstacle actually affect delivery, and does removing it produce the expected change? In experiment design, pay special attention to comparability conditions: are the two tasks similar in complexity, are participants different, was extra coaching given, did the new approach receive resources the old one didn't? If these conditions shift, narrow the conclusion's scope. If test tasks are all well-defined, complete-information samples, the conclusion only covers that subset. When facing missing data, conflicting rules, or undecidable situations, how the system exits and who takes over must also be in the validation scope. Limited experiments have limits: they may miss low-frequency issues, benefit from extra team attention (Hawthorne effect), or fail to capture collaboration costs at scale. Therefore, after one experiment, the author continues to limit rollout scope and keeps rollback conditions. The author is willing to pay for these steps. Compared to a one-time full process replacement, phased validation provides a chance to correct judgment before investment expands.

Delegation and Intervention

Discovering problems can create an impulse: since I see it, handling it myself is fastest. In emergencies that may be appropriate. But if every complex problem eventually returns to the manager, the manager must examine their intervention style. The author separates three stages: risk identification, decision making, execution. They can be owned by different people. A manager who spots a risk in a release pipeline can demand additional evidence, assign an owner, and set a deadline. Unless the risk exceeds the team's current capacity or requires crossing existing authority boundaries, there is no need to take over the entire implementation. When delegating, the manager communicates the goal, constraints, available resources, and escalation criteria. Implementation details are left to the owner. There is a real cost: the team's solution may differ from the manager's habit, and local implementation may be slower than if the manager did it personally. If the manager reclaims the task every time such a difference appears, members never accumulate the experience of making complete judgments. But delegation cannot become absence of oversight. The author sets checkpoints based on risk: easily reversible local changes allow the owner to move fast; changes involving data destruction, cross-team commitments, or hard-to-rollback migrations require earlier review of the plan and premises. Checkpoint frequency varies with risk, avoiding uniform management intensity for all tasks. The manager also personally participates in some field work: reading an incident timeline, following a long-blocked task, or checking the actual usage of a change. If information only comes from layered reports, it's hard to judge which costs were omitted. Direct contact calibrates judgment; afterward, the designated owner still completes the work and acceptance. After discovering a problem, the manager still owes the team a decision: who owns it, how much to invest, when to check, and what else will not be done as a result.

Let Risks Surface

If the manager wants the team to report problems early, they must examine how they react to bad news. When a member raises a risk, they are immediately asked to prove it will definitely happen. When they admit schedule uncertainty, they are labeled as lacking accountability. When they uncover a historical defect, the first question is why they didn't find it earlier. Faced with these responses, a rational person waits for more evidence before speaking up. By the time evidence is sufficient, the room for adjustment may have shrunk. The author allows risk reports to retain uncertainty but requires specificity: what was observed, what might be impacted, what information is missing, what support is needed. "This project will definitely delay" and "The key dependency is unconfirmed; if it's not done this week we need to adjust the subsequent integration schedule" carry different information content. The latter lets the manager make a decision and allows later facts to correct the judgment. For good-faith risks that ultimately don't materialize, the author does not judge the reporter as wrong based solely on outcome. Instead, they review whether the information at the time supported the concern, whether the verification actions were reasonable, and whether early handling prevented escalation. Similarly, the author does not reward evidence-free persistent pessimism. Once a risk is raised, someone must gather information, narrow the scope, and retract the judgment if needed. Only flagging all possibilities without participating in verification shifts the filtering cost to the whole team. Evaluation of preventive work also needs evidence. The author checks: was the risk path validated, does the change cover it, were rollback and recovery tested, is ongoing maintenance ownership assigned? As for "how much loss was avoided," without solid basis, no inflated number is written. For repetitive manual labor, compare before/after handling counts and time. For recovery capability, record actual operations and duration in drills. For collaboration improvement, check whether wait times and rework changed. Evaluation should land on records that can be rechecked. This adds some recording cost. The author keeps the parts relevant to decisions and acceptance, not requiring a full report for every improvement.

Check What's Already Done

One class of problems is easily overlooked in technical management: the task is complete, but the investment didn't produce the expected result. The system launched on time, features passed acceptance, interfaces matched the design, the team made no obvious mistakes. Months later, usage is low, the original manual process continues, and the new system adds a maintenance burden. If the manager's check stops at launch, this problem never enters view. The author agrees on a post-investment review at project initiation. What to check depends on the problem the project aimed to solve: were manual steps reduced, can users now complete previously difficult tasks, did the business actually adopt the new flow, did the old system retire on schedule? "Already launched" cannot substitute for answers to these questions. If results are absent, the author re-examines the original assumptions. Perhaps the obstacle was misidentified, the solution only addressed a small part, or business conditions changed. Before adding more features, confirm whether further investment still has a basis. Stopping the project is also a choice that needs serious evaluation. It involves sunk costs, staffing adjustments, and commitment changes — often messier than approving the next phase. But maintaining a system without a usage rationale also continuously consumes resources. Discovering problems therefore includes checking whether the problem the team is solving still exists, whether the solution still fits, and whether the goal has changed. This changes how the author allocates time. They still participate in incident response, understand technical details, and own delivery results. Simultaneously, they reserve fixed time to examine repetitive labor, unverified judgments, and post-launch actual results. For exposed problems, demand closure; for signals, schedule limited validation; for accepted risks, leave re-check conditions. The author also regularly reviews their own raised problems: how many gained evidence support, how many were refuted, which linger undecided and keep consuming the team, which actions created new burdens. Otherwise, "discovering problems" easily becomes the manager continuously adding tasks to the team. Moving from solving problems to discovering problems means an expanded scope for the author: participating in definition, making trade-offs, and verifying judgments after results appear. Early visibility only gives more handling space. Whether that space is used well ultimately comes back to concrete actions and results.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Risk ManagementEngineering ManagementTeam LeadershipProblem DiscoveryDelegationTechnical Decision MakingPost-Implementation ReviewPreventive Maintenance
Architecture and Beyond
Written by

Architecture and Beyond

Focused on AIGC SaaS technical architecture and tech team management, sharing insights on architecture, development efficiency, team leadership, startup technology choices, large‑scale website design, and high‑performance, highly‑available, scalable solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.