Mission Engineering: Who Decides If the Coffee Is Good? Defining Success Metrics Beyond Timing
This article uses a morning coffee scenario to illustrate mission engineering's three-tier metric framework (MOS, MOE, MOP), showing how to define verifiable success criteria, separate stakeholder acceptance from process completion, and create traceable metric definitions that link overall outcomes to specific performance measurements.
Clarify "Satisfaction" Before Deciding What to Record
If you only say "make a good-tasting coffee" beforehand, afterward each person has a different standard. You think it's aromatic; your family thinks it's bitter. You think delivering by 7:28 is enough; they want time to sit and drink.
Ask two questions first: "You said not too bitter — can you accept a little bitterness, or was last time's taste unacceptable?" "Is 7:28 to 7:35 enough time to drink?" No need for professional flavor descriptors; write down what matters in language both understand. Don't immediately translate "bitter" into a recipe parameter; how to adjust comes later, after reviewing records.
Assume you agree on two outcomes: taste acceptable, and enough time to drink without delaying departure. You add a third item — recording the process — so you can repeat or adjust next time. The first two judge whether the family's needs are met; record completeness is checked separately.
This cup being satisfying and being able to reproduce it next time are separate concerns. Complete records only mean you have conditions to try again; they don't prove the method is stable.
Roles become clear: family judges taste, you record operations and timing, both verify whether enough drinking time was left.
Look at the Whole Picture, and at Which Step Actually Mattered
Suppose you notice water boiled faster than before. It's tempting to conclude: "This time we'll definitely serve earlier."
But while water boiled you might have been grinding beans. Water finished early, but you weren't done, so serving time may not shift earlier. Conversely, if you were truly waiting for water, faster boiling helps.
The same improvement must be viewed within the entire arrangement. The Mission Engineering Guide splits evaluation into three interrelated levels: MOS, MOE, and MOP. In this series, understand them as:
Overall result (MOS – Measures of Success): Did the family accept the taste? Was there time to drink without delaying departure?
Effect of completing the task (MOE – Measures of Effectiveness): Did preparation bring materials and tools to a ready state? How long did brewing wait due to missing prep?
Specific equipment or human performance (MOP – Measures of Performance): Under same water volume, starting temperature, and target conditions, how long does this kettle take to heat?
This borrows coffee to explain the guide's layers. For example, beans weighed but filter missing — you can't just tick "preparation complete." Part 04 will break down tasks and list exactly what must be ready before brewing.
Record completeness is checked separately for now. Later, if you want to teach the method to someone else, you'll check not only whether items are missing, but whether that person can follow the records and succeed. What you evaluate depends on the job at hand.
The three layers must be traceably connected. Example: kettle heating time may affect when materials/tools are ready, which affects when coffee is served. To verify that link, record when water finished, when other prep finished, when brewing started. The bottleneck shows up in the records.
Don't layer by metric name alone. Both called "duration," but total time from prep to serving versus kettle heating time under specified conditions evaluate different objects.
Give "A Bit Faster" a Verifiable Definition
Start with the most easily confused timing metric, not a full table.
Reusing Part 02's timing scope, define "elapsed time from morning prep start to coffee served." First prep step could be checking notes, inspecting materials, or grabbing tools — timing doesn't wait for water to boil. Serving moment must be recorded because family cares how much drinking time remains.
For the next trial, write this metric card first. The code CF comes from Coffee, T from Time, 01 is sequence; this is the series' lookup convention, not a Mission Engineering standard encoding.
Metric ID & Version: CF-T01, v1
What to measure: Elapsed time from morning first prep step to coffee served
Purpose: Check feasibility of morning schedule, provide duration basis for "on time"
Applicable scenario: CF-WD v1, weekday morning single pour-over
Start/end points: Morning first prep start to coffee served; prior-night prep recorded separately
Recording method: You record two timestamps, subtract start from end
Unit: seconds; display can convert to minutes:seconds
Current target: Start at 7:20, serve no later than 7:28, trial limit 480 seconds
Actual result: Fill after next trial
Evidence location: Link to this brew ID, keep raw timestamps
Missing data: If either timestamp missing, cannot calculate accurately; recalled estimates noted separately480 seconds comes from this example's schedule, not a recommended brew duration.
Duration answers "how long it took"; serving timestamp answers "did it meet the agreed time."
Same seven minutes, different start times, different outcomes:
On plan: Prep start 7:20, duration 420 sec, served at 7:27 — under 480 sec, also before 7:28.
Started 3 min late: Prep start 7:23, duration 420 sec, served at 7:30 — still under 480 sec, but missed 7:28 target.
So both metrics must be kept. The current 480 seconds is derived from 7:20 to 7:28; don't add a separate "over eight minutes = fail" rule. If later you agree on both "max eight minutes" and "latest 7:28 serve," check them separately. The table only verifies time; taste and drinking time still need family confirmation.
This metric includes any interruptions (e.g., a two-minute phone call) in the elapsed time; root-cause analysis separates them later. Prior-night prep is recorded separately for method comparison. Therefore this number shows how much morning time was consumed, not the total time you spent on this coffee.
Metrics first make the ruler clear: what to measure, from where to where, how to measure. Target values belong with the scenario and mutual agreement. In Part 05 these agreements become requirements stating who must achieve what under what conditions, referencing the same metric definitions.
Example: "serving timestamp" tells you what to record; "under agreed weekday conditions, you must serve coffee by 7:28" states who does what. Requirements can also specify which actions to perform and which records to keep — not just drawing a pass/fail line on numbers.
Record Both Outcomes and Diagnostic Clues Together
Put the agreed items into a register, each with an ID for later lookup; actual records note which version of the definition was used.
In the IDs, S, E, P come from Success, Effectiveness, Performance — the three layers above. T = Time, R = Record (for lookup). T and R are not new evaluation layers and don't automatically belong to MOE. A metric's layer depends on what it measures and what judgment it supports.
Taste accepted? CF-S01 — Family tastes, keeps original words and timestamp. Initial agreement: Acceptable / Not acceptable / Not yet confirmed; adjustment wishes recorded separately to avoid treating "want a tweak" as rejection. Belongs to MOS; cannot be replaced by temperature or recipe completion.
On time? CF-S02 — You record serve time, family confirms if drinking time sufficient; cross-check if departure delayed. Initial agreement: Hope to serve by 7:28; whether leftover time is enough needs family confirmation. Belongs to MOS; CF-T01 helps explain duration, early serve doesn't guarantee enough drinking time.
Morning prep-to-serve duration CF-T01 — Record two timestamps per definition above. Initial agreement: Trial target per example; actual result pending. Helps check if time target feasible; prior-night prep recorded separately.
Post-prep ready state CF-E01 — You check off materials/tools list, record when all ready, and any wait caused by missing items. Initial agreement: List beans, filter, water, required tools first; exact amounts and conditions per chosen recipe. Belongs to MOE; checks if prep supports brewing start, identifies what was missing and how long waited.
This trial's record gaps CF-R01 — You compare against pre-trial checklist, list missing items. Initial agreement: First record bean batch, recipe version, actual amounts, key operation timestamps, ad-hoc changes. Auxiliary check for retrospective; completeness doesn't prove stable reproducibility.
Kettle heating performance CF-P01 — When suspecting wait-for-water, you record heating time under same volume, start temp, target conditions. Initial agreement: No pass/fail threshold set, not required every trial. Belongs to MOP; used to verify if boiling affects prep progress, not pre-assumed as cause.
The table has no actual results yet. Just recording duration and spotting unprepared items lets you start finding problems — no need to measure every kitchen tool to fill all three layers first.
Assume "taste acceptable" and "no departure delay" are both required; check each separately. Faster boiling, more records — neither cancels out family rejecting the taste. Summing items into a single score may hide which specific thing failed.
If multiple methods meet basic agreements, then compare which saves more time or is easier to record. Agree on what matters before comparing, to avoid cherry-picking favorable metrics after the fact.
"Tastes Good" Can Be Recorded Seriously Without Pretending Full Objectivity
Family says "a bit bitter, but I can accept it" — keep the exact words, then separately log "acceptable" and "next time prefer lighter." If unclear, ask "Is this cup drinkable?" then "What would you change next time?" A single "pass" loses adjustment suggestions; "unsatisfactory" distorts their meaning.
At minimum note who drank, which coffee, when, what they said. If first sip differs from later impression, record both separately — don't treat them as conflicting answers.
This is this person's judgment under these conditions. Family liking it doesn't mean others will; this time liking it doesn't mean switching beans lets you reuse the conclusion.
If family rushes out without tasting, log "not yet evaluated." If they drank but left no feedback, log "feedback not recorded." Neither counts as satisfied or unsatisfied; when following up, they know which cup is being asked about.
Finish One Trial, Leave an Unambiguous Conclusion
Back to the 7:27 serve, "bitter" cup. Record like this:
Served at 7:27, family says drinking time okay but hasn't noted if departure delayed. Taste "a bit bitter, next time hope lighter" — whether this cup is acceptable needs follow-up. Notes only have total duration, missing specific method; next adjustment lacks basis.
After follow-up: if family explicitly accepts and no departure delay, log "this cup met agreements, taste still has adjustment suggestion"; record gaps listed separately. If family explicitly rejects, log "taste missed agreement"; on-time serve result still stands.
If requirements change later, write why changed and from which trial new version applies; old results stay interpreted under old agreements. Metric definitions versioned — never let the same ID's "seven minutes" silently change its start/end scope.
To apply to your own work, pick a target from previous part's scenario card, write a metric like CF-T01. Then check: who can confirm result? Can existing records compute the value or support the agreed evaluation? Will missing data be defaulted to success? Can a different method still be compared with the same ruler?
Getting one metric writable so others can record it is more useful than listing a dozen undefined "efficiency, quality, satisfaction" metrics.
Summary
Family says "a bit bitter" — first clarify: unacceptable or just wants adjustment next time? Whether time was met, whether method was fully recorded — capture each answer separately so next time you know what to fix.
Once you know what to watch, you can dig deeper: where did time go? Which prep step fell short? Which tasks truly can't run in parallel? Next part starts task decomposition, putting these questions back into concrete actions.
Reference
U.S. Department of Defense, Mission Engineering Guide, Version 2.0, 2023-10-01, Section 4.2, pp. 14-17. Used to explain MOS, MOE, MOP meanings, layer relationships, and metric selection principles. Chinese explanations, coffee examples, metric IDs, evaluation options, and target values in this article are teaching designs, not the guide's prescribed coffee evaluation standards. URL: https://ac.cto.mil/wp-content/uploads/2023/11/MEG_2_Oct2023.pdf
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
