R&D Management 20 min read

Borrowing Aerospace Discipline for AI Workflows: From ASD-STE100 to MISRA Rules

The article explores how safety-critical industry standards like ASD-STE100, DO-178C, and MISRA can be applied to AI-assisted development workflows, detailing three skills for requirements, review, and code safety that mirror aerospace and automotive engineering practices.

BirdNest Tech Talk
BirdNest Tech Talk
BirdNest Tech Talk
Borrowing Aerospace Discipline for AI Workflows: From ASD-STE100 to MISRA Rules

Karpathy's Tweet: The Output Format Ladder

On October 2, Andrej Karpathy posted a widely viewed tweet arguing that as LLMs grow stronger, more human time will be spent understanding their outputs. He proposed a progression of output formats: first, ask the LLM to write in ASD-STE100, a controlled language from aerospace maintenance documentation; second, request diagrams instead of text; third, generate interactive HTML web pages; and finally, produce custom explainer videos (e.g., in 3Blue1Brown style with ElevenLabs narration). Karpathy's concluding insight: LLMs will handle more legwork autonomously, shifting human work up the abstraction ladder to oversight and understanding — provided we ask for "big, custom, disposable software artifacts" like web pages and videos.

ASD-STE100: Controlled Language for Aerospace

ASD-STE100 (Simplified Technical English) originated in the 1970s when AECMA (European Association of Aerospace Industries) found that non-native English speakers maintaining aircraft worldwide struggled with standard industrial English — long sentences, abstract nouns, and synonym rotation created ambiguity that could become safety risks. The first AECMA Simplified English Guide appeared in 1986; after AECMA merged into ASD (AeroSpace and Defence Industries Association of Europe) in 2004, the standard was renamed ASD-STE100 in 2005 and is maintained freely.

The specification has two parts:

Part 1 – Writing Rules : nine chapters governing grammar, sentence patterns, paragraphs, punctuation, and procedural writing. Key quantitative constraints: procedural sentences ≤ 20 words, descriptive sentences ≤ 25 words; noun clusters ≤ 3 words (e.g., "hydraulic reservoir"); one instruction per sentence; active voice and imperative mood for procedures; no continuous (-ing), perfect tenses, or passive voice in procedures; paragraphs ≤ 6 sentences, one topic per paragraph.

Part 2 – Dictionary : ~900 approved words, each with exactly one part of speech and one meaning; an unapproved-words list provides approved replacements. Examples: commence → START, ensure → MAKE SURE, prior to → BEFORE, replenish → FILL, utilize → USE, approximately → ABOUT, in order to → TO. A word like CLOSE is only allowed as a verb ("close the valve"); "near" must use NEAR.

Warnings and cautions are standardized: WARNING for injury risk, CAUTION for damage risk, never mixed.

Rewrite example — Standard industrial English:

It is imperative that the operator ensures the hydraulic reservoir is replenished prior to commencing operation.

(13 words, multiple unapproved terms). ASD-STE100 version:

Make sure that the hydraulic reservoir is full before you start the operation.

(13 words, all approved vocabulary, no meaning loss).

LLMs excel at STE because its rules are verifiable constraints (word counts, banned words, one-word-one-meaning, active voice) — far more effective than vague prompts like "write clearly." Karpathy notes the spec is strict, so he often asks for "80% of the way to ASD-STE100."

Applying Traditional Discipline to goal-workflow

The author maps three safety-critical industry practices into the open-source goal-workflow ( github.com/smallnest/goal-workflow) as three skills:

Requirements expression : EARS syntax + Design by Contract → /prd Acceptance verification : DO-178C bidirectional traceability + formal verification → /review-it Forbidden zones : MISRA hard prohibitions →

/smell

/prd: EARS Requirements with Contracts

In DO-178C (airborne software) and ISO 26262 (automotive functional safety), prose requirements are banned because natural-language ambiguity propagates into product defects. EARS (Easy Approach to Requirements Syntax) restricts requirements to five fixed patterns:

Universal : "The system shall …" — functions that must hold at all times.

Event-driven : "When [trigger], the system shall …" — user actions or external signals.

State-driven : "While [state], the system shall …" — only valid in a specific state.

Unwanted : "If [anomaly], the system shall …" — failure paths, defensive behavior.

Complex : "When [event] and [state], the system shall …" — combinations using AND/OR.

Syntax alone is insufficient; semantic precision comes from Design by Contract (DbC). Every functional requirement must carry three checkable clauses: precondition (what must be true before), postcondition (what must be true after), and invariant (what always holds). A requirement without a verifiable contract is rejected.

Generated example (FR-3: Modify Task Priority) :

Pattern: Event-driven

Statement: "When the user selects a new priority in the edit dialog, the system shall immediately save the change."

Precondition: Task exists; user has edit permission; new value ∈ {high, medium, low}

Postcondition: Task priority updated to new value; badge color updates accordingly

Invariant: Task always has exactly one legal priority

The precondition clause new value ∈ {high, medium, low} blocks illegal inputs at the door. The principle: "The more structured the requirement, the smaller the hallucination space." An agent implementing a contracted FR-3 produces more predictable output than one given "handle the request correctly."

/review-it: DO-178C Acceptance Gates

DO-178C closes a software change with two disciplines: bidirectional traceability and formal verification .

Bidirectional traceability matrix links Requirements ↔ Code ↔ Verification. Forward traceability ensures every requirement has implementation and test; backward traceability ensures every code change traces to a requirement or rationale. Untraced changes — "orphan changes" — are a common AI problem where agents add out-of-scope implementations.

Formal verification demands evidence, not opinion. A review finding must include one of three evidence types:

Executable reproduction : minimal test script that has actually run.

Rigorous argument : for hard-to-reproduce scenarios (e.g., concurrency), a step-by-step deduction with inputs, state, and control flow.

Tool verdict : output from go vet, tsc --noEmit, static analyzers, cross-checked in context.

"This might be a problem" is not a finding; "This crashes on input X, here is the reproducing test" is. Conversely, "looks fine" is not a clean bill; clean requires test or proof backing.

Five explicit acceptance criteria at close-out:

1. Every requirement implemented — evidenced by diff traceability to code.

2. Every implemented behavior verified — each behavior has test or proof.

3. No orphan changes — every change traces to requirement or rationale.

4. Every finding dispositioned — issue log with fix/defer/reject and reason.

5. Residual risks recorded — risk register in final report.

The final report mirrors DO-178C's Software Accomplishment Summary: traceability matrix, issue log, risk register, conformance statement. Passing requires evidence for every criterion; absence of findings does not constitute evidence. Explicit user waiver is the only exception.

/smell: MISRA Hard Prohibitions

MISRA (Motor Industry Software Reliability Association) governs C/C++ in automotive embedded systems. Unlike heuristic "code smells" that flag candidates for context review, MISRA rules are hard prohibitions : violation = confirmed finding, no candidate phase. The only gate is a safety-criticality scope check — code in embedded, aerospace, automotive, medical, payments, real-time routing, or core infrastructure defaults to Critical/Warning; explicitly out-of-scope code downgrades to advisory.

13 rules (S1–S13) each remove a class of unprovable behavior. Examples:

S1 Unbounded recursion : recursive calls without depth limit — termination unprovable, stack overflow risk.

S3 Pointer/reference chain depth : >2 levels of dereferencing or raw pointer arithmetic on containers — type system cannot guarantee memory safety.

S5 Loop nesting limit : >3 nested loops in one function.

S7 Reliance on undefined behavior : signed overflow, uninitialized reads, evaluation-order dependence — language spec gives no guarantees.

S8 Implicit conversions : implicit narrowing/widening, signed/unsigned mixing — value semantics not preserved, silent truncation possible.

S10 Swallowed error codes : unchecked return values, exceptions swallowed on critical paths — failures become silent, callers cannot decide.

Each prohibition's rationale is "deterministic reasoning": recursion and unbounded loops break termination proofs; dynamic allocation and global state make execution depend on heap/history, not just inputs; pointer arithmetic, aliasing, unions, and implicit conversions break memory safety and value semantics guaranteed by the type system; swallowed error codes make failure invisible.

Summary

Traditional safety-critical industries spent decades answering: how to prevent capable but fallible writers from making mistakes in systems where errors cost lives? Their answer is three layers: requirements layer — controlled expression eliminates ambiguity (EARS, contracts); verification layer — traceability and evidence enforce closure (DO-178C); execution layer — hard prohibitions seal off dangerous features (MISRA).

AI agents occupy the same role: strong coding ability, systematic errors — misreading requirements, injecting orphan code, swallowing errors, depending on undefined behavior. As generation cost drops, the value of the oversight layer rises. This is exactly Karpathy's "oversight and understanding."

The three skills map to the three oversight checkpoints: /prd governs requirements — answers "what is needed?" /review-it governs acceptance — answers "is it correct?" /smell governs forbidden zones — answers "any landmines?"

This approach does not manage AI like a human. It acknowledges AI errs like traditional writers, then uses institutional guardrails proven over decades. Controlled language, acceptance gates, and hard prohibitions transfer almost unchanged into AI workflows. Upgrading an AI workflow by borrowing a discipline package from aerospace and automotive offers high ROI. As Karpathy closed: "Push the boundaries here and you'll be surprised" — true for workflows as well. goal-workflow is open source: github.com/smallnest/goal-workflow. Install with: npx skills add smallnest/goal-workflow In Claude Code, invoke /prd, /review-it, /smell. Documentation:

https://goal.rpcx.io
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Design by ContractMISRAAI workflowsgoal-workflowEARSDO-178CASD-STE100safety-critical systems
BirdNest Tech Talk
Written by

BirdNest Tech Talk

Author of the rpcx microservice framework, original book author, and chair of Baidu's Go CMC committee.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.