Why OpenAI Sends Engineers On-Site: The FDE Playbook for Enterprise AI
OpenAI's Forward Deployed Engineering lead Colin Jarvis reveals why stronger models require on-site engineers to bridge the gap between demo capabilities and production trust, detailing eval-driven development, deterministic-probabilistic boundaries, and the 'Eat Pain → Excrete Product' methodology across Morgan Stanley, semiconductor, supply chain, and Klarna case studies.
Introduction
This article captures a 2025 conversation between former Palantir Forward Deployed Engineer (FDE) Apoorv Agrawal and OpenAI FDE lead Colin Jarvis. They explore why OpenAI maintains a growing FDE team despite increasingly capable models, and how the FDE model differs from traditional consulting.
01 Why Stronger Models Demand More FDEs
Colin Jarvis joined OpenAI in November 2022, weeks before ChatGPT's release. Early enterprise projects (Morgan Stanley, Klarna) revealed that simply handing over APIs and documentation rarely moved customers to production. The reliable path: engineers embed on-site to understand data, processes, and how frontline users actually work.
Morgan Stanley Case Study
In 2023, Morgan Stanley wanted GPT-4 to surface internal research for wealth advisors — a classic RAG scenario before RAG had matured engineering patterns. The technical pipeline (retrieval tuning, guardrails) was functional in 6–8 weeks. The subsequent 4 months were spent on pilot, annotation, eval-driven iteration, and building advisor trust. Because GPT-4 is probabilistic, the team could not guarantee deterministic outputs like traditional software. They had advisors use the system live, annotate results, accumulate eval sets, and iterate on real failures while preserving the ability to verify answers against source documents.
Result: ~98% advisor adoption and ~3× increase in research-report usage. Apoorv's takeaway: "Build trust through iteration, not just technical excellence." Production readiness requires answering: Can the answer be verified? Are critical operations guarded? Can model errors be traced? Will users trust the system with their workflow? FDE work lives in the gap between "model can do it" and "enterprise dares to use it."
02 Deploying Agents: Not Autonomous from Day One
Semiconductor Verification Agent
At a European semiconductor firm, OpenAI FDEs spent weeks mapping the chip R&D value chain and focused on Verification, where engineers spend an estimated 70–80% of their time running tests, fixing bugs, and handling legacy compatibility. The initial agent (using Codex) did not modify code; it investigated failures — reading logs, examining code, generating tickets with hypothesized root causes. Only after engineers trusted the investigations did the team add an execution environment so the agent could run tests, observe results, and iterate fixes. The tool evolved into a Debug Investigation and Triage Agent, aiming to have simple issues resolved and complex ones pre-investigated by morning.
Eval-Driven Development
Colin emphasized that real engineering workflows involve 20+ sequential actions across systems. OpenAI first builds real trajectories with domain experts, creates labeled eval sets, then develops against them. LLM-driven capability is not considered complete until eval validates it.
Automotive Supply-Chain Orchestrator
For an APAC automaker, OpenAI left data in existing systems and built APIs for an LLM orchestrator. In a demo scenario (25% tariff on China→Korea exports), the system identified affected parts, analyzed cost propagation, and re-optimized sourcing. Crucially, hard business rules (e.g., dual-sourcing requirements, lead-time limits) were encoded as deterministic constraints, not prompts. Colin's principle: "Whenever you can use determinism, do it." Ambiguous semantic judgments go to the LLM; rules requiring 100% correctness go to deterministic code. Verification paths let users inspect underlying tables. Optimization used a simulator: the model adjusted parameters, ran multiple scenarios, compared cost/delivery trade-offs, and recommended a plan that still passed deterministic validation.
Apoorv summarized two pillars: Use eval-driven development and Trade off determinism and probabilism deliberately . Agent maturity is not measured by autonomous step count but by mechanisms that add a layer of verifiability with each expanded permission.
03 Klarna: From Custom Project to Reusable Product
Apoorv raised the classic FDE critique: isn't this just consulting? Colin answered with the Klarna customer-service project. Early on, hundreds of policies made hand-written prompts and tool calls unmaintainable. The team parameterized instructions and tools, backed each intent with evals, and scaled from ~20 to 400+ policies.
After solving Klarna's problem, they extracted a general framework (later open-sourced as Swarm). One success wasn't enough; they validated the primitives on larger, more complex engagements (e.g., T-Mobile). Only then did FDE and product teams productize into the Agents SDK and AgentKit. OpenAI positions FDE as a 0→1 team: dive into the hardest concrete problems, then hand off to scaling teams or partners.
Apoorv's "Eat Pain → Excrete Product" framework: first engagement yields ~20% reusable code; after 2–3 similar engagements, ~50% reusable; only then push to a platform team. The danger is generalizing too early — building a generic solution before discovering the real problem. Colin warns founders: clarify whether FDE exists for services revenue or for product discovery. Mixing both leads to becoming a consulting shop because short-term service revenue is addictive.
OpenAI targets two problem classes: (1) clear product hypotheses (e.g., customer service, document generation) needing design partners; (2) extremely hard domain problems (semiconductors, life sciences) that may feed research. Economic value must be in the $10M–$1B+ range. As of the interview, OpenAI FDE grew from 2 to 39 people, targeting 52 by year-end, with no intention of becoming a thousands-person implementation arm.
04 Beyond MCP: The Missing Logic / Metadata Translation Layer
Asked about the next big opportunity, Colin pointed not to new model architectures but to translating raw enterprise data and hidden business logic into something LLMs can actually use. Data sits in warehouses, SharePoint, internal systems; logic lives in undocumented workflows. An MCP connector lets a model call a system, but doesn't tell it when to combine warehouse data with SharePoint docs or how to apply implicit business rules. FDEs repeatedly build a lightweight "Logic Layer" or "Metadata Translation Layer" to bridge this gap.
Apoorv connected this to Palantir's Ontology. Colin agreed the challenge isn't new — data engineering has long debated centralization vs. virtual access layers — but the consumer has changed: LLMs now write queries and combine data, so metadata, semantics, and data organization directly determine agent correctness.
Future: Fine-Tuning for Specialized Domains
Colin's long/short bet: fine-tuning. As agent orchestration, tool use, eval, annotation, and execution environments mature, teams can more easily generate training data, annotate rapidly, build training sets, and fine-tune models for specific GenAI use cases (chip design, drug discovery). Previously, these domains lacked sufficient real-task data; now agent systems can record processes, evaluate outcomes, and curate datasets. The next phase may not just be more complex agent workflows but distilling some capabilities back into the model via fine-tuning.
Across Morgan Stanley, semiconductors, supply chain, Klarna, and T-Mobile, the pattern is consistent: models hit enterprise constraints — trust, eval, data, business rules, environmental limits — not raw intelligence. FDE enters the most concrete problems, solves them, then extracts primitives that become products, whose production loops may eventually feed model improvement. As Apoorv closed: "Customer projects aren't the endpoint; they're the starting point for product discovery."
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
