OpenAI API Evolution: Completions to Responses — Why Open Source Still Uses Chat Completions
This article traces OpenAI's API evolution across six generations from Completions to Responses API, explains why the open-source ecosystem remains anchored to Chat Completions despite official advances, compares architectural trade-offs between stateless and stateful paradigms, and provides a practical selection guide for engineers building heterogeneous model gateways or agent systems.
Six Generations of OpenAI API Evolution
From GPT-3's launch in 2020 to the present, OpenAI's LLM API has undergone six key paradigm shifts.
OpenAI API six-generation evolution map
1. First Generation: Stateless Text Completion (/v1/completions)
From 2020 to 2022, the primary endpoint was POST /v1/completions for models like text-davinci-003 and code-davinci-002. The design mirrored the raw autoregressive language model: given a prefix, predict the next most probable token.
{
"model": "text-davinci-003",
"prompt": "请解释什么是量子计算:",
"max_tokens": 256,
"temperature": 0.7
}Pain points: No concept of conversation, roles, or system instructions. Multi-turn dialogue required manual concatenation of history into a single string, which was fragile — extra User: tokens or misconfigured stop sequences could break the dialogue state.
2. Second Generation: Role Semantics and Industry De Facto Standard (/v1/chat/completions)
On March 1, 2023, with gpt-3.5-turbo, OpenAI introduced POST /v1/chat/completions. The breakthrough was ChatML (Chat Markup Language) and a structured messages array with a role system:
{
"model": "gpt-3.5-turbo",
"messages": [
{"role": "system", "content": "你是一位资深的分布式系统架构师。"},
{"role": "user", "content": "什么是 Raft 协议?"}
]
}Core value: Role-based decoupling — system prompts, user queries, and assistant replies gained clear semantic boundaries, greatly improving control and safety.
Streaming standard: stream: true with Server-Sent Events (SSE) established the industry-wide token-by-token streaming format ( data: {"choices":[{"delta":{"content":"..."}}]}).
Industry status: This endpoint became the de facto "HTTP communication protocol" for open-source models and heterogeneous inference backends.
3. Third Generation: From Conversation to Action (Functions and Tools)
To let models drive external software, OpenAI upgraded the protocol twice:
June 2023 (Function Calling): Added top-level functions parameter using JSON Schema to describe external function signatures. Models could return structured function_call (name + JSON arguments) instead of plain text.
November 2023 (Tools API unification): At DevDay, functions evolved into an extensible tools list with tool_choice and single-turn parallel tool calling . Models could emit multiple tool_calls in one inference, executed concurrently by the client:
{
"model": "gpt-4-turbo",
"messages": [...],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"}
},
"required": ["city"]
}
}
}
]
}This completed the leap from "chat assistant" to "agent brain".
4. Fourth Generation: Rigid Format Guarantees and Cloud State Exploration
Enterprise agents faced two bottlenecks: JSON syntax errors/hallucinations, and high barriers for multi-turn memory and retrieval. OpenAI responded with:
Structured Outputs (August 2024): In /v1/chat/completions, a strict response_format with json_schema and strict: true enforces 100% schema compliance via grammar-based masking during token sampling.
Assistants API rise and fall (/v1/assistants, /v1/threads, /v1/runs): OpenAI's first attempt to manage the full agent lifecycle — conversation history (Threads), file search, code interpreter — all hosted in OpenAI's cloud. Clients only started a Run and polled status. However, opaque state, high async latency, and poor local tool integration led to heavy criticism. OpenAI deprecated it after launching Responses API and fully shut it down on August 26, 2026 .
5. Fifth Generation: Unified Agent Runtime (/v1/responses)
On March 11, 2025, OpenAI released POST /v1/responses, designed to merge Chat Completions' simplicity with Assistants API's power :
Top-level parameter normalization: Replaced confusing system role with top-level instructions; unified user input and multimodal content in input.
Lightweight state persistence: store: true plus previous_response_id links multi-turn context without re-uploading thousands of history tokens.
Native agent tool loop: Built-in web search, file retrieval, Python code interpreter, and deep support for Anthropic's open MCP (Model Context Protocol).
Native reasoning model support: Structured thought-process output for models like o1 and o3-mini, separating reasoning from final reply.
6. Sixth Generation: Open Open Standard (Open Responses, January 2026)
Responses API remains a proprietary endpoint. On January 15, 2026, OpenAI, Hugging Face, OpenRouter, Vercel, and others launched Open Responses ( openresponses.org), an open, vendor-neutral specification inheriting Responses API's agentic-loop core:
Vendor neutrality: Same client code connects to OpenAI, Anthropic, Google Gemini, and local open-source models.
Atomic Item architecture: Context decomposed into clear Item units for consistent state updates, tool-call traces, and reasoning streaming.
Semantic streaming: Replaces Chat Completions' crude text deltas with structured event streams.
Deep Dive: Why Open Source Sticks with Chat Completions
Despite OpenAI's 2025 move to Responses API, "OpenAI-compatible" in open source and third-party gateways still points to /v1/chat/completions 95%+ of the time . This isn't sluggishness; it's dictated by inference and agent software architecture fundamentals.
Technical selection showdown: open-source de facto vs official integrated evolution
1. Architectural Orthogonality: Stateless Inference vs Stateful Runtime
High-performance inference engines (vLLM, SGLang, TGI, Ollama) focus on squeezing hardware throughput — efficient matrix multiplication and memory management (PagedAttention, RadixAttention) .
Chat Completions = pure stateless computation: Input tokens in, output tokens out. Compute finishes, memory and context released immediately. No server-side user sessions, DB connections, or file storage.
Responses API = stateful SaaS runtime: Requires HA relational DB for session state, object storage for uploaded files, sandbox containers for Python execution, vector DB for retrieval.
Demanding vLLM or Ollama to implement full Responses API is like asking the Linux kernel to run e-commerce microservices and databases — it breaks layering orthogonality.
2. Control Ownership: Should Orchestration Live in Client or Server?
Enterprise architectures follow "model does compute, application layer does control flow" :
Session state (history) must stay in the company's own PostgreSQL/Redis for audit, compliance, and archiving.
Tool execution and business system integration (internal ERP, payment gateways) must run in controlled on-prem environments — never expose internal credentials or DB permissions to a public cloud model service.
Agent reasoning loops should be governed by flexible application frameworks (Claude Code, Antigravity, LangChain, LlamaIndex, etc.).
Thus open source treats the model as a "pure compute endpoint" via standardized /v1/chat/completions, not a cloud-hosted business state manager.
3. Anti-Vendor Lock-in and Heterogeneous Model Hot-Swapping
Chat Completions' greatest contribution is erasing protocol gaps between heterogeneous models .
Production systems dynamically route by task complexity: simple intent classification → local lightweight model; complex reasoning → DeepSeek-R1; multimodal → Qwen2.5-VL; general generation → GPT-4o. Because all share the same messages and choices protocol, the upper layer only changes base_url and model for seamless sub-second switching. Deep binding to a vendor-specific stateful endpoint makes migration costs exponential.
Detail Showdown: Chat Completions vs Responses API Core Evolution
Comparing execution loop, request payload, state management, and output consumption.
Execution loop analysis: client-side multi-hop vs server-side atomic
1. Execution Loop: Client Multi-Hop vs Server Atomic Loop
Chat Completions (Mode A): Classic client-orchestrated multi-hop loop . Model returns tool_calls; client executes local tools; client must repackage original conversation, model's tool_calls, and tool results in full for a second inference round. Two full network round-trips, second round re-transmits all history tokens.
Responses API (Mode B): Server-side unified atomic loop . Client sends one request; server completes reasoning, tool execution, and answer assembly in its cloud sandbox, returning final output_text and execution trace in a single call.
2. Request Payload: From messages to instructions + input
Old Chat Completions mixes system persona into the first message:
# Chat Completions (/v1/chat/completions)
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "你是一位资深的金融风控分析师。"},
{"role": "user", "content": "请分析苹果公司最新的资产负债表。"}
]
)Responses API lifts system persona to a top-level constant, flattens input:
# Responses API (/v1/responses)
response = client.responses.create(
model="gpt-4o",
instructions="你是一位资深的金融风控分析师。",
input="请分析苹果公司最新的资产负债表。"
)3. Session State and Token Optimization
In Chat Completions, client manages all context. By turn 10, the client must re-send thousands of tokens from prior turns, consuming upload bandwidth and adding client-side sliding-window complexity.
Responses API with persistence:
# Turn 1: enable cloud persistence
resp1 = client.responses.create(
model="gpt-4o",
input="你好,我正在设计高并发分布式锁方案。",
store=True
)
# Turn 2: reference previous response ID, no history re-transmission
resp2 = client.responses.create(
model="gpt-4o",
input="如果发生网络分区导致脑裂,该如何防范?",
previous_response_id=resp1.id
)[!NOTE] Billing and context cost note: previous_response_id eliminates client upload bandwidth and latency, but OpenAI still bills historical context as input tokens . When multi-turn hits cloud prefix caching, Prompt Caching discounted rates apply.
4. Structured Output: response_format vs text.format
Chat Completions declares structured output via top-level response_format:
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "audit_report",
"strict": true,
"schema": { ... }
}
}Responses API consolidates into text.format:
response = client.responses.create(
model="gpt-4o",
input="提取合同文本中的交易双方与签约金额",
text={
"format": {
"type": "json_schema",
"name": "contract_info",
"strict": True,
"schema": { ... }
}
}
)5. Response Consumption: From choices Unwrapping to output_text
Chat Completions requires deep unwrapping:
# Old Chat Completions
reply_text = response.choices[0].message.contentResponses API provides convenient output_text and structured output array for tool traces:
# New Responses API direct reply text
reply_text = response.output_text
# Inspect agent internal execution:
for item in response.output:
if item.type == "message":
print("Text message:", item.content)
elif item.type == "web_search_call":
print("Triggered web search:", item.action)Open-Source Compatibility Future: From Chat to Open Responses
Status quo: Chat Completions remains unshakable. In 2026, vLLM, Ollama, SGLang, DeepSeek, Qwen, etc., still prioritize /v1/chat/completions with the most complete docs. Open-source frameworks use built-in Jinja2 Chat Templates to render standard messages into model-specific prompts (Llama-3, Qwen, DeepSeek templates) and implement Function Calling via regex/JSON parsing.
Bottleneck: Protocol limits for complex agents. As agents move from single-step Q&A to multi-step planning, long-chain reasoning, and complex tool use, Chat Completions' simple choices[0].delta streaming shows limits: cannot structurally expose reasoning chains, tool and text streams conflate, client network retransmission overhead huge.
Breakthrough: Open Responses building cross-vendor open bridge. The 2026 Open Responses spec is being adopted by OpenRouter, LiteLLM, Hugging Face. Its goal is not to burden open-source inference engines with heavy DB/SaaS duties, but to define a lightweight, agent-loop-oriented standard protocol object — preserving Responses API's atomicity and semantic streaming advantages while keeping open source's anti-lock-in tradition.
Engineering Landing: Architect's Technical Selection Guide
Facing "open-source de facto standard" vs "official integrated endpoint", how should engineers decide?
Selection Matrix
Scenario 1: Enterprise heterogeneous model gateway & private deployment
Recommended: Chat Completions (/v1/chat/completions)
Engines: vLLM, Ollama, SGLang, DeepSeek, Qwen
Core need: Absolute vendor lock-in elimination, private data compliance
State ownership: Client-side DB (PostgreSQL/Redis)
Architecture advice: Unify on this endpoint as internal bus; business layer builds own loop
Scenario 2: Cloud-heavy native agents
Recommended: Responses API (/v1/responses)
Engines: OpenAI official GPT-4o, o1, o3-mini
Core need: Dev efficiency first, avoid building cloud execution sandboxes
State ownership: OpenAI cloud-hosted ( store: true)
Architecture advice: Leverage built-in WebSearch / MCP / Prompt Cache
Scenario 3: Future cross-vendor multi-agent architecture
Recommended: Open Responses spec adoption
Engines: OpenRouter, LiteLLM, mixed multi-vendor gateways
Core need: Balance agent-loop capabilities with multi-model hot-swapping
State ownership: Application-layer state hub, standardized Item interaction
Architecture advice: Build protocol adaptation layer based on Open Responses for gradual evolution
Selection Rules Detail
Firmly choose Chat Completions when: Building private deployments for finance, healthcare, government, or heavily relying on DeepSeek/self-hosted inference clusters. Chat Completions is the only choice combining ecosystem maturity with heterogeneous compatibility.
Try Responses API when: Agile team, 100% dependent on OpenAI closed models, need agents with real-time web search, Python plotting, or remote MCP services. Responses API saves massive effort building Docker sandboxes and complex retry loops.
Watch Open Responses when: Building a platform-level agent framework, want atomic Item streaming and structured reasoning-chain output without locking into a single closed platform. Design the bottom protocol adaptation layer against the Open Responses spec.
Summary
Protocol evolution is never "new version out, old version immediately obsolete" — it follows a clear "dual-track" pattern:
One track: Universal interconnect spec for foundational infrastructure . Like HTTP/1.1 remaining the internet's bedrock for 20+ years, /v1/chat/completions — with its extreme simplicity and pure stateless mathematical nature — has firmly established the compute foundation of the open-source LLM world.
Other track: Vertical integration of agent operating systems . OpenAI's push for Responses API and the industry's Open Responses effort essentially raise the API abstraction level from "text completion machine" to "general-purpose agent runtime".
Understanding the essential watershed between these two tracks, and clarifying what "OpenAI-compatible" truly means in different contexts, enables the clearest, most robust architectural decisions amid the noisy technology wave.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Development & AI Practice
DevSecOps engineer sharing experiences and insights on AI, Web3, and Claude code development. Aims to help solve technical challenges, improve development efficiency, and grow through community interaction. Feel free to comment and discuss.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
