Cloud Native 17 min read

Apache APISIX Evolution: From Disrupting Kong to Premier AI Gateway

This article traces Apache APISIX's four-phase evolution: its etcd-based architecture achieving millisecond config updates versus Kong's seconds, its 9-month Apache graduation, multi-language plugin runner and Wasm support, Kubernetes-native ingress controller, and transformation into an AI Gateway with token-based rate limiting, protocol normalization, and SSE streaming optimization.

Ops Development & AI Practice
Ops Development & AI Practice
Ops Development & AI Practice
Apache APISIX Evolution: From Disrupting Kong to Premier AI Gateway

Phase 1: Birth and Breakthrough — Architectural Assault on Traditional Gateway Weaknesses

In early 2019, microservice gateways were dominated by Kong and raw Nginx. Spring Cloud Zuul 1's blocking I/O exhausted thread pools under high concurrency, while Nginx required nginx -s reload for every route or upstream change, causing CPU spikes, memory surges, and forced long‑connection drops.

Kong's Legacy Burdens

Heavy relational DB dependency: Kong stored routes, consumers, and plugin configs in PostgreSQL or Cassandra. Every gateway node start or config refresh hit the DB, turning the database into a single‑point bottleneck at scale.

Second‑level config propagation latency: To avoid per‑request DB queries, Kong cached locally and polled for changes. New routes or auth plugins took seconds to reach all nodes — unacceptable for cloud‑native real‑time elasticity.

Routing algorithm degradation: Early Kong used linear lookup and regex scanning; latency grew O(N) with route count, dragging down data‑plane throughput.

Technical Breakthrough: Drop Relational DB, Embrace etcd Watch

Raft‑based strong consistency and decentralized HA: etcd clusters provide high‑throughput reads/writes with strong consistency, eliminating traditional master‑slave replication ops.

gRPC bidirectional streaming + Watch: APISIX workers establish HTTP/2 long‑connections to etcd, watching prefixes like /apisix/routes/ and /apisix/upstreams/. Config changes are pushed incrementally in milliseconds.

Zero‑reload, in‑process hot update: Workers apply updates directly in Lua memory — no Nginx reload, no worker restart. Existing long connections stay intact; config propagation latency drops from seconds to under 2 ms .

Radix tree (Radixtree) ultra‑fast routing: The custom lua-resty-radixtree (C‑implemented prefix tree) keeps route lookup at O(K) where K is URI path depth, eliminating CPU linear growth even at 100k routes.

These advantages gave APISIX a decisive “millisecond hot‑deploy, zero‑reload, order‑of‑magnitude performance” edge over early Kong.

Phase 2: Open Source & Record‑Fast Apache Graduation

June 2019: APISIX open‑sourced on GitHub. Founders Ming Wen and Yuan Sheng launched API7.ai and donated all IP to the Apache Software Foundation.

Oct 2019: Entered Apache Incubator.

Apache Way adherence: Transparent governance, async mailing‑list discussions, open committer promotion. Global contributors added enterprise plugins and performance patches.

July 2020: Graduated as a Top‑Level Project (TLP) in just 9 months — fastest Chinese‑origin project at the time — establishing global credibility and neutral governance.

Phase 3: Ecosystem Expansion — Multi‑Language Decoupling & Cloud‑Native Deepening

2020‑2023: Kubernetes became the orchestration standard; service mesh (Envoy/Istio) rose. APISIX broke out of the Nginx/Lua comfort zone.

Breaking the Lua Barrier: Plugin Runner & WebAssembly

Plugin Runner (cross‑language IPC): Spawns a sidecar child process supporting Go, Java, Python . Workers communicate via Unix Domain Socket (UDS) with a lightweight binary RPC. Request context is passed to the runner for custom auth, audit, or crypto logic. Process‑level isolation prevents plugin crashes from taking down the data plane.

WebAssembly (Wasm) sandbox runtime: Integrates Proxy‑Wasm standard. Developers write high‑performance plugins in Rust, C++, TinyGo , compile to Wasm bytecode, and inject directly into APISIX. Removes IPC overhead, keeps millisecond speed, and gains Wasm memory safety.

Kubernetes‑Native Integration: apisix‑ingress‑controller

Official apisix-ingress-controller outperforms default Ingress‑Nginx:

CRD‑based fine‑grained traffic control: Beyond standard Ingress, provides native ApisixRoute, ApisixUpstream, ApisixPluginConfig, ApisixConsumer CRDs for canary, blue‑green, traffic mirroring, dynamic rate limiting via declarative YAML.

Zero‑reload Pod IP direct connection: On Endpoint/Pod IP changes, controller writes incremental updates to etcd; APISIX data plane refreshes upstream nodes in milliseconds, eliminating packet loss and traffic jitter during scaling.

This architecture — extreme performance plus agile multi‑language extension — made APISIX the go‑to replacement for F5 hardware, Zuul, and self‑managed Nginx clusters.

Phase 4: LLM Era Leap — Evolving into an AI Gateway

Since 2023, generative AI/LLM traffic has surged, posing new challenges:

Long‑lived connections & Time‑to‑First‑Token (TTFT): Models stream via Server‑Sent Events (SSE). Traditional HTTP response buffering chunks output, destroying the typewriter effect.

Governance shift from QPS to Tokens: Single requests vary wildly in token count; QPS limits can't prevent burst token spikes that trigger vendor throttling or budget overruns. Need TPM (Tokens Per Minute) and RPM (Requests Per Minute) token‑bucket metering.

Multi‑vendor protocol fragmentation & high failure rates: OpenAI, Anthropic Claude, Google Gemini, Azure OpenAI each have distinct APIs/auth. Public clouds frequently return 429, 504, or partial outages.

APISIX adapted rapidly, becoming an enterprise‑grade AI Gateway :

1. ai-proxy Plugin: Unified Protocol & Smart Routing

Protocol abstraction: Clients call standard OpenAI /v1/chat/completions; APISIX bidirectionally translates to target vendor protocols — no per‑model SDKs needed.

Dynamic weight & cost routing: Distributes requests based on vendor pricing, latency, and model capability.

Auto failover & seamless retry: On 429/5xx from primary model, gateway retries backup vendors or local private clusters (Ollama, vLLM) within milliseconds, boosting SLA.

2. Deep SSE Streaming Optimization

Zero‑buffer pass‑through: Detects text/event-stream header and disables Nginx response buffering, forwarding each token chunk in microseconds.

Streaming audit & usage parsing: Without blocking the stream, captures usage fields ( prompt_tokens, completion_tokens) at connection end, emitting to Prometheus, OpenTelemetry, or log centers for per‑tenant cost accounting and compliance.

Today APISIX, alongside Alibaba's Higress, LiteLLM, and Envoy LLM extensions, forms the core network foundation for production‑grade AI applications.

Summary: From Technical Confrontation to Enterprise Infrastructure Standard

At inception: Challenged relational‑DB dogma, introduced etcd Watch + radix tree, delivering “millisecond hot‑deploy + high throughput” to pierce legacy gateway weaknesses.

Growth phase: Embraced Apache Way, built global community, pioneered Plugin Runner + Wasm to shatter the Lua‑only silo.

Cloud‑native & AI wave: Delivered K8s Ingress controller, then seamlessly morphed into an AI Gateway with token‑level governance, protocol unification, and failover.

For lightweight proxies or small pilots, architectural differences may seem abstract. But at Walmart, China Mobile, or major financial institutions — facing million‑QPS concurrency, cross‑DC HA, zero‑downtime hot deploy, strict compliance audits and SLA guarantees — any Nginx reload jitter or config inconsistency causes direct business loss. That is why battle‑tested, architecturally solid open‑source infrastructure like Apache APISIX remains irreplaceable in enterprise tech selection and senior architect role requirements.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

cloud-nativeKubernetesWebAssemblyAPI GatewayetcdApache APISIXAI GatewayPlugin Runner
Ops Development & AI Practice
Written by

Ops Development & AI Practice

DevSecOps engineer sharing experiences and insights on AI, Web3, and Claude code development. Aims to help solve technical challenges, improve development efficiency, and grow through community interaction. Feel free to comment and discuss.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.