Rack Awareness: From Zero to High‑QPS – Boost Availability, Cut Cross‑AZ Traffic
A real‑world rack‑power outage showed that three‑replica Kafka clusters can lose all replicas when brokers share a failure domain, prompting a deep dive into rack awareness—how fault‑domain tags are injected, replica‑placement and leader‑distribution algorithms, consumer‑proximity reads, bandwidth costs, failure scenarios, and the stepwise evolution from hundred‑thousand to ten‑million QPS deployments.
