NAT Gateway Async Logging: DPDK Lock-Free Queue Benchmarks & Optimization

This article benchmarks DPDK lock-free queue modes for asynchronous logging in a high-concurrency NAT gateway, showing single-core throughput of ~800k logs/sec and multi-core SPSC design achieving 4.6M logs/sec, while analyzing compiler optimization impact and producer-consumer scaling behavior.

360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
NAT Gateway Async Logging: DPDK Lock-Free Queue Benchmarks & Optimization

Background

The NAT gateway in the ops environment must log every connection establishment. Under high concurrency, direct I/O on forwarding cores degrades packet forwarding performance. The solution is to offload logging to dedicated non-forwarding cores via asynchronous logging. The same requirement exists for aisa traffic logging with even higher volume. DPDK lock-free queues, built on hugepages, convert I/O into memory reads/writes, boosting forwarding core performance. However, a rate mismatch exists: forwarding cores write to memory, while log-handling cores read from memory and perform I/O, making the consumer slower.

Test Methodology

The test simulates the extreme case: a master core continuously writes to the lock-free queue (blocking when full), while a log consumer core reads and guarantees no loss. Log entry length and format match the ops connection-log scenario. DPDK lock-free queues support four modes: single-producer single-consumer (SPSC), multi-producer single-consumer (MPSC), single-producer multi-consumer (SPMC), and multi-producer multi-consumer (MPMC). SPSC has no resource contention and serves as the performance baseline. SPMC is irrelevant because forwarding cores are multiple producers. GCC optimization levels (-O3 vs -O0) are also compared.

Single-Core Benchmark Results

With only the master core writing (regardless of queue mode), the lock-free queue limit is approximately 800k logs/sec. -O3 optimization yields ~36% improvement over -O0. Single-producer modes slightly outperform multi-producer modes, but the gap is small. Writing to stdout instead of a file drastically reduces throughput (23,749 logs/sec for MPMC at -O0), confirming file I/O is not the bottleneck in the primary tests.

Multi-Core Write Analysis

Realistic test: 8 forwarding cores writing concurrently, one consumer core. MPSC mode at -O3 reaches 2,601,137 logs/sec vs 803,651 for single-producer -O3 — a >3x increase. The author explains: when the queue is full, a producer spins on a continue loop; the gap between the queue becoming non-full and the producer re-checking creates idle time. With multiple producers, another core may find the queue non-full during that gap, reducing wasted cycles. Thus more producers better utilize the consumer's drain rate.

Log Optimization Design

The community dpvs version uses a single log-output core to avoid file locks and preserve log order. For ops, connections are sharded across forwarding cores (lock-free), so all logs for a given connection stay on one core. This allows pairing each forwarding core with a dedicated log-output core, each writing to its own file. The lock-free queue can then use SPSC mode — the highest-throughput mode — with independent, contention-free pipelines.

Per-core log pipeline architecture
Per-core log pipeline architecture

Multi-Core SPSC Benchmark Results

With 8 cores (8 SPSC pipelines), -O3 file output achieves 4,638,113 logs/sec; 4 cores achieve 4,615,216 logs/sec; single core achieves 1,021,821 logs/sec. -O0 yields 3,891,251 (8 cores), 3,873,803 (4 cores), 826,875 (1 core). Scaling plateaus at 4 cores, suggesting system I/O becomes the bottleneck.

Conclusions

SPSC mode delivers peak lock-free queue performance; MPSC/MPMC incur modest overhead.

Throughput scales near-linearly with core count until system I/O saturates.

-O3 compilation improves performance ~30%; -g debug symbols do not affect runtime performance and are recommended even for release builds to aid debugging.

The community async logger (~2M logs/sec) suffices for ops (typically 10k-20k new connections/sec); multi-file output is unnecessary unless required.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

performance benchmarkcompiler optimizationDPDKlock-free queueasync loggingMPSCNAT gatewaySPSC
360 Zhihui Cloud Developer
Written by

360 Zhihui Cloud Developer

360 Zhihui Cloud is an enterprise open service platform that aims to "aggregate data value and empower an intelligent future," leveraging 360's extensive product and technology resources to deliver platform services to customers.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.