Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks
The DSP advertising data pipeline was re‑architected by splitting three distinct links—NRR, BT, and CT—and assigning DLF Paimon for massive low‑frequency data, EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and unified service‑layer merging, achieving 2‑minute freshness, sub‑5 ms point queries, ~60% storage cost reduction, and improved fault isolation.
Background and Challenges
In DSP advertising, a single data link must support online serving, billing, ad‑hoc analysis, and BI reporting, each with vastly different freshness, latency, back‑fill, and cost requirements. The existing StarRocks All‑in‑One architecture caused low‑frequency cold data to occupy expensive local storage, high‑frequency writes to affect query stability, and long‑window back‑fills to threaten SLA compliance. Data volume reached 9.6 PB with a daily increment of ~26 TB, and the target was sub‑2‑minute freshness with millisecond‑level online point‑lookup latency.
From All‑in‑One to Differentiated Links
The solution replaces the monolithic pipeline with three specialized links:
NRR link : ~25 TB daily increment, low freshness requirement (hour‑level).
BT link : ~1 TB daily increment, highest demands for real‑time, back‑fill (30 days), and online point‑lookup.
CT link : Small volume, used for merging and supplementing via primary‑key tables.
All links share a unified Kafka multi‑topic entry but diverge in storage, compute, and query paths.
NRR Link – DLF Paimon for Large Low‑Frequency Data
NRR data is written via EMR Serverless Spark batch or Flink micro‑batch to DLF Paimon on object storage, then queried through EMR Serverless StarRocks External Catalog. This design yields two benefits:
Low‑frequency data no longer occupies costly StarRocks local disks, reducing overall storage cost by ~60%.
StarRocks can still query Paimon tables without data migration, preserving second‑level ad‑hoc query capability.
BT Link – EMR Serverless StarRocks Primary Tables for Low‑Latency Queries
BT data is streamed by Flink into StarRocks primary tables with a 60 s checkpoint interval. The primary‑key model supports high‑frequency updates and sub‑5 ms P99 point‑lookup latency for online serving and billing. The link also serves:
Online serving & billing : Direct PK point‑lookup, P99 < 5 ms.
Ad‑hoc analysis : Same real‑time table provides ~2 minute data freshness.
BI reporting : Asynchronous materialized views refresh every ~10 minutes, decoupling heavy aggregation from the online path.
This separation ensures that online point‑lookups remain unaffected by BI aggregation load.
2‑Minute Freshness and the Role of Materialized Views
Instead of relying on materialized views for all real‑time needs, the design leverages the primary‑key table’s inherent real‑time capability to achieve 2‑minute freshness for online and ad‑hoc scenarios. Materialized views are relegated to BI aggregation, where a 10‑minute asynchronous refresh is sufficient, dramatically lowering write and refresh pressure on StarRocks.
Architecture Gains
Query freshness improved from 10 minutes to 2 minutes.
Online point‑lookup P99 latency consistently < 5 ms.
Storage cost reduced by ~60% thanks to object‑storage‑based Paimon for low‑frequency data.
Fault isolation enhanced: each link can evolve and be optimized independently, limiting failures to a single link.
Conclusion – Collaborative Role of Alibaba Cloud Big‑Data Stack
The transformation does not replace a single component but redefines responsibilities across the data pipeline: NRR uses DLF Paimon for massive low‑frequency ingestion, BT relies on EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and CT merges supplemental data while materialized views accelerate BI aggregation. The combined use of Flink, DLF Paimon, EMR Serverless Spark, and StarRocks demonstrates a sustainable near‑real‑time data‑warehouse architecture for large‑scale advertising workloads.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Alibaba Cloud Big Data AI Platform
The Alibaba Cloud Big Data AI Platform builds on Alibaba’s leading cloud infrastructure, big‑data and AI engineering capabilities, scenario algorithms, and extensive industry experience to offer enterprises and developers a one‑stop, cloud‑native big‑data and AI capability suite. It boosts AI development efficiency, enables large‑scale AI deployment across industries, and drives business value.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
