Big Data 12 min read

Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks

The DSP advertising data pipeline was re‑architected by splitting three distinct links—NRR, BT, and CT—and assigning DLF Paimon for massive low‑frequency data, EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and unified service‑layer merging, achieving 2‑minute freshness, sub‑5 ms point queries, ~60% storage cost reduction, and improved fault isolation.

Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks

Background and Challenges

In DSP advertising, a single data link must support online serving, billing, ad‑hoc analysis, and BI reporting, each with vastly different freshness, latency, back‑fill, and cost requirements. The existing StarRocks All‑in‑One architecture caused low‑frequency cold data to occupy expensive local storage, high‑frequency writes to affect query stability, and long‑window back‑fills to threaten SLA compliance. Data volume reached 9.6 PB with a daily increment of ~26 TB, and the target was sub‑2‑minute freshness with millisecond‑level online point‑lookup latency.

From All‑in‑One to Differentiated Links

The solution replaces the monolithic pipeline with three specialized links:

NRR link : ~25 TB daily increment, low freshness requirement (hour‑level).

BT link : ~1 TB daily increment, highest demands for real‑time, back‑fill (30 days), and online point‑lookup.

CT link : Small volume, used for merging and supplementing via primary‑key tables.

All links share a unified Kafka multi‑topic entry but diverge in storage, compute, and query paths.

NRR Link – DLF Paimon for Large Low‑Frequency Data

NRR data is written via EMR Serverless Spark batch or Flink micro‑batch to DLF Paimon on object storage, then queried through EMR Serverless StarRocks External Catalog. This design yields two benefits:

Low‑frequency data no longer occupies costly StarRocks local disks, reducing overall storage cost by ~60%.

StarRocks can still query Paimon tables without data migration, preserving second‑level ad‑hoc query capability.

BT Link – EMR Serverless StarRocks Primary Tables for Low‑Latency Queries

BT data is streamed by Flink into StarRocks primary tables with a 60 s checkpoint interval. The primary‑key model supports high‑frequency updates and sub‑5 ms P99 point‑lookup latency for online serving and billing. The link also serves:

Online serving & billing : Direct PK point‑lookup, P99 < 5 ms.

Ad‑hoc analysis : Same real‑time table provides ~2 minute data freshness.

BI reporting : Asynchronous materialized views refresh every ~10 minutes, decoupling heavy aggregation from the online path.

This separation ensures that online point‑lookups remain unaffected by BI aggregation load.

2‑Minute Freshness and the Role of Materialized Views

Instead of relying on materialized views for all real‑time needs, the design leverages the primary‑key table’s inherent real‑time capability to achieve 2‑minute freshness for online and ad‑hoc scenarios. Materialized views are relegated to BI aggregation, where a 10‑minute asynchronous refresh is sufficient, dramatically lowering write and refresh pressure on StarRocks.

Architecture Gains

Query freshness improved from 10 minutes to 2 minutes.

Online point‑lookup P99 latency consistently < 5 ms.

Storage cost reduced by ~60% thanks to object‑storage‑based Paimon for low‑frequency data.

Fault isolation enhanced: each link can evolve and be optimized independently, limiting failures to a single link.

Conclusion – Collaborative Role of Alibaba Cloud Big‑Data Stack

The transformation does not replace a single component but redefines responsibilities across the data pipeline: NRR uses DLF Paimon for massive low‑frequency ingestion, BT relies on EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and CT merges supplemental data while materialized views accelerate BI aggregation. The combined use of Flink, DLF Paimon, EMR Serverless Spark, and StarRocks demonstrates a sustainable near‑real‑time data‑warehouse architecture for large‑scale advertising workloads.

Architecture Overview
Architecture Overview
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

FlinkStarRocksDSPEMR ServerlessDLF PaimonReal‑time Data Warehouse
Alibaba Cloud Big Data AI Platform
Written by

Alibaba Cloud Big Data AI Platform

The Alibaba Cloud Big Data AI Platform builds on Alibaba’s leading cloud infrastructure, big‑data and AI engineering capabilities, scenario algorithms, and extensive industry experience to offer enterprises and developers a one‑stop, cloud‑native big‑data and AI capability suite. It boosts AI development efficiency, enables large‑scale AI deployment across industries, and drives business value.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.