Tagged articles

StarRocks

262 articles · Page 1 of 3
Lakehouse Research Base
Lakehouse Research Base
Sep 6, 2026 · Big Data

StarRocks Query Acceleration on Paimon: Production Tuning & Best Practices

This article details production practices for accelerating StarRocks queries on Paimon external tables, covering architecture, table design (partitioning, bucketing, compaction), dirty data handling, catalog configuration, predicate pushdown verification, SQL optimization with partition/bucket pruning, query method selection (direct, async materialized views, hot/cold tiering), and BE-level tuning.

Bucket PruningLakehouseMaterialized View
0 likes · 22 min read
StarRocks Query Acceleration on Paimon: Production Tuning & Best Practices
StarRocks
StarRocks
Sep 3, 2026 · Databases

Querying Paimon Semi-Structured Data with StarRocks: Variant, Shredding & SQL Practices

This article explains how StarRocks queries Paimon Variant data, covering the differences between regular Parquet JSON and Variant, the three-phase read architecture, Shredding optimization for hot paths, practical SQL examples using get_variant_* functions, and a decision framework for choosing between JSON, Plain Variant, Shredded Variant, and formal typed columns.

LakehousePaimonParquet
0 likes · 21 min read
Querying Paimon Semi-Structured Data with StarRocks: Variant, Shredding & SQL Practices
StarRocks
StarRocks
Sep 2, 2026 · Databases

Why AI-Generated SQL Isn't Enough: Semantic Views Bridge the Business Logic Gap

This article explains why syntactically correct AI-generated SQL often fails to reflect business logic, introduces StarRocks' Semantic View as an executable 'information contract' that defines metrics, relationships, and synonyms, and demonstrates with a fishing gear example how it enables deterministic SQL expansion for both AI agents and BI tools.

AI-generated SQLData ModelingInformation Contract
0 likes · 15 min read
Why AI-Generated SQL Isn't Enough: Semantic Views Bridge the Business Logic Gap
Lakehouse Research Base
Lakehouse Research Base
Aug 25, 2026 · Big Data

StarRocks on Paimon: Morning Fast, Daytime Slow - Root Cause & Layered Optimization

This article analyzes why identical StarRocks queries on Paimon external tables run fast in early morning but slow dramatically during daytime peaks, identifying five layered root causes from file fragmentation to storage pressure, and provides a systematic troubleshooting methodology plus a four-layer optimization framework covering source governance, compute caching, storage scaling, and business scheduling.

Delete VectorLakehousePaimon
0 likes · 21 min read
StarRocks on Paimon: Morning Fast, Daytime Slow - Root Cause & Layered Optimization
Lakehouse Research Base
Lakehouse Research Base
Aug 19, 2026 · Big Data

How 1565PB of High-Quality Datasets Are Reshaping Lakehouse Architecture for AI

China's high-quality dataset count hit 120,000 totaling 1,565 PB with 60% quarterly growth, exposing four systemic gaps in traditional BI-oriented lakehouse architectures — storage, governance, compute performance, and security — and driving an AI-native reference architecture built on Paimon, StarRocks, tiered storage, operator-level lineage, and multi-modal retrieval.

AI training dataApache PaimonData Quality
0 likes · 43 min read
How 1565PB of High-Quality Datasets Are Reshaping Lakehouse Architecture for AI
Lakehouse Research Base
Lakehouse Research Base
Aug 17, 2026 · Big Data

StarRocks Real-Time Wide Tables: Direct Write vs Paimon Lakehouse Architecture Comparison

This article deeply compares two StarRocks real-time wide table architectures—direct primary key table writes versus Paimon lakehouse external tables—across seven dimensions including query performance, write latency, data reuse, resource isolation, operational complexity, storage cost, and evolution potential, providing banking scenario selection criteria and a recommended Flink+Fluss+Paimon+StarRocks streaming lakehouse architecture.

FlinkLakehousePaimon
0 likes · 21 min read
StarRocks Real-Time Wide Tables: Direct Write vs Paimon Lakehouse Architecture Comparison
Ray's Galactic Tech
Ray's Galactic Tech
Aug 7, 2026 · Databases

OLAP Engine Selection for Billion‑Row Analytics: Doris, ClickHouse, StarRocks

This guide compares Doris, ClickHouse, and StarRocks for billion‑row real‑time analytics, outlining their strengths and weaknesses across write throughput, update handling, join performance, concurrency, cost, and operational complexity, and provides concrete modeling, ingestion, query governance, and fault‑tolerance recommendations to help teams choose the most suitable OLAP platform.

ClickHouseDorisFlink
0 likes · 32 min read
OLAP Engine Selection for Billion‑Row Analytics: Doris, ClickHouse, StarRocks
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Aug 1, 2026 · Big Data

Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks

The DSP advertising data pipeline was re‑architected by splitting three distinct links—NRR, BT, and CT—and assigning DLF Paimon for massive low‑frequency data, EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and unified service‑layer merging, achieving 2‑minute freshness, sub‑5 ms point queries, ~60% storage cost reduction, and improved fault isolation.

DLF PaimonDSPEMR Serverless
0 likes · 12 min read
Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 29, 2026 · Big Data

Exploring EMR Serverless StarRocks AI Functions: Multimodal Embedding, Semantic Aggregation, and Mixed Retrieval

The article analyzes the newly released AI Function suite in Alibaba Cloud EMR Serverless StarRocks, detailing multimodal embedding, AI‑driven aggregation, semantic filtering, mixed vector‑full‑text search, architectural advantages such as SQL‑native execution, async pipelines, bounded resources, and real‑world use cases in advertising, gaming, and finance.

AI FunctionSQLStarRocks
0 likes · 16 min read
Exploring EMR Serverless StarRocks AI Functions: Multimodal Embedding, Semantic Aggregation, and Mixed Retrieval
StarRocks
StarRocks
Jul 15, 2026 · Big Data

Building a Lake‑Stream Integrated Data Pipeline with Fluss, Paimon, and StarRocks

The article details how Taotian Group’s senior data engineer Zhu Ao designed a lake‑stream integrated architecture using Fluss for second‑level real‑time storage, Paimon for minute‑level lakehouse persistence, and StarRocks as a unified OLAP query layer, achieving over 50% faster development and more than 80% cost reduction.

Data IntegrationFlussLakehouse
0 likes · 14 min read
Building a Lake‑Stream Integrated Data Pipeline with Fluss, Paimon, and StarRocks
DataFunTalk
DataFunTalk
Jul 6, 2026 · Artificial Intelligence

How to Overcome Metric Definitions, Real‑Time Data, Knowledge, and Permission Challenges in Enterprise Data Agents

The talk presents a three‑layer architecture for enterprise Data Agents, explains how StarRocks‑based AI‑native real‑time data foundations, Semantic View, and context layers work together, and outlines practical checks and feedback loops to address metric, permission, latency, and governance bottlenecks.

AI Native Data PlatformData AgentEnterprise AI
0 likes · 8 min read
How to Overcome Metric Definitions, Real‑Time Data, Knowledge, and Permission Challenges in Enterprise Data Agents
StarRocks
StarRocks
Jun 25, 2026 · Databases

StarRocks 4.1 Enables Faster Iceberg Queries While Preserving Data Freshness

StarRocks 4.1 introduces an incremental materialized view for Apache Iceberg that ties refresh cost to data changes instead of table size, dramatically cutting refresh time, maintaining low latency, and keeping query results fresh even as tables scale to terabytes or petabytes, with a fallback to partition refresh when needed.

Apache IcebergData FreshnessIncremental Materialized View
0 likes · 8 min read
StarRocks 4.1 Enables Faster Iceberg Queries While Preserving Data Freshness
StarRocks
StarRocks
Jun 17, 2026 · Databases

How StarRocks 4.1 Simplifies Operations and Boosts Production Performance

StarRocks 4.1 introduces automatic multi‑tenant data management, large‑capacity tablets, second‑level schema evolution, enhanced cache observability, and deeper Iceberg support, addressing static data distribution, data skew, high repair costs and expertise requirements while delivering up to 1.86× higher throughput and dramatically lower latency in production workloads.

Cache ObservabilityData DistributionFast Schema Evolution
0 likes · 13 min read
How StarRocks 4.1 Simplifies Operations and Boosts Production Performance
StarRocks
StarRocks
Jun 12, 2026 · Big Data

Building a Millisecond-Responsive Real-Time Data Engine with StarRocks, Fluss, and Paimon

This article presents a lake‑stream integrated solution that combines Apache Fluss, Apache Paimon, and StarRocks to achieve second‑level data freshness, tenfold storage cost reduction, and a single‑query access pattern for both real‑time and historical data, detailing its architecture, advantages, query modes, and future roadmap.

FlussLakehousePaimon
0 likes · 13 min read
Building a Millisecond-Responsive Real-Time Data Engine with StarRocks, Fluss, and Paimon
StarRocks
StarRocks
Jun 4, 2026 · Databases

How StarRocks and Iceberg Enable Federated Queries: A Practical Walkthrough

This article details Fresha's real‑world integration of StarRocks with Apache Iceberg, covering metadata planning, distributed execution, adaptive metadata retrieval, hot‑cold data layering, missing statistics handling, catalog configuration, and performance optimizations that together demonstrate how federated queries can be efficiently executed over data‑lake tables.

Apache IcebergData LakeFederated Query
0 likes · 14 min read
How StarRocks and Iceberg Enable Federated Queries: A Practical Walkthrough
StarRocks
StarRocks
May 28, 2026 · Industry Insights

How Fresha Built a Modern Real‑Time Analytics Stack with AutoMQ and StarRocks

Fresha replaced its Postgres‑Snowflake‑MSK pipeline with an AutoMQ‑based Diskless Kafka message layer and StarRocks for real‑time analytics, cutting storage costs 17‑20×, dropping query latency from seconds to sub‑second, and migrating ~1,000 topics in a week with zero downtime.

AutoMQCost OptimizationData Pipeline
0 likes · 24 min read
How Fresha Built a Modern Real‑Time Analytics Stack with AutoMQ and StarRocks
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
May 23, 2026 · Cloud Computing

Best Practice: Using EMR Serverless StarRocks AI Function for Financial Text Classification

This article demonstrates how to leverage StarRocks AI Function on EMR Serverless to perform sentiment analysis, intelligent classification, information extraction, and PII redaction on financial text entirely within SQL, eliminating data export, reducing latency, and ensuring compliance while providing concrete code examples, performance benchmarks, and best‑practice recommendations.

AI FunctionEMR ServerlessFinancial NLP
0 likes · 25 min read
Best Practice: Using EMR Serverless StarRocks AI Function for Financial Text Classification
StarRocks
StarRocks
May 21, 2026 · Databases

Say Goodbye to Repeated Pitfalls with Our Open‑Source AI Skill for Database Troubleshooting

The article introduces starrocks‑debug‑skills, an open‑source, three‑layer knowledge base (Skills, Cases, Tools) that captures real‑world StarRocks troubleshooting experience, shows how AI assistants can use it to diagnose issues such as import timeouts, version errors, and compaction slowdowns, and explains how to contribute new cases.

AIDatabase TroubleshootingStarRocks
0 likes · 13 min read
Say Goodbye to Repeated Pitfalls with Our Open‑Source AI Skill for Database Troubleshooting
StarRocks
StarRocks
May 20, 2026 · Big Data

How StarRocks, Paimon, and Fluss Enable Multimodal Fusion Search in a Lakehouse

The Streaming Lakehouse Meetup (May 27) explores breaking data silos by unifying structured tables, images, video, audio, and high‑dimensional vectors through StarRocks‑Paimon‑Fluss integration, covering multimodal fusion retrieval, vector search internals, native reader/writer performance gains, and real‑world ANN indexing practices.

FlussLakehouseMultimodal
0 likes · 5 min read
How StarRocks, Paimon, and Fluss Enable Multimodal Fusion Search in a Lakehouse
StarRocks
StarRocks
Apr 16, 2026 · Databases

Why Traditional Databases Stall AI Agents—and How StarRocks Overcomes the Bottleneck

Traditional databases were built for low‑frequency, human‑driven queries, but AI agents generate dozens of concurrent, sub‑second queries that expose architectural limits, and StarRocks addresses these challenges with self‑healing optimization, real‑time data pipelines, extreme concurrency handling, and seamless lakehouse access.

Database ConcurrencyLakehouseQuery Optimization
0 likes · 13 min read
Why Traditional Databases Stall AI Agents—and How StarRocks Overcomes the Bottleneck
vivo Internet Technology
vivo Internet Technology
Mar 25, 2026 · Industry Insights

How Vivo Scaled Marketing Automation with Presto, Bitmap, and StarRocks

This case study details how Vivo’s marketing automation platform evolved its data‑driven architecture—from a Presto‑based wide‑table design, through a Bitmap optimization, to a StarRocks migration—addressing performance bottlenecks, reducing resource costs, and enhancing data security.

BitmapOLAPPerformance Optimization
0 likes · 11 min read
How Vivo Scaled Marketing Automation with Presto, Bitmap, and StarRocks
StarRocks
StarRocks
Mar 11, 2026 · Databases

How StarRocks Supercharges Real‑Time Ad Funnel Monitoring and Creative Optimization

This article dissects the full advertising funnel, explains why CTR and eCPM are critical, and demonstrates how StarRocks combined with Flink can deliver minute‑level real‑time monitoring, material selection, anomaly alerts, A/B testing, and a successful migration from Druid for massive ad‑tech workloads.

AdvertisingMaterialized ViewsReal-time Analytics
0 likes · 20 min read
How StarRocks Supercharges Real‑Time Ad Funnel Monitoring and Creative Optimization
StarRocks
StarRocks
Mar 5, 2026 · Big Data

How Fanatics Scaled to PB‑Level Data with StarRocks & Apache Iceberg Lakehouse

Fanatics unified its fragmented data stack by building a StarRocks‑powered Lakehouse on Apache Iceberg, replacing Redshift, Snowflake, Athena, and Druid, which cut costs by up to 95%, delivered sub‑second dashboard queries on petabyte‑scale data, and enabled real‑time and historical analytics on a single platform.

Apache IcebergFanaticsLakehouse
0 likes · 10 min read
How Fanatics Scaled to PB‑Level Data with StarRocks & Apache Iceberg Lakehouse
StarRocks
StarRocks
Feb 11, 2026 · Big Data

How StarRocks and Apache Paimon Build a True Lakehouse Native Engine

This article details the deep integration of StarRocks with Apache Paimon, describing the unified architecture, version evolution, performance enhancements, time‑travel queries, native readers/writers, distributed planning, and future roadmap for achieving lakehouse‑native analytics at scale.

Apache PaimonData LakeLakehouse
0 likes · 10 min read
How StarRocks and Apache Paimon Build a True Lakehouse Native Engine
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Feb 4, 2026 · Big Data

How Paimon + StarRocks Power Real‑Time OLAP for Double‑11 Mega‑Sales

During Double‑11 mega‑sales, Taobao Group faced exploding OLAP query traffic, costly data sync pipelines, and slow near‑real‑time analytics, so they unified real‑time and batch data in Paimon, leveraged StarRocks for high‑performance lake queries, tuned cluster settings, and saved nearly ten‑million yuan annually while cutting refresh latency by 80%.

Data LakeOLAPPaimon
0 likes · 22 min read
How Paimon + StarRocks Power Real‑Time OLAP for Double‑11 Mega‑Sales
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Feb 2, 2026 · Big Data

How We Built a Scalable Lakehouse Architecture with StarRocks, Paimon, and Flink

This article details the evolution of a data warehouse at RenliJia from a MaxCompute‑centric setup to a modern lakehouse using StarRocks, Paimon, Flink, and Fluss, describing design goals, technical evaluations, implementation steps for offline, OLAP, and real‑time workloads, and the challenges and future plans that emerged.

FlinkLakehousePaimon
0 likes · 25 min read
How We Built a Scalable Lakehouse Architecture with StarRocks, Paimon, and Flink
Lakehouse Research Base
Lakehouse Research Base
Jan 27, 2026 · Big Data

Big Data Expert's Failure in Traditional Enterprise: Heavy Tech Stack Wastes Resources on Small Data

A big data expert from a large tech company builds a full Hadoop/Spark/CDH stack for a traditional enterprise with only hundreds of thousands of daily records, causing high maintenance costs, half-hour query delays, and eventual departure; the case underscores the importance of matching technology to business scale, with StarRocks proposed as a lightweight alternative.

CDHHadoopSpark
0 likes · 11 min read
Big Data Expert's Failure in Traditional Enterprise: Heavy Tech Stack Wastes Resources on Small Data
StarRocks
StarRocks
Jan 22, 2026 · Big Data

How Paimon + StarRocks Accelerates Double‑11 OLAP Queries by 80% Refresh Speed

This article explains how Taotian Group unified real‑time and offline data using Paimon as lake storage and StarRocks for high‑performance OLAP, eliminating costly sync pipelines, cutting refresh time by about 80%, saving nearly ten million yuan annually, and detailing the architecture, cluster safeguards, configuration tweaks, monitoring, and future roadmap for large‑scale promotional events.

OLAPPaimonReal-time Analytics
0 likes · 24 min read
How Paimon + StarRocks Accelerates Double‑11 OLAP Queries by 80% Refresh Speed
Lakehouse Research Base
Lakehouse Research Base
Jan 15, 2026 · Databases

StarRocks Temporary Partitions: Complete Guide to Viewing & Safe Deletion

This guide explains why StarRocks temporary partitions block DDL operations due to metadata consistency, locking, and atomicity constraints, and provides step-by-step commands for viewing partitions via SHOW TEMPORARY PARTITIONS and information_schema, plus safe deletion using ALTER TABLE DROP TEMPORARY PARTITION with critical precautions.

DDL operationsData ValidationDatabase Administration
0 likes · 10 min read
StarRocks Temporary Partitions: Complete Guide to Viewing & Safe Deletion
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jan 8, 2026 · Big Data

How Gaode Maps Built a Real‑Time Lakehouse for Billion‑Scale Trajectory Data

This article details Gaode Maps' end‑to‑end lakehouse solution for massive, high‑frequency trajectory data, covering the challenges of real‑time visibility, query performance, and storage cost, and explaining how a hot‑warm‑cold tiering architecture built on Apache Flink, Paimon, StarRocks, Redis and Lindorm delivers millisecond‑level queries while cutting storage expenses.

Apache FlinkApache PaimonLakehouse
0 likes · 19 min read
How Gaode Maps Built a Real‑Time Lakehouse for Billion‑Scale Trajectory Data
Lakehouse Research Base
Lakehouse Research Base
Jan 8, 2026 · Big Data

SME Data Architecture Selection: Lightweight Lakehouse Implementation Guide

This article provides a practical framework for SMEs to select and implement data platform architectures, comparing traditional warehouses, lakehouse, and multimodal data lakes across cost, ROI, and operational efficiency, with scenario-based recommendations and a lightweight lakehouse implementation guide using open-source stack Flink, Paimon, StarRocks, and MinIO.

Data LakeFlinkLakehouse
0 likes · 21 min read
SME Data Architecture Selection: Lightweight Lakehouse Implementation Guide
StarRocks
StarRocks
Jan 7, 2026 · Big Data

How Gaode Maps Built a Real‑Time Lakehouse for Billion‑Scale Trajectory Data

This article details Gaode Maps' end‑to‑end lakehouse solution for handling high‑frequency, high‑volume trajectory data, covering the challenges of real‑time visibility, multi‑scenario queries, storage cost, and data silos, and describing the layered storage architecture, performance validation, and future expansion plans.

Apache FlinkLakehouseReal-time Data
0 likes · 21 min read
How Gaode Maps Built a Real‑Time Lakehouse for Billion‑Scale Trajectory Data
Lakehouse Research Base
Lakehouse Research Base
Jan 2, 2026 · Databases

StarRocks Production Incident: Continuous Report Refresh Triggers Cluster Crash & Recovery

This article details a StarRocks cluster crash caused by users continuously refreshing reconciliation dashboards, the emergency response including query timeout reduction and compute node scaling, root cause analysis highlighting storage-compute separation benefits, and long-term preventive measures like business reporting processes, Multi-Warehouse isolation, and query governance.

Multi-WarehouseSQL optimizationStarRocks
0 likes · 17 min read
StarRocks Production Incident: Continuous Report Refresh Triggers Cluster Crash & Recovery
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Dec 30, 2025 · Big Data

How StarRocks and Apache Paimon Unite to Build a True Lakehouse Native Engine

StarRocks and Apache Paimon have been progressively integrated across multiple releases, enabling a unified lakehouse architecture that supports multi-source federated analysis, time-travel queries, native readers/writers, distributed planning, and advanced profiling, while delivering performance gains that bring Paimon query speed on par with native StarRocks tables.

Apache PaimonData IntegrationLakehouse
0 likes · 9 min read
How StarRocks and Apache Paimon Unite to Build a True Lakehouse Native Engine
StarRocks
StarRocks
Dec 25, 2025 · Big Data

How dbt, DataOps, and StarRocks Combine to Accelerate Real‑Time Data Modeling

This article explains how dbt drives automated data modeling and governance, how DataOps practices bring agility and control to data projects, and how StarRocks’ lakehouse architecture enables real‑time and batch analytics, illustrated with concrete workflows, version‑control conventions, and enterprise case studies.

Data ModelingDataOpsELT
0 likes · 14 min read
How dbt, DataOps, and StarRocks Combine to Accelerate Real‑Time Data Modeling
Lakehouse Research Base
Lakehouse Research Base
Dec 21, 2025 · Databases

StarRocks Join Performance Gap: Paimon/Iceberg vs Internal Tables - Root Causes & Fixes

This article analyzes why StarRocks multi-table joins on Paimon/Iceberg tables remain significantly slower than internal tables even after cache warming, detailing six root causes across storage format, data layout, caching, execution engine, metadata, and version management, plus four optimization strategies to narrow the gap.

Join PerformanceLakehouseOLAP
0 likes · 17 min read
StarRocks Join Performance Gap: Paimon/Iceberg vs Internal Tables - Root Causes & Fixes
StarRocks
StarRocks
Dec 18, 2025 · Databases

How Fresha Scaled Real‑Time Analytics with StarRocks: A Deep Dive into Their Hybrid Architecture

Facing Postgres overload and costly Snowflake queries, Fresha rebuilt its analytics platform by introducing StarRocks as a unified SQL entry point, combining federated lakehouse queries with high‑performance internal tables, which reduced homepage query latency to around 200 ms and achieved minute‑level data freshness across real‑time, historical, and search workloads.

Compute-Storage SeparationData PipelineHybrid Architecture
0 likes · 20 min read
How Fresha Scaled Real‑Time Analytics with StarRocks: A Deep Dive into Their Hybrid Architecture
StarRocks
StarRocks
Dec 11, 2025 · Databases

How StarRocks Redesigns Bulk Import to Cut Small Files and Boost Throughput

This article explains how StarRocks mitigates the hidden risks of massive one‑time data imports in a storage‑compute separated architecture by redesigning the write path to spill to local disk, merge centrally, and write to object storage, resulting in fewer small files, higher write throughput, and more stable query performance.

Bulk ImportS3StarRocks
0 likes · 12 min read
How StarRocks Redesigns Bulk Import to Cut Small Files and Boost Throughput
Ctrip Technology
Ctrip Technology
Nov 27, 2025 · Big Data

How Ctrip Cut Query Latency by 85% with StarRocks’ Compute‑Storage Separation

Ctrip migrated its massive User Behavior Tracking system from ClickHouse to a compute‑storage separated StarRocks cluster on Kubernetes, achieving millisecond‑level query latency, halving storage usage, reducing node count, and sustaining millions‑of‑rows‑per‑second write throughput while simplifying scaling and operations.

ClickHouseCompute-Storage SeparationKubernetes
0 likes · 15 min read
How Ctrip Cut Query Latency by 85% with StarRocks’ Compute‑Storage Separation
StarRocks
StarRocks
Nov 18, 2025 · Databases

StarRocks Beats ClickHouse, Snowflake, and Databricks in Coffee‑Shop Benchmark – Up to 10× Faster and Cheaper

A reproducible evaluation of StarRocks using the open‑source Coffee‑shop Benchmark shows that across 500 M, 1 B and 5 B row scales, StarRocks completes 17 complex join and aggregation queries 2–10× faster and with significantly lower cost than ClickHouse, Snowflake and Databricks, demonstrating superior performance and cost efficiency for analytical workloads.

Coffee-shop BenchmarkJOIN optimizationStarRocks
0 likes · 11 min read
StarRocks Beats ClickHouse, Snowflake, and Databricks in Coffee‑Shop Benchmark – Up to 10× Faster and Cheaper
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Nov 15, 2025 · Big Data

From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse

This article chronicles a ten‑year evolution of a self‑built big data platform—detailing early Hadoop clusters, successive migrations to Spark, Hive, Hudi, and StarRocks, the operational challenges encountered, and the comprehensive shift to Alibaba Cloud EMR Serverless that delivered significant cost, performance, and stability gains while outlining future intelligent‑ecosystem plans.

Data LakeEMR ServerlessSpark
0 likes · 17 min read
From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse
Lakehouse Research Base
Lakehouse Research Base
Nov 15, 2025 · Big Data

StarRocks Lakehouse: Internal Tables vs. Paimon/Iceberg – 5-10x Performance Gains

The article benchmarks StarRocks internal tables against Paimon and Iceberg lake tables using TPCH 100G, shows internal tables excel at real-time analytics while lake tables optimize cold storage cost and multi-engine sharing, and details optimization strategies (metadata, data cache, predicate pushdown) achieving 5-10x speedup, illustrated by a retail case study cutting query latency from seconds to milliseconds and storage cost by 65%.

Data CacheHot-Cold TieringLakehouse
0 likes · 9 min read
StarRocks Lakehouse: Internal Tables vs. Paimon/Iceberg – 5-10x Performance Gains
StarRocks
StarRocks
Oct 28, 2025 · Databases

How Cisco Migrated from Pinot to StarRocks and Boosted Query Performance by Up to 70%

This article details Cisco Webex's migration from a complex Pinot‑Trino OLAP stack to StarRocks, covering the challenges of the legacy system, the step‑by‑step migration process—including storage, compute, and SQL dialect transformation—and the resulting performance gains, cost reductions, and operational improvements.

OLAPPinotStarRocks
0 likes · 23 min read
How Cisco Migrated from Pinot to StarRocks and Boosted Query Performance by Up to 70%
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Oct 18, 2025 · Big Data

Alibaba Cloud EMR’s AI Evolution: Accelerating Big Data Performance

Since its 2016 launch, Alibaba Cloud EMR has transformed from a basic open‑source Hadoop service into a high‑performance, AI‑enabled big‑data platform, delivering optimized I/O, vectorized processing, and integrated AI functions such as natural‑language SQL, StarRocks and Spark enhancements, while supporting diverse industry workloads.

Cloud ComputingEMRSpark
0 likes · 9 min read
Alibaba Cloud EMR’s AI Evolution: Accelerating Big Data Performance
StarRocks
StarRocks
Oct 14, 2025 · Big Data

How Ctrip Scaled UBT Analytics by Migrating from ClickHouse to StarRocks

Ctrip's User Behavior Tracking (UBT) system, handling 30 TB of daily data, moved from ClickHouse to StarRocks' compute‑storage separated architecture, cutting average query latency from 1.4 seconds to 203 ms, halving storage, reducing nodes from 50 to 40, and boosting write throughput to 3 million rows per second.

ClickHouseFlinkKafka
0 likes · 15 min read
How Ctrip Scaled UBT Analytics by Migrating from ClickHouse to StarRocks
StarRocks
StarRocks
Sep 23, 2025 · Databases

How Zepto Scaled Real‑Time Brand Analytics with StarRocks: From Postgres MVP to Sub‑Second Queries

Zepto transformed its brand‑analytics platform from a Postgres MVP into a production‑grade, sub‑second real‑time analytics solution by adopting StarRocks, redesigning its data pipeline with Databricks, Kafka, and Flink, and choosing a storage‑compute architecture that supports massive joins and rapid insights.

Data PipelineDatabricksFlink
0 likes · 14 min read
How Zepto Scaled Real‑Time Brand Analytics with StarRocks: From Postgres MVP to Sub‑Second Queries
Lakehouse Research Base
Lakehouse Research Base
Sep 16, 2025 · Databases

StarRocks 3.3 Indexes: When to Use Bitmap, Bloom Filter, N-Gram & Full-Text Search

This guide details the four index types in StarRocks 3.3—Bitmap, Bloom Filter, N-Gram Bloom Filter, and Full-Text Inverted Index—covering their core principles, supported data types, creation syntax, ideal use cases (low/high cardinality, fuzzy/long-text queries), anti-patterns, and real-world optimization examples with performance metrics.

Bitmap IndexBloom FilterN-Gram Bloom Filter
0 likes · 27 min read
StarRocks 3.3 Indexes: When to Use Bitmap, Bloom Filter, N-Gram & Full-Text Search
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Sep 15, 2025 · Big Data

How a FinTech Firm Boosted Real‑Time Decision Making with StarRocks Data Warehouse

This case study details how Shuhe Technology, a leading fintech company, overcame data redundancy, low resource utilization, and slow reporting by adopting Alibaba Cloud EMR Serverless StarRocks for a unified, real‑time data warehouse, achieving standardized data pipelines, cost savings, and minute‑level decision latency.

Real-time Data WarehouseStarRocksbig data
0 likes · 8 min read
How a FinTech Firm Boosted Real‑Time Decision Making with StarRocks Data Warehouse
StarRocks
StarRocks
Sep 9, 2025 · Big Data

From Hadoop to StarRocks: Revamping a Government Procurement Data Platform

Facing massive data volumes, complex component dependencies, high TCO, and real‑time processing limits, the政采云 platform replaced its Hadoop stack with StarRocks’ minimalist, decoupled architecture, achieving lower costs, elastic scaling, faster queries, easier operations, and robust fault tolerance across diverse government procurement workloads.

Cost OptimizationHadoop migrationStarRocks
0 likes · 16 min read
From Hadoop to StarRocks: Revamping a Government Procurement Data Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Sep 5, 2025 · Big Data

How StarRocks + Paimon Powered Real‑Time Analytics for Alibaba’s Taobao Flash Sale

Facing minute‑level decision demands and billions of marketing events during Taobao's Flash Sale, the Ele.me data team built a real‑time lakehouse with StarRocks and Paimon, leveraging asynchronous materialized views, RoaringBitmap de‑duplication, and resource isolation to achieve sub‑second query latency, lower storage costs, and stable high‑concurrency.

LakehouseMaterialized ViewsPaimon
0 likes · 25 min read
How StarRocks + Paimon Powered Real‑Time Analytics for Alibaba’s Taobao Flash Sale
Lakehouse Research Base
Lakehouse Research Base
Sep 4, 2025 · Databases

Recovering Deleted Databases, Tables, and Partitions in StarRocks

This guide explains how to recover accidentally dropped databases, tables, and partitions in StarRocks using the RECOVER command, covering the 1-day default retention window, configuration options, limitations like TRUNCATE and FORCE drops, and practical step-by-step examples with renaming strategies.

DROP TABLERECOVER commandStarRocks
0 likes · 5 min read
Recovering Deleted Databases, Tables, and Partitions in StarRocks
StarRocks
StarRocks
Sep 2, 2025 · Big Data

How StarRocks + Paimon Powered Real‑Time Analytics for Alibaba’s Flash Sale

Faced with billions of marketing events and minute‑level decision requirements during Taobao's flash‑sale campaign, the e‑commerce data team built a real‑time lakehouse using StarRocks and Paimon, leveraged asynchronous materialized views and RoaringBitmap deduplication, and achieved sub‑second query latency, massive cost savings, and stable high‑concurrency performance.

LakehouseMaterialized ViewsPaimon
0 likes · 26 min read
How StarRocks + Paimon Powered Real‑Time Analytics for Alibaba’s Flash Sale
Lakehouse Research Base
Lakehouse Research Base
Aug 28, 2025 · Databases

Why Stream Load Beats INSERT 30x in StarRocks Compute-Storage Separation

In StarRocks compute-storage separation clusters, Stream Load achieves 20-30x faster bulk data ingestion than INSERT by leveraging batch-oriented design, direct compute-node access, single-transaction metadata handling, large-file storage writes, and concentrated resource scheduling, avoiding INSERT's per-row overheads.

Compute-Storage SeparationINSERTStarRocks
0 likes · 11 min read
Why Stream Load Beats INSERT 30x in StarRocks Compute-Storage Separation
Big Data Technology Tribe
Big Data Technology Tribe
Aug 22, 2025 · Backend Development

How StarRocks Keeps Metadata Consistent Across FE Nodes

This article explains the roles of StarRocks FE and BE nodes, details the metadata stored in FE, describes the leader‑follower‑observer architecture, and shows how BDB JE replication, journal logs, and checkpoint mechanisms ensure metadata synchronization and durability even after node failures.

BDB JEReplicationStarRocks
0 likes · 17 min read
How StarRocks Keeps Metadata Consistent Across FE Nodes
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Aug 21, 2025 · Big Data

How Hypergryph Built a High‑Performance Real‑Time Analytics Platform with StarRocks

This case study details how Hypergryph leveraged Alibaba Cloud EMR Serverless StarRocks, Flink, and Kafka to replace a ClickHouse data warehouse with a high‑performance, elastic, and easy‑to‑operate real‑time analytics platform that dramatically improved query speed, stability, operational efficiency, and cost for their gaming business.

Cloud ComputingData PipelineFlink
0 likes · 8 min read
How Hypergryph Built a High‑Performance Real‑Time Analytics Platform with StarRocks
StarRocks
StarRocks
Aug 19, 2025 · Big Data

How Joydata Scaled to 150 Billion Daily Events with StarRocks: A Data Architecture Journey

Facing daily data growth from millions to 150 billion records, Joydata‑U transformed its analytics platform through three architectural stages—Hadoop, Hadoop + Trino, and finally StarRocks—introducing resource isolation, Flat JSON acceleration, and Bitmap indexing to cut query latency by up to seven times and achieve sub‑2‑minute data freshness across BI, ad‑tech, game analytics, and CRM workloads.

Bitmap IndexFlat JSONFlink
0 likes · 12 min read
How Joydata Scaled to 150 Billion Daily Events with StarRocks: A Data Architecture Journey
StarRocks
StarRocks
Aug 6, 2025 · Databases

How Qunar Migrated to StarRocks: Architecture, Performance Gains & Best Practices

This article details Qunar's transition to StarRocks as a unified OLAP engine, covering the business background, engine evaluation, architecture redesign, observability, high‑availability strategies, query‑performance optimizations, real‑world application cases, community contributions, and future plans.

OLAPStarRocksdata platform
0 likes · 21 min read
How Qunar Migrated to StarRocks: Architecture, Performance Gains & Best Practices
StarRocks
StarRocks
Jul 23, 2025 · Big Data

How StarRocks Powers Intelligent BI with AI‑Native Lakehouse Architecture

This article explores the evolution of business intelligence toward intelligent BI, detailing traditional BI limitations, agile BI improvements, and how StarRocks' MPP lakehouse engine combined with large language models enables natural‑language analytics, real‑time performance, AI‑driven insights, and scalable enterprise deployments.

AI IntegrationIntelligent BILakehouse
0 likes · 19 min read
How StarRocks Powers Intelligent BI with AI‑Native Lakehouse Architecture
StarRocks
StarRocks
Jul 16, 2025 · Cloud Native

Build a Decoupled Storage‑Compute Data Platform with StarRocks and MinIO

This step‑by‑step tutorial shows how to deploy StarRocks and MinIO in a decoupled storage‑compute architecture using Docker Compose and Kubernetes, configure local caching, create storage volumes, load public datasets, and run SQL queries to explore the combined data.

Decoupled StorageDocker ComposeKubernetes
0 likes · 14 min read
Build a Decoupled Storage‑Compute Data Platform with StarRocks and MinIO
StarRocks
StarRocks
Jul 9, 2025 · Big Data

How Shopee Built a Near‑Real‑Time Data Warehouse with Paimon and StarRocks

Shopee combined the Paimon data lake with StarRocks and Flink to create a quasi‑real‑time warehouse, enabling fast task diagnostics and a high‑performance financial reconciliation system while dramatically reducing storage costs and latency through innovative ODS, snapshot, and branch table techniques.

FlinkPaimonReal-time Data Warehouse
0 likes · 13 min read
How Shopee Built a Near‑Real‑Time Data Warehouse with Paimon and StarRocks
StarRocks
StarRocks
Jul 1, 2025 · Big Data

How StarRocks Boosted Suixingfu’s Real‑Time Data Platform: 3× Faster Queries & 10× Faster Analytics

Suixingfu rebuilt its payment data pipeline by replacing a fragmented Lambda stack with a unified Porter CDC + StarRocks + Elasticsearch architecture, achieving three‑fold query speed, ten‑fold analytics efficiency, 20% storage reduction, and sub‑second data‑capture latency across high‑concurrency, ad‑hoc, and batch workloads.

CDCFlinkReal-time Analytics
0 likes · 14 min read
How StarRocks Boosted Suixingfu’s Real‑Time Data Platform: 3× Faster Queries & 10× Faster Analytics
StarRocks
StarRocks
Jun 26, 2025 · Databases

What’s New in StarRocks 3.5? Snapshot Backup, Bulk Load, Partition & Transaction Enhancements

StarRocks 3.5 introduces a cluster‑level Snapshot backup for fast recovery, a bulk‑load optimization that reduces small files and compaction cost, smarter partition management with time‑based merging and TTL, multi‑statement transactions with full ACID guarantees, low‑cardinality dictionary support for lake tables, and several security and performance upgrades.

ACID TransactionsData LakeLow Cardinality Dictionary
0 likes · 17 min read
What’s New in StarRocks 3.5? Snapshot Backup, Bulk Load, Partition & Transaction Enhancements
StarRocks
StarRocks
Jun 17, 2025 · Databases

How to Ace the StarRocks SRCA Certification: Key Topics and Study Strategies

This guide outlines the StarRocks SRCA certification exam, highlights essential study resources, breaks down the core topics such as architecture, data import/export, SQL optimization and performance tuning, and offers practical tips, mock‑exam details, and personal experience to help candidates succeed.

Database CertificationSQL optimizationSRCA
0 likes · 9 min read
How to Ace the StarRocks SRCA Certification: Key Topics and Study Strategies
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
May 21, 2025 · Big Data

How Alibaba’s A+ Traffic Analysis Achieved Sub‑Second Log Queries with StarRocks & Paimon

This article details how Alibaba's A+ traffic analysis platform tackled trillion‑row log ingestion and high‑concurrency queries by redesigning storage with Paimon, leveraging Flink for real‑time ingestion, and using StarRocks for fast lake analytics, ultimately reducing query latency from minutes to seconds.

FlinkLog AnalyticsPaimon
0 likes · 15 min read
How Alibaba’s A+ Traffic Analysis Achieved Sub‑Second Log Queries with StarRocks & Paimon
StarRocks
StarRocks
May 13, 2025 · Artificial Intelligence

How StarRocks MCP Server Enables LLMs to Query Databases Without Custom Plugins

StarRocks MCP Server provides a universal adapter that lets large language models like Claude, OpenAI, and Gemini execute SQL queries directly against StarRocks, simplifying data Q&A, intelligent analysis, and automated reporting by eliminating the need for bespoke plugins or complex prompt engineering.

AI agentsData AnalyticsLLM
0 likes · 14 min read
How StarRocks MCP Server Enables LLMs to Query Databases Without Custom Plugins
Lakehouse Research Base
Lakehouse Research Base
May 6, 2025 · Databases

StarRocks THRIFT_EAGAIN Timeout: Root Cause Analysis & Fix

This article details troubleshooting StarRocks write failures caused by THRIFT_EAGAIN timeouts, analyzing JVM memory patterns, Thrift RPC behavior, TRUNCATE TABLE resource overhead, and resolving the issue by increasing thrift_rpc_timeout_ms from 5s to 30s while planning long-term architectural improvements.

JVM memoryStarRocksTHRIFT_EAGAIN
0 likes · 13 min read
StarRocks THRIFT_EAGAIN Timeout: Root Cause Analysis & Fix
Lakehouse Research Base
Lakehouse Research Base
May 5, 2025 · Big Data

DataOps Best Practices on Alibaba Cloud: Architecture, CI/CD, Quality & Self-Healing

This article details a comprehensive DataOps implementation using Alibaba Cloud's DataWorks, Flink, StarRocks, MaxCompute, and Yunxiao pipeline, covering four-layer architecture, standardized development, automated CI/CD with gray releases, full-chain quality monitoring achieving 80% self-healing, cross-team collaboration, cost optimization, and a phased rollout roadmap delivering 30% faster delivery and 30% cost reduction.

Alibaba CloudCI/CDData Quality
0 likes · 18 min read
DataOps Best Practices on Alibaba Cloud: Architecture, CI/CD, Quality & Self-Healing
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Apr 27, 2025 · Big Data

Scaling Property Services: StarRocks‑Powered Storage‑Compute Separation for 8000+ Communities

Facing a flood of data from over 8,000 communities, the Bifeng service team migrated from a monolithic storage‑compute architecture to a StarRocks‑based storage‑compute separation solution, achieving lower costs, higher resource utilization, faster queries, and improved SLA across their property management platform.

Infrastructure MigrationOLAPStarRocks
0 likes · 11 min read
Scaling Property Services: StarRocks‑Powered Storage‑Compute Separation for 8000+ Communities
StarRocks
StarRocks
Apr 24, 2025 · Databases

Inside StarRocks Optimizer: Architecture, Multi‑Stage Optimization, and Advanced Features

This article provides a comprehensive technical overview of StarRocks' query optimizer, covering its evolution, core architecture, multi‑stage optimization pipeline, key optimizations such as multi‑join colocate, low‑cardinality global dictionary, MV union rewrite, and advanced mechanisms like cost‑estimation fixes, query feedback, adaptive execution, runtime filters, join‑reorder strategies, and SQL plan management.

Cost-Based OptimizationMaterialized ViewsOLAP
0 likes · 26 min read
Inside StarRocks Optimizer: Architecture, Multi‑Stage Optimization, and Advanced Features
Big Data Technology & Architecture
Big Data Technology & Architecture
Apr 22, 2025 · Artificial Intelligence

Introduction to Retrieval‑Augmented Generation (RAG) and Vector Indexing with StarRocks and DeepSeek

This article explains the fundamentals of Retrieval‑Augmented Generation, demonstrates how to create and query vector indexes using StarRocks, shows how DeepSeek provides embeddings and answer generation, and walks through a complete end‑to‑end RAG pipeline with code examples and a web UI.

AIDeepSeekPython
0 likes · 20 min read
Introduction to Retrieval‑Augmented Generation (RAG) and Vector Indexing with StarRocks and DeepSeek
Lakehouse Research Base
Lakehouse Research Base
Apr 6, 2025 · Databases

Accelerating StarRocks Node Decommissioning via tablet_sched_slot_num_per_path Tuning

The author decommissioned three StarRocks nodes with over 1M tablets each; default replica migration moved only 200k replicas in 24 hours, so they increased tablet_sched_slot_num_per_path from 2 to 32 per disk during off-peak, achieving 800k+ replicas migrated per node in 8 hours while monitoring disk I/O.

Database AdministrationStarRocksnode decommissioning
0 likes · 4 min read
Accelerating StarRocks Node Decommissioning via tablet_sched_slot_num_per_path Tuning
StarRocks
StarRocks
Mar 27, 2025 · Databases

How JD Logistics Boosted Query Speed and Cut Costs with StarRocks Storage‑Compute Separation

JD Logistics transformed its one‑stop self‑service analytics platform, UData, by migrating from an integrated storage‑compute architecture to a storage‑compute separated design powered by StarRocks, achieving sub‑10‑second P95/P99 query latency, reducing storage costs by 90%, and cutting compute expenses around 30% while supporting massive data volumes.

KubernetesPerformance OptimizationStarRocks
0 likes · 20 min read
How JD Logistics Boosted Query Speed and Cut Costs with StarRocks Storage‑Compute Separation
vivo Internet Technology
vivo Internet Technology
Mar 26, 2025 · Big Data

Reading Encrypted ORC Files in StarRocks: Architecture and Implementation Details

The article details how StarRocks extends the Apache ORC C++ library to decrypt column‑level encrypted ORC files, describing the file hierarchy, AES‑128‑CTR key handling, the query‑time master‑key retrieval, a decorator‑based decryption/decompression pipeline, and the block‑skip‑read mechanism that enables efficient predicate push‑down.

ORCStarRocksbig data
0 likes · 19 min read
Reading Encrypted ORC Files in StarRocks: Architecture and Implementation Details
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Mar 20, 2025 · Big Data

How to Read and Write StarRocks Data with EMR Serverless Spark

This step‑by‑step guide explains how to use EMR Serverless Spark together with the StarRocks Spark Connector to create a workspace, upload the connector JAR, configure network connections, create databases and tables in StarRocks, and perform read/write operations via SQL sessions, Notebook sessions, or batch Spark jobs, complete with code examples and UI screenshots.

Data IntegrationEMR ServerlessSpark
0 likes · 14 min read
How to Read and Write StarRocks Data with EMR Serverless Spark
StarRocks
StarRocks
Feb 27, 2025 · Big Data

How iQIYI Boosted Ad Query Performance 400% with StarRocks – A Deep Dive into OLAP Evolution

This article details iQIYI's transition from Impala+Kudu and ClickHouse to StarRocks, describing the OLAP architecture, performance gains of up to 400% in advertising workloads, the technical challenges of data consistency, lake‑warehouse fusion, operational scaling, and the step‑by‑step migration process using a dual‑run platform.

ClickHouseFlinkOLAP
0 likes · 15 min read
How iQIYI Boosted Ad Query Performance 400% with StarRocks – A Deep Dive into OLAP Evolution