Tagged articles

data engineering

321 articles · Page 3 of 4
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 24, 2021 · Big Data

Big Data Industry Trends and Career Advice for Data Developers

The article analyzes recent Q3 financial reports of major internet companies, discusses the uneven development of data engineering talent, examines the challenges of data platforms and middle‑office services, and offers practical advice for developers to broaden technical depth, improve soft skills, and increase resilience in a tightening market.

advertising revenuecareer advicedata engineering
0 likes · 11 min read
Big Data Industry Trends and Career Advice for Data Developers
DataFunTalk
DataFunTalk
Nov 24, 2021 · Big Data

Tencent Game Big Data Analysis Engine: Architecture, Practices, and Future Plans

This article presents Tencent's game big‑data analysis platform, detailing its background, the architecture of the iData engine—including offline multi‑dimensional analysis (TGMars), online portrait analysis (TGFace), and real‑time multi‑dimensional analysis (TGDruid)—application scenarios, performance insights, and future ecosystem and open‑source plans.

Game AnalyticsOLAPReal-time Analytics
0 likes · 15 min read
Tencent Game Big Data Analysis Engine: Architecture, Practices, and Future Plans
dbaplus Community
dbaplus Community
Nov 21, 2021 · Big Data

How Small Companies Can Break Into Big Data Projects and Master High‑Concurrency Architecture

This article explores why small and medium enterprises struggle with big‑data adoption, proposes partnership‑based strategies to gain access to large datasets, and offers concrete technical roadmaps—including distributed storage, streaming pipelines, and query stacks—to help engineers practice high‑concurrency big‑data systems.

SME Strategydata engineeringhigh concurrency
0 likes · 9 min read
How Small Companies Can Break Into Big Data Projects and Master High‑Concurrency Architecture
DataFunTalk
DataFunTalk
Nov 20, 2021 · Big Data

How to Build a Big Data Platform from Zero to One: Architecture, Components, and Best Practices

This article provides a comprehensive guide to designing and implementing a big‑data platform, covering architecture overview, data ingestion with Flume, storage on HDFS/Hive/HBase, processing engines such as Hive, Spark and Flink, scheduling solutions like Azkaban and Airflow, and the construction of self‑service analytics systems.

ETLHadoopSelf‑service analytics
0 likes · 29 min read
How to Build a Big Data Platform from Zero to One: Architecture, Components, and Best Practices
21CTO
21CTO
Nov 1, 2021 · Big Data

Essential Data Engineering Roadmap: Skills, Tools, and Technologies to Master

This guide outlines the fast‑growing data engineering career path, covering essential Linux fundamentals, programming languages, testing, database concepts, data warehouses, processing frameworks, messaging systems, cluster computing, workflow scheduling, monitoring, infrastructure as code, and CI/CD tools.

big datadata engineeringdata pipelines
0 likes · 5 min read
Essential Data Engineering Roadmap: Skills, Tools, and Technologies to Master
Big Data Technology & Architecture
Big Data Technology & Architecture
Oct 14, 2021 · Big Data

Overview of Big Data Architecture Trends and Curated Resources

This article, discovered on the Yunqi community site, provides a system‑architecture perspective overview of current big‑data architecture hotspots, development trajectories, emerging trends, and unresolved challenges, while highlighting the field’s rapid evolution and recommending a curated list of in‑depth resources for further study.

Resourcesdata architecturedata engineering
0 likes · 5 min read
Overview of Big Data Architecture Trends and Curated Resources
Airbnb Technology Team
Airbnb Technology Team
Sep 27, 2021 · Big Data

Midas Certification: Airbnb’s End-to-End Data Quality Framework

Airbnb’s Midas certification establishes a company‑wide, multi‑dimensional golden‑standard for data quality—covering accuracy, consistency, timeliness, cost, and completeness—by requiring collaborative design, automated health checks, and four review stages, ensuring certified data is reliable, well‑documented, and ready for reporting, experimentation, and machine‑learning.

AirbnbData QualityMidas Certification
0 likes · 12 min read
Midas Certification: Airbnb’s End-to-End Data Quality Framework
DataFunTalk
DataFunTalk
Sep 10, 2021 · Big Data

Presto High‑Performance Engine Practice at Meitu: Technical Selection, HA Design, and Cross‑Cluster Scheduling

This article details Meitu's adoption of the Presto ad‑hoc ROLAP engine, comparing it with Hive on Spark and Impala, describing enhancements for coordinator high‑availability, and explaining a cross‑cluster scheduling strategy that leverages idle Presto resources to improve overall big‑data workload efficiency.

Cross-Cluster SchedulingHAHigh Performance
0 likes · 16 min read
Presto High‑Performance Engine Practice at Meitu: Technical Selection, HA Design, and Cross‑Cluster Scheduling
ByteDance ADFE Team
ByteDance ADFE Team
Aug 31, 2021 · Big Data

Evolution of the Big Data Technology Stack Over the Past Five Years

This article reviews the evolution of big data technologies in the last five years, covering streaming and batch processing frameworks, column‑store NoSQL databases, programming language trends, the cloud‑native multi‑model database Lindorm, and practical Flink/Blink usage with code examples.

FlinkLindormSQL
0 likes · 24 min read
Evolution of the Big Data Technology Stack Over the Past Five Years
Volcano Engine Developer Services
Volcano Engine Developer Services
Aug 3, 2021 · Big Data

Inside ByteDance’s Traffic Platform: Powering Trillions of Real‑Time Events

This article, compiled from a Volcano Engine meetup, explains how ByteDance’s unified traffic platform designs, governs, and processes massive event‑tracking data in real time, covering embedding content solutions, link architecture, dynamic processing engines, and data‑governance practices that support trillions of daily events.

big datadata engineeringdata governance
0 likes · 16 min read
Inside ByteDance’s Traffic Platform: Powering Trillions of Real‑Time Events
Airbnb Technology Team
Airbnb Technology Team
Jul 29, 2021 · Big Data

Airbnb’s Data Quality Improvement Plan: Organizational, Architectural, and Governance Practices

Airbnb’s 2019 Data Quality Improvement Plan reorganized its data‑engineering workforce, introduced a dedicated data‑engineer role, adopted a decentralized Minerva‑based architecture with Spark pipelines, instituted rigorous testing, governance, and certification processes, and established SLAs and monitoring to ensure timely, trustworthy, well‑documented data across the enterprise.

AirbnbData Qualitybig data
0 likes · 13 min read
Airbnb’s Data Quality Improvement Plan: Organizational, Architectural, and Governance Practices
TAL Education Technology
TAL Education Technology
Jul 1, 2021 · Big Data

Optimization of A/B Test Metric Computation Using Spark and ClickHouse

This article details the design and multi‑stage optimization of an A/B testing metric system, describing its product architecture, Spark‑based computation engine, ClickHouse OLAP layer, cumulative calculation improvements, and batch processing techniques that reduced processing time from hours to a few minutes for hundreds of experiments and metrics.

A/B testingClickHouseSpark
0 likes · 8 min read
Optimization of A/B Test Metric Computation Using Spark and ClickHouse
Zhongtong Tech
Zhongtong Tech
May 31, 2021 · Big Data

How Zhongtong Express Built a Robust Big Data Quality Assurance System

At the 2021 QECon conference in Shenzhen, Zhongtong Express senior architect Wu Da detailed the design and evolution of their big data quality assurance framework, covering six key layers and highlighting future trends in predictive analytics and deep business integration.

data engineeringquality assurance
0 likes · 4 min read
How Zhongtong Express Built a Robust Big Data Quality Assurance System
ITFLY8 Architecture Home
ITFLY8 Architecture Home
May 26, 2021 · Databases

How to Store Billions of IDs in Redis Without Running Out of Memory

This article examines the challenges of storing massive DMP ID mappings in Redis—including memory fragmentation, expansion, and latency constraints—and presents eviction, bucket‑hashing, and fragmentation‑reduction techniques to achieve efficient, real‑time, large‑scale key‑value storage.

Key-value hashingMemory OptimizationRedis
0 likes · 11 min read
How to Store Billions of IDs in Redis Without Running Out of Memory
DeWu Technology
DeWu Technology
May 22, 2021 · Big Data

Unified Semantic Layer for Data Development: Addressing Pain Points and Optimizing Queries

A unified semantic layer for data development solves metric‑change ripple effects, developer burden, and large‑scale query performance problems by offering consistent metric definitions, multi‑view access, concise auto‑generated SQL, instant propagation of updates, and engine‑driven optimal query selection, thereby bridging business and engineering and cutting maintenance effort.

OLAPQuery OptimizationSemantic Layer
0 likes · 5 min read
Unified Semantic Layer for Data Development: Addressing Pain Points and Optimizing Queries
Tencent Cloud Developer
Tencent Cloud Developer
May 18, 2021 · Big Data

Latest ClickHouse Technologies and Practical Applications

ClickHouse, born from Yandex’s Metrica and now a top‑50 open‑source analytics engine, achieves exceptional speed through a vectorized compute engine, column‑store architecture, and an active community, powering real‑time workloads at companies like Tencent Music, Sina, Bilibili, and Suning while introducing features such as column merging, projections, and storage‑compute separation for future scalability.

ClickHouseColumnar DatabaseOLAP
0 likes · 17 min read
Latest ClickHouse Technologies and Practical Applications
DataFunTalk
DataFunTalk
May 11, 2021 · Big Data

Design and Practice of Baixin Bank's Flink‑Based Real‑Time Computing Platform and Hudi‑Powered Real‑Time Data Lake

This article details Baixin Bank's construction of a Flink‑driven real‑time computing platform integrated with Hudi as a real‑time data lake, covering background, architecture, data collection, transformation, storage layers, technical challenges, future roadmap, and practical lessons for similar big‑data initiatives.

FlinkHudiReal-time Data Lake
0 likes · 12 min read
Design and Practice of Baixin Bank's Flink‑Based Real‑Time Computing Platform and Hudi‑Powered Real‑Time Data Lake
Meituan Technology Team
Meituan Technology Team
Apr 15, 2021 · Big Data

Data Governance Practices at Meituan Hotel & Travel Platform

Meituan’s hotel‑travel platform tackled exploding data‑quality, cost, efficiency, and security issues by establishing a full‑link governance framework—standardized processes, a Data Management Committee, and unified “One Model, One Logic, One Service, One Portal” systems—that cut per‑unit costs by ~40%, boosted engineer productivity over 60%, eliminated major security incidents, and set the stage for autonomous, AI‑driven data governance.

Data QualityMeituanbig data
0 likes · 32 min read
Data Governance Practices at Meituan Hotel & Travel Platform
iQIYI Technical Product Team
iQIYI Technical Product Team
Apr 9, 2021 · Big Data

Real-Time Data Warehouse at iQIYI Video Production Using Spark and ClickHouse

To meet iQIYI video production’s thousands‑QPS, petabyte‑scale, frequently‑updated data and large‑table join requirements, the team built a Spark‑plus‑ClickHouse real‑time warehouse that streams Kafka changes, joins HBase dimensions, and writes to ClickHouse, reducing reporting development time from days to hours while supporting both offline and real‑time analytics.

ClickHouseHBaseKafka
0 likes · 12 min read
Real-Time Data Warehouse at iQIYI Video Production Using Spark and ClickHouse
DataFunTalk
DataFunTalk
Mar 27, 2021 · Big Data

Kuaishou's HDFS Architecture, Scale, Challenges, and Practices

This article presents an in‑depth technical overview of Kuaishou's massive HDFS deployment, detailing its architecture, petabyte‑scale data and thousands‑of‑node clusters, the key scalability challenges faced, and the custom solutions—including FixedOrder, RBF balancer, observer read, slow‑node mitigation, and tiered protection—implemented to keep the system performant and reliable.

HDFSKuaishouPerformance Optimization
0 likes · 12 min read
Kuaishou's HDFS Architecture, Scale, Challenges, and Practices
21CTO
21CTO
Feb 22, 2021 · Artificial Intelligence

How to Strengthen an Algorithm Engineer’s Real‑World Impact: Tech, Business, and Soft Skills

The article outlines a three‑dimensional framework—technical, business, and soft‑skill competencies—that algorithm engineers need to master in order to successfully deliver machine‑learning solutions in production environments, offering practical advice on data handling, model evaluation, stakeholder communication, and personal development.

Business Analysisdata engineeringmachine learning
0 likes · 15 min read
How to Strengthen an Algorithm Engineer’s Real‑World Impact: Tech, Business, and Soft Skills
DevOps
DevOps
Feb 9, 2021 · Operations

Choosing Between DataOps, MLOps, and AIOps: A Guide for Data Teams

The article examines how data teams can select the appropriate Ops framework—DataOps, MLOps, or AIOps—by comparing their origins, principles, responsibilities, and tooling, and stresses that cultural principles outweigh technology choices for efficient delivery of data and machine‑learning products.

AIOpsDataOpsDevOps
0 likes · 12 min read
Choosing Between DataOps, MLOps, and AIOps: A Guide for Data Teams
DataFunTalk
DataFunTalk
Feb 5, 2021 · Big Data

Design and Implementation of Beike's Data Management Platform (DMP)

This article details how Beike built a comprehensive Data Management Platform (DMP) that integrates user behavior and business data across multiple apps, outlines its five‑layer architecture, discusses data collection, processing, storage, real‑time profiling, and presents performance results and future optimization directions.

DMPHiveReal-time Analytics
0 likes · 20 min read
Design and Implementation of Beike's Data Management Platform (DMP)
TAL Education Technology
TAL Education Technology
Jan 28, 2021 · Big Data

Batch-Stream Fusion in Education: TAL’s Real-Time Data Platform Practices

This article, presented by senior data platform engineer Mao Xiangyi of TAL Education, details the design and implementation of the company’s real‑time T‑Streaming platform, covering its three‑layer data architecture, batch‑stream integration techniques, ODS layer real‑timeization, Flink SQL development workflow, hybrid‑cloud deployment, and a case study of K‑12 renewal reporting.

Batch-Stream IntegrationEducation AnalyticsFlink
0 likes · 18 min read
Batch-Stream Fusion in Education: TAL’s Real-Time Data Platform Practices
Xueersi Online School Tech Team
Xueersi Online School Tech Team
Jan 15, 2021 · Artificial Intelligence

Recommendation System Architecture and Engineering Overview

This article presents a comprehensive overview of a recommendation system, covering its business background, purpose, detailed engineering architecture—including data sources, computation, storage, online learning, service and access layers—and discusses key challenges, module design, and practical reflections.

AB TestingTensorFlowdata engineering
0 likes · 14 min read
Recommendation System Architecture and Engineering Overview
DataFunTalk
DataFunTalk
Nov 28, 2020 · Artificial Intelligence

Building Fast-Iterating Machine Learning Systems at Tubi: A/B Testing, Simple Models, and Embedding Strategies

This article shares Tubi's practical experience in rapidly iterating machine‑learning systems, emphasizing the early importance of simple end‑to‑end A/B testing platforms, clear launch plans, heat‑based and embedding‑based ranking models, and a culture of fast experimentation over complex deep‑learning research.

A/B testingArtificial Intelligencedata engineering
0 likes · 8 min read
Building Fast-Iterating Machine Learning Systems at Tubi: A/B Testing, Simple Models, and Embedding Strategies
Beike Product & Technology
Beike Product & Technology
Nov 13, 2020 · Big Data

Beike One‑Stop Big Data Development Platform: Architecture, Evolution, and Future Outlook

The article summarizes Beike's one‑stop big data development platform, describing its data business background, the evolution from a simple Hadoop‑Kafka‑Hive stack to a metadata‑driven, asset‑oriented platform, and outlines current capabilities in data management, integration, scheduling, quality, openness, and future plans.

ETLbig datadata engineering
0 likes · 11 min read
Beike One‑Stop Big Data Development Platform: Architecture, Evolution, and Future Outlook
dbaplus Community
dbaplus Community
Oct 13, 2020 · Big Data

How to Build a Real‑Time Data Warehouse with Flink: Principles, Architecture, and Best Practices

This article explains why real‑time data warehouses are needed, outlines their core principles, compares them with offline warehouses, describes typical use cases such as real‑time OLAP, dashboards, feature generation and monitoring, and provides a step‑by‑step guide to designing, implementing, and operating a Flink‑based streaming warehouse with Kafka, HBase, and metadata management.

FlinkKafkaOLAP
0 likes · 29 min read
How to Build a Real‑Time Data Warehouse with Flink: Principles, Architecture, and Best Practices
DataFunTalk
DataFunTalk
Oct 7, 2020 · Big Data

Yanxuan Data Warehouse: Architecture, Standards, and Evaluation Framework

This article outlines the Yanxuan data warehouse’s layered architecture, the offline and real‑time development platforms, the comprehensive standards for metric definition, model design, and SQL development, and proposes a six‑dimensional evaluation system covering data norms, security, quality, stability, continuous improvement, and development efficiency.

SQL Standardsbig datadata engineering
0 likes · 12 min read
Yanxuan Data Warehouse: Architecture, Standards, and Evaluation Framework
Youku Technology
Youku Technology
Sep 18, 2020 · Big Data

Digitalization of Youku Long‑Video Content Supply Chain: Practices and Architecture

Youku’s digital content‑supply‑chain system transforms long‑video production by introducing a three‑stage framework—structured evaluation of talent and scripts, information‑driven production management, and a unified demand‑aligned content strategy—that curtails delays, mitigates risk, and saves over 100 million RMB while scaling to billions of data records daily.

Artificial IntelligenceContent Supply Chainbig data
0 likes · 11 min read
Digitalization of Youku Long‑Video Content Supply Chain: Practices and Architecture
Beike Product & Technology
Beike Product & Technology
Aug 17, 2020 · Big Data

Bitmap-Based User Segmentation in a DMP Platform Using ClickHouse

This article describes how a data management platform (DMP) at Beike leverages ClickHouse bitmap structures and Spark pipelines to generate global numeric user IDs, design tag-specific bitmap rules for enum, continuous, and date attributes, handle boundary cases, and produce high‑performance bitmap SQL for real‑time user group estimation and complex segment logic.

BitmapClickHouseDMP
0 likes · 17 min read
Bitmap-Based User Segmentation in a DMP Platform Using ClickHouse
Big Data Technology Architecture
Big Data Technology Architecture
Aug 12, 2020 · Big Data

Overview of New Features and Improvements in Apache Spark 3.0

Apache Spark 3.0 introduces a suite of performance enhancements, richer APIs, improved monitoring, SQL compatibility, new data sources, and ecosystem extensions, including Adaptive Query Execution, Dynamic Partition Pruning, Join Hints, pandas UDF improvements, and accelerator‑aware scheduling, to boost scalability and ease of use for big‑data workloads.

Apache SparkPerformance OptimizationSpark 3.0
0 likes · 15 min read
Overview of New Features and Improvements in Apache Spark 3.0
Qunar Tech Salon
Qunar Tech Salon
Jul 15, 2020 · Artificial Intelligence

Qunar Technology Carnival: Interviews on Search Optimization, AIOps Fault Localization, and Revenue Management

The Qunar Technology Carnival features in‑depth interviews with experts Wang Mingyou, He Yang, and Jia Ziyan who share practical experiences on search ranking improvements, AIOps‑driven fault localization, and data‑driven revenue management, highlighting challenges, solutions, and future directions in AI‑powered systems.

AIOpsQunarRevenue Management
0 likes · 10 min read
Qunar Technology Carnival: Interviews on Search Optimization, AIOps Fault Localization, and Revenue Management
Xianyu Technology
Xianyu Technology
Jul 9, 2020 · Product Management

Xianyu Product Structuring: Evolution, Current Strategies, and Future Directions

Xianyu’s product‑information structuring has progressed from simple text mining to multimodal AI pipelines that now boost coverage by nearly 50 %, while facing precision and engineering hurdles, and it plans to adopt a standardized VID attribute system, plug‑in multimodal models, and rule‑based input assistance to enable seamless, photo‑driven publishing.

Rule Enginedata engineeringe-commerce
0 likes · 10 min read
Xianyu Product Structuring: Evolution, Current Strategies, and Future Directions
Big Data Technology Architecture
Big Data Technology Architecture
Jun 29, 2020 · Big Data

Real‑time Data Warehouse Construction: Goals, Architecture, and Best Practices with Apache Flink

This article summarizes the objectives, design principles, application scenarios, layer‑by‑layer construction methods, quality assurance mechanisms, and supporting tools for building a real‑time data warehouse using Apache Flink, providing practical guidance for data engineers and architects.

Apache FlinkData QualityFlink
0 likes · 24 min read
Real‑time Data Warehouse Construction: Goals, Architecture, and Best Practices with Apache Flink
DataFunTalk
DataFunTalk
Jun 14, 2020 · Big Data

Designing an Offline Big Data Processing Architecture Based on Object Storage

This article presents a comprehensive offline big‑data processing framework that leverages scalable object storage for PB‑level data, details storage and compute engine requirements, compares cost options, describes data pipeline design, and showcases an e‑commerce case study with Spark‑driven analytics.

Cost OptimizationData PipelineObject Storage
0 likes · 19 min read
Designing an Offline Big Data Processing Architecture Based on Object Storage
Beike Product & Technology
Beike Product & Technology
Jun 12, 2020 · Big Data

Design and Implementation of SQL on Streaming (SQL 1.0 → SQL 2.0) in a Real‑Time Computing Platform

This article describes the evolution of a real‑time computing platform from SQL 1.0 built on Spark Structured Streaming to SQL 2.0 powered by Flink‑SQL, covering dynamic tables, continuous queries, dimension‑table joins, cache optimization, DDL extensions, platformization, operational challenges and future roadmap.

Dimension TableFlinkReal-Time Computing
0 likes · 19 min read
Design and Implementation of SQL on Streaming (SQL 1.0 → SQL 2.0) in a Real‑Time Computing Platform
iQIYI Technical Product Team
iQIYI Technical Product Team
Jun 12, 2020 · Artificial Intelligence

Deepthought: An End‑to‑End Machine Learning Platform at iQIYI

Deepthought is iQIYI’s end‑to‑end machine‑learning platform that unifies distributed frameworks, decouples pipeline stages, integrates with Tongtian Tower, and offers visual drag‑and‑drop configuration, evolving from a fraud‑detection prototype to a generic system with real‑time inference, automated hyper‑parameter optimization, and support for large‑scale data across anti‑fraud, recommendation, and analytics workloads.

AI platformAutoMLSpark
0 likes · 13 min read
Deepthought: An End‑to‑End Machine Learning Platform at iQIYI
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Jun 2, 2020 · Artificial Intelligence

How to Transition into AI: Real Stories, Tools, and a Practical Roadmap

This article shares personal journeys of five Huawei AI experts, recommends essential AI books, walks through setting up PyCharm with ModelArts for hands‑on model training, and outlines a three‑stage AI career roadmap—from practical coding to mastering principles and deploying inference—offering actionable guidance for anyone looking to break into artificial intelligence.

AI DeploymentAI career transitionModelArts
0 likes · 34 min read
How to Transition into AI: Real Stories, Tools, and a Practical Roadmap
dbaplus Community
dbaplus Community
Apr 26, 2020 · Big Data

Evolving from Data Warehouses to Data Middle Platforms: Architecture & Practices

This talk reviews China's big‑data evolution from early enterprise data warehouses to modern data middle platforms, outlines core architectural components, technology selections, data development practices, lifecycle and quality management, and shares practical Q&A insights for building scalable, cost‑effective data infrastructures.

big datadata architecturedata engineering
0 likes · 28 min read
Evolving from Data Warehouses to Data Middle Platforms: Architecture & Practices
DataFunTalk
DataFunTalk
Apr 22, 2020 · Big Data

Didi's Real-Time Computing Practices with Apache Flink: Architecture, StreamSQL, and Operational Insights

Senior Didi technology expert Liang Li-yin shares how Didi leverages Apache Flink for large‑scale real‑time computing, covering service architecture, StreamSQL advantages, multi‑cluster management, task control, monitoring, meta‑store integration, challenges, and future plans such as high availability, real‑time ML, and unified batch‑stream processing.

Apache FlinkReal-Time ComputingStreamSQL
0 likes · 14 min read
Didi's Real-Time Computing Practices with Apache Flink: Architecture, StreamSQL, and Operational Insights
Qudian (formerly Qufenqi) Technology Team
Qudian (formerly Qufenqi) Technology Team
Mar 4, 2020 · Artificial Intelligence

How Intelligent Marketing Leverages AI and Big Data to Boost Conversion Rates

This article explains how intelligent marketing transforms traditional, labor‑intensive strategies into data‑driven, AI‑powered systems by detailing the multi‑layer architecture, data pipelines, machine‑learning models such as LR and GBDT+LR, and future directions like personalized copy generation and deep‑learning enhancements.

AIdata engineeringfeature engineering
0 likes · 8 min read
How Intelligent Marketing Leverages AI and Big Data to Boost Conversion Rates
Alibaba Cloud Developer
Alibaba Cloud Developer
Mar 3, 2020 · Artificial Intelligence

How Alibaba Turns AI, Deep Learning, and Big Data into Enterprise Power

Jia Yangqing’s talk from the Alibaba CIO Academy explains what artificial intelligence is, its applications, the challenges of perception and decision making, the evolution of deep‑learning models, the need for massive compute power, and how enterprises can strategically adopt AI and big‑data technologies to drive innovation.

AI platformsCloud Computingdata engineering
0 likes · 16 min read
How Alibaba Turns AI, Deep Learning, and Big Data into Enterprise Power
dbaplus Community
dbaplus Community
Jan 14, 2020 · Big Data

How OPPO Built a Real‑Time Data Warehouse with Flink SQL

This article details{32-64 words} OPPO's evolution from an offline data warehouse to a real‑time platform, describing the business scale, data‑mid platform architecture, migration strategy using Flink SQL, extensions like AthenaX, and practical use cases such as real‑time ETL, CTR calculation, and tag import.

ETLFlinkReal-time Data Warehouse
0 likes · 18 min read
How OPPO Built a Real‑Time Data Warehouse with Flink SQL
Bitu Technology
Bitu Technology
Dec 20, 2019 · Big Data

Building a Model‑Driven Data Platform at Tubi: From Data Warehouse to Automated Machine Learning

The article describes how Tubi, North America’s largest free‑streaming service, built a model‑driven data platform using a high‑quality data warehouse, DBT‑based transformations, Kubernetes‑hosted JupyterHub, low‑latency Scala/Akka services, and automated machine‑learning pipelines to accelerate experimentation and decision‑making.

data engineeringdata platformdbt
0 likes · 11 min read
Building a Model‑Driven Data Platform at Tubi: From Data Warehouse to Automated Machine Learning
Youzan Coder
Youzan Coder
Nov 20, 2019 · Big Data

Understanding Youzan's Data Middle Platform: Architecture, Challenges, and Construction

He Fei explains how Youzan built a two‑layer data middle platform—combining a technology stack of offline, online and streaming components with an asset layer for cataloguing, quality, lineage and unified APIs—to tackle diverse business demands, technical complexity, and to enable cost‑optimized, reusable real‑time data services.

data engineeringdata platform
0 likes · 15 min read
Understanding Youzan's Data Middle Platform: Architecture, Challenges, and Construction
FunTester
FunTester
Oct 22, 2019 · Backend Development

How to Scrape 7.2 Million Historical Weather Records with Groovy

This article explains how to use a Groovy script to crawl over 7 million historical weather entries for 3,200 cities spanning 2011‑2019, process the JSON responses, and store the cleaned data into a MySQL table, while sharing practical tips and code snippets.

GroovyJavaMySQL
0 likes · 7 min read
How to Scrape 7.2 Million Historical Weather Records with Groovy
iQIYI Technical Product Team
iQIYI Technical Product Team
Oct 11, 2019 · Artificial Intelligence

Insights into iQIYI's Recommendation Platform Architecture and Practices

iQIYI’s recommendation middle‑platform consolidates content, behavior, and machine‑learning services into a modular architecture that lets any front‑end business connect to a unified recommendation engine, cutting integration time from weeks to days, boosting development efficiency by over 30 % while simplifying maintenance and future upgrades.

AIPlatform ArchitectureTech Platform
0 likes · 11 min read
Insights into iQIYI's Recommendation Platform Architecture and Practices
Big Data Technology & Architecture
Big Data Technology & Architecture
Oct 3, 2019 · Big Data

Data Development Interview Tips and Career Guidance

This article offers practical advice for data development job interviews, explaining why Java is essential, comparing Java and Python, outlining required backend framework knowledge, discussing the role of SQL and data warehousing, and addressing work‑life concerns such as overtime and company size choices.

JavaPythonSQL
0 likes · 4 min read
Data Development Interview Tips and Career Guidance
Big Data Technology & Architecture
Big Data Technology & Architecture
Aug 26, 2019 · Big Data

Comprehensive Collection of Apache Flink Learning Resources

This article compiles a curated list of the most reliable and official Apache Flink learning materials—including beginner tutorials, source‑code walkthroughs, advanced topics, community articles, real‑world case studies, and downloadable resources—providing a one‑stop reference for developers and researchers interested in stream processing and big‑data analytics.

Apache FlinkResourcesbig data
0 likes · 10 min read
Comprehensive Collection of Apache Flink Learning Resources
Meituan Technology Team
Meituan Technology Team
Aug 15, 2019 · Big Data

Inconsistent Predictions in XGBoost on Spark Due to Different Missing Value Handling

The discrepancy between XGBoost’s Java engine and Spark arose because XGBoost4j treats zero as the default missing value while Spark’s sparse vectors use NaN, causing inconsistent predictions, and was resolved by explicitly setting Float.NaN as the missing value or converting sparse vectors to dense so both engines handle zeros uniformly.

SparkSparseVectorXGBoost
0 likes · 13 min read
Inconsistent Predictions in XGBoost on Spark Due to Different Missing Value Handling
DataFunTalk
DataFunTalk
Aug 15, 2019 · Artificial Intelligence

Intelligent Customer Acquisition System Practice at Du Xiaoman Financial

This article presents a comprehensive overview of Du Xiaoman Financial's intelligent customer acquisition system, covering acquisition channels, efficiency improvements through multi‑stage models, data understanding with deepFM, the platform architecture, and related recruitment for senior machine‑learning engineers.

AICustomer AcquisitionFinancial Technology
0 likes · 9 min read
Intelligent Customer Acquisition System Practice at Du Xiaoman Financial
DataFunTalk
DataFunTalk
Aug 14, 2019 · Artificial Intelligence

Understanding Recommendation Systems: From Information Overload to Personalized AI Solutions

The article explores how the rapid growth of the internet has created information overload, discusses the challenges of recommendation systems such as sparsity and timeliness, outlines a four‑step personalized content pipeline, and highlights the interdisciplinary nature of building effective AI‑driven recommendation solutions.

AIbig datadata engineering
0 likes · 16 min read
Understanding Recommendation Systems: From Information Overload to Personalized AI Solutions
21CTO
21CTO
Aug 6, 2019 · Databases

Why SQL Is Making a Comeback: From NoSQL’s Rise to the New Data Era

This article explores the resurgence of SQL, tracing its historical roots, the rise and limitations of NoSQL, and how modern cloud and NewSQL solutions are re‑establishing SQL as the universal interface for data storage, processing, and analysis.

NewSQLNoSQLSQL
0 likes · 14 min read
Why SQL Is Making a Comeback: From NoSQL’s Rise to the New Data Era
dbaplus Community
dbaplus Community
Jul 24, 2019 · Big Data

Essential Open-Source Tools Every Big Data Engineer Should Know

This article compiles a comprehensive list of common open‑source tools for big data platforms—covering programming languages, data collection, ETL, storage, analysis, query, management, and monitoring—to help learners and practitioners quickly locate and understand the technologies they need.

ETLHadoopSpark
0 likes · 15 min read
Essential Open-Source Tools Every Big Data Engineer Should Know
Big Data Technology & Architecture
Big Data Technology & Architecture
Jun 20, 2019 · Big Data

Comprehensive Guide to Flink SQL: Background, New Features, Programming Model, Operators, Functions, and a Practical NBA Scoring Leader Example

This article provides an in‑depth overview of Flink SQL, covering its origins, the latest 1.7.0 and 1.8.0 enhancements, the underlying programming model, common operators and built‑in functions, and a complete end‑to‑end example that analyzes NBA scoring‑leader data using Flink SQL.

Apache FlinkFlink SQLSQL
0 likes · 27 min read
Comprehensive Guide to Flink SQL: Background, New Features, Programming Model, Operators, Functions, and a Practical NBA Scoring Leader Example
Architects' Tech Alliance
Architects' Tech Alliance
Apr 20, 2019 · Industry Insights

Why Data Middle Platforms Are the New Production Lines for Data Products

The article examines how data middle platforms transform raw, fragmented enterprise data into valuable data products through a supply‑chain approach, outlining their origins, core processes, deep‑processing techniques, and the essential capabilities needed for successful implementation.

Data Supply Chaindata engineeringdata platform
0 likes · 13 min read
Why Data Middle Platforms Are the New Production Lines for Data Products
Big Data Technology & Architecture
Big Data Technology & Architecture
Feb 15, 2019 · Big Data

Big Data Mastery Roadmap

This article outlines a comprehensive series of over 500 planned tutorials covering Java advanced features, distributed theory, Hadoop, Spark, Flink, and various big‑data storage and processing technologies, designed to guide engineers transitioning into big‑data development from fundamentals to expert level.

FlinkHadoopJava
0 likes · 4 min read
Big Data Mastery Roadmap
Big Data Technology & Architecture
Big Data Technology & Architecture
Feb 12, 2019 · Big Data

Big Data Mastery Roadmap – Series Overview

An extensive roadmap series titled “Big Data Mastery Roadmap” outlines essential topics—from Java advanced features and JVM internals to Hadoop, Spark, Flink, and big-data algorithms—guiding engineers transitioning to big data development with curated references, updates, and author insights.

Learning Pathdata engineeringdistributed systems
0 likes · 5 min read
Big Data Mastery Roadmap – Series Overview
Efficient Ops
Efficient Ops
Jan 3, 2019 · Operations

Building a Scalable AIOps Platform from Zero: A Guide for Small Teams

This article outlines how to design and implement a large‑scale, AI‑driven operations platform—from defining goals with the 5W‑1H method, through data collection, storage, and processing, to building the three‑horse‑power components of monitoring, alerting, and CI/CD—targeted especially at small‑to‑mid‑size enterprises.

AIOpsArtificial Intelligencedata engineering
0 likes · 17 min read
Building a Scalable AIOps Platform from Zero: A Guide for Small Teams
21CTO
21CTO
Dec 21, 2018 · Big Data

What Does a Data Engineer Do? Skills, Certifications, and Career Path

This article explains the role of a data engineer, outlines essential big‑data architecture tools, key technical skills, differences from data scientists, and offers guidance on certifications and learning paths to launch a successful data‑engineering career.

Skillscertificationsdata engineering
0 likes · 7 min read
What Does a Data Engineer Do? Skills, Certifications, and Career Path
Zhuanzhuan Tech
Zhuanzhuan Tech
Dec 21, 2018 · Big Data

Design and Implementation of a Scalable User Profiling System Using Hive and SQL Templates

To meet the growing demands of precise, cost‑effective user operations, the article outlines a lightweight, flexible profiling system built on Hive that uses SQL templates, custom UDFs, and set‑operation logic to enable attribute‑based user segmentation, batch processing, and seamless integration with downstream services.

SQL templatesSet Operationsdata engineering
0 likes · 11 min read
Design and Implementation of a Scalable User Profiling System Using Hive and SQL Templates
DataFunTalk
DataFunTalk
Nov 24, 2018 · Big Data

The Evolution of iQIYI's Big Data Analytics Platform

This article chronicles iQIYI’s journey from a simple Hive‑based data pipeline to the sophisticated, multi‑engine “Tongtian Tower” platform, detailing the development of the Magic Mirror system, the Gear workflow manager, BabelBD, the Monet visual analytics tool, and the integrated BI ecosystem that now supports billions of daily users.

BIbig datadata engineering
0 likes · 18 min read
The Evolution of iQIYI's Big Data Analytics Platform
DataFunTalk
DataFunTalk
Nov 21, 2018 · Artificial Intelligence

Personalized Recommendation System of 51 Credit Card: Architecture, Challenges, and Growth Cases

This article details how 51 Credit Card leverages artificial intelligence to build a personalized recommendation system, covering business pain points, technical challenges, a three‑layer tagging architecture from bill and app data, model deployment pipelines, and real‑world growth case studies that boosted conversion and ROI.

AIdata engineeringfinance
0 likes · 14 min read
Personalized Recommendation System of 51 Credit Card: Architecture, Challenges, and Growth Cases
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Nov 20, 2018 · Big Data

A Decade of Alibaba's Big Data Platform Evolution Through Double 11

The article chronicles Alibaba's ten‑year journey of building and scaling its big data platform—from early Oracle clusters and Hadoop‑based Cloud‑Ladder 1 to the self‑developed ODPS/MaxCompute, real‑time Blink engine, and the unified DataWorks ecosystem—highlighting key technical milestones, performance breakthroughs, and operational challenges that powered successive Double 11 shopping festivals.

AlibabaMaxComputedata engineering
0 likes · 22 min read
A Decade of Alibaba's Big Data Platform Evolution Through Double 11
Programmer DD
Programmer DD
Nov 7, 2018 · Big Data

Choosing the Right SQL Engine for Big Data: A Practical Guide

This article explores various SQL engines and storage options for big‑data workloads, compares their performance and capabilities, shows practical code examples, and offers guidance on writing efficient SQL in complex data environments.

HiveSQLSQL Engines
0 likes · 6 min read
Choosing the Right SQL Engine for Big Data: A Practical Guide
Architect's Tech Stack
Architect's Tech Stack
Oct 23, 2018 · Fundamentals

Common Data Collection Challenges in Startups and Practical Solutions

The article examines three typical data collection problems faced by startups—unclear collection methods, chaotic tracking points, and poor collaboration between data and engineering teams—and offers practical strategies such as adopting full‑event models, appointing data architects, and securing top‑down support to achieve reliable, comprehensive analytics.

Analyticsdata collectiondata engineering
0 likes · 10 min read
Common Data Collection Challenges in Startups and Practical Solutions
Meituan Technology Team
Meituan Technology Team
Oct 18, 2018 · Big Data

Building a Real-Time Data Warehouse with Flink at Meituan

Meituan replaced its Storm‑based pipeline with a four‑layer real‑time data warehouse powered by Flink, using hybrid storage (Cellar KV, Elasticsearch, Druid, MySQL) to deliver low‑latency, high‑throughput services, dramatically simplifying SQL‑driven development, unifying metrics, cutting compute costs, and paving the way for offline‑grade accuracy and reliability.

FlinkMeituanReal-time Data Warehouse
0 likes · 16 min read
Building a Real-Time Data Warehouse with Flink at Meituan
Beike Product & Technology
Beike Product & Technology
Sep 28, 2018 · Databases

Using ClickHouse for Large‑Scale User Behavior Analysis at Beike Zhaofang

This article details how Beike Zhaofang leveraged the ClickHouse columnar OLAP database for large‑scale user behavior analysis, covering its architecture, key features, performance benchmarks against other engines, data ingestion pipelines, custom UDFs for funnel and retention metrics, deployment setup, and future enhancements.

ClickHouseFunnel AnalysisOLAP
0 likes · 13 min read
Using ClickHouse for Large‑Scale User Behavior Analysis at Beike Zhaofang
DataFunTalk
DataFunTalk
Sep 2, 2018 · Artificial Intelligence

From Zero to One: Building and Deploying Knowledge Graphs at Beike Real Estate

This article details the evolution, architecture, and practical applications of knowledge graphs at Beike Real Estate, covering their historical background, five‑view advantages, data pipelines, ontology construction, intelligent search, recommendation, and chatbot integration, while also discussing challenges and future directions.

Artificial IntelligenceIntelligent AssistantNLP
0 likes · 13 min read
From Zero to One: Building and Deploying Knowledge Graphs at Beike Real Estate
Meitu Technology
Meitu Technology
Aug 14, 2018 · Big Data

Meitu Data Platform Architecture and Practices

Meitu’s data platform, serving dozens of apps with 500 million monthly active users and billions of daily events, combines the Arachnia log‑collection system, Kafka ingestion, multi‑layer storage (HDFS, MongoDB, HBase, Elasticsearch), offline Hive/MapReduce processing and real‑time Storm/Flink/Naix pipelines, supported by data‑workshop tools, staged evolution for scalability, and robust security and query‑validation mechanisms.

ETLHiveStreaming
0 likes · 16 min read
Meitu Data Platform Architecture and Practices
Meitu Technology
Meitu Technology
Aug 11, 2018 · Big Data

Meitu Technology Salon: Evolution of the Big Data Platform, Distributed Bitmap (Naix), and Apache Kylin

At Meitu’s Technology Salon, senior big‑data experts detailed the end‑to‑end architecture and stability measures of Meitu’s large‑scale data platform, introduced the high‑performance distributed bitmap solution Naix, showcased the evolution of Meizu’s user‑insight system, and highlighted Apache Kylin’s OLAP capabilities and Superset integration for scalable, real‑time analytics.

Apache KylinData Analyticsbig data
0 likes · 9 min read
Meitu Technology Salon: Evolution of the Big Data Platform, Distributed Bitmap (Naix), and Apache Kylin
Architecture Digest
Architecture Digest
Jul 6, 2018 · Backend Development

Essential Backend Infrastructure and Services for Java Applications

This article outlines the fundamental backend components, frameworks, and services—including API gateways, authentication centers, configuration management, service governance, scheduling, logging, data pipelines, and monitoring—required to build robust, scalable Java business applications for both online and internal use.

API GatewayBackendJava
0 likes · 20 min read
Essential Backend Infrastructure and Services for Java Applications
AntTech
AntTech
Apr 28, 2018 · Artificial Intelligence

The Future of AI: Intelligent Infrastructure, Engineering Challenges, and the Limits of Current Approaches

The article examines Michael I. Jordan’s critique of current AI research, highlights the need for intelligent infrastructure that integrates computing, data, and physical systems across domains, and uses real‑world examples to argue for a new engineering discipline beyond narrow deep‑learning hype.

AIdata engineeringfuture of AI
0 likes · 23 min read
The Future of AI: Intelligent Infrastructure, Engineering Challenges, and the Limits of Current Approaches
Meituan Technology Team
Meituan Technology Team
Jan 26, 2018 · Big Data

Design and Implementation of a Real-Time Data Processing System at Meituan

Meituan designed a Storm‑based real‑time data processing platform that guarantees at‑least‑once delivery and high availability, employs a custom spout, regression‑driven traffic smoothing, and a low‑latency KV store with atomic operations, persisting results in Kafka, MySQL and Cellar to power merchant dashboards and heat‑tag analytics, while planning broader real‑time analytics expansion.

Real-time DataStormbig data
0 likes · 10 min read
Design and Implementation of a Real-Time Data Processing System at Meituan
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Jan 5, 2018 · Big Data

What Drove Big Data’s 2017 Surge and What’s Next? Insights & Predictions

Analyzing 2017’s big data boom, the article explores how the 4V characteristics—volume, variety, velocity, and value—spurred innovations like distributed storage, NoSQL, real‑time stream processing, and AI integration, and predicts future hotspots such as SQL resurgence, cloud‑based platforms, and AI‑driven analytics.

Artificial Intelligencebig datadata engineering
0 likes · 11 min read
What Drove Big Data’s 2017 Surge and What’s Next? Insights & Predictions
Alibaba Cloud Developer
Alibaba Cloud Developer
Oct 23, 2017 · Big Data

How Alibaba Built Its Full‑Domain Data Platform: Architecture & Lessons

In this detailed account, Alibaba senior technologist Zhang Lei explains the concept of a full‑domain data platform, its “four‑horizontal and three‑vertical” architecture, the OneData ecosystem, cost‑saving strategies, data quality tools, and practical challenges of building and operating massive big‑data infrastructure across the Alibaba ecosystem.

Alibabadata engineeringdata middle platform
0 likes · 15 min read
How Alibaba Built Its Full‑Domain Data Platform: Architecture & Lessons