Tagged articles

big data

3804 articles · Page 35 of 39
dbaplus Community
dbaplus Community
Dec 14, 2017 · Big Data

Scaling Vipshop’s Big Data Platform: Monitoring, Multi‑HDFS, Yarn Optimization & Capping

In 2017 Vipshop’s senior big‑data architect shares how the company grew its Hadoop‑based platform from zero to a thousand‑node cluster, detailing cluster health monitoring, multi‑HDFS deployment via Hive, Yarn container allocation improvements, and a hook‑driven Capping resource‑control system to boost stability and efficiency.

HDFSbig datacapping
0 likes · 15 min read
Scaling Vipshop’s Big Data Platform: Monitoring, Multi‑HDFS, Yarn Optimization & Capping
Qunar Tech Salon
Qunar Tech Salon
Dec 14, 2017 · Databases

TiDB Architecture, Deployment, and Monitoring Practices at Qunar

This article explains Qunar's transition from MySQL, Redis, and HBase to TiDB, detailing the background of distributed databases, TiDB's architecture, hardware selection, deployment automation, monitoring setup, and real‑world usage scenarios to address scalability and high‑availability challenges.

TiDBbig datadatabase architecture
0 likes · 14 min read
TiDB Architecture, Deployment, and Monitoring Practices at Qunar
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Dec 11, 2017 · Artificial Intelligence

How AI and Big Data Are Transforming Urban Traffic Management

The 2017 12th China Intelligent Transportation Conference highlighted system thinking, AI, and innovation as key drivers for smarter city traffic, outlining a three‑step top‑level design, AI‑powered applications, and intersection innovations that together promise safer, more efficient, and fully automated urban mobility.

AIIntelligent TransportationTraffic Management
0 likes · 8 min read
How AI and Big Data Are Transforming Urban Traffic Management
AntTech
AntTech
Dec 11, 2017 · Artificial Intelligence

How AI and Big Data Transform the Insurance Industry: Differentiated Pricing, Smart Claims, Risk Control, and Operations

The article examines how emerging AI and big‑data technologies are reshaping insurance by enabling differentiated pricing, automating claims and customer service, strengthening fraud detection, and improving personalized product recommendation and operational efficiency across the sector.

Artificial IntelligenceDifferentiated PricingInsurtech
0 likes · 13 min read
How AI and Big Data Transform the Insurance Industry: Differentiated Pricing, Smart Claims, Risk Control, and Operations
Efficient Ops
Efficient Ops
Dec 7, 2017 · Operations

How Multi-Dimensional Root Cause Analysis Boosts Monitoring Efficiency with AI

This article introduces the challenges of multi-dimensional monitoring, explains the limitations of traditional alerting, and presents the MDRCA algorithm—combining K‑means clustering, Explanatory Power, and Surprise metrics—to pinpoint root causes efficiently, while sharing practical AI integration experiences for large‑scale monitoring platforms.

AIKMeansMultidimensional
0 likes · 15 min read
How Multi-Dimensional Root Cause Analysis Boosts Monitoring Efficiency with AI
Meituan Technology Team
Meituan Technology Team
Dec 1, 2017 · Big Data

Metric Logic Tree: Automated Anomaly Analysis for Business Metrics

The Metric Logic Tree automates business metric anomaly analysis by integrating heterogeneous data sources (Kylin, MySQL, Elasticsearch, Druid) with a three‑layer architecture—metric calculation, algorithmic analysis (waterfall and Gini‑coefficient methods), and a master‑worker computation service—that parallelizes queries, delivers immediate conclusions, and shortens decision cycles, as demonstrated in Meituan‑Dianping’s hotel‑travel operations.

Data Pipelinealgorithmanomaly detection
0 likes · 7 min read
Metric Logic Tree: Automated Anomaly Analysis for Business Metrics
AntTech
AntTech
Dec 1, 2017 · Big Data

Insights and Paper Summaries from KDD 2017 Conference

The article provides a comprehensive overview of KDD 2017, including acceptance statistics, best paper awards, Ant Group's contributions, detailed discussions on AB testing, graph mining, and selected research papers across data mining, machine learning, and anomaly detection, offering valuable insights for practitioners and researchers.

AB TestingKDDbig data
0 likes · 30 min read
Insights and Paper Summaries from KDD 2017 Conference
Efficient Ops
Efficient Ops
Nov 27, 2017 · Operations

How Facebook Scales to Billions: Disaggregated Networks, Storage, and Warm Spark

Facebook’s journey from early startup ops to supporting over 2 billion monthly users reveals how disaggregated network, storage, and warm‑storage‑enabled Spark architectures overcome scalability bottlenecks, illustrating the operational strategies and design principles that power massive, reliable data‑center services.

Cloud InfrastructureOperationsbig data
0 likes · 12 min read
How Facebook Scales to Billions: Disaggregated Networks, Storage, and Warm Spark
iQIYI Technical Product Team
iQIYI Technical Product Team
Nov 24, 2017 · Information Security

Risk Control System for Live Streaming: Real‑time Interception (Pluto) and Big Data Analysis (Mars)

iQIYI’s live‑stream risk‑control platform combines the real‑time interception engine Pluto with the big‑data analytics system Mars to curb black‑market registration fraud and red‑packet abuse, processing over a billion daily requests through adaptive filters, Kafka‑Spark pipelines, and clustering algorithms that now limit fake popularity to 10‑30 % and red‑packet capture to under 3 %.

MarsPlutobig data
0 likes · 11 min read
Risk Control System for Live Streaming: Real‑time Interception (Pluto) and Big Data Analysis (Mars)
Suning Technology
Suning Technology
Nov 20, 2017 · Big Data

How ZEUS Turns Monitoring Data into Automated Decisions for Enterprise Systems

ZEUS, Suning’s decision analysis platform, integrates monitoring data from tools like Baymax and HIRO, applies CEP aggregation and Drools rule evaluation, and leverages big‑data storage and machine‑learning models to automatically identify root causes, provide real‑time alerts, and enable self‑healing in large‑scale distributed systems.

Rule EngineSelf-Healingbig data
0 likes · 14 min read
How ZEUS Turns Monitoring Data into Automated Decisions for Enterprise Systems
Architects' Tech Alliance
Architects' Tech Alliance
Nov 16, 2017 · Operations

Understanding AIOps: How AI‑Driven Operations Transform IT Management

The article explains how AIOps—an AI‑powered IT operations platform that combines big‑data analytics, machine learning, and automation—revolutionizes traditional IT Ops by enabling rapid, accurate incident detection, root‑cause analysis, and self‑healing, thereby freeing CIOs to focus on strategic business value.

AIOpsDigital Transformationautomation
0 likes · 8 min read
Understanding AIOps: How AI‑Driven Operations Transform IT Management
Efficient Ops
Efficient Ops
Nov 15, 2017 · Big Data

How Tencent Built a 10 TB‑Per‑Day Full‑Link Log Monitoring Platform

This article explains how Tencent's ZhiYun full‑link log monitoring platform handles massive daily logs, overcomes challenges of diverse log formats, high throughput, fault‑tolerant design, and provides scalable storage, query, and alerting capabilities for distributed micro‑service environments.

Data PipelineLog Monitoringbig data
0 likes · 10 min read
How Tencent Built a 10 TB‑Per‑Day Full‑Link Log Monitoring Platform
Suning Technology
Suning Technology
Nov 13, 2017 · Backend Development

How Suning Scaled Its Membership System for Double‑11: From Legacy POS to Multi‑Active Architecture

This article examines Suning's evolution of its membership platform—from an early offline POS system to a vertically split, cloud‑native architecture—detailing capacity planning, performance testing, data migration with Spark, multi‑active deployment, and future plans for cross‑region high availability.

Cloud NativeSystem Architecturebig data
0 likes · 15 min read
How Suning Scaled Its Membership System for Double‑11: From Legacy POS to Multi‑Active Architecture
21CTO
21CTO
Nov 11, 2017 · Big Data

How We Built a Scalable Seller Log System with Kafka, Storm, ES & HBase

This article explains the design and implementation of a unified seller‑operation logging platform that uses Kafka for ingestion, Storm for real‑time processing, Elasticsearch for hot‑data search, and HBase for cold‑data storage, detailing the challenges faced and the optimizations applied.

ElasticsearchHBaseKafka
0 likes · 12 min read
How We Built a Scalable Seller Log System with Kafka, Storm, ES & HBase
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Nov 8, 2017 · Operations

Inside Ctrip’s Evolving Architecture: Ops, Frameworks, and Big Data Insights

This article explores Ctrip’s continuously evolving architecture, detailing its three-layer composition of operations, frameworks, and applications, and examines real-world case studies of its release system, configuration management, SOA, and a massive User Profile big‑data project, highlighting key innovations and lessons learned.

CtripSOASystem Architecture
0 likes · 11 min read
Inside Ctrip’s Evolving Architecture: Ops, Frameworks, and Big Data Insights
Tencent Cloud Developer
Tencent Cloud Developer
Nov 3, 2017 · Industry Insights

How Tencent Cloud’s Big Data Platform Ranked in China’s Fifth Evaluation

China’s Data Center Alliance released its fifth big‑data product evaluation, testing 17 solutions from 16 vendors across SQL, NoSQL, and machine‑learning workloads, with Tencent Cloud’s platform achieving top rankings in NoSQL tests and highlighting the nation’s push toward standardized, high‑performance big‑data infrastructure.

Data PlatformsIndustry BenchmarkTencent Cloud
0 likes · 5 min read
How Tencent Cloud’s Big Data Platform Ranked in China’s Fifth Evaluation
Alibaba Cloud Developer
Alibaba Cloud Developer
Nov 3, 2017 · Big Data

How Alibaba Built an EB-Scale, Real-Time Big Data Platform

Alibaba’s senior data expert Yao Bin Hui explains how the company constructed a standardized, end-to-end big-data ecosystem—from low-level data collection and AI algorithms to data services and product platforms—enabling petabyte-scale integration and second-level response times that power both internal operations and millions of external users.

AlibabaData Servicesbig data
0 likes · 10 min read
How Alibaba Built an EB-Scale, Real-Time Big Data Platform
dbaplus Community
dbaplus Community
Oct 30, 2017 · Big Data

How to Build a Real‑Time Spam Monitoring System with Apache Storm

This article walks through the design, deployment, and code implementation of a real‑time spam detection pipeline using Apache Storm, comparing it with Hadoop, detailing cluster setup, topology components, data flow, and how to package and run the solution on a distributed Storm cluster.

Apache StormHibernateJava
0 likes · 13 min read
How to Build a Real‑Time Spam Monitoring System with Apache Storm
21CTO
21CTO
Oct 26, 2017 · Backend Development

From Data Platform Battles to AI Dreams: A Senior Engineer’s 3‑Year Journey at Alibaba

A senior Alibaba engineer reflects on three years of building a large‑scale data platform, tackling distributed rate‑limiting challenges, leading cross‑regional projects, and pursuing AI research, while sharing personal insights on career growth, technical problem‑solving, and the value of continuous learning.

AI learningbig datacareer reflections
0 likes · 11 min read
From Data Platform Battles to AI Dreams: A Senior Engineer’s 3‑Year Journey at Alibaba
Liulishuo Tech Team
Liulishuo Tech Team
Oct 22, 2017 · Big Data

Data-CI: A SQL-Based Data Unit Testing Framework for ETL

The article introduces data-ci, a SQL‑driven unit testing framework that lets engineers write, organize, and automate data validation tests for ETL pipelines, providing assertions, failure callbacks, coverage reporting, and CI integration to improve data quality and reliability.

Data QualityETLSQL
0 likes · 9 min read
Data-CI: A SQL-Based Data Unit Testing Framework for ETL
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Oct 21, 2017 · Big Data

Deploy Hadoop CDH5.4 on CentOS 6: Install HDFS, YARN, and WebHDFS

This guide walks through preparing three CentOS 6.9 nodes, configuring hostnames, time sync, password‑less SSH, disabling IPv6, installing JDK, downloading CDH 5.4, setting up core‑site and hdfs‑site XML files, formatting the NameNode, starting HDFS services, configuring YARN and MapReduce, and verifying the installations via the Web UI.

CDHCentOSHDFS
0 likes · 18 min read
Deploy Hadoop CDH5.4 on CentOS 6: Install HDFS, YARN, and WebHDFS
Efficient Ops
Efficient Ops
Oct 18, 2017 · Operations

How Bilibili Scaled Its Log System to 10TB Daily with Elastic Stack

This article details Bilibili's Billions log platform—from its fragmented origins and design goals to the elastic‑stack‑based architecture, shard management, log sampling, custom Go splitters, and monitoring enhancements—highlighting the challenges faced and the roadmap for future improvements.

Elastic StackLog ManagementOperations
0 likes · 17 min read
How Bilibili Scaled Its Log System to 10TB Daily with Elastic Stack
dbaplus Community
dbaplus Community
Oct 15, 2017 · Big Data

How JD Built a Scalable Seller Log Platform with Kafka, Storm, ES & HBase

This article details JD's end‑to‑end seller log system architecture, explaining why Kafka, Storm, Elasticsearch and HBase were chosen, the challenges faced during scaling, and the practical solutions implemented to achieve a unified, high‑throughput logging platform for merchants and operations.

ElasticsearchHBaseKafka
0 likes · 13 min read
How JD Built a Scalable Seller Log Platform with Kafka, Storm, ES & HBase
Alibaba Cloud Developer
Alibaba Cloud Developer
Oct 15, 2017 · Information Security

How Alibaba’s Data Security Maturity Model (DSMM) Is Shaping China’s Data Protection Landscape

The article explains Alibaba's Data Security Maturity Model (DSMM), its partnership program, the involvement of 17 leading security firms, and how the model aims to improve data security capabilities across industries by establishing standardized assessment criteria and fostering ecosystem collaboration.

AlibabaDSMMbig data
0 likes · 10 min read
How Alibaba’s Data Security Maturity Model (DSMM) Is Shaping China’s Data Protection Landscape
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Oct 12, 2017 · Backend Development

How Taobao Scaled Its Backend Architecture Over Time

This article outlines Taobao's learning objectives, traces the evolution of its backend architecture from V1.0 to V3.0, highlights the technical challenges faced at each stage, and explains the architectural decisions—such as modularization, service‑oriented frameworks, distributed storage, and large‑scale monitoring—that enabled massive scalability, reliability, and performance improvements.

ArchitectureBackendbig data
0 likes · 6 min read
How Taobao Scaled Its Backend Architecture Over Time
Baidu Intelligent Testing
Baidu Intelligent Testing
Oct 9, 2017 · Big Data

User Behavior Analysis: From Data Acquisition to Funnel Insights

The article explains how to move beyond macro app metrics by collecting offline and real‑time user data, storing it in HDFS, processing it with Spark, visualizing behavior paths as state‑machine trees, and performing branch‑funnel analysis to uncover conversion bottlenecks and improve product quality.

AnalyticsFunnel Analysisbig data
0 likes · 5 min read
User Behavior Analysis: From Data Acquisition to Funnel Insights
ITPUB
ITPUB
Sep 30, 2017 · Big Data

Designing Scalable Open‑Source ETL Systems: Lessons from Baidu Waimai

This talk details Baidu Waimai's end‑to‑end ETL design, covering demand sources, data flow patterns, multi‑stage system evolution, storage choices, scheduling architecture, configuration‑driven processing, quality monitoring, and how data lineage enables transparent, self‑service data delivery.

Data PipelineData QualityETL
0 likes · 25 min read
Designing Scalable Open‑Source ETL Systems: Lessons from Baidu Waimai
Tongcheng Travel Technology Center
Tongcheng Travel Technology Center
Sep 29, 2017 · Big Data

Evolution of Monitoring Architecture and Traffic Alert Algorithms at Tongcheng Travel

This article describes how Tongcheng Travel’s monitoring system evolved from a monolithic design to a distributed and big‑data‑based architecture, introducing real‑time processing with Storm, machine‑learning‑enhanced alerts, and a multivariate linear regression model that dramatically improves traffic anomaly detection accuracy.

architecture evolutionbig datamachine learning
0 likes · 10 min read
Evolution of Monitoring Architecture and Traffic Alert Algorithms at Tongcheng Travel
ITPUB
ITPUB
Sep 29, 2017 · Big Data

Designing an Open ETL System: Baidu Waimai’s Scalable Data Pipeline Practices

In this talk, a Baidu Waimai engineer explains the motivations, requirements, and architectural choices behind their open‑source ETL platform, covering data flow patterns, logical mappings, storage options, scheduling, metadata management, and quality monitoring to achieve scalable, transparent, and explainable data delivery.

Data PipelineETLbig data
0 likes · 26 min read
Designing an Open ETL System: Baidu Waimai’s Scalable Data Pipeline Practices
21CTO
21CTO
Sep 25, 2017 · Big Data

How Meitu Scaled Its Billion-User Data Analytics: Architecture Evolution and Lessons

This article explains how Meitu built and evolved a large‑scale data statistics platform to handle billions of users, detailing the challenges of growing data volume, the architectural shifts from simple scripts to Hadoop, and the design of modular components for job management, scheduling, execution, and future expansion.

HadoopHiveJob Scheduling
0 likes · 16 min read
How Meitu Scaled Its Billion-User Data Analytics: Architecture Evolution and Lessons
ITPUB
ITPUB
Sep 22, 2017 · Big Data

How Baidu Waimai Scaled Traffic Analysis with Apache Kylin: A Deep Dive

This article presents a detailed case study of Baidu Waimai's traffic analysis platform, outlining the data challenges of high dimensionality and volume, the evaluation of OLAP engines, the adoption of Apache Kylin for pre‑computation, the end‑to‑end data modeling, cube construction, incremental builds, and integration with Saiku‑Mondrian reporting, while sharing practical lessons and performance gains.

Apache KylinOLAPPrecomputation
0 likes · 29 min read
How Baidu Waimai Scaled Traffic Analysis with Apache Kylin: A Deep Dive
Meituan Technology Team
Meituan Technology Team
Sep 21, 2017 · Big Data

Feature Production Scheduling: Architecture Evolution and Core Technologies

Using Meituan‑Dianping’s hospitality online feature system as a case study, the article describes how feature production scheduling evolved from offline batch ETL to automated, metadata‑driven pipelines and sub‑second streaming, detailing the underlying architecture, incremental updates, storage abstraction, write‑shaving, atomicity, and recovery mechanisms.

Data PipelineSystem Architecturebig data
0 likes · 23 min read
Feature Production Scheduling: Architecture Evolution and Core Technologies
Ctrip Technology
Ctrip Technology
Sep 20, 2017 · Big Data

Building a Real‑Time Computing Platform with Spark Streaming at Ctrip: Design, Implementation, and Lessons Learned

This article describes how Ctrip migrated its large‑scale real‑time platform from JStorm to Spark Streaming, detailing the architectural design, the Muise Spark Core encapsulation, operational metrics, encountered pitfalls, and future plans to adopt Flink and Beam for streaming workloads.

Exactly-OnceSpark StreamingYARN
0 likes · 22 min read
Building a Real‑Time Computing Platform with Spark Streaming at Ctrip: Design, Implementation, and Lessons Learned
Alibaba Cloud Developer
Alibaba Cloud Developer
Sep 19, 2017 · Artificial Intelligence

Inside Alibaba’s 2017 Tech Forum: AI, Big Data, and Cloud Innovations Unveiled

At the inaugural 2017 Alibaba Technology Forum held at Hong Kong University of Science and Technology, senior executives highlighted Alibaba’s cutting‑edge AI, machine learning, big‑data, and cloud breakthroughs, showcasing how data‑driven technologies power billions of users across e‑commerce, finance, logistics, healthcare, and entertainment.

Cloud Computingbig data
0 likes · 6 min read
Inside Alibaba’s 2017 Tech Forum: AI, Big Data, and Cloud Innovations Unveiled
MaGe Linux Operations
MaGe Linux Operations
Sep 11, 2017 · Big Data

How Big Data Can Revolutionize Operations Monitoring

This article explores applying big‑data thinking and platforms—such as Flume, Spark Streaming, and HBase—to operations monitoring, detailing data sources, metric categories, architecture design, implementation steps, and the benefits of a scalable, low‑code monitoring platform.

ArchitectureOperationsSpark Streaming
0 likes · 10 min read
How Big Data Can Revolutionize Operations Monitoring
21CTO
21CTO
Sep 5, 2017 · Big Data

Build a PHP Word Count with Hadoop MapReduce: Step-by-Step Guide

This article explains what MapReduce is, when to use it, and how to implement a PHP word‑count and a gold‑price average calculation on an Apache Hadoop cluster, covering installation hints, mapper and reducer scripts, testing commands, and visualizing results with gnuplot.

Data ProcessingGnuplotHadoop
0 likes · 10 min read
Build a PHP Word Count with Hadoop MapReduce: Step-by-Step Guide
MaGe Linux Operations
MaGe Linux Operations
Sep 4, 2017 · Fundamentals

The Ultimate Technical Knowledge Map: 50+ Skill Charts for Architects & Developers

This article presents a comprehensive collection of technical knowledge maps compiled over years, covering architecture, Java, microservices, consistency, big data, cloud computing, mobile development, front‑end, back‑end, DevOps, and more, aiming to help engineers and architects master essential skills and best practices.

ArchitectureCloud ComputingJava
0 likes · 6 min read
The Ultimate Technical Knowledge Map: 50+ Skill Charts for Architects & Developers
Tencent IMWeb Frontend Team
Tencent IMWeb Frontend Team
Sep 3, 2017 · Frontend Development

What’s Hot This Week in Web Tech? Apple Event, KSQL, Polymer 3, and More

This week’s IMWeb Frontend Community roundup highlights the Apple September event details, introduces KSQL for Apache Kafka, previews Polymer 3.0’s shift to ES6 modules, discusses the Ayo.js Node.js fork, ASP.NET Core 2 Razor pages, VS 2017 preview, container adoption trends, and Oracle’s cloud database innovations.

Technology Newsbig datacloud
0 likes · 6 min read
What’s Hot This Week in Web Tech? Apple Event, KSQL, Polymer 3, and More
Architecture Digest
Architecture Digest
Sep 2, 2017 · Big Data

Designing a High‑Availability, High‑Efficiency Distributed Scheduling Platform for Big Data

This article examines the principles, features, and implementation details of distributed scheduling for big‑data ETL pipelines, covering decentralised schedulers, host selection strategies, fault‑tolerance, operator abstraction, elasticity, trigger mechanisms, visual monitoring, alarm handling, data fan‑in/fan‑out, parameter consistency, real‑time quality checks, lineage tracking, and field‑level traceability.

Data PipelineETLbig data
0 likes · 23 min read
Designing a High‑Availability, High‑Efficiency Distributed Scheduling Platform for Big Data
21CTO
21CTO
Aug 27, 2017 · Big Data

Uncovering Ghost Bikes: How to Crawl and Analyze Mobike Data in Chengdu

This article details the process of capturing Mobike's public API data, building a high‑performance Python crawler with proxy rotation, storing the results in databases, and performing large‑scale analysis to reveal stationary bikes, travel distances, usage frequency, and urban development patterns in Chengdu.

Mobikebig databike sharing
0 likes · 13 min read
Uncovering Ghost Bikes: How to Crawl and Analyze Mobike Data in Chengdu
Meituan Technology Team
Meituan Technology Team
Aug 25, 2017 · Big Data

Data Platform Integration and Multi‑Data‑Center Architecture at Meituan‑Dianping

After Meituan merged with Dianping, engineers unified two massive Hadoop ecosystems across Beijing and Shanghai by breaking the project into four phases—unify, copy, switch, fuse—standardizing versions, implementing zone‑aware transfers, cross‑realm Kerberos, and federated metadata to achieve a single, reliable multi‑data‑center platform.

Cluster FusionDistcpHadoop
0 likes · 32 min read
Data Platform Integration and Multi‑Data‑Center Architecture at Meituan‑Dianping
21CTO
21CTO
Aug 21, 2017 · Big Data

Rethinking Hadoop: When to Use It and How Cloud Computing Changes the Game

This article reviews when Hadoop is appropriate, outlines its core features and limitations, explains cloud computing concepts and service models, and highlights the benefits of pre‑built Hadoop images for accelerating big‑data projects.

HadoopPre-built Imagesbig data
0 likes · 13 min read
Rethinking Hadoop: When to Use It and How Cloud Computing Changes the Game
Architecture Digest
Architecture Digest
Aug 15, 2017 · Artificial Intelligence

Why AI Engineers Must Understand Basic Infrastructure: From Big Data to Deep Learning

The article explains why AI engineers need foundational infrastructure knowledge—covering big‑data processing, cloud services, containerization, MapReduce, and deep‑learning platforms—to effectively solve real‑world problems, collaborate with teams, and build scalable, maintainable AI solutions.

AI InfrastructureCloud ComputingMachine Learning Engineering
0 likes · 14 min read
Why AI Engineers Must Understand Basic Infrastructure: From Big Data to Deep Learning
21CTO
21CTO
Aug 14, 2017 · Big Data

Unveiling Flink’s Multi‑Layer Execution Graph: From StreamGraph to Physical Deployment

This article explains Flink’s architecture, detailing the roles of Client, JobManager and TaskManager, walks through a SocketTextStreamWordCount example, and clarifies the four‑layer graph model—StreamGraph, JobGraph, ExecutionGraph, and the physical execution graph—highlighting why each layer exists.

Execution GraphFlinkJobManager
0 likes · 9 min read
Unveiling Flink’s Multi‑Layer Execution Graph: From StreamGraph to Physical Deployment
Alibaba Cloud Developer
Alibaba Cloud Developer
Aug 10, 2017 · Big Data

Alibaba’s HBase Innovations: Powering Big Data at Scale – HBaseCon 2017 Asia Insights

At HBaseCon 2017 Asia, Alibaba showcased a series of groundbreaking HBase enhancements—including strong synchronous replication, SQL-on-HBase capabilities, cross‑cluster range data copy, and read/write path optimizations—that dramatically improve performance, reliability, and usability for large‑scale big‑data storage.

HBasePerformanceReplication
0 likes · 10 min read
Alibaba’s HBase Innovations: Powering Big Data at Scale – HBaseCon 2017 Asia Insights
High Availability Architecture
High Availability Architecture
Aug 8, 2017 · Big Data

Practical Big Data Architecture Evolution and Lessons Learned

The article reviews the evolution of big‑data architectures from a simple RDB‑centric pipeline to a SaaS‑based solution, highlighting common bottlenecks such as scaling, integration, cost, and operational complexity, and shares practical experiences and best‑practice recommendations for building efficient, maintainable data platforms.

ArchitectureData PipelineSaaS
0 likes · 12 min read
Practical Big Data Architecture Evolution and Lessons Learned
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jul 26, 2017 · Big Data

Inside Taobao’s Massive Data Architecture: From Hadoop “Cloud Ladder” to Real‑Time “Galaxy”

This article details Taobao’s multi‑layer massive data platform, covering its five‑tier architecture, the 1500‑node Hadoop “Cloud Ladder” for batch processing, the low‑latency “Galaxy” stream engine, MySQL‑based MyFOX, HBase‑based Prom storage, the glider middle‑layer, and sophisticated caching strategies that together support petabytes of data and millions of daily queries.

CachingHBaseHadoop
0 likes · 16 min read
Inside Taobao’s Massive Data Architecture: From Hadoop “Cloud Ladder” to Real‑Time “Galaxy”
21CTO
21CTO
Jul 22, 2017 · Big Data

Why Every Company Needs a Chief Data Officer to Unlock Data Value

The article explains the strategic importance of the Chief Data Officer role, outlining how CDOs drive data‑driven innovation through a four‑stage data supply chain—data supply, logistics, science, and execution—to create competitive advantage and business growth.

Chief Data OfficerData Supply Chainbig data
0 likes · 14 min read
Why Every Company Needs a Chief Data Officer to Unlock Data Value
Architecture Digest
Architecture Digest
Jul 22, 2017 · Big Data

Popular Big Data Tools and Their Descriptions

This article provides an extensive overview of more than ninety open‑source and commercial big‑data tools—including ETL platforms, resource managers, storage systems, messaging queues, processing engines, and visualization libraries—detailing their core functions, typical use cases, and notable adopters.

AnalyticsData IntegrationETL
0 likes · 26 min read
Popular Big Data Tools and Their Descriptions
High Availability Architecture
High Availability Architecture
Jul 19, 2017 · Artificial Intelligence

Weiflow: A Scalable Machine Learning Workflow Framework for Sina Weibo

The article introduces Weiflow, a dual‑layer DAG‑based machine‑learning workflow framework designed for Sina Weibo, and explains how its modular XML configuration, Scala implementation, and integration with Spark, TensorFlow, Hive, Storm, and Flink improve development efficiency, scalability, and execution performance across the entire ML pipeline.

DAGScalaSpark
0 likes · 16 min read
Weiflow: A Scalable Machine Learning Workflow Framework for Sina Weibo
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jul 17, 2017 · Big Data

Mastering Data Sync, Real‑Time Processing, and Scalable Storage for Modern Systems

This article explores practical techniques for synchronizing heterogeneous data sources, performing batch and incremental analytics with Hadoop and Spark, designing low‑latency real‑time computation pipelines, implementing push notifications, and choosing appropriate storage solutions—from in‑memory caches to distributed databases—while addressing performance, reliability, and scalability challenges.

big datadata synchronizationdatabases
0 likes · 25 min read
Mastering Data Sync, Real‑Time Processing, and Scalable Storage for Modern Systems
Alibaba Cloud Developer
Alibaba Cloud Developer
Jul 17, 2017 · Artificial Intelligence

How Alibaba Turns Big Data into ‘Data New Energy’ with Automated Tagging and Distributed Knowledge Graphs

Alibaba's senior algorithm expert Yang Hongxia explains how the company fuses massive, heterogeneous data sources into a unified platform, builds automated tag‑production pipelines and large‑scale distributed knowledge graphs, and applies these technologies to drive smarter business decisions and AI‑enabled services.

Alibabaautomated taggingbig data
0 likes · 14 min read
How Alibaba Turns Big Data into ‘Data New Energy’ with Automated Tagging and Distributed Knowledge Graphs
Efficient Ops
Efficient Ops
Jul 16, 2017 · Cloud Computing

Why PB‑Level Object Storage Is Essential and How to Choose the Right Solution

With data volumes soaring to petabyte scales, the article explains why object storage is the only viable solution for massive storage needs, outlines procurement considerations, design principles, and operational challenges, and offers practical guidance for building, evaluating, and scaling PB‑level storage systems.

Cloud ComputingObject Storagebig data
0 likes · 38 min read
Why PB‑Level Object Storage Is Essential and How to Choose the Right Solution
Architecture Digest
Architecture Digest
Jul 13, 2017 · Operations

Comprehensive Architecture and DevOps Tool Knowledge Map

This article compiles an extensive collection of architecture knowledge maps and a detailed overview of DevOps tools, categorizing them by development, deployment, and maintenance functions while also presenting related big‑data and cloud‑computing skill maps for engineers seeking a holistic view of modern software infrastructure.

ArchitectureCloud ComputingDevOps
0 likes · 9 min read
Comprehensive Architecture and DevOps Tool Knowledge Map
High Availability Architecture
High Availability Architecture
Jul 12, 2017 · Artificial Intelligence

Machine Learning Platform and Risk‑Control Applications at DianRong Net

The article presents a comprehensive overview of DianRong Net's in‑house machine‑learning platform built on Spark, its workflow, pain points it addresses, risk‑control case studies using graph mining, and practical tips for improving model performance through data, algorithms, hyper‑parameter tuning and ensemble methods.

Sparkbig datagraph mining
0 likes · 14 min read
Machine Learning Platform and Risk‑Control Applications at DianRong Net
dbaplus Community
dbaplus Community
Jul 10, 2017 · Big Data

Master Apache Storm: Real‑Time Stream Processing from Basics to Word‑Count and Call‑Log Examples

This tutorial explains Apache Storm’s core principles, architecture, and development workflow, covering its relationship with Hadoop, key concepts such as spouts, bolts, tuples, and topologies, and provides step‑by‑step code examples for a word‑count program and a call‑log analysis application.

Apache Stormbig datacall log analysis
0 likes · 14 min read
Master Apache Storm: Real‑Time Stream Processing from Basics to Word‑Count and Call‑Log Examples
Efficient Ops
Efficient Ops
Jul 9, 2017 · Cloud Native

How Goldwind Accelerated Wind Energy Management with Cloud‑Native Microservices

Goldwind transformed its global wind‑farm operations by adopting a cloud‑native, container‑based microservice architecture that tackles iteration speed, hybrid‑cloud deployment, and IoT big‑data challenges, enabling faster development, cost reduction, and advanced energy‑forecasting capabilities.

DevOpsIoTMicroservices
0 likes · 8 min read
How Goldwind Accelerated Wind Energy Management with Cloud‑Native Microservices
21CTO
21CTO
Jul 7, 2017 · Big Data

How to Kickstart Your Big Data Career: A Complete Learning Roadmap

This guide walks beginners through the vast big data landscape, helping them choose the right role, understand essential terminology, plan a learning path, and access curated resources for becoming a data engineer or analyst, all illustrated with clear diagrams.

Data AnalysisLearning Pathbig data
0 likes · 16 min read
How to Kickstart Your Big Data Career: A Complete Learning Roadmap
Meituan Technology Team
Meituan Technology Team
Jul 6, 2017 · Backend Development

Online Feature System: Architecture, Storage, and High‑Concurrency Techniques

Using Meituan’s hotel‑travel platform as a case study, the article details a scalable online feature system architecture that combines layered storage, efficient compression, and robust synchronization to meet extreme concurrency, throughput, terabyte‑scale data, and sub‑10 ms latency demands for AI‑driven strategy services.

big datadata compressiondistributed storage
0 likes · 23 min read
Online Feature System: Architecture, Storage, and High‑Concurrency Techniques
Alibaba Cloud Developer
Alibaba Cloud Developer
Jul 5, 2017 · Artificial Intelligence

Is This the New Golden Age of Visual AI? Insights from Alibaba Cloud

The article reviews the three historic AI booms, explains why today’s cloud‑based visual intelligence represents a distinct era, outlines five key factors for successful visual AI, and showcases real‑world Alibaba Cloud applications such as product search, city‑wide monitoring, medical diagnosis, and visual advertising.

AI applicationsAlibaba CloudCloud Computing
0 likes · 18 min read
Is This the New Golden Age of Visual AI? Insights from Alibaba Cloud
Tencent Advertising Technology
Tencent Advertising Technology
Jul 3, 2017 · Artificial Intelligence

Tencent Social Advertising College Algorithm Contest

Tencent's social advertising team hosts an algorithm contest for college students, leveraging big data and machine learning to develop innovative solutions for social advertising scenarios, inviting participants to submit algorithmic approaches to real-world advertising challenges.

Academic CompetitionAlgorithm ContestSocial Advertising
0 likes · 2 min read
Tencent Social Advertising College Algorithm Contest
21CTO
21CTO
Jul 3, 2017 · Big Data

Inside the World’s Best Data Architectures: Netflix, Facebook, Airbnb, Pinterest

This article explores the cutting‑edge data pipelines of Netflix, Facebook, Airbnb and Pinterest, detailing the massive event volumes they handle, the core technologies such as Kafka, Spark, Presto and Hadoop, and how these giants design scalable, real‑time analytics infrastructures.

AirbnbFacebookNetflix
0 likes · 6 min read
Inside the World’s Best Data Architectures: Netflix, Facebook, Airbnb, Pinterest
21CTO
21CTO
Jul 1, 2017 · Operations

How Ctrip Scales Its Architecture: Ops, Release, and Big Data Insights

This article outlines Ctrip’s evolving architecture—covering its operational backbone, framework components, release system, configuration management, SOA evolution, and the massive UserProfile big‑data platform—offering practical insights from a senior developer on how the company achieves high availability and scalability.

ArchitectureOperationsSOA
0 likes · 12 min read
How Ctrip Scales Its Architecture: Ops, Release, and Big Data Insights
Java High-Performance Architecture
Java High-Performance Architecture
Jun 29, 2017 · Big Data

Master Apache Storm: Core Concepts, Real‑Time Word Count & Call Log Analytics

This tutorial introduces Apache Storm’s fundamental principles and development workflow, providing a PDF guide and source code for two practical examples—real‑time word‑count and call‑record aggregation—while covering its definition, use cases, relationship with Hadoop, core concepts, cluster architecture, and step‑by‑step usage.

Apache Stormbig datacall log analysis
0 likes · 1 min read
Master Apache Storm: Core Concepts, Real‑Time Word Count & Call Log Analytics
Efficient Ops
Efficient Ops
Jun 27, 2017 · Big Data

How a Leading Bank Evolved Its Big Data Platform Architecture

This talk outlines how China’s Guangfa Bank built, refined, and scaled its big‑data platform since 2014, covering data positioning, system architecture optimization, delivery model improvements, team restructuring, and real‑world use cases that demonstrate the platform’s impact on risk control, marketing and operational efficiency.

MicroservicesPlatform Architecturebanking
0 likes · 14 min read
How a Leading Bank Evolved Its Big Data Platform Architecture
21CTO
21CTO
Jun 20, 2017 · Artificial Intelligence

How Toutiao’s AI Powers Personalized News Recommendations

This article examines Toutiao’s rapid rise as a personalized news platform, detailing its AI‑driven recommendation pipeline, web‑crawling infrastructure, similarity‑matrix algorithms, A/B testing, and the role of human moderation in delivering highly targeted content to billions of users.

A/B testingAIbig data
0 likes · 16 min read
How Toutiao’s AI Powers Personalized News Recommendations
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 19, 2017 · Cloud Computing

How Alibaba Built a Cloud‑Native HR System That Cut Costs 100× and Boosted Speed 6×

This article details Alibaba's migration from Oracle PeopleSoft HCM to a self‑developed, cloud‑native eHR platform, describing the technical challenges, phased development using Groovy and MaxCompute, and the resulting six‑fold speed increase, hundred‑fold cost reduction, and enhanced employee experience.

Cloud ComputingGroovyHR system
0 likes · 11 min read
How Alibaba Built a Cloud‑Native HR System That Cut Costs 100× and Boosted Speed 6×
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jun 18, 2017 · Cloud Computing

Inside Alibaba’s Middleware: Career Paths, Tech Stack, and Architecture Challenges

This article explores why Alibaba's middleware is dubbed the architect's cradle, outlines career development routes within the team, details the extensive technology stack, and examines the major technical challenges such as massive data processing, real‑time analytics, and large‑scale deployment during peak events.

Cloud Computingbig datacareer development
0 likes · 25 min read
Inside Alibaba’s Middleware: Career Paths, Tech Stack, and Architecture Challenges
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Jun 16, 2017 · Big Data

How TDH Dominated the TPCx‑HS 10TB Benchmark: Strategies and Results

The article details how StarRocks and Cisco’s joint TPCx‑HS 10TB benchmark placed the TDH platform at the top of the performance ranking, explains the test setup, describes the pre‑ and post‑optimization strategies for TeraGen and TeraSort, and outlines the hardware configuration and key tuning parameters.

HadoopPerformance OptimizationTDH
0 likes · 10 min read
How TDH Dominated the TPCx‑HS 10TB Benchmark: Strategies and Results
Ctrip Technology
Ctrip Technology
Jun 13, 2017 · Operations

Evolution and Architecture of Ctrip's System: Operations, Frameworks, and Big Data

This article presents a comprehensive overview of Ctrip's evolving system architecture, detailing its operational strategies, framework components such as SOA and release systems, and the large‑scale UserProfile big‑data platform, illustrating how each iteration addressed prior challenges while introducing new capabilities.

CtripOperationsbig data
0 likes · 13 min read
Evolution and Architecture of Ctrip's System: Operations, Frameworks, and Big Data
21CTO
21CTO
Jun 9, 2017 · Big Data

From Hadoop to Spark: A Complete Roadmap to Becoming a Big Data Architect

This guide walks beginners through the essential big‑data ecosystem—from understanding Hadoop’s core components and mastering MapReduce, to using Hive, SparkSQL, Kafka, and real‑time frameworks like Storm, while also covering data ingestion, export, scheduling, and introductory machine‑learning techniques.

HiveSparkbig data
0 likes · 20 min read
From Hadoop to Spark: A Complete Roadmap to Becoming a Big Data Architect
Suning Technology
Suning Technology
Jun 9, 2017 · Big Data

How Suning’s AI‑Powered Smart Replenishment Turns Retail from B2C to C2B

Suning’s smart replenishment system showcased at CES Asia 2017 leverages big‑data analytics and machine‑learning models—linear regression, random forest, and XGBoost—to predict sales, optimize inventory across multiple warehouses, and shift retail from traditional B2C to a data‑driven C2B approach.

big datainventory optimizationmachine learning
0 likes · 5 min read
How Suning’s AI‑Powered Smart Replenishment Turns Retail from B2C to C2B
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 8, 2017 · Big Data

Flink Forward 2017: Stream Processing Insights from Alibaba, Uber & Netflix

The article recounts the 2017 Flink Forward conference in San Francisco, highlighting key sessions from DataArtisans, Uber, Netflix and Alibaba, and discusses real‑time stream processing use cases, large‑scale deployments, runtime and TableAPI/SQL improvements, and the growing adoption of Flink in the industry.

Apache FlinkFlinkReal-time Analytics
0 likes · 16 min read
Flink Forward 2017: Stream Processing Insights from Alibaba, Uber & Netflix
StarRing Big Data Open Lab
StarRing Big Data Open Lab
May 27, 2017 · Big Data

Simplify Big Data Governance with Data Lineage & Impact Analysis

Enterprise big‑data platforms face massive scale and complex metadata relationships, but using Transwarp Governor’s data lineage and impact analysis graphs enables precise tracing of data origins, rapid error localization, and prediction of downstream effects, dramatically improving data quality and governance efficiency.

Transwarp Governorbig datadata governance
0 likes · 8 min read
Simplify Big Data Governance with Data Lineage & Impact Analysis
MaGe Linux Operations
MaGe Linux Operations
May 26, 2017 · Big Data

How Big Data Transforms Everyday Life: From Finance to Healthcare

This article explains what big data is, outlines its 5V characteristics, and showcases numerous real‑world applications such as personal finance monitoring, tax fraud detection, healthcare prediction, public opinion tracking, precise marketing, product development, traffic planning, strategic decision‑making, and credit scoring.

ApplicationsData Analyticsbig data
0 likes · 4 min read
How Big Data Transforms Everyday Life: From Finance to Healthcare
Architecture Digest
Architecture Digest
May 25, 2017 · Big Data

Designing Data Warehouse Layers: Principles, Models, and Practical Practices

This article explains why data warehouses should be layered, describes the classic ODS‑DW‑APP model, details each layer’s purpose and implementation techniques, presents an improved layering scheme with dimension and temporary tables, and answers common questions about parallel DWS and DWD processing.

ETLbig datadata architecture
0 likes · 17 min read
Designing Data Warehouse Layers: Principles, Models, and Practical Practices
Alibaba Cloud Developer
Alibaba Cloud Developer
May 25, 2017 · Big Data

How Alibaba’s Blink Engine Redefines Real‑Time Big Data Processing

This article explains how Alibaba’s Blink, built on Apache Flink, transforms batch‑oriented big‑data platforms into a unified, high‑performance real‑time computing engine, detailing its architecture, state management, checkpointing, and successful deployment in e‑commerce, search, recommendation, and online machine‑learning scenarios.

AlibabaFlinkReal-Time Computing
0 likes · 17 min read
How Alibaba’s Blink Engine Redefines Real‑Time Big Data Processing
Alibaba Cloud Developer
Alibaba Cloud Developer
May 20, 2017 · Artificial Intelligence

How Alibaba’s AI‑Driven Information Retrieval Is Shaping E‑Commerce Futures

The second “Frontiers and Future of Information Retrieval” forum, co‑hosted by the Chinese Computer Society, Alibaba and academic committees, showcased how massive, structured e‑commerce data and AI algorithms are revolutionizing search, customer service, and research collaborations across the industry.

AlibabaInformation Retrievalbig data
0 likes · 4 min read
How Alibaba’s AI‑Driven Information Retrieval Is Shaping E‑Commerce Futures
Alibaba Cloud Developer
Alibaba Cloud Developer
May 17, 2017 · Databases

How Alibaba Tackles the Massive Challenges of Time‑Series Data Storage

This article details Alibaba's middleware team's exploration of time‑series data characteristics, real‑world monitoring scenarios, the limitations of traditional databases, and the evolution of their custom HiTSDB solution that combines inverted indexing, high‑compression algorithms, and distributed aggregation to meet massive write and query demands.

AlibabaHiTSDBStorage
0 likes · 25 min read
How Alibaba Tackles the Massive Challenges of Time‑Series Data Storage
MaGe Linux Operations
MaGe Linux Operations
May 17, 2017 · Big Data

How Big Data Turns Raw Information into Resource Optimization

The article explains that the ultimate value of big data lies in optimizing resource allocation by first crowdsourcing massive data, then fully mining it to uncover truth, and finally using those insights across industries such as transportation, advertising, finance, and more.

Resource Optimizationbig datacrowdsourcing
0 likes · 7 min read
How Big Data Turns Raw Information into Resource Optimization