Tagged articles

big data

3804 articles · Page 36 of 39
Baidu Waimai Technology Team
Baidu Waimai Technology Team
May 16, 2017 · Big Data

Analysis of OLTP/OLAP Integrated Solutions: Apache Phoenix, Apache Trafodion, and Splice Machine

This article examines the convergence of OLTP and OLAP by introducing Apache Phoenix, Apache Trafodion, and Splice Machine, compares their technical features, and describes how Baidu Waimai adopted a Phoenix‑based solution to address scalability and performance challenges in its operational data store.

Apache PhoenixApache TrafodionOLAP
0 likes · 12 min read
Analysis of OLTP/OLAP Integrated Solutions: Apache Phoenix, Apache Trafodion, and Splice Machine
Qunar Tech Salon
Qunar Tech Salon
May 16, 2017 · Artificial Intelligence

Personalized Recommendation Systems: Applications, User Profiling, Algorithms, and Optimization

This article presents a comprehensive overview of personalized recommendation systems, covering their application scenarios and value, user profiling, core algorithms such as content‑based and collaborative filtering, system architecture, performance and effect optimization techniques, and practical Q&A insights.

AIbig datacollaborative filtering
0 likes · 18 min read
Personalized Recommendation Systems: Applications, User Profiling, Algorithms, and Optimization
MaGe Linux Operations
MaGe Linux Operations
May 15, 2017 · Databases

Top 10 Must‑Know Data Storage Tools for Java Developers

Facing ever‑growing complexity, Java developers can streamline their projects by mastering a curated list of essential data storage and processing tools—including MongoDB, Elasticsearch, Cassandra, Redis, Hazelcast, EHCache, Hadoop, Solr, Spark, and Memcached—each offering distinct strengths for modern big‑data applications.

Data ProcessingJavaNoSQL
0 likes · 8 min read
Top 10 Must‑Know Data Storage Tools for Java Developers
ITPUB
ITPUB
May 8, 2017 · Big Data

Master Spark Performance: Practical Tuning Tips and Real‑World Examples

This article explains essential Spark concepts, illustrates common performance bottlenecks, and provides concrete tuning strategies for memory, CPU, serialization, data locality, file I/O, and shuffle reduction, backed by real‑world examples and visual metrics.

CPU OptimizationConfigurationSpark
0 likes · 19 min read
Master Spark Performance: Practical Tuning Tips and Real‑World Examples
Architects' Tech Alliance
Architects' Tech Alliance
May 7, 2017 · Big Data

Building a Complete Big Data Platform: From Hadoop Basics to Real‑Time Analytics

This guide walks beginners through the entire big‑data ecosystem—explaining the 4V characteristics, listing essential open‑source components, teaching Hadoop setup, Hive and SparkSQL usage, data ingestion with Sqoop, Flume and Kafka, task scheduling with Oozie, and real‑time processing with Storm and Spark Streaming.

Data PipelineHadoopHive
0 likes · 20 min read
Building a Complete Big Data Platform: From Hadoop Basics to Real‑Time Analytics
MaGe Linux Operations
MaGe Linux Operations
May 7, 2017 · Artificial Intelligence

Big Data & Machine Learning: Core Definitions and Essential Algorithms

This article explains what big data and machine learning are, their interrelationship, various big‑data analysis approaches, core machine‑learning concepts, and details ten fundamental algorithms—including regression, neural networks, SVM, clustering, dimensionality reduction, and recommendation—while highlighting their roles in modern data‑driven applications.

ClusteringSVMbig data
0 likes · 24 min read
Big Data & Machine Learning: Core Definitions and Essential Algorithms
MaGe Linux Operations
MaGe Linux Operations
May 4, 2017 · Big Data

How to Process 100GB Logs and Massive Datasets with Hash Partitioning and Bloom Filters

This article explains the definition and 4V characteristics of big data and presents practical algorithms—including hash partitioning, min‑heap top‑K selection, bitmap extensions, and Bloom filter techniques—to efficiently handle ultra‑large log files, integer sets, and keyword searches within strict memory limits.

BitmapBloom FilterHash Partitioning
0 likes · 12 min read
How to Process 100GB Logs and Massive Datasets with Hash Partitioning and Bloom Filters
Efficient Ops
Efficient Ops
May 3, 2017 · Operations

How Tencent Scales NBA Live Streams to Millions: Behind the Tech and Operations

This article details Tencent's large‑scale live streaming architecture for NBA games, covering the rapid growth of live video, key technical features, network transmission challenges, multi‑angle production, CDN deployment, monitoring, big‑data processing, and strategies for ensuring low latency and high reliability for millions of concurrent viewers.

CDNOperationsbig data
0 likes · 25 min read
How Tencent Scales NBA Live Streams to Millions: Behind the Tech and Operations
Baidu Waimai Technology Team
Baidu Waimai Technology Team
Apr 28, 2017 · Big Data

Recap of Baidu Waimai Tech Team’s “Code Talk” Session on Data Platform Architecture and Big Data Practices

The article summarizes Baidu Waimai’s recent “Code Talk” event, highlighting the speaker’s overview of the company’s big‑data platform evolution, its technical architecture, practical challenges such as data security and accuracy, and a lively Q&A covering storm, high availability, and metric management.

Baidu WaimaiTech Talkbig data
0 likes · 6 min read
Recap of Baidu Waimai Tech Team’s “Code Talk” Session on Data Platform Architecture and Big Data Practices
Architects' Tech Alliance
Architects' Tech Alliance
Apr 27, 2017 · Big Data

Curated List of Big Data Learning Resources from w3cschool

This article presents a comprehensive, Chinese‑language collection of big‑data resources—including relational databases, distributed file systems, key‑value stores, distributed programming tools, file data models, and key‑map frameworks—compiled by w3cschool to help programmers deepen their understanding of big data technologies.

LearningResourcesbig data
0 likes · 6 min read
Curated List of Big Data Learning Resources from w3cschool
Architecture Digest
Architecture Digest
Apr 24, 2017 · Big Data

Understanding and Solving Data Skew in Hadoop and Spark

This article explains what data skew is, why it occurs in large‑scale Hadoop and Spark jobs, illustrates typical symptoms, and presents practical strategies—including business‑level adjustments, code tweaks, and platform‑specific tuning—to mitigate and resolve skew in big‑data processing.

Data SkewHadoopSpark
0 likes · 11 min read
Understanding and Solving Data Skew in Hadoop and Spark
21CTO
21CTO
Apr 21, 2017 · R&D Management

How to Turn Technical Experience into Personal Value: Lessons from Outsourcing to Big Data

The author shares a candid journey from low‑paid outsourcing coding to senior roles in design, analysis, and big‑data architecture, revealing how understanding value networks, leveraging cloud and data trends, and expanding beyond pure coding can dramatically increase a technologist’s personal and market value.

Cloud ComputingR&D managementbig data
0 likes · 34 min read
How to Turn Technical Experience into Personal Value: Lessons from Outsourcing to Big Data
Alibaba Cloud Developer
Alibaba Cloud Developer
Apr 21, 2017 · Big Data

How Alibaba Tackles Real-Time Stream and Graph Computing at Scale

In his ASPLOS keynote, Alibaba’s Vice President Zhou Jingren detailed the company’s large‑scale stream and graph computing platforms, highlighting fault‑tolerance innovations, real‑time data challenges, and upcoming advances in graph analytics and massive machine‑learning workloads.

AIAlibababig data
0 likes · 7 min read
How Alibaba Tackles Real-Time Stream and Graph Computing at Scale
Baidu Waimai Technology Team
Baidu Waimai Technology Team
Apr 20, 2017 · Databases

Greenplum (GPDB) Architecture, Features, and Operational Tools Overview

This article explains Greenplum's MPP architecture, master‑segment design, high‑availability, interconnect network, rich management tools, parallel query planning, data loading techniques, and additional capabilities such as LDAP authentication and resource queues, demonstrating why it is a strong next‑generation big‑data query engine.

GreenplumMPPOperations
0 likes · 15 min read
Greenplum (GPDB) Architecture, Features, and Operational Tools Overview
Baidu Waimai Technology Team
Baidu Waimai Technology Team
Apr 18, 2017 · Industry Insights

Baidu Waimai’s Cloud Migration, AI Logistics, and Architecture – QCon 2017

At QCon Beijing 2017, three senior Baidu Waimai engineers detailed the company’s year‑long migration from IDC to cloud using custom operation platforms, described the AI‑driven, data‑rich logistics scheduling system that outperforms manual dispatch, and shared architectural evolutions that enabled rapid, zero‑downtime scaling of the fast‑growing delivery business.

AI logisticsOperationsarchitecture scaling
0 likes · 5 min read
Baidu Waimai’s Cloud Migration, AI Logistics, and Architecture – QCon 2017
Meituan Technology Team
Meituan Technology Team
Apr 14, 2017 · Big Data

Practical Experience of HDFS Federation at Meituan: Challenges, Improvements, and Automation

Meituan‑Dianping migrated its 2,000‑node HDFS cluster to Federation by fixing ViewFs compatibility, simplifying mount points, leveraging FastCopy for massive data moves, improving token handling, and automating split‑workflow steps, thereby overcoming single‑NameNode bottlenecks and providing a practical blueprint for large‑scale Hadoop deployments.

FastCopyFederationHDFS
0 likes · 22 min read
Practical Experience of HDFS Federation at Meituan: Challenges, Improvements, and Automation
MaGe Linux Operations
MaGe Linux Operations
Apr 13, 2017 · Big Data

How to Choose the Right Language for Your Big Data Project

This article compares R, Python, Scala, and Java for big‑data projects, outlining each language’s strengths and weaknesses, and offers guidance on selecting the most suitable language based on project requirements, team expertise, and production needs.

JavaLanguage SelectionPython
0 likes · 8 min read
How to Choose the Right Language for Your Big Data Project
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Apr 9, 2017 · Fundamentals

Understanding Bloom Filters: Fast, Space-Efficient Membership Tests

Bloom filters are highly space-efficient probabilistic data structures that quickly test set membership using multiple hash functions, guaranteeing no false negatives while allowing a small false positive rate, making them ideal for large-scale applications such as email blacklists and massive URL deduplication.

Bloom Filterbig datamembership testing
0 likes · 5 min read
Understanding Bloom Filters: Fast, Space-Efficient Membership Tests
21CTO
21CTO
Apr 4, 2017 · Artificial Intelligence

How Vipshop Evolved Its Real-Time Personalized Recommendation Engine

This article recounts Wu Guanlin’s presentation on the evolution of Vipshop’s personalized recommendation system, detailing the technical challenges of real‑time predictions, the three generations of architecture, the four‑stage recommendation engine, and the VRE platform’s design for scalability and low latency.

System Architecturebig datamachine learning
0 likes · 10 min read
How Vipshop Evolved Its Real-Time Personalized Recommendation Engine
Meituan Technology Team
Meituan Technology Team
Mar 24, 2017 · Artificial Intelligence

Tourism Recommendation System: Strategy Iterations, Architecture, and Future Challenges

The article outlines Meituan‑Dianping’s tourism recommendation system, detailing its evolution from simple hot‑sale recall to sophisticated decay‑based, GPS‑aware, collaborative filtering and XGBoost reranking pipelines, the four‑layer architecture supporting dozens of travel scenarios, and future plans to broaden recall, adopt deep models, and expand multimodal travel recommendations.

ArchitectureTourismbig data
0 likes · 26 min read
Tourism Recommendation System: Strategy Iterations, Architecture, and Future Challenges
Tongcheng Travel Technology Center
Tongcheng Travel Technology Center
Mar 24, 2017 · Operations

Evolution of Tongcheng Log System Architecture

The article chronicles the development of Tongcheng's centralized log system from early file‑based logging through a MongoDB‑based solution to the current multi‑layer architecture using Flume, Elasticsearch, and Hadoop, highlighting design decisions, challenges, and future improvement plans.

Flumebig datalog system
0 likes · 7 min read
Evolution of Tongcheng Log System Architecture
ITPUB
ITPUB
Mar 22, 2017 · Big Data

Why Spark Beats MapReduce: The RDD Story and Spark SQL Evolution

This article walks through Spark’s origins, its core RDD concept, how it improves on Hadoop’s MapReduce, the role of in‑memory processing, functional programming support, and the emergence of Spark SQL with DataFrames and the Catalyst optimizer.

DataFrameDistributed ComputingMapReduce
0 likes · 25 min read
Why Spark Beats MapReduce: The RDD Story and Spark SQL Evolution
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Mar 21, 2017 · Big Data

How Real-Time Data Streaming Is Transforming Industries Today

This article explains how real‑time data streaming turns massive, continuously growing datasets into actionable insights across finance, energy, and e‑commerce, showcasing early adopters like ConocoPhillips and DHL while urging businesses to rethink models for the next wave of data management.

Data StreamingReal-time Analyticsbig data
0 likes · 7 min read
How Real-Time Data Streaming Is Transforming Industries Today
Qunar Tech Salon
Qunar Tech Salon
Mar 12, 2017 · Big Data

Essential Skills and Career Paths for Data Professionals: From Big Data Platforms to AI

The article outlines the key competencies, responsibilities, and career development advice for data professionals across the entire data stack—from building big‑data platforms and data warehouses to visualization, analysis, algorithm engineering, and deep‑learning applications—emphasizing the importance of creating business value with data.

Data AnalystData Visualizationbig data
0 likes · 15 min read
Essential Skills and Career Paths for Data Professionals: From Big Data Platforms to AI
21CTO
21CTO
Mar 10, 2017 · Big Data

Inside Tencent Analytics: How TA Handles TB‑Scale Real‑Time Web Data

Tencent Analytics (TA) is a free web analytics platform that processes terabytes of daily data in real time, using a custom architecture featuring JavaScript collection, event streaming, in‑memory computation, and NoSQL storage with Redis and LevelDB, offering site owners instant insights and high availability.

LevelDBStreamingWeb Analytics
0 likes · 12 min read
Inside Tencent Analytics: How TA Handles TB‑Scale Real‑Time Web Data
Efficient Ops
Efficient Ops
Mar 7, 2017 · Big Data

How Tencent Scaled Its TDW to 8,800 Nodes and Mastered Cross-City Data Migration

Tencent’s senior engineer explains how the TDW (Tencent Distributed Data Warehouse) grew from a few hundred to thousands of nodes, the challenges of cross‑city migration, and the modeling, relationship‑chain, dual‑write tables, and platform strategies they built to ensure seamless, low‑impact data and task migration.

Cloud OperationsTDWbig data
0 likes · 26 min read
How Tencent Scaled Its TDW to 8,800 Nodes and Mastered Cross-City Data Migration
Alibaba Cloud Developer
Alibaba Cloud Developer
Mar 7, 2017 · Big Data

Unified Data Platforms: How UMENG+ Redefines Big Data Strategy

The article explores the evolution of big‑data applications in China, from Oracle’s trend report and the concept of "omni‑domain data" to UMENG+’s technical architecture, unified tech stack, AI integration, and future directions for delivering real customer value.

Data AnalyticsData Integrationbig data
0 likes · 12 min read
Unified Data Platforms: How UMENG+ Redefines Big Data Strategy
Efficient Ops
Efficient Ops
Mar 6, 2017 · Operations

Tencent Game Ops: Turning Service Delivery into Smart, Automated Microservices

This article details how Tencent's game operations team redefined operational services, introduced micro‑service architecture, applied big‑data driven recommendations, and built intelligent, automated pipelines for server opening, merging, version releases, and download services, achieving significant efficiency and cost gains.

Cloud NativeMicroservicesautomation
0 likes · 26 min read
Tencent Game Ops: Turning Service Delivery into Smart, Automated Microservices
Meituan Technology Team
Meituan Technology Team
Mar 2, 2017 · Big Data

Meituan Waimai Feature Archive Platform: Architecture, Tag System, and Data Processing

Meituan Waimai’s Feature Archive platform processes billions of daily orders by managing ~200 user and 400 merchant tags through a three‑layer architecture—Hive, Elasticsearch, HBase, and MySQL—offering visual tag selection, instant self‑service queries, full data extraction, and a predicate‑logic query language, while supporting future extensibility.

Data PipelineElasticsearchHBase
0 likes · 14 min read
Meituan Waimai Feature Archive Platform: Architecture, Tag System, and Data Processing
Qunar Tech Salon
Qunar Tech Salon
Mar 1, 2017 · Big Data

Building Prism: Qunar’s Real‑Time Data Platform and DevOps Journey

The article describes how Qunar designed and evolved its Prism real‑time data platform—leveraging ELK, Kafka, Spark, Docker, and Mesos—to improve data collection, monitoring, and analysis, reduce deployment time, and support scalable DevOps operations across the company.

ELKKafkaReal-time Data
0 likes · 11 min read
Building Prism: Qunar’s Real‑Time Data Platform and DevOps Journey
AntTech
AntTech
Feb 28, 2017 · Artificial Intelligence

Key Computing Capabilities Driving the Evolution of Digital Financial Services

The talk outlines nine essential computing capabilities—transaction processing, system robustness, connectivity, decision-making, data insight, intelligent services, biometric authentication, blockchain trust, and immersive integration—that have transformed Ant Financial over the past decade and outlines the challenges and strategies for the next ten years.

Artificial IntelligenceCloud Computingbig data
0 likes · 16 min read
Key Computing Capabilities Driving the Evolution of Digital Financial Services
Architecture Digest
Architecture Digest
Feb 28, 2017 · Big Data

Architecture and Real‑Time Processing Design of Tencent Analytics (TA)

This article explains the architecture, real‑time computation framework, and storage solutions of Tencent Analytics, detailing how massive TB‑level web‑traffic data are collected via JavaScript, processed in memory‑centric streaming components, and stored using Redis and LevelDB to achieve second‑level updates.

LevelDBNoSQLRedis
0 likes · 13 min read
Architecture and Real‑Time Processing Design of Tencent Analytics (TA)
Nightwalker Tech
Nightwalker Tech
Feb 27, 2017 · Big Data

Community Discussion on Learning Paths, Tools, and Applications in Big Data

A diverse group of practitioners share recommendations for books, technologies, real‑world use cases, and practical challenges when learning and applying big‑data processing, covering Hadoop, Spark, data visualization, ETL, and the relationship between data, algorithms, and business value.

Data AnalysisHadoopbig data
0 likes · 16 min read
Community Discussion on Learning Paths, Tools, and Applications in Big Data
Qunar Tech Salon
Qunar Tech Salon
Feb 26, 2017 · Big Data

Comparative Analysis of Big Data Storage and Query Solutions

This article reviews major big‑data storage and query architectures—including HBase, Dremel/Parquet, pre‑aggregation systems, Lucene, and the custom Tindex solution—evaluating their strengths, weaknesses, and suitability for real‑time, high‑volume analytical workloads.

HBaseLuceneParquet
0 likes · 20 min read
Comparative Analysis of Big Data Storage and Query Solutions
Efficient Ops
Efficient Ops
Feb 26, 2017 · Operations

How Alibaba Scales Massive Data Platforms: Lessons in Automated Operations

This article explores the challenges of operating Alibaba's large‑scale data platforms, describes the automation platform built to address them, and shares data‑driven, fine‑grained operational practices that enable stable, efficient, and cost‑effective service delivery.

OperationsPlatformautomation
0 likes · 22 min read
How Alibaba Scales Massive Data Platforms: Lessons in Automated Operations
Architects' Tech Alliance
Architects' Tech Alliance
Feb 23, 2017 · Big Data

Analysis of 2017 Chinese Spring Festival Migration Trends Using Big Data

Using big‑data visualizations, this article examines the 2017 Chinese Spring Festival migration, revealing early peak travel, the top twenty outbound cities accounting for over 40 % of movement, net flow patterns, and demographic profiles such as age, education and industry distribution across major urban centers.

ChinaPopulationSpring Festival
0 likes · 5 min read
Analysis of 2017 Chinese Spring Festival Migration Trends Using Big Data
Qunar Tech Salon
Qunar Tech Salon
Feb 22, 2017 · Big Data

Understanding Ctrip Flight Ticket Tracking System (UBT) and Its Key Metrics

This article explains Ctrip's flight ticket tracking framework (UBT), detailing client‑side and server‑side event collection methods, the purpose and trade‑offs of each tracking type, metric definitions, data association challenges, common pitfalls, and best practices for reliable data‑driven analysis.

AnalyticsCtripbig data
0 likes · 20 min read
Understanding Ctrip Flight Ticket Tracking System (UBT) and Its Key Metrics
Alibaba Cloud Developer
Alibaba Cloud Developer
Feb 22, 2017 · Artificial Intelligence

How Alibaba’s AI Powers Real‑Time Customer Segmentation and Personalized Shopping

This article explains how Alibaba leverages AI, big‑data analytics, and advanced recommendation algorithms to enable real‑time visitor clustering, personalized storefronts, and tailored content across its Customer Operation Platform, Double 11 promotion pages, QianNiu headlines, and service market, delivering significant conversion and engagement gains.

AIbig datae-commerce
0 likes · 18 min read
How Alibaba’s AI Powers Real‑Time Customer Segmentation and Personalized Shopping
Nightwalker Tech
Nightwalker Tech
Feb 20, 2017 · Backend Development

Career Development and Technology Trends for PHP Engineers

The discussion explores how PHP engineers can advance their careers by embracing new technologies such as Go, Python, big data, AI, and cloud computing, while also emphasizing soft‑skill growth, project management, and strategic decision‑making based on business trends and personal goals.

Artificial IntelligenceBackendbig data
0 likes · 9 min read
Career Development and Technology Trends for PHP Engineers
Meituan Technology Team
Meituan Technology Team
Feb 17, 2017 · Big Data

User Profiling and Machine Learning Practices for Food Delivery O2O Platforms

Meituan Delivery’s rapid expansion across multiple categories relies on detailed user profiling and machine‑learning models—such as high‑potential customer prediction, churn risk regression and Cox survival analysis—to personalize acquisition, retention, and scenario‑based cross‑selling, while addressing sparse behavior, unstructured data, and geographic context challenges.

O2Obig datachurn prediction
0 likes · 13 min read
User Profiling and Machine Learning Practices for Food Delivery O2O Platforms
21CTO
21CTO
Feb 15, 2017 · Fundamentals

How Twitter Evolved Its Search Engine: From MySQL to Earlybird and Beyond

This article explains the fundamentals of search engine architecture, covering text collection, indexing, ranking and evaluation, and then traces Twitter's internal search evolution from MySQL full‑text search to the Earlybird index server, Blender aggregation, and smart memory‑SSD strategies.

Information RetrievalTwitterbig data
0 likes · 8 min read
How Twitter Evolved Its Search Engine: From MySQL to Earlybird and Beyond
Architecture Digest
Architecture Digest
Feb 11, 2017 · Big Data

LeKe Sports Big Data Platform Evolution: From Early ETL Reporting to 2.0 Streaming Architecture

The article describes how LeKe Sports built and continuously upgraded its Hadoop‑based big data platform—from a manual ETL‑to‑Elasticsearch reporting system to a 2.0 architecture featuring Spark Streaming, SQL‑based query layers, Elasticsearch indexing, and cloud‑native storage and backup solutions—to meet rapidly growing PB‑scale data demands.

ETLHadoopSpark
0 likes · 5 min read
LeKe Sports Big Data Platform Evolution: From Early ETL Reporting to 2.0 Streaming Architecture
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Feb 7, 2017 · Big Data

What’s New in Apache CarbonData 1.0.0? 80+ Features Boost Big Data Performance

Apache CarbonData 1.0.0, now an Apache incubating project, adds over 80 new features and bug fixes—including a new data loading solution, Spark 2.1 integration, update/delete SQL support, adaptive compression for numeric types, B‑Tree LRU cache, V2 format for faster first‑query performance, vectorized reader, bucket‑table joins, off‑heap memory, single‑pass loading, and pre‑generated dictionaries—aimed at delivering faster, more flexible, and efficient columnar storage for big‑data workloads.

Apache CarbonDataSpark Integrationbig data
0 likes · 8 min read
What’s New in Apache CarbonData 1.0.0? 80+ Features Boost Big Data Performance
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Jan 24, 2017 · Big Data

Why Hadoop Remains the Backbone of Big Data: Core Modules, Tools, and Trends

This article provides a comprehensive overview of Hadoop as the leading open‑source platform for big‑data processing, detailing its core components HDFS and MapReduce, the evolution to Hadoop 2.0/YARN, and the extensive ecosystem of tools and commercial solutions that enable scalable storage, analysis, and machine‑learning on massive data sets.

Data ProcessingDistributed ComputingHDFS
0 likes · 18 min read
Why Hadoop Remains the Backbone of Big Data: Core Modules, Tools, and Trends
21CTO
21CTO
Jan 18, 2017 · Big Data

Build a Lightweight, High‑Availability Real‑Time Stream Processing System

Learn how to construct a simple, high‑availability real‑time stream processing platform using lightweight components such as Kafka, Zookeeper, Thrift/Avro, and optional storage like MongoDB or Elasticsearch, offering a practical alternative to heavyweight frameworks like Storm and Spark Streaming for small‑to‑medium enterprises.

Kafkabig datalightweight architecture
0 likes · 5 min read
Build a Lightweight, High‑Availability Real‑Time Stream Processing System
dbaplus Community
dbaplus Community
Jan 16, 2017 · Backend Development

Scaling a FinTech Platform to $100B Transactions with Four Overhauls

Over three years, a small fintech company transformed its platform from a single‑server PHP/Java stack to a micro‑service‑based Spring Cloud architecture, undergoing four major upgrades that introduced distributed systems, SOA governance, big‑data pipelines, MongoDB replication, Redis caching, and open‑source tools, enabling transaction volumes exceeding one hundred billion.

ArchitectureMicroservicesSpring Cloud
1 likes · 15 min read
Scaling a FinTech Platform to $100B Transactions with Four Overhauls
Alibaba Cloud Developer
Alibaba Cloud Developer
Jan 11, 2017 · R&D Management

How Taobao’s Beehive Platform Powers Content‑Driven Shopping During Double 11

The article explains how Taobao’s content‑centric strategy, embodied in the Beehive platform, builds an end‑to‑end content chain—from creator tools and health scoring to personalized distribution and commerce mechanisms—enabling massive, efficient content production and monetization during the Double 11 shopping festival.

Taobaobig datacontent platform
0 likes · 17 min read
How Taobao’s Beehive Platform Powers Content‑Driven Shopping During Double 11
Alibaba Cloud Developer
Alibaba Cloud Developer
Jan 9, 2017 · Big Data

How Alibaba Scaled Real‑Time Data Processing for Double 11: Architecture & Lessons

This article details Alibaba's real‑time computing architecture for the 2016 Double 11 event, covering background, core components such as DRC, TT, Galaxy, OTS, XTool and OneService, and explains optimization techniques, fault‑tolerance strategies, stress‑testing practices, and future upgrade plans to handle massive streaming data workloads.

ArchitecturePerformance OptimizationReal-Time Computing
0 likes · 14 min read
How Alibaba Scaled Real‑Time Data Processing for Double 11: Architecture & Lessons
dbaplus Community
dbaplus Community
Dec 26, 2016 · Big Data

Why Data Lakes Are Redefining Enterprise Data Architecture

This article explains the origins, core features, logical architecture, and advantages of data lakes, contrasts them with traditional data warehouses, outlines a modern data architecture that combines lakes and warehouses, and introduces the DCE intelligent data lake platform with practical Q&A.

Cloud ComputingData Lakebig data
0 likes · 14 min read
Why Data Lakes Are Redefining Enterprise Data Architecture
Tencent Cloud Developer
Tencent Cloud Developer
Dec 23, 2016 · Databases

Analysis of HBase Write-Ahead Log (WAL) Mechanism and Source Code Call Chain

The article explains HBase’s write‑ahead‑log architecture, detailing how client put/delete requests travel through RPC to the RegionServer, are processed by MultiRowMutationService, written to the WAL via FSHLog.append and sync, and finally stored in MemStore, while describing durability options and the underlying source‑code call chain.

HBaseJavaWAL
0 likes · 10 min read
Analysis of HBase Write-Ahead Log (WAL) Mechanism and Source Code Call Chain
Hulu Beijing
Hulu Beijing
Dec 20, 2016 · Big Data

How Hulu Supercharges OLAP Queries with CarbonData: Real‑World Optimizations

This article describes Hulu’s real‑world OLAP query optimization, covering the fundamentals of OLAP, comparisons of row‑ and column‑based storage formats, detailed indexing mechanisms of Parquet, ORC and CarbonData, and the specific schema, shuffle, block size, speculation and GC tuning techniques that enabled CarbonData to dramatically accelerate wide‑table queries on SparkSQL.

CarbonDataOLAPSparkSQL
0 likes · 17 min read
How Hulu Supercharges OLAP Queries with CarbonData: Real‑World Optimizations
Meituan Technology Team
Meituan Technology Team
Dec 9, 2016 · Big Data

Memory Usage Analysis of HDFS NameNode Core Data Structures

The article quantitatively breaks down HDFS NameNode memory consumption, showing that the Namespace tree and BlocksMap together dominate heap usage (≈53 GB in large clusters), provides detailed per‑object size estimates for NetworkTopology, INode and block structures, and proposes a simple formula to predict total heap requirements and tuning recommendations.

HDFSNameNodePerformance Optimization
0 likes · 13 min read
Memory Usage Analysis of HDFS NameNode Core Data Structures
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 8, 2016 · Artificial Intelligence

How AI Powers Data‑Driven Merchant Success

In this Alibaba Tech Forum talk, senior expert Wei Hu explains how machine learning and big‑data technologies are used to empower merchants with personalized storefronts, intelligent posters, and AI‑driven headlines, boosting their efficiency and sales performance.

AIAlibababig data
0 likes · 2 min read
How AI Powers Data‑Driven Merchant Success
Ctrip Technology
Ctrip Technology
Dec 2, 2016 · Big Data

Design and Architecture of Ctrip's Aegis Risk Control System

This article presents a comprehensive overview of Ctrip's Aegis risk control system, detailing its modular architecture, rule engine, data service layer, Chloro analytics platform, and future directions, while highlighting the use of streaming, big‑data processing, and machine‑learning models for real‑time fraud detection.

Rule Enginebig datamachine learning
0 likes · 13 min read
Design and Architecture of Ctrip's Aegis Risk Control System
Meitu Technology
Meitu Technology
Dec 1, 2016 · Big Data

Multi-dimensional Analysis Platform Based on User Portrait Data

Tencent's Glacier multi‑dimensional analysis platform combines massive user‑portrait tags with routine analytical reports, delivering fast, accurate real‑time queries across countless dimensional combinations, enabling analysts and operators to perform targeted operations and insights as product data continuously evolves.

GlacierMulti-dimensional AnalysisTencent
0 likes · 1 min read
Multi-dimensional Analysis Platform Based on User Portrait Data
Meitu Technology
Meitu Technology
Dec 1, 2016 · Big Data

Meitu Internet Technology Salon: Big Data Architecture Evolution and Practice, and Tencent Multi‑Dimensional Analysis Platform

At Meitu’s third Internet Technology Salon in Xiamen on November 26 2016, over 150 senior engineers heard Meitu’s Lu Rongbin detail the company’s progression from simple rsync scripts to a scalable mobile data and open statistical platform, while Tencent’s Zhao Shiyuan showcased the Glacier multi‑dimensional analysis system for fast, tag‑driven queries, underscoring collaborative technical exchange in South China.

AnalyticsMeituTencent
0 likes · 6 min read
Meitu Internet Technology Salon: Big Data Architecture Evolution and Practice, and Tencent Multi‑Dimensional Analysis Platform
Alibaba Cloud Developer
Alibaba Cloud Developer
Nov 30, 2016 · Big Data

How Alibaba’s Double 11 Turned Big Data into a Global E‑Commerce Game‑Changer

MIT Technology Review reports that Alibaba’s 2022 Double 11 shopping festival set new e‑commerce records while showcasing the company’s advanced big‑data, AI, and cloud‑computing technologies, highlighting massive transaction volumes, high‑quality data processing, robust security measures, and the strategic push toward global digital infrastructure.

Cloud Computingbig datadata security
0 likes · 11 min read
How Alibaba’s Double 11 Turned Big Data into a Global E‑Commerce Game‑Changer
Architects' Tech Alliance
Architects' Tech Alliance
Nov 28, 2016 · Big Data

User Profiling: Concepts, Stages, and Data Modeling Methods

This article explains the concept of user profiling, outlines its four-stage construction process, discusses the significance of tagging users, and details practical data modeling techniques—including static and dynamic data sources, weight calculations, and real‑world examples—aimed at improving precision marketing and recommendation systems.

behavior analysisbig datadata modeling
0 likes · 44 min read
User Profiling: Concepts, Stages, and Data Modeling Methods
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Nov 20, 2016 · Big Data

Alibaba’s Big Data Applications in Urban Governance and Social Risk Prevention

The article describes how Alibaba leverages big data, cloud computing and AI through its “City Brain” project and security platforms to improve urban traffic management, public safety, anti‑fraud measures and e‑commerce risk control, illustrating the transformative impact of data‑driven technologies on modern social governance.

AICloud Computingbig data
0 likes · 11 min read
Alibaba’s Big Data Applications in Urban Governance and Social Risk Prevention
Architecture Digest
Architecture Digest
Nov 11, 2016 · Backend Development

High‑Availability Architecture Sessions at the China Software Developers Conference (Nov 18‑20)

The conference featured a series of high‑availability architecture talks covering performance‑driven design, RPC framework resilience, big‑data platform evolution, MySQL cluster consistency, and cloud infrastructure best practices, presented by experts from 58.com, Alibaba, Tencent, Baidu, and others.

Cloud ComputingRPCbackend architecture
0 likes · 10 min read
High‑Availability Architecture Sessions at the China Software Developers Conference (Nov 18‑20)
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Nov 11, 2016 · Big Data

Why SQL Still Rules Big Data—and How NoSQL & NewSQL Fit In

The article explores the evolution of data processing from Hadoop and Spark to modern SQL, NoSQL, and NewSQL solutions, comparing their architectures, performance trade‑offs, and use‑cases, while illustrating concepts with examples like MapReduce, Hive, Impala, and streaming platforms such as Storm.

HadoopNewSQLNoSQL
0 likes · 14 min read
Why SQL Still Rules Big Data—and How NoSQL & NewSQL Fit In
Architects' Tech Alliance
Architects' Tech Alliance
Nov 8, 2016 · Cloud Computing

12 Notable Data Storage Startups to Watch in 2016

Amid rising data‑storage complexity, twelve innovative startups emerged in 2015‑2016, leveraging flash, disk, and cloud technologies to improve data mobility and management across hierarchical storage tiers, offering solutions ranging from cloud‑native storage networks to SAN arrays and virtualization platforms.

Cloud StorageSANbig data
0 likes · 15 min read
12 Notable Data Storage Startups to Watch in 2016
MaGe Linux Operations
MaGe Linux Operations
Nov 7, 2016 · Big Data

How HDFS Achieves Low Cost, High Reliability, and Fault Tolerance

This article explains how HDFS, inspired by Google’s GFS, provides a low‑cost, highly reliable, fault‑tolerant, and high‑performance distributed file system for big‑data workloads by using replication, standby NameNodes, block storage, rack awareness, and compute‑close‑to‑data strategies.

HDFSbig datadata replication
0 likes · 7 min read
How HDFS Achieves Low Cost, High Reliability, and Fault Tolerance
Architecture Digest
Architecture Digest
Nov 6, 2016 · Big Data

Evolution of Taobao’s Big Data Platform: From RAC to MaxCompute

The article chronicles Taobao’s 13‑year evolution of its big data platform, detailing three phases—from a single‑node Oracle setup and the Tianwang scheduler, through a Hadoop‑based “Cloud Ladder 1” architecture with real‑time analytics, to the current MaxCompute/ODPS era with cross‑region projects and advanced data services.

HadoopMaxComputeTaobao
0 likes · 11 min read
Evolution of Taobao’s Big Data Platform: From RAC to MaxCompute
Architects' Tech Alliance
Architects' Tech Alliance
Nov 4, 2016 · Big Data

The Seven Camps of the Global Big Data Ecosystem

The article outlines how mobile Internet merges the data‑driven society with the physical world to create a new big‑data architecture and describes the seven distinct camps—Infrastructure, Analytics, Applications, Cross‑Domain Architecture, Open‑Source, Data Sources & APIs, and Incubator & Training—that together form a comprehensive end‑to‑end big‑data solution ecosystem.

APIAnalyticsApplications
0 likes · 3 min read
The Seven Camps of the Global Big Data Ecosystem
Meituan Technology Team
Meituan Technology Team
Nov 4, 2016 · Big Data

Design and Implementation of a Low-Latency App Exception Monitoring Platform Using Spark Streaming, Kafka, and Elasticsearch

The paper presents a production‑grade, low‑cost mobile‑app exception monitoring platform built on Spark Streaming, Kafka, and Elasticsearch that achieves high availability through exactly‑once processing and checkpointing, minute‑level latency by decoupling raw and symbolized logs, high throughput via reservoir sampling, and dynamic scalability without code changes.

ElasticsearchException MonitoringKafka
0 likes · 11 min read
Design and Implementation of a Low-Latency App Exception Monitoring Platform Using Spark Streaming, Kafka, and Elasticsearch
Architects' Tech Alliance
Architects' Tech Alliance
Nov 3, 2016 · Industry Insights

Scaling Billion‑Level Ads: Architecture Lessons from Sogou’s Senior Engineer

In this interview, Sogou architect Liu Jian shares how his team built a highly available, scalable commercial advertising platform, discusses the evolution of its infrastructure, offers practical advice for engineers aspiring to become architects, and reflects on emerging technologies and time‑management strategies.

ArchitecturePlatform EngineeringSogou
0 likes · 10 min read
Scaling Billion‑Level Ads: Architecture Lessons from Sogou’s Senior Engineer
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Oct 31, 2016 · Cloud Computing

How Taobao Scaled from LAMP to Cloud: Lessons in Cloud Migration Architecture

This article examines the evolution of Taobao's technical architecture—from a LAMP stack through Oracle‑based mainframes to a cloud‑native platform—highlighting the performance, scalability, and cost challenges of traditional IT and offering best‑practice strategies for migrating enterprise systems to the cloud.

Cloud ComputingOperationsarchitecture migration
0 likes · 15 min read
How Taobao Scaled from LAMP to Cloud: Lessons in Cloud Migration Architecture