Tagged articles

big data

3804 articles · Page 34 of 39
Efficient Ops
Efficient Ops
Jul 18, 2018 · Operations

How Alibaba Scales Real‑Time Computing: Evolution of Its Operations Architecture

This article details Alibaba's real‑time computing platform, outlining its operational challenges, the unified automation platform Aquila, proactive fault‑elimination strategies, and ongoing moves toward intelligent, data‑driven management to support massive workloads during events like Double‑11.

AlibabaAquilaReal-Time Computing
0 likes · 15 min read
How Alibaba Scales Real‑Time Computing: Evolution of Its Operations Architecture
Didi Tech
Didi Tech
Jul 17, 2018 · Artificial Intelligence

Didi Showcases AI‑Driven Intelligent Transportation Research at ACM SIGIR 2018

At ACM SIGIR 2018, Didi presented AI‑driven intelligent‑transportation research—including a ride‑sharing preference prediction paper, keynote insights on smart dispatch, maps and traffic, collaborations with over twenty cities and numerous universities, open data initiatives, and plans for new thematic research programs.

Artificial IntelligenceIndustry-Academia CollaborationIntelligent Transportation
0 likes · 9 min read
Didi Showcases AI‑Driven Intelligent Transportation Research at ACM SIGIR 2018
360 Tech Engineering
360 Tech Engineering
Jul 13, 2018 · Big Data

Titan 2.0 Big Data Processing Platform: Architecture Evolution and Practice

The article describes the evolution of 360's Titan big‑data processing platform through three architectural stages, details its functional modules, explains the DITTO component framework, context and rule‑engine abstractions, and shares practical case studies and personal insights on building a flexible, self‑service data platform.

DITTOETLbig data
0 likes · 12 min read
Titan 2.0 Big Data Processing Platform: Architecture Evolution and Practice
High Availability Architecture
High Availability Architecture
Jul 12, 2018 · Information Security

Evolution of Zhihu’s Anti‑Cheat System “Wukong”: Architecture, Strategies, and Lessons Learned

This article chronicles the three‑generation evolution of Zhihu’s anti‑cheat platform Wukong, detailing its business context, spam taxonomy, multi‑layered control methods, architectural redesigns, strategy language improvements, graph‑based risk analysis, and the continuous integration of big‑data and machine‑learning techniques to combat content and behavior spam.

anti-cheatbig datagraph-analysis
0 likes · 23 min read
Evolution of Zhihu’s Anti‑Cheat System “Wukong”: Architecture, Strategies, and Lessons Learned
Ctrip Technology
Ctrip Technology
Jul 3, 2018 · Big Data

Ctrip's Presto Engine: Challenges, Improvements, and Upgrade Roadmap

This article details Ctrip's experience with the Presto distributed SQL engine, outlining the initial performance and stability issues, the comprehensive enhancements made in security, resource control, compatibility, and monitoring, and the multi‑stage upgrade plan that guides its future evolution.

KerberosPerformance OptimizationQuery Engine
0 likes · 11 min read
Ctrip's Presto Engine: Challenges, Improvements, and Upgrade Roadmap
AntTech
AntTech
Jul 3, 2018 · Backend Development

Evolution of Financial‑Grade Message Queues at Ant Financial

The article reviews the ten‑year evolution of Ant Financial's message queue, detailing its core reliability, consistency, availability and performance requirements, the architectural mechanisms built to meet them, the shift to pull‑mode and API‑mode designs, and the recent integration of compute capabilities to create a smart data transmission platform.

Streamingbig datadistributed systems
0 likes · 13 min read
Evolution of Financial‑Grade Message Queues at Ant Financial
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jul 2, 2018 · Artificial Intelligence

How JD.com Built a Multi‑Screen Personalized Recommendation Engine

This article explains how JD.com evolved its recommendation system from simple product suggestions to a sophisticated, multi‑screen, multi‑type personalized engine using big‑data collection, real‑time behavior tracking, machine‑learning models, and a modular architecture that boosts conversion and user experience.

big datae-commercemachine learning
0 likes · 14 min read
How JD.com Built a Multi‑Screen Personalized Recommendation Engine
Baidu Intelligent Testing
Baidu Intelligent Testing
Jun 29, 2018 · Product Management

Baidu Product Evaluation Framework and Common Assessment Methods

This article outlines Baidu's comprehensive product evaluation framework, describing its multi‑layer assessment system, the combination of subjective and objective metrics, and a suite of common evaluation methods such as indicator analysis, AB testing, user feedback, behavior analysis, big‑data profiling, and competitor comparison.

AB TestingQAbig data
0 likes · 16 min read
Baidu Product Evaluation Framework and Common Assessment Methods
58 Tech
58 Tech
Jun 27, 2018 · Big Data

Overview of the 58 User Profile System Architecture and Data Processing

The article describes the design, data integration, ID mapping, tag generation, and application scenarios of the 58 user profiling platform, which aggregates billions of user IDs across multiple business lines to provide online and offline persona data for personalization, analytics, and AI modeling.

Data IntegrationID mappingbig data
0 likes · 12 min read
Overview of the 58 User Profile System Architecture and Data Processing
DataFunTalk
DataFunTalk
Jun 24, 2018 · Big Data

OPPO Big Data Platform Operations and R&D Practices: Architecture, Scaling, and Monitoring

This article summarizes OPPO's rapid growth of its big‑data platform, detailing the three‑layer architecture, the evolution from Flume‑Kafka to NiFi for data ingestion, the upgrade of the OFlow task scheduler, comprehensive monitoring of data, resources and task SLA, and the development of a self‑service analytics tool called InnerEye to ensure stability, efficiency, and security.

AirflowNiFiOperations
0 likes · 10 min read
OPPO Big Data Platform Operations and R&D Practices: Architecture, Scaling, and Monitoring
Architecture Digest
Architecture Digest
Jun 18, 2018 · Operations

Design and Optimization of Large‑Scale Log Systems

This article examines the challenges of handling massive log data in high‑traffic e‑commerce platforms and presents a comprehensive architecture, optimization strategies, and practical implementations—including Rsyslog, Kafka, Fluentd, and the ELK stack—to improve scalability, performance, and reliability of log management systems.

ELKFluentdKafka
0 likes · 17 min read
Design and Optimization of Large‑Scale Log Systems
Didi Tech
Didi Tech
Jun 16, 2018 · Artificial Intelligence

AI and Big Data in Didi’s Mapping Services – Insights from WGDC 2018

At WGDC 2018, Didi’s mapping division revealed how its AI‑driven platform leverages massive real‑time travel data, machine‑learning and deep‑learning models—including a new ETA estimator, demand‑supply forecasting, and reinforcement‑learning order allocation—to deliver ultra‑accurate pick‑up points, route planning, and destination predictions, while opening de‑identified data and research topics to academia.

AIETAMapping
0 likes · 6 min read
AI and Big Data in Didi’s Mapping Services – Insights from WGDC 2018
Tencent Cloud Developer
Tencent Cloud Developer
Jun 11, 2018 · Cloud Computing

Tencent Cloud's Government Cloud Strategy and Digital Guangdong Practice

Tencent Cloud’s government‑cloud strategy, showcased by Guangdong’s “粤省事” platform, leverages WeChat as a single access point and a partner‑driven backend of AI, big‑data and IoT services to digitize certificates, streamline workflows for citizens, businesses and officials, and address low public‑service satisfaction by redesigning processes rather than merely automating them.

AIDigital TransformationWeChat mini-program
0 likes · 12 min read
Tencent Cloud's Government Cloud Strategy and Digital Guangdong Practice
Efficient Ops
Efficient Ops
Jun 6, 2018 · Big Data

How Tencent’s Multi‑Dimensional Monitoring Turns Big Data Into Real‑Time Business Insights

This article explains how Tencent’s ZhiYun multi‑dimensional monitoring system evolves from the Mobile Monitor platform, outlines its design principles, data‑factory capabilities, storage choices, and intelligent features, and demonstrates how it enables real‑time, multi‑dimensional analysis and alerting for large‑scale business operations.

Data PipelineDruidStorm
0 likes · 11 min read
How Tencent’s Multi‑Dimensional Monitoring Turns Big Data Into Real‑Time Business Insights
ITPUB
ITPUB
Jun 4, 2018 · Big Data

Is Hadoop Really Declining? Expert Insights Show Why the Ecosystem Stays Strong

Despite Gartner's 2017 claim that Hadoop is nearing the end of its production maturity, a series of interviews with Chinese big‑data experts reveal that Hadoop's ecosystem remains robust, with core components like HDFS, YARN, Spark, and HBase continuing to dominate the market.

GartnerHadoopHive
0 likes · 9 min read
Is Hadoop Really Declining? Expert Insights Show Why the Ecosystem Stays Strong
ITPUB
ITPUB
Jun 3, 2018 · Big Data

Spark vs Hadoop: Which Distributed System Fits Your Data Needs?

An in‑depth comparison of Hadoop and Spark examines their architectures, performance, cost, security, and machine‑learning capabilities, helping readers decide which open‑source distributed processing platform best matches their batch, streaming, and analytical workloads.

HadoopPerformanceSpark
0 likes · 13 min read
Spark vs Hadoop: Which Distributed System Fits Your Data Needs?
ITPUB
ITPUB
Jun 2, 2018 · Big Data

Mastering Spark: Core Concepts, Architecture, Streaming & Performance Tuning

This comprehensive guide explains Spark's ecosystem, execution principles, key features, deployment architectures, core concepts like RDD, Transformations, Actions, Jobs, Stages, Shuffle and Cache, as well as Spark Streaming mechanics and practical resource‑tuning tips for optimal big‑data processing.

ClusterRDDSpark
0 likes · 15 min read
Mastering Spark: Core Concepts, Architecture, Streaming & Performance Tuning
Tencent Cloud Developer
Tencent Cloud Developer
Jun 1, 2018 · Backend Development

Building Tencent Xinge: Architecture and Practices for Massive Mobile Push Service

The talk details Tencent Xinge’s architecture and cloud‑native practices that enable hundred‑billion‑level mobile push, combining terminal integration, real‑time backend filtering, distributed bitmap selection, precise‑push AI models, and DevOps pipelines to deliver fast, scalable, data‑driven notifications with effect tracking.

Real-time Analyticsbackend architecturebig data
0 likes · 18 min read
Building Tencent Xinge: Architecture and Practices for Massive Mobile Push Service
ITPUB
ITPUB
May 31, 2018 · Big Data

Mastering Spark on DataMagic: Fast‑Track Your Big Data Skills

This article explains Spark's role in the DataMagic platform, outlines four practical steps to quickly master Spark, details key configuration and parallelism settings, shows how to modify Spark code, and provides operational tips for cluster management and job troubleshooting.

ConfigurationDataMagicParallelism
0 likes · 10 min read
Mastering Spark on DataMagic: Fast‑Track Your Big Data Skills
dbaplus Community
dbaplus Community
May 30, 2018 · Big Data

Understanding Spark Executor Memory Management: On‑Heap, Off‑Heap, and Unified Strategies

This article explains Spark's executor memory architecture, covering on‑heap and off‑heap allocation, static versus unified memory managers, storage and execution memory handling, RDD persistence levels, eviction policies, and shuffle memory usage, providing practical formulas and configuration tips for optimal performance.

ExecutorOff‑HeapSpark
0 likes · 23 min read
Understanding Spark Executor Memory Management: On‑Heap, Off‑Heap, and Unified Strategies
Architecture Digest
Architecture Digest
May 27, 2018 · Big Data

Installing Elasticsearch and Performing Data Aggregation Queries

This article walks through installing Elasticsearch 5.6.9, configuring system limits, creating indices, inserting and deleting documents, executing complex aggregation queries, and integrating Elasticsearch with Java using the TransportClient, providing a practical guide for building analytics on large‑scale data.

AnalyticsElasticsearchbig data
0 likes · 12 min read
Installing Elasticsearch and Performing Data Aggregation Queries
DataFunTalk
DataFunTalk
May 22, 2018 · Information Security

Designing a Credit-Based Content Management System: Strategies, Risk Assessment, and AI Techniques

The article outlines how to build a credit‑based content management platform by describing the evolution of security practices, defining user‑generated, professional‑generated, and occupational content models, proposing a credit‑audit workflow with risk assessment, and presenting AI‑driven text classification and anti‑cheat methods to balance traffic, quality, and trust.

Artificial Intelligencebig datacontent moderation
0 likes · 12 min read
Designing a Credit-Based Content Management System: Strategies, Risk Assessment, and AI Techniques
Architects' Tech Alliance
Architects' Tech Alliance
May 14, 2018 · Big Data

Understanding Hadoop MapReduce Architecture and YARN: Components, Workflow, and Optimization

This article explains Hadoop's distributed storage and processing framework, details the MapReduce programming model, describes the classic JobTracker/TaskTracker architecture, outlines the shuffle and combine phases, and introduces YARN as a scalable replacement with its ResourceManager, ApplicationMaster, and NodeManager components.

HadoopMapReduceYARN
0 likes · 13 min read
Understanding Hadoop MapReduce Architecture and YARN: Components, Workflow, and Optimization
Alibaba Cloud Developer
Alibaba Cloud Developer
May 3, 2018 · Artificial Intelligence

How Alibaba’s City Brain Uses AI to Transform Urban Management

In a recent Cloud Xi conference, Hua Xiansheng, deputy director of Alibaba’s DAMO Academy Machine Intelligence Lab, presented the City Brain initiative, unveiling three AI-powered products—Tianyao, Tianying, and Tianji—that leverage massive video and sensor data to achieve real‑time perception, decision‑making, prediction, and intervention for smarter urban governance.

AIbig datasmart city
0 likes · 10 min read
How Alibaba’s City Brain Uses AI to Transform Urban Management
Suning Technology
Suning Technology
Apr 28, 2018 · Artificial Intelligence

How AI, Cloud, and Big Data Power Smart Retail: Insights from Suning’s Digital Strategy

Suning’s executive vice president outlines how digital empowerment—through AI, cloud computing, and big‑data analytics—creates a new smart‑retail ecosystem, emphasizing user‑centric services, data‑driven operations, and technology platforms that blend efficiency with a human touch.

AICloud ComputingDigital Transformation
0 likes · 9 min read
How AI, Cloud, and Big Data Power Smart Retail: Insights from Suning’s Digital Strategy
Beike Product & Technology
Beike Product & Technology
Apr 26, 2018 · Big Data

Chain Home's OLAP Platform and Kylin Usage

This article details Chain Home's OLAP platform architecture and Kylin usage, covering the evolution from early ROLAP to MOLAP multi-dimensional engine, Kylin's basic principles, platform structure, application scenarios, usage specifications, capability extensions, and middleware development.

Apache KylinChain HomeCube
0 likes · 11 min read
Chain Home's OLAP Platform and Kylin Usage
dbaplus Community
dbaplus Community
Apr 24, 2018 · Databases

Scaling Baidu’s TSDB to Trillions of Points: Elastic, High‑Performance Architecture

Baidu’s TSDB processes over 20 million data points per second per node and tens of thousands of queries per second cluster‑wide by employing a stateless read/write‑separated elastic architecture, multi‑layer storage across Redis, HBase and Hadoop, minute‑level geo‑redundant self‑healing, and a modified Gorilla compression that cuts storage by 80% with minimal CPU overhead.

TSDBTime Series Databasebig data
0 likes · 8 min read
Scaling Baidu’s TSDB to Trillions of Points: Elastic, High‑Performance Architecture
dbaplus Community
dbaplus Community
Apr 23, 2018 · Operations

Insights and Highlights from the 2018 Gdevops Global Agile Ops Summit

The 2018 Gdevops Global Agile Operations Summit in Chengdu gathered industry experts who shared practical insights on AIOps implementation, sharding database ecosystems, DevOps adoption in traditional enterprises, large‑scale data management, ElasticSearch clustering, AWS blue‑green deployments, cloud database operations, Alibaba's double‑11 ops platform, 58 delivery mini‑program architecture, and scalable game service design.

AIOpsDevOpsbig data
0 likes · 13 min read
Insights and Highlights from the 2018 Gdevops Global Agile Ops Summit
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Apr 19, 2018 · Information Security

How Suning Built a Comprehensive Information Security Architecture

This article outlines Suning's evolution from a basic network operations unit to a sophisticated, multi‑layered security architecture that integrates organizational structure, protection platforms, risk management, big‑data threat perception, and continuous improvement to safeguard e‑commerce operations.

Security Architecturebig datainformation security
0 likes · 10 min read
How Suning Built a Comprehensive Information Security Architecture
Architecture Digest
Architecture Digest
Apr 19, 2018 · Cloud Computing

Understanding the Relationship Between Cloud Computing, Big Data, and Artificial Intelligence

This article explains how cloud computing, big data, and artificial intelligence are interrelated, describing the evolution from physical resource management to virtualized, elastic services, the roles of IaaS, PaaS, and SaaS, and how each technology benefits the others in modern applications.

Artificial IntelligenceCloud ComputingIaaS
0 likes · 36 min read
Understanding the Relationship Between Cloud Computing, Big Data, and Artificial Intelligence
UCloud Tech
UCloud Tech
Apr 18, 2018 · Big Data

How Elasticsearch Powers Billion‑Record Log Analysis and Full‑Text Search

This article explains how Elasticsearch and the ELK stack address challenges of storing, securing, retrieving, and analyzing massive data volumes by providing distributed real‑time search, log collection, visualization, and even serving as a NoSQL alternative for large‑scale applications.

ELKElasticsearchbig data
0 likes · 7 min read
How Elasticsearch Powers Billion‑Record Log Analysis and Full‑Text Search
Architecture Digest
Architecture Digest
Apr 18, 2018 · Databases

Understanding Distributed Architecture and Its Applications in MySQL and Large‑Scale Systems

The article explains the concept of distributed architecture, its key characteristics such as cohesion and transparency, showcases how MySQL and middleware like Mycat are used in e‑commerce platforms, and outlines the evolution, practical implementations, and challenges of building scalable distributed database systems.

MySQLMycatbig data
0 likes · 15 min read
Understanding Distributed Architecture and Its Applications in MySQL and Large‑Scale Systems
Qunar Tech Salon
Qunar Tech Salon
Apr 10, 2018 · Big Data

Design and Implementation of Meituan's Traffic Compass Data Warehouse for Hotel‑Travel Business

The article presents Meituan's Traffic Compass—a data‑warehouse‑driven traffic analysis platform for the hotel‑travel business—detailing its background, challenges, architectural layers, dimensional modeling, Kylin‑based query engine, configuration mechanisms, performance metrics, and future optimization plans.

AnalyticsKylinMeituan
0 likes · 14 min read
Design and Implementation of Meituan's Traffic Compass Data Warehouse for Hotel‑Travel Business
Qunar Tech Salon
Qunar Tech Salon
Apr 9, 2018 · Big Data

Analysis of Apache Spark 2.2.1 Memory Management Model

This article examines Spark's unified memory manager in version 2.2.1, detailing on‑heap and off‑heap memory regions, the four on‑heap memory pools, dynamic execution‑storage memory sharing, task memory accounting, and provides concrete calculation examples to explain UI discrepancies and runtime memory limits.

ExecutorOff‑HeapSpark
0 likes · 13 min read
Analysis of Apache Spark 2.2.1 Memory Management Model
ITPUB
ITPUB
Mar 29, 2018 · Big Data

Demystifying Hadoop: MapReduce, Shuffle, and YARN Architecture

This article explains Hadoop’s core components, the MapReduce programming model, the detailed shuffle and merge processes, and how YARN replaces the classic JobTracker/TaskTracker design to improve scalability and resource utilization in large‑scale data processing clusters.

HadoopMapReduceShuffle
0 likes · 15 min read
Demystifying Hadoop: MapReduce, Shuffle, and YARN Architecture
Architecture Digest
Architecture Digest
Mar 26, 2018 · Operations

Alipay’s Double 11 Architecture: Logical Data Centers, Distributed Transactions, and High‑Availability Strategies

The article details Alipay’s comprehensive architecture for the Double 11 shopping festival, covering its three‑layer IAAS/PAAS/SAAS model, logical data‑center design, multi‑active disaster‑recovery, blue‑green deployment, distributed data sharding, transaction processing, and the Ant Credit Pay service’s performance and risk‑control mechanisms.

AlipayArchitecturebig data
0 likes · 16 min read
Alipay’s Double 11 Architecture: Logical Data Centers, Distributed Transactions, and High‑Availability Strategies
MaGe Linux Operations
MaGe Linux Operations
Mar 23, 2018 · Cloud Computing

Why Cloud Computing, Big Data, and AI Are Inseparable: A Beginner’s Guide

This article explains the origins and goals of cloud computing, how virtualization adds flexibility in time and space, the evolution from physical servers to public and private clouds, the role of IaaS, PaaS, and SaaS, and how big data and artificial intelligence intertwine with cloud services to enable modern intelligent applications.

Artificial IntelligenceCloud ComputingIaaS
0 likes · 37 min read
Why Cloud Computing, Big Data, and AI Are Inseparable: A Beginner’s Guide
Meituan Technology Team
Meituan Technology Team
Mar 22, 2018 · Big Data

High-Performance User Behavior Analysis Solution for Massive Data

The paper describes a high‑performance user‑behavior analysis system that processes hundreds of billions of daily logs for Meituan‑Dianping, using an inverted‑index structure with bitmap UUID sets and timestamp sequences, combined with Spark, Spring and Alluxio optimizations to cut query times from hours to under five seconds.

Distributed ComputingOLAP analysisbig data
0 likes · 14 min read
High-Performance User Behavior Analysis Solution for Massive Data
dbaplus Community
dbaplus Community
Mar 18, 2018 · Databases

What’s New in the Database World? March 2018 Release Roundup

The March 2018 DBAplus Newsletter compiles the latest releases across RDBMS, NoSQL, NewSQL, time‑series and big‑data ecosystems, highlighting new features, performance improvements, compatibility updates and key technical links for Oracle, MySQL, MariaDB, SQL Server, DB2, PostgreSQL, TiDB, CockroachDB, InfluxDB, Hadoop and several Chinese‑made databases.

Database ReleasesHTAPNewSQL
0 likes · 21 min read
What’s New in the Database World? March 2018 Release Roundup
MaGe Linux Operations
MaGe Linux Operations
Mar 17, 2018 · Operations

From Manual Ops to Automated Cloud: A 7‑Year Journey of a Game Ops Team

This article chronicles a game company's operations team evolution over seven years, detailing how it grew from a tiny manual crew to a large, automated, cloud‑native organization that built its own CDN, monitoring, and platform solutions while tackling scaling, reliability, and service‑orientation challenges.

CDNCloud Computingbig data
0 likes · 21 min read
From Manual Ops to Automated Cloud: A 7‑Year Journey of a Game Ops Team
Architecture Digest
Architecture Digest
Mar 14, 2018 · Big Data

Attributes Matrix and Data Flow Models of Apache Streaming Platforms

This article presents a comprehensive attributes matrix and data‑flow model overview for major Apache streaming platforms, comparing versions, sponsors, event handling, fault tolerance, processing order, latency, resource management, APIs, and supported connectors to aid practical technology selection.

Apacheattributes matrixbig data
0 likes · 16 min read
Attributes Matrix and Data Flow Models of Apache Streaming Platforms
Beike Product & Technology
Beike Product & Technology
Mar 9, 2018 · Big Data

How Lianjia Built a Low‑Latency Real‑Time Data Platform with Spark Streaming

This article details Lianjia's journey of designing and implementing a low‑latency, stable real‑time computing platform using Spark Streaming on YARN, covering technical selection, architecture components, version compatibility challenges, exactly‑once semantics, graceful shutdown, Kafka tuning, and future enhancements.

Exactly-OnceKafkaSpark Streaming
0 likes · 11 min read
How Lianjia Built a Low‑Latency Real‑Time Data Platform with Spark Streaming
Suning Technology
Suning Technology
Mar 9, 2018 · Big Data

How Suning Built a Scalable Real-Time Log Analysis Platform with Spark Streaming

Suning’s real‑time log analysis system integrates Flume, Kafka, Storm and Spark Streaming to collect, cleanse, and compute metrics like NDCG, ensuring low latency, high throughput, exact‑once processing, and robust data safety while supporting multi‑dimensional analytics on massive online‑offline traffic.

Data PipelineData QualityNDCG
0 likes · 12 min read
How Suning Built a Scalable Real-Time Log Analysis Platform with Spark Streaming
Ctrip Technology
Ctrip Technology
Mar 8, 2018 · Big Data

Ctrip Wireless APM Platform: Architecture, Metrics, and Technical Details

The article describes the evolution of Ctrip's wireless APM platform from the early UBT-based monitoring to a globally‑oriented, metric‑rich system that processes over 100 billion data points daily using Storm and Elasticsearch, detailing its design, key performance dimensions, data‑volume trade‑offs, and implementation choices.

APMCtripElasticsearch
0 likes · 12 min read
Ctrip Wireless APM Platform: Architecture, Metrics, and Technical Details
dbaplus Community
dbaplus Community
Mar 7, 2018 · Big Data

Taming Massive HDFS Data Growth: Monitoring, Capacity Planning & Hive Optimization

The article outlines a systematic approach for large‑scale Hadoop clusters to monitor daily data growth, identify abnormal paths, manage rapid expansion, clean unused cold data, and implement capacity forecasts, while providing concrete daily and quarterly actions, Hive‑specific strategies, and practical examples to keep storage under control.

Data GrowthHDFSHadoop
0 likes · 17 min read
Taming Massive HDFS Data Growth: Monitoring, Capacity Planning & Hive Optimization
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Mar 2, 2018 · Cloud Computing

How Cloud, Big Data, and AI Converge to Transform Enterprise Data Strategies

The article explores how the integration of cloud computing, big data, and artificial intelligence is reshaping enterprise data platforms, outlining a multi‑stage evolution from data unification to ecosystem building and forecasting the strategic importance of data in future business transformation.

Artificial IntelligenceData EcosystemEnterprise Data
0 likes · 9 min read
How Cloud, Big Data, and AI Converge to Transform Enterprise Data Strategies
Hulu Beijing
Hulu Beijing
Feb 28, 2018 · Big Data

How Hulu’s Nesto Engine Delivers Near‑Real‑Time OLAP on TB‑Scale Data

This article introduces Hulu's in‑house OLAP engine Nesto, detailing its near‑real‑time data ingestion, nested data model, TB‑level storage using HBase and Parquet, MPP query execution, custom predicate library, and the overall architecture that enables sub‑second ad‑hoc queries for user analytics.

HBaseOLAPQuery Engine
0 likes · 22 min read
How Hulu’s Nesto Engine Delivers Near‑Real‑Time OLAP on TB‑Scale Data
JD Tech
JD Tech
Feb 28, 2018 · Operations

CallGraph: JD.com's Distributed Tracing and Service Governance Platform

CallGraph is JD.com's internally developed distributed tracing and service governance platform that addresses the challenges of monitoring complex microservice architectures by providing low‑intrusion, low‑latency tracing, real‑time analytics, configurable sampling, and integration with JMQ, Storm, Spark, HBase, and JimDB for both operational insight and performance optimization.

MicroservicesReal-time AnalyticsService Governance
0 likes · 12 min read
CallGraph: JD.com's Distributed Tracing and Service Governance Platform
21CTO
21CTO
Feb 20, 2018 · Big Data

Why Real-Time Streaming Is the Next Big Data Revolution for Developers

This article explains how real-time streaming has evolved from batch Hadoop systems through Lambda architecture to modern Kappa-style pipelines, highlighting its growing importance for developers, enterprises, and the integration of streaming with microservices, AI, and cloud-native technologies.

AI IntegrationKappa ArchitectureLambda Architecture
0 likes · 8 min read
Why Real-Time Streaming Is the Next Big Data Revolution for Developers
Architecture Digest
Architecture Digest
Feb 11, 2018 · Artificial Intelligence

Recent Advances in Bayesian Machine Learning: Foundations, Non‑Parametric Methods, and Large‑Scale Applications

This article reviews recent progress in Bayesian machine learning, covering foundational theory, non‑parametric approaches such as Dirichlet and Indian buffet processes, regularized Bayesian inference, and scalable techniques for big‑data environments including stochastic variational methods, distributed algorithms, and hardware acceleration.

Monte Carlobayesian learningbig data
0 likes · 23 min read
Recent Advances in Bayesian Machine Learning: Foundations, Non‑Parametric Methods, and Large‑Scale Applications
Java Backend Technology
Java Backend Technology
Feb 6, 2018 · Artificial Intelligence

How JD Built a Scalable AI-Powered Recommendation Engine for E‑Commerce

This article details JD's evolution from rule‑based recommendations to a multi‑screen, AI‑driven personalization platform, describing its system architecture, data pipelines, feature services, and key technologies that enable real‑time, user‑centric product suggestions across the e‑commerce ecosystem.

Artificial Intelligencebig datae-commerce
0 likes · 20 min read
How JD Built a Scalable AI-Powered Recommendation Engine for E‑Commerce
Architecture Digest
Architecture Digest
Feb 1, 2018 · Fundamentals

How Search Engines Work: Building Inverted Indexes

This article explains the core of search engine technology by describing what an inverted index is, how it is built using single‑pass memory and multi‑way merge methods, how indexes can be partitioned and incrementally updated, and how Hadoop can be used for large‑scale indexing.

HadoopInformation Retrievalbig data
0 likes · 10 min read
How Search Engines Work: Building Inverted Indexes
iQIYI Technical Product Team
iQIYI Technical Product Team
Jan 31, 2018 · Big Data

Evolution of iQIYI Real-Time Big Data Collection System

iQIYI’s big‑data collection system has progressed from simple HTTP log uploads to a Flume‑Kafka pipeline and finally to a custom Venus‑Agent architecture with centralized configuration, persistent offsets, dual‑Kafka streams and Flink processing, now handling tens of millions of queries per second and over three hundred billion records daily to power its AI‑driven services.

FlinkFlumeKafka
0 likes · 15 min read
Evolution of iQIYI Real-Time Big Data Collection System
Meituan Technology Team
Meituan Technology Team
Jan 26, 2018 · Big Data

Design and Implementation of a Real-Time Data Processing System at Meituan

Meituan designed a Storm‑based real‑time data processing platform that guarantees at‑least‑once delivery and high availability, employs a custom spout, regression‑driven traffic smoothing, and a low‑latency KV store with atomic operations, persisting results in Kafka, MySQL and Cellar to power merchant dashboards and heat‑tag analytics, while planning broader real‑time analytics expansion.

Real-time DataStormbig data
0 likes · 10 min read
Design and Implementation of a Real-Time Data Processing System at Meituan
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Jan 18, 2018 · Big Data

Smart Flood Control: Donghua Software’s IoT, Big Data & Cloud Solution

The case study details Donghua Software’s smart flood‑control and drainage solution, which integrates IoT sensors, NB‑IoT/eLTE networks, Huawei’s FusionSphere cloud platform, big‑data analytics, and GIS to provide real‑time monitoring, predictive warnings, automated gate control, and efficient emergency dispatch for urban water management.

Cloud ComputingFlood ManagementIoT
0 likes · 12 min read
Smart Flood Control: Donghua Software’s IoT, Big Data & Cloud Solution
Efficient Ops
Efficient Ops
Jan 16, 2018 · Operations

How Tencent Secures Game Operations: Real Cases, Challenges, and Data‑Driven Solutions

This article shares a comprehensive overview of game operation security at Tencent, covering personal background, real‑world incident cases, the inherent challenges of large‑scale game services, past monitoring efforts, and a new data‑driven alerting framework that dramatically reduces false alarms while protecting game economies.

Game SecurityOperationsalerting
0 likes · 25 min read
How Tencent Secures Game Operations: Real Cases, Challenges, and Data‑Driven Solutions
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jan 15, 2018 · Backend Development

Inside the Architecture of the World’s Biggest Websites: Wikipedia, Facebook, YouTube, and More

This article surveys the technical architectures of major web platforms—including Wikipedia, Facebook, Yahoo! Mail, Twitter, Google App Engine, Amazon, and Youku—highlighting their design patterns, scaling techniques, storage solutions, and caching strategies to reveal how massive online services are built and operated.

ArchitectureBackendbig data
0 likes · 10 min read
Inside the Architecture of the World’s Biggest Websites: Wikipedia, Facebook, YouTube, and More
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Jan 5, 2018 · Big Data

What Drove Big Data’s 2017 Surge and What’s Next? Insights & Predictions

Analyzing 2017’s big data boom, the article explores how the 4V characteristics—volume, variety, velocity, and value—spurred innovations like distributed storage, NoSQL, real‑time stream processing, and AI integration, and predicts future hotspots such as SQL resurgence, cloud‑based platforms, and AI‑driven analytics.

Artificial Intelligencebig datadata engineering
0 likes · 11 min read
What Drove Big Data’s 2017 Surge and What’s Next? Insights & Predictions
AntTech
AntTech
Jan 4, 2018 · Databases

Report on VLDB 2017 Conference: Insights and Highlights from Database Research

Attending VLDB 2017 in Munich, the report summarizes the conference’s broad coverage of database research—from new hardware‑accelerated prototypes and Spark‑based big‑data processing to Oracle and SAP HANA case studies, keynotes, notable papers, and reflections on industry trends and Chinese contributions.

VLDBbig datadatabase systems
0 likes · 22 min read
Report on VLDB 2017 Conference: Insights and Highlights from Database Research
dbaplus Community
dbaplus Community
Jan 1, 2018 · Big Data

How Vipshop Leverages Data Processing, Analytics, and Mining for Smarter Ops

This article summarizes Wu Xiaoguang's talk at Gdevops 2017, detailing how Vipshop integrates data processing, analysis, and mining technologies—such as Flume, Kafka, Spark, and custom scheduling—to improve operational decision‑making, performance monitoring, root‑cause analysis, and predictive modeling across its e‑commerce platform.

Data AnalyticsData ProcessingOperations
0 likes · 23 min read
How Vipshop Leverages Data Processing, Analytics, and Mining for Smarter Ops
Tencent Architect
Tencent Architect
Dec 30, 2017 · Databases

An Overview of Time Series Databases and Tencent CTSDB

This article introduces the concept, characteristics, and use cases of time series databases, explains the data model and challenges of traditional solutions, and provides a detailed overview of Tencent's Cloud Time Series Database (CTSDB) along with performance comparisons against InfluxDB.

CTSDBTime Series Databasebig data
0 likes · 12 min read
An Overview of Time Series Databases and Tencent CTSDB
Architects' Tech Alliance
Architects' Tech Alliance
Dec 28, 2017 · Operations

Intelligent Operations: Machine‑Learning‑Based AIOps – Lecture Summary by Prof. Pei Dan

In this lecture, Prof. Pei Dan of Tsinghua University outlines the evolution of intelligent operations from rule‑based automation to machine‑learning‑driven AIOps, discusses data, feedback loops, and practical challenges, and calls for stronger collaboration between industry and academia to accelerate research and deployment.

AIOpsCloud Computingbig data
0 likes · 10 min read
Intelligent Operations: Machine‑Learning‑Based AIOps – Lecture Summary by Prof. Pei Dan
Meituan Technology Team
Meituan Technology Team
Dec 28, 2017 · Big Data

Design and Implementation of a Scalable Scenario Query System for Meituan

Meituan built a scalable scenario‑query platform that unifies traffic, activity and investment data by layering RPC services, a Storm‑driven pre‑computation tree stored in Redis/Tair, and a middle‑platform API with circuit‑breaker logic, cutting response times from seconds to under one second while dramatically reducing code coupling and simplifying future feature development.

Apache StormNoSQLPrecomputation
0 likes · 12 min read
Design and Implementation of a Scalable Scenario Query System for Meituan
Architecture Digest
Architecture Digest
Dec 27, 2017 · Backend Development

Handling Transactions, Failover, and Exactly‑Once Semantics in Distributed Systems

This article explores how distributed systems determine node liveness, manage failover and recovery, and implement at‑most‑once, at‑least‑once, and exactly‑once processing guarantees—including opaque transactions and two‑phase commit—using examples from Kafka, Zookeeper, and big‑data pipelines.

Exactly-OnceZooKeeperbig data
0 likes · 15 min read
Handling Transactions, Failover, and Exactly‑Once Semantics in Distributed Systems
dbaplus Community
dbaplus Community
Dec 26, 2017 · Big Data

Turning Raw Logs into Structured Data with DBus Visual Rule Operators

This article explains how the open‑source DBus platform, combined with the Wormhole streaming engine, captures raw application logs, lets users configure visual rule operators, and transforms the unstructured message part into schema‑driven, Kafka‑ready data for downstream analytics.

DBusLog ProcessingStructured Logging
0 likes · 15 min read
Turning Raw Logs into Structured Data with DBus Visual Rule Operators
Architecture Digest
Architecture Digest
Dec 22, 2017 · Big Data

Redesign and Optimization of the WeChat Pay Transaction Record System

This article presents a comprehensive case study of how WeChat Pay rebuilt its transaction record storage system to handle massive data volumes, improve performance, ensure data completeness, support flexible queries, and strengthen security through distributed key‑value storage, data partitioning, and operational safeguards.

Data PartitioningWeChat Paybig data
0 likes · 11 min read
Redesign and Optimization of the WeChat Pay Transaction Record System
Qunar Tech Salon
Qunar Tech Salon
Dec 21, 2017 · Big Data

Experience and Optimization Strategies for Apache Kylin in Real-Time OLAP

This article shares a data engineer's three‑year experience using Apache Kylin for real‑time OLAP on petabyte‑scale data, describing the business background, challenges of pre‑computation, cube modeling, dimension reduction, and various optimization techniques such as hierarchy, mandatory, and joint dimensions, as well as precise count‑distinct handling.

Apache KylinCube OptimizationOLAP
0 likes · 13 min read
Experience and Optimization Strategies for Apache Kylin in Real-Time OLAP
Meitu Technology
Meitu Technology
Dec 19, 2017 · Industry Insights

Inside Meitu’s In‑House Log Collection System Arachnia: Design, Challenges, and Core Mechanisms

This article introduces Meitu’s self‑developed log collection system Arachnia, explaining why a custom solution was needed for massive server‑side user‑behavior logs, the key requirements such as reliability and real‑time throughput, and the core architectural mechanisms that address those challenges.

ArachniaData PipelineLog Collection
0 likes · 2 min read
Inside Meitu’s In‑House Log Collection System Arachnia: Design, Challenges, and Core Mechanisms
Meitu Technology
Meitu Technology
Dec 19, 2017 · Big Data

Meitu Internet Technology Salon Session 7: Practices in Recommendation Algorithms, Big Data, and Personalized Recommendation

At Meitu’s seventh Internet Technology Salon in Xiamen, over a hundred experts discussed recommendation algorithms and big‑data solutions, with talks on the Arachnia log‑collection system, the Naix distributed bitmap service, Meitu’s personalized recommendation pipeline challenges, and novel data‑missing‑theory models for improved performance.

big datadata collectiondistributed bitmap
0 likes · 8 min read
Meitu Internet Technology Salon Session 7: Practices in Recommendation Algorithms, Big Data, and Personalized Recommendation
Architecture Digest
Architecture Digest
Dec 16, 2017 · Big Data

Performance Comparison of Apache Flink and Apache Storm for Real‑Time Stream Processing

This report presents a systematic performance evaluation of Apache Flink and Apache Storm across multiple real‑time processing scenarios, measuring throughput, latency, message‑delivery semantics, and state‑backend effects, and provides recommendations for selecting the most suitable engine based on the observed results.

FlinkReal-time AnalyticsStorm
0 likes · 21 min read
Performance Comparison of Apache Flink and Apache Storm for Real‑Time Stream Processing
58 Tech
58 Tech
Dec 15, 2017 · Big Data

Design and Architecture of WMDA: A Comprehensive User Behavior Analysis Platform

The article details WMDA, a no‑code and manual‑code data collection platform for PC, mobile and app that supports real‑time and offline user behavior analysis, describing its functional model, behavior taxonomy, five‑layer architecture, tracking techniques, circle‑selection, data services, streaming and batch processing pipelines, and related technologies such as Storm, Spark, Druid and Roaring Bitmap.

DruidRoaring Bitmapbig data
0 likes · 18 min read
Design and Architecture of WMDA: A Comprehensive User Behavior Analysis Platform
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Dec 15, 2017 · Operations

Automated Fault Recovery Architecture for Alibaba's Network during Double Eleven

The article describes Alibaba's end‑to‑end automated fault recovery system for its massive network, covering extensive data collection, Spark‑based event processing, flexible alerting with Siddhi, alert convergence using PageRank, and scripted recovery actions to achieve high availability during the Double Eleven traffic surge.

Operationsautomationbig data
0 likes · 9 min read
Automated Fault Recovery Architecture for Alibaba's Network during Double Eleven