Tagged articles

big data

3804 articles · Page 37 of 39
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Oct 27, 2016 · Big Data

Inside Taobao’s Massive Data Architecture: How 1.5 PB Daily Is Processed and Served

The article explains Taobao’s five‑layer data product architecture—covering data sources, compute, storage, query, and product layers—and describes how massive volumes of data are ingested, processed in batch and streaming, stored in MySQL and HBase clusters, and served efficiently through a unified middle‑layer and sophisticated caching mechanisms.

CachingHBaseHadoop
0 likes · 15 min read
Inside Taobao’s Massive Data Architecture: How 1.5 PB Daily Is Processed and Served
21CTO
21CTO
Oct 21, 2016 · Artificial Intelligence

How Toutiao Dominated Chinese News with AI‑Powered Personalization

This article examines Toutiao’s evolution from a simple news aggregator to a 600‑billion‑RMB valued AI‑driven recommendation platform, detailing its market growth, data‑driven personalization, product features, business model, talent philosophy, and future outlook.

AIRecommendation EngineToutiao
0 likes · 10 min read
How Toutiao Dominated Chinese News with AI‑Powered Personalization
Efficient Ops
Efficient Ops
Oct 20, 2016 · Operations

Transforming Business Operations with Cloud, Big Data, and Integrated IT Management

The article explains how modern business operation management integrates cloud computing, big data analytics, and proactive IT monitoring to shift from traditional infrastructure‑centric maintenance to a user‑experience‑driven, data‑powered approach that boosts performance, accelerates growth, and supports digital transformation.

Cloud ComputingDigital TransformationIT monitoring
0 likes · 8 min read
Transforming Business Operations with Cloud, Big Data, and Integrated IT Management
Liulishuo Tech Team
Liulishuo Tech Team
Oct 17, 2016 · Big Data

Practical Tips and Common Pitfalls for Tuning Apache Spark Performance

This article shares hands‑on experience from Spark Summit attendees, covering why Spark is powerful, common performance problems such as slow jobs, OOM, data skew, excessive partitions, and provides concrete tuning advice on executors, cores, memory, and debugging techniques.

Apache SparkData SkewExecutor Configuration
0 likes · 11 min read
Practical Tips and Common Pitfalls for Tuning Apache Spark Performance
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Oct 17, 2016 · Artificial Intelligence

Wang Jian’s Keynote at the 2016 Hangzhou Yunqi Conference: Data Brain, AI, and Cloud Computing

In his 2016 Yunqi Conference keynote, Wang Jian highlighted how Alibaba’s cloud and AI technologies transform city traffic by linking surveillance cameras to traffic lights, discussed the evolution from Deep Blue to AlphaGo, and reflected on the broader impact of data-driven innovation on society.

AICloud Computingbig data
0 likes · 13 min read
Wang Jian’s Keynote at the 2016 Hangzhou Yunqi Conference: Data Brain, AI, and Cloud Computing
Qunar Tech Salon
Qunar Tech Salon
Oct 17, 2016 · Information Security

Design and Implementation of a Cloud‑Based Web Application Firewall at Ctrip

This article describes Ctrip's challenges with web security, evaluates hardware and commercial cloud WAF shortcomings, and presents a low‑cost, low‑risk cloud‑based WAF solution that leverages DNS redirection, closed‑loop rule management, Lua/Tengine deployment, supervised machine‑learning log analysis, and big‑data streaming for real‑time threat detection and mitigation.

Distributed ArchitectureWAFWeb Security
0 likes · 9 min read
Design and Implementation of a Cloud‑Based Web Application Firewall at Ctrip
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Oct 16, 2016 · Big Data

Mastering Data Sync, Real-Time Analytics, and Scalable Storage for Modern Systems

This article explains how to design and implement heterogeneous data synchronization, leverage batch and stream processing frameworks like Hadoop and Storm for large‑scale analysis, and choose appropriate storage solutions—from in‑memory databases to distributed column‑family stores—while addressing performance, reliability, and monitoring in complex distributed environments.

big datadata synchronizationdatabases
0 likes · 26 min read
Mastering Data Sync, Real-Time Analytics, and Scalable Storage for Modern Systems
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Oct 15, 2016 · Artificial Intelligence

Computing Power as the Engine of Digitalization and Intelligent Manufacturing – Insights from Alibaba Cloud Conference

In his keynote at the Alibaba Cloud Conference, CTO Zhang Jianfeng explains how advances in computing power, AI, big data, and IoT are driving the digital transformation of retail, manufacturing, and services, enabling smarter products, personalized experiences, and a fully connected intelligent world.

Cloud ComputingIoTbig data
0 likes · 19 min read
Computing Power as the Engine of Digitalization and Intelligent Manufacturing – Insights from Alibaba Cloud Conference
Alibaba Cloud Developer
Alibaba Cloud Developer
Oct 14, 2016 · Artificial Intelligence

How Alibaba’s CTO Envisions AI‑Driven Smart Manufacturing and the Future of Digitalized Worlds

In his Yunqi Conference keynote, Alibaba CTO Zhang Jianfeng explains how soaring computing power, digitalization, AI, IoT and immersive technologies will transform retail, manufacturing and services into intelligent, personalized ecosystems, illustrating the vision with examples like a smart golf club and a city‑wide data brain.

AIIoTbig data
0 likes · 21 min read
How Alibaba’s CTO Envisions AI‑Driven Smart Manufacturing and the Future of Digitalized Worlds
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Oct 8, 2016 · Big Data

Evolving Data Warehouses with Hadoop & Spark: Core Technologies

Data warehouses centralize and transform enterprise data for multidimensional analysis, and modern demands have spawned four types—traditional, real‑time, associative discovery, and data marts—each with distinct technical requirements, while Hadoop‑based solutions like Transwarp Data Hub address challenges of scale, variety, latency, and security.

Distributed ComputingHadoopReal-time Analytics
0 likes · 21 min read
Evolving Data Warehouses with Hadoop & Spark: Core Technologies
Java High-Performance Architecture
Java High-Performance Architecture
Sep 27, 2016 · Big Data

Build a Hadoop Cluster with Docker: Step‑by‑Step Guide

Learn how to quickly set up a multi‑node Hadoop cluster on a single machine using Docker containers, covering image preparation, SSH configuration, fixed IP assignment with pipework, and building custom Hadoop images, enabling a lightweight, cost‑effective big‑data environment for development and testing.

CentOSClusterDocker
0 likes · 9 min read
Build a Hadoop Cluster with Docker: Step‑by‑Step Guide
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Sep 26, 2016 · Operations

Automate Cluster Health Checks with Koalas: Cutting Big Data Downtime

The article introduces Koalas, an automated distributed diagnostic tool for TDH clusters that identifies and resolves computing environment issues—such as network, platform, and system problems—through one‑click checks, detailed reporting, and both preventive and diagnostic use cases.

Cluster MonitoringNetwork ThroughputPerformance Optimization
0 likes · 8 min read
Automate Cluster Health Checks with Koalas: Cutting Big Data Downtime
dbaplus Community
dbaplus Community
Sep 12, 2016 · Big Data

Apache Flume Quickstart: Log Collection and Kafka Integration

This article introduces Apache Flume, explains its design goals of reliability, scalability, manageability and extensibility, outlines core concepts and architecture, provides step‑by‑step configuration using the first mode, demonstrates integration with Zookeeper, Kafka and a shell script, and shows how to launch and verify the agent.

Apache FlumeKafka Integrationbig data
0 likes · 7 min read
Apache Flume Quickstart: Log Collection and Kafka Integration
Ctrip Technology
Ctrip Technology
Sep 10, 2016 · Artificial Intelligence

Deep Learning Anti‑Scam Guide: An Informal Introduction to Neural Networks, Training, and Practical Applications

This article provides a light‑hearted yet thorough overview of deep learning, covering neural network fundamentals, layer construction, back‑propagation, ResNet shortcuts, encoder‑decoder structures, PU‑learning for unlabeled data, GPU acceleration, and practical advice on data size, frameworks, and deployment in financial scenarios.

BackpropagationPu-Learningbig data
0 likes · 27 min read
Deep Learning Anti‑Scam Guide: An Informal Introduction to Neural Networks, Training, and Practical Applications
Meituan Technology Team
Meituan Technology Team
Aug 26, 2016 · Big Data

Memory Architecture and Analysis of Hadoop HDFS NameNode

The article dissects Hadoop 2.4.1’s HDFS NameNode memory architecture, detailing how the Namespace, BlockManager, NetworkTopology, and LeaseManager consume the heap, exposing scaling problems when metadata reaches hundreds of millions of inodes and blocks, and recommending file merging, block‑size tuning, federation, or external KV stores to mitigate heap pressure.

HDFSNameNodebig data
0 likes · 17 min read
Memory Architecture and Analysis of Hadoop HDFS NameNode
Ctrip Technology
Ctrip Technology
Aug 26, 2016 · Information Security

Ctrip Information Security Salon Summary – Cloud WAF, Big Data Analysis, ELK Monitoring, and Recruitment Highlights

The Ctrip Information Security Salon held on August 20 in Shanghai featured expert talks on cloud‑based WAF, big‑data security analytics, ELK‑driven monitoring, product security practices, and concluded with a recruitment drive for security engineers, showcasing practical implementations and industry challenges.

Cloud WAFCtripELK
0 likes · 8 min read
Ctrip Information Security Salon Summary – Cloud WAF, Big Data Analysis, ELK Monitoring, and Recruitment Highlights
MaGe Linux Operations
MaGe Linux Operations
Aug 23, 2016 · Big Data

Step-by-Step Guide to Building a Hadoop Cluster on CentOS 6.5

This article provides a comprehensive, hands‑on tutorial for setting up a Hadoop 2.6.4 cluster on a CentOS 6.5 development server, covering SSH password‑less login, user/group creation, DNS configuration, JDK installation, environment variables, Hadoop installation, HDFS and YARN configuration, and troubleshooting native library warnings.

CentOSCluster SetupHDFS
0 likes · 12 min read
Step-by-Step Guide to Building a Hadoop Cluster on CentOS 6.5
Ctrip Technology
Ctrip Technology
Aug 19, 2016 · Big Data

Ctrip's Big Data Architecture and Personalized Recommendation System

This article describes how Ctrip transformed its traditional application architecture into a high‑concurrency, big‑data‑driven platform, detailing storage, compute, and business‑layer redesigns that enable massive data ingestion, real‑time user‑intent services, and a scalable personalized recommendation system.

CtripHadoopSpark
0 likes · 14 min read
Ctrip's Big Data Architecture and Personalized Recommendation System
Architects' Tech Alliance
Architects' Tech Alliance
Aug 18, 2016 · Cloud Computing

Understanding the Evolution and Competition Between Traditional and Emerging Storage Technologies

The article analyzes how cloud-native, distributed, and software‑defined storage solutions are reshaping the enterprise storage market, compares them with traditional high‑reliability systems, and offers guidance on selecting, integrating, and migrating storage technologies based on business scenarios, cost, and performance considerations.

Cloud StorageSDSbig data
0 likes · 10 min read
Understanding the Evolution and Competition Between Traditional and Emerging Storage Technologies
Ctrip Technology
Ctrip Technology
Aug 12, 2016 · Big Data

Ctrip's Real-Time Data Platform: Architecture, Practices, and Lessons Learned

This article details Ctrip's journey building a unified real-time data platform—covering business motivations, architectural requirements, technology choices like Kafka and Storm, implementation of Avro schemas, monitoring, alerting, operational lessons, and future explorations such as Streaming CQL and JStorm.

KafkaPlatform ArchitectureReal-time Data
0 likes · 15 min read
Ctrip's Real-Time Data Platform: Architecture, Practices, and Lessons Learned
Meituan Technology Team
Meituan Technology Team
Aug 5, 2016 · Big Data

Meituan Delivery Big Data: Full‑Chain Application Insights

The talk details Meituan‑Dianping’s end‑to‑end delivery big‑data system, highlighting mobile‑ and local‑centric usage, a two‑stage forecasting pipeline that combines autoregressive baselines with a boosting‑based multiplier model, layered log‑collect‑process‑serve architecture, sophisticated feature engineering, real‑time inference, and strategies for logistics constraints and cold‑start merchants.

Meituanbig datadelivery
0 likes · 8 min read
Meituan Delivery Big Data: Full‑Chain Application Insights
Meituan Technology Team
Meituan Technology Team
Aug 5, 2016 · Big Data

Meituan-Dianping Tech Salon: Full‑Chain Application of Food‑Delivery Big Data – User Profiling, Marketing Strategies, and Predictive Modeling

The Meituan‑Dianping tech salon detailed how food‑delivery big data drives full‑chain marketing, using RFM‑based user segmentation, rich demographic and behavior profiles, churn‑prediction and survival models, and scenario‑driven expansion tactics to acquire, retain, and grow customers across the order lifecycle.

MeituanRFM Segmentationbig data
0 likes · 9 min read
Meituan-Dianping Tech Salon: Full‑Chain Application of Food‑Delivery Big Data – User Profiling, Marketing Strategies, and Predictive Modeling
Architecture Digest
Architecture Digest
Aug 4, 2016 · Big Data

Heron vs. Storm: Architecture, Performance Evaluation, and Design Lessons

The article provides a comprehensive overview of Twitter's Heron stream processing system, comparing its architecture, design principles, back‑pressure mechanisms, resource utilization, and performance test results with Storm/JStorm, and concludes with practical insights for large‑scale deployments.

ArchitectureHeronStorm
0 likes · 23 min read
Heron vs. Storm: Architecture, Performance Evaluation, and Design Lessons
ITPUB
ITPUB
Jul 19, 2016 · Big Data

From Traditional Data Warehouses to Big Data: Practical Techniques and Migration Insights

The talk shares hands‑on experiences and best‑practice methods for traditional data‑warehouse processing, public and behavioral data handling in big‑data environments, and practical guidance for migrating legacy warehouses to modern Hadoop‑based platforms, emphasizing data governance, security, and performance optimization.

ETLHadoopSpark
0 likes · 13 min read
From Traditional Data Warehouses to Big Data: Practical Techniques and Migration Insights
Architect
Architect
Jul 14, 2016 · Big Data

Understanding Custom Stream IDs and Topology Building in Apache Storm

This article explains how to construct Apache Storm topologies with custom stream IDs, demonstrates the classic WordCountTopology example, and provides detailed Java code snippets illustrating spout and bolt configurations, stream declarations, and grouping strategies for real‑time stream processing.

Apache StormCustom Stream IDJava
0 likes · 8 min read
Understanding Custom Stream IDs and Topology Building in Apache Storm
Baidu Intelligent Testing
Baidu Intelligent Testing
Jul 13, 2016 · Artificial Intelligence

Detecting Offline Merchant Service Issues Using Machine Learning and Big Data at Nuomi

The article describes how Nuomi analyzes refund and complaint data with machine‑learning and big‑data techniques, extracts features for single‑ and multi‑store scenarios, builds decision‑tree models with regional adjustments, and creates an online workflow to promptly intervene on merchants that fail to serve customers.

big datacustomer experiencedecision tree
0 likes · 5 min read
Detecting Offline Merchant Service Issues Using Machine Learning and Big Data at Nuomi
Efficient Ops
Efficient Ops
Jul 11, 2016 · Operations

How Tencent's Intelligent Monitoring Transforms Ops Automation

Leveraging Tencent's extensive experience in social platform operations, this talk explores intelligent monitoring practices—covering active, passive, and side‑channel techniques, full‑link observability, data processing pipelines, and alert convergence—to enhance reliability, availability, and user experience while reducing noise for ops teams.

Operationsalert managementautomation
0 likes · 22 min read
How Tencent's Intelligent Monitoring Transforms Ops Automation
ITPUB
ITPUB
Jul 10, 2016 · Big Data

Can Spark Really Process Hundreds of Terabytes Interactively?

This article examines Apache Spark's interactive mode performance, revealing that while small datasets respond within seconds, processing beyond about 1 TB dramatically increases latency, and it discusses practical limits, hardware considerations, and the need to preload large datasets from disk.

Apache SparkPerformanceResponse Time
0 likes · 5 min read
Can Spark Really Process Hundreds of Terabytes Interactively?
Architecture Digest
Architecture Digest
Jul 5, 2016 · Big Data

Why Map‑Reduce Is Not the Solution to Your Big Data Problem – A Critical Look at Hadoop

The article reviews Hadoop’s origins from Google’s pioneering papers, explains its architecture and ecosystem, evaluates its strengths such as scalability and benchmarks, discusses current limitations like single‑point failures and complex programming, and outlines upcoming improvements including HDFS Federation and next‑generation MapReduce.

Distributed ComputingFutureHDFS
0 likes · 14 min read
Why Map‑Reduce Is Not the Solution to Your Big Data Problem – A Critical Look at Hadoop
ITPUB
ITPUB
Jun 29, 2016 · Big Data

Why OLTP Falls Short for Big Data: OLAP, Hadoop & MPP Explained

The article explains how traditional OLTP systems cannot satisfy modern big‑data analytics needs and compares OLAP, Hadoop, and MPP architectures, highlighting their data processing models, scalability, cloud‑based managed services, and practical recommendations for building effective data warehouses.

HadoopMPPOLAP
0 likes · 21 min read
Why OLTP Falls Short for Big Data: OLAP, Hadoop & MPP Explained
Qunar Tech Salon
Qunar Tech Salon
Jun 24, 2016 · Backend Development

Overview of Alibaba's Open Source Projects

This article provides a comprehensive overview of Alibaba's numerous open‑source projects, ranging from high‑performance service frameworks and databases to messaging middleware, frontend tools, testing platforms, and infrastructure utilities, highlighting their key features and typical use cases.

AlibabaBackendbig data
0 likes · 22 min read
Overview of Alibaba's Open Source Projects
Efficient Ops
Efficient Ops
Jun 19, 2016 · Operations

How Real‑Time Log Analysis Is Revolutionizing IT Operations

This article summarizes a 2016 Global Operations conference talk that explains the concept of IT Operations Analytics (ITOA), its four data sources, the evolution of log management from databases to real‑time search engines, and real‑world case studies demonstrating how fast, large‑scale log analysis improves monitoring, security, and business insight.

IT Operationsbig datalog-analysis
0 likes · 25 min read
How Real‑Time Log Analysis Is Revolutionizing IT Operations
21CTO
21CTO
Jun 18, 2016 · Databases

Unlock Ultra‑High Compression with HiStore’s Knowledge‑Grid Columnar Database

HiStore, Alibaba’s columnar database built on a patented Knowledge‑Grid, delivers ultra‑high compression (over 10:1, up to 40:1), low‑cost storage, rapid query performance, linear scalability, and seamless MySQL compatibility, making it ideal for massive OLAP workloads and real‑time analytics across diverse industries.

Columnar DatabaseOLAPbig data
0 likes · 8 min read
Unlock Ultra‑High Compression with HiStore’s Knowledge‑Grid Columnar Database
21CTO
21CTO
Jun 17, 2016 · Fundamentals

2016 Programmer Salary Survey: Who Earns the Most and Emerging Tech Trends

The 2016 programmer salary report reveals that front‑end, back‑end and mobile developers dominate the workforce, big‑data engineers command the highest pay, senior engineers see sharp salary jumps, and emerging technologies like Swift, WeChat, and Python shape future career choices.

BackendCareer TrendsMobile Development
0 likes · 8 min read
2016 Programmer Salary Survey: Who Earns the Most and Emerging Tech Trends
21CTO
21CTO
Jun 15, 2016 · Big Data

Choosing the Right Data Ingestion Tool: Flume, Fluentd, Logstash, and More

This article reviews major data collection platforms—including Apache Flume, Fluentd, Logstash, Chukwa, Scribe, and Splunk Forwarder—explaining their architectures, strengths, and limitations to help engineers select the most reliable and scalable solution for big‑data pipelines.

Apache FlumeData IngestionFluentd
0 likes · 10 min read
Choosing the Right Data Ingestion Tool: Flume, Fluentd, Logstash, and More

How BitMap Accelerates Active-Day Distribution Calculations in Big Data

BitMap, a space‑saving bit‑array structure, can replace costly I/O‑heavy Spark jobs for computing user active‑day distributions by converting joins and distinct operations into fast bitwise logic, enabling efficient 30‑day rolling metrics with minimal memory and superior performance, as demonstrated by real‑world benchmarks.

Active DaysPerformanceSpark
0 likes · 8 min read
How BitMap Accelerates Active-Day Distribution Calculations in Big Data
ITPUB
ITPUB
Jun 11, 2016 · Big Data

How 58 Daojia Leverages User Portraits to Boost Operations and Fight Fraud

This article details 58 Daojia's data‑driven approach to building user‑portrait tags, covering tag construction, evaluation, and practical applications such as personalized recommendations, anti‑fraud measures, coupon distribution, and dynamic pricing, while outlining the underlying big‑data architecture and technical challenges.

anti-fraudbig datadata mining
0 likes · 18 min read
How 58 Daojia Leverages User Portraits to Boost Operations and Fight Fraud
Architecture Digest
Architecture Digest
Jun 9, 2016 · Databases

Understanding HBase Architecture and Core Principles

This article provides a comprehensive overview of HBase, covering its distributed architecture, component roles, data organization, read/write mechanisms, and best practices for schema and region design to ensure efficient big‑data storage and retrieval.

ArchitectureData StorageHBase
0 likes · 17 min read
Understanding HBase Architecture and Core Principles
Hulu Beijing
Hulu Beijing
May 31, 2016 · Big Data

What’s New in Hadoop 3.0? Key Features and Improvements Explained

Hadoop 3.0, built on JDK 1.8, adds erasure‑coded HDFS, multi‑NameNode support, native MapReduce task optimizations, cgroup‑based YARN memory and disk isolation, and container resizing, with an alpha slated for summer and a GA release expected in November or December.

HDFSHadoopMapReduce
0 likes · 5 min read
What’s New in Hadoop 3.0? Key Features and Improvements Explained
dbaplus Community
dbaplus Community
May 26, 2016 · Big Data

Mastering Apache Parquet: Columnar Storage, Nested Data, and Performance Gains

This article explains Apache Parquet’s columnar storage format, its support for nested data models, the underlying striping/assembly algorithm, file structure, push‑down optimizations, and performance advantages within the Hadoop ecosystem, providing a comprehensive guide for big‑data practitioners.

Apache ParquetHadoopPerformance Optimization
0 likes · 22 min read
Mastering Apache Parquet: Columnar Storage, Nested Data, and Performance Gains
Architect
Architect
May 25, 2016 · Big Data

How Flink Manages Memory to Overcome JVM Limitations

The article explains how Flink tackles JVM memory challenges by using proactive memory management, a custom serialization framework, cache‑friendly binary operations, and off‑heap memory techniques to reduce GC pressure, avoid OOM, and improve performance in big‑data workloads.

FlinkJVMOff‑Heap
0 likes · 17 min read
How Flink Manages Memory to Overcome JVM Limitations
Architecture Digest
Architecture Digest
May 25, 2016 · Big Data

Advanced Spark Performance Optimization: Data Skew and Shuffle Tuning

This article provides a comprehensive guide on tackling Spark performance bottlenecks by diagnosing data skew, locating the offending stages and operators, and applying a range of practical solutions—including Hive pre‑processing, key filtering, shuffle parallelism, two‑stage aggregation, map‑join, and combined strategies—followed by an in‑depth discussion of shuffle manager evolution and key configuration parameters for fine‑tuning.

Data SkewShuffle OptimizationShuffleManager
0 likes · 35 min read
Advanced Spark Performance Optimization: Data Skew and Shuffle Tuning
58UXD
58UXD
May 18, 2016 · Product Management

Why Companies Doubt User Research—and How to Make It Truly Valuable

This article examines why many enterprises view user research as ineffective, outlines the four biggest challenges—defining clear goals, cultivating insight, building capable teams, and adopting the right mindset—and offers practical strategies for making research results actionable, integrating them into product development, and evolving the role of user researchers.

AgileUXbig data
0 likes · 14 min read
Why Companies Doubt User Research—and How to Make It Truly Valuable
Meituan Technology Team
Meituan Technology Team
May 13, 2016 · Big Data

Spark Performance Optimization Guide: Data Skew and Shuffle Tuning

This advanced Spark performance guide explains how data skew arises during shuffles and presents eight practical solutions—including Hive preprocessing, key filtering, increased shuffle parallelism, two‑stage aggregation, map joins, sampling, random prefixes, and combined strategies—while also detailing key shuffle‑tuning parameters such as spark.shuffle.file.buffer, spark.reducer.maxSizeInFlight, and spark.shuffle.manager to improve memory usage and execution speed.

Data SkewPerformance OptimizationShuffle Tuning
0 likes · 33 min read
Spark Performance Optimization Guide: Data Skew and Shuffle Tuning
Qunar Tech Salon
Qunar Tech Salon
May 13, 2016 · Big Data

Overview and Architecture of Hadoop Distributed File System (HDFS)

This article provides a comprehensive overview of Hadoop Distributed File System (HDFS), detailing its design goals, architecture components such as NameNode, DataNode and SecondaryNameNode, data block handling, replication strategies, communication protocols, and the read, write, and delete processes.

HDFSHadoopNameNode
0 likes · 18 min read
Overview and Architecture of Hadoop Distributed File System (HDFS)
Efficient Ops
Efficient Ops
May 12, 2016 · Operations

How Big Data Powers Precise IT Operations for Modern Enterprises

This article explains what big data is, outlines its four V characteristics, and describes how precise IT operations—aligning services with business needs—leverage big data analytics to improve service quality, predict user behavior, and enhance competitiveness for both traditional and internet enterprises.

Digital TransformationEnterprise ITIT Operations
0 likes · 15 min read
How Big Data Powers Precise IT Operations for Modern Enterprises
Architect
Architect
May 11, 2016 · Big Data

Comprehensive Guide to Hadoop MapReduce Job Execution, Scheduling, and Optimization

This article provides an in‑depth explanation of Hadoop MapReduce architecture, covering the roles of JobClient, JobTracker, TaskTracker and HDFS, the complete job lifecycle from submission to completion, scheduling strategies, shuffle and sort mechanisms, fault tolerance, and performance tuning techniques.

HadoopJobTrackerMapReduce
0 likes · 20 min read
Comprehensive Guide to Hadoop MapReduce Job Execution, Scheduling, and Optimization
Architecture Digest
Architecture Digest
May 7, 2016 · Fundamentals

Overview of Alibaba Open‑Source Projects and Tools

This article provides a comprehensive overview of numerous Alibaba open‑source projects, ranging from service frameworks like Dubbo and database tools such as Druid and OceanBase to front‑end libraries, distributed systems, testing platforms, and cloud utilities, each briefly described with links for further reference.

AlibabaJavabig data
0 likes · 27 min read
Overview of Alibaba Open‑Source Projects and Tools
Architect
Architect
May 6, 2016 · Big Data

Integrating Kylin, Mondrian, and Saiku to Build an OLAP Analysis Tool

This article describes how the Youzan data team combined Apache Kylin, Mondrian, and Saiku into a three‑layer OLAP system, covering background, component overviews, technical architecture, schema integration challenges, count‑distinct handling, Kylin‑specific SQL quirks, and practical solutions.

HBaseHiveKylin
0 likes · 12 min read
Integrating Kylin, Mondrian, and Saiku to Build an OLAP Analysis Tool
Baidu Intelligent Testing
Baidu Intelligent Testing
May 4, 2016 · Big Data

Understanding Big Data: The Importance of Data Breadth and User Profiling for Precise Marketing and Product Optimization

The article explains the core concepts of big data, emphasizing data breadth across product lines, illustrates how comprehensive user profiling can drive personalized marketing and product improvements, and provides practical examples of cross‑product data analysis in e‑commerce, finance, travel, and gaming contexts.

big datacross‑product analysisdata breadth
0 likes · 5 min read
Understanding Big Data: The Importance of Data Breadth and User Profiling for Precise Marketing and Product Optimization
Meituan Technology Team
Meituan Technology Team
Apr 29, 2016 · Big Data

Introduction to Spark in Big Data

Apache Spark, a versatile big‑data platform supporting batch processing, SQL queries, real‑time streaming, and machine‑learning workloads, dramatically accelerates data‑intensive jobs, as demonstrated by Meituan‑Dianping, where its high‑performance engine reduces execution times and enhances scalability across diverse analytical and operational pipelines.

Batch ProcessingSparkStreaming
0 likes · 1 min read
Introduction to Spark in Big Data
Architecture Digest
Architecture Digest
Apr 25, 2016 · Big Data

Curated Learning Resources for Spark and Scala Beginners

This article compiles a comprehensive list of tutorials, books, online courses, and tools to help beginners get started with Apache Spark and the Scala programming language, including setup instructions, code snippets, and links to free and paid learning materials.

Learning ResourcesScalaSpark
0 likes · 7 min read
Curated Learning Resources for Spark and Scala Beginners
ITPUB
ITPUB
Apr 24, 2016 · Big Data

12 Essential Hive Performance Tips for Faster Hadoop Queries

This guide presents twelve practical Hive tuning techniques—including avoiding MapReduce, limiting string concatenation, steering clear of subqueries, choosing the right file formats, managing vectorization, sizing containers, enabling statistics, and optimizing joins—to dramatically improve query speed on Hadoop.

HadoopHiveOptimization
0 likes · 7 min read
12 Essential Hive Performance Tips for Faster Hadoop Queries
Big Data and Microservices
Big Data and Microservices
Apr 21, 2016 · Information Security

How Can Banks Secure Big Data? Key Strategies for Protecting Customer Information

In the era of big data, banks face unprecedented information security challenges due to massive, valuable, and highly damaging data breaches, and must adopt encryption, flexible access control, rigorous auditing, DLP solutions, strict data management, and robust outsourcing controls to safeguard customer information.

Access ControlDLPbanking
0 likes · 10 min read
How Can Banks Secure Big Data? Key Strategies for Protecting Customer Information
Big Data and Microservices
Big Data and Microservices
Apr 19, 2016 · Industry Insights

Designing a Scalable Real‑Time Stock Prediction Architecture with Open‑Source Tools

This article outlines a reference architecture for a low‑latency, horizontally scalable real‑time stock prediction system built with open‑source components such as Spring Cloud Data Flow, Apache Geode, Spark MLlib, and Hadoop, and discusses data flow steps, simplified deployment, and algorithm choices for market forecasting.

big datamachine learningopen-source architecture
0 likes · 7 min read
Designing a Scalable Real‑Time Stock Prediction Architecture with Open‑Source Tools
Efficient Ops
Efficient Ops
Apr 17, 2016 · Operations

How CIOs Can Navigate Massive Technological and Industry Shifts

In this speech, former Chinese Ministry of Industry and Information Technology deputy minister Yang Xueshan outlines six strategic principles for CIOs—understanding major technological and industry trends, focusing on internal data, embracing fusion, connectivity, platforms, CPS, and intelligence, and taking practical, grounded actions to stay relevant.

CIODigital Transformationbig data
0 likes · 18 min read
How CIOs Can Navigate Massive Technological and Industry Shifts
21CTO
21CTO
Apr 16, 2016 · Databases

Optimizing HBase Log Queries: Index Design and RowKey Strategies

This article examines the challenges of storing and querying log data in HBase, outlines the drawbacks of custom indexing, and presents practical rowKey design, filter usage, and integration with external search engines to improve query performance.

HBaseLog StorageNoSQL
0 likes · 15 min read
Optimizing HBase Log Queries: Index Design and RowKey Strategies
21CTO
21CTO
Apr 14, 2016 · Big Data

How Meituan’s Data Architecture Powers Precise Mobile Marketing

This article details Meituan Dianping's data‑driven approach to precise marketing, describing the O2O marketing framework, a layered pyramid data system, profiling techniques, budget monitoring, and two real‑world case studies that together illustrate how big‑data technologies boost marketing efficiency on mobile platforms.

big datadata architecturemachine learning
0 likes · 12 min read
How Meituan’s Data Architecture Powers Precise Mobile Marketing
Efficient Ops
Efficient Ops
Apr 14, 2016 · Big Data

Why Big Data May Not Be the Gold Mine You Expect: Insights and Pitfalls

The article examines what big data really means, its core 4 V characteristics, current limitations in China, the overhyped value of data, the importance of business‑driven applications, and why starting from small, relevant data is essential for true predictive power.

Business IntelligenceData Analysisbig data
0 likes · 13 min read
Why Big Data May Not Be the Gold Mine You Expect: Insights and Pitfalls
Architect
Architect
Apr 10, 2016 · Big Data

Introduction to Flume NG: Architecture, Components, Configuration, and Best Practices

This article provides a comprehensive overview of Flume NG, covering its architecture, core components (source, channel, sink), reliability mechanisms, common deployment scenarios, installation steps, configuration examples, compilation instructions, and practical best‑practice recommendations for building robust log‑collection pipelines.

ApacheConfigurationData Ingestion
0 likes · 16 min read
Introduction to Flume NG: Architecture, Components, Configuration, and Best Practices
Architecture Digest
Architecture Digest
Apr 9, 2016 · Big Data

Practical Experience of Using Spark at Meituan: Platformization, ETL Templates, Feature Platform, Data Mining, and Real‑World Applications

This article describes how Meituan migrated from Hive‑SQL and MapReduce to Spark on YARN, built an interactive Zeppelin‑based development platform, created reusable ETL templates, constructed a Spark‑driven feature and data‑mining platform, and applied Spark to interactive user‑behavior analysis and large‑scale SEM services, highlighting performance gains and operational benefits.

Distributed ComputingETLMeituan
0 likes · 19 min read
Practical Experience of Using Spark at Meituan: Platformization, ETL Templates, Feature Platform, Data Mining, and Real‑World Applications
Big Data and Microservices
Big Data and Microservices
Apr 7, 2016 · Big Data

Turning Big Data into Actionable Security Visualizations: Process & Real‑World Cases

This article explains how to transform massive security‑related big data into clear visual insights, covering storytelling, data processing, visual encoding, design workflow, and two real‑world case studies that illustrate vulnerability mapping and internal traffic analysis for improved threat awareness.

Data Visualizationbig datadesign process
0 likes · 10 min read
Turning Big Data into Actionable Security Visualizations: Process & Real‑World Cases
dbaplus Community
dbaplus Community
Apr 6, 2016 · Fundamentals

Essential Open‑Source Technologies Every Engineer Should Know

This article provides a comprehensive, curated overview of the most influential open‑source software across the full technology stack—including operating systems, web servers, programming languages, frameworks, databases, big‑data tools, and development utilities—offering practical insights for engineers seeking to understand and adopt proven solutions.

big datadatabasesopen-source
0 likes · 24 min read
Essential Open‑Source Technologies Every Engineer Should Know
21CTO
21CTO
Apr 4, 2016 · Big Data

How Asana Scaled Its Data Infrastructure: From MySQL to Redshift & Hadoop

This article details Asana's evolution from a simple Python‑MySQL setup to a robust, scalable data platform using Redshift, Hadoop, Luigi, and modern BI tools, highlighting challenges, solutions, and lessons learned for building reliable data pipelines in fast‑growing startups.

Data InfrastructureETLHadoop
0 likes · 15 min read
How Asana Scaled Its Data Infrastructure: From MySQL to Redshift & Hadoop
dbaplus Community
dbaplus Community
Apr 3, 2016 · Big Data

How Asana Scaled Its Data Infrastructure: From MySQL to Redshift & Beyond

Facing rapid growth, Asana overhauled its data infrastructure—from a single‑machine MySQL setup to a Redshift‑backed warehouse, Hadoop‑based log processing, Luigi orchestration, and self‑service BI tools—highlighting the challenges, solutions, and future plans for scalable, reliable analytics.

Business IntelligenceData InfrastructureETL
0 likes · 16 min read
How Asana Scaled Its Data Infrastructure: From MySQL to Redshift & Beyond
Architect
Architect
Apr 3, 2016 · Big Data

Apache Flume NG Architecture, Core Concepts, and Practical Configuration Guide

This article introduces Apache Flume NG, a distributed and reliable log collection system, explains its core architecture components such as Event, Flow, Agent, Source, Channel, and Sink, and provides detailed configuration examples for various pipelines, including load‑balancing, failover, and integration with HDFS.

Apache FlumeConfigurationData Ingestion
0 likes · 12 min read
Apache Flume NG Architecture, Core Concepts, and Practical Configuration Guide
21CTO
21CTO
Mar 31, 2016 · Big Data

Inside Airbnb’s Massive Big Data Platform: Architecture, Lessons & Scaling Secrets

Airbnb’s engineering team outlines the evolution of its big‑data platform, detailing the philosophy behind its architecture, the dual “gold” and “silver” Hive clusters, migration to Mesos, use of Presto, Airpal, Airflow, and the performance and cost gains achieved through these design choices.

AirbnbAirflowbig data
0 likes · 11 min read
Inside Airbnb’s Massive Big Data Platform: Architecture, Lessons & Scaling Secrets
Big Data and Microservices
Big Data and Microservices
Mar 30, 2016 · Industry Insights

How Text Mining is Transforming the Securities Industry: Trends and Challenges

This article examines the rapid growth of structured and unstructured data in the securities sector, outlines text mining fundamentals, explores key algorithms and tools, and analyzes current industry services, investment communities, and professional solutions while highlighting existing challenges and future opportunities.

Natural Language ProcessingSentiment Analysisbig data
0 likes · 32 min read
How Text Mining is Transforming the Securities Industry: Trends and Challenges
Architect
Architect
Mar 29, 2016 · Big Data

Understanding Apache Storm Architecture, Stream Groupings, and the Acker Mechanism

This article provides a comprehensive overview of Apache Storm’s architecture, including the roles of Nimbus, Supervisor, and ZooKeeper, explains various stream groupings, details the Acker mechanism, and describes task execution, parallelism calculation, and internal data flow within the Storm cluster.

Apache StormReal-time Analyticsbig data
0 likes · 19 min read
Understanding Apache Storm Architecture, Stream Groupings, and the Acker Mechanism
Architecture Digest
Architecture Digest
Mar 28, 2016 · Big Data

Overview of the Hadoop Ecosystem and Modern Big Data Technologies

This article provides a comprehensive overview of Hadoop and its surrounding ecosystem, detailing core components, storage principles, key algorithms, and a wide range of modern big‑data technologies such as Spark, Flink, Kafka, NoSQL databases, and cloud‑based processing platforms.

Data ProcessingHadoopKafka
0 likes · 11 min read
Overview of the Hadoop Ecosystem and Modern Big Data Technologies
Big Data and Microservices
Big Data and Microservices
Mar 23, 2016 · Industry Insights

Inside the Securities Tech Revolution: Cloud, Microservices, and Big Data

The article examines the paradox of the Chinese securities industry—high demand for cutting‑edge trading, quantitative and high‑frequency systems versus outdated IT—while detailing the team’s FinTech startup approach, their Node.js/Docker/MongoDB stack, a cloud‑native trading platform, microservice architecture, big‑data pipelines, performance tuning, and DevOps practices.

Cloud ComputingDevOpsMicroservices
0 likes · 21 min read
Inside the Securities Tech Revolution: Cloud, Microservices, and Big Data
ITPUB
ITPUB
Mar 19, 2016 · Big Data

Inside HDFS: How NameNode and DataNode Manage Big Data Writes and Reads

This article explains the fundamentals of distributed file systems, focusing on Hadoop’s HDFS architecture, the separation of metadata and data via NameNode and DataNode, and detailed step‑by‑step write and read processes, including replication, fault recovery, and block splitting across nodes.

DataNodeHDFSNameNode
0 likes · 8 min read
Inside HDFS: How NameNode and DataNode Manage Big Data Writes and Reads
21CTO
21CTO
Mar 16, 2016 · Big Data

Inside Uber’s Tech: How Data, AI, and Cloud Power Ride‑Sharing in China

Uber’s CTO Thuan Pham revealed at a Chinese tech salon how the company’s global architecture, data‑center strategy, cloud partnership with Baidu, anti‑fraud machine‑learning models, map localization and big‑data analytics together enable a unified yet locally adapted ride‑sharing platform across China and the world.

Artificial IntelligenceCloud ComputingUber
0 likes · 17 min read
Inside Uber’s Tech: How Data, AI, and Cloud Power Ride‑Sharing in China
Architect
Architect
Mar 10, 2016 · Big Data

Analysis and Practice of a Real-Time Hadoop Data Security Solution

The article presents a detailed technical overview of Apache Eagle's real-time Hadoop data security architecture, covering distributed data collection, stream processing, metadata‑driven policy enforcement, machine‑learning‑based anomaly detection, and integration with Hadoop ecosystem components such as HBase, Kafka, and Storm.

Apache EagleHadoopbig data
0 likes · 25 min read
Analysis and Practice of a Real-Time Hadoop Data Security Solution
Architect
Architect
Mar 8, 2016 · Big Data

In‑Depth Analysis of Apache Kafka: Architecture, Core Concepts, and Benchmark

This article provides a comprehensive technical overview of Apache Kafka, covering its architecture, core concepts, design goals, comparison with other message queues, replication, consumer groups, delivery guarantees, and performance benchmarking, making it a valuable resource for big‑data engineers.

KafkaReplicationStreaming
0 likes · 30 min read
In‑Depth Analysis of Apache Kafka: Architecture, Core Concepts, and Benchmark
Architect
Architect
Mar 6, 2016 · Big Data

Clustering Geolocated User Events with DBSCAN and Spark

This article explains how to apply the DBSCAN clustering algorithm to geolocated user event data and leverage Apache Spark’s distributed processing with PairRDDs to efficiently identify frequent user regions, detect outliers, and build location‑based services such as personalized recommendations and security alerts.

ClusteringDBSCANSpark
0 likes · 8 min read
Clustering Geolocated User Events with DBSCAN and Spark