Tagged articles

Data Integration

293 articles · Page 3 of 3
DataFunTalk
DataFunTalk
May 19, 2022 · Big Data

SeaTunnel: Distributed Data Integration Platform and Its Application in Traffic Management

This article introduces Apache SeaTunnel, a distributed, high‑performance data integration platform built on Spark and Flink, outlines its technical features, workflow, and plugin ecosystem, and details a concrete traffic‑management use case involving incremental Oracle‑to‑warehouse data synchronization with Spark resources and scheduled shell scripts.

Apache FlinkApache SparkBig Data
0 likes · 12 min read
SeaTunnel: Distributed Data Integration Platform and Its Application in Traffic Management
Top Architect
Top Architect
May 11, 2022 · Databases

An Introduction to Change Data Capture (CDC) Practices

This article introduces the concept and practice of Change Data Capture (CDC), explaining why CDC is needed for real‑time analytics, how it works by capturing DML changes, modern approaches using transaction logs, and key considerations for building a production‑ready CDC system.

CDCChange Data CaptureData Integration
0 likes · 8 min read
An Introduction to Change Data Capture (CDC) Practices
Alibaba Cloud Native
Alibaba Cloud Native
Apr 20, 2022 · Cloud Native

How to Seamlessly Integrate Cloud Services with Alibaba EventBridge

This guide walks through Alibaba Cloud EventBridge’s event standardization, shows step‑by‑step how to integrate OSS events for automatic file unzipping, demonstrates custom event source filtering to store data in RDS, and explains using EventBridge event streams to route RocketMQ messages to MNS, complete with code snippets and configuration details.

Alibaba CloudCloud NativeData Integration
0 likes · 9 min read
How to Seamlessly Integrate Cloud Services with Alibaba EventBridge
DataFunTalk
DataFunTalk
Apr 20, 2022 · Big Data

OpenMLDB Pulsar Connector: A Real‑time Data Integration Guide

This article presents a step‑by‑step tutorial on using the OpenMLDB Pulsar Connector to stream real‑time data from Apache Pulsar into OpenMLDB, covering connector architecture, key features, Docker‑based installation, sink configuration, schema registration, message production, verification queries, and future roadmap details.

Apache PulsarConnectorData Integration
0 likes · 13 min read
OpenMLDB Pulsar Connector: A Real‑time Data Integration Guide
Alibaba Cloud Developer
Alibaba Cloud Developer
Apr 2, 2022 · Big Data

What’s New in Flink CDC 2.2? A Deep Dive into Added Sources and Core Features

The article introduces Flink CDC 2.2, highlighting its expanded support for twelve data sources—including OceanBase, PolarDB‑X, SqlServer, and TiDB—while detailing core features such as the incremental snapshot framework, multi‑version Flink compatibility, dynamic table addition, and numerous bug fixes and performance improvements.

Apache FlinkChange Data CaptureConnector
0 likes · 9 min read
What’s New in Flink CDC 2.2? A Deep Dive into Added Sources and Core Features
Architects Research Society
Architects Research Society
Mar 31, 2022 · R&D Management

The Strategic Role of Enterprise Architects and Their Key Focus Areas

Enterprise architects align IT strategy with business goals by managing application portfolios, technology risk, IT operations, security, data integration, and financial considerations, while balancing long‑term strategic planning with tactical execution in a rapidly changing environment.

Application Portfolio ManagementData IntegrationEnterprise Architecture
0 likes · 5 min read
The Strategic Role of Enterprise Architects and Their Key Focus Areas

Data Lake Construction and Practice at NetEase Yanxuan

NetEase Yanxuan replaced its cumbersome data‑warehouse with a flexible Delta‑Lake/Iceberg data lake, creating a unified metadata layer and real‑time ingestion pipelines that cut latency from nightly batches to seconds, slashed compute and storage costs, supported diverse business scenarios and machine‑learning feature engineering, and set the stage for broader future expansion.

Data IntegrationData LakeDelta Lake
0 likes · 16 min read
Data Lake Construction and Practice at NetEase Yanxuan
IT Xianyu
IT Xianyu
Mar 3, 2022 · Databases

Introducing SPL: An Open‑Source Structured Data Processing Language with Full SQL‑92 Capabilities

SPL is an open‑source structured data processing language that extends full SQL‑92 functionality to a wide range of data sources—including CSV, Excel, JSON, NoSQL and Hadoop—allowing developers to perform complex queries, multi‑step calculations, and mixed‑source analytics without a traditional relational database.

Big DataData IntegrationQuery Language
0 likes · 14 min read
Introducing SPL: An Open‑Source Structured Data Processing Language with Full SQL‑92 Capabilities
vivo Internet Technology
vivo Internet Technology
Feb 23, 2022 · Big Data

Kafka-based Real-Time Data Warehouse: Architecture and Practice for Search

The article explains how Kafka serves as the core of a real‑time data warehouse for search, detailing its advantages over traditional databases, integration with Flink for low‑latency stream processing, architectural patterns such as Lambda/Kappa, scaling challenges, and comprehensive monitoring using Kafka Eagle.

Apache KafkaData IntegrationFlink
0 likes · 15 min read
Kafka-based Real-Time Data Warehouse: Architecture and Practice for Search
DataFunTalk
DataFunTalk
Jan 28, 2022 · Big Data

Real-Time Customer Data Platform (RT‑CDP) Architecture and Implementation at iFanFan

This article explains the concept, challenges, and key business goals of a real‑time Customer Data Platform, details the technology stack selection—including Nebula Graph, Apache Flink, Apache Beam, Kudu, and Doris—and describes the modular architecture, data model, identity service, streaming computation, storage layers, rule engine, operational results, and future directions.

Big DataCDPData Integration
0 likes · 43 min read
Real-Time Customer Data Platform (RT‑CDP) Architecture and Implementation at iFanFan
Baidu Geek Talk
Baidu Geek Talk
Jan 26, 2022 · Big Data

How a Real‑Time CDP Solves Data Silos: Architecture, Tech Choices & Lessons

This article examines the design and implementation of a tenant‑level real‑time Customer Data Platform, detailing CDP fundamentals, business and technical challenges, key architectural components, technology selections such as graph databases, stream processing, storage engines, and the operational practices that enable high‑throughput, low‑latency data integration and analytics.

CDPData IntegrationFlink
0 likes · 42 min read
How a Real‑Time CDP Solves Data Silos: Architecture, Tech Choices & Lessons
IT Architects Alliance
IT Architects Alliance
Jan 25, 2022 · Operations

Design and Architecture of a Shared Resource Platform and Its Technical System

This document outlines the logical and technical architecture of a government shared resource platform, describing application system upgrades, data collection and analysis, multi‑layer system design, standards compliance, interface management, and overall system integration for improved service quality and decision support.

Big DataData IntegrationGovernment IT
0 likes · 23 min read
Design and Architecture of a Shared Resource Platform and Its Technical System
DataFunTalk
DataFunTalk
Jan 22, 2022 · Big Data

Alibaba Cloud Data Integration (DataX) Architecture, Design Principles, and Solution Overview

This presentation details Alibaba Cloud DataWorks Data Integration (DataX), covering its architecture, core design principles, offline and real‑time synchronization mechanisms, deployment modes, product positioning, use‑case scenarios, and its role within the broader DataWorks ecosystem, highlighting its capabilities for large‑scale data movement and processing.

Alibaba CloudBig DataData Integration
0 likes · 19 min read
Alibaba Cloud Data Integration (DataX) Architecture, Design Principles, and Solution Overview
21CTO
21CTO
Jan 6, 2022 · R&D Management

CTO’s Three Phases: Why Architecture, Data, and Management Can’t Be Unified

The article reflects on a CTO’s evolution through three stages, examines why architecture, data, servers and R&D management cannot be fully unified in large organizations, and offers practical guidance on innovation, team growth, and establishing a coherent R&D culture.

CTOData IntegrationInnovation
0 likes · 13 min read
CTO’s Three Phases: Why Architecture, Data, and Management Can’t Be Unified
DataFunTalk
DataFunTalk
Jan 6, 2022 · Artificial Intelligence

Deep Application‑Driven Construction of Medical Knowledge Graphs

This article presents a comprehensive overview of medical knowledge graph development, covering global and domestic progress, domain characteristics, a detailed seven‑piece ontology and "Huizhi" graph construction process, platform support, and real‑world healthcare applications such as intelligent alerts, guideline recommendations, and data reporting.

Data IntegrationMedical Knowledge Graphhealthcare
0 likes · 11 min read
Deep Application‑Driven Construction of Medical Knowledge Graphs
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 31, 2021 · Big Data

Apache SeaTunnel Joins the Apache Incubator: Overview, Features, and Real‑World Use Cases

SeaTunnel, the China‑originated data‑integration platform built on Spark and Flink, has been accepted into the Apache Incubator, and this article introduces its history, architecture, plugin ecosystem, deployment requirements, and numerous enterprise deployments across batch and streaming big‑data scenarios.

Big DataData IntegrationETL
0 likes · 7 min read
Apache SeaTunnel Joins the Apache Incubator: Overview, Features, and Real‑World Use Cases
DataFunSummit
DataFunSummit
Dec 28, 2021 · Artificial Intelligence

Deep Application‑Driven Construction of Medical Knowledge Graphs: Methods, Models, and Case Studies

This article presents a comprehensive overview of medical knowledge graph development, covering global and domestic progress, domain characteristics, a six‑step construction workflow—including schema design, ontology term set creation, and graph building—and showcases practical applications such as intelligent alerts, guideline recommendations, and data direct reporting.

Big DataData IntegrationMedical Knowledge Graph
0 likes · 11 min read
Deep Application‑Driven Construction of Medical Knowledge Graphs: Methods, Models, and Case Studies
Architects Research Society
Architects Research Society
Dec 23, 2021 · Fundamentals

Enterprise Integration: Challenges, Models, and Techniques

The article explains enterprise integration as the essential practice of connecting applications, data, and devices across distributed, cloud‑native environments, covering its evolution, key challenges, and core techniques such as messaging, application connectors, data flow platforms, integration patterns, and APIs.

APIsCloud NativeData Integration
0 likes · 7 min read
Enterprise Integration: Challenges, Models, and Techniques
Ctrip Technology
Ctrip Technology
Dec 16, 2021 · Big Data

Data Standard Management Practices in Ctrip Vacation Data Governance

This article outlines Ctrip Vacation's data standard management approach, covering why standards are needed, the three‑element framework of scope, tools, and policies, and detailed practices for data integration, production change handling, metadata governance, portal dashboard standardization, and self‑service query templating.

Big DataData GovernanceData Integration
0 likes · 12 min read
Data Standard Management Practices in Ctrip Vacation Data Governance
DataFunTalk
DataFunTalk
Dec 9, 2021 · Big Data

Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices

This article introduces the cloud‑native LakeHouse solution from China Mobile Cloud, covering its lake‑warehouse integration concept, overall architecture, core functions such as storage‑compute separation, one‑click data ingestion, intelligent metadata discovery, serverless execution, JDBC support, incremental updates, and typical application scenarios in public and private clouds.

Big DataCloud NativeData Integration
0 likes · 17 min read
Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices
Efficient Ops
Efficient Ops
Dec 6, 2021 · Operations

How Scenario‑Based AIOps Transforms IT Operations: Insights from GOPS 2023

The article summarizes a GOPS conference presentation by Dingmao Technology on AIOps scenario‑driven construction, detailing challenges, definition of scenarios, technical methods, roadmap planning, and future prospects, while showcasing practical examples and supporting technologies for intelligent IT operations.

AIOpsArtificial IntelligenceData Integration
0 likes · 8 min read
How Scenario‑Based AIOps Transforms IT Operations: Insights from GOPS 2023
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 8, 2021 · Big Data

Understanding Flink CDC 2.0: Core Design, Snapshot & Incremental Reading, and Code Walkthrough

This article introduces Flink CDC 2.0, explains its distributed full‑load and incremental reading mechanisms, details the slice partitioning, snapshot correction, and binlog handling logic, and provides a complete Java example that demonstrates how to configure Flink SQL, MySQL source, and Kafka sink.

Big DataCDCData Integration
0 likes · 29 min read
Understanding Flink CDC 2.0: Core Design, Snapshot & Incremental Reading, and Code Walkthrough
DataFunTalk
DataFunTalk
Sep 11, 2021 · Cloud Computing

Industrial Data Cloud Migration: Architecture, Core Technologies, and Case Studies with Alibaba Cloud IoT

This article explains the background, challenges, overall architecture, core technology optimizations, edge‑computing integration, data modeling, serialization, and real‑world case studies of moving industrial IoT data to Alibaba Cloud, illustrating how cloud‑native solutions enable digital transformation in manufacturing.

Big DataCloud ComputingData Integration
0 likes · 16 min read
Industrial Data Cloud Migration: Architecture, Core Technologies, and Case Studies with Alibaba Cloud IoT
Architects' Tech Alliance
Architects' Tech Alliance
Sep 2, 2021 · Big Data

Core Technologies and Architecture of a Big Data Platform

The article outlines a typical big data platform architecture, detailing its core layers—data collection, storage and analysis, sharing, application, real-time computation, and task scheduling—while describing key technologies such as Flume, DataX, HDFS, Hive, Spark, Spark Streaming, and Redis.

Data IntegrationHadoopSpark
0 likes · 9 min read
Core Technologies and Architecture of a Big Data Platform
Big Data Technology Architecture
Big Data Technology Architecture
Aug 17, 2021 · Big Data

Detailed Overview of Flink CDC 2.0: Architecture, Features, and Future Roadmap

This article provides an in‑depth technical overview of Flink CDC 2.0, covering its CDC fundamentals, comparison of query‑based and log‑based approaches, the new lock‑free chunk algorithm, FLIP‑27 based parallel snapshot reading, performance benchmarks, documentation improvements, and future roadmap for stability and ecosystem integration.

Change Data CaptureData IntegrationDebezium
0 likes · 16 min read
Detailed Overview of Flink CDC 2.0: Architecture, Features, and Future Roadmap
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 28, 2021 · Big Data

Understanding Customer Data Platforms (CDP): Why They’re Needed and How to Build One

The article explains what a Customer Data Platform (CDP) is, why businesses need it to overcome fragmented multi‑channel data, enable fine‑grained operations and data‑driven growth, and outlines the key steps for building a CDP, including data collection, OneID unification, tagging, lifecycle management, and marketing execution.

CDPCustomer Data PlatformData Integration
0 likes · 10 min read
Understanding Customer Data Platforms (CDP): Why They’re Needed and How to Build One
DataFunTalk
DataFunTalk
Jul 7, 2021 · Big Data

Solving Data Island Challenges and Enabling Advanced OLAP Analysis on Heterogeneous Big Data Platforms – Kyligence Solution Overview

This article explains the growing analytical demands in the big‑data era, the limitations of traditional OLAP, and how Kyligence’s distributed OLAP engine addresses data‑island issues, multi‑dimensional and many‑to‑many analysis, unified security, and performance optimization with MDX on Spark, delivering a seamless Excel‑like experience.

Big DataData IntegrationDistributed Computing
0 likes · 9 min read
Solving Data Island Challenges and Enabling Advanced OLAP Analysis on Heterogeneous Big Data Platforms – Kyligence Solution Overview
Laravel Tech Community
Laravel Tech Community
Jul 1, 2021 · Frontend Development

How to Build a Visual Dashboard (Large Screen) Using FineReport

This article walks through the complete process of creating a high‑impact visual dashboard for large‑screen displays using FineReport, covering tool selection, data preparation, report creation, design, visual polishing, and adding dynamic effects.

DashboardData IntegrationFineReport
0 likes · 5 min read
How to Build a Visual Dashboard (Large Screen) Using FineReport
Code Ape Tech Column
Code Ape Tech Column
Jul 1, 2021 · Backend Development

Master Spring Batch: Core Concepts, Architecture, and Practical Tips

This article provides a comprehensive guide to Spring Batch, covering its purpose, architecture, core components such as Job, Step, ItemReader/Writer/Processor, chunk processing, skip strategies, configuration tips, and common memory issues, all illustrated with code examples and diagrams.

Batch ProcessingChunk ProcessingData Integration
0 likes · 19 min read
Master Spring Batch: Core Concepts, Architecture, and Practical Tips
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jun 21, 2021 · Big Data

What Is a Big Data Platform and How to Design Its Architecture?

This article explains what a big data platform is, outlines its seven‑component overall architecture, details the technical stack from data sources to applications, and describes the key subsystems such as catalog management, data integration, governance, storage, processing, sharing, development, and analysis.

Data GovernanceData Integrationdistributed systems
0 likes · 11 min read
What Is a Big Data Platform and How to Design Its Architecture?
DataFunTalk
DataFunTalk
Jun 5, 2021 · Big Data

Building and Evolving a Data Service Platform for NetEase Cloud Music

The article details how NetEase Cloud Music co‑built a unified data service platform with NetEase YouShu, describing its architecture, phased development from internal use to online high‑concurrency services, feature enhancements such as API marketplace, multi‑source support, parameter conversion, and future roadmap for broader data products.

API PlatformBig DataData Integration
0 likes · 16 min read
Building and Evolving a Data Service Platform for NetEase Cloud Music
IT Architects Alliance
IT Architects Alliance
May 30, 2021 · Big Data

NetEase Game Streaming ETL Architecture and Practices Based on Flink

This article presents NetEase Game's Flink‑based streaming ETL system, detailing business background, log classifications, specialized and generic ETL services, Python UDF integration, runtime optimizations, HDFS write tuning, SLA metrics, fault‑tolerance mechanisms, and future roadmap for unified data lakes and PyFlink support.

Big DataData IntegrationETL
0 likes · 19 min read
NetEase Game Streaming ETL Architecture and Practices Based on Flink
DataFunTalk
DataFunTalk
Apr 23, 2021 · Big Data

Building and Evolving Zhihu’s Flink‑Based Data Integration Platform

This article details Zhihu’s transition from a Sqoop‑driven data integration system to a Flink‑centric platform, covering business scenarios, historical architecture, design goals, technology choices, performance optimizations, and future plans for unified streaming‑batch processing across diverse storage systems.

Batch ProcessingBig DataData Integration
0 likes · 14 min read
Building and Evolving Zhihu’s Flink‑Based Data Integration Platform
Big Data Technology & Architecture
Big Data Technology & Architecture
Mar 2, 2021 · Big Data

An Introduction to Kafka Connect: Architecture, Components, and Hands‑On Setup

This article introduces Kafka Connect, explaining its purpose as a scalable and reliable tool for moving data between Apache Kafka and external systems, detailing its core concepts, architecture, deployment modes, configuration files, and a step‑by‑step example that streams data from a file source to a file sink.

Data IntegrationETLKafka Connect
0 likes · 12 min read
An Introduction to Kafka Connect: Architecture, Components, and Hands‑On Setup
DataFunSummit
DataFunSummit
Dec 13, 2020 · Big Data

Data Services: Definition, Value, Lifecycle, Classification and Construction Guidelines

The article explains how traditional point‑to‑point data integration leads to data quality, consistency and cost issues, introduces the concept of data services as a unified, reusable way to provide data, outlines their benefits, lifecycle stages, classification into data‑set and API services, and presents Huawei’s practical construction strategy and the “Three‑1s” supply‑chain goals.

Data Integrationdata servicesservice lifecycle
0 likes · 23 min read
Data Services: Definition, Value, Lifecycle, Classification and Construction Guidelines
Architects Research Society
Architects Research Society
Dec 13, 2020 · Backend Development

Understanding Ballerina’s Native Data Types, Parallel Processing, and Development Tools

This article introduces Ballerina’s unique language features, including native XML/JSON data types, datatable handling, inline definitions, parallel processing with workers and fork‑join, and the comprehensive development toolset such as Composer, Testerina, connectors, and editor plugins, illustrating code examples throughout.

BallerinaData IntegrationFork‑Join
0 likes · 11 min read
Understanding Ballerina’s Native Data Types, Parallel Processing, and Development Tools
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 28, 2020 · Big Data

ETL Fundamentals and Introduction to Kettle (Pentaho Data Integration)

This article provides an in-depth overview of ETL concepts, including extraction, transformation, loading, data warehouse architecture, and detailed discussion of Kettle (Pentaho Data Integration) features, design principles, components, transformations, jobs, database connections, metadata management, and practical examples for building robust data integration pipelines.

Data IntegrationData WarehouseETL
0 likes · 57 min read
ETL Fundamentals and Introduction to Kettle (Pentaho Data Integration)
DataFunTalk
DataFunTalk
Nov 17, 2020 · Artificial Intelligence

Alink: A Flink‑Based Machine Learning Platform – Overview, Features, and Quick‑Start Guide

This article introduces Alink, Alibaba's open‑source machine‑learning platform built on Flink, explains its core algorithms, performance comparison with Spark ML, version‑wise feature evolution, and provides practical quick‑start instructions for both Java (Maven) and Python (PyAlink) users, including data source handling, type conversion components, unified file‑system operations, and an overview of its FM algorithm implementation.

AlinkBatch ProcessingData Integration
0 likes · 13 min read
Alink: A Flink‑Based Machine Learning Platform – Overview, Features, and Quick‑Start Guide
360 Tech Engineering
360 Tech Engineering
Nov 6, 2020 · Big Data

Guide to Flink SQL: Features, Scenarios, and Productization

Flink SQL, the high‑level SQL interface for Apache Flink, offers language‑independent, dependency‑free, easy‑to‑use stream processing with advanced features such as DDL, UDFs, time semantics, windowing, pattern matching, and built‑in connectors, supporting data synchronization, batch‑stream fusion, Hive integration, and various product enhancements.

Data IntegrationFlinkHive
0 likes · 11 min read
Guide to Flink SQL: Features, Scenarios, and Productization
Architects Research Society
Architects Research Society
Aug 20, 2020 · Big Data

Differences Between Talend and Pentaho ETL Tools

The article explains the fundamentals of ETL, compares Talend and Pentaho in terms of openness, connectivity, support, performance, GUI usability, deployment flexibility, and cost, and concludes with guidance on choosing the appropriate tool based on specific business and technical requirements.

ComparisonData IntegrationETL
0 likes · 7 min read
Differences Between Talend and Pentaho ETL Tools
Qunar Tech Salon
Qunar Tech Salon
Jun 3, 2020 · Fundamentals

Optimizing International Hotel Data Aggregation Algorithms at Qunar

The article outlines Qunar’s challenges in aggregating international hotel data, analyzes issues such as localized address formats and limited text similarity parsing, and presents a pattern‑matching and weighted scoring approach that improves aggregation accuracy across multiple countries.

Data IntegrationPattern Matchingalgorithm optimization
0 likes · 7 min read
Optimizing International Hotel Data Aggregation Algorithms at Qunar
Meituan Technology Team
Meituan Technology Team
May 28, 2020 · Big Data

Design and Implementation of Meituan Delivery A/B Testing Platform and Evaluation System

The article details Meituan Delivery’s A/B testing platform and evaluation system, explaining its closed‑loop design, multi‑strategy traffic allocation with AA grouping, comprehensive metric hierarchy, statistical rigor, data integration, and implementation architecture, and outlines future tools for traffic‑volume recommendation.

A/B testingData Integrationevaluation system
0 likes · 20 min read
Design and Implementation of Meituan Delivery A/B Testing Platform and Evaluation System
Big Data Technology & Architecture
Big Data Technology & Architecture
May 20, 2020 · Big Data

Technical Overview of Real-time Data Platform (RTDP) Architecture and Component Selection

This article presents a comprehensive technical overview of the Real-time Data Platform (RTDP), detailing its overall architecture, component selection—including DBus, Kafka, Wormhole, Moonbox, and Davinci—design philosophies, functional features, and various deployment patterns such as synchronous, stream-processing, rotation, and intelligent modes.

Data GovernanceData Integration
0 likes · 26 min read
Technical Overview of Real-time Data Platform (RTDP) Architecture and Component Selection
Amap Tech
Amap Tech
Apr 10, 2020 · Backend Development

Platformization of POI Deep Information Integration at Amap: Design and Implementation

Amap transformed its fragmented POI deep‑information pipelines into a unified platform that automates data acquisition, parsing, dimension alignment, specification mapping, and lifecycle management across billions of records, enabling product managers to integrate, debug, and scale diverse content‑provider feeds with real‑time, end‑to‑end control.

Big DataConversion EngineData Integration
0 likes · 13 min read
Platformization of POI Deep Information Integration at Amap: Design and Implementation
DataFunTalk
DataFunTalk
Apr 2, 2020 · Artificial Intelligence

Building and Applying an Industry Knowledge Graph: Lessons from Beike Real Estate

The article explains how Beike Real Estate constructs an industry knowledge graph by integrating internal and external data, outlines the technical framework and data processing steps, and demonstrates its AI-driven applications such as intelligent Q&A, recommendation, and decision support for the real‑estate market.

AI ApplicationsData IntegrationKnowledge Graph
0 likes · 8 min read
Building and Applying an Industry Knowledge Graph: Lessons from Beike Real Estate
21CTO
21CTO
Feb 19, 2020 · Big Data

Building an Open-Source Big Data Analytics Stack: Challenges & Benefits

The article explains why modern companies rely on data‑driven decisions, outlines the two main challenges of tracking data and connecting it to BI, describes the three‑step analytics stack (integration, warehouse, analysis), and highlights the cost, flexibility, and security advantages of open‑source tools.

Big DataData AnalyticsData Integration
0 likes · 5 min read
Building an Open-Source Big Data Analytics Stack: Challenges & Benefits
Java High-Performance Architecture
Java High-Performance Architecture
Jan 7, 2020 · Backend Development

How to Build a Scalable Reporting Service in a Microservice Architecture

To generate a user‑enriched order report in a microservice system, the article compares four approaches—direct DB access, REST data aggregation, batch pulling, and an event‑driven model—highlighting their trade‑offs in coupling, performance, scalability, and resilience, and recommends the event‑push solution.

Data IntegrationKafkaarchitecture
0 likes · 5 min read
How to Build a Scalable Reporting Service in a Microservice Architecture
HomeTech
HomeTech
Dec 12, 2019 · Big Data

Architecture and Design of the Home Data Integration Governance Platform

The article describes the background, architecture, and design principles of a unified big‑data scheduling and data‑exchange platform, detailing its data ingestion “direct‑train”, centralized scheduling engine, and DataX‑based data‑exchange components along with monitoring, alerting, and security features.

Big DataData IntegrationDataX
0 likes · 7 min read
Architecture and Design of the Home Data Integration Governance Platform
Architects Research Society
Architects Research Society
Oct 23, 2019 · Big Data

Talend Performance Tuning Strategy: Identifying and Eliminating Bottlenecks

This article presents a structured, repeatable approach for Talend data‑integration jobs that guides readers through pinpointing performance bottlenecks, testing individual pipeline stages, and applying targeted optimizations to sources, targets, and transformations to achieve higher throughput and more reliable ETL processes.

Data IntegrationETLTalend
0 likes · 9 min read
Talend Performance Tuning Strategy: Identifying and Eliminating Bottlenecks
YooTech Youzu Tech Team
YooTech Youzu Tech Team
Oct 16, 2019 · Product Management

How I Built an Automated Financial Reporting System for Global Game Platforms

This article details the end‑to‑end design and implementation of a custom tool—named “Crystal Palace”—that automates financial reporting across App Store, Google Play, Facebook and Amazon, turning a tedious manual reconciliation process into a scalable, data‑driven solution for game publishers.

AutomationData Integrationfinancial reporting
0 likes · 6 min read
How I Built an Automated Financial Reporting System for Global Game Platforms
Snowball Engineer Team
Snowball Engineer Team
Sep 24, 2019 · Big Data

Snowball Data Middle Platform (AIBO): Architecture, Capabilities, and Future Outlook

The article introduces Snowball's AIBO data middle platform, detailing its storage‑compute separation architecture, core capabilities such as data integration, catalog, tagging, analysis tools, micro‑service data APIs, and outlines future enhancements for security, lineage, and continuous business‑driven iteration.

Big DataData AnalysisData Integration
0 likes · 12 min read
Snowball Data Middle Platform (AIBO): Architecture, Capabilities, and Future Outlook
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Sep 3, 2019 · Big Data

QuickSQL: 360’s Unified Multi-Source Query Engine Explained

This article outlines how 360’s data center built QuickSQL, a federated SQL engine that unifies queries across heterogeneous sources such as Hive, MySQL, and Elasticsearch, detailing the business challenges, architectural design, performance benchmarks, and future roadmap for multi‑source data analysis.

Big DataData IntegrationFederated Query
0 likes · 12 min read
QuickSQL: 360’s Unified Multi-Source Query Engine Explained
Beike Product & Technology
Beike Product & Technology
Aug 29, 2019 · Big Data

TiSpark Integration with TiDB/TiKV for Efficient Data Synchronization and OLAP in the Databus Project

This article introduces TiSpark—an extension of Spark that tightly integrates with TiDB/TiKV to enable high‑performance, scalable data synchronization and OLAP queries, details its architecture, key configuration, performance advantages over Spark SQL and Sqoop, and outlines its role in the Databus data‑integration platform.

Big DataData IntegrationSpark
0 likes · 10 min read
TiSpark Integration with TiDB/TiKV for Efficient Data Synchronization and OLAP in the Databus Project
360 Tech Engineering
360 Tech Engineering
Aug 27, 2019 · Databases

Quicksql: A Unified Cross‑Data‑Source SQL Query Engine

Quicksql is an open‑source, cross‑data‑source SQL engine built on Apache Calcite that provides a unified, safe, and fast SQL interface, enabling users to query heterogeneous storage systems such as Hive, MySQL, Elasticsearch, and Druid through command‑line, API, or JDBC connections.

Apache CalciteData IntegrationOpen Source
0 likes · 6 min read
Quicksql: A Unified Cross‑Data‑Source SQL Query Engine
Architects' Tech Alliance
Architects' Tech Alliance
Aug 5, 2019 · Industry Insights

Why Customer Data Platforms Are Redefining Modern Marketing

The article examines how fragmented SaaS marketing stacks limit real‑time data use, explains the evolution from early CRM to marketing automation, highlights the shortcomings of MQL models, and shows how Customer Data Platforms (CDPs) restore data continuity, scalability, and campaign effectiveness.

Customer Data PlatformData IntegrationDigital Marketing
0 likes · 9 min read
Why Customer Data Platforms Are Redefining Modern Marketing
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 2, 2019 · Big Data

Integrating Apache Flink with Apache Pulsar for Scalable Elastic Data Processing

This article explains how Apache Pulsar and Apache Flink can be combined to provide a unified, scalable, and fault‑tolerant data processing platform, covering Pulsar's architecture, its differences from other messaging systems, various integration patterns, and concrete code examples for stream and batch workloads.

Apache FlinkApache PulsarBig Data
0 likes · 13 min read
Integrating Apache Flink with Apache Pulsar for Scalable Elastic Data Processing
Architecture Digest
Architecture Digest
May 13, 2019 · Artificial Intelligence

Enterprise Knowledge Graphs: Development Trends, Use Cases, Database Selection, and Implementation Practices

This article outlines the evolution of knowledge graphs, describes typical enterprise application scenarios, compares graph database options such as Neo4j, Cayley and Dgraph, and presents a six‑step methodology for building, storing, and applying knowledge graphs in large‑scale business environments.

Data IntegrationEnterprise AIGraph Database
0 likes · 13 min read
Enterprise Knowledge Graphs: Development Trends, Use Cases, Database Selection, and Implementation Practices
37 Interactive Technology Team
37 Interactive Technology Team
Mar 28, 2019 · Big Data

Approaches to Building a Basic Data Platform

To handle terabytes of daily data and diverse business needs, the company built a three‑layer basic data platform—collection/computation/storage, unified data management, and API‑driven services—augmented by a standardized collection system, a robust Domino scheduler, and a self‑service analysis tool, aiming to evolve into a full data‑middle‑office for end‑to‑end intelligence.

Data IntegrationSchedulingSelf‑service analytics
0 likes · 8 min read
Approaches to Building a Basic Data Platform
Beike Product & Technology
Beike Product & Technology
Feb 21, 2019 · Big Data

DATABUS Data Integration Platform: Architecture, Capabilities, and TiDB Ecosystem

The article presents an in‑depth overview of the DATABUS data integration platform, detailing its background, current challenges, core capabilities such as data syncing, metadata automation, real‑time subscriptions, and its reliance on TiDB, TiSpark, Hudi, and related big‑data technologies to enable near‑real‑time data warehousing.

Big DataData IntegrationHive
0 likes · 13 min read
DATABUS Data Integration Platform: Architecture, Capabilities, and TiDB Ecosystem
360 Tech Engineering
360 Tech Engineering
Dec 28, 2018 · Databases

Quicksql: A Unified, Secure, and Fast Cross-Data-Source SQL Query Engine

Quicksql is an open‑source unified SQL query engine that simplifies and secures cross‑data‑source queries by providing a consistent ANSI‑based language, automatic engine selection, and support for mixed queries across Hive, MySQL, Elasticsearch, and other platforms, reducing learning and integration costs.

Data IntegrationOpen SourceSQL Engine
0 likes · 6 min read
Quicksql: A Unified, Secure, and Fast Cross-Data-Source SQL Query Engine
Efficient Ops
Efficient Ops
Dec 24, 2018 · Operations

How Baidu’s Noah Platform Unifies Ops Data with Pull, Push, and Lazy ETL

This article explains how Baidu Cloud's Noah intelligent operations product builds a unified operations knowledge base by categorizing metadata, status, and event data and applying three ETL approaches—Pull, Push, and Lazy—to handle offline, near‑line, and real‑time data integration.

Cloud ComputingData IntegrationETL
0 likes · 8 min read
How Baidu’s Noah Platform Unifies Ops Data with Pull, Push, and Lazy ETL
Youzan Coder
Youzan Coder
Aug 31, 2018 · Big Data

Evolution of Youzan Search Platform Architecture: From 1.0 to 4.0

The Youzan Search Platform evolved from a simple Elasticsearch cluster in 2015 to a modular, message‑driven architecture with proxy validation, caching, and management tools, and now plans a cloud‑native, Kubernetes‑based 4.0 version that automates data sync, isolates workloads, and scales elastically to support billions of records.

Data IntegrationElasticsearchSearch Architecture
0 likes · 14 min read
Evolution of Youzan Search Platform Architecture: From 1.0 to 4.0
Zhongtong Tech
Zhongtong Tech
Aug 31, 2018 · Databases

How Aries Uses MySQL GTID Binlog to Power Real‑Time Data Sync at Scale

Aries, an internally built MySQL incremental log distribution platform, leverages GTID‑based binlog dumping to achieve stable, consistent, and real‑time data synchronization across heterogeneous systems, supporting use cases such as Elasticsearch sync, cache updates, archiving, and live statistics.

Data IntegrationDatabaseGTID
0 likes · 7 min read
How Aries Uses MySQL GTID Binlog to Power Real‑Time Data Sync at Scale
dbaplus Community
dbaplus Community
Aug 8, 2018 · Big Data

How to Build a Real‑Time Data Platform: Tech Stack & Design Patterns

This article explains the architecture of a Real‑Time Data Platform (RTDP), details the technical selection of core components such as DBus, Kafka, Wormhole, Moonbox and Davinci, and discusses data management, security, operations, and four deployment modes—synchronization, flow, rotation and intelligent—illustrating how each fits different business scenarios.

Big Data ArchitectureData IntegrationKafka
0 likes · 24 min read
How to Build a Real‑Time Data Platform: Tech Stack & Design Patterns
58 Tech
58 Tech
Jun 27, 2018 · Big Data

Overview of the 58 User Profile System Architecture and Data Processing

The article describes the design, data integration, ID mapping, tag generation, and application scenarios of the 58 user profiling platform, which aggregates billions of user IDs across multiple business lines to provide online and offline persona data for personalization, analytics, and AI modeling.

Big DataData IntegrationID mapping
0 likes · 12 min read
Overview of the 58 User Profile System Architecture and Data Processing
ITPUB
ITPUB
Nov 23, 2017 · Big Data

7 Typical Big Data Projects Every Hadoop Engineer Should Know

The article outlines seven common big‑data initiatives—data integration, specialized analytics, Hadoop‑as‑a‑service, stream processing, complex event handling, ETL pipelines, and SAS replacement—explaining their goals, typical technologies such as HDFS, Hive, Spark, Storm, Kafka, and practical considerations for enterprises adopting Hadoop ecosystems.

Data IntegrationHadoopStream Processing
0 likes · 8 min read
7 Typical Big Data Projects Every Hadoop Engineer Should Know
Efficient Ops
Efficient Ops
Sep 25, 2017 · Operations

How Qunar Scaled Application Ops Automation from Hundreds to Tens of Thousands of Servers

This article details Qunar's journey of automating application operations, covering the evolution of their host‑management system, unified monitoring/alert platform, and data‑interchange mechanisms that enabled the company to grow from a few hundred to over ten thousand servers with a stable six‑person ops team.

Data IntegrationMonitoringOperations Automation
0 likes · 25 min read
How Qunar Scaled Application Ops Automation from Hundreds to Tens of Thousands of Servers
StarRing Big Data Open Lab
StarRing Big Data Open Lab
Jul 28, 2017 · Big Data

How Transwarp Transporter Enables Near‑Real‑Time ETL in Big Data Pipelines

The article introduces Transwarp Transporter, a near‑real‑time ETL tool for TDH 5.x, explains its architecture, visual dashboard, drag‑and‑drop data‑flow design, debugging features, parameter management, and highlights how it empowers business users to achieve fast, reliable data migration in big‑data environments.

Data IntegrationData PipelineETL
0 likes · 7 min read
How Transwarp Transporter Enables Near‑Real‑Time ETL in Big Data Pipelines
Architecture Digest
Architecture Digest
Jul 22, 2017 · Big Data

Popular Big Data Tools and Their Descriptions

This article provides an extensive overview of more than ninety open‑source and commercial big‑data tools—including ETL platforms, resource managers, storage systems, messaging queues, processing engines, and visualization libraries—detailing their core functions, typical use cases, and notable adopters.

Big DataData IntegrationETL
0 likes · 26 min read
Popular Big Data Tools and Their Descriptions
Alibaba Cloud Developer
Alibaba Cloud Developer
Mar 7, 2017 · Big Data

Unified Data Platforms: How UMENG+ Redefines Big Data Strategy

The article explores the evolution of big‑data applications in China, from Oracle’s trend report and the concept of "omni‑domain data" to UMENG+’s technical architecture, unified tech stack, AI integration, and future directions for delivering real customer value.

Big DataData AnalyticsData Integration
0 likes · 12 min read
Unified Data Platforms: How UMENG+ Redefines Big Data Strategy

The Growing Role of Apache Kafka in Modern Big Data Architectures

The article explains how Apache Kafka has become a pivotal, high‑scalable publish‑subscribe system in the big‑data ecosystem, addressing the limitations of traditional databases, enabling real‑time data integration across specialized distributed systems, and shaping future data‑governance practices.

Apache KafkaData IntegrationStreaming
0 likes · 7 min read
The Growing Role of Apache Kafka in Modern Big Data Architectures
Architect
Architect
Jul 6, 2015 · Big Data

Understanding Logs: The Core of Distributed Systems and Data Integration

This article explains how logs—simple, append‑only, time‑ordered records—serve as the fundamental abstraction behind databases, distributed systems, data integration pipelines, and stream‑processing platforms like Kafka and Hadoop, illustrating their role in ordering, replication, scalability, and real‑time analytics.

Data IntegrationHadoopKafka
0 likes · 48 min read
Understanding Logs: The Core of Distributed Systems and Data Integration