Tagged articles

metadata

217 articles · Page 2 of 3
DataFunSummit
DataFunSummit
Apr 3, 2023 · Big Data

Evolution and Architecture of Data Lineage in Volcano Engine DataLeap

This article outlines the background, development stages, architectural evolution, key features such as incremental updates and quality metrics, and future directions of the data lineage capability within Volcano Engine's DataLeap big‑data governance platform.

DataLeapbig datametadata
0 likes · 18 min read
Evolution and Architecture of Data Lineage in Volcano Engine DataLeap
DataFunTalk
DataFunTalk
Mar 15, 2023 · Big Data

Evolution of Next‑Generation Cloud Data Platform Architecture

This technical presentation reviews the historical development of big data platforms, outlines the four generations of cloud data platform architectures, details the modern cloud‑native stack—including unified metadata, scheduling, and integration systems—and showcases a real‑world industrial manufacturing case with a Q&A session.

Cloud Data Platformdata architecturemetadata
0 likes · 23 min read
Evolution of Next‑Generation Cloud Data Platform Architecture
Big Data Technology & Architecture
Big Data Technology & Architecture
Mar 14, 2023 · Big Data

Comprehensive Guide to Data Lineage: Model Design, Optimization, and Use Cases at ByteDance

This article presents an in‑depth overview of data lineage at ByteDance, detailing the design of storage, display, abstraction, implementation, and storage layers, optimization techniques for real‑time updates and queries, open export methods, practical use cases across asset, development, governance, and security domains, and future directions.

Apache AtlasJanusGraphdata lineage
0 likes · 20 min read
Comprehensive Guide to Data Lineage: Model Design, Optimization, and Use Cases at ByteDance
DataFunSummit
DataFunSummit
Feb 11, 2023 · Big Data

Intelligent Metadata Governance for Power Data: Background, Solution, Value and Case Studies

This article presents a comprehensive overview of the intelligent metadata‑driven data governance framework implemented by Southern Power Grid Yunnan, detailing its background, challenges, architectural design, key AI‑enabled technologies, practical case studies, and the resulting business value for the power industry.

AIData Qualityelectric power
0 likes · 14 min read
Intelligent Metadata Governance for Power Data: Background, Solution, Value and Case Studies
DataFunTalk
DataFunTalk
Jan 19, 2023 · Big Data

Data Governance Strategies: Concepts, Practices, and Case Studies

The article explains the importance of data governance for organizations handling big data, outlines narrow and broad governance approaches, presents strategic design principles, and shares practical case studies from leading companies, while also offering a downloadable ebook of governance strategies.

Case Studiesdata managementdata security
0 likes · 7 min read
Data Governance Strategies: Concepts, Practices, and Case Studies
DataFunTalk
DataFunTalk
Jan 13, 2023 · Big Data

Data Governance Strategies and Practices: Insights from Leading Companies

The article explains the importance of data governance for organizations handling big data, distinguishes narrow and broad governance approaches, outlines strategic principles, and presents case studies from companies like Tencent, SF Tech, Huolala, and NetEase to illustrate effective governance practices.

Data QualityEnterprise Datacase study
0 likes · 8 min read
Data Governance Strategies and Practices: Insights from Leading Companies
DataFunTalk
DataFunTalk
Dec 5, 2022 · Big Data

Data Governance Practices at ZTO Express: Challenges, Solutions, and Future Plans

The article details ZTO Express's data governance journey, covering company background, drivers and goals, challenges such as data asset inventory, standardization, quality, and modeling, and outlines their multi‑layered governance framework, practical implementations in data quality, model and metadata, and future plans.

data platformlogisticsmetadata
0 likes · 17 min read
Data Governance Practices at ZTO Express: Challenges, Solutions, and Future Plans
DeWu Technology
DeWu Technology
Nov 30, 2022 · Big Data

Fundamentals and Implementation of Data Lineage in Big Data Environments

Data lineage in big‑data environments tracks how data moves and transforms—from source tables through SQL processing to final storage—enabling management tasks such as domain segmentation, performance tuning, anomaly detection, and dependency verification, with implementations ranging from simple regex extraction to robust AST parsing and optimization, as used by tools like Alibaba DataWorks and Apache Atlas.

ASTHiveSQL parsing
0 likes · 7 min read
Fundamentals and Implementation of Data Lineage in Big Data Environments
DataFunSummit
DataFunSummit
Nov 24, 2022 · Big Data

Metadata Management and Governance Practices at Wing Payment: Architecture, Techniques, and Future Outlook

This article explains how Wing Payment uses metadata as the foundation of its data‑governance practice, describing the challenges of data quality, efficiency, cost and security, the four‑step governance framework, the design of its metadata platform, and future directions such as multi‑source management and intelligent recommendation.

data securitymaster datametadata
0 likes · 18 min read
Metadata Management and Governance Practices at Wing Payment: Architecture, Techniques, and Future Outlook
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 22, 2022 · Big Data

Comprehensive Guide to Metadata Management, Data Quality, and Optimization in Big Data Systems

This article provides an in-depth overview of metadata concepts, their technical and business classifications, value in data management, applications such as data profiling and lineage, optimization techniques for compute and storage, lifecycle management, and comprehensive data quality assurance practices within large‑scale big data environments.

Optimizationbig-datadata-quality
0 likes · 38 min read
Comprehensive Guide to Metadata Management, Data Quality, and Optimization in Big Data Systems
Data Thinking Notes
Data Thinking Notes
Nov 16, 2022 · Big Data

Why Metadata Management Is Essential for Data Warehouses

This article explains the concept of metadata, its role in data warehouses, why managing metadata is critical for building, maintaining, and scaling data warehouse systems, and outlines practical steps, use cases, and tools for effective metadata management.

ETLdata governancedata warehouse
0 likes · 15 min read
Why Metadata Management Is Essential for Data Warehouses
Past Memory Big Data
Past Memory Big Data
Nov 15, 2022 · Big Data

How Uber Accelerated Presto Queries with Alluxio Local Cache

Uber processes over 500,000 daily Presto queries across 20 clusters handling more than 50 PB of data, and by deploying Alluxio Local Cache on NVMe disks they raised cache‑hit rates from roughly 65% to over 90% while addressing real‑time partition updates, node churn, and cache‑size constraints.

AlluxioPerformance Optimizationbig data
0 likes · 15 min read
How Uber Accelerated Presto Queries with Alluxio Local Cache
DataFunSummit
DataFunSummit
Nov 11, 2022 · Big Data

Tencent Oula Data Governance Platform: Architecture, Practices, and Solutions

The article presents an in‑depth overview of Tencent's Oula data governance platform, describing its construction goals, core capabilities, DataOps‑driven development workflow, unified metric store, data map services, and practical Q&A on asset health scoring and data lineage, illustrating a comprehensive end‑to‑end big‑data governance solution.

DataOpsTencentbig data platform
0 likes · 17 min read
Tencent Oula Data Governance Platform: Architecture, Practices, and Solutions
DataFunTalk
DataFunTalk
Nov 2, 2022 · Big Data

Tencent Oula Data Governance Platform: Architecture, Practices, and Solutions

Tencent's Oula platform, launched in 2019, provides a DataOps‑driven, end‑to‑end data governance solution covering data discovery, asset factory, metric platform, and governance engine, and the talk details its construction goals, data development governance, unified metric system, data map, and Q&A on asset health and lineage.

DataOpsMetric Platformdata platform
0 likes · 17 min read
Tencent Oula Data Governance Platform: Architecture, Practices, and Solutions
ITPUB
ITPUB
Oct 20, 2022 · Big Data

Will HDFS Be Replaced? Analyzing Its Drawbacks and Future Alternatives

The article examines why Hadoop's Distributed File System may become obsolete by detailing its three main shortcomings—deployment complexity, metadata memory limits, and high replication overhead—and explores how newer architectures and erasure coding could address these issues.

HDFSbig datadistributed file system
0 likes · 8 min read
Will HDFS Be Replaced? Analyzing Its Drawbacks and Future Alternatives
Past Memory Big Data
Past Memory Big Data
Oct 9, 2022 · Operations

How Cloud Music Scaled Data Governance: Practices, Metrics, and Lessons Learned

The article details Cloud Music’s data‑governance journey, covering early modeling standards, self‑service data tools, quality and metadata management, asset‑reuse improvements, and cost‑saving Spark optimizations, while sharing concrete metrics, processes, and the team’s systematic methodology.

Cost OptimizationQuality Managementcloud music
0 likes · 18 min read
How Cloud Music Scaled Data Governance: Practices, Metrics, and Lessons Learned
IT Services Circle
IT Services Circle
Oct 5, 2022 · Databases

Debugging a NullPointerException During ShardingSphere Startup Caused by TiDB View Metadata

The article details a NullPointerException that occurs when ShardingSphere starts, explains how null values in TiDB‑generated view metadata trigger the error, describes the step‑by‑step investigation and reproduction using TiUP, and discusses several possible fixes and a call for community contributions.

JavaNullPointerExceptionShardingSphere
0 likes · 4 min read
Debugging a NullPointerException During ShardingSphere Startup Caused by TiDB View Metadata
AsiaInfo Technology: New Tech Exploration
AsiaInfo Technology: New Tech Exploration
Sep 23, 2022 · Industry Insights

How Domain Model Metadata Boosts Business System Reuse and Efficiency

This article explores how structured domain model metadata, derived from domain‑driven design principles, can standardize business component descriptions, enable visual UML modeling, support low‑code and no‑code generation, and ultimately reduce development costs while accelerating delivery of enterprise support systems.

Domain-Driven DesignUMLlow-code
0 likes · 20 min read
How Domain Model Metadata Boosts Business System Reuse and Efficiency
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Sep 23, 2022 · Databases

How Baidu’s TafDB Achieves Trillion‑Scale Metadata Storage with Near‑Zero Latency

This article explores the design and engineering of Baidu’s TafDB, a distributed metadata database that powers cloud object and file storage, detailing its architecture, namespace evolution, transaction optimizations, garbage collection strategies, and clock mechanisms that enable trillion‑scale metadata and millions of QPS.

cloud storagemetadatascalability
0 likes · 19 min read
How Baidu’s TafDB Achieves Trillion‑Scale Metadata Storage with Near‑Zero Latency
ShiZhen AI
ShiZhen AI
Sep 7, 2022 · Big Data

Getting Started with DataHub: A One‑Stop Guide to Metadata Governance

This article walks you through the fundamentals of data governance, explains metadata management concepts, compares traditional tools with DataHub, and provides a step‑by‑step tutorial for installing Docker, Python, and DataHub 0.8.20 on CentOS 7, ingesting MySQL metadata, and exploring the UI.

DataHubDockerPython
0 likes · 19 min read
Getting Started with DataHub: A One‑Stop Guide to Metadata Governance
GuanYuan Data Tech Team
GuanYuan Data Tech Team
Aug 4, 2022 · Cloud Native

What Is a Cloud‑Native Data Platform? Architecture, Components, and Best Practices

This article explores the evolution and architecture of cloud‑native data platforms, covering their historical roots, modern components such as storage layers, ingestion, processing, metadata, and consumption, and offers practical guidance on selecting tools, designing pipelines, and implementing best‑practice strategies for scalable, flexible data infrastructure.

big-datacloud-nativedata architecture
0 likes · 41 min read
What Is a Cloud‑Native Data Platform? Architecture, Components, and Best Practices
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 18, 2022 · Big Data

Systematic Data Governance Practices in Meituan Accommodation Business

This article details Meituan's accommodation data governance team's evolution toward an automated, systematic, and standardized governance framework, covering background challenges, the conceptualization of a comprehensive governance system, its practical implementation across standardization, digitization, and systematization, and the resulting operational benefits and future directions.

Standardizationautomationmetadata
0 likes · 30 min read
Systematic Data Governance Practices in Meituan Accommodation Business
Big Data Technology Architecture
Big Data Technology Architecture
Jun 7, 2022 · Big Data

Multi-Modal Index in Apache Hudi 0.11.0: Design, Implementation, and Performance Benefits

This article explains the motivation, design principles, implementation details, and performance improvements of the new multi‑modal indexing subsystem introduced in Apache Hudi 0.11.0 for Lakehouse architectures, covering scalable metadata, ACID updates, fast lookups, file listing, data skipping, upsert performance, and future work.

Apache HudiIndexingmetadata
0 likes · 19 min read
Multi-Modal Index in Apache Hudi 0.11.0: Design, Implementation, and Performance Benefits
MaGe Linux Operations
MaGe Linux Operations
Jun 6, 2022 · Databases

What Is a Schema? From Databases to Kubernetes Explained

This article explains the concept of a schema—from its Greek origins and psychological meaning to its role as metadata in databases and Kubernetes—detailing different database schema models, Kubernetes resource definitions, and how to extend and register custom schemas in Go.

YAMLdatabasegolang
0 likes · 8 min read
What Is a Schema? From Databases to Kubernetes Explained
Laravel Tech Community
Laravel Tech Community
May 30, 2022 · Backend Development

Highlights of Apache Pulsar 2.10.0 Release: New Features and Bug Fixes

The Apache Pulsar 2.10.0 release introduces automatic cluster failover, lazy‑loading producers, new TableView support, enhanced broker interceptors, enriched client authentication, Etcd metadata storage, and numerous bug fixes, offering developers and operators a more flexible and performant messaging platform.

Apache PulsarClientbroker
0 likes · 7 min read
Highlights of Apache Pulsar 2.10.0 Release: New Features and Bug Fixes
Architect
Architect
May 25, 2022 · Big Data

Metadata Infrastructure and Governance in Bilibili's Data Platform

The article details how Bilibili built a unified metadata infrastructure—including a URN‑based model, collection pipelines, quality assurance, storage in TiDB/ES/HugeGraph, and query services—to support data discovery, lineage, impact analysis, and governance across its growing data platform.

ETLbig datadata catalog
0 likes · 21 min read
Metadata Infrastructure and Governance in Bilibili's Data Platform
Architects Research Society
Architects Research Society
May 24, 2022 · Big Data

Understanding Data Fabric: Key Pillars for Data & Analytics Leaders

The article explains the emerging concept of Data Fabric (data weaving), its design principles, how it integrates metadata, knowledge graphs, and AI/ML to automate data integration across hybrid and multi‑cloud environments, and outlines four essential pillars that leaders must master to deliver business value.

AI/MLData Fabricknowledge graph
0 likes · 8 min read
Understanding Data Fabric: Key Pillars for Data & Analytics Leaders
Bilibili Tech
Bilibili Tech
May 24, 2022 · Big Data

Metadata Infrastructure and Governance in Bilibili Data Platform

Bilibili’s data platform consolidates scattered metadata into a unified URN‑based model stored across TiDB, Elasticsearch, and HugeGraph, offering batch‑pull and embedded collection, flexible SQL‑like queries, comprehensive lineage mapping, and powering data‑map, lineage‑map, and impact‑analysis tools while planning expanded quality assurance and self‑service dictionaries.

SQL parsingdata governancedata lineage
0 likes · 21 min read
Metadata Infrastructure and Governance in Bilibili Data Platform
Liangxu Linux
Liangxu Linux
Apr 28, 2022 · Fundamentals

What Is an Inode? Understanding Linux File Metadata and Links

This article explains the concept of inodes in Unix/Linux filesystems, detailing their structure, stored metadata, size calculations, inode numbers, directory handling, hard and soft links, and the special behaviors that arise from separating file names from inode identifiers.

Hard LinkInodefilesystem
0 likes · 9 min read
What Is an Inode? Understanding Linux File Metadata and Links
Python Programming Learning Circle
Python Programming Learning Circle
Apr 8, 2022 · Fundamentals

A Comprehensive Guide to Python Decorators and Aspect-Oriented Programming (AOP)

This article explains the concept of Aspect‑Oriented Programming (AOP) and demonstrates how Python decorators—both function‑based and class‑based—can be used to implement AOP features such as pre‑ and post‑execution logic, handling arguments, preserving metadata with functools.wraps, and stacking multiple decorators.

AOPPythonaspect-oriented-programming
0 likes · 18 min read
A Comprehensive Guide to Python Decorators and Aspect-Oriented Programming (AOP)
dbaplus Community
dbaplus Community
Dec 22, 2021 · Fundamentals

How Xiaomi Built a Scalable Metadata Platform for Data Governance

This article details Xiaomi's end‑to‑end metadata platform, covering its three‑layer architecture, the evolution of full‑domain metadata, real‑time lineage, precise measurement, and how these capabilities enable data map, governance, cost control, and quality improvements for future business empowerment.

Data QualityXiaomidata governance
0 likes · 20 min read
How Xiaomi Built a Scalable Metadata Platform for Data Governance
Top Architect
Top Architect
Dec 13, 2021 · Big Data

Design and Implementation of BanYu's Big Data Access Control System

This article describes the evolution from an unsecured data warehouse to a comprehensive big‑data access control system at BanYu, detailing the background, data access methods, design goals, authentication and authorization mechanisms, policy configuration, integration with Metabase, and the overall workflow that balances security with efficiency.

Access ControlHiveLDAP
0 likes · 15 min read
Design and Implementation of BanYu's Big Data Access Control System
DataFunTalk
DataFunTalk
Dec 9, 2021 · Big Data

Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices

This article introduces the cloud‑native LakeHouse solution from China Mobile Cloud, covering its lake‑warehouse integration concept, overall architecture, core functions such as storage‑compute separation, one‑click data ingestion, intelligent metadata discovery, serverless execution, JDBC support, incremental updates, and typical application scenarios in public and private clouds.

Data IntegrationKubernetesServerless
0 likes · 17 min read
Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices
DataFunTalk
DataFunTalk
Nov 27, 2021 · Big Data

iQIYI Data Middle Platform: Architecture, Data Governance Practices, and Future Plans

The article details iQIYI’s data middle platform architecture and its comprehensive data governance practices, covering platform overview, data flow, unified standards, metadata management, production quality assurance, and future AI‑driven enhancements, illustrating how centralized data services improve reliability, efficiency, and security.

Data Qualitybig datadata governance
0 likes · 27 min read
iQIYI Data Middle Platform: Architecture, Data Governance Practices, and Future Plans
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 10, 2021 · Fundamentals

What Is Data Governance and How to Implement It: Concepts, Goals, Methodology, Tools, and Case Studies

This article explains data governance—from its definition and why it’s needed, to its goals, core components, PDCA‑based methodology, essential tools, and real‑world implementations at Meituan and Ant Financial—providing a comprehensive guide for organizations seeking to manage and leverage their data assets effectively.

data securitymetadata
0 likes · 25 min read
What Is Data Governance and How to Implement It: Concepts, Goals, Methodology, Tools, and Case Studies
Aikesheng Open Source Community
Aikesheng Open Source Community
Oct 14, 2021 · Databases

Understanding reload @@config_all vs reload @@metadata in Dble and How to Resolve Metadata Issues

This article explains the differences between the Dble commands reload @@config_all and reload @@metadata, analyzes why metadata errors occur after table changes without configuration updates, and provides step‑by‑step guidance to synchronize metadata correctly using these commands and check full @@metadata.

metadatareload
0 likes · 7 min read
Understanding reload @@config_all vs reload @@metadata in Dble and How to Resolve Metadata Issues
Architects' Tech Alliance
Architects' Tech Alliance
Sep 11, 2021 · Big Data

Understanding Data Warehouses: Definitions, Differences, Architecture, Modeling, and Best Practices

This article explains what a data warehouse is, contrasts it with traditional databases, outlines how to design and build a warehouse—including model selection, subject‑area definition, bus matrix, layering, and data quality—while also covering related concepts such as data middle platforms, data lakes, metadata, and modeling techniques.

Data ModelingData QualityETL
0 likes · 16 min read
Understanding Data Warehouses: Definitions, Differences, Architecture, Modeling, and Best Practices
Ctrip Technology
Ctrip Technology
Sep 9, 2021 · Big Data

Building Data Lineage at Ctrip: Architecture, Implementation, and Real‑World Applications

This article describes how Ctrip built a data lineage system for its big data platform, covering the concept of data lineage, collection methods, open‑source tools such as Apache Atlas and DataHub, the in‑house table‑level and field‑level solutions, implementation details for Hive, Spark and Presto, storage in JanusGraph, and practical applications in data governance, metadata management, scheduling and sensitivity labeling.

HiveJanusGraphKafka
0 likes · 16 min read
Building Data Lineage at Ctrip: Architecture, Implementation, and Real‑World Applications
IT Architects Alliance
IT Architects Alliance
Sep 1, 2021 · Big Data

Understanding Data Middle Platform Architecture and Its Core Components

The article explains the concept of a data middle platform, describing its architecture, the essential big‑data foundation, metadata management, data service components such as BI and tag systems, and how these layers together enable unified data access, governance, and business intelligence across enterprises.

Business IntelligenceTag Managementdata architecture
0 likes · 14 min read
Understanding Data Middle Platform Architecture and Its Core Components
Efficient Ops
Efficient Ops
Aug 31, 2021 · Cloud Computing

Why Object Storage Is the New Backbone of Cloud Data Management

This article explains how object storage emerged as a cloud-native solution that surpasses traditional DAS, SAN, and NAS architectures by offering virtually unlimited capacity, robust metadata handling, and simple RESTful APIs for modern applications and large‑scale data workloads.

Object Storagecloud storagedata architecture
0 likes · 11 min read
Why Object Storage Is the New Backbone of Cloud Data Management
Alibaba Cloud Developer
Alibaba Cloud Developer
Aug 23, 2021 · Databases

How MySQL 8.0’s Data Dictionary Eliminates Metadata Redundancy and Boosts Performance

MySQL 8.0 replaces duplicated server‑level and engine‑level metadata with a unified data dictionary stored in InnoDB, introduces a two‑level cache (local and shared) built on templated hash maps, and provides atomic DDL operations, dramatically improving metadata consistency, performance, and management simplicity.

DDLMySQLcache architecture
0 likes · 21 min read
How MySQL 8.0’s Data Dictionary Eliminates Metadata Redundancy and Boosts Performance
IT Architects Alliance
IT Architects Alliance
Aug 9, 2021 · Big Data

Data Warehouse Architecture Overview: Layers, Sources, Modeling, Storage, and Management

This article explains the logical layered architecture of modern data warehouses, covering data sources, ODS, DW/DWS layers, collection, storage on HDFS, synchronization tools, dimensional modeling (star, snowflake, constellation), metadata management, and task scheduling and monitoring, highlighting best practices for scalable big‑data solutions.

ETLdata warehousemetadata
0 likes · 12 min read
Data Warehouse Architecture Overview: Layers, Sources, Modeling, Storage, and Management
Top Architect
Top Architect
Aug 3, 2021 · Fundamentals

Design and Considerations of Distributed File Systems

This article provides a comprehensive overview of distributed file systems, covering their historical evolution, essential requirements such as POSIX compliance, persistence, scalability, and security, and comparing centralized (e.g., GFS) and decentralized (e.g., Ceph) architectures along with strategies for high availability, performance optimization, and data consistency.

consistencydistributed file systemmetadata
0 likes · 19 min read
Design and Considerations of Distributed File Systems
Tencent Cloud Middleware
Tencent Cloud Middleware
Jun 28, 2021 · Big Data

Getting Started with Kafka’s New KRaft Mode: A Step‑by‑Step Guide

This article introduces Apache Kafka’s KRaft (Kafka Raft) mode, explains its architectural differences from ZooKeeper‑based deployments, details essential configuration parameters, and provides a complete step‑by‑step procedure—including commands and utility tools—to set up and operate a KRaft cluster.

ConfigurationKRaftKafka
0 likes · 14 min read
Getting Started with Kafka’s New KRaft Mode: A Step‑by‑Step Guide
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jun 1, 2021 · Fundamentals

How Huawei Built a Comprehensive Data Governance Framework for Digital Transformation

Huawei’s 2017 digital‑transformation vision led to a five‑step data‑governance blueprint that evolved through two phases, defining a detailed data‑classification framework, structured and unstructured data management methods, metadata governance, and compliance‑driven external data handling to support enterprise‑wide intelligent operations.

data classificationdata governancemetadata
0 likes · 20 min read
How Huawei Built a Comprehensive Data Governance Framework for Digital Transformation
58 Tech
58 Tech
May 19, 2021 · Mobile Development

Hooking Swift Functions by Modifying the Virtual Table (VTable)

This article explains a novel Swift hooking technique that modifies the virtual function table (VTable) to replace method implementations, detailing Swift's runtime structures such as TypeContext, Metadata, OverrideTable, and providing concrete ARM64 assembly and Swift code examples.

HookingSwiftTypeContext
0 likes · 19 min read
Hooking Swift Functions by Modifying the Virtual Table (VTable)
Big Data Technology & Architecture
Big Data Technology & Architecture
May 19, 2021 · Big Data

Comprehensive Guide to Data Governance: Metadata, Data Quality, Standards, and Asset Management

This article provides an extensive overview of data governance in the big‑data era, covering common pitfalls, the role of metadata, data quality management, data standardization, and data asset management, and offers practical recommendations for organizations to implement effective governance practices.

Data Qualitybig datadata asset management
0 likes · 42 min read
Comprehensive Guide to Data Governance: Metadata, Data Quality, Standards, and Asset Management
Meituan Technology Team
Meituan Technology Team
May 6, 2021 · Backend Development

GraphQL‑Based BFF Architecture with Metadata‑Driven Data Aggregation

The article describes a backend‑for‑frontend architecture that pushes GraphQL into the BFF layer, separates fetch and display units, drives execution with metadata, unifies query models, applies caching and parallel processing optimizations, and demonstrates over 50 % logic reuse and doubled development efficiency in production.

BFFGraphQLPerformance Optimization
0 likes · 35 min read
GraphQL‑Based BFF Architecture with Metadata‑Driven Data Aggregation
DataFunTalk
DataFunTalk
Apr 30, 2021 · Cloud Native

JuiceFS: A Cloud‑Native Distributed File System for Big Data and AI Workloads

This article presents JuiceFS, an open‑source cloud‑native distributed file system that addresses the limitations of object storage for big‑data and AI workloads by providing strong consistency, high‑performance metadata, multi‑protocol support, small‑file management, and deep Kubernetes integration.

Artificial Intelligencecloud-nativedistributed file system
0 likes · 13 min read
JuiceFS: A Cloud‑Native Distributed File System for Big Data and AI Workloads
Big Data Technology Architecture
Big Data Technology Architecture
Apr 19, 2021 · Big Data

Reframing Apache Hudi as a Data Lake Platform: Vision, Capabilities, and Future Directions

Apache Hudi is being re‑positioned from a simple table format to a full‑featured data lake platform, offering transactional storage, MVCC concurrency, metadata services, Deltastreamer ingestion, and plans for cache and timeline metadata services, aligning its vision with modern lakehouse architectures.

Apache HudiTransactional Storagemetadata
0 likes · 5 min read
Reframing Apache Hudi as a Data Lake Platform: Vision, Capabilities, and Future Directions
DataFunTalk
DataFunTalk
Apr 14, 2021 · Big Data

Beike's Data Development Platform: Evolution, Architecture, and Future Outlook

The talk by Beike senior engineer Yang Zongqiang details the evolution of the company's data development platform, covering background, three architecture upgrades, platform features such as metadata management, data integration, scheduling, quality assurance, and future directions for building an enterprise‑grade big‑data system.

Data Qualitydata platformmetadata
0 likes · 21 min read
Beike's Data Development Platform: Evolution, Architecture, and Future Outlook
Big Data Technology Architecture
Big Data Technology Architecture
Apr 5, 2021 · Big Data

Understanding Apache Iceberg: Table Format Architecture, Comparison with Hive Metastore, and Business Benefits

This article introduces Apache Iceberg as an open table format for massive analytic datasets, explains its underlying concepts such as schema, partitioning, statistics, and read/write APIs, compares it with Hive Metastore, outlines its ACID commit process, highlights the performance and operational advantages for big‑data workloads, and previews upcoming community features.

ACIDApache IcebergParquet
0 likes · 19 min read
Understanding Apache Iceberg: Table Format Architecture, Comparison with Hive Metastore, and Business Benefits
21CTO
21CTO
Feb 14, 2021 · Cloud Computing

How Metadata‑Driven Multi‑Tenant Architecture Powers Scalable SaaS Platforms

This article explains how a metadata‑driven multi‑tenant data model decouples logical and physical schemas, enabling rapid SaaS product rollout, seamless scaling, fine‑grained customization, and zero‑downtime schema changes while ensuring data isolation, security, and high performance across millions of tenants.

SaaSmetadatamulti-tenant
0 likes · 45 min read
How Metadata‑Driven Multi‑Tenant Architecture Powers Scalable SaaS Platforms
DataFunTalk
DataFunTalk
Feb 2, 2021 · Big Data

Metadata Management: Concepts, Architecture, and Applications in Data Warehousing

This article explains the fundamentals and value of metadata, describes a comprehensive metadata management system and its layered architecture, outlines key technologies such as automatic SQL metadata extraction, and showcases practical applications like metadata query, impact analysis, data lineage, and business‑driven data needs within modern data warehouses.

SQL parsingdata lineagedata warehouse
0 likes · 17 min read
Metadata Management: Concepts, Architecture, and Applications in Data Warehousing
DataFunTalk
DataFunTalk
Jan 29, 2021 · Artificial Intelligence

Content Embedding Practices and Challenges at Hulu

This article presents Hulu's multi‑layered approach to content understanding and embedding, describing tag‑based graph embeddings, metadata‑BERT enhancements, multimodal video/audio feature aggregation, and various applications such as similarity search, ranking, cold‑start retrieval, and collection modeling, while also discussing current limitations and open research questions.

Hulucontent embeddinggraph embeddings
0 likes · 12 min read
Content Embedding Practices and Challenges at Hulu
360 Smart Cloud
360 Smart Cloud
Jan 28, 2021 · Big Data

Overview of the Qirin Big Data Platform: Architecture, Modules, and Capabilities

The article provides a comprehensive overview of the Qirin big‑data platform, detailing its architecture, core modules such as resource management, metadata, data ingestion, task development, interactive query, and self‑service analysis, and outlines future development plans for the system.

Data Ingestiondata platforminteractive query
0 likes · 12 min read
Overview of the Qirin Big Data Platform: Architecture, Modules, and Capabilities
DataFunTalk
DataFunTalk
Jan 21, 2021 · Big Data

Kuaishou Metadata Platform: Evolution, Architecture, and Application Scenarios

This article introduces the development history, current architecture, abstraction methods, and key application scenarios of Kuaishou's metadata platform, highlighting challenges such as heterogeneous data integration, large-scale asset management, and the platform's role in data search, lineage, governance, and future enhancements.

Kuaishoudata lineagemetadata
0 likes · 16 min read
Kuaishou Metadata Platform: Evolution, Architecture, and Application Scenarios
360 Tech Engineering
360 Tech Engineering
Jan 7, 2021 · Big Data

Overview of the Qirin Big Data Platform Architecture and Core Modules

The article introduces the Qirin big data platform—a one‑stop solution covering resource management, metadata, data ingestion, task development, interactive querying, and self‑service analysis—detailing its modular architecture, typical processing workflow, and future development plans for enterprise‑wide data services.

Data Ingestionbig datadata platform
0 likes · 11 min read
Overview of the Qirin Big Data Platform Architecture and Core Modules
Youzan Coder
Youzan Coder
Dec 25, 2020 · Big Data

Metadata Governance and Collection in a Data Asset Platform

The platform implements comprehensive metadata governance by extracting, standardizing, and ingesting basic, trend, resource, lineage, and task metadata from offline and real‑time systems via a Kafka‑based SDK, enabling unified storage, monitoring, alerts, and future automation to improve data asset visibility and quality.

SDKbig datadata collection
0 likes · 18 min read
Metadata Governance and Collection in a Data Asset Platform
DataFunTalk
DataFunTalk
Dec 19, 2020 · Big Data

Evolution of iQIYI Data Warehouse from 1.0 to 2.0: Architecture, Modeling Practices, and Future Directions

This article details iQIYI's transition from a fragmented Data Warehouse 1.0 to a unified, standardized Data Warehouse 2.0, covering layered architecture, dimension and metric design, modeling workflows, metadata management, data lineage, and upcoming intelligent and automated data platform initiatives.

Data Modelingdata lineagedata warehouse
0 likes · 25 min read
Evolution of iQIYI Data Warehouse from 1.0 to 2.0: Architecture, Modeling Practices, and Future Directions
Sohu Tech Products
Sohu Tech Products
Dec 2, 2020 · Big Data

Optimizing Hive SQL Lineage Parsing: Techniques, Implementation, and Practical Insights

This article presents a comprehensive overview of Hive SQL lineage parsing, detailing the challenges of data provenance in large‑scale data warehouses, introducing ANTLR‑based parsing techniques, and describing a series of optimizations—including AST pruning, CTE handling, UDF registration, and metadata service integration—to improve both table‑level and column‑level lineage extraction and visualization.

ANTLRHiveSQL lineage
0 likes · 18 min read
Optimizing Hive SQL Lineage Parsing: Techniques, Implementation, and Practical Insights
MaGe Linux Operations
MaGe Linux Operations
Dec 1, 2020 · Fundamentals

Why Journaling Keeps File Systems Safe: Write-Ahead Logging Explained

File systems risk data corruption during power loss or crashes because writes are not atomic, so journaling—recording intended operations in a write‑ahead log before committing them—ensures metadata and user data consistency, with variations like data journaling and ordered (metadata) journaling improving performance and reliability.

Write-Ahead Loggingdata integrityjournaling
0 likes · 6 min read
Why Journaling Keeps File Systems Safe: Write-Ahead Logging Explained
Big Data Technology Architecture
Big Data Technology Architecture
Nov 21, 2020 · Big Data

Multi-Engine Support and Future Directions of Alibaba Cloud Data Lake Building Service

The article explains how Alibaba Cloud's Data Lake Building Service enables fine‑grained lake management by integrating multiple compute engines—including EMR, MaxCompute, Blink, Hologres, PAI, and open‑source Hive, Spark, and Presto—through unified metadata and OSS storage, while outlining current features, special format support, and planned future enhancements.

Alibaba CloudEMROSS
0 likes · 9 min read
Multi-Engine Support and Future Directions of Alibaba Cloud Data Lake Building Service
iQIYI Technical Product Team
iQIYI Technical Product Team
Nov 13, 2020 · Big Data

Evolution of iQIYI Data Warehouse from 1.0 to 2.0: Architecture, Modeling, Metadata, and Data Lineage

The talk chronicles iQIYI’s shift from a fragmented five‑layer Data Warehouse 1.0 to a unified 2.0 architecture featuring a central Dimension Layer, business‑focused data marts, and subject‑oriented warehouses, while detailing platform services, rigorous metadata management, lineage tracking, and future goals of intelligent, automated, service‑oriented, model‑driven data governance.

Data Modelingdata lineageiQIYI
0 likes · 23 min read
Evolution of iQIYI Data Warehouse from 1.0 to 2.0: Architecture, Modeling, Metadata, and Data Lineage
Efficient Ops
Efficient Ops
Nov 4, 2020 · Fundamentals

How Journal File Systems Prevent Data Corruption After Crashes

Journal file systems use write‑ahead logging to record each write operation as a transaction, ensuring that after power loss or crashes the system can replay logs and maintain metadata and user‑data consistency, avoiding corruption and space waste through techniques like data, ordered, and metadata journaling.

Data ConsistencyFile SystemWrite-Ahead Logging
0 likes · 8 min read
How Journal File Systems Prevent Data Corruption After Crashes
JavaEdge
JavaEdge
Sep 15, 2020 · Backend Development

How Kafka Uses ZooKeeper for Metadata Management and Client Coordination

This article explains how Kafka relies on ZooKeeper to store cluster metadata, detailing the ZK node hierarchy, the process by which clients locate brokers, the broker‑side handling of metadata requests, and recommended practices for large‑scale deployments.

Kafkabackend developmentmetadata
0 likes · 8 min read
How Kafka Uses ZooKeeper for Metadata Management and Client Coordination
Youku Technology
Youku Technology
Aug 17, 2020 · Backend Development

Improving Development Efficiency with a Metadata Center: Architecture, Implementation, and Performance

The Metadata Center, built by Alibaba Entertainment, streamlines development by offering a searchable Data Source Plaza and a configurable Custom Interface engine that abstracts service calls, adds unified monitoring and circuit‑breaker safeguards, and leverages optimized scripting, cutting upfront coding effort and accelerating feature delivery across Youku applications.

Circuit BreakerGroovyService Integration
0 likes · 12 min read
Improving Development Efficiency with a Metadata Center: Architecture, Implementation, and Performance
Efficient Ops
Efficient Ops
Aug 5, 2020 · Cloud Computing

Why Object Storage Is the Next Big Thing in Cloud Computing

This article explains the fundamentals of object storage, compares it with block and file storage, outlines its architecture, components, advantages, use cases, and limitations, showing why it has become the dominant storage model in modern cloud environments.

Object Storagecloud storagedata architecture
0 likes · 11 min read
Why Object Storage Is the Next Big Thing in Cloud Computing
Ctrip Technology
Ctrip Technology
Jul 30, 2020 · Backend Development

Design and Implementation of Tripyun Ctrip Cloud Metadata and SPU Management Platform

This article describes the background, metadata concepts, SPU and category‑attribute modeling, MySQL tree‑storage techniques, dynamic business configuration, rule‑engine and workflow‑engine solutions adopted in Tripyun‑Ctrip Cloud to build a reusable, scalable middle‑platform for the tourism industry.

BackendData ModelingRule Engine
0 likes · 21 min read
Design and Implementation of Tripyun Ctrip Cloud Metadata and SPU Management Platform
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 12, 2020 · Big Data

Common Metadata Management Patterns in Storage Systems

This article explains why metadata management is crucial for storage systems and reviews four typical approaches—initial external‑DB storage, in‑memory loading, partitioned services with a proxy layer, and tiered caching/persistence—illustrated with diagrams and real‑world examples.

Storage Systemsmetadatatiered architecture
0 likes · 5 min read
Common Metadata Management Patterns in Storage Systems
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 11, 2020 · Big Data

Alluxio Tiered Metadata Management and Asynchronous Cache Eviction Implementation

The article explains Alluxio's tiered metadata management architecture, describing how the system separates hot and cold metadata into cached and persisted layers, and details the custom asynchronous eviction thread and cache implementation that replace Guava cache for efficient large‑scale metadata handling.

Alluxiocachedistributed storage
0 likes · 15 min read
Alluxio Tiered Metadata Management and Asynchronous Cache Eviction Implementation
Architects Research Society
Architects Research Society
Jun 15, 2020 · Databases

Overview of Data Modeling, Architecture, Master Data Management, Metadata, and Data Quality

This article explains the concepts of data modeling and architecture, including logical data, process, and rule modeling, various data model types, master data management principles, metadata categories, and data quality management practices, highlighting their roles in enterprise information systems.

Data ModelingData QualityMaster Data Management
0 likes · 9 min read
Overview of Data Modeling, Architecture, Master Data Management, Metadata, and Data Quality
Big Data Technology & Architecture
Big Data Technology & Architecture
May 24, 2020 · Big Data

Data Governance Core Areas and Practices for Banking

The article provides a comprehensive overview of banking data governance, covering core domains such as data models, metadata, standards, quality, lifecycle, distribution, exchange, security, and services, and explains how big‑data techniques can improve risk control, product innovation, and operational efficiency.

Data Qualitybankingdata security
0 likes · 16 min read
Data Governance Core Areas and Practices for Banking
Youzan Coder
Youzan Coder
Mar 18, 2020 · Big Data

The Evolution of Youzan’s Data Warehouse in a Big Data Environment

The article traces Youzan’s data warehouse from its chaotic early days lacking structure, through a 2016 Airflow‑driven construction phase that introduced layered ODS/DW/Data Mart architecture and naming standards, to a mature stage focused on efficiency, security, SparkSQL, dimensional modeling, metadata, and ongoing real‑time and governance challenges.

AirflowETLHive
0 likes · 20 min read
The Evolution of Youzan’s Data Warehouse in a Big Data Environment
Meituan Technology Team
Meituan Technology Team
Mar 12, 2020 · Big Data

Data Governance Practices in Meituan Delivery: Architecture, Standards, and Security

Meituan Delivery’s data‑governance framework combines a four‑layer warehouse architecture with comprehensive business, technical, security, and resource‑management standards, continuous metadata and security controls, and tools such as Wherehows and QuickSight, delivering standardized, secure, and easily shareable data while guiding future optimization and emerging‑technology adoption.

big datadata architecturedata governance
0 likes · 27 min read
Data Governance Practices in Meituan Delivery: Architecture, Standards, and Security
ITPUB
ITPUB
Jan 10, 2020 · Fundamentals

Understanding Inodes: How Unix/Linux Stores File Metadata

This article explains Unix/Linux inodes—the metadata structures that store file information—covering their purpose, contents, size considerations, inode numbers, directory handling, hard and soft links, and special inode-related operations, with practical command examples and visual illustrations.

File SystemHard LinkInode
0 likes · 10 min read
Understanding Inodes: How Unix/Linux Stores File Metadata
vivo Internet Technology
vivo Internet Technology
Dec 18, 2019 · Big Data

Comprehensive Overview of Big Data Architecture, Lambda/Kappa Models, and End-to-End Data Platform Design

The article surveys modern big‑data architecture, contrasting Lambda and Kappa models, highlights common governance and integration pain points, and proposes an end‑to‑end platform featuring unified metadata, stream‑batch processing, one‑click ingestion, standardized modeling, intelligent query abstraction, and a comprehensive development IDE.

Data ModelingETLKappa Architecture
0 likes · 13 min read
Comprehensive Overview of Big Data Architecture, Lambda/Kappa Models, and End-to-End Data Platform Design
Programmer DD
Programmer DD
Nov 7, 2019 · Backend Development

Master Spring Boot Configuration Processor to Generate Accurate Metadata

This tutorial explains how to use Spring Boot's Configuration Processor to generate JSON metadata for configuration properties, covering dependency setup, Java bean definitions, property files, tests, and how the resulting metadata improves IDE auto‑completion and documentation.

@ConfigurationPropertiesConfiguration ProcessorJava
0 likes · 10 min read
Master Spring Boot Configuration Processor to Generate Accurate Metadata
DevOps Cloud Academy
DevOps Cloud Academy
Aug 11, 2019 · Big Data

Overview of MFS Distributed File System Architecture Similar to GoogleFS

The article explains the MFS distributed file system, detailing its four components—Master, Metalogger, Chunkserver, and Client—along with hardware recommendations, metadata handling, replication strategies, and FUSE‑based client mounting, providing a comprehensive guide to building a GoogleFS‑like storage cluster.

MFSbig datachunkserver
0 likes · 5 min read
Overview of MFS Distributed File System Architecture Similar to GoogleFS
21CTO
21CTO
Jul 2, 2019 · Backend Development

Designing a Scalable Feed Stream System for Billions of Users

This article explains how to design a high‑performance feed‑stream architecture—including product definition, data modeling, storage choices, synchronization modes, metadata handling, commenting, likes, sorting, search, and deletion—so that a system can support tens of millions to billions of users while remaining reliable and scalable.

Storagefeed streammetadata
0 likes · 21 min read
Designing a Scalable Feed Stream System for Billions of Users