Tagged articles

data engineering

321 articles · Page 1 of 4
Ray's Galactic Tech
Ray's Galactic Tech
Sep 29, 2026 · Backend Development

Production-Grade LLM Annotation Pipeline: Pydantic, PostgreSQL & Kafka Patterns

The article details how a SaaS team evolved a fragile LLM ticket-classification script into a production-grade pipeline using strict Pydantic contracts, PostgreSQL state machines with atomic leasing, Kafka outbox patterns, immutable training snapshots, and metric-gated rollouts to handle model drift, replay, and human review safely.

KafkaLLM data pipelinePostgreSQL
0 likes · 27 min read
Production-Grade LLM Annotation Pipeline: Pydantic, PostgreSQL & Kafka Patterns
Lakehouse Research Base
Lakehouse Research Base
Aug 31, 2026 · Big Data

Factory Lines & Data Pipelines: The Shared Orchestration DNA

This article reveals the deep structural isomorphism between manufacturing assembly lines and data pipeline orchestration, mapping five shared design philosophies, a 15-dimension logical correspondence, bidirectional lessons, and six practical mental models for data engineers to adopt factory-floor thinking.

assembly linedata engineeringdata orchestration
0 likes · 22 min read
Factory Lines & Data Pipelines: The Shared Orchestration DNA
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Aug 28, 2026 · Industry Insights

Why Palantir’s Ontology Grows Instead of Being Designed

The article explains how Palantir’s ontology is not pre‑designed but gradually emerges from a massive, long‑running data‑engineering foundation that continuously ingests structured business systems, accumulates common patterns, and turns raw data into a platform‑wide semantic layer that drives actions.

Enterprise Knowledge GraphFoundryOntology
0 likes · 10 min read
Why Palantir’s Ontology Grows Instead of Being Designed
IT Services Circle
IT Services Circle
Aug 27, 2026 · Big Data

10 Python Libraries That Actually Make Data Professionals Stronger

The article presents a curated list of ten Python libraries—Polars, Pandera, DuckDB, Rich, Pydantic, MLflow, tqdm, pyinstrument, RapidFuzz, and sqlite-utils—each chosen for its ability to eliminate specific bottlenecks, reduce errors, and turn guesswork into verifiable steps in data workflows, complete with concrete code examples and practical trade‑offs.

DuckDBMLflowPolars
0 likes · 15 min read
10 Python Libraries That Actually Make Data Professionals Stronger
Ray's Galactic Tech
Ray's Galactic Tech
Aug 26, 2026 · Big Data

Layered Real‑Time Data Warehouse with Flink CDC, Kafka & Doris

The article explains why real‑time data‑warehouse projects often fail in production and presents a complete, production‑ready solution that layers ODS, DWD, DWS and ADS using Flink CDC to capture MySQL changes, Kafka for buffering and replay, and Doris for OLAP storage, with detailed guidance on architecture, state handling, fault‑tolerance and operations.

CDCDorisFlink
0 likes · 35 min read
Layered Real‑Time Data Warehouse with Flink CDC, Kafka & Doris
ITPUB
ITPUB
Aug 22, 2026 · Databases

How DBAs Can Transform with AI: Insights from DTCC 2026

The 17th China Database Technology Conference showcased DBA‑AI transformation strategies, highlighting five‑stage evolution paths, lightweight big‑data + AI architectures, and the emerging Agentic Data Stack, while warning of AI hallucination risks and emphasizing human‑AI collaboration for future data engineering.

AIAgentic AIDBA
0 likes · 12 min read
How DBAs Can Transform with AI: Insights from DTCC 2026
Niu Liu
Niu Liu
Aug 21, 2026 · Big Data

Real-Time Data Warehouse Evolution: Flink Writes, Paimon Stores, Doris Queries, One Platform Manages

This article details a lightweight real-time data warehouse architecture using Apache Flink 2.0+ for streaming writes, Apache Paimon for lakehouse storage, Apache Doris for interactive queries, and a custom RT-DWH management platform to unify metadata, quality, permissions, and observability — reducing component count and operational complexity for small-to-medium teams.

Apache DorisApache FlinkApache Paimon
0 likes · 21 min read
Real-Time Data Warehouse Evolution: Flink Writes, Paimon Stores, Doris Queries, One Platform Manages
Architecture Digest
Architecture Digest
Aug 20, 2026 · Artificial Intelligence

7 Tough RAG Interview Questions ByteDance Asked – Why Most Candidates Fail the First Three

The article breaks down the seven RAG interview questions ByteDance uses, detailing data source classification, multi‑layer cleaning pipelines, PDF parsing strategies, knowledge extraction, incremental updates, conflict resolution, and version control, and explains what interviewers are really looking for.

Incremental UpdateKnowledge ExtractionRAG
0 likes · 26 min read
7 Tough RAG Interview Questions ByteDance Asked – Why Most Candidates Fail the First Three
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Aug 9, 2026 · Artificial Intelligence

Why Enterprise AI Needs All Three Legs: Data, Agent, and FDE

The article explains how a large enterprise succeeded in AI‑enabled sales by cleaning five years of data, deploying a dedicated AI agent for each of eleven sales stages, and using Front‑end Deployment Engineers to translate expert knowledge into repeatable processes, showing that missing any of these three components makes the system limp.

AI DeploymentAgent ArchitectureEnterprise AI
0 likes · 9 min read
Why Enterprise AI Needs All Three Legs: Data, Agent, and FDE
Big Data Technology & Architecture
Big Data Technology & Architecture
Aug 5, 2026 · Big Data

12 Tough Data‑AI Interview Questions from Leading Companies – Can You Master the Latest Trends? (Part 1)

This article (part 1) presents twelve high‑level interview questions from top tech firms covering the shift from AI‑Ready to Data‑Agent‑Ready data warehouses, AI‑centric metadata and lineage, lakehouse formats like Paimon, and key Iceberg features that empower AI‑driven analytics.

AILakehousedata engineering
0 likes · 8 min read
12 Tough Data‑AI Interview Questions from Leading Companies – Can You Master the Latest Trends? (Part 1)
Data Bricklaying Diary
Data Bricklaying Diary
Aug 1, 2026 · Big Data

AI Data Engineering: The Data Supply System for the Agent Era

This article defines AI Data Engineering as a data supply system for large models and agents, extending traditional data engineering with semantic modeling, RAG, controlled data services, permission governance, and feedback loops to make data understandable, retrievable, callable, and auditable for reliable enterprise AI deployment.

AI data engineeringAgentsData Services
0 likes · 11 min read
AI Data Engineering: The Data Supply System for the Agent Era
DataFunSummit
DataFunSummit
Jul 21, 2026 · Industry Insights

When Models Get Cheaper, Who’s Making Money?

The article argues that as large‑language‑model costs plunge, profit shifts from model providers to companies that prepare, govern, and route data for AI, citing Databricks’ $3 billion raise, Anthropic’s data‑engineered accuracy jump, and the emerging European compliance market.

AIAI GovernanceBusiness Models
0 likes · 12 min read
When Models Get Cheaper, Who’s Making Money?
DataFunTalk
DataFunTalk
Jul 18, 2026 · Industry Insights

Why Data Agent Stalls at 70% and Hits 95% Only With a Semantic Layer

The Data for AI Beijing meetup revealed that Data Agents plateau at about 70% accuracy without a well‑defined semantic (context) layer, but can reach the 95% production threshold once that layer is built, highlighting a shift from engine‑centric to metadata‑centric architectures, six‑round convergence practices, and large‑scale metadata deployments.

AI infrastructureData AgentGravitino
0 likes · 23 min read
Why Data Agent Stalls at 70% and Hits 95% Only With a Semantic Layer
Data Party THU
Data Party THU
Jul 12, 2026 · Big Data

Why Polars Beats Pandas for Massive Data Processing: A Deep Dive

Polars outperforms Pandas on large‑scale ETL by using a multithreaded lazy execution model, columnar Arrow storage, and query optimizations, delivering up to 94× speedups on 10 GB workloads, while Pandas remains suitable for smaller datasets and tight ML ecosystem integration.

ETLLazy ExecutionPolars
0 likes · 15 min read
Why Polars Beats Pandas for Massive Data Processing: A Deep Dive
Machine Heart
Machine Heart
Jul 8, 2026 · Artificial Intelligence

How LingBot‑VLA 2.0 Powers 20 Robot Configurations with an Open‑Source Embodied Brain

LingBot‑VLA 2.0 introduces a token‑level loss‑free MoE, dual‑query distillation, and a 60k‑hour heterogeneous dataset to achieve cross‑embodiment visual‑language‑action capabilities across 20 robot morphologies, delivering superior benchmark performance and sub‑130 ms inference while being fully open‑sourced.

Embodied AIFuture PredictionMixture of Experts
0 likes · 16 min read
How LingBot‑VLA 2.0 Powers 20 Robot Configurations with an Open‑Source Embodied Brain
Data Integration and Governance
Data Integration and Governance
Jun 30, 2026 · Fundamentals

9 Essential Data Cleaning Techniques Every Analyst Must Master

Before visualizing or modeling, analysts must first resolve common data quality issues—duplicate records, inconsistent formats, missing or abnormal values, and mismatched definitions—by applying nine systematic cleaning steps that turn raw, chaotic data into reliable, comparable, and reusable information.

Data Validationdata cleaningdata engineering
0 likes · 20 min read
9 Essential Data Cleaning Techniques Every Analyst Must Master
dbaplus Community
dbaplus Community
Jun 28, 2026 · Operations

Why Tencent Music Rejects AI Hype: Building an OpenClaw‑Powered Intelligent Ops Ecosystem

The article details Tencent Music's step‑by‑step evolution from manual alert handling to a three‑layer cloud‑native AIOps platform, describing data pipelines, dynamic 3‑sigma alerts, full‑link observability, and the OpenClaw sandbox with multi‑agent architecture that prioritises scenario‑driven, safe AI integration.

AIAIOpsOpenClaw
0 likes · 17 min read
Why Tencent Music Rejects AI Hype: Building an OpenClaw‑Powered Intelligent Ops Ecosystem
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 25, 2026 · Big Data

Taobao Live’s Shift from ETL to Managed Development with DataWorks Data Agent

The article details how Taobao Live’s data engineering team replaced traditional ETL bottlenecks with a three‑layer, AI‑native architecture built on DataWorks Data Agent, using NL2DSL2SQL, ontology‑driven knowledge bases, and multi‑agent collaboration to achieve near‑100% code generation and higher accuracy.

AI-nativeData AgentDataWorks
0 likes · 10 min read
Taobao Live’s Shift from ETL to Managed Development with DataWorks Data Agent
DataFunSummit
DataFunSummit
Jun 14, 2026 · Artificial Intelligence

How cz-cli Empowers Data Engineers by Giving AI Real Understanding of Data Warehouses

The article analyzes how data engineers lose focus to repetitive tasks, describes the design journey from generic LLM usage to the specialized cz-cli agent, details its 37 skills and typical scenarios such as lineage analysis and incremental pipelines, and shows how the tool returns attention control to engineers while also enabling business users to self‑serve data.

AI AgentsLLMautomation
0 likes · 13 min read
How cz-cli Empowers Data Engineers by Giving AI Real Understanding of Data Warehouses
Frontend AI Walk
Frontend AI Walk
Jun 14, 2026 · Industry Insights

Redefining the Career Track: The Forward Deployed Engineer Blueprint

The article defines the Forward Deployed Engineer (FDE) role as a bridge between software engineers and customers, outlines its core duties, compares it with Sales Engineer and Solutions Architect, presents market data, a detailed skill framework, a four‑pillar self‑assessment, and a step‑by‑step transition roadmap for aspiring engineers.

Forward Deployed Engineercareer transitioncloud platforms
0 likes · 19 min read
Redefining the Career Track: The Forward Deployed Engineer Blueprint
Fei's Miscellaneous Talks
Fei's Miscellaneous Talks
Jun 12, 2026 · Artificial Intelligence

Online Decision-Making vs. Offline Learning in Recommendation System Architecture

This article outlines the industrial‑grade offline architecture of a recommendation system, detailing how online services handle real‑time decisions while a comprehensive offline pipeline processes data, builds user profiles, extracts features, and trains models to continuously improve personalization.

Feature PipelineOffline Architecturedata engineering
0 likes · 21 min read
Online Decision-Making vs. Offline Learning in Recommendation System Architecture
Machine Heart
Machine Heart
Jun 2, 2026 · Artificial Intelligence

When AI Becomes Its Own Data Engineer: Inside DataMaster

DataMaster introduces an autonomous AI data engineer that automatically searches, cleans, combines, and reuses data, enabling fixed models and training pipelines to achieve substantial performance gains across benchmarks such as MLE‑Bench Lite and PostTrainBench, including a 31.0% GPQA score.

AI researchData-Centric AIDataMaster
0 likes · 11 min read
When AI Becomes Its Own Data Engineer: Inside DataMaster
BanTech Think Tank
BanTech Think Tank
May 29, 2026 · Artificial Intelligence

From Chaotic Transaction Data to Intelligent Insight: AI Raises Classification Accuracy by 20%

The article recounts how the Postal Savings Bank of China’s software R&D team built a domain‑specific corpus, applied a three‑stage BERT training pipeline and knowledge‑distillation to lift transaction‑type classification accuracy above 95%, delivering finer customer profiling, scenario‑based marketing calendars and a roadmap for future data‑driven banking.

AIBERTNLP
0 likes · 9 min read
From Chaotic Transaction Data to Intelligent Insight: AI Raises Classification Accuracy by 20%
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
May 13, 2026 · Artificial Intelligence

How to Explain a Jump from 71% to 94% Tool‑Calling Accuracy in a JD Interview

The article walks through a JD interview scenario where a candidate explains how a tool‑calling accuracy metric rose from 71% to 94% by detailing the full SFT data‑engineering pipeline, teacher‑model trajectory generation, quality validation, evaluation methodology, and interview‑ready talking points.

Function CallingInterview PreparationLLM fine-tuning
0 likes · 19 min read
How to Explain a Jump from 71% to 94% Tool‑Calling Accuracy in a JD Interview
DataFunTalk
DataFunTalk
May 2, 2026 · Big Data

Building a One-Person Data Team: Core Skills of a Full‑Stack Data Engineer

The article examines why a single data engineer can run an end‑to‑end data team, outlines the essential abilities—semantic ownership, building an agentic data stack, and leveraging historical context—while discussing ChatBI’s limits, validation loops, and the open‑source Datus 0.3 harness for practical implementation.

Agentic AIChatBIDatus
0 likes · 14 min read
Building a One-Person Data Team: Core Skills of a Full‑Stack Data Engineer
Cloud Architecture
Cloud Architecture
Apr 14, 2026 · Big Data

Spark SQL Deep Dive: From API to Core for Real‑Time Billion‑Row Processing

This article provides a comprehensive, step‑by‑step guide to mastering Spark SQL in high‑concurrency, billion‑row scenarios, covering the execution chain, architecture layers, Delta Lake integration, performance tuning, production‑grade streaming and batch pipelines, Kubernetes deployment, parameter management, observability, and real‑world case studies.

Delta LakeKubernetesSpark SQL
0 likes · 44 min read
Spark SQL Deep Dive: From API to Core for Real‑Time Billion‑Row Processing
Smart Sea Tide
Smart Sea Tide
Apr 10, 2026 · Big Data

Datus-agent: AI‑Native Data Engineering Agent for Context‑Aware Analytics

Facing fragmented workflows and tool overload in modern data engineering, Datus-agent introduces an AI‑native, open‑source agent that builds evolving contextual maps, encapsulates domain‑specific sub‑agents, and automates the full data‑pipeline lifecycle—from natural‑language queries to SQL generation, quality checks, and task scheduling—enabling analysts and engineers to collaborate efficiently.

AIContextual AIData Ops
0 likes · 10 min read
Datus-agent: AI‑Native Data Engineering Agent for Context‑Aware Analytics
Big Data Tech Team
Big Data Tech Team
Apr 9, 2026 · Industry Insights

Why Data Engineers Are the New AI Powerhouses: 4 Core Reasons & Actionable Tips

The article analyzes why data development engineers are becoming more valuable in the AI era, outlining four core reasons—including data‑driven AI limits, the rise of RAG architectures, heightened data compliance, and a talent shortage—while offering concrete advice on mastering real‑time pipelines, unstructured data, and AI infrastructure.

AI infrastructureRAGbig data
0 likes · 8 min read
Why Data Engineers Are the New AI Powerhouses: 4 Core Reasons & Actionable Tips
Data Integration and Governance
Data Integration and Governance
Apr 8, 2026 · Big Data

Master the Four Data Integration Patterns in One Guide

This article explains the four common data integration patterns—ETL, ELT, API‑based, and message‑queue approaches—detailing their core workflows, suitable scenarios, advantages, and trade‑offs so readers can choose the method that best fits their business and technical constraints.

APIELTETL
0 likes · 10 min read
Master the Four Data Integration Patterns in One Guide
Big Data Technology & Architecture
Big Data Technology & Architecture
Apr 3, 2026 · Industry Insights

Why Daft, Ray, and Lance Are Redefining Multimodal Data Pipelines

This article analyzes how the Daft‑Ray‑Lance stack tackles the challenges of multimodal AI workloads by offering a high‑performance Rust engine, adaptive back‑pressure, seamless Ray‑based distributed scheduling, and a storage format optimized for random access, vector indexing, and zero‑copy schema evolution, complete with benchmark comparisons and practical deployment guidance.

BenchmarkDaftLance
0 likes · 21 min read
Why Daft, Ray, and Lance Are Redefining Multimodal Data Pipelines
Big Data Tech Team
Big Data Tech Team
Apr 1, 2026 · Big Data

Why Your 2026 Big Data Resume Is Being Ignored and How to Fix It

In the 2026 spring hiring season, many big‑data job seekers see their resumes disappear because they still focus on offline batch processing, while employers now demand real‑time streaming, AI‑driven data pipelines, and cloud‑native deployment skills such as Flink, vector databases, and Kubernetes.

AI IntegrationFlinkbig data
0 likes · 7 min read
Why Your 2026 Big Data Resume Is Being Ignored and How to Fix It
dbaplus Community
dbaplus Community
Mar 22, 2026 · Industry Insights

Will Data Engineers Vanish by 2030? A Bold Forecast for the Future of Data Stacks

The article predicts that by 2030 the traditional data‑engineer role and modern data‑stack components will collapse into a few unified, HTAP‑capable databases, semantic layers, and AI agents, reshaping pipelines, warehouses, and even edge computing while urging engineers to pivot toward semantic modeling and AI orchestration.

AIHTAPSemantic Layer
0 likes · 19 min read
Will Data Engineers Vanish by 2030? A Bold Forecast for the Future of Data Stacks
Alibaba Cloud Developer
Alibaba Cloud Developer
Mar 6, 2026 · Big Data

How DataWorks Turns Data Quality Rules into Code with Data Contracts

This article explains how DataWorks integrates data quality specifications directly into the SQL development workflow using Data Contracts, addressing governance lag, versioning gaps, and trust issues while providing a unified, version‑controlled, and automated quality assurance process for offline data pipelines.

Data QualityDataWorksSQL
0 likes · 12 min read
How DataWorks Turns Data Quality Rules into Code with Data Contracts
SuanNi
SuanNi
Feb 23, 2026 · Artificial Intelligence

How FireRed-Image-Edit Sets New Standards for AI-Powered Image Editing

FireRed-Image-Edit, an open‑source instruction‑driven diffusion model, combines massive high‑quality data, a dual‑stream multimodal architecture, progressive training, and a comprehensive multi‑dimensional benchmark to achieve unprecedented pixel‑level control and human‑like editing performance across diverse visual tasks.

AITraining Strategiesdata engineering
0 likes · 12 min read
How FireRed-Image-Edit Sets New Standards for AI-Powered Image Editing
BanTech Think Tank
BanTech Think Tank
Feb 11, 2026 · Industry Insights

Building High-Quality Data Sets for Commercial Bank AI: Exploration and Practice

The article examines how commercial banks, driven by national policies and AI advances, develop high‑quality data sets—detailing their strategic importance, regulatory background, challenges in systematization, engineering, and operation, and showcasing ICBC’s layered data architecture, knowledge‑engine pipelines, and future data‑driven competitiveness.

AI modelsArtificial IntelligenceKnowledge Engineering
0 likes · 13 min read
Building High-Quality Data Sets for Commercial Bank AI: Exploration and Practice
DataFunSummit
DataFunSummit
Feb 1, 2026 · Artificial Intelligence

How AI Agents Are Redefining Data Engineering: Expert Insights and Real‑World Practices

In a deep‑dive roundtable, three data‑engineering veterans discuss the rise of AI agents, the importance of data context, memory mechanisms, workflow versus agent trade‑offs, and the future of database intelligence, offering practical strategies and architectural philosophies for building smarter data pipelines.

Context EngineeringDatabase IntelligenceImmersive Analytics
0 likes · 24 min read
How AI Agents Are Redefining Data Engineering: Expert Insights and Real‑World Practices
Fun with Large Models
Fun with Large Models
Jan 12, 2026 · Artificial Intelligence

Why You Should Master Large‑Model Training: A Full‑Process Practical Guide

The article explains why mastering large‑model training is crucial for professionals, researchers, and enterprises, outlines the end‑to‑end pipeline—from data preparation and pre‑training to instruction fine‑tuning and RLHF alignment—compares training with RAG, and presents a structured learning roadmap.

AI AgentsPyTorchRAG
0 likes · 14 min read
Why You Should Master Large‑Model Training: A Full‑Process Practical Guide
Big Data Tech Team
Big Data Tech Team
Dec 29, 2025 · Big Data

Master Big Data Development: A Complete Roadmap from Beginner to Expert

This guide presents a comprehensive big‑data development roadmap, detailing industry opportunities, a six‑module technology stack, four progressive learning stages, hands‑on project ideas, interview question strategies, common pitfalls, and curated resources, helping aspiring engineers become proficient and interview‑ready while avoiding common mistakes.

Interview PreparationLearning Pathbig data
0 likes · 11 min read
Master Big Data Development: A Complete Roadmap from Beginner to Expert
Big Data Tech Team
Big Data Tech Team
Dec 26, 2025 · Interview Experience

How to Nail a 2‑Minute Data Engineer Self‑Introduction

This guide outlines a concise, 1.5‑2‑minute self‑introduction for data engineering interviews, highlighting essential personal details, technical stack, project achievements, business impact, and common pitfalls to avoid, with a concrete example and actionable tips.

big datacareer advicedata engineering
0 likes · 5 min read
How to Nail a 2‑Minute Data Engineer Self‑Introduction
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 16, 2025 · Artificial Intelligence

How We Built an AI‑Powered Data Agent to Automate Data Retrieval at Scale

This article details the design and implementation of Matra, an AI‑driven data assistant for a large e‑commerce platform, covering the challenges of legacy data assets, knowledge‑base construction, GraphRAG integration, multi‑stage agent frameworks, practical results, and future plans for continuous improvement.

AIData RetrievalLLM
0 likes · 22 min read
How We Built an AI‑Powered Data Agent to Automate Data Retrieval at Scale
StarRocks
StarRocks
Dec 11, 2025 · Databases

How StarRocks Redesigns Bulk Import to Cut Small Files and Boost Throughput

This article explains how StarRocks mitigates the hidden risks of massive one‑time data imports in a storage‑compute separated architecture by redesigning the write path to spill to local disk, merge centrally, and write to object storage, resulting in fewer small files, higher write throughput, and more stable query performance.

Bulk ImportS3StarRocks
0 likes · 12 min read
How StarRocks Redesigns Bulk Import to Cut Small Files and Boost Throughput
Data Integration and Governance
Data Integration and Governance
Nov 28, 2025 · Big Data

Why Data Quality Fails and How to Build a High‑Quality Dataset in 4 Steps

The article explains common data‑quality pitfalls—such as inconsistent source entry, multi‑source integration issues, changing business rules, and ETL errors—then defines six concrete quality dimensions and presents a repeatable four‑step workflow, plus practical tool recommendations, for creating reliable datasets.

Data QualityData ValidationETL
0 likes · 11 min read
Why Data Quality Fails and How to Build a High‑Quality Dataset in 4 Steps
DataFunSummit
DataFunSummit
Nov 27, 2025 · Big Data

How BMW Turned Data Into Growth: A Sensors Data Case Study

This article details BMW's digital transformation journey using Sensors Data, covering the background of rapid app growth, the cross‑regional data collection challenges, the systematic solution architecture—including mapping, preprocessing, and historical data migration—and the resulting business impact and future AI‑driven roadmap.

Analyticsbig datadata engineering
0 likes · 13 min read
How BMW Turned Data Into Growth: A Sensors Data Case Study
Data STUDIO
Data STUDIO
Nov 25, 2025 · Big Data

Why Parquet Is the Faster, Lighter, Safer Alternative to CSV in Python

The article explains why CSV becomes a bottleneck for large‑scale data, demonstrates how Parquet’s columnar, typed, and compressed format dramatically reduces storage, speeds up reads, and improves data safety, and provides step‑by‑step Python code for migrating and benchmarking the switch.

CSVDuckDBParquet
0 likes · 18 min read
Why Parquet Is the Faster, Lighter, Safer Alternative to CSV in Python
PMTalk Product Manager Community
PMTalk Product Manager Community
Nov 23, 2025 · Artificial Intelligence

Essential Strategies for Building Successful AI Products

This guide outlines a step‑by‑step framework for creating AI products, covering problem discovery, user‑centric motivation analysis, compliance and ethics, defining a Minimum Viable Intelligent Product, assembling multidisciplinary teams, leveraging data and model selection, designing trustworthy UX, go‑to‑market tactics, moat building, and continuous monitoring for improvement.

AIEthicsMVP
0 likes · 17 min read
Essential Strategies for Building Successful AI Products
Ctrip Technology
Ctrip Technology
Nov 20, 2025 · Big Data

How Ctrip Achieved Minute‑Level Real‑Time Analytics with Flink CDC & Apache Paimon

Ctrip transformed its traditional T+1 offline warehouse into a near‑real‑time lakehouse by integrating Flink CDC with Apache Paimon, designing a two‑stage CDC ingestion, optimizing performance, implementing dynamic updates, and deploying the solution across multiple business scenarios, achieving minute‑level latency, reduced costs, and faster data‑driven decisions.

CDCFlinkPaimon
0 likes · 27 min read
How Ctrip Achieved Minute‑Level Real‑Time Analytics with Flink CDC & Apache Paimon
Past Memory Big Data
Past Memory Big Data
Nov 12, 2025 · Big Data

How Uber Upgraded Over 2 Million Spark Jobs from 2.4 to 3.3

Uber migrated more than two million daily Spark applications from version 2.4 to 3.3, detailing the motivations, architecture, four-step migration process, custom tools like Polyglot Piranha and Iron Dome, and the resulting performance, cost, and productivity gains.

Apache SparkIron DomeKubernetes
0 likes · 11 min read
How Uber Upgraded Over 2 Million Spark Jobs from 2.4 to 3.3
JD Cloud Developers
JD Cloud Developers
Nov 10, 2025 · Artificial Intelligence

How an AI‑Powered Experiment Analysis Agent Transforms Data Insights

This document outlines the background, design, architecture, workflow, and large‑model integration of an AI‑driven Experiment Analysis Agent, detailing how it consolidates data, automates analysis via modular pipelines, leverages DeepSeek models, and enhances user experience through unified front‑end forms and intelligent messaging.

Workflow Automationdata engineering
0 likes · 15 min read
How an AI‑Powered Experiment Analysis Agent Transforms Data Insights
Alimama Tech
Alimama Tech
Oct 15, 2025 · Artificial Intelligence

How Alibaba’s Taobao Starry Model Delivers Precise, Consistent E‑commerce Image Edits

Alibaba’s Taobao Starry Image Editing model tackles the e‑commerce challenge of maintaining visual consistency by introducing a high‑fidelity, plug‑in architecture, a million‑scale consistency dataset, and multi‑stage multilingual training, enabling precise, controllable edits without altering product layout or background.

Plug‑in Architectureconsistencydata engineering
0 likes · 10 min read
How Alibaba’s Taobao Starry Model Delivers Precise, Consistent E‑commerce Image Edits
AI2ML AI to Machine Learning
AI2ML AI to Machine Learning
Sep 28, 2025 · Artificial Intelligence

Core Metrics for Enterprise Large‑Model Engineering

The article outlines the five essential engineering domains—application, model, compute, knowledge, and data—in the era of large models, and details concrete scale, efficiency, service, value, quality, and security metrics that enterprises should track to drive intelligent outcomes.

AI EngineeringKnowledge ManagementModel Performance
0 likes · 7 min read
Core Metrics for Enterprise Large‑Model Engineering
Huolala Tech
Huolala Tech
Sep 19, 2025 · Big Data

How We Migrated 40PB of Offline Big Data Across Clouds with Zero Downtime

Over a year after completing a five‑month, cross‑cloud migration of Huolala’s 40 PB offline big‑data platform—spanning storage, compute, services, and infrastructure—the team details the architecture, verification methods, high‑throughput migration tools, network isolation strategies, and lessons learned to guide similar large‑scale data migrations.

Data Validationautomationcloud migration
0 likes · 16 min read
How We Migrated 40PB of Offline Big Data Across Clouds with Zero Downtime
DataFunTalk
DataFunTalk
Sep 15, 2025 · Artificial Intelligence

How AI+Data Agents Are Transforming the Automotive Industry’s Digital Leap

In an interview, Di Xingxing of Autohome details their AI+Data framework—unified lake‑warehouse, intelligent engine, and agent services—that breaks data silos, blends traditional models with LLMs, leverages causal inference and RAG knowledge bases, and uses continuous feedback to build explainable, evolving data agents for accurate sales forecasting, competitive analysis, and end‑to‑end business automation in the automotive industry.

AIAutomotiveCausal Inference
0 likes · 10 min read
How AI+Data Agents Are Transforming the Automotive Industry’s Digital Leap
Data Party THU
Data Party THU
Sep 6, 2025 · Big Data

From Data Chaos to Predictive Insight: My Solo Journey in the 2025 Big Data Competition

An individual participant recounts their journey in the 2025 China University Computer Competition Big Data Challenge, detailing data cleaning, feature engineering, model building on 300‑stock historical prices, and insights gained from solo competition experience, highlighting challenges, lessons, and future directions in financial AI.

Competitionbig datadata engineering
0 likes · 4 min read
From Data Chaos to Predictive Insight: My Solo Journey in the 2025 Big Data Competition
DataFunSummit
DataFunSummit
Jul 20, 2025 · Big Data

Why Incremental Computing Is Replacing Lambda Architecture in Modern Big Data Platforms

This interview with Yunqi Technology CTO Guan Tao explains how the traditional Lambda architecture’s triple‑system complexity drives costs and operational pain, and why the company’s General Incremental Computing (GIC) approach offers a unified, cost‑effective Kappa‑style solution for real‑time, batch, and interactive analytics.

Kappa ArchitectureLambda Architecturedata engineering
0 likes · 13 min read
Why Incremental Computing Is Replacing Lambda Architecture in Modern Big Data Platforms
DataFunTalk
DataFunTalk
Jul 18, 2025 · Artificial Intelligence

How Alibaba Tackles Low-Resource Language Data for Multilingual LLMs

Alibaba International’s senior data science expert explains a systematic five‑strategy solution—data acquisition, augmentation, quality optimization, engineering pipeline, and evaluation loop—to overcome data scarcity, high annotation cost, and processing challenges for low‑resource languages in multilingual large language models.

AIdata engineeringlow-resource languages
0 likes · 13 min read
How Alibaba Tackles Low-Resource Language Data for Multilingual LLMs
JD Retail Technology
JD Retail Technology
Jun 18, 2025 · Artificial Intelligence

How JD’s Tech Teams Power 618: AI, Logistics, and Voice Innovations

The article explores how JD’s engineers across retail, logistics, and AI divisions use model distillation, data selection, intelligent routing, and advanced voice recognition to improve the 618 shopping festival experience, highlighting real‑world technical challenges, solutions, and the company’s talent development programs.

AIdata engineeringlogistics
0 likes · 16 min read
How JD’s Tech Teams Power 618: AI, Logistics, and Voice Innovations
Lakehouse Research Base
Lakehouse Research Base
May 20, 2025 · Big Data

Why DataOps Jobs Will Explode: The Chemistry of Analytics, Lean, Agile & DevOps

This article analyzes DataOps as a deep fusion of data analytics, lean thinking, agile practices, and DevOps culture, contrasting it with DevOps, detailing its four genetic components with case studies, explaining its emergence from technical, business, and organizational drivers, examining current hiring fragmentation and talent gaps, and predicting explosive job growth driven by policy, tooling maturity, and talent supply reforms.

AgileData PipelineData Quality
0 likes · 20 min read
Why DataOps Jobs Will Explode: The Chemistry of Analytics, Lean, Agile & DevOps
Full-Stack Internet Architecture
Full-Stack Internet Architecture
May 20, 2025 · Big Data

Why Learn Kafka? Core Benefits, Use Cases, and a Summary

This article explains why Kafka is widely adopted by top companies, outlines its high throughput, scalability, and durability, and describes key real‑time data pipeline, stream processing, and big‑data integration scenarios, concluding that mastering Kafka is essential for modern backend and data engineering roles.

Kafkadata engineeringreal-time processing
0 likes · 4 min read
Why Learn Kafka? Core Benefits, Use Cases, and a Summary
Alibaba Cloud Native
Alibaba Cloud Native
May 18, 2025 · Cloud Native

Airflow vs Argo Workflows: Which Cloud‑Native Scheduler Wins for Data Engineering?

This comprehensive guide compares Apache Airflow and Argo Workflows—two leading cloud‑native distributed task schedulers—by examining their core features, architectures, DAG handling, performance, language support, big‑data and AI integrations, and provides practical selection advice for data engineers and DevOps teams.

AirflowArgo Workflowsdata engineering
0 likes · 23 min read
Airflow vs Argo Workflows: Which Cloud‑Native Scheduler Wins for Data Engineering?
Fighter's World
Fighter's World
May 17, 2025 · Industry Insights

Hidden Roadblocks That Sabotage B2B Large Model Products

The article dissects why many B2B GenAI projects fail to scale despite heavy investment, highlighting overlooked challenges in data preparation, model specialization, product integration, user experience, and organizational culture, and proposes concrete ways to bridge these gaps.

B2BGenAIdata engineering
0 likes · 21 min read
Hidden Roadblocks That Sabotage B2B Large Model Products
DevOps Engineer
DevOps Engineer
Apr 25, 2025 · Big Data

Reflections on PyCon LT 2025 Data Day: Sessions on Static Code Analysis, Data Warehouses, Pipelines, and Data Science Tools

The author recounts attending PyCon LT 2025 Data Day, summarizing talks on building a simple static code analyzer with AST, challenges of data warehouses versus data lakes, cloud cost‑scraping pipelines, A/B testing libraries, privacy‑enhancing data processing, and tools like Panel and Dagster, while noting the inspiring presence of female speakers.

DagsterData SciencePanel
0 likes · 7 min read
Reflections on PyCon LT 2025 Data Day: Sessions on Static Code Analysis, Data Warehouses, Pipelines, and Data Science Tools
Big Data Tech Team
Big Data Tech Team
Apr 20, 2025 · Industry Insights

Essential Skills & Tech Stacks for Every Data Team Role

This guide breaks down the main positions in a data team— from data development and analysis engineers to product managers and operations specialists—detailing each role’s key responsibilities, essential skill sets, and the typical technology stack they rely on.

Data Analyticsbig datadata engineering
0 likes · 7 min read
Essential Skills & Tech Stacks for Every Data Team Role
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Apr 15, 2025 · Big Data

Boosting Game Data Engineering with Alibaba Cloud EMR Serverless Spark

Yingjiao Network transformed its game data platform by adopting Alibaba Cloud EMR Serverless Spark, addressing previous architecture pain points, enhancing data collection, offline scheduling, and online analytics, which led to higher development speed, 50% faster compute, and improved stability for global game operations.

Cloud Computingdata engineeringgaming analytics
0 likes · 9 min read
Boosting Game Data Engineering with Alibaba Cloud EMR Serverless Spark
Kuaishou Tech
Kuaishou Tech
Apr 2, 2025 · Big Data

Apache Hudi Asia Summit Successfully Held

The first Apache Hudi Asia Summit in Beijing attracted over 230 attendees, featuring technical discussions on data lake optimization and case studies from companies like Fastly and Meituan.

Apache HudiData LakeTechnical Conference
0 likes · 12 min read
Apache Hudi Asia Summit Successfully Held
Baidu Geek Talk
Baidu Geek Talk
Mar 24, 2025 · Big Data

How Turing Data Finder Transforms Growth Analysis with a Unified Data Platform

The article provides a detailed technical overview of the Turing Data Finder (TDF) platform, describing its background, core components, data schema, ingestion workflow, and a suite of growth‑analysis features such as event, retention, funnel, path, component, distribution, and attribution analysis, while also outlining performance‑optimisation techniques and future development directions.

SQL optimizationTuring Data Finderbig data
0 likes · 17 min read
How Turing Data Finder Transforms Growth Analysis with a Unified Data Platform
Big Data Technology & Architecture
Big Data Technology & Architecture
Mar 3, 2025 · Big Data

The Turning Point for Data Development: From Traditional Data Engineering to AI Data Engineering

The article analyzes how the rapid rise of open‑source large‑model AI in 2025 is reshaping the data development profession, urging developers to transition from specialized data‑engineer roles to full‑stack AI data engineering skills such as distributed computing, lake‑house architectures, and model tuning.

AIDistributed ComputingFlink
0 likes · 7 min read
The Turning Point for Data Development: From Traditional Data Engineering to AI Data Engineering
ITPUB
ITPUB
Feb 11, 2025 · Operations

Why Your Monitoring Fails and How to Build Effective Observability Data

Many companies deploy fragmented monitoring and observability tools yet still struggle to pinpoint incidents; this article analyzes the root causes—under‑utilized tools and scenario‑agnostic data—and offers practical steps to organize metrics, build layered insights, and improve fault‑resolution efficiency.

SREdata engineeringincident response
0 likes · 12 min read
Why Your Monitoring Fails and How to Build Effective Observability Data
Big Data Technology Architecture
Big Data Technology Architecture
Feb 8, 2025 · Big Data

How AI Can Accelerate Data Engineering: Practical DeepSeek Use Cases and Tips

This article shows how AI tools like DeepSeek can dramatically speed up data‑engineering tasks—such as fixing long‑running SQL queries, building real‑time data pipelines with Flink, and deciphering legacy stored procedures—while offering concrete prompts, real‑world case studies, and five time‑saving techniques.

DeepSeekSQL optimizationautomation
0 likes · 6 min read
How AI Can Accelerate Data Engineering: Practical DeepSeek Use Cases and Tips
DataFunSummit
DataFunSummit
Feb 5, 2025 · Artificial Intelligence

Exploration and Practice of Large‑Model Data Construction

This presentation details engineering‑focused approaches to building, mixing, and filtering data for large language models, covering data preparation, pre‑training mix strategies such as DoReMi, DoGE and online sampling, post‑training data quality selection methods, and practical Q&A on scaling laws and PDF processing.

AIData MixingPretraining
0 likes · 15 min read
Exploration and Practice of Large‑Model Data Construction
21CTO
21CTO
Feb 4, 2025 · Big Data

Why Python Beats Java and Scala for Modern Data Engineering

The article compares Java, Scala, SQL, and Python for data‑engineering tasks, arguing that Python’s versatility, rich ecosystem, and ease of use make it the preferred language for both small‑scale and massive Spark workloads despite its performance trade‑offs.

SQLScalaSpark
0 likes · 7 min read
Why Python Beats Java and Scala for Modern Data Engineering
Big Data Technology & Architecture
Big Data Technology & Architecture
Jan 15, 2025 · Big Data

From Operations to Data Engineering: A Student’s Real‑World Journey and Practical Guide

This article shares a data‑engineering student’s personal experience—from a misaligned operations role to mastering big‑data technologies, building a portfolio, crafting a targeted resume, and navigating multi‑stage interviews—offering concrete advice and a structured learning roadmap for aspiring data professionals.

Interview PreparationLearning PathResume Writing
0 likes · 14 min read
From Operations to Data Engineering: A Student’s Real‑World Journey and Practical Guide
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jan 6, 2025 · Cloud Native

How Fluid Enables Seamless Dynamic Dataset Mounting for Cloud‑Native AI Development

PAI‑DSW leverages the Fluid project to provide a cloud‑native AI development platform where data scientists can dynamically mount and unmount OSS datasets on running Kubernetes pods without restarting, improving workflow efficiency and addressing the challenges of heterogeneous data source management in AI engineering.

AI developmentFluidKubernetes
0 likes · 18 min read
How Fluid Enables Seamless Dynamic Dataset Mounting for Cloud‑Native AI Development
JD Tech
JD Tech
Dec 30, 2024 · Big Data

Techniques for Writing Elegant and Efficient SQL in Big Data Environments

The article shares practical methods and code examples for making SQL both readable and high‑performing in large‑scale data platforms, covering predicate push‑down with subqueries, deduplication strategies, bucket utilization, and Python‑driven job parameter handling.

HivePerformanceSQL
0 likes · 14 min read
Techniques for Writing Elegant and Efficient SQL in Big Data Environments
dbaplus Community
dbaplus Community
Dec 24, 2024 · Big Data

How Bilibili Scaled Its Tag System for Massive Data and Real‑Time Accuracy

The article details Bilibili's comprehensive redesign of its tag system—including background challenges, architectural layers, technical upgrades like Iceberg integration and shard‑based ClickHouse writes, crowd selection methods, online service guarantees, performance metrics, and future plans—showcasing a data‑driven solution that boosts stability, speed, and business coverage.

ClickHouseDistributed ComputingOnline Service
0 likes · 24 min read
How Bilibili Scaled Its Tag System for Massive Data and Real‑Time Accuracy
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Dec 5, 2024 · Big Data

Interview with Jianchen: Journey from Open Source Contributor to Data Engineer at Xiaohongshu

In this interview, Xiaohongshu data engineer Jianchen recounts his evolution from a computer‑science student discovering open‑source through MIT6.824 to contributing to SOFAJRaft and Apache RocketMQ, detailing his OSPP projects, the decision to join Xiaohongshu, and his work on a cloud‑native Kafka engine that cut storage and compute usage by half.

Apache RocketMQKafkaSOFAJRaft
0 likes · 11 min read
Interview with Jianchen: Journey from Open Source Contributor to Data Engineer at Xiaohongshu
DataFunSummit
DataFunSummit
Dec 5, 2024 · Big Data

Ping An Financial Services' Big Data Platform Construction and Data Governance Practices

This article details Ping An Financial Services' journey in building a comprehensive big‑data platform, addressing fragmentation, low data timeliness, processing limits, and governance challenges through a four‑stage technical evolution, modular tool development, and a systematic data‑governance framework to support its digital transformation.

Financial Servicesdata engineeringdata governance
0 likes · 16 min read
Ping An Financial Services' Big Data Platform Construction and Data Governance Practices
ByteDance Data Platform
ByteDance Data Platform
Nov 6, 2024 · Big Data

How Douyin’s Data Platform Overcomes EB‑Scale Metric Challenges

This article explains how Douyin Group tackles massive data volume, quality, and efficiency issues by building a four‑layer intelligent platform, standardizing metric management, automating metric decomposition, and creating reusable metric services that boost agility, stability, and cross‑team collaboration.

Data Qualitybig datadata engineering
0 likes · 20 min read
How Douyin’s Data Platform Overcomes EB‑Scale Metric Challenges
Bilibili Tech
Bilibili Tech
Oct 25, 2024 · Big Data

DataFunSummit2024: Next-Generation Data Architecture Technology Summit

DataFunSummit2024, co-hosted by Bilibili, convenes industry experts, scholars, and enterprise leaders across six forums to discuss next‑generation data architecture, showcasing Bilibili’s Iceberg‑based stream‑batch innovations, AI‑BI analytics, NoETL practices, and emerging alternatives to Lambda architecture.

AI+BILambda ArchitectureNoETL
0 likes · 3 min read
DataFunSummit2024: Next-Generation Data Architecture Technology Summit
Baobao Algorithm Notes
Baobao Algorithm Notes
Oct 7, 2024 · Artificial Intelligence

Mastering LLM Supervised Fine‑Tuning: Practical Tips, Data Strategies, and Debugging

This article provides a comprehensive, experience‑driven guide to supervised fine‑tuning (SFT) of large language models, covering special tokens, latency considerations, data diversity and production, training frameworks and hyper‑parameters, over‑/under‑fitting diagnostics, and evaluation metrics such as helpfulness, honesty, and harmlessness.

AILLMSFT
0 likes · 40 min read
Mastering LLM Supervised Fine‑Tuning: Practical Tips, Data Strategies, and Debugging
AntData
AntData
Sep 26, 2024 · Artificial Intelligence

DB-GPT: Open-Source AI-Native Data Application Development Framework

DB‑GPT is an open‑source AI‑native data‑application framework that provides multi‑model management, Text‑to‑SQL optimization, RAG, multi‑agent collaboration, and intelligent workflow orchestration, enabling developers to build scalable large‑model database applications, with proven enterprise adoption, community growth, and academic publications.

AIRAGdata engineering
0 likes · 6 min read
DB-GPT: Open-Source AI-Native Data Application Development Framework
JD Retail Technology
JD Retail Technology
Sep 25, 2024 · Big Data

From a Personal Journey to Data Platform Architecture: Insights on Big Data, Cloud Computing, and System Design

The article narrates the author’s 30‑year programming career and shares technical reflections on building business‑agnostic, configurable data platforms, covering batch, streaming, interactive computing, big‑data sharding, Spark, Flink, cloud migration, and the philosophy of software architecture.

Batch ProcessingCloud ComputingSystem Design
0 likes · 23 min read
From a Personal Journey to Data Platform Architecture: Insights on Big Data, Cloud Computing, and System Design
AntTech
AntTech
Sep 10, 2024 · Big Data

From DATA for AI to AI for DATA: Evolution of Ant Group’s Intelligent Data System

The talk reviews the rapid evolution of data technologies—from early database foundations and big‑data breakthroughs to the rise of generative AI—highlighting how Ant Group’s data platform is shifting from a cost‑efficiency focus to a value‑centric, multimodal, AI‑driven ecosystem.

Artificial IntelligenceData PlatformsMultimodal Data
0 likes · 17 min read
From DATA for AI to AI for DATA: Evolution of Ant Group’s Intelligent Data System