DataFunSummit
Author

DataFunSummit

Official account of the DataFun community, dedicated to sharing big data and AI industry summit news and speaker talks, with regular downloadable resource packs.

1.8k
Articles
0
Likes
8.2k
Views
0
Comments
Recent Articles

Latest from DataFunSummit

100 recent articles max
DataFunSummit
DataFunSummit
Jun 23, 2026 · Artificial Intelligence

Financial Large Language Models: Architecture Shifts, Engineering Lessons, and Cutting‑Edge Agent Strategies

The article analyzes how strict compliance, data‑security, and rigorous business requirements reshape financial large‑model deployments, detailing a PageIndex‑based retrieval architecture, engineering pitfalls such as rule explosion and prompt bloat, model‑selection trade‑offs, and forward‑looking agent‑centric designs.

Model SelectionRetrieval Augmentationagentic AI
0 likes · 11 min read
Financial Large Language Models: Architecture Shifts, Engineering Lessons, and Cutting‑Edge Agent Strategies
DataFunSummit
DataFunSummit
Jun 22, 2026 · Artificial Intelligence

Building DataFlow: An Industrial‑Grade LLM Data Pipeline from Documents to Training

The article presents DataFlow, an open‑source, GPU‑centric data‑engineering framework that tackles LLM data‑preparation bottlenecks by defining a two‑level operator taxonomy, a LLM‑driven WebAgent for automatic crawling, a PDF‑to‑Markdown MinerU, a Ray‑based distributed runtime, and extensive multimodal extensions, and validates the design with quantitative experiments showing significant quality gains across math, code, and reasoning benchmarks.

Data PipelineDataFlowLLM
0 likes · 14 min read
Building DataFlow: An Industrial‑Grade LLM Data Pipeline from Documents to Training
DataFunSummit
DataFunSummit
Jun 22, 2026 · Industry Insights

From Old Wine to AI‑Native Teams: The Truth of Ontology Governance in AI

During a DataFunTalk roundtable, industry veterans from Huawei, Ping An and a startup dissected ontology as a management challenge, exposed the paradox that modeling pains business more than IT, warned of hidden technical debt in flashy AI projects, and shared hard‑won lessons on building AI‑Native organizations from the ground up.

AI Native OrganizationAI governanceOntology
0 likes · 17 min read
From Old Wine to AI‑Native Teams: The Truth of Ontology Governance in AI
DataFunSummit
DataFunSummit
Jun 21, 2026 · Artificial Intelligence

Unified Scheduling Optimization for xLLM in Complex Business Scenarios

This article analyzes how the xLLM open‑source LLM inference engine tackles the coexistence of multiple priority levels and strict SLO latency targets by introducing a dynamic, SLO‑aware batch scheduler and a PD‑separation architecture that improve throughput and SLO satisfaction across diverse workloads.

Hierarchical Block ManagerKV cachePD Separation
0 likes · 13 min read
Unified Scheduling Optimization for xLLM in Complex Business Scenarios
DataFunSummit
DataFunSummit
Jun 21, 2026 · Artificial Intelligence

How OpenClaw Transforms Traditional Enterprise Data Asset Architecture

The article analyzes the limitations of conventional data asset architectures for AI, introduces OpenClaw's layered, operator‑driven platform design, details the three components of high‑quality datasets, and shares practical implementation insights and challenges from a real‑world deployment.

AI data architectureAgentData Governance
0 likes · 13 min read
How OpenClaw Transforms Traditional Enterprise Data Asset Architecture
DataFunSummit
DataFunSummit
Jun 20, 2026 · Artificial Intelligence

Harness Engineering: Execution Control, Safety Boundaries, Human‑AI Collaboration, and Multi‑Agent Design

In a 90‑minute DataFunTalk live session, experts Huang Jia, Qu Xiangmou and Yao Binbin dissect ten critical challenges of moving AI agents from demo to production—covering sandbox vs permission boundaries, checkpoint design, rollback strategies, tool‑call safety, multi‑agent coordination, human‑in‑the‑loop control, observability, and memory management—to illustrate how rigorous engineering, not just model capability, enables trustworthy, controllable agents.

AI AgentsExecution ControlObservability
0 likes · 18 min read
Harness Engineering: Execution Control, Safety Boundaries, Human‑AI Collaboration, and Multi‑Agent Design
DataFunSummit
DataFunSummit
Jun 20, 2026 · Big Data

Building an Agentic Analytics Platform for the Gaming Industry with SelectDB

The article analyzes the fourfold challenges of game‑industry data analysis—high timeliness, massive concurrency, heterogeneous sources, and petabyte‑scale volumes—and explains how SelectDB’s evolution to an AI‑Ready, Agentic platform with MCP and a semantic layer addresses these issues through real‑time OLAP, multimodal processing, and autonomous decision loops.

AI-ReadyBig DataGame Data Analytics
0 likes · 16 min read
Building an Agentic Analytics Platform for the Gaming Industry with SelectDB
DataFunSummit
DataFunSummit
Jun 19, 2026 · Artificial Intelligence

Mastering Data Acquisition for AI Agents: From Crawler Pitfalls to MCP Browser Control

The article distills three Bright Data webinars, detailing how to overcome traditional web‑crawling challenges with an adaptive Crawler API, integrate the Model Context Protocol (MCP) for human‑like browser control, and build a LangGraph‑powered AI search engine while addressing compliance, billing, and scaling considerations.

AI AgentsAPI billingBright Data
0 likes · 15 min read
Mastering Data Acquisition for AI Agents: From Crawler Pitfalls to MCP Browser Control
DataFunSummit
DataFunSummit
Jun 19, 2026 · Artificial Intelligence

Why Memory Bottlenecks AI Agents: Inside MemOS Architecture and 200% Cloud Usage Surge

The article analyzes how memory has become the critical bottleneck for AI agents, compares model‑driven and application‑driven memory approaches, details the five‑layer MemOS framework, reports cloud service call growth of over 200% and token‑cost reductions of up to 72%, and shows real‑world enterprise deployments such as OpenClaw and ClawForce.

AI AgentsAI memoryCloud Services
0 likes · 16 min read
Why Memory Bottlenecks AI Agents: Inside MemOS Architecture and 200% Cloud Usage Surge
DataFunSummit
DataFunSummit
Jun 19, 2026 · Big Data

Near‑Real‑Time Data Warehousing with Yunqi Lakehouse: Cases from Xiaohongshu, Kuaishou, Meituan

The article examines how Xiaohongshu, Kuaishou and Meituan adopted Yunqi Lakehouse’s General Incremental Computing and Single‑Engine architecture to achieve near‑real‑time data warehouses, cutting resource usage to as low as 1/20 of full‑batch jobs, reducing data latency from days to minutes, and improving query performance.

Big DataGeneral Incremental ComputingReal-Time Data Warehouse
0 likes · 12 min read
Near‑Real‑Time Data Warehousing with Yunqi Lakehouse: Cases from Xiaohongshu, Kuaishou, Meituan