Tagged articles

Semi-structured Data

9 articles · Page 1 of 1
StarRocks
StarRocks
Sep 3, 2026 · Databases

Querying Paimon Semi-Structured Data with StarRocks: Variant, Shredding & SQL Practices

This article explains how StarRocks queries Paimon Variant data, covering the differences between regular Parquet JSON and Variant, the three-phase read architecture, Shredding optimization for hot paths, practical SQL examples using get_variant_* functions, and a decision framework for choosing between JSON, Plain Variant, Shredded Variant, and formal typed columns.

LakehousePaimonParquet
0 likes · 21 min read
Querying Paimon Semi-Structured Data with StarRocks: Variant, Shredding & SQL Practices
LuTiao Programming
LuTiao Programming
Apr 26, 2026 · Industry Insights

Is JSON Quietly Fading? How AI Is Redefining Data Interaction

The article examines how JSON, once the universal data language for backend systems, is being pushed back from the interaction layer in AI‑driven applications, outlines its five key shortcomings, and explores emerging alternatives such as natural‑language interfaces, Markdown, YAML, and function calling.

AIData InterchangeFunction Calling
0 likes · 9 min read
Is JSON Quietly Fading? How AI Is Redefining Data Interaction
Data Integration and Governance
Data Integration and Governance
Dec 29, 2025 · Fundamentals

Master Structured, Semi‑Structured, and Unstructured Data: A Beginner’s Guide to Data Governance

The article explains the three core data categories—structured, semi‑structured, and unstructured—illustrates their characteristics with real manufacturing examples, compares suitable storage and processing tools, and outlines three practical principles for effective data integration and governance.

Data IntegrationETLSemi-structured Data
0 likes · 9 min read
Master Structured, Semi‑Structured, and Unstructured Data: A Beginner’s Guide to Data Governance
Data Integration and Governance
Data Integration and Governance
Dec 8, 2025 · Fundamentals

Structured, Semi‑Structured, and Unstructured Data: Choosing the Right Storage for IoT Sensor Streams

The article explains the definitions and practical differences of structured, semi‑structured, and unstructured data, compares suitable storage and processing technologies, discusses governance challenges, and offers a step‑by‑step strategy—including a data‑integration tool example—to help IoT projects select the optimal data architecture.

Data StorageIoTSemi-structured Data
0 likes · 10 min read
Structured, Semi‑Structured, and Unstructured Data: Choosing the Right Storage for IoT Sensor Streams
Big Data Technology & Architecture
Big Data Technology & Architecture
Oct 21, 2024 · Big Data

Key New Features of Apache Doris 3.0: Storage‑Compute Separation, Lakehouse Integration, Semi‑Structured Data, ETL Enhancements, Materialized Views, and Java UDTF

Apache Doris 3.0 introduces storage‑compute separation, native lakehouse write‑back, optimized Variant handling for semi‑structured data, stronger ETL transaction support, enhanced multi‑table materialized views, and Java UDTF capabilities, providing developers with more flexible, cost‑effective, and high‑performance analytics solutions.

Apache DorisETLJava UDTF
0 likes · 7 min read
Key New Features of Apache Doris 3.0: Storage‑Compute Separation, Lakehouse Integration, Semi‑Structured Data, ETL Enhancements, Materialized Views, and Java UDTF
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jan 31, 2024 · Artificial Intelligence

Advanced RAG with Semi‑Structured Data Using LangChain, Unstructured, and ChromaDB

This tutorial demonstrates how to build an advanced Retrieval‑Augmented Generation (RAG) system for semi‑structured PDF data by leveraging LangChain, the unstructured library, ChromaDB vector store, and OpenAI models, covering installation, PDF partitioning, element classification, summarization, and query execution.

AIChromaDBLangChain
0 likes · 11 min read
Advanced RAG with Semi‑Structured Data Using LangChain, Unstructured, and ChromaDB
DataFunTalk
DataFunTalk
Jan 1, 2024 · Big Data

MaxCompute Semi-Structured Data: Concepts, Solutions, and Benefits

This article explains the nature of semi‑structured data, compares traditional schema‑on‑read and schema‑on‑write approaches, and details MaxCompute's columnar storage solution that balances flexibility, performance, and cost for large‑scale data warehouses.

MaxComputeSemi-structured Databig data
0 likes · 19 min read
MaxCompute Semi-Structured Data: Concepts, Solutions, and Benefits
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Sep 14, 2023 · Big Data

How MaxCompute Turns Semi‑Structured Data into High‑Performance Columnar Storage

This article explains the nature of semi‑structured data, compares schema‑on‑read and schema‑on‑write approaches, and shows how Alibaba Cloud MaxCompute leverages columnar storage and dynamic parsing to achieve low‑cost, high‑performance analytics for large‑scale data workloads.

MaxComputeSemi-structured Datacolumnar storage
0 likes · 20 min read
How MaxCompute Turns Semi‑Structured Data into High‑Performance Columnar Storage
DataFunSummit
DataFunSummit
Sep 7, 2023 · Big Data

MaxCompute Semi-Structured Data Solutions: Architecture, Comparison, and Performance Benefits

This article explains the concepts of semi‑structured data, compares traditional schema‑on‑read and schema‑on‑write approaches, and details MaxCompute's columnar storage solution—including AliORC, adaptive query processing, and handling of dirty or sparse data—to achieve high performance and low cost in big‑data warehousing.

MaxComputeSemi-structured Data
0 likes · 20 min read
MaxCompute Semi-Structured Data Solutions: Architecture, Comparison, and Performance Benefits