Big Data 15 min read

Data Planning Blueprint: From Real Needs to DRD and Collection

This article outlines a comprehensive methodology for data planning, covering historical context, requirement analysis with subtractive thinking, Data Requirements Document (DRD) design, data flow diagrams, data dictionaries, and internal versus external data collection strategies including web scraping considerations.

Smart Sea Tide
Smart Sea Tide
Smart Sea Tide
Data Planning Blueprint: From Real Needs to DRD and Collection

Historical Context of Data Evolution

Since the 1980s, mature database technology and high-performance storage have driven exponential data growth. Futurist John Naisbitt noted we are "drowning in information but starved for knowledge." Early enterprise data mining (1994 Donnelley Marketing survey: 56% of US retailers/manufacturers had marketing databases) led to Business Intelligence (BI), defined by Gartner Group in 1996 as fact-based decision support using data warehouses, OLAP, and data mining. By 1998, SGI chief scientist John Mashey coined "Big Data" to describe challenges of volume, velocity, variety, and veracity. Tu Zipei emphasized big data's value lies in exchanging, integrating, and analyzing massive datasets to create "big knowledge, big technology, big profit, and big development."

Analyzing Requirements: Uncover Real Needs

Find the Requestor and Interpret Current State

Many data products fail due to misidentified needs. The first step is locating the direct requestor to avoid distortion through layers. Customers often articulate solutions ("faster horse") rather than underlying needs (speed). Requirement analysis begins with interpreting the client's basic profile and current dissatisfaction. Two practical methods: collect stakeholder complaints to find common pain points, and ask for critiques of existing solutions to reveal priorities.

Subtractive Thinking: Focus on Core Value

Avoid three pitfalls: feature stuffing, blindly copying competitors, and forcing immature cutting-edge technologies. In big data projects, atomic capabilities are similar; differentiation comes from scenario penetration, technical integration, and sustainable upgradability. Subtractive thinking means removing unnecessary or mismatched tasks to increase efficiency. Contrast with additive thinking (bucket theory) which spreads limited energy across weaknesses. Instead, prioritize the most productive area — what the customer is willing to pay for. List solution modules, iteratively cut the least important; the last remaining module indicates highest customer value, enabling phased planning.

Data Inventory: Prepare Data Requirements

Data Requirements Document (DRD)

Unlike traditional fixed business logic, big data product value depends on data quality ("garbage in, garbage out"). DRD serves as a communication artifact for development teams, comprising three parts:

Source : Which system, interface, update frequency.

Measures : Indicator definitions and calculation logic.

Dimensions : Data elements for slicing metrics, defining granularity.

DRD produces two complementary deliverables: data flow diagrams and data dictionaries.

Data Flow Diagram (DFD)

DFD (Data Flow Diagram) graphically describes system functions, inputs, outputs, and data stores logically. Components:

Data Flow : Fixed-composition data moving between components; must be named except to/from data stores.

Process : Transforms input flows to output flows; each has a name and hierarchical ID.

Data Store : Temporary data at rest; named.

External Entity : Outside persons/organizations that originate or receive data.

Data Flow Diagram components
Data Flow Diagram components

Data Dictionary

Defines every DFD element. In database design, it also describes table schemas: field names, data types, primary/foreign keys.

Data Dictionary example
Data Dictionary example

Data Collection: Internal and External Sources

Internal Data

Generated from organizational business processes and operations: purchase records, logistics, reviews. Internet companies also capture implicit feedback via event tracking (buried points) for user behavior.

External Data

Data generated outside the organization: macroeconomic indices, census, industry prices. Valuable when internal data is insufficient (startups), poor quality, or when internal metrics change without clear drivers. Sources include public data portals and web scraping (crawlers), though scraping carries legal risks.

Internal vs External Data Comparison
Internal vs External Data Comparison
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

big datadata collectionrequirement analysisdata dictionarydata flow diagramdata planningDRDsubtractive thinking
Smart Sea Tide
Written by

Smart Sea Tide

Sharing cutting‑edge big data and AI technologies, with occasional lifestyle insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.