Data Planning Blueprint: From Real Needs to DRD and Collection
This article outlines a comprehensive methodology for data planning, covering historical context, requirement analysis with subtractive thinking, Data Requirements Document (DRD) design, data flow diagrams, data dictionaries, and internal versus external data collection strategies including web scraping considerations.
Historical Context of Data Evolution
Since the 1980s, mature database technology and high-performance storage have driven exponential data growth. Futurist John Naisbitt noted we are "drowning in information but starved for knowledge." Early enterprise data mining (1994 Donnelley Marketing survey: 56% of US retailers/manufacturers had marketing databases) led to Business Intelligence (BI), defined by Gartner Group in 1996 as fact-based decision support using data warehouses, OLAP, and data mining. By 1998, SGI chief scientist John Mashey coined "Big Data" to describe challenges of volume, velocity, variety, and veracity. Tu Zipei emphasized big data's value lies in exchanging, integrating, and analyzing massive datasets to create "big knowledge, big technology, big profit, and big development."
Analyzing Requirements: Uncover Real Needs
Find the Requestor and Interpret Current State
Many data products fail due to misidentified needs. The first step is locating the direct requestor to avoid distortion through layers. Customers often articulate solutions ("faster horse") rather than underlying needs (speed). Requirement analysis begins with interpreting the client's basic profile and current dissatisfaction. Two practical methods: collect stakeholder complaints to find common pain points, and ask for critiques of existing solutions to reveal priorities.
Subtractive Thinking: Focus on Core Value
Avoid three pitfalls: feature stuffing, blindly copying competitors, and forcing immature cutting-edge technologies. In big data projects, atomic capabilities are similar; differentiation comes from scenario penetration, technical integration, and sustainable upgradability. Subtractive thinking means removing unnecessary or mismatched tasks to increase efficiency. Contrast with additive thinking (bucket theory) which spreads limited energy across weaknesses. Instead, prioritize the most productive area — what the customer is willing to pay for. List solution modules, iteratively cut the least important; the last remaining module indicates highest customer value, enabling phased planning.
Data Inventory: Prepare Data Requirements
Data Requirements Document (DRD)
Unlike traditional fixed business logic, big data product value depends on data quality ("garbage in, garbage out"). DRD serves as a communication artifact for development teams, comprising three parts:
Source : Which system, interface, update frequency.
Measures : Indicator definitions and calculation logic.
Dimensions : Data elements for slicing metrics, defining granularity.
DRD produces two complementary deliverables: data flow diagrams and data dictionaries.
Data Flow Diagram (DFD)
DFD (Data Flow Diagram) graphically describes system functions, inputs, outputs, and data stores logically. Components:
Data Flow : Fixed-composition data moving between components; must be named except to/from data stores.
Process : Transforms input flows to output flows; each has a name and hierarchical ID.
Data Store : Temporary data at rest; named.
External Entity : Outside persons/organizations that originate or receive data.
Data Dictionary
Defines every DFD element. In database design, it also describes table schemas: field names, data types, primary/foreign keys.
Data Collection: Internal and External Sources
Internal Data
Generated from organizational business processes and operations: purchase records, logistics, reviews. Internet companies also capture implicit feedback via event tracking (buried points) for user behavior.
External Data
Data generated outside the organization: macroeconomic indices, census, industry prices. Valuable when internal data is insufficient (startups), poor quality, or when internal metrics change without clear drivers. Sources include public data portals and web scraping (crawlers), though scraping carries legal risks.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Smart Sea Tide
Sharing cutting‑edge big data and AI technologies, with occasional lifestyle insights.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
