Fundamentals 6 min read

Why Averages Lie: Simpson's Paradox Explained with a Concrete Example

This article explains Simpson's Paradox through a numerical example where Plan A outperforms Plan B in both simple and difficult tasks individually, but Plan B appears better when aggregated due to differing task distributions, illustrating how hidden variables can reverse aggregate conclusions and how to avoid such pitfalls.

Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Why Averages Lie: Simpson's Paradox Explained with a Concrete Example

Introduction

We often use averages to compare two plans, groups, or strategies. Averages appear objective, concise, and easily fit into headlines and conclusions. However, sometimes the direction of the comparison reverses when data are combined versus when they are examined separately. This does not necessarily mean the data are wrong; rather, an important variable may be hidden behind the average.

01 | Both Tables Are Correct, Yet Conclusions Are Opposite

Suppose two plans handle simple tasks and difficult tasks. The success rates are:

Simple tasks: Plan A 90/100 = 90%; Plan B 80/100 = 80%

Difficult tasks: Plan A 20/1000 = 2%; Plan B 1/100 = 1%

In each subgroup, Plan A has a higher success rate than Plan B. Yet when all tasks are merged:

Plan A: 110 successes out of 1100 attempts ≈ 10%

Plan B: 81 successes out of 200 attempts ≈ 40.5%

Plan A is better in every subgroup but worse in the overall average. This is the famous Simpson's Paradox .

Table showing subgroup and aggregated success rates for Plan A and Plan B
Table showing subgroup and aggregated success rates for Plan A and Plan B

02 | What Changes the Conclusion Is Not the Result, But the Distribution

Examining the data reveals that the two plans face completely different task structures. Plan A handles a large volume of difficult tasks (1000), while Plan B handles mostly simple tasks (100). Difficult tasks are inherently harder to succeed at. When different plans have different task compositions, directly comparing overall averages mixes the "task difficulty" variable into the result.

An average only tells us the final outcome; it does not actively reveal which groups the data come from, what proportion each group occupies, or whether the two plans actually faced the same problem. When this information is hidden, a seemingly fair average may lose its comparative meaning.

Illustration of differing task distributions between Plan A and Plan B
Illustration of differing task distributions between Plan A and Plan B

03 | Group First, Then Compare

When encountering an average, one can first ask three questions (illustrated in the article's diagram). If the overall conclusion contradicts the subgroup conclusions, further investigation is needed to check for hidden variables. Many disputes arise not because someone calculated incorrectly, but because the two sides are effectively comparing different sets of objects.

Diagram suggesting three questions to ask when seeing an average
Diagram suggesting three questions to ask when seeing an average

04 | How to Avoid Being Misled by Averages

A more reliable reporting approach should retain at least three types of information: subgroup-level results, group sizes or proportions, and the context defining the groups. When condition differences are large, further methods such as standardization, stratified comparison, or sensitivity analysis can be used to test whether the conclusion depends on a particular sample distribution.

Data analysis is not about finding the most convenient number; it is about confirming that the number has not averaged away important information.

Summary of recommended reporting practices to avoid Simpson's Paradox
Summary of recommended reporting practices to avoid Simpson's Paradox

Conclusion

The average does not deceive us. What misleads us is forgetting to ask: what is this average composed of? Good data analysis does not only ask whose average is higher, but continues to ask: what exactly has the average hidden?

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

StatisticsData AnalysisSimpson's ParadoxData InterpretationAveragesHidden VariablesStratified Comparison
Network Intelligence Research Center (NIRC)
Written by

Network Intelligence Research Center (NIRC)

NIRC is based on the National Key Laboratory of Network and Switching Technology at Beijing University of Posts and Telecommunications. It has built a technology matrix across four AI domains—intelligent cloud networking, natural language processing, computer vision, and machine learning systems—dedicated to solving real‑world problems, creating top‑tier systems, publishing high‑impact papers, and contributing significantly to the rapid advancement of China's network technology.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.