SIMPSON’S PARADOX : THE MIRAGE OF DATA SCIENCE
The Great War of Mathematics and Simpson’s Paradox
Simpson’s Paradox : The Mirage Of Data Science
The Great War of Mathematics
Let’s journey back in time to the year 1711.
You find yourself in the midst of The Great War of Mathematics, a fierce intellectual battle fought between two legendary nemeses: Isaac Newton and Gottfried Wilhelm Leibniz. Don’t worry — you’re perfectly safe here (no arrows of logic or razor-sharp deductions will harm you).
Both titans wield armies of unimaginable power, consisting of axioms and theorems, ready to clash for mathematical supremacy. Have you ever wondered what a true battlefield of mathematics might look like? If not, let’s imagine it today!
Since childhood, I’ve dreamed of having a hypothesis named after me. Guess what? I’ve concocted this completely nonsensical one here 😅.
With that, I invite you to embark on a journey through my hypothetical world, where math is not just theory — it’s warfare.
The Setup
- Axioms are like fearless infantry, forming the foundation of the army.
- Theorems are the cavalry, swift and powerful, riding in to secure victory.
Newton’s army boasts a formidable force of infantry N_i and cavalry N_c. Meanwhile, Leibniz’s troops are equally valiant, with infantry L_i and cavalry L_c. Both armies are of the same total strength, setting the stage for a balanced confrontation.
The battlefield is set. The war rages on for five long years, filled with intellectual maneuvers, strategic deductions, and relentless battles of wits. But in 1716, just as Leibniz’s forces seem poised for victory, disaster strikes. The great Leibniz — not by logic or a rival theorem, but by a cruel irony — misplaces his spectacles and vanishes from the battlefield, leaving his ideas to falter.
Both armies suffer tremendous losses over the course of this great war of ideas. As a renowned mathematician in this hypothetical world, I now turn to the statistics that reveal the true scale of these intellectual casualties. Let’s examine the graph below, where we begin to unravel the mystery behind the staggering losses on both sides.
Figure 1
The front that suffered fewer casualties was victorious.
Which side do you think suffered greater casualties?
Many of you might assume that Newton’s side took the brunt of the losses, especially since both his infantry and cavalry suffered a higher percentage of casualties. The evidence seems clear — Newton’s forces appeared to be in far worse shape.
But Newton did not lose!!!!
This revelation struck me like a thunderclap. I, too, was shaken by the truth of it. Yet, the data does not lie. When I examined the overall statistics, the outcome was irrefutable. Those still skeptical of my words, I invite you to glance through the following graph.
Figure 2
“Overall casualty rate of Newton could be less than that of Leibniz if at least one of the Newton’s regiment’s casualty rate had been less than that of Leibniz . There was no way that the weighted average of casualty rates could show such contrary results.”
— Common Sense
Now, I ask you to take a moment and reflect on where the anomaly might lie. Why does our common sense seem to betray us? How could Newton’s side, despite higher casualties in both infantry and cavalry, emerge victorious?
The timing is perfect for a revelation!
Brace yourself — here’s the formal definition of Simpson’s Paradox:
Simpson’s paradox is a phenomenon in probability and statistics in which a trend appears in several groups of data but disappears or reverses when the groups are combined.
While you ponder this, let us dive deeper into the demographics of both armies. Understanding the makeup of Newton and Leibniz’s forces may provide the key to unraveling this perplexing paradox.
Notice the stark difference between the two armies. After examining these pie charts, the entire paradox unraveled before my eyes, reducing itself to nothing more than a simple oversight — confounding variables. Do you see it? Let me explain.
Confounding variables are factors that influence other variables in such a way that they create misleading or distorted associations between them.
Earlier, we failed to account for a crucial variable: the composition of the army. Once we incorporate the effect of this, the mystery dissolves, and we arrive at a remarkably intuitive explanation for the Simpson’s Paradox we encountered earlier.
WAS IT REALLY A PARADOX ?
Although Newton’s infantry suffered a higher casualty rate than Leibniz’s, it made up only 20% of Newton’s army. In other words, while nearly all soldiers in Newton’s infantry regiment perished, the regiment itself was relatively small.
In contrast, Leibniz’s army had a much larger proportion of infantry, with a significant casualty rate of around 50%. This combination of a larger infantry proportion and a higher casualty rate contributed to the overall casualty rate of Leibniz’s army surpassing that of Newton’s.
Lets put the Mathematics going behind in front of you.
From the graphs, the following observations are very straightforward:

And you see, the mathematical results are consistent with the figure 2 .
If data can be partitioned in such a way that a single Simpson’s reversal emerges, then, logically, it seems inevitable that this reversal will continue as the data is divided even further.
This very phenomenon is vividly illustrated in Figure 3, where the casualty rates of four distinct military units are laid bare: Infantry I (the swordsmen of the first battalion), Infantry II (the archers of the second battalion), Cavalry I (the nimble, lightly armored cavalry), and Cavalry II (the mighty, heavily armored cavalry). With this partitioning, the data splits into four subsets, each telling its own tale. In every single one, Leibniz’s battalion suffers greater casualties than Newton’s — at least, that’s what the numbers suggest.
Figure 3
Yet, as we zoom out and aggregate the data to the level of regiments, the tides begin to shift, and we encounter Figure 1. Here, a stunning reversal occurs — Leibniz’s army now appears to have outperformed Newton’s! But the twists do not end there. As we continue to aggregate the data, adding more layers, the final conclusion arrives, and once again, it’s Newton’s army that emerges victorious considering the overall casualty rate.
Simpson’s paradox says however that this can occur; it is possible for any relation between two variables to be reversed in every component subpopulation
This chain of reversals begs the ultimate question: Which army truly proved superior?
I now present you with even more detailed demographics of both armies and leave it to you to verify the trend reversal we’ve just witnessed.
Do you want to hear one more truth?
Newton and Leibniz never participated in any actual wars 😜 . While their intellectual battle over the discovery of calculus is legendary, they never fought on a physical battlefield. This fact serves as a powerful reminder to snap us out of the hypothetical world we’ve been exploring.
Time to be serious
It’s time to set aside our imaginative scenarios and consider the true implications of Simpson’s Paradox. You might be wondering: How does Simpson’s Paradox manifest itself in the real world?
It turns out, this paradox isn’t just a theoretical quirk — it appears in practical situations more often than you might think. Even during the COVID-19 pandemic, Simpson’s Paradox emerged: when comparing Case Fatality Rates (CFR) between China and Italy, Italy initially seemed to have a higher survival rate, but when the data was split by age group, the conclusion flipped.
For now, let’s focus on one of the most well-known examples — a study comparing two approaches for treating kidney stones.
A study by statistician Steven A. Julious and senior statistical programmer Mark A. Mullee compares two kidney stone treatments: one cutting-edge and minimally invasive, the other a reliable, long-standing method. The success rates, sorted by stone size, suggest a clear winner.
But is there more to consider?
Reference:
Steven A. Julious and Mark A. Mullee (3 December 1994). Confounding and Simpson’s paradox. BMJ, 309(6967): 1480–1481. doi: 10.1136/bmj.309.6967.1480. PMC: 2541623. PMID: 7804052.
Unraveling the Mystery: The Simpson’s Paradox at Work
Let’s break this down step by step. At first glance, the data seems pretty straightforward. When you look at the treatments for small and large stones separately, the outcome is clear — treatment A consistently outperforms treatment B for both small and large stones. So far, so good, right? Treatment A is the obvious winner in each case.
But here’s where things take an unexpected turn: when you combine the data for both small and large stones together, suddenly treatment B looks like the better treatment overall.
This seeming contradiction is rooted in the distribution of the data.
Here’s what happened:
More patients with small stones, (which are generally easier to treat) were assigned treatment B. At the same time, a larger proportion of patients with large stones, (which are much harder to treat successfully) were assigned treatment A. The size of the kidney stone, it turns out, has a much greater impact on treatment success than the choice of treatment itself. Smaller stones have a much higher chance of success, regardless of which treatment is used.
When we combine the groups, it creates an illusion that treatment B is more effective because more patients with small stones received it, and those patients had higher success rates. But if the distribution of patients between the two treatments had been more balanced, this paradox would not have appeared.
This is exactly why you can’t always trust the overall numbers without considering how the data is divided. Without factoring in the size of the kidney stone, you might draw the wrong conclusions.
This lesson extends beyond kidney stone treatments — it’s a reminder that data can be misleading if not carefully interpreted. The numbers may seem clear at first, but they can hide deeper truths. Simpson’s Paradox teaches us to question assumptions and dig deeper before drawing conclusions.
**Do still you believe the statistical phenomenon described above should be considered a paradox?**
Although criticized, Simpson’s paradox remains a fascinating and frequently explored topic in statistics and data analysis. Its ongoing study across diverse disciplines highlights the importance of nuanced statistical methods and the dangers of oversimplifying data interpretations.
References:
- Simpson’s Paradox and Investment Management : Research paper by *Gilbert W. Bassett Jr.*
- **Simpson’s Paradox in Clinical Research.**
- **Wikipedia**
Written by: Naman Patidar , Rupant Dixit, Swarnim Verma
메타데이터
- post_id
- 561978fc087c
- slug
- simpsons-paradox-the-mirage-of-data-science-561978fc087c
- url
- https://medium.com/stamatics-iit-kanpur/simpsons-paradox-the-mirage-of-data-science-561978fc087c
- canonical_url
- https://medium.com/stamatics-iit-kanpur/simpsons-paradox-the-mirage-of-data-science-561978fc087c
- author_url
- https://medium.com/@swarnimverma0
- status
- ok
- fetched_at
- 2026-06-14 13:58:26