Onpode
Cover art for Confounding variables: why correlation breaks causation in epidemiology

Confounding variables: why correlation breaks causation in epidemiology

August 25, 2026 · 15 min

Hugo Vance & Lila Soto

For roughly twenty years, coffee appeared to raise heart disease risk — until researchers controlled for smoking, a confounding variable, and the finding completely inverted. One unmeasured third variable can reverse an entire conclusion. Residual confounding from unmeasured lifestyle factors means most observational nutrition findings carry irreducible uncertainty.

Observational epidemiology detects statistical associations between exposures and health outcomes by studying populations in real-world settings without experimental intervention.

0:0015:00
Get the next episode on Health

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Health

About this episode

For roughly twenty years, the epidemiological consensus said coffee raised the risk of heart disease. The finding was real — built on cohort data, tracking people over time. It was also wrong, in the way that a lot of careful science can be wrong: one unmeasured variable, smoking, was doing all the causal work. Control for it, and coffee flips from suspect to protective. This episode uses that reversal as a doorway into something harder. The coffee story is usually told as a triumph — error caught, adjustment made, science self-corrects. But the researchers who formalized confounding saw it differently. The incomparability between exposed and unexposed groups isn't a methodological slip. It's the structural default of observational research, because people choose their own exposures, and that choice is tangled up with everything else about how they live. The beta-carotene case makes the stakes concrete: observational studies suggested supplementation cut cancer risk, randomized trials later found harm in some populations, and the gap between those findings wasn't a variable anyone forgot to include. It was the entire life arrangement — walking habits, stress, financial security — of people who ate enough vegetables to register in the data. The episode then traces the partial escapes: randomized trials, natural experiments like the Dutch Hunger Winter, Mendelian randomization. Each one is real. Each one carries its own assumptions. When those assumptions break, the causal inference problem relocates rather than disappears. What to do with that — especially when silence is also a public health decision — is where the conversation ends up.

Frequently asked

What is a confounding variable in epidemiology?

A confounding variable is a third variable that sits alongside both the exposure and the outcome, making an association look causal when it isn't. In 1950s–70s cohort studies, smoking was the confounder between coffee and heart disease — coffee drinkers smoked more, so coffee looked harmful until smoking was controlled for.

Why did early studies wrongly link coffee to heart disease?

Coffee drinkers in 1950s–70s cohort studies smoked at higher rates. Smoking drove the heart disease risk, not coffee. Once researchers controlled for smoking, the finding completely inverted — coffee went from appearing harmful to modestly protective. The error persisted for roughly twenty years before the confounder was properly addressed.

What happened with beta-carotene supplements and cancer risk?

Observational studies through the 1980s suggested beta-carotene supplementation reduced cancer risk. Randomized trials then showed the opposite — harm in some populations, not merely a null result. The confounder was the entire lifestyle profile of people who ate enough vegetables to register in cohort studies, which no adjustment fully captured.

What is a directed acyclic graph (DAG) in epidemiology?

A DAG is a diagram epidemiologists use to visualize confounding. A third variable L is drawn with arrows pointing to both exposure A and outcome Y, forming a fork. Until that fork is blocked by measuring and holding L fixed, the A-to-Y path cannot be read as causal — the shape makes the problem visible.

Can randomized controlled trials eliminate confounding?

Randomization distributes both measured and unmeasured confounders equally across groups by chance, neutralizing the fork before it forms — something statistical adjustment cannot do. However, RCTs are impractical for most long-term public health questions like lifetime diet, and methods such as Mendelian randomization carry their own breakable assumptions.

Grounded in 12 sources
Causal Reasoning and Inference in Epidemiology · link.springer.com
Confounding: a routine concern in the interpretation of epidemiological studies - Statistical Methods in Cancer Research Volume V: Bias Assessment in Case–Control and Cohort Studies for Hazard Identif · ncbi.nlm.nih.gov
Methods for Evaluating Causality in Observational Studies: Part 27 of a Series on Evaluation of Scientific Publications - PMC · pmc.ncbi.nlm.nih.gov
Observational Research Rigor Alone Does Not Justify Causal Inference · pmc.ncbi.nlm.nih.gov
Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): Explanation and Elaboration - PMC · pmc.ncbi.nlm.nih.gov
Epidemiology, genetic epidemiology and Mendelian ... - PMC · pmc.ncbi.nlm.nih.gov
Perspective: Are Large, Simple Trials the Solution for Nutrition ... · pmc.ncbi.nlm.nih.gov
Observational Studies - PMC · pmc.ncbi.nlm.nih.gov
Limiting Dependence on Nonrandomized Studies and ... · sciencedirect.com
Observational Studies - an overview · sciencedirect.com
Necessary conditions for valid causal inference from ... · tandfonline.com
A Fresh Look at Problem Areas in Research Methodology in Nutrition · mdpi.com
Read transcript

Hugo Vance: Lila, good to be back — though I'll say, the thing I've been turning over before we sat down is one of those reversals that feels almost embarrassing when you say it out loud.

Lila Soto: Oh, I know the one — yeah, let's go there.

Hugo Vance: From the nineteen-fifties through the early nineteen-seventies, the prevailing epidemiological view — built on real cohort data, observational studies tracking people over time — was that coffee raised your risk of heart disease. That was the finding. That was what got communicated.

Lila Soto: And then it just... flipped.

Hugo Vance: Completely inverted — not softened, inverted — once researchers properly accounted for smoking. Coffee drinkers in those cohorts smoked at higher rates. Smoking is a confounder: a third variable sitting alongside both the exposure and the outcome, making the association look causal when it isn't. Control for it, and coffee goes from suspect to modestly protective.

Lila Soto: The confounder is doing the whole job, and coffee is just — standing nearby, looking guilty.

Hugo Vance: You see, that's actually the precise issue. It's like noticing that everyone who carries an umbrella gets wet, and concluding that umbrellas cause rain. The umbrella arrived because of the rain; it didn't produce it. The association is real — umbrella and wetness genuinely co-occur — but the causal story runs the other direction entirely.

Lila Soto: Oh — and that's the difference between association and causation, right there, in one image.

Hugo Vance: Precisely. Association means two variables co-occur statistically. Causation means changing one produces a change in the other. Observational studies — by design, because there's no random assignment — can demonstrate the former and cannot, on their own, establish the latter.

Lila Soto: And we ran the wrong story for, what — twenty years?

Hugo Vance: Roughly twenty years. Which is the part that doesn't let you settle. One measured variable — smoking — flipped the entire finding. The question that I think closes this doorway and opens a much larger one is: what are the unmeasured confounders doing right now, to the findings we currently accept as settled?

Lila Soto: Yeah — and that's exactly where I want to go, because I don't think most people listening have any idea that 'we controlled for some things' and 'we controlled for all the things' are completely different claims.

Hugo Vance: Well, and that gap — between 'some things' and 'all things' — is actually where I'd want to push back on the coffee story, because I think the way we tell it is too comfortable. The implication is: we caught the error, we adjusted, science worked. But that frames it as a blunder someone made. Kish — who's the person who actually gave 'confounding' its formal meaning — he was pointing at something structural. The incomparability between exposed and unexposed groups isn't a mistake. It's the default condition the moment you don't randomly assign people.

Lila Soto: The default condition — so it's not that we did observational studies badly.

Hugo Vance: By design. Because people choose their exposures. Coffee drinkers in 1960 were not randomly drawn from a hat. They were a type of person, with a type of life. And that life co-varies with other characteristics — unmeasured ones — by default.

Lila Soto: And that's the lifestyle clustering thing — it's almost a personality type inside the data. The person who's careful about one thing tends to be careful about other things, and you can't always measure which 'other things' are doing the work.

Hugo Vance: Yes. And epidemiologists now have a geometric way to see it — the directed acyclic graph, the DAG. You draw a fork: a third variable L with arrows pointing to both the exposure A and the outcome Y. That fork is confounding, made visible. Until you block L — meaning measure it, hold it fixed — you cannot read the A-to-Y path as causal. It's not philosophy. It's the shape of the problem.

Lila Soto: Huh — and the fork just... sits there in the diagram whether you want to see it or not.

Hugo Vance: Indeed. And Greenland and Robins in 1986 gave that fork its real teeth — counterfactual conditions. The question becomes: what would this person's outcome have been under the opposite exposure? You cannot observe that. A coffee drinker cannot also be, simultaneously, a non-coffee-drinker in the same life. So you're always filling in an absence.

Lila Soto: Wait, so the 1986 paper — that's not just a refinement of the coffee story, that's the thing that says the coffee fix wasn't actually a fix?

Hugo Vance: That's — yes, precisely that. Smoking was measurable. It was culturally salient, enormous in the data. So we could block that particular L. But residual confounding — the confounders we haven't measured, don't know to measure — those are still in the fork. Adjustment doesn't dissolve them. It addresses the ones you named.

Lila Soto: Which means — okay, I mean, this is the part that kind of unsettles me — the coffee reversal wasn't a triumph. It was a lucky case where the confounder happened to be something we could see.

Hugo Vance: That is exactly the caution I'd raise. And the beta-carotene findings make it concrete — observational studies through the eighties suggested supplementation reduced cancer risk. Large studies, careful adjustment. Randomized trials then showed the opposite: harm, in some populations. The confounder wasn't smoking that time. It was something in the profile of people who ate vegetables and lived differently. Unmeasured.

Lila Soto: So the question isn't just 'what else did we miss in coffee' — it's how many findings are walking around right now with an invisible fork in their diagram that nobody's named yet.

Hugo Vance: And that invisible fork is, I think, the thing the beta-carotene case makes undeniable. Because the confounding there wasn't a variable anyone forgot to measure. It was the entire life arrangement of the person eating vegetables. The person who ate enough beta-carotene to register in those cohort studies — that person walked to work, slept differently, had lower financial stress, possibly. Matching on smoking and diet doesn't reach all of that. It cannot.

Lila Soto: And the adjustment looked rigorous.

Hugo Vance: Completely rigorous by the standards available. Matching, stratification, propensity scores — you name it. And yet there's Margaret, a retiree in 1990, whose doctor read those cohort studies and said, the evidence suggests supplementation cuts cancer risk. Margaret starts taking beta-carotene daily. The randomized trial results arrive years later: harm in some populations. Not neutral. Harm. She was acting on a finding that was directionally wrong.

Lila Soto: Wait — the trial said harm, not just 'inconclusive'?

Hugo Vance: Harm. In specific populations. That's — yes, that's the word. Which is what makes it a landmark rather than a footnote.

Lila Soto: And the tools — matching, propensity scores — they could only reach what was actually measured going in. So if nobody logged 'how far does this person walk each day,' that's just... permanently outside the adjustment.

Hugo Vance: That is residual confounding, precisely defined. Whatever persists after adjustment because it was never measured — or mismeasured, or simply unknown — that's what's left, and it is invisible by construction. You cannot adjust for a variable that isn't in your dataset.

Lila Soto: Which means the more sophisticated the adjustment, kind of — the more it might look like the problem is solved when it isn't.

Hugo Vance: That's the uncomfortable implication, yes. We become very sophisticated at hiding our uncertainty without reducing it. The lifestyle clustering is what makes this structural rather than fixable — the careful person is careful across domains simultaneously, and those domains co-vary in ways no questionnaire fully maps.

Lila Soto: Hm. And it's almost a character variable — like, carefulness itself is the confounder, and you can't put carefulness in a regression.

Hugo Vance: Well — and genuinely, there are harder aspects to this. There are partial escapes from it. The randomized controlled trial, natural experiments, Mendelian randomization. Each one carries its own assumptions, and when those assumptions break, the same problem re-enters through a different door. That part, I think, deserves its own accounting.

Lila Soto: Yeah — because if even the escape routes have trapdoors, that changes what we can actually claim to know.

Hugo Vance: And the RCT is where that escape actually starts — because randomization does something adjustment never can. By the laws of chance, it distributes both measured and unmeasured confounders equally across groups. You don't have to name the confounder. You don't have to measure it. The fork is neutralized before it forms.

Lila Soto: Wait — so randomization doesn't identify the confounder, it just... makes it not matter?

Hugo Vance: By design, not by measurement. That is the structural difference. And it is why the RCT sits where it sits — not cultural prestige, structural superiority.

Lila Soto: Okay but — and this is the part that feels like relief and then immediately doesn't — you can't run an RCT on most of the things we actually care about. Like, decades of diet. Prenatal exposure. Lifelong behavior. You can't randomize someone's childhood nutrition.

Hugo Vance: Which is exactly where natural experiments enter. And the Dutch Hunger Winter of 1944 is the canonical case. A Nazi blockade cut off food supply to the Netherlands. It was not designed as a study — it was a famine. But researchers later used the accident of that history as a birth cohort natural experiment, tracing long-term health consequences of prenatal nutritional deprivation. One of the first major applications of the method.

Lila Soto: So history becomes the randomizer.

Hugo Vance: Quasi-randomizer. The blockade didn't assign women to famine by any logic related to their health. So the exposure is — not perfectly, but meaningfully — detached from the lifestyle clustering problem.

Lila Soto: Hm — and then Mendelian randomization is kind of the same move but pushed even earlier? Like, your alleles are assigned at conception, before any behavior begins, so the genetic variant mimics randomization without an actual experiment.

Hugo Vance: That's the instrumental variable logic, yes. The genetic variant affects the exposure but — and this is the load-bearing assumption — has no direct effect on the outcome except through the exposure. No other pathway.

Lila Soto: And that 'no other pathway' assumption — that's pleiotropy, right? I mean, that's the name for when it breaks?

Hugo Vance: Pleiotropy — when the genetic variant affects the outcome through pathways other than your exposure. And genes are promiscuous in that sense. A variant associated with, say, alcohol metabolism may also influence behavior through neurological routes that have nothing to do with the enzyme it encodes. The moment that happens, the instrument is broken.

Lila Soto: Oh — so when pleiotropy holds, you're back to the same causal inference problem you started with. The escape route loops around.

Hugo Vance: Through a different door. That's — yes, that's the honest accounting. The RCT is impractical for most long-term public health questions. The natural experiment depends on history producing the right accident. Mendelian randomization requires no pleiotropy and no population stratification — and when those assumptions fail, which is not rare, the method reproduces the very problem it was built to solve. The epistemological gap doesn't close. It relocates.

Lila Soto: So the partial escapes are real — but they're not exits. They're better-lit tunnels with their own walls at the end. And for most of the questions that actually reach the public — long-term nutrition, behavior, environment — we're permanently in that tunnel.

Hugo Vance: And that's — I think that's where I keep sitting. If the tunnel has no exit, then almost every piece of public health guidance built on observational associations carries irreducible uncertainty. And we almost never say that out loud. We say 'coffee is associated with lower heart disease rates in populations where smoking has been controlled for' — in the methods section. The headline reads like a verdict.

Lila Soto: The word 'association' is doing so much work — or it should be. When we drop it, we're making an epistemological claim the data literally cannot support.

Hugo Vance: Yes. And the other position — the one I take seriously — is that causation is a matter of carefully stated and defensible assumptions, not design type alone. Some methodologists would say: articulate your causal claim, name your required assumptions, show how the study meets them. That's a real position. I'm cautious about it, though. Defended assumptions are still assumptions. What would convince me is a demonstration that unmeasured confounding was reliably absent — not rare. Absent. I don't think we have that case.

Lila Soto: But then the alternative is silence — right? Waiting for RCT-level proof on diet, on exercise, on a lifetime of behavior. That's decades of nothing reaching the public.

Hugo Vance: And that silence is itself a decision with consequences. You don't get to be neutral by waiting. The delay is a public health choice — and probably a harmful one. I don't think there's a clean exit from that.

Lila Soto: So the question I can't shake — and maybe it's just where this lands for me — is whether epidemiology has actually changed how it communicates findings, or whether it just quietly apologizes in the methods section while the headline reads like a verdict.

Hugo Vance: I think, honestly, the methods section. That's the uncomfortable answer. The word 'association' does real epistemic work — and it keeps getting buried.

Lila Soto: That's an uneasy place to stop. But I think it's the right one. Thank you for sitting in it with me.