Correlation vs Causation
Correlation measures the strength and direction of a linear relationship between two variables. Causation means one variable directly causes a change in another. A correlation between two variables does not prove that one causes the other — there may be confounding variables or the relationship may be coincidental.
Key Takeaways
- Correlation (r) measures the strength and direction of a linear association between two quantitative variables. It ranges from -1 to +1.
- Causation means a change in one variable directly produces a change in another. Establishing causation requires a well-designed experiment with random assignment.
- Confounding variables (lurking variables) are unmeasured factors that affect both variables, creating a misleading association.
- Only randomized experiments can establish causation. Observational studies can only show association (correlation).
Why Correlation Doesn't Imply Causation
"Correlation does not imply causation" is one of the most important principles in statistics — and one of the most frequently tested on the AP exam. Understanding why requires knowing the difference between observational studies and experiments.
First, let's clarify correlation. The correlation coefficient (r) measures how closely two quantitative variables follow a linear pattern. If r is close to +1, there's a strong positive linear relationship (as one variable increases, so does the other). If r is close to -1, there's a strong negative linear relationship. If r is near 0, there's no linear relationship.
But here's the problem: just because two variables are correlated doesn't mean one causes the other. The classic example: ice cream sales and drowning deaths are positively correlated. Does ice cream cause drowning? No — a confounding variable (hot weather) drives both. People eat more ice cream when it's hot, and they also swim more when it's hot.
This is why study design matters. In an observational study, you simply observe and record data without manipulating anything. You might find that students who eat breakfast score higher on tests. But you can't conclude that breakfast causes better scores — maybe students who eat breakfast also come from wealthier families, get more sleep, or have more involved parents.
In a randomized experiment, you randomly assign subjects to treatment and control groups. Random assignment is the key — it distributes all confounding variables (both known and unknown) roughly equally across groups. Any difference in outcomes can then be attributed to the treatment. This is why randomized experiments are the gold standard for establishing causation.
There are three possible explanations for a correlation between X and Y: (1) X causes Y, (2) Y causes X (reverse causation), or (3) a third variable Z causes both X and Y (confounding). Only a well-designed experiment with random assignment can rule out explanations 2 and 3.
Correlation vs Causation on the AP Statistics Exam
The AP Statistics exam tests this concept in multiple contexts. You'll see it in questions about study design (observational vs experimental), regression interpretation, and inference.
The most common question format: you're given a description of a study and asked whether you can conclude causation. The answer depends entirely on study design. If subjects were randomly assigned to groups → you can make causal claims. If it was an observational study → you can only claim association, not causation.
In regression contexts, always use the language of association, not causation: "There is a positive association between study hours and test scores" rather than "Studying more hours causes higher test scores" (unless the data came from an experiment).
For FRQs, be prepared to: (1) identify confounding variables in a given scenario, (2) explain why an observational study cannot establish causation, and (3) describe how you would design an experiment with random assignment to test a causal claim.
Remember: r² (the coefficient of determination) tells you the percentage of variation in Y that is explained by the linear relationship with X. But "explained" doesn't mean "caused" — it's a statistical relationship, not necessarily a causal one.
Common Mistakes Students Make
- Using causal language for observational studies. Never say "X causes Y" or "X leads to Y" based on observational data. Use "X is associated with Y" or "there is a correlation between X and Y."
- Thinking a strong correlation (r close to ±1) proves causation. The strength of the correlation doesn't change the fundamental issue. Even a perfect correlation (r = 1) doesn't establish causation without experimental evidence.
- Forgetting to identify confounding variables. When the AP exam asks why you can't conclude causation, always name at least one specific confounding variable that could explain the observed relationship.
Related Topics
Frequently Asked Questions
A confounding variable (also called a lurking variable) is a variable that is associated with both the explanatory variable and the response variable, creating a misleading appearance of a direct relationship. For example, if studying the relationship between shoe size and reading ability in children, age is a confounding variable — older children have bigger feet AND read better.
The primary way to establish causation is through a well-designed randomized controlled experiment. Key elements: (1) random assignment of subjects to treatment and control groups, (2) a control group for comparison, (3) sufficient sample size, and (4) controlling for potential confounding variables. Observational studies, no matter how large, cannot establish causation.