Quantitative Methods
Correlation, Confounding, and the Causal Question Underneath
Learn to state the causal question hiding inside a correlation, test it cheaply by stratifying, and name the assumption every method needs before you spend a budget on a marker.
- Intermediate
- 12 min total
- 14 chapters
What decision this helps you make: Whether to fund a programme aimed at a correlate, and what evidence you require before treating a relationship in your data as something you can act on.
- Related calculator: A/B Test Sample Size Calculator
What this topic is
A correlation is a fact about two columns of data. A causal claim is a statement about what would happen if you intervened: if you set the value of one thing, holding the rest of the world as it is. The two are different objects. Most business decisions are causal claims, and most business evidence is correlational, and the gap between them is where budgets go to die.
Why it matters
Almost every growth initiative is built on a correlation. Customers who use the app churn less, so push app installs. Accounts with a success manager renew more, so hire success managers. People who read three emails convert, so send more emails. In each case the correlation is real and the intervention may do nothing, because the behaviour was a marker of a customer who was already going to stay rather than a cause of their staying. Stratifying the data before committing budget takes an afternoon and routinely reverses the decision.
Who should learn it
Anyone proposing or approving a programme justified by a relationship in the data: growth leads, customer success, marketing, product, and finance teams evaluating the business case.
What you will understand
- The four explanations for any correlation, and how to tell which one you are looking at
- How to stratify a relationship and read how much of it was composition
- The exact assumption each causal method rests on, and what happens when it fails
- Why some methods answer only a local question, and when a local answer is enough
Prerequisites
Common misconception
"We controlled for the obvious factors, so what is left must be causal." Controlling handles the confounders you measured. It does nothing about the ones you did not, and there is no diagnostic in your data that tells you whether an unmeasured one exists. The residual looks identical either way. Worse, controlling for the wrong variable actively creates bias: adjust for something that sits on the causal path and you erase the real effect, and adjust for a variable that both the cause and the outcome influence and you manufacture a relationship that was never there.