What is the difference between correlation and causation?

“Correlation does not imply causation” is one of the most repeated phrases in research methods. Most students can quote it. Yet causal language still appears regularly in theses and papers based on correlational data: “X leads to Y”, “X improves Y”, “X is a key driver of Y”. Examiners and reviewers notice this immediately, and it is one of the most common reasons for requested revisions.

This post explains the difference between correlation and causation, why correlation alone cannot prove cause, what conditions are needed for causal claims, which research designs support stronger causal inference, and how to write about your findings accurately.

What is correlation?

A correlation is a statistical association between two variables: as one changes, the other tends to change in a consistent way.

  • Positive correlation: both variables tend to increase together. For example, hours of study and exam scores.
  • Negative correlation: as one increases, the other tends to decrease. For example, workload and job satisfaction.
  • No correlation: no consistent linear relationship.

Correlation is often measured with a coefficient such as Pearson's r or Spearman's rho, which ranges from −1 to +1. The closer the value is to −1 or +1, the stronger the association.

Importantly, correlation describes co-variation. It tells you that two variables move together, but not why.

What is causation?

Causation means that a change in one variable produces a change in another. If X causes Y, then changing X (while other factors stay the same) will lead to a change in Y.

Causal claims are powerful because they suggest what will happen if we intervene. If a training programme causes better performance, organisations can introduce it to improve performance. If the relationship is only correlational, introducing the programme might not have the expected effect.

Why correlation does not prove causation

There are several alternative explanations for a correlation between X and Y.

1. Reverse causation

Y might cause X, rather than X causing Y. For example, a correlation between job satisfaction and performance could mean satisfied employees perform better, or that high performers become more satisfied because they receive recognition.

2. Confounding (third variables)

A third variable, Z, might cause both X and Y. A classic illustration: ice cream sales and drowning incidents are correlated, but hot weather increases both. Neither causes the other.

In research, confounders are often less obvious. For example, a correlation between using a learning app and higher grades might be explained by student motivation: motivated students both use the app more and study harder in general.

3. Coincidence or chance

With many variables and large datasets, some correlations will appear by chance alone. There are well-known humorous examples of unrelated time series that happen to move together. Statistical significance testing helps reduce this risk, but does not remove it, especially when many tests are run.

4. Selection effects

The way participants are selected can create or distort correlations. For example, if a study only includes people admitted to a selective programme, relationships among variables within that group may differ from those in the general population.

5. Measurement artefacts

If both variables are measured using the same method, such as self-report on the same survey, some of the correlation may reflect shared method variance rather than a true relationship.

What is needed to support a causal claim?

Philosophers and methodologists have discussed causation for centuries. In research methods, three conditions are commonly used as a basic framework:

  1. Covariation: X and Y must be associated. Correlation provides evidence for this condition only.
  2. Temporal precedence: the cause must come before the effect in time.
  3. Elimination of alternative explanations: other plausible causes, such as confounders or reverse causation, must be ruled out or controlled.

In epidemiology, the Bradford Hill considerations (1965) are often cited when evaluating causal relationships from observational data. They include strength of association, consistency across studies, specificity, temporality, biological gradient (dose–response), plausibility, coherence, experimental evidence and analogy. They are considerations to guide judgement, not a checklist that proves causation.

Research designs and causal inference

Different designs provide different levels of support for causal claims.

Randomised controlled experiments

In a true experiment, participants are randomly assigned to conditions, and the researcher manipulates the independent variable. Random assignment balances known and unknown confounders across groups on average. This makes experiments the strongest design for causal inference, although they have their own limitations, such as artificial settings or ethical constraints.

Quasi-experimental designs

When random assignment is not possible, quasi-experimental designs, such as non-equivalent control groups, interrupted time series, difference-in-differences or regression discontinuity, can strengthen causal inference if carefully designed and analysed.

Longitudinal designs

Measuring variables at multiple time points helps establish temporal precedence and allows analysis of change. Cross-lagged panel models, for example, can examine whether X at time 1 predicts Y at time 2 more strongly than Y at time 1 predicts X at time 2. They still cannot eliminate all confounding.

Cross-sectional correlational designs

Data collected at one time point, such as a typical survey, can show associations and test whether data are consistent with a theoretical model. However, they provide limited evidence for causation because they cannot establish temporal order and can only control for measured confounders.

Statistical control

Regression and structural equation modelling allow you to control for measured variables. This can reduce, but not eliminate, confounding. Unmeasured confounders can still bias results. Advanced approaches such as instrumental variables and propensity score methods aim to strengthen causal inference in observational data, but they rely on assumptions that must be justified.

Using diagrams to think about causation

A simple but powerful way to think through causal questions is to draw a diagram of the variables and the possible arrows between them. In epidemiology and increasingly in the social sciences, these are formalised as directed acyclic graphs (DAGs), but even an informal sketch helps.

For example, suppose you are studying whether remote work (X) is related to employee wellbeing (Y). Draw X and Y, then ask:

  • What factors might affect both remote work and wellbeing? Job type, seniority, caring responsibilities and commuting distance are possibilities. These are potential confounders that should be measured and controlled if possible.
  • What factors might lie on the path from remote work to wellbeing? Work–life balance or reduced commuting stress might be mediators. Controlling for them would remove part of the effect you want to study.
  • Could wellbeing influence remote work? Employees with poorer wellbeing might request remote work, creating reverse causation.
  • Are there variables that are caused by both X and Y? Controlling for these (sometimes called colliders) can create misleading associations.

This exercise helps you choose control variables for sound reasons rather than simply adding everything available. It also gives you a clear way to explain your analytic choices, and their limits, in your methodology and discussion chapters.

Questions to ask before making a causal claim

  • Did I manipulate the independent variable, or only measure it?
  • Were participants randomly assigned to conditions?
  • Do I know that the presumed cause came before the effect?
  • Have I measured and controlled the most plausible confounders?
  • Could the relationship run in the opposite direction?
  • Is the finding consistent with previous studies using different designs?
  • Does my language in the abstract, results and discussion match the strength of my evidence?

Writing about your findings accurately

If your design is correlational, use language that matches your evidence.

Avoid (causal language):

  • “X causes Y.”
  • “X leads to Y.”
  • “X increases/improves/reduces Y.”
  • “The effect of X on Y.” (This is common in regression language but can imply causation. Use with care and clarify.)
  • “X is a driver of Y.”

Prefer (associational language):

  • “X was positively associated with Y.”
  • “X was related to Y.”
  • “Higher levels of X were linked to higher levels of Y.”
  • “X predicted Y.” (Statistically, prediction does not imply causation, but explain this to readers.)
  • “The findings are consistent with the proposed model in which X influences Y.”

Discuss limitations explicitly

In your limitations section, acknowledge that the design cannot establish causal direction and describe plausible alternative explanations. Suggest designs that future research could use, such as longitudinal or experimental studies.

Implications and recommendations

Be careful when writing practical recommendations. Instead of “Organisations should increase X to improve Y”, consider “Organisations may consider initiatives that support X, although experimental or longitudinal research is needed to confirm whether changes in X lead to improvements in Y.”

An example

A cross-sectional survey of 400 teachers finds a positive correlation between professional development hours and teaching self-efficacy (r = .38, p < .001). (Numbers are illustrative.)

Overstated conclusion: “Professional development increases teaching self-efficacy. Schools should require more professional development hours.”

Accurate conclusion: “Teachers who reported more professional development hours also reported higher teaching self-efficacy. This association may reflect an effect of professional development on self-efficacy, but it could also mean that more confident teachers seek out more development opportunities, or that school support influences both. Longitudinal or experimental designs are needed to test causal direction.”

Common mistakes to avoid

  • Using causal verbs in the abstract, results or discussion of a correlational study.
  • Assuming that controlling for a few variables in regression proves causation.
  • Treating mediation analysis on cross-sectional data as proof of a causal mechanism.
  • Ignoring reverse causation in the discussion.
  • Making strong policy recommendations from correlational evidence alone.

Final thoughts

Correlation tells you that two variables move together. Causation tells you that one makes the other change. The gap between them is filled by research design, temporal order, and the careful elimination of alternative explanations. If your study is correlational, that is perfectly acceptable, but your language and conclusions must match your evidence. Precise, honest wording shows methodological maturity and protects your research from one of the most common criticisms.

Have you struggled with causal language in your thesis? Share an example sentence in the comments and we can help you rephrase it.