Why should I report effect sizes along with p-values?

Imagine two studies. In Study A, a new tutoring programme improves exam scores, p < .001. In Study B, a different programme also improves scores, p = .02. Which programme is more effective?

Based on p-values alone, you cannot tell. Study A may simply have had thousands of participants, making even a tiny improvement statistically significant. Study B may have had a much larger improvement but a small sample. To compare them meaningfully, you need to know how big the effect was. That is what an effect size tells you.

Reporting effect sizes is now expected by many journals, professional bodies and examiners. The APA Publication Manual, for example, recommends reporting effect sizes and confidence intervals. This post explains what effect sizes are, why they matter, which ones to use for common analyses, how to interpret them and how to report them.

What is an effect size?

An effect size is a quantitative measure of the magnitude of a phenomenon: the size of a difference between groups, the strength of a relationship between variables, or the amount of variance explained by a model.

Effect sizes can be:

  • Unstandardised: expressed in the original units, such as “a 5-point increase on a 100-point test” or “a 2.3 kg reduction in weight”. These are often the most meaningful when units are familiar.
  • Standardised: expressed in unit-free terms, such as Cohen's d or a correlation coefficient. These allow comparison across studies using different measures.

Why p-values are not enough

1. p-values depend heavily on sample size

With a very large sample, trivial differences become statistically significant. With a small sample, important differences may not reach significance. The p-value mixes effect size and sample size into a single number, so it cannot tell you about magnitude on its own.

2. Significance does not equal importance

A statistically significant result may be too small to matter in practice. Decision-makers, such as teachers, clinicians, managers and policymakers, need to know whether an effect is big enough to justify action or cost.

3. Non-significant results can still be informative

If a study finds a moderate effect size that is not statistically significant because of a small sample, the effect size shows that the finding may be worth investigating further. Reporting only “not significant” hides this.

4. Effect sizes allow comparison and synthesis

Meta-analyses combine effect sizes across studies. If you report effect sizes, your study can contribute to future syntheses. Effect sizes also allow you to compare your findings with previous research.

5. Effect sizes support power analysis

To plan sample size for future studies, researchers need estimates of expected effect sizes. Your reported effect sizes help others design better studies.

Common effect sizes and when to use them

Comparing two means: Cohen's d (and Hedges' g)

Cohen's d expresses the difference between two means in standard deviation units:

d = (Mean₁ − Mean₂) ÷ pooled standard deviation

For example, d = 0.50 means the two group means differ by half a standard deviation. Hedges' g is a similar measure with a correction for small-sample bias, often preferred when samples are small.

Cohen's widely cited benchmarks for d are approximately 0.2 (small), 0.5 (medium) and 0.8 (large).

Comparing more than two means (ANOVA): eta squared and partial eta squared

  • Eta squared (η²): the proportion of total variance in the outcome explained by a factor.
  • Partial eta squared (η²p): the proportion of variance explained by a factor, excluding variance explained by other factors in the model. SPSS reports this by default in its general linear model procedures.
  • Omega squared (ω²): a less biased alternative, particularly useful in small samples.

Commonly cited benchmarks for η² are approximately .01 (small), .06 (medium) and .14 (large). Be careful: partial eta squared values are not directly comparable across different designs.

Correlation: r

The correlation coefficient itself is an effect size. Cohen's benchmarks are approximately .10 (small), .30 (medium) and .50 (large). Squaring r gives the proportion of shared variance.

Regression: R², adjusted R², f² and standardised coefficients

  • R²: proportion of variance in the outcome explained by all predictors.
  • Adjusted R²: R² adjusted for the number of predictors, useful when comparing models.
  • ΔR²: the change in R² when adding predictors in hierarchical regression.
  • Cohen's f²: effect size for a set of predictors, with benchmarks of approximately .02 (small), .15 (medium) and .35 (large).
  • Standardised coefficients (β): indicate the relative strength of individual predictors.

Categorical data: phi, Cramér's V and odds ratios

  • Phi (φ): for 2 × 2 tables.
  • Cramér's V: for larger contingency tables. Interpretation depends on the number of categories.
  • Odds ratio (OR): the odds of an outcome in one group compared with another. An OR of 1 means no difference; values further from 1 indicate larger effects. Widely used in health research and logistic regression.
  • Risk ratio and risk difference: often more intuitive for readers in health contexts.

Non-parametric tests

For tests such as Mann–Whitney U or Wilcoxon signed-rank, effect sizes such as r (calculated from the z statistic) or rank-biserial correlation are commonly reported.

Interpreting effect sizes: beyond small, medium and large

Cohen's benchmarks are useful as a starting point, but Cohen himself cautioned that they were rough conventions for the behavioural sciences, to be used only when better information is not available. A “small” effect can be very important if:

  • The outcome is serious, such as mortality or safety incidents.
  • The intervention is cheap and easy to implement at scale.
  • The effect accumulates over time.

Conversely, a “large” effect in a small, unrepresentative sample may be unstable and exaggerated.

The best approach is to interpret effect sizes in context:

  • Compare with effect sizes from previous studies in your field.
  • Consider practical meaning in real units, such as minutes, marks, dollars or percentage points.
  • Consider the cost, feasibility and consequences of the effect.

Confidence intervals for effect sizes

An effect size from a sample is an estimate. A confidence interval (CI) shows the precision of that estimate. For example, d = 0.45, 95% CI [0.10, 0.80] tells you the data are compatible with effects ranging from small to large. A wide interval signals uncertainty, often due to a small sample.

Reporting CIs helps readers judge how much confidence to place in your estimate and avoids the false certainty of a single number.

A worked comparison: same conclusion, very different effects

Let us return to the two tutoring studies from the introduction and add some illustrative details.

  • Study A: 4,000 students in each group. The tutoring group scored 1 point higher on a 100-point exam (SD = 15). The difference is statistically significant, p ≈ .003, but Cohen's d is only about 0.07. In practical terms, one extra mark out of 100 is unlikely to change a student's grade.
  • Study B: 40 students in each group. The tutoring group scored 8 points higher (SD = 15). The difference is also statistically significant, p ≈ .02, and Cohen's d is about 0.53, a medium effect. Eight marks could move many students up a grade band.

Looking only at p-values, Study A appears to provide “stronger” evidence because its p-value is smaller. Looking at effect sizes, Study B's programme appears far more effective, although its estimate is less precise because of the smaller sample. Its confidence interval would be much wider than Study A's.

A decision-maker choosing between the two programmes would need both pieces of information: the size of the effect and the precision of the estimate. This is exactly why effect sizes and confidence intervals should always accompany p-values.

Effect sizes and sample size planning

Effect sizes also play a key role before data collection. To calculate the sample size needed for adequate statistical power, you need an expected effect size. Sources for this estimate include previous studies, meta-analyses and pilot data. A tool such as G*Power can then calculate the required sample size for your chosen test, alpha level and desired power. Reporting this power analysis in your methodology chapter shows that your sample size was planned rather than arbitrary.

How to report effect sizes

Include the effect size alongside the test statistic and p-value, with a confidence interval where possible. Examples (illustrative numbers):

  • t-test: “Participants in the mindfulness group reported lower stress (M = 2.8, SD = 0.9) than those in the control group (M = 3.3, SD = 1.0), t(98) = 2.62, p = .010, d = 0.52, 95% CI [0.12, 0.92].”
  • ANOVA: “There was a significant effect of teaching method on test scores, F(2, 147) = 6.84, p = .001, η² = .09.”
  • Correlation: “Workload was positively correlated with burnout, r(198) = .41, p < .001, 95% CI [.29, .52].”
  • Regression: “The model explained 27% of the variance in turnover intention, R² = .27, F(4, 245) = 22.6, p < .001.”
  • Chi-square: “Employment status was associated with training participation, χ²(2, N = 300) = 14.2, p < .001, Cramér's V = .22.”

Getting effect sizes from software

  • SPSS: reports partial eta squared in ANOVA (request “Estimates of effect size” in Options), R² in regression, and phi/Cramér's V in Crosstabs (Statistics button). Recent versions also report Cohen's d for t-tests.
  • R and Jamovi: packages and modules such as effectsize or the built-in options in Jamovi provide many effect sizes with confidence intervals.
  • Online calculators: can compute effect sizes from reported statistics, but check the formulas they use.

Common mistakes to avoid

  • Reporting only p-values.
  • Labelling effects as “small” or “large” without considering context.
  • Comparing partial eta squared values across different designs as if they were equivalent.
  • Reporting effect sizes without specifying which measure was used.
  • Ignoring effect sizes for non-significant results.
  • Omitting confidence intervals.

Final thoughts

p-values tell you whether an effect is unlikely under the null hypothesis. Effect sizes tell you how large the effect is. Confidence intervals tell you how precise that estimate is. Together, they give a complete and honest picture of your findings. Make it a habit to report effect sizes for every main analysis, interpret them in context, and discuss their practical significance. Your thesis will be stronger, and your findings more useful to others.

Not sure which effect size fits your analysis? Ask in the comments below.