What are Type I and Type II errors in hypothesis testing?

Every time you run a statistical test, you make a decision: reject the null hypothesis or do not reject it. Because this decision is based on sample data rather than the whole population, it can be wrong. There are two ways to be wrong, and they are known as Type I and Type II errors.

Understanding these errors helps you choose an appropriate significance level, plan your sample size, interpret non-significant results sensibly and explain the limitations of your study. They are also a favourite topic in viva examinations. This post explains both errors in plain language, how they relate to significance and statistical power, and how to reduce them in your research.

The four possible outcomes of a hypothesis test

In reality, the null hypothesis (H₀) is either true or false. Your test either rejects it or does not. This creates four possible outcomes:

H₀ is actually trueH₀ is actually false
Reject H₀Type I error (false positive)
Probability = α
Correct decision (true positive)
Probability = 1 − β (power)
Do not reject H₀Correct decision (true negative)
Probability = 1 − α
Type II error (false negative)
Probability = β

Type I error: a false positive

A Type I error occurs when you reject a true null hypothesis. In other words, you conclude there is an effect or difference when, in reality, there is none.

Example: A study concludes that a new teaching app improves exam performance, but in reality the app has no effect. The observed difference occurred because of random sampling variation.

The probability of a Type I error, when the null hypothesis is true, is set by your significance level, alpha (α). If α = .05, you accept a 5% risk of a false positive for each test where the null is true.

Consequences

Type I errors can lead to false claims in the literature, wasted resources on ineffective interventions, and failed replication attempts. In medicine, a Type I error might mean approving an ineffective treatment.

Type II error: a false negative

A Type II error occurs when you fail to reject a false null hypothesis. You conclude there is no effect when, in reality, there is one.

Example: A study concludes that a stress-reduction programme has no effect, but in reality it does reduce stress. The study failed to detect it, perhaps because the sample was too small.

The probability of a Type II error is called beta (β). It depends on the true effect size, the sample size, the variability of the data and the alpha level.

Consequences

Type II errors can cause useful interventions to be abandoned, important relationships to be overlooked, and research directions to be dropped prematurely.

A simple analogy: the courtroom

A criminal trial offers a helpful analogy. The null hypothesis is “the defendant is innocent”.

  • Type I error: convicting an innocent person.
  • Type II error: acquitting a guilty person.

Legal systems usually treat convicting an innocent person as the more serious error, so they require strong evidence (“beyond reasonable doubt”). This is similar to setting a low alpha level. But making it harder to convict also increases the chance of acquitting guilty people. The same trade-off exists in statistics.

Another common analogy is a medical test: a false positive tells a healthy person they have a disease (Type I), while a false negative tells a sick person they are healthy (Type II).

The trade-off between Type I and Type II errors

For a fixed sample size, reducing the risk of one error increases the risk of the other.

  • Lowering alpha from .05 to .01 reduces Type I errors but makes it harder to detect real effects, increasing Type II errors.
  • Raising alpha makes it easier to detect effects but increases false positives.

The main way to reduce both errors at the same time is to increase the sample size or reduce measurement error, which increases the precision of your estimates.

Statistical power

Power is the probability of correctly rejecting a false null hypothesis, that is, detecting an effect that really exists. Power equals 1 − β.

A widely used convention, associated with Jacob Cohen, is to aim for power of at least .80, meaning an 80% chance of detecting a true effect of a specified size, and therefore a 20% risk of a Type II error. Some fields aim for .90 or higher.

What affects power?

  • Effect size: larger true effects are easier to detect.
  • Sample size: larger samples increase power.
  • Alpha level: higher alpha increases power (but also Type I error risk).
  • Variability: less noise in the data increases power. Reliable measures help.
  • Study design: within-subjects designs and well-chosen covariates can increase power.
  • One- vs two-tailed tests: one-tailed tests have more power for effects in the predicted direction, but should only be used with strong justification.

Power analysis: planning your sample size

A power analysis, done before data collection, estimates the sample size needed to detect an expected effect size with a chosen alpha and power. Tools such as G*Power (free software) and functions in R make this straightforward.

For example, to detect a medium effect (d = 0.5) in an independent-samples t-test with α = .05 (two-tailed) and power = .80, you need roughly 64 participants per group. To detect a small effect (d = 0.2) with the same settings, you would need roughly 394 per group. For a large effect (d = 0.8), about 26 per group would be enough.

These numbers show why many small studies are underpowered for detecting small or moderate effects.

Where do effect size estimates come from?

  • Previous studies or meta-analyses on similar variables.
  • Pilot studies (with caution, because small pilot samples give imprecise estimates).
  • The smallest effect size that would be practically meaningful in your context.

Running a power analysis in G*Power: a quick example

Suppose you plan a survey study using multiple regression with five predictors, and previous research suggests a medium effect for the overall model. A typical G*Power procedure would be:

  1. Choose Test family: F tests.
  2. Choose Statistical test: Linear multiple regression: Fixed model, R² deviation from zero.
  3. Choose Type of power analysis: A priori.
  4. Enter the expected effect size, for example f² = 0.15 (Cohen's medium benchmark), α = .05, power = .80 and number of predictors = 5.
  5. Click Calculate. G*Power reports the required total sample size, which for these settings is a little under 100 participants.

If you expect a small effect (f² = 0.02), the required sample size increases dramatically, to several hundred participants. This is why justifying your expected effect size carefully matters so much.

In your thesis, you might write: “An a priori power analysis using G*Power 3.1 indicated that a minimum sample of 92 participants was required to detect a medium effect (f² = 0.15) with five predictors, α = .05 and power = .80. The final sample of 214 therefore provided adequate power for the main analysis.” (Check the exact figure in the software for your own settings.)

Why underpowered studies are a problem for everyone

Low power does not only increase the risk of missing real effects. It also affects how significant results should be interpreted. When studies are underpowered, the significant results that do appear are more likely to overestimate the true effect size, because only unusually large sample estimates cross the significance threshold. In a literature full of small studies, this can make effects seem larger and more consistent than they really are, contributing to problems with replication. Planning for adequate power is therefore part of responsible research practice, not just a technical requirement.

Multiple comparisons and inflated Type I error

Each test carries its own Type I error risk. When you run many tests, the chance of at least one false positive increases. With 10 independent tests at α = .05 where all null hypotheses are true, the probability of at least one false positive is about 40%.

Ways to manage this include:

  • Limiting tests to those that address your research questions.
  • Using corrections such as Bonferroni, Holm or false discovery rate procedures where appropriate.
  • Using planned comparisons rather than testing everything.
  • Clearly labelling exploratory analyses as exploratory.

Interpreting non-significant results

If your result is not statistically significant, remember that a Type II error is possible, especially in small samples. Instead of concluding “there is no effect”, consider:

  • What was the effect size and its confidence interval?
  • Was the study adequately powered to detect a meaningful effect?
  • Could measurement error have reduced power?

A useful sentence: “The difference was not statistically significant; however, given the modest sample size and the confidence interval ranging from a small negative to a moderate positive effect, a meaningful effect cannot be ruled out.”

Reporting in your thesis

  • State your alpha level and whether tests were one- or two-tailed.
  • Report your a priori power analysis in the methodology chapter, including the expected effect size, alpha, desired power and resulting sample size.
  • Describe any corrections for multiple comparisons.
  • Discuss the risk of Type I and Type II errors in your limitations section, especially if your sample was smaller than planned.

Common mistakes to avoid

  • Treating non-significant results as proof of no effect.
  • Running many tests without considering inflated Type I error.
  • Skipping power analysis, then collecting a sample too small to detect expected effects.
  • Calculating “observed power” after the study from the obtained effect size, which is widely criticised because it adds no information beyond the p-value.
  • Confusing alpha (Type I error rate) with beta (Type II error rate).

Final thoughts

Type I errors are false positives; Type II errors are false negatives. Both are part of any decision based on sample data, and they trade off against each other. You manage them by choosing a sensible alpha, planning sample size with power analysis, using reliable measures, limiting unnecessary tests and interpreting results cautiously. Understanding these errors will help you design stronger studies and discuss your findings with the nuance examiners expect.

Have you run a power analysis for your study? Share your questions in the comments below.