What is a p-value, and what does "statistically significant" really mean?

Few numbers in research carry as much weight, or cause as much confusion, as the p-value. Students often feel relief when they see “p < .05” and disappointment when they see “p = .08”. Some treat significance as proof that a hypothesis is true; others treat non-significance as proof that there is no effect. Neither interpretation is correct.

Understanding what a p-value really means will help you interpret your results accurately, write about them responsibly, and answer examiner questions with confidence. This post explains the p-value in plain language, the logic of significance testing, common misinterpretations, and good practice for reporting.

The logic of hypothesis testing

Most statistical tests in thesis research follow the logic of null hypothesis significance testing (NHST).

  • Null hypothesis (H₀): usually a statement of “no effect” or “no difference”. For example, “There is no difference in mean exam scores between students taught with method A and method B.”
  • Alternative hypothesis (H₁): a statement that there is an effect or difference. For example, “Mean exam scores differ between the two methods.”

You collect data, calculate a test statistic (such as t, F or χ²), and then ask: if the null hypothesis were true, how surprising would data like mine be?

What is a p-value?

The p-value is the probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is true (and assuming the statistical model's other assumptions hold).

In simpler terms: if there were really no effect, how likely is it that you would see a difference or relationship as large as yours, or larger, just by random sampling variation?

  • A small p-value (for example, .01) means that data like yours would be quite unusual if the null hypothesis were true.
  • A large p-value (for example, .45) means that data like yours would be quite common even if the null hypothesis were true.

What does “statistically significant” mean?

Before analysis, researchers choose a significance level, called alpha (α), often .05. If the p-value is less than alpha, the result is called statistically significant, and the null hypothesis is rejected.

So “statistically significant at α = .05” means: if the null hypothesis were true, a result this extreme would occur less than 5% of the time. It is a statement about how compatible the data are with the null hypothesis.

Importantly, “significant” in statistics does not mean important, large or meaningful. It only indicates that the result is unlikely under the null hypothesis, given the assumptions.

Why .05?

The .05 threshold is a convention, popularised in the early twentieth century, largely through the influence of Ronald Fisher. It has no special scientific status. Some fields use stricter levels, such as .01 or .005, and particle physics famously uses much stricter thresholds for discovery claims. What matters is to state your alpha level in advance and apply it consistently.

An intuitive example: the suspicious coin

Imagine a friend claims a coin is fair. You decide to test this by flipping it 20 times. Your null hypothesis is that the coin is fair (the probability of heads is 0.5).

  • If you get 11 heads, that is very ordinary for a fair coin. The p-value would be large, and you would have no reason to doubt the coin.
  • If you get 16 heads, you might start to feel suspicious. For a fair coin, getting 16 or more heads, or 4 or fewer heads (a two-sided test), happens only about 1% of the time. The p-value is about .01, so the result is statistically significant at α = .05.

Notice what the p-value of about .01 does and does not mean. It means that if the coin were fair, a result this extreme would be rare. It does not mean there is a 1% chance the coin is fair. If you had strong prior reasons to believe the coin was fair (for example, it came from a bank and looks completely normal), you might reasonably want more flips before concluding it is biased. And even if the coin is biased, the p-value alone does not tell you how biased it is; for that, you would estimate the probability of heads and its confidence interval.

This simple example captures the core logic of every significance test you run in your thesis, whether it is a t-test, a chi-square test or a regression coefficient.

One-tailed and two-tailed tests

A two-tailed test looks for an effect in either direction; a one-tailed test looks for an effect in one pre-specified direction only. One-tailed tests give smaller p-values for effects in the predicted direction, which makes them tempting. However, they should only be used when there is a strong, pre-registered or clearly justified directional prediction, and when an effect in the opposite direction would be treated the same as no effect. Many journals and examiners expect two-tailed tests by default. Choosing a one-tailed test after seeing the data is not acceptable.

Common misinterpretations of p-values

In 2016, the American Statistical Association (ASA) issued a formal statement on p-values to address widespread misunderstanding. Several of its key points are reflected below.

Misinterpretation 1: “p is the probability that the null hypothesis is true.”

No. The p-value is calculated assuming the null hypothesis is true. It cannot tell you the probability that the null hypothesis is true or false. For example, p = .03 does not mean there is a 3% chance there is no effect.

Misinterpretation 2: “p is the probability that the result happened by chance.”

This is a common but misleading shortcut. The p-value is the probability of data this extreme if chance (the null model) were the only explanation. It is not the probability that chance produced your result.

Misinterpretation 3: “A significant result proves my hypothesis.”

No. Significance indicates the data are unlikely under the null hypothesis, but it does not prove the alternative hypothesis, the theory behind it, or a causal relationship. Other explanations, such as confounding, bias or chance in a single study, remain possible.

Misinterpretation 4: “A non-significant result proves there is no effect.”

No. Absence of evidence is not evidence of absence. A non-significant result may reflect a small sample, low statistical power, measurement error or a small effect. Look at the effect size and confidence interval to see what range of effects is consistent with your data.

Misinterpretation 5: “A smaller p-value means a larger or more important effect.”

Not necessarily. p-values depend on both effect size and sample size. With a very large sample, a tiny, practically trivial effect can have a very small p-value. With a small sample, a large and important effect might not reach significance.

Misinterpretation 6: “p = .049 and p = .051 are fundamentally different.”

These two results provide almost identical evidence. Treating .05 as a sharp dividing line between “real” and “not real” findings is misleading. This is why many researchers recommend reporting exact p-values and avoiding over-reliance on the threshold.

Statistical significance vs. practical significance

Suppose a study with 20,000 participants finds that a new teaching method improves test scores by 0.3 points on a 100-point scale, with p < .001. The result is statistically significant but probably of little practical importance.

Now suppose a small pilot study with 20 participants finds an improvement of 12 points, with p = .09. The result is not statistically significant, but the effect may be practically important and worth investigating further with a larger sample.

This is why effect sizes and confidence intervals are essential companions to p-values.

What influences the p-value?

  • Effect size: larger true effects tend to produce smaller p-values.
  • Sample size: larger samples make smaller effects detectable and produce smaller p-values for the same effect size.
  • Variability: more variability (noise) in the data tends to increase p-values.
  • Test choice and assumptions: violating test assumptions can make p-values inaccurate.

The problem of multiple testing

If you run many tests at α = .05, some will be significant just by chance. For example, with 20 independent tests where all null hypotheses are true, you would expect about one to be significant by chance alone. This increases the risk of false positives (Type I errors).

Ways to address this include:

  • Planning a limited number of theory-driven tests in advance.
  • Using corrections such as Bonferroni or Holm when appropriate.
  • Distinguishing confirmatory from exploratory analyses in your reporting.
  • Avoiding “p-hacking”: trying many analyses and reporting only the significant ones.

Good practice for reporting p-values

  • Report exact p-values (for example, p = .032) rather than only “p < .05”, except for very small values, which are reported as p < .001.
  • Never report “p = .000”. Software rounds very small values; report them as p < .001.
  • Report the test statistic and degrees of freedom, such as t(58) = 2.41, p = .019.
  • Report effect sizes and confidence intervals alongside p-values.
  • State your alpha level and whether tests were one- or two-tailed.
  • Report all planned tests, including non-significant ones.
  • Use careful language: “statistically significant” rather than just “significant”, to avoid confusion with everyday meaning.

Example

“Students taught with the flipped classroom approach scored higher on the final exam (M = 72.4, SD = 9.8) than those taught with traditional lectures (M = 67.1, SD = 10.3), t(118) = 2.89, p = .005, d = 0.53, 95% CI for the mean difference [1.67, 8.93].” (Numbers are illustrative.)

Writing about non-significant results

Instead of “There was no difference between groups”, write: “The difference between groups was not statistically significant, t(58) = 1.21, p = .231, d = 0.31, 95% CI [−1.2, 4.9]. The confidence interval indicates that the data are compatible with a range of effects from a small negative difference to a moderate positive difference, so the study cannot rule out a meaningful effect.”

Final thoughts

A p-value is a useful tool, but only when you understand what it does and does not tell you. It measures how compatible your data are with the null hypothesis under specific assumptions. It does not tell you whether your hypothesis is true, how large or important an effect is, or whether a finding will replicate. Report exact p-values alongside effect sizes and confidence intervals, avoid treating .05 as a magic line, and interpret your results in the context of your study design and existing evidence. That is what good statistical reasoning looks like.

Do you find p-values confusing? Share your questions in the comments below.