P-Values: What They Mean and What They Don't
Intro statistics / AP Statistics · 12 flashcards · 7 quiz questions · updated 2026-08-19
A p-value answers one narrow question: if the null hypothesis were true, how often would we see data at least this extreme? It says nothing about the probability that the null is true, and nothing about how large or important an effect is.
Almost every mark lost on this topic is a misreading rather than a calculation, so the wording of your conclusion matters as much as the arithmetic.
The definition, said carefully
p = P(data this extreme or more extreme | H₀ true).
Read the vertical bar out loud as 'given that'. The condition is the null hypothesis being true — the p-value is computed inside a world where it holds.
What a p-value is not
- •Not the probability that H₀ is true. That would require reversing the conditional, which needs Bayes' theorem and a prior.
- •Not the probability that your result happened by chance.
- •Not a measure of effect size — a trivial difference reaches p < 0.001 with a large enough sample.
- •Not evidence that H₀ is true when p is large. Failing to reject is not accepting.
Stating the hypotheses
- •H₀ always contains equality: μ = 100, or p₁ = p₂.
- •H₁ (the alternative) carries the claim being tested: μ > 100 (one-tailed) or μ ≠ 100 (two-tailed).
- •Choose one- or two-tailed before seeing the data. Switching afterwards doubles your effective Type I error rate.
- •Hypotheses are about population parameters, never about sample statistics — write μ, not x̄.
The decision and the two errors
| H₀ actually true | H₀ actually false | |
|---|---|---|
| Reject H₀ | Type I error (probability α) | Correct — power (1 − β) |
| Fail to reject H₀ | Correct | Type II error (probability β) |
- •α is chosen before the test — 0.05 by convention, not by law.
- •Lowering α makes Type I errors rarer and Type II errors more common.
- •Power rises with a larger sample, a larger true effect, less variability, or a higher α.
Writing a conclusion that earns full marks
A complete conclusion has four parts: the decision, the significance level, the context, and the parameter.
'Since p = 0.03 < α = 0.05, we reject H₀. There is convincing evidence that the mean recovery time for the new treatment is less than 7 days.'
And when it goes the other way: 'Since p = 0.21 > α = 0.05, we fail to reject H₀. There is not convincing evidence that the mean differs from 7 days.' Never 'we accept H₀', and never 'the treatment has no effect'.
Confidence intervals say more
A 95% confidence interval and a two-tailed test at α = 0.05 agree: if the null value sits outside the interval, the test rejects. The interval adds what the p-value withholds — the size and direction of the effect, with its uncertainty.
Report both where you can. 'p = 0.03' tells a reader the result is unlikely under H₀; '2.1 days faster, 95% CI [0.3, 3.9]' tells them whether it matters.
Common mistakes
- ✗Writing 'the probability the null is true is 3%'. It is the probability of the data given the null.
- ✗Saying 'we accept the null hypothesis' rather than 'we fail to reject'.
- ✗Treating p = 0.049 and p = 0.051 as categorically different results.
- ✗Reading a small p-value as a large effect.
- ✗Choosing a one-tailed test after seeing which direction the data went.
Flashcards
Tap a card to reveal the answer.
Define the p-value.⌄
The probability of observing data at least as extreme as the sample, assuming H₀ is true.
Does p = 0.04 mean H₀ has a 4% chance of being true?⌄
No. It conditions on H₀ being true; it does not give the probability of H₀.
What is a Type I error?⌄
Rejecting a true null hypothesis. Its probability is α.
What is a Type II error?⌄
Failing to reject a false null hypothesis. Its probability is β.
Define power.⌄
1 − β: the probability of correctly rejecting a false null.
Name three ways to increase power.⌄
Increase the sample size, increase α, or reduce variability (a larger true effect also helps).
Which hypothesis contains the equality?⌄
The null hypothesis, always.
How do a 95% CI and a two-tailed α = 0.05 test relate?⌄
They agree: rejecting corresponds to the null value falling outside the interval.
What does a large p-value tell you?⌄
The data are consistent with H₀ — not that H₀ is true.
Why does a huge sample produce tiny p-values?⌄
Standard error shrinks with n, so even trivial differences become statistically detectable.
Correct phrasing after rejecting?⌄
'There is convincing evidence that…' followed by the alternative in context.
What does α represent?⌄
The threshold for the p-value, and the long-run rate of Type I errors when H₀ is true.
Practice quiz
Answer first, then open the explanation.
1. A p-value of 0.02 means:
- A. There is a 2% chance the null hypothesis is true
- B. There is a 2% chance the result occurred by chance
- C. If the null were true, data this extreme would occur about 2% of the time
- D. The effect is small
Show answer
C. If the null were true, data this extreme would occur about 2% of the time
The p-value is conditional on the null being true; it is not a probability about the hypothesis.2. Rejecting a true null hypothesis is:
- A. A Type I error
- B. A Type II error
- C. Correct
- D. Impossible
Show answer
A. A Type I error
That is the definition of a Type I error, with probability α.3. Which change increases the power of a test?
- A. A smaller sample
- B. A smaller α
- C. A larger sample
- D. More variability
Show answer
C. A larger sample
More data reduces standard error, making a real effect easier to detect.4. With p = 0.30 at α = 0.05, the correct conclusion is:
- A. Accept H₀ — there is no effect
- B. Fail to reject H₀ — there is not convincing evidence of an effect
- C. Reject H₀
- D. The test is invalid
Show answer
B. Fail to reject H₀ — there is not convincing evidence of an effect
Failing to reject is a statement about insufficient evidence, not proof of no effect.5. A 95% confidence interval for a mean difference is [1.2, 4.6]. A two-tailed test of 'no difference' at α = 0.05 would:
- A. Fail to reject
- B. Reject
- C. Be inconclusive
- D. Require a larger sample
Show answer
B. Reject
Zero is outside the interval, so the test rejects at the matching significance level.6. Hypotheses should be stated in terms of:
- A. Sample statistics like x̄
- B. Population parameters like μ
- C. The p-value
- D. The confidence level
Show answer
B. Population parameters like μ
Tests make inferences about populations; the sample is the evidence, not the claim.7. Two studies report p = 0.049 and p = 0.051. The right reading is:
- A. The first found an effect, the second did not
- B. They provide almost identical evidence
- C. The second is wrong
- D. Neither is interpretable
Show answer
B. They provide almost identical evidence
The α = 0.05 threshold is a convention, not a cliff in the underlying evidence.
Turn your own lecture into this
Record the class, upload the slides, or paste a YouTube link — Almanac writes the notes, the flashcards and the quiz from your own course material, in 100+ languages.