P-Values: What They Mean and What They Don't

Intro statistics / AP Statistics · 12 flashcards · 7 quiz questions · updated 2026-08-19

A p-value answers one narrow question: if the null hypothesis were true, how often would we see data at least this extreme? It says nothing about the probability that the null is true, and nothing about how large or important an effect is.

Almost every mark lost on this topic is a misreading rather than a calculation, so the wording of your conclusion matters as much as the arithmetic.

The definition, said carefully

p = P(data this extreme or more extreme | H₀ true).

Read the vertical bar out loud as 'given that'. The condition is the null hypothesis being true — the p-value is computed inside a world where it holds.

What a p-value is not

Stating the hypotheses

The decision and the two errors

H₀ actually trueH₀ actually false
Reject H₀Type I error (probability α)Correct — power (1 − β)
Fail to reject H₀CorrectType II error (probability β)

Writing a conclusion that earns full marks

A complete conclusion has four parts: the decision, the significance level, the context, and the parameter.

'Since p = 0.03 < α = 0.05, we reject H₀. There is convincing evidence that the mean recovery time for the new treatment is less than 7 days.'

And when it goes the other way: 'Since p = 0.21 > α = 0.05, we fail to reject H₀. There is not convincing evidence that the mean differs from 7 days.' Never 'we accept H₀', and never 'the treatment has no effect'.

Confidence intervals say more

A 95% confidence interval and a two-tailed test at α = 0.05 agree: if the null value sits outside the interval, the test rejects. The interval adds what the p-value withholds — the size and direction of the effect, with its uncertainty.

Report both where you can. 'p = 0.03' tells a reader the result is unlikely under H₀; '2.1 days faster, 95% CI [0.3, 3.9]' tells them whether it matters.

Common mistakes

Flashcards

Tap a card to reveal the answer.

Practice quiz

Answer first, then open the explanation.

  1. 1. A p-value of 0.02 means:

    • A. There is a 2% chance the null hypothesis is true
    • B. There is a 2% chance the result occurred by chance
    • C. If the null were true, data this extreme would occur about 2% of the time
    • D. The effect is small
    Show answer

    C. If the null were true, data this extreme would occur about 2% of the time
    The p-value is conditional on the null being true; it is not a probability about the hypothesis.

  2. 2. Rejecting a true null hypothesis is:

    • A. A Type I error
    • B. A Type II error
    • C. Correct
    • D. Impossible
    Show answer

    A. A Type I error
    That is the definition of a Type I error, with probability α.

  3. 3. Which change increases the power of a test?

    • A. A smaller sample
    • B. A smaller α
    • C. A larger sample
    • D. More variability
    Show answer

    C. A larger sample
    More data reduces standard error, making a real effect easier to detect.

  4. 4. With p = 0.30 at α = 0.05, the correct conclusion is:

    • A. Accept H₀ — there is no effect
    • B. Fail to reject H₀ — there is not convincing evidence of an effect
    • C. Reject H₀
    • D. The test is invalid
    Show answer

    B. Fail to reject H₀ — there is not convincing evidence of an effect
    Failing to reject is a statement about insufficient evidence, not proof of no effect.

  5. 5. A 95% confidence interval for a mean difference is [1.2, 4.6]. A two-tailed test of 'no difference' at α = 0.05 would:

    • A. Fail to reject
    • B. Reject
    • C. Be inconclusive
    • D. Require a larger sample
    Show answer

    B. Reject
    Zero is outside the interval, so the test rejects at the matching significance level.

  6. 6. Hypotheses should be stated in terms of:

    • A. Sample statistics like x̄
    • B. Population parameters like μ
    • C. The p-value
    • D. The confidence level
    Show answer

    B. Population parameters like μ
    Tests make inferences about populations; the sample is the evidence, not the claim.

  7. 7. Two studies report p = 0.049 and p = 0.051. The right reading is:

    • A. The first found an effect, the second did not
    • B. They provide almost identical evidence
    • C. The second is wrong
    • D. Neither is interpretable
    Show answer

    B. They provide almost identical evidence
    The α = 0.05 threshold is a convention, not a cliff in the underlying evidence.

Turn your own lecture into this

Record the class, upload the slides, or paste a YouTube link — Almanac writes the notes, the flashcards and the quiz from your own course material, in 100+ languages.

Keep going