Learning center

A route through statistical reasoning

Do not memorize a gallery of charts. Work through a question, change the assumptions in its lab, and finish by explaining where the conclusion stops.

Start where the claim breaks

Each route contains three guide-and-lab pairs, a concrete practice task, and two answers you can check without creating an account.

4guided routes
12guide + lab pairs
16defined concepts
0sign-ups required

Choose a route

Four questions, twelve connected lessons

Routes are ordered, but they are not locked. Begin with the question closest to a claim you need to evaluate.

PATH 01

Describe before you summarize

What does one headline number leave out?

Begin with center, denominators, and random variation. This route is designed for readers who want a reliable vocabulary before moving into inference.

Outcome: Choose and describe a summary without hiding its denominator, distribution, or scale.

  1. 1.1

    Mean, median, and an extreme value

    Practice: Move one delivery time and explain which center changed, by how much, and why.

  2. 1.2

    Percentage points and relative change

    Practice: Translate one rate change into points, relative percent, and changed outcomes.

  3. 1.3

    Runs in a fair random sequence

    Practice: Separate a surprising-looking run from evidence that the next outcome changed.

Check your explanation

1. A rate rises from 4% to 6%. What are the two valid change statements?

It rose by 2 percentage points and by 50% relative to the 4% baseline. The first is an absolute difference; the second divides that difference by the starting rate.

2. Can a run of four heads make the next fair, independent flip more likely to be tails?

No. The run changes how the observed sequence looks, but independence keeps the next-flip probability at one half.

PATH 02

Samples, bias, and uncertainty

When does more data fail to mean better evidence?

Follow the path from who enters a sample, through repeated sample means, to the long-run behavior of confidence intervals.

Outcome: Distinguish selection quality, random error, standard error, and interval coverage.

  1. 2.1

    A large sample centered on the wrong target

    Practice: Hold response bias fixed while increasing sample size and describe what narrows and what does not.

  2. 2.2

    The sampling distribution of a mean

    Practice: Compare n = 1 with n = 30 without claiming that the source population became normal.

  3. 2.3

    Coverage across repeated intervals

    Practice: Find a correctly calculated interval that misses and explain why this does not contradict 95% coverage.

Check your explanation

1. Why can ten times as many voluntary responses make an estimate more precise but not more representative?

More responses reduce random variation around the voluntary-response mechanism. They do not change who was more likely to respond, so the estimate can become tightly concentrated around a biased target.

2. What does 95% describe in a conventional confidence procedure?

It describes the long-run share of intervals from repeated uses of the procedure that cover the fixed target, not a 95% probability assigned to a parameter after one interval has been calculated.

PATH 03

Compare like with like

Which hidden structure can reverse or weaken a comparison?

Use shape, subgroup composition, and repeated measurement to diagnose conclusions that a single total or coefficient cannot support.

Outcome: Audit comparisons for nonlinearity, unequal mixtures, and selection on extreme measurements.

  1. 3.1

    Four shapes behind one coefficient

    Practice: Match the same correlation summary to different point patterns and name what the coefficient misses.

  2. 3.2

    A subgroup comparison that reverses in total

    Practice: Reweight both options to the same group mix before interpreting the aggregate difference.

  3. 3.3

    Extreme selection and retest movement

    Practice: Change reliability and selection fraction, then predict the direction of the retest mean.

Check your explanation

1. Can a correlation near zero prove that two variables are unrelated?

No. Pearson correlation summarizes linear association. A strong U-shaped relationship can have a coefficient near zero, so the point pattern and measurement context remain essential.

2. Why can option B win in every subgroup but lose overall?

The overall rates use different subgroup weights. If B contains many more observations from the lower-rate subgroup, its weighted total can fall below A even while B has the higher rate inside each comparable subgroup.

PATH 04

Evaluate a statistical claim

What context must travel with a striking result?

Finish with three common publication failures: ignoring prevalence, cropping a visual scale, and searching many hypotheses without reporting the search.

Outcome: Ask for the prior rate, graphical baseline, effect size, denominator, and full analysis family.

  1. 4.1

    Rare events and false alerts

    Practice: Build the four-cell count table before interpreting an alert percentage.

  2. 4.2

    One dataset on two vertical scales

    Practice: Describe the numerical change separately from its pixel amplification.

  3. 4.3

    One small p-value among many tests

    Practice: Record the full family of tests and compare the unadjusted and family-wise thresholds.

Check your explanation

1. A detector has 99% sensitivity and specificity. Is an alert 99% likely to be correct?

Not necessarily. Alert precision also depends on prevalence. At a 1% base rate, 99 true alerts and 99 false alerts occur per 10,000 cases, so only half of alerts are true.

2. What should accompany a highlighted p-value from a broad search?

Report the number and definition of tests, the selection rule, effect sizes, uncertainty, and any multiplicity adjustment or preregistered primary outcome.

A repeatable study cycle

Use a lab as evidence, not decoration

Every experiment exposes the controls, seed, generated values, and a paired explanation. This four-step routine turns a visual change into a claim you can defend.

  1. 01Predict

    Write down what you expect before moving a control.

  2. 02Change one assumption

    Keep other inputs and the seed fixed where possible.

  3. 03Inspect the values

    Open the table or CSV rather than reading shape alone.

  4. 04State the boundary

    Name the model condition that could make your conclusion fail.

Concept index

Definitions with somewhere to go next

These are short working definitions. Each term links to a guide that derives or tests it in context.

Base rate
The frequency of an event before new evidence is considered. It is required when reversing a conditional probability. Open the guide
Bias
A systematic difference between the quantity an estimator targets on average and the quantity a study intends to describe. Open the guide
Confidence coverage
The long-run fraction of intervals from a procedure that contain the fixed target under its stated model. Open the guide
Correlation
A standardized summary of linear association. It does not by itself describe curvature, clusters, or causation. Open the guide
Denominator
The reference count beneath a rate. A percentage is not interpretable until its denominator and population are defined. Open the guide
Independence
A model condition in which learning one outcome does not change the probability distribution of another. Open the guide
Mean
The arithmetic total divided by the count. Every value influences it, including unusually large or small observations. Open the guide
Median
The middle ordered value, or midpoint of the two middle values. It is less responsive than the mean to extreme observations. Open the guide
Multiplicity
The increase in opportunities for a chance result when many hypotheses, outcomes, groups, or models are searched. Open the guide
Percentage point
An absolute difference between two rates expressed in percent, such as 6% minus 4% equals 2 points. Open the guide
Prevalence
The share of a defined population that has a condition or event. In an alert problem, it supplies the base rate. Open the guide
Regression to the mean
The expected inward movement of a group selected for extreme noisy measurements when it is measured again. Open the guide
Sampling distribution
The distribution of a statistic over repeated samples generated by the same design and population model. Open the guide
Specificity
The share of condition-absent cases correctly cleared by a binary detector; one minus specificity is its false-positive rate. Open the guide
Standard error
The standard deviation of an estimator across repeated samples. It describes sampling variation, not data quality by itself. Open the guide
Standardization
Reweighting group-specific results to a common composition so aggregate comparisons use like-for-like weights. Open the guide

Put the routes to work

Audit a complete case file

Three synthetic datasets turn the same ideas into an alert review, a segmented rollout, and a survey response audit.

Open the casebook