Lab 09 / Sampling

When more answers miss the target

Can a larger sample be more precise and still be wrong?

7 min synthetic data fixed seed

Repeat a probability sample and a tilted voluntary sample, then separate shrinking random variation from persistent selection bias. Change one control at a time, then open the values under the chart before interpreting the picture.

Interactive synthetic lab
runs in this browser

View the values behind this chart

Loading the deterministic simulation...

SEED

What this display answers

A larger sample reduces random variation around the target created by its own sampling mechanism. It does not make that mechanism representative.

Here, equal-probability samples center on the programmed town rate of 60%. Voluntary samples draw 90% of responses from the higher-rate group, so they become increasingly stable around 76% instead.

What to notice

  1. Start with 50 observations. Both distributions are wide, but they have different centers.
  2. Increase the sample size toward 500. Both distributions narrow while the gap between their centers remains.
  3. Lower the voluntary pool's Group A share to 50%. The selection bias disappears because the response mix now matches the population mix.

The model behind it

The synthetic population is half Group A and half Group B. Their programmed positive-response rates are 80% and 40%, so the population target is 60%.

The probability mechanism samples the population target directly. The tilted mechanism changes only the group mix among available responses. Repeated draws are independent Bernoulli samples from those two stated rates.

Where the result stops

Real probability surveys may use weights, clusters, strata, and nonresponse adjustments that require design-specific uncertainty calculations.

A response rate or a large count alone does not reveal the direction or size of bias. That depends on how selection relates to the outcome.