Sampling Bias LabStatLens

Pick a "representative" sample of words from the Gettysburg Address by eye, then compare it to a random sample. The passage is our population of N = 268 words; the variable is each word's length (letters). The true mean is marked on both charts.

Simulate many:

Click words in the passage to build a by-eye sample of n, then “Add my by-eye sample.” “Simulate many” models the tendency to pick longer, important-looking words.

The population: Lincoln's Gettysburg Address

“By eye” samples (biased)

Random samples (unbiased)

Add a by-eye sample and a random sample to compare them.

Sampling Bias Lab

This is the classic "Sampling Words from the Gettysburg Address" activity. People asked to pick a representative sample tend to choose longer, memorable words, so their sample mean word-length runs too high — a sampling bias. A random sample has no such tendency: its mean centers on the true population mean.

Try this: add several by-eye samples and several random samples. The by-eye dots pile up to the right of the true mean; the random dots straddle it. Then increase n: the random distribution narrows on the truth, but the by-eye distribution narrows on the wrong value — bias doesn't go away with a bigger sample.

How to use it

Compare the two charts: the by-eye distribution centers to the right of μ (biased), while the random distribution centers on μ (unbiased).

True mean word length ≈ 4.3 letters (marked μ on both charts).