Read complete course notes

All guides

Statistics

Ten lessons on describing data, z-scores, probability and Bayes, the binomial, sampling distributions, confidence intervals, hypothesis tests, t procedures, chi-square and regression, with sliders and worked examples.

A study guide to a first college statistics course, not a full course. Topics follow the standard introductory sequence. Practice values in the labs are invented. Work extra problems from a textbook.

['Algebra']

Course outline

  1. Describing data: center and spread

    Choose and compute a sensible center and spread, and see how outliers move them.

  2. Distributions and z-scores

    Standardize a value and read how unusual it is.

  3. Probability rules and Bayes

    Combine probabilities correctly and update with new evidence.

  4. Random variables and the binomial

    Compute the mean and spread of counts of successes.

  5. Sampling distributions and the central limit theorem

    Explain why sample means vary less than individual values.

  6. Confidence intervals for a proportion

    Build and read an interval with its margin of error.

  7. Hypothesis tests and p-values

    Run a one-proportion test and say what a p-value does and does not mean.

  8. The t distribution and inference for a mean

    Test a mean when the population sd is unknown.

  9. Chi-square tests for categorical data

    Compare observed counts with expected counts.

  10. Correlation and regression

    Read a fitted line and avoid causal mistakes.

Sources and curriculum note

Reviewed October 8, 2026.

Complete course reading notes

Read every lesson below. The interactive reader above contains the same explanations, with visual tools and quizzes.

1. Describing data: center and spread

Learning goal: Choose and compute a sensible center and spread, and see how outliers move them.

A data set has a center and a spread. The mean is the balance point: add the values and divide by how many there are. The median is the middle value after sorting. For a symmetric set they agree. For a skewed set or a set with an outlier, the mean is pulled toward the tail and the median barely moves.

Spread says how far values sit from the center. The range uses only the two extremes. The standard deviation measures a typical distance from the mean, so one extreme value can inflate it. The interquartile range, the width of the middle half of the data, resists outliers the way the median does.

Always describe shape, center, spread and unusual values together. Report the mean and standard deviation for roughly symmetric data. Report the median and the interquartile range when the data are skewed or contain outliers, such as house prices or incomes.

Worked example

Find the mean and median of 2, 3, 4, 5, 26.

  1. Add: 2+3+4+5+26 = 40
  2. Divide by 5: mean = 8
  3. Sorted middle value is 4: median = 4
  4. The outlier 26 pulls the mean above the median
Practice problem and solution

Find the mean of 2, 3, 4, 5 and 26. Enter a number.

Sum is 40 and there are 5 values, so the mean is 8.

Mental model: The mean is the balance point and the median is the middle; skew separates them.

Common trap: Using the mean and standard deviation on skewed data with outliers.

2. Distributions and z-scores

Learning goal: Standardize a value and read how unusual it is.

A z-score says how many standard deviations a value sits above or below the mean: z equals value minus mean, divided by standard deviation. A z of 0 is average, +1 is one standard deviation above, and -2 is two below. Because the score is unitless, you can compare an exam score with a height or a price.

For a roughly bell-shaped (normal) distribution, about 68 percent of values fall within 1 standard deviation of the mean, about 95 percent within 2, and about 99.7 percent within 3. A value with |z| above 2 is unusual, and above 3 is very unusual.

Z-scores describe where a value lies in its own distribution. They say nothing about whether the distribution is normal. Check the shape before you attach a normal-curve percentage to a z-score.

Worked example

A score is 85 on a test with mean 70 and sd 10. Find z.

  1. Subtract the mean: 85 - 70 = 15
  2. Divide by sd: 15 / 10 = 1.5
  3. z = 1.5, one and a half sd above average
Practice problem and solution

A score of 62 on a test with mean 70 and sd 4. Find z. Enter a number.

(62 - 70) / 4 = -2.

Mental model: z equals (value minus mean) over sd; |z| above 2 is unusual.

Common trap: Applying normal-curve percentages to data that are not bell shaped.

3. Probability rules and Bayes

Learning goal: Combine probabilities correctly and update with new evidence.

Probability runs from 0 to 1. The complement rule says P(not A) equals 1 minus P(A). For either of two events, P(A or B) equals P(A) + P(B) - P(A and B); subtract the overlap so it is not counted twice. For independent events, P(A and B) equals P(A) times P(B).

Conditional probability P(A given B) restricts attention to cases where B happened: P(A given B) equals P(A and B) over P(B). It is not the same as P(B given A). Confusing the two is the most common error in reading medical tests and court statistics.

Bayes' rule turns a test result into an updated probability. Start with the base rate, how common the condition is. A positive result from a good test can still be probably wrong when the condition is rare, because false positives from the many healthy people outnumber true positives from the few sick ones.

Worked example

Base rate 10 percent, test finds 90 percent of cases and clears 90 percent of healthy people. P(sick given positive)?

  1. Out of 1000 people: 100 sick, 900 healthy
  2. True positives: 90 of 100
  3. False positives: 10 percent of 900 = 90
  4. Sick given positive = 90 / (90 + 90) = 50 percent
Practice problem and solution

Base rate 10 percent, sensitivity 90 percent, specificity 90 percent. P(sick given positive) as a percent? Enter a number.

90 true positives and 90 false positives out of 1000 people gives 50 percent.

Mental model: Update a base rate with Bayes; rare conditions produce many false positives.

Common trap: Treating P(A given B) as if it equals P(B given A).

4. Random variables and the binomial

Learning goal: Compute the mean and spread of counts of successes.

A random variable assigns a number to each outcome of a chance process. Its expected value is the long-run average, found by weighting each value by its probability. Its standard deviation measures typical distance from that average.

A binomial setting has a fixed number n of trials, each with two outcomes, the same success probability p, and independent trials. The number of successes then has mean n times p and standard deviation the square root of n times p times (1 - p). The chance of no success at all is (1 - p) raised to the power n.

Check the conditions before you use the binomial. Drawing cards without replacement is not independent, and a probability that changes from trial to trial breaks the model. A small sample from a large population is an acceptable approximation.

Worked example

Ten shots at 30 percent each. Find the mean number of hits.

  1. n = 10, p = 0.3
  2. Mean = n x p = 3
  3. SD = sqrt(10 x 0.3 x 0.7) = 1.45
Practice problem and solution

Twenty trials with success chance 0.4. Mean number of successes? Enter a number.

20 x 0.4 = 8.

Mental model: Binomial counts have mean np and spread sqrt(np(1-p)).

Common trap: Using the binomial when trials are dependent or p changes.

5. Sampling distributions and the central limit theorem

Learning goal: Explain why sample means vary less than individual values.

A statistic such as a sample mean changes from sample to sample. Its sampling distribution describes that variation. The standard deviation of the sampling distribution is called the standard error. For a mean it equals the population standard deviation divided by the square root of n.

The central limit theorem says that for large enough n the sampling distribution of the mean is approximately normal, whatever the shape of the population. Larger samples shrink the standard error, but slowly: to cut it in half you need four times the sample size.

A bigger sample reduces random error, not bias. A biased sampling method, such as asking only volunteers, stays biased at any size. Random selection is what makes the standard error meaningful.

Worked example

Population sd 20, n = 25. Find the standard error.

  1. SE = sd / sqrt(n)
  2. sqrt(25) = 5
  3. SE = 20 / 5 = 4
Practice problem and solution

Population sd 30, n = 36. Standard error? Enter a number.

30 / 6 = 5.

Mental model: SE is sd over root n; bigger samples reduce random error, not bias.

Common trap: Thinking a huge sample cures a biased sampling method.

6. Confidence intervals for a proportion

Learning goal: Build and read an interval with its margin of error.

A confidence interval gives a range of plausible values for a population parameter. For a proportion it is the sample proportion plus or minus a margin of error. The margin equals a critical value times the standard error, and for 95 percent confidence the critical value is about 1.96.

The standard error of a sample proportion is the square root of p-hat times (1 - p-hat) over n. The margin is widest when p-hat is near one half and shrinks as n grows, but only with the square root of n.

The 95 percent describes the method: if you repeated the sampling many times, about 95 percent of intervals built this way would contain the true value. It does not say there is a 95 percent chance this one interval does. Check that the sample is random and that there are at least 10 successes and 10 failures.

Worked example

n = 400, p-hat = 0.5. Find the 95 percent margin.

  1. SE = sqrt(0.5 x 0.5 / 400) = 0.025
  2. Margin = 1.96 x 0.025 = 0.049
  3. About 4.9 points
Practice problem and solution

n = 100, p-hat = 0.5. 95 percent margin in percentage points? Enter a number to one decimal place.

1.96 x sqrt(0.25/100) = 1.96 x 0.05 = 0.098, which is 9.8 points.

Mental model: Margin = 1.96 x SE; it shrinks with root n.

Common trap: Reading 95 percent as the chance that this one interval holds the truth.

7. Hypothesis tests and p-values

Learning goal: Run a one-proportion test and say what a p-value does and does not mean.

A hypothesis test starts with a null hypothesis, a specific claim such as p equals 0.5, and asks how surprising the data would be if it were true. The test statistic measures the distance between the sample result and the null value in standard errors. For a proportion, z equals (p-hat - p0) over sqrt(p0(1 - p0)/n).

The p-value is the probability, assuming the null is true, of a result at least as extreme as the one observed. A small p-value, commonly below 0.05, is evidence against the null. A large p-value does not prove the null; it says the data are consistent with it.

Two errors are possible. A Type I error rejects a true null, and its chance is the significance level. A Type II error fails to reject a false null. Statistical significance is not the same as practical importance, since a huge sample can make a trivial difference significant.

Worked example

60 successes in 100, null p = 0.5. Find z.

  1. SE = sqrt(0.5 x 0.5 / 100) = 0.05
  2. p-hat - p0 = 0.6 - 0.5 = 0.1
  3. z = 0.1 / 0.05 = 2
  4. 2 exceeds 1.96, so reject at the 5 percent level
Practice problem and solution

58 successes in 100, null p = 0.5. Find z. Enter a number to one decimal place.

(0.58 - 0.5) / 0.05 = 1.6.

Mental model: Test statistic = distance in standard errors; p-value is about the data, not the null.

Common trap: Reading the p-value as the probability that the null is true.

8. The t distribution and inference for a mean

Learning goal: Test a mean when the population sd is unknown.

When the population standard deviation is unknown you estimate it with the sample standard deviation s, and the standardized mean follows a t distribution rather than a normal one. The statistic is t equals (sample mean - null mean) over (s over sqrt(n)).

The t distribution has heavier tails than the normal, which pays for the extra uncertainty from estimating s. Its shape depends on degrees of freedom, n - 1. As n grows it approaches the normal curve. A 95 percent interval for the mean is the sample mean plus or minus a t critical value times s over sqrt(n).

Conditions: the sample is random, and the data are roughly normal or n is large (about 30 or more). Plot the data first. Strong skew or outliers with a small sample make t procedures unreliable.

Worked example

n = 36, mean 106, s = 12, null mean 100. Find t.

  1. SE = 12 / sqrt(36) = 2
  2. Difference = 106 - 100 = 6
  3. t = 6 / 2 = 3
  4. df = 35; t = 3 is strong evidence against the null
Practice problem and solution

n = 49, mean 52, s = 14, null mean 50. Find t. Enter a number to one decimal place.

SE = 14/7 = 2; t = (52 - 50)/2 = 1.

Mental model: Replace sigma with s and use t with n-1 df.

Common trap: Using a z critical value with a small sample and unknown sd.

9. Chi-square tests for categorical data

Learning goal: Compare observed counts with expected counts.

For counts in categories, the chi-square statistic adds up, over every category, (observed - expected) squared divided by expected. A small total means the data sit close to the claim. A large total means at least one category is far off.

In a goodness-of-fit test the null gives the expected share of each category, and degrees of freedom equal the number of categories minus 1. For a fair six-sided die rolled 60 times, each face is expected 10 times and df is 5. The 5 percent critical value for df 5 is about 11.07.

Use counts, not percentages, and require each expected count to be at least 5. The test says whether the pattern differs from the null, not which category is responsible; compare the individual terms. A test of independence in a two-way table works the same way with df of (rows - 1) times (columns - 1).

Worked example

Fair die, 60 rolls, observed 16, 4, 10, 10, 10, 10. Find chi-square.

  1. Expected each = 10
  2. (16-10)^2/10 = 3.6
  3. (4-10)^2/10 = 3.6
  4. The other four terms are 0
  5. Total = 7.2, below 11.07: no evidence the die is unfair
Practice problem and solution

Observed 15, 5, 10, 10, 10, 10 against expected 10 each. Find chi-square. Enter a number.

(15-10)^2/10 + (5-10)^2/10 = 2.5 + 2.5 = 5.

Mental model: Chi-square sums squared gaps over expected; df is categories minus 1.

Common trap: Computing with percentages or with expected counts below 5.

10. Correlation and regression

Learning goal: Read a fitted line and avoid causal mistakes.

The correlation r runs from -1 to 1 and measures the strength and direction of a linear relationship. It has no units and is easily moved by outliers. A curved pattern can have r near 0 while the variables are strongly related.

The least-squares line predicts y from x. Its slope equals r times (sd of y over sd of x), and it passes through the point of the two means. R squared, the square of r, is the share of the variation in y that the line accounts for. An r of 0.8 gives an R squared of 0.64.

Do not predict far outside the range of x, an error called extrapolation. Check the residual plot for curves or changing spread. Correlation does not prove causation: a lurking variable, such as temperature driving both ice cream sales and sunburns, can create an association with no direct link.

Worked example

r = 0.8, sd y = 12, sd x = 4. Find slope and R squared.

  1. Slope = r x (sy / sx) = 0.8 x 3 = 2.4
  2. R squared = 0.8^2 = 0.64
  3. The line explains 64 percent of the variation in y
Practice problem and solution

r = 0.5, sd y = 12, sd x = 4. Find the slope. Enter a number.

0.5 x 12 / 4 = 1.5.

Mental model: Slope is r times sy over sx; R squared is r squared; association is not cause.

Common trap: Treating a strong correlation as proof that one variable causes the other.