Inference for Categorical Data: Chi-Square
What this unit covers
The topics below follow the published Statistics course framework for Unit 8. This unit is worth 2–5% of the exam, so budget your time against that rather than against how long the unit takes to teach.
Lessons in this unit
- Chi-Square Goodness of Fit15 min · 3 objectivesState hypotheses for a goodness-of-fit test · Compute expected counts and the chi-square statistic · Determine degrees of freedom and interpret the result
- Expected Counts in Two-Way Tables13 min · 3 objectivesCompute expected counts for a two-way table under the null hypothesis · Explain the logic behind the expected-count formula · Verify the Large Counts condition for two-way tables
- Chi-Square Test for Independence14 min · 3 objectivesState hypotheses for a test of independence · Compute the chi-square statistic and degrees of freedom for a two-way table · Interpret the conclusion in context
- Independence vs. Homogeneity13 min · 3 objectivesDistinguish a test of independence from a test of homogeneity · Match the test to the sampling design · State appropriate hypotheses for each
Formulas in Unit 8
Every term in Unit 8
All 13 terms we publish for Inference for Categorical Data: Chi-Square, with definitions. Reading them through is the fastest way to find the ones you cannot define — then drill those in cram mode until you can produce them without the prompt.
- Chi-square statistic
- χ² = Σ(observed − expected)²/expected, summed over every cell. Always non-negative, and larger values mean worse fit to the null.
- Goodness-of-fit test
- Compares one categorical variable's observed counts to a claimed distribution. df = categories − 1.
- Test for homogeneity
- Compares the distribution of one categorical variable across several populations or treatments. df = (rows − 1)(columns − 1).
- Test for independence
- Tests whether two categorical variables are associated within one population. Same statistic and df as homogeneity; the difference is the sampling design.
- Conditions for chi-square
- Random sample or random assignment, 10% condition, and every EXPECTED count at least 5 — expected, not observed.
- Calculating expected counts
- For a two-way table, (row total × column total)/grand total.
- Chi-square distribution shape
- Right-skewed, becoming more symmetric as degrees of freedom rise. Only large values give small p-values, so the test is inherently one-sided.
- Which chi-square test to use
- One sample and one variable is goodness-of-fit; several samples and one variable is homogeneity; one sample and two variables is independence.
- Follow-up after a significant chi-square
- Identify the cells with the largest contributions to χ² and describe how observed differs from expected there, in context.
- Stating chi-square hypotheses
- Goodness-of-fit states a claimed distribution; homogeneity and independence state no difference and no association respectively, always in context.
- Why expected counts, not observed
- The condition guards the approximation of the sampling distribution, which depends on what the null predicts rather than on what was seen.
- Combining categories
- When an expected count is under 5, adjacent categories may be merged — reducing degrees of freedom accordingly.
- Interpreting a large contribution to chi-square
- That cell is where observed and expected diverge most, and it is where the description of the association should focus.
What examiners penalize here
- Expected counts are almost always decimals (like 18.4) — do not round them to whole numbers before computing χ². And check Large Counts on the **expected** counts, not the observed ones: every expected count must be at least 5.
- Decide independence vs. homogeneity by counting samples: **one** random sample with two variables → independence; **several** samples/groups compared on one variable → homogeneity. State the matching hypotheses — "independent" vs. "same distribution across groups" — to earn the setup point.
Practice Statistics
Our practice bank is drawn from across the whole course rather than filtered to one unit, which is closer to how the exam asks anyway — it will not tell you which unit a question is testing.
Questions about this unit
How much of the AP Statistics exam is Unit 8?
Unit 8, Inference for Categorical Data: Chi-Square, is worth 2–5% of the Statistics multiple-choice section according to the published course framework. Across all 9 units that makes it one of the lighter units, so it is not where a review phase should start.
What topics are covered in Statistics Unit 8?
Inference for Categorical Data: Chi-Square covers Goodness of fit, Independence, Homogeneity and Expected counts. We publish 13 terms with definitions for this unit, all of them on this page.
How should I study Statistics Unit 8?
Read the 4 lessons below first — about 55 minutes — then drill the 13 terms in cram mode until you can produce each definition from memory rather than just recognize it. Recognition is what makes a unit feel finished when it is not. Finish with practice questions and read the explanation for every one you get right by elimination as well as the ones you miss.
All 9 units of AP Statistics
- Unit 1 · Exploring One-Variable Data
- Unit 2 · Exploring Two-Variable Data
- Unit 3 · Collecting Data
- Unit 4 · Probability & Random Variables
- Unit 5 · Sampling Distributions
- Unit 6 · Inference for Proportions
- Unit 7 · Inference for Means
- Unit 8 · Inference for Categorical Data: Chi-Square
- Unit 9 · Inference for Quantitative Data: Slopes
Unit names, topics and exam weights follow the published College Board course framework for AP Statistics. AP® is a trademark registered by the College Board, which does not endorse this site.