Collecting Data
What this unit covers
The topics below follow the published Statistics course framework for Unit 3. This unit is worth 12–15% of the exam, so budget your time against that rather than against how long the unit takes to teach.
Lessons in this unit
- Sampling Methods13 min · 3 objectivesDistinguish a population from a sample and a parameter from a statistic · Describe simple random, stratified, cluster, and systematic sampling · Explain why random sampling supports generalization to a population
- Experiments & Their Design15 min · 3 objectivesIdentify the treatments, experimental units, explanatory and response variables · Apply the three principles of experimental design: control, randomization, replication · Explain the purpose of a control group and a placebo
- Bias in Data Collection12 min · 3 objectivesIdentify undercoverage, nonresponse, and response bias · Explain how question wording can bias survey results · Distinguish bias from sampling variability
- Randomization, Blocking & Scope of Conclusions13 min · 3 objectivesDescribe randomized block and matched-pairs designs and their purpose · Explain how blocking reduces variability · State the scope of conclusions permitted by a study’s design
Formulas in Unit 3
Every term in Unit 3
All 24 terms we publish for Collecting Data, with definitions. Reading them through is the fastest way to find the ones you cannot define — then drill those in cram mode until you can produce them without the prompt.
- Observational study vs experiment
- An experiment imposes a treatment; an observational study only records. Only an experiment can establish causation.
- Blocking
- Grouping similar units and randomizing within each block, to remove a known source of variability. The experimental analogue of stratifying.
- Population vs sample
- The population is everyone you want to describe; the sample is who you actually measure. Inference generalises from one to the other.
- Parameter vs statistic
- A parameter describes a population (μ, p, σ); a statistic describes a sample (x̄, p̂, s). Statistics estimate parameters.
- Simple random sample
- Every group of the given size has an equal chance of being chosen — a stronger condition than every individual having an equal chance.
- Stratified random sample
- Divide the population into similar groups and sample within each. Reduces variability when strata differ from each other.
- Cluster sample
- Divide into groups, then randomly select whole groups. Cheaper than stratifying, and used when clusters resemble the population.
- Systematic sample
- Select every kth individual after a random start. Valid unless there is a periodic pattern matching the interval.
- Convenience and voluntary response samples
- Both are biased — voluntary response over-represents people with strong opinions. Neither supports inference.
- Undercoverage
- Some part of the population has no chance of being selected, so the sample cannot represent it however large it is.
- Nonresponse bias
- Selected individuals who do not respond may differ systematically from those who do.
- Response bias
- The way a question is asked, or who asks it, changes the answers given. Distinct from nonresponse.
- Confounding variable
- A variable associated with both the explanatory variable and the response, so their effects cannot be separated.
- Three principles of experimental design
- Comparison with a control group, random assignment of treatments, and replication with enough experimental units.
- Random assignment vs random selection
- Random assignment permits causal conclusions; random selection permits generalization to the population. They answer different questions.
- Blinding and placebo effect
- Single-blind hides treatment from subjects, double-blind from subjects and those measuring. Controls for expectation, which is a real physiological effect.
- Matched pairs design
- Each subject receives both treatments, or subjects are paired by a similar characteristic. Analyzed with a one-sample t procedure on the differences.
- Describing a sampling procedure fully
- Name the population, how individuals are numbered, how the random mechanism is applied, and what to do about repeats.
- Why a large biased sample is worse than a small random one
- Bias does not shrink with sample size; only variability does. A million-person voluntary poll is still unrepresentative.
- Scope of inference
- Random selection lets you generalize to the population; random assignment lets you claim causation. Both together permit both conclusions.
- Experimental units vs subjects
- Subjects are human experimental units. Naming them correctly matters when describing replication.
- Completely randomized vs block design
- Completely randomized assigns treatments at random to all units; blocking groups similar units first and randomises within each block.
- Why blocking is not the same as stratifying
- Blocking happens in an experiment to control a nuisance variable; stratifying happens in a sample to reduce variability of an estimate.
- Placebo and control group
- A control group provides a comparison baseline; a placebo controls for the psychological effect of receiving a treatment. They are not synonyms.
What examiners penalize here
- Keep two "randoms" distinct: random **selection** (sampling) lets you *generalize* to a population; random **assignment** (experiments) lets you conclude *causation*. A study can have one, both, or neither — and that determines exactly what conclusions are allowed.
- When asked to identify bias, name the *type* (undercoverage, nonresponse, response, or wording) **and** state its likely *direction* — will the estimate be too high or too low? AP rubrics reward explaining how the flaw pushes the result.
- A classic free-response ending: "Can we conclude the treatment *caused* the difference, and can we generalize to all ___?" Answer both parts using the design — cite random *assignment* for causation and random *selection* for generalization. Missing either random feature limits the claim.
Practice Statistics
Our practice bank is drawn from across the whole course rather than filtered to one unit, which is closer to how the exam asks anyway — it will not tell you which unit a question is testing.
Questions about this unit
How much of the AP Statistics exam is Unit 3?
Unit 3, Collecting Data, is worth 12–15% of the Statistics multiple-choice section according to the published course framework. Across all 9 units that makes it a substantial share — heavier than an even split would give it.
What topics are covered in Statistics Unit 3?
Collecting Data covers Sampling, Experiments, Bias and Randomization. We publish 24 terms with definitions for this unit, all of them on this page.
How should I study Statistics Unit 3?
Read the 4 lessons below first — about 55 minutes — then drill the 24 terms in cram mode until you can produce each definition from memory rather than just recognize it. Recognition is what makes a unit feel finished when it is not. Finish with practice questions and read the explanation for every one you get right by elimination as well as the ones you miss.
All 9 units of AP Statistics
- Unit 1 · Exploring One-Variable Data
- Unit 2 · Exploring Two-Variable Data
- Unit 3 · Collecting Data
- Unit 4 · Probability & Random Variables
- Unit 5 · Sampling Distributions
- Unit 6 · Inference for Proportions
- Unit 7 · Inference for Means
- Unit 8 · Inference for Categorical Data: Chi-Square
- Unit 9 · Inference for Quantitative Data: Slopes
Unit names, topics and exam weights follow the published College Board course framework for AP Statistics. AP® is a trademark registered by the College Board, which does not endorse this site.