All 9 Statistics units
📊
AP Statistics · Unit 2 of 9

Exploring Two-Variable Data

5–7% of the exam4 lessons · 54 min19 terms

What this unit covers

The topics below follow the published Statistics course framework for Unit 2. This unit is worth 5–7% of the exam, so budget your time against that rather than against how long the unit takes to teach.

ScatterplotsCorrelationRegressionResiduals

Lessons in this unit

Formulas in Unit 2

Correlation coefficient
r = (1/(n−1)) · Σ [ (x_i − x-bar)/s_x ] · [ (y_i − y-bar)/s_y ]
r is the average product of the z-scores of x and y. Because it uses standardized values, r has no units and is unaffected by changes of scale.
Regression line and slope
y-hat = a + bx · b = r · (s_y / s_x) · a = y-bar − b·x-bar
The slope equals r times the ratio of the standard deviations. The line passes through (x-bar, y-bar), which is how the intercept formula is derived.
Coefficient of determination
r² = (fraction of variation in y explained by the linear model on x)
r² ranges from 0 to 1. If r² = 0.64, then 64% of the variation in the response is accounted for by the linear relationship with x. Square r to get it.
Residual
residual = y − y-hat = observed − predicted
A positive residual means the actual value lies above the line (model underpredicted); a negative residual means it lies below (model overpredicted).

Every term in Unit 2

All 19 terms we publish for Exploring Two-Variable Data, with definitions. Reading them through is the fastest way to find the ones you cannot define — then drill those in cram mode until you can produce them without the prompt.

Residual
Observed minus predicted, y − ŷ. A positive residual means the model underestimated that observation.
Explanatory vs response variable
The explanatory variable goes on the x-axis and is thought to explain changes in the response on the y-axis.
Describing a scatterplot
Direction, form, strength and unusual features, in context. All four are required for full credit.
Correlation coefficient r
Measures the strength and direction of a LINEAR relationship, between −1 and 1. It has no units and is unaffected by which variable is x.
What r does not tell you
It does not establish causation, does not detect curved relationships, and is not resistant to outliers.
Least-squares regression line
ŷ = a + bx, minimizing the sum of squared residuals. It always passes through the point (x̄, ȳ).
Interpreting the slope
"For each additional one-unit increase in x, the predicted y increases by b units" — predicted, in context, with units.
Interpreting the y-intercept
The predicted response when x = 0. Often meaningless in context, and worth saying so.
Coefficient of determination r²
The percentage of variation in y explained by the linear relationship with x. Always report it as a percentage of variation.
Residual plot
A random scatter about zero supports a linear model; a curved pattern means the relationship is not linear and a different model is needed.
Influential point vs outlier
An outlier lies far from the pattern; an influential point substantially changes the regression line, usually because its x-value is extreme.
Extrapolation
Predicting outside the range of the observed x-values. Unreliable, because there is no evidence the pattern continues.
Transforming to achieve linearity
Taking logs of y linearises exponential relationships; logs of both linearises power relationships.
Interpreting r in context
State the direction, the strength and that the relationship is linear, naming both variables. r alone is not an interpretation.
Predicting with a regression equation
Substitute and label the answer as a predicted value. Saying it will happen overstates what a model provides.
Residual for a specific point
Observed minus predicted at that x. A positive residual means the model underpredicted that observation.
Why r² is not r squared conceptually
r² is the proportion of variation in the response explained by the model, which is a different statement from the strength of the linear association.
Effect of removing an influential point
Recompute and compare. An influential point that pulls the line toward itself will change the slope noticeably when removed.
Choosing a transformation
A curved residual plot means the linear model is wrong. Try log(y) for exponential growth and log-log for a power relationship.

What examiners penalize here

Practice Statistics

Our practice bank is drawn from across the whole course rather than filtered to one unit, which is closer to how the exam asks anyway — it will not tell you which unit a question is testing.

Questions about this unit

How much of the AP Statistics exam is Unit 2?

Unit 2, Exploring Two-Variable Data, is worth 5–7% of the Statistics multiple-choice section according to the published course framework. Across all 9 units that makes it one of the lighter units, so it is not where a review phase should start.

What topics are covered in Statistics Unit 2?

Exploring Two-Variable Data covers Scatterplots, Correlation, Regression and Residuals. We publish 19 terms with definitions for this unit, all of them on this page.

How should I study Statistics Unit 2?

Read the 4 lessons below first — about 55 minutes — then drill the 19 terms in cram mode until you can produce each definition from memory rather than just recognize it. Recognition is what makes a unit feel finished when it is not. Finish with practice questions and read the explanation for every one you get right by elimination as well as the ones you miss.

All 9 units of AP Statistics

  1. Unit 1 · Exploring One-Variable Data
  2. Unit 2 · Exploring Two-Variable Data
  3. Unit 3 · Collecting Data
  4. Unit 4 · Probability & Random Variables
  5. Unit 5 · Sampling Distributions
  6. Unit 6 · Inference for Proportions
  7. Unit 7 · Inference for Means
  8. Unit 8 · Inference for Categorical Data: Chi-Square
  9. Unit 9 · Inference for Quantitative Data: Slopes

Unit names, topics and exam weights follow the published College Board course framework for AP Statistics. AP® is a trademark registered by the College Board, which does not endorse this site.