Describing Distributions
- Classify variables as categorical or quantitative and choose an appropriate graph
- Describe a distribution by its shape, center, spread, and unusual features in context
- Read shape from a graph, including skew direction and modality
Variables come in two kinds
Every column of data is a variable, and it is either categorical (values are labels or groups — eye color, brand, yes/no) or quantitative (values are numbers you can meaningfully average — height, income, test score). The kind decides the graph: categorical data go in bar charts and two-way tables, quantitative data go in dotplots, stemplots, histograms, and boxplots. A common trap is a number that is really a label — a zip code or a jersey number is categorical even though it looks numeric.
Shape, Center, Spread, and unusual features
To describe a quantitative distribution, hit four things — remember SOCS: Shape (symmetric, skewed left, or skewed right; one peak or several), Outliers/unusual features (gaps, clusters, extreme values), Center (a typical value, mean or median), and Spread (how much values vary, range/IQR/standard deviation). On the AP exam, description only earns credit in context — name the variable and its units, not just the numbers.
Reading skew
A distribution is skewed right when it has a long tail stretching toward the high values (most data bunched low, tail to the right) and skewed left when the long tail stretches toward the low values. The name always follows the tail, not the peak. Incomes are famously skewed right — most people earn modest amounts while a few earn enormous sums that pull the tail rightward.
Say the four descriptors out loud as SOCS — Shape, Outliers, Center, Spread — and attach the variable name to each. "The distribution of daily rainfall in mm is skewed right, centered near 3 mm, spread from 0 to 40 mm, with a possible high outlier" earns far more than "it goes up then down."
A histogram of the sale prices of 200 homes has most bars clustered between $150k and $350k, with a few bars trailing out to $1.2 million and no homes below $150k. Describe the shape and explain which direction it is skewed.
- 1.Locate the bulk of the data: most homes fall in the $150k–$350k range, so the peak is on the low end.
- 2.Locate the tail: a few very expensive homes stretch far out toward $1.2 million, on the high (right) side.
- 3.The long tail points toward the large values, and skew is named for the tail.
Which of the following variables is categorical?
Numbers are not automatically quantitative. Ask whether an average of the values would mean anything. The mean of a set of area codes or player numbers is nonsense — those are categorical variables that merely happen to use digits.
A distribution of exam scores has a long tail stretching toward the low scores, while most students scored high. How is this distribution best described?
Answer the 2 checkpoints as you read.
Sign in to save your progress