Bias in Data Collection
- Explain how the way data is collected can introduce bias
- Describe how a non-representative sample skews conclusions
- Recognize privacy and consent issues in gathering data
Where data bias begins
Bias in data is a systematic slant that makes the data unrepresentative of reality, and it often creeps in at the moment of collection. If the people or situations sampled do not reflect the whole population you care about, every conclusion drawn from that data inherits the distortion — no amount of clever analysis afterward can fully remove it. Bias is dangerous precisely because the numbers still look objective; the flaw is in what was measured and who was included, not in the arithmetic.
The unrepresentative sample
A sample is the subset of a population you actually collect data from, and it should mirror the whole population. A sampling bias occurs when some groups are systematically over- or under-represented. Polling only smartphone-app users about a city policy ignores residents without smartphones, who may be older or lower-income; a survey answered only by people who feel strongly over-represents extremes. The result is a confident conclusion about "everyone" that really describes only the slice that was reachable.
Consent, privacy, and how data is gathered
How data is gathered raises ethical questions beyond accuracy. People generate data constantly — searches, locations, purchases — and much is collected without their clear awareness. Consent (did the person agree to this use?) and privacy (is personal information exposed or combined in revealing ways?) are central concerns. Data collected without consent can violate expectations even when it is technically legal, and combining separate data sets can re-identify people who thought they were anonymous.
A company measures customer satisfaction by emailing a survey link and analyzing only the responses it receives. Why might its results be biased?
- 1.Identify who is sampled: only customers who both received the email and chose to respond.
- 2.Consider who is left out: customers without email on file, and anyone who ignored the survey — often those who are merely content, or those with no time.
- 3.Consider who over-responds: people with intense opinions (very happy or very angry) are the most motivated to reply.
- 4.So the sample over-represents strong opinions and excludes non-email customers, making it unrepresentative of all customers.
A researcher wants to know the favorite sport of all students in a large high school but surveys only students leaving basketball tryouts. Why are the results biased?
Bias introduced during collection cannot be fixed by analyzing the data harder. If the sample was not representative, the conclusion is skewed no matter how careful the math is.
A fitness app quietly sells users’ location histories to advertisers, a use buried deep in its terms of service. Which pair of concerns does this most directly raise?
When a scenario describes collecting data, scan for two things: is the sample representative, and was there consent/privacy protection? Those are the two questions the exam repeatedly rewards you for asking.
Answer the 2 checkpoints as you read.
Sign in to save your progress