← Back to course

Sampling, Validity & Reliability

You’ll be able to

Populations and samples

You rarely study everyone; you study a sample meant to represent a larger population. How you select that sample shapes what your results can claim. Random sampling, where every member of the population has an equal chance of selection, best supports generalizing findings back to the population. Convenience sampling (whoever is easy to reach) is common in student research but limits generalization, because the sample may differ systematically from the population. Purposive sampling deliberately selects information-rich cases and is standard in qualitative work, where representativeness is not the goal. Naming your sampling method — and its limits — is part of honest design.

Validity: are you measuring the right thing?

Validity asks whether your study actually measures and reflects what it claims to. Internal validity is whether your design supports the causal or descriptive conclusion you draw — free of confounds and alternative explanations. External validity is whether your findings generalize beyond your specific sample and setting. Construct validity is whether your instrument truly captures the abstract concept you intend (does a five-question quiz really measure "anxiety"?). A study can be executed flawlessly yet be invalid if it measures the wrong thing or if a confound explains the result.

Reliability: are your results consistent?

Reliability asks whether your measurement is consistent and repeatable — would the same procedure produce the same results again? A scale that gives a different weight each time is unreliable. Reliability and validity are distinct: a measure can be reliable but not valid (a mis-set scale that consistently reads five pounds heavy — repeatable, but wrong), and a valid measure must first be reliable. In qualitative research the parallel concerns are credibility and dependability, often strengthened by triangulation and clear documentation. Good design protects both: you want the right measurement, taken consistently.

Worked example

A student surveys 30 friends about study habits and concludes, "High schoolers who listen to music study longer." Critique the design’s sampling and validity.

  1. 1.Examine the sample: 30 friends is a convenience sample, not random, and friends likely resemble the researcher — so it may not represent all high schoolers.
  2. 2.Assess external validity: because the sample is unrepresentative, generalizing to "high schoolers" broadly is not justified.
  3. 3.Assess internal validity: the conclusion implies causation, but a self-report survey at one time cannot rule out confounds — perhaps motivated students both study more and choose to play music.
  4. 4.Assess construct validity: does "study longer" (self-reported minutes) truly capture effective studying? Time is not the same as learning.
  5. 5.Recommend fixes: broaden and randomize the sample, avoid causal language for correlational data, and refine the measure of "studying."
Answer: The study uses a convenience sample of 30 friends, so its findings cannot generalize (weak external validity), and its causal claim is unsupported because confounds like motivation are not controlled (weak internal validity); "study longer" may also not validly capture effective studying (construct validity). Fixes include a broader, randomized sample, correlational (not causal) language, and a better-defined measure.
Checkpoint

A bathroom scale consistently reads exactly five pounds heavier than a person’s true weight every time. This measurement is best described as:

Watch out

Reliability does not guarantee validity. A consistently wrong instrument is reliably wrong. Always ask both questions separately: "Does it measure the right thing?" (validity) and "Does it measure consistently?" (reliability).

Checkpoint

A researcher recruits participants by posting in an online gaming forum, then claims the results represent "all teenagers." What is the main threat to the study’s external validity?

On the exam

The AP Research rubric expects you to name your sampling method and openly discuss threats to validity and reliability in your Limitations. Reviewers reward researchers who anticipate their own design’s weaknesses — hiding them reads as not understanding them.

Answer the 2 checkpoints as you read.

Sign in to save your progress