← Back to course

Operant Conditioning & Observational Learning

You’ll be able to

Reinforcement and punishment

B. F. Skinner’s operant conditioning shapes voluntary behavior through its consequences. Reinforcement increases a behavior; punishment decreases it. Crucially, positive means adding a stimulus and negative means removing one — neither means "good" or "bad." So positive reinforcement adds something desirable (a treat), negative reinforcement removes something aversive (taking aspirin ends a headache — the relief reinforces taking it), positive punishment adds something unpleasant (a scolding), and negative punishment removes something pleasant (losing phone privileges). Shaping builds new behaviors by reinforcing successive approximations toward a goal.

Schedules of reinforcement

How often reinforcement is delivered shapes behavior. Continuous reinforcement (reward every time) produces fast learning but fast extinction. Partial (intermittent) schedules resist extinction. Ratio schedules reward after a number of responses; interval schedules reward after time. Fixed schedules are predictable, variable unpredictable. Variable-ratio schedules (rewarding after an unpredictable number of responses, like a slot machine) produce the highest, most persistent response rates. Fixed-interval schedules produce a scalloped pattern — a burst of responding just before the reward is due (like studying right before a scheduled exam).

Observational learning: Bandura

Albert Bandura showed we also learn by watching — observational learning, or modeling. In his famous Bobo doll experiment, children who watched an adult aggressively hit an inflatable doll later imitated that aggression, while those who saw a gentle model did not. Learning occurred with no direct reinforcement of the child — vicarious consequences (seeing the model rewarded or punished) were enough. This work grounds social learning theory and drives concern about media violence, showing that behavior spreads through imitation of models, especially admired or similar ones.

Worked example

A teenager’s parents take away her car keys whenever she comes home past curfew. Her late arrivals decrease. Then they let her sleep in on weekends when she does her chores, and chores increase. Classify each consequence precisely.

  1. 1.First case — direction: coming home late decreases, so the consequence is a punishment.
  2. 2.First case — add or remove: the parents remove something desirable (the car keys), so it is negative punishment.
  3. 3.Second case — direction: doing chores increases, so the consequence is a reinforcement.
  4. 4.Second case — add or remove: the parents remove something aversive (the obligation to wake early / add the reward of sleeping in). Removing the early wake-up as relief makes it negative reinforcement; if you frame sleeping in as an added reward, it is positive reinforcement — the key is that the behavior increases.
Answer: Taking away the keys to reduce lateness is negative punishment (removing something pleasant to decrease behavior). Rewarding chores by letting her sleep in increases the behavior, so it is reinforcement — negative reinforcement if framed as removing an aversive early wake-up, or positive reinforcement if framed as adding a reward. Direction (increase vs. decrease) is what fixes reinforcement vs. punishment.
Checkpoint

A driver buckles their seatbelt to stop the car’s annoying beeping. Over time they buckle up faster and faster. Which process best describes why the buckling behavior increased?

Watch out

The classic trap: negative reinforcement is not punishment. Reinforcement always increases behavior; "negative" only means a stimulus is removed. Negative reinforcement (removing something aversive) strengthens a behavior, while punishment weakens one. Decide increase vs. decrease first, then add vs. remove.

Checkpoint

A gambler keeps pulling a slot-machine lever, which pays out after an unpredictable number of pulls. This schedule produces very high, steady responding that is hard to extinguish. Which schedule is it?

On the exam

For schedules, decode the two words: first word = predictability (fixed = set, variable = unpredictable), second word = trigger (ratio = number of responses, interval = time). Variable-ratio = fastest, most persistent responding (gambling).

Answer the 2 checkpoints as you read.

Sign in to save your progress