§12 Reward Learning — Rescorla–Wagner and TD(λ)
RW: ΔV = α · (λ - V)
TD: δ = r + γ · V(s') - V(s)
TD: δ = r + γ · V(s') - V(s)
The Rescorla–Wagner model describes classical conditioning: the value V of a stimulus converges asymptotically to the reward λ. The TD(λ) model generalizes this to sequential decisions, producing a reward prediction error signal δ that closely matches the firing patterns of dopamine neurons. The blocking effect is a key prediction of the RW model.
— ✦ —