The Fundamental Problem of RL
In supervised learning, the loss function is differentiable with respect to model parameters. In reinforcement learning, the “loss” (negative reward) depends on actions sampled from the policy. You cannot differentiate through a sample.
An agent with policy generates a trajectory and receives total reward . We want to maximize:
The expectation is over trajectories sampled from the policy. The policy determines the distribution, the reward is evaluated on the sample, and we need .
1. The Log-Derivative Trick
The key identity: for any distribution ,
Proof (one line):
We multiplied and divided by . The ratio is the score function (of the parameter, not the data).
This transforms the gradient of an expectation into an expectation of a gradient—which we can estimate by Monte Carlo sampling.
2. REINFORCE
Applying the log-derivative trick to the RL objective:
The gradient estimate from a single trajectory:
Interpretation: Increase the log-probability of actions that led to high reward. Decrease it for low reward. The gradient is proportional to both the reward and the score of each action.
This is REINFORCE (Williams, 1992). It is unbiased: . But it has a devastating problem.
3. The Variance Problem
REINFORCE has enormous variance. Two sources:
Credit assignment: The total reward multiplies every action’s score, even though action had nothing to do with reward . All actions are “credited” equally for the total outcome.
Scale sensitivity: If is always positive (e.g., rewards in ), the gradient always increases the probability of every action taken. The learning signal comes only from the relative magnitude of the reward, which is buried under a large mean.
Both issues inflate the variance of the gradient estimate, requiring many samples per update.
4. Baselines: Variance Reduction Without Bias
Key insight: For any function that depends only on the state (not the action):
This is because has zero mean under (a consequence of ).
Therefore, subtracting from the reward does not change the expected gradient:
But it can dramatically reduce the variance. The optimal baseline is —the expected return from state . This is the value function .
With baseline, the gradient becomes:
The term is the advantage: how much better this action was compared to the average. Actions better than expected get reinforced; worse-than-expected get suppressed.
Connection to martingales: The advantage is a martingale difference sequence—its conditional expectation given is zero. This orthogonality (cf. Martingales) is precisely why the baseline reduces variance without introducing bias.
5. From REINFORCE to Actor-Critic
| Method | Gradient estimate | Variance | Bias |
|---|---|---|---|
| REINFORCE | Very high | None | |
| + Baseline | High | None | |
| Actor-Critic | Low | Some (from ) | |
| GAE () | Tunable | Tunable |
Actor-Critic replaces the Monte Carlo return with the one-step TD error . This introduces bias (because is approximate) but drastically reduces variance (only one step of randomness instead of the full trajectory).
GAE (Generalized Advantage Estimation) interpolates between the two extremes via a parameter :
- : Full Monte Carlo return (unbiased, high variance)
- : One-step TD (biased, low variance)
This is exactly the bias-variance tradeoff in a new guise: more bootstrapping (lower ) reduces variance at the cost of bias from the value function approximation.
The Takeaway
Policy gradient is built on a single algebraic trick—the log-derivative identity—that converts an intractable gradient-through-sampling into a tractable expectation. Everything that follows (baselines, advantages, actor-critic, GAE) is variance reduction: the ongoing struggle to extract a clean learning signal from the inherent noise of sampled trajectories. The tools are orthogonality (baselines), bootstrapping (TD learning), and the bias-variance tradeoff—the same ideas that recur across all of statistical learning.