The Fundamental Problem of RL

In supervised learning, the loss function is differentiable with respect to model parameters. In reinforcement learning, the “loss” (negative reward) depends on actions sampled from the policy. You cannot differentiate through a sample.

An agent with policy πθ(a∣s)\pi_\theta(a|s) generates a trajectory τ=(s0,a0,r0,s1,a1,r1,…)\tau = (s_0, a_0, r_0, s_1, a_1, r_1, \ldots) and receives total reward R(τ)R(\tau). We want to maximize:

J(θ)=Eτ∼πθ[R(τ)]J(\theta) = \mathbb{E}_{\tau \sim \pi_\theta}[R(\tau)]

The expectation is over trajectories sampled from the policy. The policy determines the distribution, the reward is evaluated on the sample, and we need ∇θJ\nabla_\theta J.

1. The Log-Derivative Trick

The key identity: for any distribution pθ(x)p_\theta(x),

∇θEx∼pθ[f(x)]=Ex∼pθ[f(x) ∇θlog⁡pθ(x)]\nabla_\theta \mathbb{E}_{x \sim p_\theta}[f(x)] = \mathbb{E}_{x \sim p_\theta}[f(x) \, \nabla_\theta \log p_\theta(x)]

Proof (one line):

∇θ∫f(x)pθ(x) dx=∫f(x)∇θpθ(x) dx=∫f(x) pθ(x)∇θpθ(x)pθ(x) dx\nabla_\theta \int f(x) p_\theta(x) \, dx = \int f(x) \nabla_\theta p_\theta(x) \, dx = \int f(x) \, p_\theta(x) \frac{\nabla_\theta p_\theta(x)}{p_\theta(x)} \, dx

We multiplied and divided by pθ(x)p_\theta(x). The ratio ∇θpθ/pθ=∇θlog⁡pθ\nabla_\theta p_\theta / p_\theta = \nabla_\theta \log p_\theta is the score function (of the parameter, not the data).

This transforms the gradient of an expectation into an expectation of a gradient—which we can estimate by Monte Carlo sampling.

2. REINFORCE

Applying the log-derivative trick to the RL objective:

∇θJ=Eτ∼πθ[R(τ)∑t=0T∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}_{\tau \sim \pi_\theta}\left[R(\tau) \sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t | s_t)\right]

The gradient estimate from a single trajectory:

g^=R(τ)∑t=0T∇θlog⁡πθ(at∣st)\hat{g} = R(\tau) \sum_{t=0}^{T} \nabla_\theta \log \pi_\theta(a_t | s_t)

Interpretation: Increase the log-probability of actions that led to high reward. Decrease it for low reward. The gradient is proportional to both the reward and the score of each action.

This is REINFORCE (Williams, 1992). It is unbiased: E[g^]=∇θJ\mathbb{E}[\hat{g}] = \nabla_\theta J. But it has a devastating problem.

3. The Variance Problem

REINFORCE has enormous variance. Two sources:

Credit assignment: The total reward R(τ)R(\tau) multiplies every action’s score, even though action a0a_0 had nothing to do with reward r99r_{99}. All actions are “credited” equally for the total outcome.

Scale sensitivity: If R(τ)R(\tau) is always positive (e.g., rewards in [0,100][0, 100]), the gradient always increases the probability of every action taken. The learning signal comes only from the relative magnitude of the reward, which is buried under a large mean.

Both issues inflate the variance of the gradient estimate, requiring many samples per update.

4. Baselines: Variance Reduction Without Bias

Key insight: For any function b(s)b(s) that depends only on the state (not the action):

Ea∼πθ[∇θlog⁡πθ(a∣s)⋅b(s)]=0\mathbb{E}_{a \sim \pi_\theta}[\nabla_\theta \log \pi_\theta(a|s) \cdot b(s)] = 0

This is because ∇θlog⁡πθ(a∣s)\nabla_\theta \log \pi_\theta(a|s) has zero mean under πθ\pi_\theta (a consequence of ∇θ∫πθ(a∣s) da=∇θ1=0\nabla_\theta \int \pi_\theta(a|s) \, da = \nabla_\theta 1 = 0).

Therefore, subtracting b(s)b(s) from the reward does not change the expected gradient:

∇θJ=E[(R(τ)−b(st))∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}\left[(R(\tau) - b(s_t)) \nabla_\theta \log \pi_\theta(a_t | s_t)\right]

But it can dramatically reduce the variance. The optimal baseline is b∗(s)=E[R(τ)∣st=s]b^*(s) = \mathbb{E}[R(\tau) | s_t = s]—the expected return from state ss. This is the value function Vπ(s)V^\pi(s).

With baseline, the gradient becomes:

g^=∑t(Rt−V(st))∇θlog⁡πθ(at∣st)\hat{g} = \sum_t (R_t - V(s_t)) \nabla_\theta \log \pi_\theta(a_t | s_t)

The term At=Rt−V(st)A_t = R_t - V(s_t) is the advantage: how much better this action was compared to the average. Actions better than expected get reinforced; worse-than-expected get suppressed.

Connection to martingales: The advantage AtA_t is a martingale difference sequence—its conditional expectation given sts_t is zero. This orthogonality (cf. Martingales) is precisely why the baseline reduces variance without introducing bias.

5. From REINFORCE to Actor-Critic

MethodGradient estimateVarianceBias
REINFORCER(τ)∇log⁡πR(\tau) \nabla \log \piVery highNone
+ Baseline(Rt−b)∇log⁡π(R_t - b) \nabla \log \piHighNone
Actor-Critic(rt+V^(st+1)−V^(st))∇log⁡π(r_t + \hat{V}(s_{t+1}) - \hat{V}(s_t)) \nabla \log \piLowSome (from V^\hat{V})
GAE (λ\lambda)A^tλ∇log⁡π\hat{A}^\lambda_t \nabla \log \piTunableTunable

Actor-Critic replaces the Monte Carlo return RtR_t with the one-step TD error δt=rt+γV^(st+1)−V^(st)\delta_t = r_t + \gamma \hat{V}(s_{t+1}) - \hat{V}(s_t). This introduces bias (because V^\hat{V} is approximate) but drastically reduces variance (only one step of randomness instead of the full trajectory).

GAE (Generalized Advantage Estimation) interpolates between the two extremes via a parameter λ∈[0,1]\lambda \in [0,1]:

  • λ=1\lambda = 1: Full Monte Carlo return (unbiased, high variance)
  • λ=0\lambda = 0: One-step TD (biased, low variance)

This is exactly the bias-variance tradeoff in a new guise: more bootstrapping (lower λ\lambda) reduces variance at the cost of bias from the value function approximation.

The Takeaway

Policy gradient is built on a single algebraic trick—the log-derivative identity—that converts an intractable gradient-through-sampling into a tractable expectation. Everything that follows (baselines, advantages, actor-critic, GAE) is variance reduction: the ongoing struggle to extract a clean learning signal from the inherent noise of sampled trajectories. The tools are orthogonality (baselines), bootstrapping (TD learning), and the bias-variance tradeoff—the same ideas that recur across all of statistical learning.