Skip to content

Original Feishu document · Source revision 26

💡

Mechanisms lesson: This lesson builds on the Bellman equation to derive TD, Q-learning, policy gradients, Advantage, and Actor-Critic, and explains the practical implications of these formulas for long-horizon, partially observable robotic tasks with high-dimensional continuous actions.

Learning Objectives ​

After completing this lesson, you should be able to derive the Bellman expectation and optimality equations; distinguish among Monte Carlo, TD, SARSA, and Q-learning; derive the policy gradient and baseline; explain the two sources of error in Actor-Critic; and design training and calibration experiments for continuous-action robotic tasks.

1. MDPs and the Policy Objective ​

An MDP is defined by . The policy generates actions:

Reading: In state s_t, action a_t is sampled from a policy distribution parameterized by theta.

Derivation: A stochastic policy maps states to probability distributions over actions; continuous control commonly uses Gaussian or other reparameterizable distributions.

The optimization objective is:

Reading: The policy objective is to maximize the expected sum of T discounted rewards over length-T trajectories generated by the policy.

Derivation: The policy and environment jointly define a trajectory distribution. Taking the expectation of trajectory returns under this distribution yields the optimization objective for parameter theta.

Robotic tasks are usually POMDPs, so the actual policy is conditioned on history or a latent state rather than the complete .

2. Bellman Expectation ​

Reading: The state value equals the expectation of the immediate reward plus the discounted value of the next state, after selecting an action according to the policy and transitioning according to the environment.

Derivation: Decompose the return into the current reward and the subsequent return, then marginalize over the current action and next state to obtain the Bellman expectation equation.

Reading: The action value equals the immediate reward plus the discounted expected Q-value when actions continue to be selected according to the current policy in the next state.

Derivation: Because the first action a is fixed by conditioning, only the next state and subsequent policy actions need to be integrated over.

These form a set of self-consistent equations: value is jointly defined by the one-step outcome and the value of the next state.

3. Monte Carlo and TD ​

MethodTargetBiasVariance
Monte CarloFull-trajectory return (Monte Carlo target)LowHigh
TD(0)One-step reward plus next-state value (TD target)Has bootstrap biasLower
n-stepn-step rewards plus terminal valueIntermediateIntermediate
TD(λ)Weighted multiscale n-step returnsAdjustableAdjustable

TD error:

Reading: The TD error is the sum of the one-step reward and the nonterminal next-state value, minus the current value prediction.

Derivation: Subtract the current estimate from the one-step Bellman target to obtain the residual; d_t prevents bootstrapping after termination, while the target network bar phi provides a more stable estimate of the next-state value.

It is both the value-update error and an approximation of how good an action is relative to expectations.

4. SARSA and Q-Learning ​

SARSA uses the action actually taken next:

Reading: The SARSA target uses the value of the action actually sampled by the current behavior policy in the next state.

Derivation: This is a single-sample estimate of the Bellman equation for policy pi. Because the next action comes from the same policy, this is an on-policy update.

Q-learning uses the maximum-value action:

Reading: The Q-learning target uses the largest target Q-value among all actions in the next state.

Derivation: Replacing the expectation over next actions under a fixed policy with the maximum over actions yields a sample target for the Bellman optimality equation. The behavior data may therefore come from a policy different from the target policy.

SARSA is on-policy and learns the value of the current behavior policy; Q-learning is off-policy and approaches the optimal greedy policy. In high-dimensional continuous action spaces, is difficult to solve directly, requiring an Actor or an optimizer.

5. Policy Gradients ​

Using the log-derivative trick:

Reading: The gradient of the trajectory probability equals the trajectory probability itself multiplied by the gradient of its log probability.

Derivation: Because nabla log p equals nabla p divided by p, multiplying both sides by p gives the result. This step rewrites the difficult gradient of a product of probabilities as a sum of log-probabilities.

When the environment dynamics are independent of the parameters:

Reading: At each step, the gradient of the action log-probability is weighted by the future return from that step onward, then summed along the trajectory and averaged in expectation.

Derivation: Apply the log-derivative trick to the trajectory expectation. Because the environment transitions do not depend on theta, only the gradients of the policy’s action probabilities remain. Using reward-to-go removes rewards that occurred before an action and are therefore unrelated to it.

The probabilities of high-return actions increase, while those of low-return actions decrease.

6. Baselines and Advantage ​

Subtracting a baseline that depends only on the state does not change the expected gradient:

Reading: At a fixed state, the action expectation of the policy log-probability gradient multiplied by any baseline that depends only on the state is zero.

Derivation: Factor b(s) out of the expectation. The remaining term is the sum of the gradients of all action probabilities, which is the gradient of the probability normalization constant 1 and is therefore zero.

Choose :

Reading: Advantage measures how much better action a performs than the policy’s average performance in state s.

Derivation: Subtracting V as a state baseline from Q preserves the relative differences among actions while reducing the variance of the policy gradient.

This yields a lower-variance policy gradient:

Reading: The Actor adjusts the probability of each action according to its Advantage: the probabilities of above-average actions increase, while those of below-average actions decrease.

Derivation: Replace Q in the policy gradient with Q minus the V baseline. The preceding equation proves that this replacement does not change the expectation.

7. Actor-Critic ​

The Critic minimizes:

Reading: The Critic minimizes the squared error between the current Q prediction and the stop-gradient TD target.

Derivation: The Bellman target serves as the regression label. Stopping the gradient prevents the label from moving through the target branch at the same time; a target network and double Q-learning can further reduce instability and overestimation.

The Actor maximizes:

Reading: The Actor minimizes negative Q, thereby making its own actions appear to have higher long-term value to the Critic.

Derivation: Maximizing expected Q is equivalent to minimizing its negative. If the Critic overestimates actions outside the data distribution, this objective will actively push the Actor into erroneous regions.

The Critic’s extrapolation error directly drives the Actor to select incorrect actions. This is a major source of instability in Actor-Critic methods for continuous control.

8. Continuous-Action Actors ​

Gaussian policy:

Reading: A Gaussian Actor transforms parameter-independent standard noise into an action using a state-dependent mean and standard deviation.

Derivation: Reparameterization isolates the randomness in epsilon, allowing the loss gradient with respect to the action to propagate to mu, sigma, and the policy parameters.

Reading: The deterministic policy gradient multiplies the Critic’s gradient with respect to the action by the gradient of the Actor’s output action with respect to its parameters.

Derivation: Apply the chain rule to the composite function Q(s,mu_theta(s)). A deterministic policy outputs actions directly. When robotic actions are bounded, they are often passed through tanh and scaled to the physical range; saturation weakens the gradient.

9. Entropy Regularization and SAC Intuition ​

Maximum-entropy objective:

Reading: The maximum-entropy objective rewards both task returns and the policy’s action entropy in each state, with alpha controlling the value assigned to stochasticity.

Derivation: This objective is obtained by adding an entropy reward to the standard return. It encourages the policy to retain multiple high-value actions, but excessive entropy on real robots manifests as jitter and unsafe exploration.

Entropy encourages exploration and action diversity. In robotics, excessive entropy may cause dangerous jitter, so action bounds, safety layers, and offline pretraining are required.

10. Partial Observability and Long-Horizon Credit Assignment ​

The same image may correspond to “just made contact” or “currently slipping.” The Critic requires history:

Reading: Under partial observability, the Critic evaluates the current action using h_t, the history of observations up to the current time and previous actions.

Derivation: When a single frame cannot distinguish velocity, contact phase, or hidden state, history provides the input for constructing an approximately sufficient state. An RNN, Transformer, or filter can perform this compression.

A final failure in a long-horizon task may have been caused by an early action. n-step returns, λ-returns, skill-level values, and stage rewards can alleviate this problem, but they may also introduce incorrect credit assignment.

11. Minimal Experiments ​

  1. Implement Monte Carlo, TD, and Actor-Critic on a two-dimensional continuous-control task.
  2. Compare different values of n-step and λ.
  3. Add observation occlusion and compare a single-frame Critic with a history-conditioned Critic.
  4. Make the Critic overestimate out-of-distribution actions and observe the Actor collapse.
  5. Add double Q-learning, a target network, and entropy, then ablate each component separately.

12. Failure Modes and Diagnostics ​

Failure ModeObservable SymptomsDiagnosis and Mitigation
Q overestimationPredicted Q continues to rise while actual returns stagnate or declineDouble Q-learning, target smoothing, and calibration against binned empirical returns
Actor exploits the CriticThe policy outputs actions rarely seen in the data and receives spuriously high Q-valuesAudit action coverage, apply behavior constraints, and use conservative value estimation
Action saturationtanh remains near its bounds for long periods, gradients approach zero, and execution produces large impactsLog pre-tanh values, rescale actions, and add smoothing and clipping
Entropy causes dangerous jitterThe policy continues random exploration during contact phasesState-dependent entropy, safety layers, offline pretraining, and staged exploration
Insufficient historyValues are multimodal for the same image and predictions are inaccurateAdd history and state estimation, and stratify evaluation by contact phase
Unstable reward scaleCritic gradients or entropy weights vary sharply across tasksReward normalization, automatic temperature tuning, and cross-task scale ablations

13. Research Landscape and Evidence Boundaries ​

WorkPaper FindingsAuthors’ InterpretationCourse Assessment
REINFORCEDirectly optimizes a stochastic policy using sampled returns and log-probability gradientsPolicy gradients can be estimated without a differentiable environmentIt is the theoretical starting point, but its high variance makes it an insufficient baseline for modern continuous-control robotics
GAEConstructs an Advantage estimator with adjustable bias and variance using exponentially weighted TD residualsIt can stabilize credit assignment in policy optimizationIt should be reported together with the horizon, termination handling, and Critic quality
DDPGApplies deterministic policy gradients, experience replay, and target networks to continuous actionsAn Actor can replace explicit maximization over continuous actionsIt is sensitive to Q overestimation and hyperparameters and should be compared against at least TD3 or SAC
TD3Uses twin Critics, delayed Actor updates, and target policy smoothingThese mechanisms can mitigate overestimation caused by function approximationImproved stability does not solve offline out-of-distribution problems; fixed-dataset settings still require conservative methods
SACUses a stochastic Actor, twin Q-functions, and a maximum-entropy objective for off-policy continuous controlEntropy regularization can improve sample efficiency and robustnessOn real robots, the evidence must account for exploration, safety layers, and action frequency rather than merely restating simulation returns

14. Exercises ​

  1. Derive the policy-gradient log-derivative.
  2. Prove that a state-dependent baseline does not change the expected gradient.
  3. Compare the data distributions of SARSA and Q-learning.
  4. Explain why continuous action spaces require an Actor.
  5. Design an experiment that distinguishes Actor errors from Critic errors.
  6. Explain how the Q-value of an Action Chunk should be defined.

Article text is licensed under the Apache License 2.0