Original Feishu document · Source revision 23
💡
Mechanism lesson: Offline RL no longer interacts with the environment online and uses only a fixed dataset. It aims to improve the policy using outcome signals from successes and failures, but it must address out-of-distribution actions, value overestimation, and insufficient data coverage.
Learning Objectives
After completing this lesson, you should be able to explain the differences among Offline RL, behavior cloning, and off-policy RL; derive out-of-distribution overestimation; understand behavior constraints, conservative Q-learning, Advantage-weighted regression, and IQL; and design quality stratification and fair ablations for offline robot data.
1. Problem Definition
Fixed dataset:
Read as: The offline dataset contains N transition samples, each consisting of a state, action, reward, next state, and termination flag.
Derivation: This defines fixed experience data. During training, data may only be read from D; no new actions can be executed to query their actual outcomes.
The data is generated by behavior policy . The objective is to learn a new policy , but the actual outcomes of out-of-distribution actions can no longer be queried from the real environment.
| Method | Uses action labels | Uses rewards | Can actively explore |
|---|---|---|---|
| Behavior cloning | Yes | No | No |
| Offline RL | Yes | Yes | No |
| Online RL | Optional | Yes | Yes |
2. Overestimation of Out-of-Distribution Actions
Q-learning target:
Read as: The Q-learning target equals the current reward plus, if the transition is nonterminal, the discounted maximum target Q-value at the next state.
Derivation: This is a single-sample estimate based on the Bellman optimality equation. The danger in the offline setting is that maximization searches over actions not covered by the data.
Maximization selects the action with the highest estimated Q-value, but the highest estimate often comes from an action never observed in the data. Let the error be:
Read as: The estimated Q-value equals the true optimal action value plus the function-approximation error epsilon.
Derivation: This is the definition of the error decomposition. Errors in data-dense regions can be corrected through supervision, whereas out-of-distribution actions lack ground-truth labels and may have larger, systematically biased errors.
Then:
Read as: Even if the Q-error for each action is individually zero-mean, taking the maximum before the expectation still yields a value no smaller than the maximum true Q-value.
Derivation: The maximum is a convex function. By Jensen's inequality, the expected maximum is no smaller than the maximum of the expectations; the optimizer tends to select actions whose values happen to be inflated by positive errors.
The Actor then moves further toward these spuriously high-valued actions, creating an “overestimation → further out of distribution → greater overestimation” feedback loop.
💡
Interactive Validation|Bellman, Value, and Offline RL Lab
Adjust the discount factor, proportion of out-of-distribution actions, and conservative penalty to observe the relationship between value backups and OOD overestimation.
Bellman, Value, and Offline RL Lab
:::3. Behavior Constraints
Constrain the new policy so that it does not deviate too far from the behavior policy:
Read as: At states in the dataset, maximize the Q-value of actions from the new policy while requiring the average distance between the new policy and the behavior policy that generated the data to be no greater than epsilon.
Derivation: This is constrained policy improvement: the value term provides the direction of improvement, while the distributional distance restricts actions to a neighborhood that can be validated by the data. An overly tight constraint degenerates into imitation, whereas an overly loose constraint reintroduces OOD risk.
The distance can be defined using KL divergence, MMD, a behavior model, or an action decoder. An excessively strong constraint degenerates into behavior cloning, whereas an excessively weak constraint again allows OOD actions to be selected.
4. Conservative Q-Learning
CQL assigns lower values to actions outside the dataset. Its abstract objective is:
Read as: In addition to the TD loss, CQL increases the log-sum-exp cost over all actions and offsets it by the mean Q-value of actions in the dataset to provide an anchor.
Derivation: Log-sum-exp is a smooth maximum. Minimizing it suppresses high Q-values that might be selected by the policy, while subtracting the Q-values of dataset actions prevents actions supported by the data from being suppressed as well. In continuous action spaces, this term is usually approximated by sampling from multiple action distributions.
It penalizes high Q-values for policy actions while increasing the relative value of dataset actions, thereby producing a conservative lower bound. Excessive conservatism may overlook rare but effective actions in the data.
5. Advantage-Weighted Regression
Rather than having the Actor directly maximize arbitrary Q-values, weighted imitation is performed on actions in the dataset:
Read as: AWR performs weighted maximum likelihood on dataset actions, with higher-weight actions exerting greater influence on policy fitting.
Derivation: Projecting the target distribution obtained through KL-regularized policy improvement back onto the parameterized policy is equivalent to performing Advantage-weighted cross-entropy regression on dataset actions.
Read as: The action weight is the exponential of the Advantage divided by the temperature beta.
Derivation: The KL-regularized optimal policy reweights the behavior policy exponentially according to Advantage. A smaller beta concentrates more strongly on high-Advantage actions; in practice, weights are often clipped to control variance.
Because it learns only from actions that appear in the dataset, it reduces OOD risk. RECAP's Advantage conditioning is closely related to this idea.
6. Implicit Q-Learning
IQL does not maximize Q over policy actions. Instead, it learns the state value through expectile regression:
Read as: IQL performs tau-expectile regression on the residual between the Q-value of a dataset action and the state value V.
Derivation: Unlike explicitly maximizing over actions, expectile regression uses an asymmetric squared loss to move V toward the upper portion of the value distribution of dataset actions, thereby forming a high-value baseline using only data-supported actions.
Expectile loss:
Read as: The expectile loss applies different weights to positive and negative residuals and then multiplies them by the squared residual.
Derivation: When u is positive, the weight is tau; when it is negative, the weight is one minus tau. A tau greater than one-half places more emphasis on samples whose Q-values exceed V, pushing V toward the upper expectile.
A larger moves V closer to the high-value portion of the dataset actions. Then:
Read as: The IQL Advantage is the Q-value of a dataset action minus the expectile value baseline for that state.
Derivation: V represents an upper baseline for the value distribution of dataset actions. A positive Advantage therefore identifies actions that are relatively better and that actually appear in the dataset.
Read as: The IQL Actor refits the actions in the offline dataset using exponential Advantage weights.
Derivation: As in AWR, this step performs policy improvement only over dataset actions and does not allow the Actor to directly search for arbitrary high-Q actions. Practical implementations usually clip the exponential weights.
7. Data Quality and Coverage
| Data type | Value | Risk |
|---|---|---|
| Expert successes | High-quality actions | Narrow state coverage |
| Autonomous successes | Matches the current policy distribution | Limited diversity of success modes |
| Failure trajectories | Covers bad states and boundaries | Lacks correct recovery actions |
| Human corrections | Directly provides recovery behaviors | Intervention-selection bias |
| Random exploration | Broad coverage | Unsafe and inefficient on real robots |
For Offline RL, more heterogeneous data is not necessarily better. Task phases, outcomes, takeovers, and control failures must be recorded to avoid treating system anomalies as evidence of policy value.
8. Sequence Models and Return Conditioning
Decision Transformer-style methods condition on the desired return:
Read as: A return-conditioned sequence model predicts the current action from the return-to-go up to the current step, the state history, and previous actions.
Derivation: After interleaving return-to-go, states, and actions into a sequence, causal sequence modeling learns the conditional action distribution. High-return conditions beyond the support of the data lack reliable supervision.
This reframes RL as sequence modeling, but a high-return condition is reliable only when the corresponding behavior exists in the dataset. Requesting a return beyond the support of the data may produce unpredictable actions.
9. Connection to Generative VLAs
An offline Critic can:
- Score and filter VLA trajectories.
- Rerank Flow/Diffusion action candidates.
- Control the policy through Advantage conditioning.
- Assign weights to post-training data.
The Critic and VLA must use consistent definitions of states, actions, and Action Chunks; otherwise, value labels cannot be aligned correctly.
10. Minimal Experiments and Fair Comparisons
- Using the same fixed dataset, compare BC, AWR/IQL, and CQL.
- Holding the algorithm fixed, progressively add success, failure, and correction data.
- Report dataset action support and the policy's OOD distance.
- Evaluate value calibration and bin results by actual success rate.
- Test on novel initial states, objects, and dynamics.
- Report safety, collisions, and human takeovers together.
11. Failure Modes and Diagnostics
| Failure mode | Diagnostic evidence | Mitigation |
|---|---|---|
| OOD Q-overestimation | Policy actions are far from the data; predicted Q-values are high but actual returns are low | Conservative Q-learning, behavior constraints, and support and calibration evaluation |
| Excessive conservatism | The policy nearly copies the behavior policy and ignores a small number of high-value actions | Sweep regularization strength, stratify by data quality, and compare against BC |
| High-return condition outside support | Decision Transformer requests returns absent from the dataset | Limit the conditioning range, report training-return support, and apply OOD detection |
| Contamination by failure data | System failures, latency, or sensor anomalies are mistaken for action value | Failure labels, time synchronization, and stratification by failure cause |
| State aliasing | The action-return distribution for the same image is multimodal and cannot be calibrated | History-conditioned Critic, state estimation, and contact-phase labels |
| Confounding from added data | Offline RL improvements coincide with the use of more or higher-quality data | Compare BC, CQL, IQL, and AWR on a fixed dataset |
12. Exercises
- Prove why the max operation produces overestimation bias.
- Explain the bias–improvement trade-off of behavior constraints.
- Derive Advantage-weighted regression.
- Use one-dimensional data to demonstrate how CQL penalizes out-of-distribution actions.
- Compare IQL with explicit Actor maximization of Q.
- Design a RECAP control experiment using equal amounts of data.
13. Paper Map and Evidence Boundaries
| Work | Paper facts | Authors' interpretation | Course assessment |
|---|---|---|---|
| BCQ / BEAR | Constrain policy actions through batch-constrained action generation and distributional distance, respectively | Constraining actions near the support of the behavior policy can reduce offline extrapolation error | The quality of the behavior model and the distance metric can themselves become new sources of error |
| CQL | Applies conservative regularization to high Q-values over a broad range of actions while anchoring the values of dataset actions | Learning conservative values can provide guarantees for offline policy improvement | The degree of conservatism, action coverage, and actual gains relative to BC should be reported; training Q-values alone are insufficient |
| IQL | Learns state values using expectile regression and performs Advantage-weighted regression on dataset actions | Stable offline learning can be achieved without explicitly querying actions outside the dataset | It can only recombine capabilities already present in the data; missing skills and recovery states still require new data |
| Decision Transformer | Conditionally models return-to-go, states, and actions as a causal sequence | Sequence modeling can unify conditional behavior prediction for Offline RL | High-return conditions must lie within training support, and sequence likelihood cannot replace real closed-loop validation |
| AWR / AWAC | Performs policy regression on dataset actions using exponential Advantage weights | Weighted supervised learning can achieve stable and conservative policy improvement | The weights depend heavily on Critic calibration and temperature; a small number of erroneously high-Advantage samples must not be allowed to dominate training |

