Skip to content

Original Feishu Document · Source Revision 20

💡

Mechanisms lesson: This lesson examines how world models can be used not only for planning, but also to generate imagined trajectories within the model and use them to train an Actor, Critic, or skill policy. The focus is on the gradient paths of imagined rollouts, model bias, and the trade-off between short and long rollouts.

Learning Objectives ​

After completing this lesson, you should be able to distinguish real rollouts from imagined rollouts; formulate the representation, dynamics, reward, value, and policy objectives of Dreamer-style methods; explain whether gradients pass through the world model; analyze how model errors contaminate the policy; and design experiments that combine real and imagined data for training.

1. Why Learn Within a Model ​

Real-robot interaction is expensive, slow, and dangerous. If a world model can predict latent states:

Interpretation: Given the current latent state and action, the world model produces a conditional distribution over the next latent state.

Derivation: This is the state-transition interface of the imagined environment. It comes from latent dynamics learned on real trajectories rather than from newly collected real states.

It can then generate rollouts internally:

Interpretation: At step k of an imagined rollout, the model advances from its current predicted state to the next state using the Actor's action.

Derivation: In the deterministic case, repeatedly composing f_phi produces a multi-step trajectory. A stochastic model must instead sample or propagate a distribution at each step, causing both error and uncertainty to accumulate along the rollout.

The policy can then be trained using predicted rewards:

Interpretation: The Actor's H-step imagination objective is the discounted sum of the first H predicted rewards plus the terminal value of the latent state at step H, with the expectation taken over the initial state, policy actions, and model stochasticity.

Derivation: The infinite-horizon return is truncated at H: the model explicitly predicts the first H steps, and the Critic approximates the remaining tail. Summing through H-1 includes exactly H rewards.

This does not create data from nothing; it uses model assumptions to interpolate between observed data. Model errors can be repeatedly amplified by imagination-based training.

2. Components of an Imagined Rollout ​

ModuleMathematical objectRole
RepresentationApproximate posterior q: infers the latent state from observation historyMaps observations to a decision state
DynamicsLatent transition distribution p: predicts the state after an actionPredicts the state after an action
RewardReward predictor: estimates task feedback for the current state and actionPredicts task evaluation
TerminationTermination predictor: estimates whether the trajectory continuesPredicts whether the task has ended
ActorActor policy: selects actions from latent statesSelects actions in imagination
CriticCritic value function: estimates future returnEstimates future return

3. Learning the Critic on Imagined Trajectories ​

Imagined TD target:

Interpretation: The one-step value target equals the predicted immediate reward plus the product of the nontermination probability, discount factor, and target value of the next state.

Derivation: Bellman recursion decomposes the long-term return into the current reward and the return from the next state. When the termination probability is one, bootstrapping stops. The target-network parameters bar xi are used to stabilize training.

Interpretation: The lambda-return interpolates between the one-step value target and a longer recursive return according to lambda, while the continuation probability suppresses returns after termination.

Derivation: Geometrically weighting returns with different n-step horizons yields this recursive form. When lambda is zero, it reduces to one-step TD; as lambda approaches one, it relies more heavily on long imagined returns.

Critic loss:

Interpretation: The Critic minimizes the mean squared error between its own value prediction and the stop-gradient lambda-return target.

Derivation: Treating the return as a regression target yields the squared loss. The stop-gradient operator sg prevents the Critic from reducing the error by changing the target branch rather than the value prediction.

If the world model makes incorrect predictions in a particular region, the Critic will treat the model's fantasies as real future returns. This can be mitigated by limiting the imagination horizon, incorporating real next states, or using ensemble uncertainty.

4. Gradient Paths for the Actor ​

4.1 Policy Gradients That Do Not Pass Through the Model ​

Treat the world model as an environment and use imagined states and the Advantage:

Interpretation: The policy gradient multiplies the gradient of each action's log-probability by that action's advantage over the average, then sums over the imagined horizon and takes the expectation.

Derivation: This follows from applying the log-derivative trick to the trajectory probability. When the world model is not differentiated through, it only generates sampled states, while the Advantage serves as a return weight that reduces variance.

The advantage is simplicity; the disadvantage is that model errors affect the policy only indirectly through the data distribution.

4.2 Pathwise Optimization Through the Model ​

If the dynamics are differentiable, return gradients can pass through the model:

Interpretation: The Actor parameters affect return both directly through the current action and indirectly by changing the next state, which recursively influences all subsequent returns.

Derivation: This uses total derivatives and the chain rule for a reparameterizable policy and differentiable dynamics. The second line explicitly preserves the state recursion, ensuring that the indirect effects of earlier actions on later states are not omitted.

This approach is sample-efficient, but small systematic errors in the model may directly become erroneous policy gradients.

Formula Visualization|How Return Gradients Pass Through Actions and the World Model ​

Course Whiteboard

5. Combining Real and Imagined Data ​

Training approachProportion of real dataAdvantagesRisks
Real only100%Not contaminated by model fantasiesHigh interaction cost
Short imaginationReal initialization, short rolloutLower bias and effective data augmentationLimited long-horizon value
Long imaginationSparse real initializationCovers long-horizon decisionsAccumulation of model error
Uncertainty gatingSelects data according to model confidenceExpands data within trusted regionsDifficult confidence calibration

6. Abstract Workflow of Dreamer-Style Methods ​

  1. Train the representation, dynamics, reward, and termination models on real trajectories.
  2. Encode latent states from real observations.
  3. Use the current Actor to generate imagined trajectories in latent space.
  4. Train the Critic using imagined rewards.
  5. Train the Actor to select actions that increase imagined return.
  6. Execute the new policy in the real environment and collect data to correct the world model.

The core of Dreamer is not to “generate realistic videos,” but to perform differentiable or approximately differentiable decision learning in a latent space that is sufficiently useful.

7. Three Ways Model Bias Propagates ​

BiasPropagation mechanismDiagnostics
State-representation biasThe Actor and Critic observe an incorrect stateOcclusion tests, history length, representation probes
Dynamics biasThe prediction error at each step becomes the input to the next stepMulti-step open-loop predictions versus real rollouts
Reward biasThe policy pursues a spurious objective defined by the modelCorrelation between model-predicted return and real success rate

8. Combining Imagination-Based Learning with VLA ​

A VLA can serve as the Actor, while the world model generates latent futures and the Critic evaluates Action Chunks:

Interpretation: The value of action chunk A_t is the expected sum of the H-step rewards predicted by the world model and the terminal value. The action chunk contains exactly H actions, from t through t+H-1.

Derivation: This extends a single-step action value to a fixed-length action chunk and marginalizes over futures generated by the model. The resulting score can be used to rerank candidate Action Chunks or provide value guidance.

This can support candidate-action reranking, value conditioning, guidance for generative policies, or low-level skill selection. It does not mean that the VLA automatically becomes a world model; the additional prediction target and training signal must be explicitly identified.

9. Minimal Experiment ​

  1. Fit one-dimensional dynamics with friction using real trajectories.
  2. Train a real-data-only Actor and an imagined Actor separately.
  3. Sweep the imagination horizon.
  4. Artificially increase dynamics bias and observe policy performance.
  5. Add ensemble uncertainty gating.
  6. Compare model-predicted return, real return, and Critic calibration.

10. Paper Landscape and Evidence Boundaries ​

The key question in imagination-based learning is not how many trajectories are generated, but in which state-action regions those trajectories are trustworthy, how the Actor's gradients depend on the model, and whether the method ultimately improves real-world closed-loop performance.

WorkPaper factsAuthors' interpretationCourse assessment
World ModelsFirst trains a visual latent-variable model and recurrent dynamics, then trains the controller within the modelCompressed dreams can support policy learningEstablished a clear paradigm, but early results on visual games do not directly demonstrate viability for robots with complex contact dynamics
MBPOGenerates short model rollouts from real states and mixes them with real data to train the policyShort rollouts can balance model bias against data augmentationAn important baseline for evaluating the imagination horizon; robotic applications must additionally address partial observability and state-estimation error
Dreamer seriesTrains the value function and Actor using lambda-returns on latent imagined trajectories from an RSSMDifferentiable latent imagination can efficiently learn long-horizon behaviorThe central approach of this lesson; model loss, value calibration, and real-task gains should all be reported rather than only the final score
TD-MPC2Combines latent dynamics, value learning, a policy prior, and online short-horizon planningModel learning and local planning can scale togetherIt is not equivalent to a purely imagined Actor and is well suited as a comparison for hybrid imagination training and online planning
RECAP / robot experience learningUses successes, failures, and corrections from real deployments to continue improving the policyReal experience can compensate for the limitations of offline imitationProvides an essential control for imagined data: without real-world feedback, model bias cannot be reliably detected and corrected

11. Exercises ​

  1. Distinguish the statistical sources of imagined states and real observations.
  2. Write the TD target with a termination probability.
  3. Explain why gradients passing through the world model may amplify model errors.
  4. Design three controlled experiments for short imagination, long imagination, and uncertainty gating.
  5. Explain what each module learns in a joint system consisting of a VLA, world model, and Critic.
  6. Map the abstract Dreamer workflow onto a real-robot task.

Article text is licensed under the Apache License 2.0