Original Feishu Document · Source Revision 16
💡
Mechanism lesson: Video generation models can predict future appearances, object changes, and visual subgoals, but “looking realistic” does not imply “physically executable.” This lesson distinguishes action-free video prediction, action-conditioned video models, visual subgoals, and video-based planning, and establishes evaluation criteria for physical consistency and control gains.
Learning Objectives
After completing this lesson, you should be able to distinguish video prediction from action-conditioned world models; explain pixel likelihood, latent video, and diffusion video; construct visual subgoals; analyze cameras, scale, contact, and uncontrollable backgrounds; and design an evidence chain from video quality to robot control gains.
1. Which Conditional Distribution Does a Video Model Learn?
1.1 Action-Free Video Prediction
How to read it: Given the observation history up to the present, the model predicts the joint distribution of future observations from the next frame through frame T.
Derivation: This is a conditional distribution over a purely observational sequence. Actions, camera motion, and external events in the future are all marginalized within the distribution. It therefore learns correlations and motion priors, but cannot by itself answer what the intervention outcome of a particular robot action would be.
It learns the observed natural temporal evolution, with actions and external factors implicit in the video. It is suitable for learning motion priors, but cannot answer what would happen if the robot selected a different action.
1.2 Action-Conditioned Video Models
How to read it: Given the observation history and the action sequence from t through T-1, the model predicts the distribution of future observations from t+1 through T.
Derivation: Each action a_k affects the next observation o_{k+1}, so the T-t future frames correspond to T-t actions. Only when the robot actively executes the actions and the data provide sufficient coverage does this conditional model more closely approximate action-conditioned dynamics suitable for counterfactual comparison.
Actions become intervention variables, allowing the model to compare the visual futures produced by different actions and making it closer to a world model.
1.3 Goal-Conditioned Video Generation
How to read it: Given the current observation history and task goal g, the model generates a possible future video that satisfies the goal.
Derivation: Conditioning on language or a goal image can concentrate generated outputs toward the goal distribution, but the actions are marginalized. The distribution therefore expresses “what kind of process might occur,” without guaranteeing that the current robot has a control sequence capable of realizing it.
Generating “the process that should occur” from language or a goal image can provide a high-level plan, but without action conditioning, it does not guarantee that the robot can realize that plan.
2. Why Is Pixel Space Difficult?
Images are high-dimensional, and the future can unfold in many different ways. Pixel MSE:
How to read it: The pixel loss is the expected per-pixel squared difference between the next ground-truth image and the predicted image.
Derivation: If, conditioned on the history, each pixel follows an independent Gaussian distribution with fixed variance and a mean given by the predicted image, maximum likelihood is equivalent to minimizing this MSE. A multimodal future is averaged into a blurry image by the conditional mean.
It averages across multiple futures, producing blurry images. Generative video models use latent representations, VAEs, Diffusion, or Flow to represent multimodal futures, but computational cost and physical consistency remain challenges.
3. Latent Video Models
First encode the frames:
How to read it: The encoder E_psi maps the observation at frame t to the latent representation z_t.
Derivation: This is a deterministic encoding definition. It replaces pixels with a lower-dimensional representation for generation and planning. Whether the representation preserves depth, contact, and object identity depends on the training objective rather than dimensionality reduction itself.
Generate the future in latent space:
How to read it: Given the latent history, future action sequence, and task goal, the model predicts the joint distribution of future latent video trajectories.
Derivation: This projects the conditional distribution over pixel videos into latent space through the encoder while preserving the one-to-one correspondence between future frames and action intervals. If goal conditioning is not used, g can be removed from the formula.
The result can then be decoded or used directly for planning. Latent representations can reduce computation and allow the model to ignore irrelevant textures, but they may also discard contact, depth, and small-object state.
| Space | Advantages | Risks |
|---|---|---|
| Pixels | Visualization and general-purpose supervision | High-dimensional; backgrounds consume capacity |
| Visual latent space | Efficient and generative | Physical variables are not explicit |
| Object state | Suitable for planning and causal reasoning | Depends on detection and tracking |
| Keypoints/trajectories | Motion is explicit | Insufficient contact and appearance information |
4. Visual Subgoals
The high-level model generates a future goal image , and the low-level policy learns:
How to read it: Based on the current observation and visual subgoal image, the low-level policy produces a conditional distribution over the current action.
Derivation: This is the visual form of a goal-conditioned policy: it concretizes a language goal as a desired scene, allowing the policy to learn to reduce the controllable difference between the current image and the goal image.
Visual subgoals contain more specific spatial layouts than language, but they must satisfy the following conditions:
- They must be consistent with the identities of objects in the current scene.
- Object positions and poses must be reachable.
- They must not violate occlusion, collision, or contact constraints.
- The low-level policy must be able to estimate the distance to the subgoal and complete it.
The visual subgoals in π0.7 can be understood along this cross-cutting path: the primary model remains a VLA, but it uses generated future frames as prompts.
5. From Video Candidates to Planning
Given multiple candidate futures , a visual goal score can be computed:
How to read it: The score of candidate future i is the negative distance between its final image and the goal image, minus the infeasibility or risk cost over the entire trajectory.
Derivation: Minimizing the visual goal distance can be written as maximizing its negative. Adding a trajectory cost weighted by the coefficient lambda allows collisions, uncertainty, and unreachability to affect the ranking. This score is a designed objective and is not equivalent to the true task reward.
measures distance to the goal, while measures collision, uncertainty, or infeasibility. The problem is that visual similarity does not imply task success: a cup may appear to be on a shelf while actually floating or resting unstably.
6. Physical Consistency
| Type of consistency | Check |
|---|---|
| Object permanence | Objects must not appear, disappear, or deform without cause |
| Geometry | Depth, occlusion, and collision must be plausible |
| Motion | Velocity and acceleration must be continuous |
| Contact | Object changes must have plausible contact causes |
| Embodiment | Robot poses must be reachable and satisfy joint limits |
| Causality | Changing the action should systematically change the future |
Object tracking, depth, keypoints, action reconstruction, or dynamics constraints can be added, but the system must ultimately be validated through execution on a real robot.
7. Human Videos and Robot Videos
Human videos provide rich tasks and interactions but lack robot actions. They can be used for:
- Pretraining video representations.
- Learning high-level events and subgoals.
- Learning object trajectories and affordances.
- Generating visual plans for robots.
Robot videos provide action conditioning and proprioceptive state, enabling the training of causal dynamics. Joint training should align objects and events rather than forcibly aligning human hands and robot arms frame by frame.
8. What Video Diffusion / Flow Generates
Video Diffusion learns to denoise corrupted latent representations:
How to read it: First, at diffusion noise time tau, the ground-truth video latent is mixed with Gaussian noise. A network is then trained to predict the added noise from the noisy latent, the noise time, and condition c.
Derivation: The forward process defines a known Gaussian perturbation, and the reverse denoising process can be trained using a noise-prediction parameterization. Here, tau refers specifically to diffusion noise time and must not be confused with video frame time t.
How to read it: Along a linear interpolation path, the Flow model learns the velocity vector that points from a noise video latent toward the ground-truth future video latent.
Derivation: Differentiating the straight-line path with respect to tau yields the constant vector “data minus noise,” making it possible to train the conditional velocity field by regression. During inference, integrating this velocity field transports noise samples to the future-video distribution.
Both Diffusion and Flow can represent multimodal futures, but planning requires repeated sampling, long-video generation, and candidate evaluation, so achieving real-time performance is generally more difficult than with direct action generation.
9. Major Failure Modes
- The generated future is visually plausible, but the required robot actions are unreachable.
- The model primarily predicts background and camera motion.
- Contact changes occur without a mechanical cause.
- The video goal is inconsistent with the low-level policy’s training distribution.
- Long videos preserve semantics but gradually drift in object identity.
- The visual evaluator rewards “appearing complete” rather than actual stable completion.
10. Evaluation Levels
| Level | Metrics | Cannot substitute for |
|---|---|---|
| Visual quality | FVD, perceptual distance, human preference | Physical correctness |
| Prediction | Object trajectories, key events, contact accuracy | Planning value |
| Counterfactual reasoning | Differences in the future after changing actions | Real-world execution |
| Planning | Correlation between candidate rankings and true returns | Closed-loop success |
| Control | Real-robot success and recovery rates | Safety and generalization |
11. Minimal Experiment
- Train an action-free video model and an action-conditioned video model.
- Provide different actions from the same initial state and check whether the predicted futures are distinguishable.
- Replace purely pixel-based metrics with object-trajectory evaluation.
- Generate visual subgoals and have the same low-level policy execute them.
- Add unreachable goal images and test the gating mechanism.
- Compare the correlations among visual quality, candidate ranking, and real-world success rate.
12. Exercises
- Write down three video conditional distributions and explain their differences.
- Design a counterexample that is visually plausible but physically incorrect.
- Explain why a video world model requires action conditioning to support counterfactual planning.
- Define reachability, completion, and safety checks for visual subgoals.
- Compare the advantages and disadvantages of pixel space, latent space, object state, and keypoints for planning.
- Explain why π0.7 visual subgoals connect VLA, world models, and hierarchical planning.
13. Paper Landscape and Evidence Boundaries
Video metrics, planning metrics, and real-world control metrics must be reported separately by level. The works below address different problems and should not be ranked solely by generation quality or a single success rate.
| Work | Facts from the paper | Authors’ interpretation | Course assessment |
|---|---|---|---|
| Visual Foresight | Uses action-conditioned video prediction and visual MPC to select robotic pushing actions | Pixel-level future prediction can directly support goal-image control without manually specified rewards | Established the visual MPC paradigm; success in 2D pushing cannot substitute for evidence involving 3D contact, force, and stable grasping |
| SV2P | Uses stochastic latent variables to represent uncertainty in video futures and evaluates them on robot interaction videos | Stochastic video models can cover multiple futures better than deterministic prediction | Diversity has control value only when it improves candidate ranking and real-world execution |
| RoboNet | Aggregates multi-robot, multi-view interaction data to study cross-robot video prediction and control | Large-scale heterogeneous video can improve the transfer of visual dynamics | A cross-cutting case between Paths E and B; differences in actions, embodiment, and cameras must be modeled explicitly |
| UniPi | Uses text-conditioned video generation to produce visual plans, which are then converted into actions through interfaces such as inverse dynamics | Video can serve as a general planning representation across tasks and embodiments | The high-level planning space is highly promising, but executability depends on low-level action recovery, state feedback, and replanning after failure |
| π0.7 visual subgoals | A general-purpose VLA can use prompts such as visual subgoals to improve task controllability | Generated intermediate visual states can facilitate hierarchical execution | Its primary classification remains VLA; this lesson examines only the cross-cutting world-model mechanism in which visual futures serve as conditions and subgoals |
| Modern embodied world models | UniSim, Genie, Cosmos, and related systems scale toward interactive environments or large-scale video generation | Larger generative models may become environments for embodied training and evaluation | A detailed comparison is provided in B6; whether they improve real-robot closed-loop performance remains the ultimate threshold |

