Original Feishu Document · Source Revision 66
💡
Core change: π0.7 no longer relies solely on a single task description. Instead, it incorporates subtasks, visual subgoals, data quality, policy type, and control mode into the Prompt, allowing a generalist policy to be precisely steered.
1. Why Generality Does Not Equal Usability
A generalist policy may know how to complete a task but use the wrong speed, grasping strategy, or subtask order. In real-world deployment, users need to control “how to do it,” not merely “what to do.” π0.7 calls this capability steerability.
💡
Chapter 1 Interactive Explainer|Steerability Lab
Adjust task proficiency, Prompt alignment, and environmental disturbances to distinguish between “being able to complete the task” and “being able to complete it in the specified manner.”
Chapter 1 Interactive Explainer|Steerability Lab
:::2. Model Architecture
π0.7 has approximately 5B parameters and consists of a 4B VLM Backbone, a video-history encoder, and an approximately 860M-parameter Action Expert. The video-history encoder allows the model to use not only the current frame but also past actions and state changes.
💡
Chapter 2 Interactive Explainer|Architecture and Information-Flow Inspector
Toggle video history, visual subgoals, Metadata, and Knowledge Insulation to inspect Token and gradient paths.
Chapter 2 Interactive Explainer|Architecture and Information-Flow Inspector
:::3. Diversified Prompts
| Condition | What It Controls |
|---|---|
| Overall task language | Final objective |
| Current subtask | Semantic action for the next stage |
| Visual subgoal | Specific visual state to be reached |
| Episode Metadata | Data quality, policy source, speed, or behavioral style |
| Control Mode | Autonomous operation, high-level guidance, or human coaching |
💡
Chapter 3 Interactive Explainer|Multimodal Prompt Composer
Toggle the five Prompt categories individually to observe condition completeness, action ambiguity, and adaptation to missing conditions.
Chapter 3 Interactive Explainer|Multimodal Prompt Composer
:::4. How Visual Subgoals Are Generated
A lightweight world model conditionally generates future visual subgoals from the current observation, current subtask, and sample metadata:
How to read it: The future visual subgoal is sampled from a conditional world model parameterized by . Conditioned on the current observation , current subtask , and metadata , the model assigns probabilities to different candidate goal images.
Derivation: The same current observation and high-level plan may correspond to multiple reasonable goal images, so the goal generator models a conditional distribution rather than a single deterministic image; psi denotes the generator’s parameters.
The low-level policy then generates actions based on the difference between the current image and the goal image. Visual goals express pose, position, contact state, and environment layout more precisely than text.
4.1 Where Visual Subgoals in Training Samples Come From
First, a robot trajectory is segmented into clips with clearly defined subtasks. For a clip starting at time and ending at , the training sample can be abstracted as:
How to read it: A world-model training sample consists of four components: the observation at the start of the clip, the current subtask text , the Episode Metadata , and the supervised visual subgoal .
Derivation: The current observation, refined language instruction, mode variables, and goal image required for goal generation are collected into a training-sample tuple z, making it possible to define the goal dataset and loss uniformly.
The paper directly uses the final frame of a high-quality subtask clip as the ground-truth visual subgoal:
How to read it: If a clip represents “open the refrigerator door,” then the image at the end of the clip, in which the refrigerator door is open, is the supervised visual subgoal for that clip.
Derivation: Select the subtask completion time t_end in a successful trajectory and use the observation at that time as the supervised goal image for the current time. This yields goal labels that can be constructed automatically from video.
This does not mean there is only one correct image for a given condition. Different demonstrations may use different arm poses, contact locations, or object paths, so the data actually defines a conditional distribution:
How to read it: Even after the current scene, current subtask, and behavioral metadata have been specified, there may still be multiple reasonable future visual states. Each training clip provides one sample from this conditional distribution.
Derivation: Goal images in the training set are drawn from real successful trajectories filtered by the current observation, refined language instruction, and mode, so they can be regarded as samples from the corresponding conditional data distribution.
4.2 Expanding the Paper’s Objective into Conditional Flow Matching
The paper describes world-model training with the following high-level objective:
How to read it: Adjust the world-model parameters psi to minimize the average conditional Flow Matching loss over the high-quality subtask dataset.
Derivation: The engineering implementation does not directly maximize image density. Instead, it constructs paths from noise to goals and minimizes the velocity-field regression error. Under ideal conditions, the resulting marginal flow transports the noise distribution to the conditional visual-goal distribution.
To expand this objective, first sample a starting point from a simple noise distribution and randomly select a flow time:
How to read it: is Gaussian noise with the same shape as the goal-image latent; is an intermediate time sampled randomly between 0 and 1.
Derivation: Conditional Flow Matching requires a random noise starting point and a random path time. A standard Gaussian provides an easy-to-sample starting point, while uniform time sampling ensures that training covers the entire path from noise to goal.
Use the simplest linear probability path to connect the noise to the ground-truth visual subgoal:
How to read it: At , the state is pure noise; at , it is the ground-truth subgoal; at intermediate time , it is a linear mixture of the two.
Derivation: Linearly interpolating the noise and goal image with one minus tau and tau automatically satisfies the boundary conditions: noise at tau=0 and the ground-truth goal at tau=1.
Differentiate with respect to flow time to obtain the target velocity along this conditional path:
How to read it: Along this linear path, the state should move at every time in the direction of “ground-truth subgoal minus starting noise.” More general probability paths may have target velocities that vary over time.
Derivation: Differentiating the linear interpolation with respect to tau, while treating the noise and goal as constants within a single training sample, gives a constant target velocity equal to the goal image minus the noise.
4.3 What the World Model Actually Learns
Let the network take the current noisy intermediate state, flow time, and all conditions as inputs. The standard velocity-regression loss can be written as:
How to read it: Feed the intermediate state to the model and ask it to predict the direction in which that state should change at the current time. The closer the predicted velocity is to the ground-truth path velocity , the lower the loss.
Derivation: Every randomly sampled intermediate point has a known target velocity, g_star minus epsilon. Taking the expectation of the squared velocity-prediction error over these samples yields the conditional Flow Matching regression loss.
From the basic result for squared-loss regression, when the dataset is infinite and model capacity is sufficient, the optimal velocity field is the conditional mean velocity:
How to read it: When the model knows only the current position, time, and Prompt, it learns the average direction of motion across all training paths that may pass through that point. Different conditions produce different velocity fields and endpoint distributions.
Derivation: After fixing the intermediate position, time, and condition, the optimal regression function under squared loss equals the conditional expectation of the target velocity. This follows by setting the derivative of the mean-squared risk with respect to the prediction to zero.
💡
Key distinction: The world model does not directly regress a unique final frame. It learns a conditional velocity field, so different noise samples can generate multiple future visual states that satisfy the same subtask.
4.4 How a Visual Subgoal Is Generated from Noise at Inference Time
At inference time, Gaussian noise is sampled again as the initial state:
How to read it: Each time a visual subgoal is generated, a new random starting point is sampled. Different starting points allow the model to produce different but reasonable goal images.
Derivation: The ground-truth goal image is no longer available during inference, so an initial noise state x0 is resampled from the base distribution used during training and serves as the starting point of the goal-generation ODE.
The ordinary differential equation is then solved under the current observation, subtask, and metadata conditions:
How to read it: Starting from , repeatedly query the world model for the direction in which the image latent should currently be modified, then integrate forward using that velocity.
Derivation: Once the local conditional velocity has been learned, it is used as the right-hand side of an ODE and integrated from tau=0 to 1, transporting noise into the goal-image space along the learned flow.
After integration reaches the endpoint, it yields one visual-subgoal sample:
How to read it: The ODE endpoint is the generated subgoal image . After marginalizing over the randomness of the initial noise, all possible endpoints collectively form the model’s conditional distribution.
Derivation: If the conditional velocity field approximates the target probability path, the distribution of the ODE endpoint x1 is the goal generator’s distribution conditioned on the current observation, refined language instruction, and mode. One integration produces one goal sample.
4.5 Why Multi-View Visual Subgoals Are Needed
π0.7 uses multi-view subgoals rather than generating only one image:
How to read it: The visual subgoal consists jointly of goal images from camera views. The -th goal corresponds to the scene that camera is expected to observe in the near future.
Derivation: A multi-camera system must simultaneously constrain the goal state in every view, so the goal images from n views are arranged in a fixed order to form a joint goal variable.
| View | What It Constrains More Easily | Typical Ambiguity |
|---|---|---|
| Robot-base or environment camera | Object position, environment layout, and the open or closed state of doors or drawers | Local contact between the gripper and object is difficult to see clearly |
| Wrist camera | Gripper pose, contact location, local alignment, and grasping method | The field of view is narrow, making the global object state difficult to determine |
| Combined multi-view | Constrains both the object outcome and robot end-effector outcome | Geometric and semantic consistency must be maintained across views |
4.6 Why Visual Subgoals Make Action Learning Easier
Without a visual subgoal, the policy faces a highly underdetermined problem:
How to read it: “Open the refrigerator” does not specify which side of the handle the robot should grasp, what wrist pose it should use, or which local state it should reach in the near term.
Derivation: A conventional VLA directly predicts actions conditioned on the task and current visual input. This is the basic computation graph of a conditional policy; the arrow represents a learned mapping rather than a physical causal equation.
After a visual subgoal is added, the problem becomes closer to inverse dynamics:
How to read it: The visual subgoal first specifies “what the world should look like next,” after which the Action Expert infers the continuous control actions required to move from the current state to the goal state.
Derivation: After goal imagery is introduced, the policy is asked to generate an action chunk that advances the current state toward the goal state. This further narrows an open-ended task condition into a local state-transition problem.
From a probabilistic perspective, the visual goal can be understood as reducing uncertainty in the action distribution:
How to read it: Beyond the current observation and task text, knowing the subtask, visual goal, and policy metadata usually reduces the conditional entropy of reasonable actions, making the action-prediction problem more specific. This is a mechanistic interpretation, not a strict inequality directly proved by the paper.
Derivation: Conditional entropy satisfies the property that adding conditions cannot increase average uncertainty. Goal images, refined language instructions, and modes provide additional information, so the residual entropy of actions under these richer conditions is no greater than when only the observation and original task are given.
4.7 Noteworthy Values in the Official Training Setup
| Setting | Value in the Paper | Purpose |
|---|---|---|
| Samples with visual subgoals | Approximately 25% of each batch | Use visual goals to accelerate training while retaining the ability to operate without them |
| Drop subtask text when a visual goal is present | 30% | Allow visual goals to replace text when expressing finer-grained spatial states |
| Drop all Episode Metadata | 15% | Allow the model to work when metadata is unavailable at test time |
| Drop individual metadata fields | 5% each | Prevent the model from overrelying on any single speed, quality, or error label |
| Control Mode dropout | Not used | Always specify whether actions use joint or end-effector control |
💡
Interactive Explainer:Adjust subtask, speed, and quality conditions to observe how visual-goal mode probabilities and latent-space probability flows change; switch tabs to inspect the conditional distribution, Flow Matching rollout, and action-policy equations.
π0.7 Visual Subgoal and Flow Matching Explainer
:::5. How the High-Level Policy Automatically Prompts the Low-Level Model
The high-level semantic policy generates the next subtask instruction from the task, current observation, and subtask history. π0.7 can therefore switch among direct human control, automatic high-level planning, and visual-goal guidance.
5.1 How the Complete Prompt Enters the Low-Level Policy
The overall task, current subtask, visual subgoal, Episode Metadata, and control mode are combined into a context:
How to read it: describes the final task, describes the current semantic subtask, describes the visual state to be reached in the near term, specifies behavioral attributes such as speed, quality, and errors, and specifies joint or end-effector control mode.
Derivation: The original task, refined language instruction, goal, mode, and other control variables are arranged in an agreed order to form the unified condition C_t, allowing the downstream policy to refer only to this condition set.
The low-level policy in π0.7 learns a conditional distribution over future action chunks:
How to read it: The policy parameterized by assigns a conditional probability density to the next -step action sequence based on the most recent steps of video and proprioceptive-state history, together with the complete Prompt .
Derivation: The policy requires both the latest T steps of visual history to infer the current dynamics and C_t to specify the task, goal, and control mode. The action-chunk distribution is therefore conditioned on both.
The paper describes the overall VLA training objective using approximate conditional log-likelihood:
How to read it: Adjust the policy parameters so that the ground-truth action chunks in the training data receive higher conditional probability under the corresponding observation history and complete Prompt.
Derivation: For every ground-truth action chunk in the training data, compute the conditional log-probability and take its expectation. Selecting theta to maximize this quantity yields maximum-likelihood behavior cloning under the extended condition C_t.
Mechanism: This is a maximum-likelihood objective at the policy-distribution level. The Flow Matching implementation does not directly compute a normalized action density or an evidence lower bound; instead, it minimizes a conditional vector-field regression loss. With sufficient model capacity and optimization, the terminal distribution of the generated flow approaches the data action distribution.
5.2 Flow Matching Formulation for the Action Expert
The π0.7 Action Expert processes a fixed sequence of 50 action tokens. The supervised action chunk can be written as:
How to read this: Each action sample is not a single-step control command, but a joint 50-step action sequence beginning at the current time, capturing temporal and multi-joint correlations.
Derivation: The paper uses the 50-step action sequence starting at t as the supervised chunk. Therefore, the actions with consecutive indices from t through t+49 are defined as the target A_t_star.
As in the image world model, Gaussian noise is sampled for the action chunk and an interpolation is constructed:
How to read this: is an intermediate state between action noise and the ground-truth action chunk. The Action Expert must learn how to continuously transport a random action vector into a coordinated action chunk from the data.
Derivation: Action Flow Matching likewise uses linear interpolation from Gaussian noise to the ground-truth action chunk. tau controls the intermediate position, and the two endpoints determine the corresponding target velocity.
The corresponding velocity-regression objective can be abstracted as:
How to read this: Conditioned on the observation history and the complete Prompt, the Action Expert predicts the direction in which the current noisy action chunk should evolve until it forms a ground-truth executable action sequence.
Derivation: The target velocity of the linear action path is A_t_star minus epsilon_A. Conditioned on the visual history and the complete condition C_t, the network regresses this velocity. Taking the expectation over data, noise, and time yields the training loss.
5.3 The Closed Loop of Automated Prompts
At runtime, the system can be understood as three-level conditional generation:
How to read this: Based on the task and recent history, the high-level semantic policy selects the subtask that should be executed now, such as “open the refrigerator door.”
Derivation: The refined language instruction is not generated by fixed rules; it is conditionally predicted from the visual history and the original task. The distribution p_phi therefore represents possible local instructions and permits sampling or selection.
How to read this: Based on the current image, subtask, and desired behavioral attributes, the world model generates the near-term multiview visual state that the system should aim to reach.
Derivation: Conditioned on the current observation, refined language instruction, and mode, the goal generator produces a distribution over goal images. A hat denotes that these are generated by the model at inference time rather than provided directly by the data.
How to read this: The low-level Action Expert combines all conditions to sample a continuous action chunk, of which only a shorter portion is executed. When a new observation arrives, the system replans, thereby forming a closed-loop controller.
Derivation: The final action policy jointly conditions on the perception history, original task, generated refined language instruction, generated goal, mode, and other control conditions, forming a controllable conditional action distribution.
5.4 Why Execution Remains Smooth Under Latency
π0.7 uses the training-time version of Real-Time Action Chunking, simulating inference delays of 0 to 12 time steps during training. The robot control frequency is 50 Hz, so the maximum simulated delay is:
How to read this: Each control cycle lasts 20 milliseconds, so 12 cycles correspond to 240 milliseconds. Exposing the model to different delays during training can reduce discontinuities and jitter when switching between old and new action chunks.
Derivation: A control frequency of 50 Hz means that each step lasts 1/50 of a second. If at most 12 steps are executed before replanning, the longest open-loop interval until the next update is 12 divided by 50 seconds, or 0.24 seconds.
💡
Chapter 5 Interactive Explainer | Hierarchical Closed-Loop Simulator
Observe how the high-level subtask, world-model visual goal, Action Expert, and real environment repeatedly replan in a loop.
Chapter 5 Interactive Explainer | Hierarchical Closed-Loop Simulator
:::6. Why Low-Quality and Failed Data Are Now More Valuable
When training samples include quality and policy metadata, the model can learn that “this action comes from a low-quality demonstration” or “this trajectory comes from an older policy.” By specifying high-quality, fast, or safe conditions at inference time, the desired behavior can be extracted from mixed data without deleting all imperfect data.
6.1 How Bad Data Contaminates the Action Distribution Without Metadata
Let denote latent behavioral modes such as demonstration quality, speed, policy source, and whether an error occurred. If is not provided during training, the model can learn only the marginal distribution:
How to read this: For the same observation and task, the action distribution is a weighted mixture of different modes, including high-quality demonstrations, slow demonstrations, failure-recovery trajectories, and trajectories from older policies. When the model does not know which mode the current sample belongs to, it can only fit all of them simultaneously.
Derivation: The mode m is an unobserved discrete variable. Applying the conditional law of total probability to it, the overall action distribution equals the weighted sum of the action distributions under each mode, with weights given by the mode probabilities.
When different modes conflict in action space, a model with finite capacity may exhibit mode confusion. For example, it may learn both rapid approach motions and the hesitation and retreat found in failure trajectories, yet be unable to deliberately select between them at inference time.
6.2 How Metadata Turns Mixed Data into Selectable Data
After metadata is added to the Prompt, the policy instead learns:
How to read this: The model no longer asks only, “How is the current task usually performed?” Instead, it asks, “How should the task be performed under the specified speed, quality, error state, and policy source?”
Derivation: Once the mode m is explicitly included in the conditioning information, action data that were previously mixed together are decomposed into more homogeneous subdistributions. The policy can therefore learn mode-controllable action generation.
At inference time, the desired conditions can be selected explicitly:
How to read this: Even if the training set contains low-quality and failed trajectories, the model can be instructed at runtime to sample from the “high-quality and error-free” conditional action distribution.
Derivation: Fixing the mode fields to high quality and no error at inference time amounts to selecting the action distribution corresponding to that data subgroup from the conditional policy, rather than averaging over all quality and error modes.
💡
Capability boundary: Metadata cannot automatically turn failure into success. This approach assumes that labels are reliable, different modes have sufficient data coverage, and the model actually uses the conditioning information. These assumptions must be validated through shuffled Metadata, counterfactual Prompts, and evaluations stratified by failure mode.
💡
Chapter 6 Interactive Explainer | Mixed-Quality Data Laboratory
Adjust the proportions of high-quality, older-policy, and failed data, and compare the marginal mixture distribution with the Metadata-conditioned distribution.
Chapter 6 Interactive Explainer | Mixed-Quality Data Laboratory
:::7. A Cautious Interpretation of Emergent Capabilities
The paper demonstrates novel task combinations and cross-scenario capabilities, but “emergence” should be understood as compositional generalization arising from diverse Prompts, data coverage, and shared representations. Without strict task-combination holdouts and repeated statistical evaluation, a single novel behavior cannot be directly equated with general reasoning ability.
💡
Chapter 7 Interactive Explainer | Emergence Evidence Evaluator
Adjust task novelty and the number of repetitions, and assess the strength of evidence using compositional holdouts, confidence intervals, and baselines.
Chapter 7 Interactive Explainer | Emergence Evidence Evaluator
:::Exercises
- Compare which control constraints are best expressed by language subtasks versus visual subgoals.
- Design an ablation to determine whether the model truly uses the video history rather than only the final frame.
- Define steerability metrics covering task success, instruction adherence, policy-switching latency, and safety constraints.

