Skip to content

Original Feishu Document · Source Revision 25

💡

Course positioning: This chapter does not treat π0 as a model to be memorized, but as a case study in VLA architecture. The central questions are: How can vision-language representations pretrained on internet-scale data serve as conditions for a continuous action distribution? What do the semantic model and action model learn respectively? How are the ground-truth actions used during training replaced by a generative process at inference time?

1. VLA Inputs, Outputs, and Learning Objective ​

Learning objectives: After completing this chapter, you should be able to verbally interpret the VLA conditional action distribution, draw the information flow between the VLM Backbone and Action Expert, explain the difference between Flow Matching training and inference, and distinguish among a VLA, a world model, and a high-level planner.

The model takes language instructions, multiple images, and robot proprioceptive state as input, and outputs a fixed-length chunk of continuous actions. Because different robots have different degrees of freedom, the training pipeline uses embodiment-specific input-output mappings to convert them into tensors that the model can process.

Denote the images, language, proprioceptive state, and history collectively as context . The VLA learns:

Reading: The policy parameterized by theta assigns probabilities to future H-step action chunks beginning at the current time, conditioned on the current multimodal context c_t.

Derivation: Language, images, and proprioceptive state are combined into the conditioning variable; the future H-step actions are treated as a joint random vector.

Vision and language are not optional embellishments; they are conditioning variables for the action distribution. Language changes the task objective, images provide evidence about the environment, and proprioceptive information constrains the actions currently feasible for the robot.

VariableContentStatistical rolePhysical meaning
lLanguage instructionConditioning variableTask and constraints
o_t^Multi-view imagesConditioning variableObservations of the external environment
q_tJoint, end-effector, and gripper statesConditioning variableRobot proprioceptive state
a_Action chunkTraining label and generated variableFuture H-step continuous control commands beginning at the current time

The training data come from the joint distribution . The model therefore first learns how contexts and actions are jointly distributed in the data, rather than deriving robot dynamics from first principles.

2. Dual-Expert Architecture ​

ModulePrimary responsibilityTraining source
VLM BackboneImages, language, task semantics, and scene representationsInternet-scale vision-language pretraining plus robot data
Action ExpertProcesses proprioceptive state, noisy actions, and flow time, and outputs a velocity fieldRobot action data; initialized from scratch

In the paper, the VLM Backbone is based on PaliGemma, while the Action Expert is a smaller Transformer. They operate jointly within the attention layers, but robot state and action tokens are routed to the Action Expert.

3. How a Training Sample Passes Through the Network ​

  1. Images are converted into visual tokens by the vision encoder, while language is tokenized into text tokens.
  2. The robot state q_t is linearly projected and passed into the Action Expert.
  3. Noise is added to the ground-truth action chunk A to obtain A^tau, which is provided as input together with the flow-time encoding.
  4. The VLM provides semantic conditioning, and the Action Expert predicts the velocity at each action position.
  5. Continuous action outputs are trained using a conditional Flow Matching loss.

4. Why Pre-training and Post-training Are Both Needed ​

Broad pretraining data provide task coverage, cross-robot transfer, and semantic generalization, but their quality is uneven. High-quality post-training data emphasize stable, smooth, fast, and recoverable behavior. The logic resembles pretraining and instruction post-training for language models: scale establishes the capability frontier, while curated data shape the behavioral distribution.

5. Why the High-Level Policy Remains Separate in π0 ​

Complex tasks such as clearing a dining table contain multiple stages. An external high-level VLM can generate the current subtask for π0, such as “pick up the plate” or “put the trash in the bin,” after which the low-level π0 executes the actions. In this configuration, the system is not a fully end-to-end hierarchical policy: high-level semantic decision-making and low-level action policy execution remain two separate model calls.

6. Training Computation Graph: How Semantic Conditioning Enters Action Generation ​

The VLM Backbone provides scene and task semantics, while the Action Expert receives the robot state, noisy actions, and flow time. They are not two unrelated models: the Action Expert reads the VLM conditioning through attention, causing the same noisy action to flow toward different correct actions under different language and image conditions.

6.1 Conditional Flow Matching Objective ​

Let denote the ground-truth action chunk and denote Gaussian noise. Construct the intermediate action:

Reading: The intermediate action x_tau is a linear interpolation between Gaussian noise x_0 and the ground-truth action chunk x_1; tau increases from zero to one.

Derivation: When tau equals zero, the path is at the noise sample; when tau equals one, it reaches the ground-truth action, defining the simplest conditional probability path.

The target velocity along the straight-line path is:

Reading: At any flow time, the target velocity along the linear path equals the ground-truth action chunk minus the initial noise.

Derivation: Differentiating the linear interpolation with respect to tau yields a constant velocity.

The model training objective is:

Reading: Sample a context and ground-truth action from robot data, sample noise from a standard Gaussian, and sample flow time uniformly from zero to one; train the model velocity to match the ground-truth velocity of the linear path.

Derivation: The ground-truth action, noise, and time are all known during training, so the intermediate state and supervised velocity can be constructed, allowing the conditional vector field to be fit with MSE.

The model does not memorize a single fixed action. Instead, it learns an action vector field jointly modulated by images, language, and proprioceptive state; the same initial noise sample flows toward different actions under different contexts.

6.2 What Capabilities Are Actually Updated by the Gradients ​

The loss backpropagates through the Action Expert and may also update VLM layers trained jointly with it. Action errors can therefore modify vision-language representations, making them more attentive to factors useful for control, such as the relative position between the gripper and an object. Joint training alone, however, does not guarantee that the model develops explicit 3D geometry or causal state representations; this must be validated using targeted representation probes and intervention experiments.

7. Inference Computation Graph: What Happens to the Ground-Truth Actions Used During Training? ​

During training, the ground-truth action is known because it is needed to construct the supervised target. During deployment, the ground-truth action is unknown, so the system can only begin from noise:

Reading: At the beginning of inference, the action state is sampled from a standard Gaussian noise distribution.

Derivation: Flow Matching inference requires a random initial state, so x(0) is drawn from an easily sampled standard Gaussian; the conditional velocity field then transports this noise to the action distribution.

The ground-truth training action x_1 is no longer available; it influences the generation process only indirectly through the parameters of the vector field.

The system then performs numerical integration:

Reading: The rate of change of the action state over flow time is given by the conditional vector field, which is integrated from noise at tau equal to zero to an action sample at tau equal to one.

Derivation: After learning a continuous velocity field, numerical integration with Euler’s method or another ODE solver transports the noise into an action chunk.

This produces the action chunk . Consequently, generating one action chunk usually requires multiple Action Expert forward passes. After generation, the robot executes part of the action chunk and then regenerates actions using new observations. Executing too many steps reduces the frequency of closed-loop corrections, while executing too few increases inference latency and action jitter.

StageInformation available to the modelWhat the model must produceMain risks
TrainingImages, language, proprioception, ground-truth actions, noise, and timeTarget velocityData bias, action misalignment, and label noise
InferenceImages, language, proprioception, noise, and timeComplete action chunkSampling error, computational latency, and closed-loop distribution shift
ExecutionGenerated actions and control periodPhysical robot motionContact, calibration, latency, and safety constraints

8. Boundaries Among VLAs, World Models, and Planners ​

ModuleTypical mathematical objectQuestion answeredRole in π0
VLA policyConditional action distributionWhat should be done now?VLM Backbone plus Action Expert
World modelAction-conditioned future distributionWhat will happen after taking this action?π0 itself has no separate explicit world model
Value functionState-action valueHow good is this action in the long term?The base π0 is trained primarily from imitation data and does not rely on an explicit critic
High-level plannerSubtask or candidate action sequenceWhat is the next stage of a long-horizon task?An external high-level VLM can provide subtasks
Low-level controllerControl law and execution constraintsHow can the action command be tracked stably?Robot control stack outside the model

Therefore, “connecting a VLM to a robot” does not mean that one model already covers all of Physical AI. A more accurate statement is that VLM representations are used as conditioning to enable a generative policy to output continuous actions; long-horizon planning, explicit future prediction, safety control, and failure recovery still require additional mechanisms or data feedback loops.

Minimum Experiments for This Chapter ​

  1. Hold the Action Expert fixed and change only the language instruction to determine whether the generated actions change systematically in a task-consistent manner.
  2. Occlude image regions relevant to the task and compare action errors to determine which visual evidence the model uses.
  3. Compare single-step actions, Action Chunks of different lengths, and different replanning frequencies.
  4. Compare an MSE point-estimate action head with a Flow Matching action head on a bimodal task.
  5. Sweep control latency and report success rate, collision rate, and action smoothness rather than reporting only mean success rate.

9. What π0 Actually Demonstrates ​

  • The large-scale, multitask, multi-robot system reported in the paper demonstrates substantial dexterity, but its capability boundaries remain constrained by data coverage and robot-specific adaptation.
  • The paper’s ablations show that the combination of VLM pretraining and the Flow Action Expert outperforms the smaller baselines evaluated in the paper. This demonstrates the effectiveness of the joint design, but does not isolate the full causal contribution of each component.
  • Action chunking and continuous generation demonstrate the feasibility of handling long action sequences and tasks involving deformable objects. Whether they outperform Token, Diffusion, or other interfaces still requires comparisons with matched data and compute.

10. What Cannot Be Directly Concluded from the Paper ​

  • The paper does not prove that the model understands real physical causality; its success may result from data coverage and closed-loop visual feedback.
  • The paper does not establish zero-shot transfer to arbitrary robots; the action space, cameras, and control frequency still require adaptation.
  • Reliability cannot be assessed from a small number of demonstration videos; the failure distribution, repeated trials, and task-stratified metrics must be examined.

Reading Check ​

  1. Draw the information flow among VLM tokens, state tokens, and noisy action tokens.
  2. Explain why the Action Expert can be much smaller than the VLM Backbone.
  3. Discuss the trade-off between open-loop execution of action chunks and closed-loop replanning at every step.

Original Sources ​

Article text is licensed under the Apache License 2.0