Original Feishu Document · Source Revision 16
💡
Course overview: This lesson develops the VLA pathway from first principles rather than organizing it around a particular company or individual paper. RT, Open X, Octo, OpenVLA, GR00T, and others constitute the mainstream lineage, while π0, FAST, π0.5, π0.6*, and π0.7 are presented as a continuous series of engineering case studies in the Paper Lab at the end of the lesson.
Learning Objectives
After completing this lesson, you should be able to explain the probabilistic objective of direct policy learning; distinguish among VLMs, VLAs, action experts, and controllers; compare regression, autoregressive tokens, Diffusion, and Flow Matching; understand multi-robot co-training and post-training; determine which capabilities VLAs do and do not possess; and place the π series within the VLA pathway rather than equating it with Physical AI as a whole.
1. The Core Assumption of the Direct Policy Pathway
Direct policy learning bypasses explicit planning and explicit system identification, directly learning:
Reading: The policy parameterized by theta assigns a conditional probability to a future H-step action chunk based on the observation history up to the current time, the current proprioceptive state, and the language goal.
Derivation: Treat the future H-step continuous action sequence as a random vector. A direct policy conditions this joint action distribution on multimodal context.
This is read as: “Given the observation history, proprioceptive state, and language goal, generate a future segment of actions.” It compresses perception, task conditioning, and action selection into a single conditional distribution.
Its advantages are direct inference and the ability to leverage large-scale demonstrations. Its limitation is that the model is primarily constrained by the data distribution and may not be able to explicitly answer questions about counterfactuals, long-term value, or dynamics parameters.
2. From Behavior Cloning to VLA
2.1 Behavior Cloning
Given robot data:
Reading: The robot dataset consists of observation histories, proprioceptive states, language tasks, and the corresponding future H-step action chunks.
Derivation: Pair each robot demonstration's observation history, proprioceptive state, task, and future action chunk at time t into a supervised sample. The collection of all such samples constitutes the robot training dataset.
A real dataset should also retain episode identifiers, timestamps, coordinate frames, control frequencies, and embodiment metadata; only the minimal mathematical fields required for training are shown here.
Behavior cloning minimizes:
Reading: The behavior cloning loss is the average negative log conditional probability of the ground-truth action chunks over the robot dataset.
Derivation: Maximize the conditional likelihood of all demonstrated action chunks; negating it yields the negative log-likelihood objective to be minimized.
2.2 Vision-Language Conditioning
A VLA introduces vision-language representations pretrained on internet-scale data into the policy's conditioning context. The VLM provides object, scene, and linguistic semantics, while the action module converts these representations into continuous control.
2.3 A VLM Is Not a VLA
| Model | Training target | Output |
|---|---|---|
| VLM | Semantic relationships between images and text | Text tokens or multimodal representations |
| VLA | Action distribution conditioned on multimodal context | Action tokens, continuous actions, or trajectories |
| Controller | Feedback error and physical execution | Motor commands or torques |
3. How to Represent an Action Distribution
3.1 Point Regression
Reading: Given context c, the model outputs a single predicted action, denoted by the conditional mean mu_theta(c).
Derivation: The negative log-likelihood of a fixed-variance Gaussian is equivalent to MSE, and the optimal deterministic prediction under MSE is the conditional mean.
It is fast, but the fixed-variance Gaussian assumption can average across multiple valid actions.
3.2 Autoregressive Action Tokens
Reading: The probability of the complete action-token sequence equals the product of the probability of each token conditioned on the preceding tokens and the context.
Derivation: Expand the joint distribution using the probability chain rule; during training, apply cross-entropy to each token.
The actions are first converted into a discrete sequence by a tokenizer, allowing the reuse of language-model cross-entropy training. However, this approach introduces quantization error and serial decoding latency.
3.3 Diffusion
The model learns a denoising direction or score at different noise levels and produces continuous actions through multi-step sampling. It is well suited to multimodal trajectories but incurs greater inference cost.
3.4 Flow Matching
The model learns a continuous-time vector field:
Reading: The rate of change of the generated state x with respect to flow time t equals the velocity field predicted by the model given the current state, time, and context.
Derivation: Flow Matching learns a vector field that transports a simple noise distribution to the conditional action distribution. During inference, this ordinary differential equation is integrated.
Starting from a simple noise distribution, conditional action samples are obtained through integration.
| Representation | Learning target | Inference | Best suited for |
|---|---|---|---|
| Regression | Conditional mean | Single forward pass | Approximately unimodal distributions and low latency |
| Action Tokens | Discrete-sequence probability | Token by token | Unified pretraining with a VLM |
| Diffusion | Score or noise | Multi-step denoising | Multimodal continuous actions |
| Flow | Probability-flow vector field | ODE integration | Continuous action chunks |
4. Why Action Chunks Became Mainstream
A single-step policy predicts only at a time, making it prone to high-frequency jitter and unable to express short-term coordination among grasping, rotation, and release. An action-chunk policy directly generates:
Reading: The action chunk A_t at time t consists of H consecutive actions beginning with the current action.
Derivation: Starting with the current action a_t, collect H consecutive actions in temporal order and concatenate them into a joint variable to obtain the action chunk A_t.
The length H selects a temporal scale: a longer action chunk increases short-term consistency, but it also extends open-loop execution and delays feedback.
Action chunks can learn short-horizon trajectory structure and reduce inference frequency, but they increase open-loop duration. In real deployments, the system typically executes only the first few steps of an action chunk and then generates a new one from updated observations.
5. Multi-Robot and Multi-Task Co-Training
A general-purpose VLA aims to use data from different robots, cameras, tasks, and control frequencies. Unified training must address at least:
- Different action dimensions and joint semantics.
- Different coordinate frames, units, and normalization schemes.
- Different camera configurations and visual distributions.
- Different control frequencies and action-chunk durations.
- Different levels of data quality, teleoperation styles, and failure rates.
“Padding tensors to the same length” does not mean that the actions are semantically aligned. A model may use an embodiment ID to learn multiple local policies, or it may learn transferable object-centric structure; dedicated experiments are required to distinguish between these cases.
6. Pretraining, Co-Training, and Post-Training
| Stage | Primary data | Primary role |
|---|---|---|
| Vision-language pretraining | Internet-scale image-text data | Semantic and visual representations |
| Robot pretraining/co-training | Multi-task, multi-robot trajectories | Broad action capabilities and conditional alignment |
| Task post-training | High-quality target-scenario data | Stability, speed, and behavioral style |
| Experience learning | Autonomous successes, failures, and corrections | Correcting the deployment distribution and improving success rates |
Scaling increases capability coverage, post-training shapes deployment behavior, and experience learning corrects the state distribution induced by the policy itself. These three effects cannot be reduced to simply “more data.”
7. What a VLA Can Learn—and What It Does Not Learn Automatically
| May learn | Cannot be guaranteed by architecture alone |
|---|---|
| Action patterns conditioned on objects and tasks | An explicit causal model of physics |
| Shared visual-action representations across tasks | Long-term optimality |
| Short-horizon action coordination and feedback correction | Recovery from out-of-distribution failures |
| Language-conditioned control of actions | Zero-shot transfer to arbitrary robots |
| Manipulation priors that recur throughout the data | Safety, contact stability, and controller tracking |
These gaps connect to other pathways: counterfactual prediction connects to world models, long-term quality connects to value learning, task decomposition connects to hierarchical planning, cross-embodiment transfer connects to data representation, and real-world execution connects to dynamics and control.
8. Placing the π Series Back Within the VLA Pathway
| Stage | What it adds to the VLA pathway | What it should not be mistaken for |
|---|---|---|
| π0 | VLM Backbone and Flow Action Expert | A complete Physical AI system |
| FAST | Action tokenization and unified cross-entropy pretraining | Proof that discrete actions are necessarily superior to continuous actions |
| π0.5 | Heterogeneous co-training, open-world tasks, and hierarchical conditioning | Generalization emerging automatically from scale alone |
| π0.6* / RECAP | Improvement using deployment experience and Advantage conditioning | A solution to general-purpose online RL |
| π0.7 | Diverse prompts, visual subgoals, and steerability | An explicit world model equivalent to a complete planner |
The π series constitutes a continuous set of engineering case studies within the VLA pathway. FAST connects to action representation, RECAP connects to value learning, and π0.7 connects to world models and hierarchical planning, but their primary classification remains VLA.
9. Training and Inference Must Be Analyzed Separately
| Stage | Available information | Must produce |
|---|---|---|
| Behavior-cloning training | Observations, language, proprioception, and ground-truth actions | Action likelihood or generative target |
| Generative-policy training | Ground-truth actions, noise, and time | Denoising direction or vector field |
| Inference | Observations, language, proprioception, and random noise | Complete action chunk |
| Execution | Generated actions and control cycle | Physical robot motion |
Ground-truth actions are available during training but not during deployment. If these two settings are not distinguished when reading a paper, it is easy to misunderstand the network inputs and the sources of supervision.
10. Decisive Evaluations for VLAs
- Task holdout: Novel task compositions rather than merely novel backgrounds.
- Object holdout: Variations in shape, material, and affordances.
- Embodiment holdout: Changes in control interfaces and kinematics.
- Closed-loop recovery: Whether the system recovers after deliberate perturbations.
- Multimodal tasks: Whether the policy covers multiple valid action modes.
- Latency and control: Action-generation latency, smoothness, and collision rate.
- Equal-data ablation: Separating gains from model mechanisms from gains due to data scale.
10.1 Minimal Reproducible Experiment
Using the same two-dimensional obstacle-avoidance dataset, hold the number of training samples and network capacity fixed, and separately train point-regression, action-token, and Flow Matching policies.
- Report offline action error, but do not treat it as the final conclusion.
- Run closed-loop execution and record goal-reaching rate, collision rate, mode coverage, and inference latency.
- Introduce variations in the starting point, obstacle positions, and language expressions, and measure out-of-distribution degradation separately.
- Change only the executed length of each action chunk and observe the trade-off between smoothness and recovery speed.
This experiment requires the reader to translate the claim that “one action representation is more powerful” into a falsifiable closed-loop claim.
11. Exercises
- Draw the boundaries among the VLM, VLA, Action Expert, and low-level controller.
- Compare regression, Tokens, Diffusion, and Flow on a bimodal obstacle-avoidance task.
- Define an action schema and embodiment metadata for cross-robot data.
- Explain why action chunks simultaneously improve smoothness and reduce the frequency of closed-loop correction.
- Explain π0.6* once from the VLA pathway and once from the value-learning pathway.
- Design an intervention experiment that can verify that a VLA uses language conditioning rather than merely recognizing the scene.
Paper Lab
The Mainstream VLA Lineage: RT, Open X, Octo, OpenVLA, GR00T, and Gemini Robotics
An In-Depth Study of the π0 Architecture
FAST and Action Representation
π0.5 and Open-World Generalization

