Skip to content

Original Feishu Document · Source Revision 16

💡

Course overview: This lesson develops the VLA pathway from first principles rather than organizing it around a particular company or individual paper. RT, Open X, Octo, OpenVLA, GR00T, and others constitute the mainstream lineage, while π0, FAST, π0.5, π0.6*, and π0.7 are presented as a continuous series of engineering case studies in the Paper Lab at the end of the lesson.

Learning Objectives ​

After completing this lesson, you should be able to explain the probabilistic objective of direct policy learning; distinguish among VLMs, VLAs, action experts, and controllers; compare regression, autoregressive tokens, Diffusion, and Flow Matching; understand multi-robot co-training and post-training; determine which capabilities VLAs do and do not possess; and place the π series within the VLA pathway rather than equating it with Physical AI as a whole.

1. The Core Assumption of the Direct Policy Pathway ​

Direct policy learning bypasses explicit planning and explicit system identification, directly learning:

Reading: The policy parameterized by theta assigns a conditional probability to a future H-step action chunk based on the observation history up to the current time, the current proprioceptive state, and the language goal.

Derivation: Treat the future H-step continuous action sequence as a random vector. A direct policy conditions this joint action distribution on multimodal context.

This is read as: “Given the observation history, proprioceptive state, and language goal, generate a future segment of actions.” It compresses perception, task conditioning, and action selection into a single conditional distribution.

Its advantages are direct inference and the ability to leverage large-scale demonstrations. Its limitation is that the model is primarily constrained by the data distribution and may not be able to explicitly answer questions about counterfactuals, long-term value, or dynamics parameters.

2. From Behavior Cloning to VLA ​

2.1 Behavior Cloning ​

Given robot data:

Reading: The robot dataset consists of observation histories, proprioceptive states, language tasks, and the corresponding future H-step action chunks.

Derivation: Pair each robot demonstration's observation history, proprioceptive state, task, and future action chunk at time t into a supervised sample. The collection of all such samples constitutes the robot training dataset.

A real dataset should also retain episode identifiers, timestamps, coordinate frames, control frequencies, and embodiment metadata; only the minimal mathematical fields required for training are shown here.

Behavior cloning minimizes:

Reading: The behavior cloning loss is the average negative log conditional probability of the ground-truth action chunks over the robot dataset.

Derivation: Maximize the conditional likelihood of all demonstrated action chunks; negating it yields the negative log-likelihood objective to be minimized.

2.2 Vision-Language Conditioning ​

A VLA introduces vision-language representations pretrained on internet-scale data into the policy's conditioning context. The VLM provides object, scene, and linguistic semantics, while the action module converts these representations into continuous control.

2.3 A VLM Is Not a VLA ​

ModelTraining targetOutput
VLMSemantic relationships between images and textText tokens or multimodal representations
VLAAction distribution conditioned on multimodal contextAction tokens, continuous actions, or trajectories
ControllerFeedback error and physical executionMotor commands or torques

3. How to Represent an Action Distribution ​

3.1 Point Regression ​

Reading: Given context c, the model outputs a single predicted action, denoted by the conditional mean mu_theta(c).

Derivation: The negative log-likelihood of a fixed-variance Gaussian is equivalent to MSE, and the optimal deterministic prediction under MSE is the conditional mean.

It is fast, but the fixed-variance Gaussian assumption can average across multiple valid actions.

3.2 Autoregressive Action Tokens ​

Reading: The probability of the complete action-token sequence equals the product of the probability of each token conditioned on the preceding tokens and the context.

Derivation: Expand the joint distribution using the probability chain rule; during training, apply cross-entropy to each token.

The actions are first converted into a discrete sequence by a tokenizer, allowing the reuse of language-model cross-entropy training. However, this approach introduces quantization error and serial decoding latency.

3.3 Diffusion ​

The model learns a denoising direction or score at different noise levels and produces continuous actions through multi-step sampling. It is well suited to multimodal trajectories but incurs greater inference cost.

3.4 Flow Matching ​

The model learns a continuous-time vector field:

Reading: The rate of change of the generated state x with respect to flow time t equals the velocity field predicted by the model given the current state, time, and context.

Derivation: Flow Matching learns a vector field that transports a simple noise distribution to the conditional action distribution. During inference, this ordinary differential equation is integrated.

Starting from a simple noise distribution, conditional action samples are obtained through integration.

RepresentationLearning targetInferenceBest suited for
RegressionConditional meanSingle forward passApproximately unimodal distributions and low latency
Action TokensDiscrete-sequence probabilityToken by tokenUnified pretraining with a VLM
DiffusionScore or noiseMulti-step denoisingMultimodal continuous actions
FlowProbability-flow vector fieldODE integrationContinuous action chunks

4. Why Action Chunks Became Mainstream ​

A single-step policy predicts only at a time, making it prone to high-frequency jitter and unable to express short-term coordination among grasping, rotation, and release. An action-chunk policy directly generates:

Reading: The action chunk A_t at time t consists of H consecutive actions beginning with the current action.

Derivation: Starting with the current action a_t, collect H consecutive actions in temporal order and concatenate them into a joint variable to obtain the action chunk A_t.

The length H selects a temporal scale: a longer action chunk increases short-term consistency, but it also extends open-loop execution and delays feedback.

Action chunks can learn short-horizon trajectory structure and reduce inference frequency, but they increase open-loop duration. In real deployments, the system typically executes only the first few steps of an action chunk and then generates a new one from updated observations.

5. Multi-Robot and Multi-Task Co-Training ​

A general-purpose VLA aims to use data from different robots, cameras, tasks, and control frequencies. Unified training must address at least:

  • Different action dimensions and joint semantics.
  • Different coordinate frames, units, and normalization schemes.
  • Different camera configurations and visual distributions.
  • Different control frequencies and action-chunk durations.
  • Different levels of data quality, teleoperation styles, and failure rates.

“Padding tensors to the same length” does not mean that the actions are semantically aligned. A model may use an embodiment ID to learn multiple local policies, or it may learn transferable object-centric structure; dedicated experiments are required to distinguish between these cases.

6. Pretraining, Co-Training, and Post-Training ​

StagePrimary dataPrimary role
Vision-language pretrainingInternet-scale image-text dataSemantic and visual representations
Robot pretraining/co-trainingMulti-task, multi-robot trajectoriesBroad action capabilities and conditional alignment
Task post-trainingHigh-quality target-scenario dataStability, speed, and behavioral style
Experience learningAutonomous successes, failures, and correctionsCorrecting the deployment distribution and improving success rates

Scaling increases capability coverage, post-training shapes deployment behavior, and experience learning corrects the state distribution induced by the policy itself. These three effects cannot be reduced to simply “more data.”

7. What a VLA Can Learn—and What It Does Not Learn Automatically ​

May learnCannot be guaranteed by architecture alone
Action patterns conditioned on objects and tasksAn explicit causal model of physics
Shared visual-action representations across tasksLong-term optimality
Short-horizon action coordination and feedback correctionRecovery from out-of-distribution failures
Language-conditioned control of actionsZero-shot transfer to arbitrary robots
Manipulation priors that recur throughout the dataSafety, contact stability, and controller tracking

These gaps connect to other pathways: counterfactual prediction connects to world models, long-term quality connects to value learning, task decomposition connects to hierarchical planning, cross-embodiment transfer connects to data representation, and real-world execution connects to dynamics and control.

8. Placing the π Series Back Within the VLA Pathway ​

StageWhat it adds to the VLA pathwayWhat it should not be mistaken for
π0VLM Backbone and Flow Action ExpertA complete Physical AI system
FASTAction tokenization and unified cross-entropy pretrainingProof that discrete actions are necessarily superior to continuous actions
π0.5Heterogeneous co-training, open-world tasks, and hierarchical conditioningGeneralization emerging automatically from scale alone
π0.6* / RECAPImprovement using deployment experience and Advantage conditioningA solution to general-purpose online RL
π0.7Diverse prompts, visual subgoals, and steerabilityAn explicit world model equivalent to a complete planner

The π series constitutes a continuous set of engineering case studies within the VLA pathway. FAST connects to action representation, RECAP connects to value learning, and π0.7 connects to world models and hierarchical planning, but their primary classification remains VLA.

9. Training and Inference Must Be Analyzed Separately ​

StageAvailable informationMust produce
Behavior-cloning trainingObservations, language, proprioception, and ground-truth actionsAction likelihood or generative target
Generative-policy trainingGround-truth actions, noise, and timeDenoising direction or vector field
InferenceObservations, language, proprioception, and random noiseComplete action chunk
ExecutionGenerated actions and control cyclePhysical robot motion

Ground-truth actions are available during training but not during deployment. If these two settings are not distinguished when reading a paper, it is easy to misunderstand the network inputs and the sources of supervision.

10. Decisive Evaluations for VLAs ​

  1. Task holdout: Novel task compositions rather than merely novel backgrounds.
  2. Object holdout: Variations in shape, material, and affordances.
  3. Embodiment holdout: Changes in control interfaces and kinematics.
  4. Closed-loop recovery: Whether the system recovers after deliberate perturbations.
  5. Multimodal tasks: Whether the policy covers multiple valid action modes.
  6. Latency and control: Action-generation latency, smoothness, and collision rate.
  7. Equal-data ablation: Separating gains from model mechanisms from gains due to data scale.

10.1 Minimal Reproducible Experiment ​

Using the same two-dimensional obstacle-avoidance dataset, hold the number of training samples and network capacity fixed, and separately train point-regression, action-token, and Flow Matching policies.

  1. Report offline action error, but do not treat it as the final conclusion.
  2. Run closed-loop execution and record goal-reaching rate, collision rate, mode coverage, and inference latency.
  3. Introduce variations in the starting point, obstacle positions, and language expressions, and measure out-of-distribution degradation separately.
  4. Change only the executed length of each action chunk and observe the trade-off between smoothness and recovery speed.

This experiment requires the reader to translate the claim that “one action representation is more powerful” into a falsifiable closed-loop claim.

11. Exercises ​

  1. Draw the boundaries among the VLM, VLA, Action Expert, and low-level controller.
  2. Compare regression, Tokens, Diffusion, and Flow on a bimodal obstacle-avoidance task.
  3. Define an action schema and embodiment metadata for cross-robot data.
  4. Explain why action chunks simultaneously improve smoothness and reduce the frequency of closed-loop correction.
  5. Explain π0.6* once from the VLA pathway and once from the value-learning pathway.
  6. Design an intervention experiment that can verify that a VLA uses language conditioning rather than merely recognizing the scene.

Paper Lab ​

The Mainstream VLA Lineage: RT, Open X, Octo, OpenVLA, GR00T, and Gemini Robotics

An In-Depth Study of the π0 Architecture

FAST and Action Representation

π0.5 and Open-World Generalization

π0.6*, RECAP, and Experience Learning

π0.7 and Steerability

Article text is licensed under the Apache License 2.0