Original Feishu Document · Source Revision 24
💡
Core course in this track: Physical AI data is not limited to robot trajectories. Human videos, hand skeletons, object changes, language descriptions, and failure records can all provide supervision, but they correspond to different mathematical objects. This course examines how to learn representations, motion, latent actions, and transferable cross-robot structure from these data sources.
Learning Objectives
After completing this course, you should be able to distinguish among observation supervision, action supervision, state-change supervision, and outcome supervision; explain what information is missing from human videos; understand motion learning through optical flow, keypoints, object-centric representations, and inverse dynamics; design latent actions and behavior tokenizers; analyze the assumptions and failure modes of cross-embodiment alignment; and design verifiable experiments for incorporating human data into VLAs, world models, and hierarchical policies.

1. What Supervision Does a Data Sample Actually Contain?
| Data | Directly observable | Not directly observable | Learnable objects |
|---|---|---|---|
| Robot trajectories | Images, proprioceptive state, actions, outcomes | Expert intent, unexecuted counterfactual actions | Policies, dynamics, value functions |
| Human videos | Images, changes in hands and objects | Robot control commands, true scale, forces, and contact | Visual representations, motion structure, latent actions |
| Language descriptions | Task and event semantics | Precise geometry and control timing | Task conditions, subgoals, event labels |
| Hand skeletons | Keypoint positions and topology | Contact forces, object weight, robot-joint mappings | Grasp types, relative motion, interaction phases |
| Failure records | States, actions, outcomes, or interventions | The complete causal chain behind a failure | Recovery policies, value functions, data reweighting |
“No action labels” does not mean “no supervision,” but the supervision often shifts from direct actions to state changes, temporal ordering, or object-interaction structure.
2. The Observation Model for Human Video
A video can be written as:
Reading: A length-T video observation consists of image frames I_1 through I_T arranged in temporal order.
Derivation: A camera initially records a sequence of pixels rather than world coordinates, forces, or robot actions. All subsequent supervision for motion, events, and latent actions must be estimated from these observations or obtained using additional sensors.
If the video contains object keypoints , pixel displacement can be calculated as:
Reading: The pixel displacement of the i-th image keypoint at time t equals its position in the next frame minus its position in the current frame.
Derivation: Keypoint tracking provides two-dimensional coordinates in adjacent frames. Taking their difference yields the direction and magnitude of motion on the image plane. This is a change in observation space and does not include depth scale, camera-motion compensation, or physical units.
However, pixel displacement is not physical velocity:
Reading: Physical velocity in world coordinates cannot be equated with pixel displacement divided by time.
Derivation: Perspective projection maps three-dimensional positions onto two-dimensional pixels. The same world-frame velocity can produce different pixel velocities because of depth, focal length, and camera motion. Pixel changes can be converted into velocity estimates with physical units only after camera calibration, depth estimation, and motion compensation.
This is because camera motion, depth, focal length, occlusion, and perspective all affect pixel changes. Monocular video usually provides only relative-motion or directional cues; it does not directly provide robot end-effector velocity or torque.
2.1 From Images to Motion Representations
| Representation | Information retained | Main ambiguities |
|---|---|---|
| Pixel optical flow | Local visual motion | Camera motion and depth |
| Human keypoints | Body and hand geometry | Occlusion, scale, and 3D depth |
| Object-centric trajectories | Relative object displacement and interaction | Object identity and contact |
| Video latent | Temporal and semantic structure | Interpretability and utility for control |
| Event labels | Phases such as contact, stable grasp, and release | Label boundaries and annotation cost |
3. Hand–Object Interaction and Object-Centric Representations
Robots need to know “what the hand did relative to the object” more than “how many pixels changed across the entire image.” An object-centric representation can be written as:
Reading: The hand–object relationship consists of relative position, relative orientation, and relative velocity expressed in the object coordinate frame.
Derivation: First, subtract the positions to obtain the hand translation relative to the object, and then multiply by the transpose of the object’s rotation matrix to express this vector in the object coordinate frame. Similarly, relative rotation is obtained by multiplying the inverse of the object orientation by the hand orientation. This removes changes in absolute scene position and more directly describes grasping, insertion, and rotation relationships.
It contains relative position, relative orientation, and relative velocity. Relative representations are more robust to camera translation and changes in absolute coordinates across scenes, and they more closely match the control variables of tasks such as grasping, insertion, and rotation.
3.1 Contact Is Not a Pixel
An interaction representation must consider at least the contact location, contact normal, normal force, tangential slip, and object deformation. Visual video often provides only weak cues about contact events and must be learned jointly with force sensing, tactile sensing, or action outcomes.
4. Learning Motion from Video Without Fabricating Robot Actions
There are four common approaches:
| Approach | Supervision objective | How robots use it | Main assumption |
|---|---|---|---|
| Video prediction | Predict the next frame or future state from past video frames | Learn future visual states and subgoals | The visual future is relevant to the task future |
| Temporal/contrastive learning | Correct ordering, neighboring clips, or cross-view consistency | Pretrain motion representations | The representation transfers to robot tasks |
| Inverse dynamics | Infer possible robot actions from changes between adjacent states | Infer latent actions from state changes | State changes are sufficient to identify actions |
| Latent actions | Infer latent behavioral variables that explain state changes from adjacent observations | Encode video changes as behavior tokens | Latent actions share semantics across embodiments |
The key distinction is that video can teach a robot “what changed” or “where it should move next,” but it does not automatically teach the robot “how much force to apply, which joint trajectory to follow, or what control frequency to use.”
5. Latent Actions: Inferring Action Variables from State Changes
If the true action is unobserved, a latent variable can be introduced:
Reading: The inference model estimates latent action z_t from adjacent observations, and the generative model then predicts the next observation from the current observation and that latent action.
Derivation: When the true action is unobserved, the factor responsible for the state change can be treated as a latent variable. The encoder infers z_t from an observation pair, while the decoder requires z_t to contain enough information to explain the next state. However, any invertible reparameterization may yield the same prediction, so objects, events, or robot actions are also needed to provide semantic constraints.
represents the latent behavior that causes the state change. Training can jointly require that:
- The next observation be predicted from the current observation and latent action.
- Similar state changes produce nearby latent representations.
- Different functional events remain distinguishable.
- Robot actions conditioned on the same latent produce corresponding changes.
However, latent actions are non-identifiable: multiple transformations of the latent variables can produce identical observation predictions. Additional constraints must be provided by objects, time, contact, task outcomes, or robot data.
6. A Behavior Tokenizer Is Not an Ordinary Compressor
A tokenizer maps a behavior trajectory to tokens:
Reading: Behavior Tokenizer T encodes continuous trajectory tau into a length-M sequence of discrete or continuous tokens.
Derivation: A trajectory contains substantial high-frequency detail, and the tokenizer compresses repeatable motion and event structure into a shorter sequence. If the tokens are also intended for planning and cross-embodiment reuse, the tokenizer cannot optimize only for compression ratio; it must also retain object changes, contact boundaries, and task functions.
Good behavior tokens must simultaneously satisfy the following requirements:
- Preserve control-relevant information when reconstructing trajectories.
- Remain reasonably stable across different speeds and scales for behaviors with the same function.
- Distinguish among different functional phases.
- Be predictable by language, vision, or a world model.
- Be executable or adaptable by the target robot after decoding.
The minimum reconstruction loss is:
Reading: The reconstruction loss is the expected squared error between the original trajectory and the trajectory reconstructed by the decoder after encoding with the tokenizer.
Derivation: T first compresses the trajectory into tokens, and D then reconstructs it. Minimizing the error prevents the tokens from discarding all motion information. However, low reconstruction error may merely preserve jitter and absolute coordinates, so semantic constraints from events, tasks, cross-view consistency, or downstream control objectives are also required.
However, low reconstruction error does not imply correct behavioral semantics. A tokenizer may precisely record wrist jitter while losing event boundaries such as “stable grasp” and “release.”
7. Cross-Embodiment Alignment: What Can Be Shared and What Differences Must Be Preserved?
| Shareable layer | Reason | Embodiment differences that must still be preserved |
|---|---|---|
| Objects, action intent, and events | Task semantics are relatively stable | Reachability and grasp geometry |
| Relative displacement and direction | More transferable than absolute coordinates | Scale, speed, and dynamics |
| Skill phases | Grasping, transport, and release share common structure | Number of joints, end-effector shape, and control interface |
| Visual scene representations | Internet-scale data is abundant | Viewpoint, sensors, and task-relevant details |
| Language conditions | Task descriptions can be shared | Robot capabilities and safety constraints |
The goal of cross-embodiment learning is not to make human and robot action values identical, but to identify intermediate variables that can be reinterpreted for different actuators.
8. How Human Video Enters VLAs, World Models, and Hierarchical Policies
| Approach | What human data provides | How it is integrated |
|---|---|---|
| VLA | Objects, language, action phases, and visual priors | Pretraining, auxiliary tasks, and latent-action conditioning |
| World model | State changes and future visual observations | Video prediction, latent dynamics, and future subgoals |
| Hierarchical policy | Task order, events, and skill boundaries | High-level planning, skill tokens, and memory |
| Value learning | Preference and outcome comparisons | Reward models, success prediction, and data filtering |
| Control | Relative motion and contact phases | Target trajectories, impedance parameters, and reference motions |
9. Minimal Experiment: Does Human Video Actually Improve Cross-Embodiment Policies?
Fix the robot policy architecture and the number of robot trajectories, and compare five training settings: robot data only, raw human video added, object-centric motion representations added, latent actions added, and behavior tokens added. Use the same amount of human-side data in each applicable setting, and hold out novel viewpoints, objects, operators, and robot embodiments.
Minimum reporting requirements: robot closed-loop success rate, few-shot adaptation curves, linear probes for objects and phases, transfer to unseen embodiments, slip rate and peak force in contact tasks, a temporal-shuffling ablation on human videos, an appearance-replacement ablation, and the net gain under an equal amount of robot data.
A human video shows “a hand moving a cup to a rack,” and the robot uses a manipulator to perform the same function.
| Baseline | Supervision | Evaluation target |
|---|---|---|
| Robot actions only | Direct behavior cloning | Baseline success rate |
| Human visual pretraining added | Temporal video representations | Novel objects and viewpoints |
| Object-relative trajectories added | Hand–object interaction intermediate variables | Transfer across cameras and scales |
| Latent actions added | State-change conditioning | Adaptation with limited robot data |
| Event/skill labels added | High-level phase supervision | Long-horizon composition and recovery |
The amount of robot data must be held equal, and tasks must be held out, to avoid mistaking greater data scale or repeated objects for gains from cross-embodiment representations.
10. Major Failure Modes
- Treating human pixel velocity directly as robot control velocity.
- Mistaking a shared vocabulary for shared action semantics.
- Latent actions encoding only video appearance rather than controllable changes.
- Correct object detection but missing contact phases and mechanical states.
- Forcing frame-by-frame alignment even though human and robot tasks proceed at different timescales.
- Reporting only representation similarity rather than closed-loop robot gains.
11. Credible Evaluation
- Linear probes of representations: whether objects, phases, and relative motion are predictable.
- Task holdouts across viewpoints, scales, and embodiments.
- Few-shot robot adaptation efficiency.
- Functional success rate of generated actions rather than reconstruction error alone.
- Recovery rate, peak force, and slip rate in contact tasks.
- Ablations that remove human video, shuffle its temporal order, or replace objects.
12. Exercises
- List the observable variables, unobservable variables, and learnable intermediate variables in human video.
- Derive the difference between the conditional probability used for inverse dynamics and that used for video prediction.
- Design an experiment that distinguishes “gains in pixel prediction” from “gains in control.”
- Define training objectives and failure modes for a cross-robot Behavior Tokenizer.
- Represent a grasping phase using the relative hand–object pose, and explain why it is more robust than absolute pixels.
- Explain the intersections among WAM-TTT, human video, and VLAs.
13. Paper Facts, Authors’ Interpretations, and Course Assessments
| Work | Paper facts | Authors’ interpretation | Course assessment |
|---|---|---|---|
| R3M | Pretrains visual representations on large-scale egocentric human video and evaluates them on multiple categories of downstream robot tasks | Interaction structure in human video can produce transferable visual representations for robotics | This demonstrates the benefits of representation pretraining; it does not mean that video directly provides robot actions or contact forces |
| MimicPlay | Uses human play videos to provide high-level guidance, then learns low-level manipulation from a small number of robot demonstrations | Human behavioral priors can support long-horizon robot imitation learning | The key bridge is the combination of high-level motion or goal conditioning with low-level robot data; the gains cannot be attributed entirely to cross-embodiment action alignment |
| XSkill | Studies the discovery of cross-embodiment skills from human and robot videos and applies them to downstream imitation | Shared skill representations can connect different acting bodies | It is necessary to verify that skills preserve task function, event boundaries, and object state rather than merely being close in embedding space |
| Open X-Embodiment | Aggregates multi-robot, multi-task data and trains cross-embodiment policy models | Standardized data formats and large-scale heterogeneous co-training can improve generalization | Shared semantics do not require unified numerical action values; robot identity, action space, and control frequency must still be explicitly preserved |
E1|Video Motion and Object-Centric Representations

