Skip to content

Original Feishu Document · Source Revision 24

💡

Core course in this track: Physical AI data is not limited to robot trajectories. Human videos, hand skeletons, object changes, language descriptions, and failure records can all provide supervision, but they correspond to different mathematical objects. This course examines how to learn representations, motion, latent actions, and transferable cross-robot structure from these data sources.

Learning Objectives ​

After completing this course, you should be able to distinguish among observation supervision, action supervision, state-change supervision, and outcome supervision; explain what information is missing from human videos; understand motion learning through optical flow, keypoints, object-centric representations, and inverse dynamics; design latent actions and behavior tokenizers; analyze the assumptions and failure modes of cross-embodiment alignment; and design verifiable experiments for incorporating human data into VLAs, world models, and hierarchical policies.

Course Canvas

1. What Supervision Does a Data Sample Actually Contain? ​

DataDirectly observableNot directly observableLearnable objects
Robot trajectoriesImages, proprioceptive state, actions, outcomesExpert intent, unexecuted counterfactual actionsPolicies, dynamics, value functions
Human videosImages, changes in hands and objectsRobot control commands, true scale, forces, and contactVisual representations, motion structure, latent actions
Language descriptionsTask and event semanticsPrecise geometry and control timingTask conditions, subgoals, event labels
Hand skeletonsKeypoint positions and topologyContact forces, object weight, robot-joint mappingsGrasp types, relative motion, interaction phases
Failure recordsStates, actions, outcomes, or interventionsThe complete causal chain behind a failureRecovery policies, value functions, data reweighting

“No action labels” does not mean “no supervision,” but the supervision often shifts from direct actions to state changes, temporal ordering, or object-interaction structure.

2. The Observation Model for Human Video ​

A video can be written as:

Reading: A length-T video observation consists of image frames I_1 through I_T arranged in temporal order.

Derivation: A camera initially records a sequence of pixels rather than world coordinates, forces, or robot actions. All subsequent supervision for motion, events, and latent actions must be estimated from these observations or obtained using additional sensors.

If the video contains object keypoints , pixel displacement can be calculated as:

Reading: The pixel displacement of the i-th image keypoint at time t equals its position in the next frame minus its position in the current frame.

Derivation: Keypoint tracking provides two-dimensional coordinates in adjacent frames. Taking their difference yields the direction and magnitude of motion on the image plane. This is a change in observation space and does not include depth scale, camera-motion compensation, or physical units.

However, pixel displacement is not physical velocity:

Reading: Physical velocity in world coordinates cannot be equated with pixel displacement divided by time.

Derivation: Perspective projection maps three-dimensional positions onto two-dimensional pixels. The same world-frame velocity can produce different pixel velocities because of depth, focal length, and camera motion. Pixel changes can be converted into velocity estimates with physical units only after camera calibration, depth estimation, and motion compensation.

This is because camera motion, depth, focal length, occlusion, and perspective all affect pixel changes. Monocular video usually provides only relative-motion or directional cues; it does not directly provide robot end-effector velocity or torque.

2.1 From Images to Motion Representations ​

RepresentationInformation retainedMain ambiguities
Pixel optical flowLocal visual motionCamera motion and depth
Human keypointsBody and hand geometryOcclusion, scale, and 3D depth
Object-centric trajectoriesRelative object displacement and interactionObject identity and contact
Video latentTemporal and semantic structureInterpretability and utility for control
Event labelsPhases such as contact, stable grasp, and releaseLabel boundaries and annotation cost

3. Hand–Object Interaction and Object-Centric Representations ​

Robots need to know “what the hand did relative to the object” more than “how many pixels changed across the entire image.” An object-centric representation can be written as:

Reading: The hand–object relationship consists of relative position, relative orientation, and relative velocity expressed in the object coordinate frame.

Derivation: First, subtract the positions to obtain the hand translation relative to the object, and then multiply by the transpose of the object’s rotation matrix to express this vector in the object coordinate frame. Similarly, relative rotation is obtained by multiplying the inverse of the object orientation by the hand orientation. This removes changes in absolute scene position and more directly describes grasping, insertion, and rotation relationships.

It contains relative position, relative orientation, and relative velocity. Relative representations are more robust to camera translation and changes in absolute coordinates across scenes, and they more closely match the control variables of tasks such as grasping, insertion, and rotation.

3.1 Contact Is Not a Pixel ​

An interaction representation must consider at least the contact location, contact normal, normal force, tangential slip, and object deformation. Visual video often provides only weak cues about contact events and must be learned jointly with force sensing, tactile sensing, or action outcomes.

4. Learning Motion from Video Without Fabricating Robot Actions ​

There are four common approaches:

ApproachSupervision objectiveHow robots use itMain assumption
Video predictionPredict the next frame or future state from past video framesLearn future visual states and subgoalsThe visual future is relevant to the task future
Temporal/contrastive learningCorrect ordering, neighboring clips, or cross-view consistencyPretrain motion representationsThe representation transfers to robot tasks
Inverse dynamicsInfer possible robot actions from changes between adjacent statesInfer latent actions from state changesState changes are sufficient to identify actions
Latent actionsInfer latent behavioral variables that explain state changes from adjacent observationsEncode video changes as behavior tokensLatent actions share semantics across embodiments

The key distinction is that video can teach a robot “what changed” or “where it should move next,” but it does not automatically teach the robot “how much force to apply, which joint trajectory to follow, or what control frequency to use.”

5. Latent Actions: Inferring Action Variables from State Changes ​

If the true action is unobserved, a latent variable can be introduced:

Reading: The inference model estimates latent action z_t from adjacent observations, and the generative model then predicts the next observation from the current observation and that latent action.

Derivation: When the true action is unobserved, the factor responsible for the state change can be treated as a latent variable. The encoder infers z_t from an observation pair, while the decoder requires z_t to contain enough information to explain the next state. However, any invertible reparameterization may yield the same prediction, so objects, events, or robot actions are also needed to provide semantic constraints.

represents the latent behavior that causes the state change. Training can jointly require that:

  • The next observation be predicted from the current observation and latent action.
  • Similar state changes produce nearby latent representations.
  • Different functional events remain distinguishable.
  • Robot actions conditioned on the same latent produce corresponding changes.

However, latent actions are non-identifiable: multiple transformations of the latent variables can produce identical observation predictions. Additional constraints must be provided by objects, time, contact, task outcomes, or robot data.

6. A Behavior Tokenizer Is Not an Ordinary Compressor ​

A tokenizer maps a behavior trajectory to tokens:

Reading: Behavior Tokenizer T encodes continuous trajectory tau into a length-M sequence of discrete or continuous tokens.

Derivation: A trajectory contains substantial high-frequency detail, and the tokenizer compresses repeatable motion and event structure into a shorter sequence. If the tokens are also intended for planning and cross-embodiment reuse, the tokenizer cannot optimize only for compression ratio; it must also retain object changes, contact boundaries, and task functions.

Good behavior tokens must simultaneously satisfy the following requirements:

  • Preserve control-relevant information when reconstructing trajectories.
  • Remain reasonably stable across different speeds and scales for behaviors with the same function.
  • Distinguish among different functional phases.
  • Be predictable by language, vision, or a world model.
  • Be executable or adaptable by the target robot after decoding.

The minimum reconstruction loss is:

Reading: The reconstruction loss is the expected squared error between the original trajectory and the trajectory reconstructed by the decoder after encoding with the tokenizer.

Derivation: T first compresses the trajectory into tokens, and D then reconstructs it. Minimizing the error prevents the tokens from discarding all motion information. However, low reconstruction error may merely preserve jitter and absolute coordinates, so semantic constraints from events, tasks, cross-view consistency, or downstream control objectives are also required.

However, low reconstruction error does not imply correct behavioral semantics. A tokenizer may precisely record wrist jitter while losing event boundaries such as “stable grasp” and “release.”

7. Cross-Embodiment Alignment: What Can Be Shared and What Differences Must Be Preserved? ​

Shareable layerReasonEmbodiment differences that must still be preserved
Objects, action intent, and eventsTask semantics are relatively stableReachability and grasp geometry
Relative displacement and directionMore transferable than absolute coordinatesScale, speed, and dynamics
Skill phasesGrasping, transport, and release share common structureNumber of joints, end-effector shape, and control interface
Visual scene representationsInternet-scale data is abundantViewpoint, sensors, and task-relevant details
Language conditionsTask descriptions can be sharedRobot capabilities and safety constraints

The goal of cross-embodiment learning is not to make human and robot action values identical, but to identify intermediate variables that can be reinterpreted for different actuators.

8. How Human Video Enters VLAs, World Models, and Hierarchical Policies ​

ApproachWhat human data providesHow it is integrated
VLAObjects, language, action phases, and visual priorsPretraining, auxiliary tasks, and latent-action conditioning
World modelState changes and future visual observationsVideo prediction, latent dynamics, and future subgoals
Hierarchical policyTask order, events, and skill boundariesHigh-level planning, skill tokens, and memory
Value learningPreference and outcome comparisonsReward models, success prediction, and data filtering
ControlRelative motion and contact phasesTarget trajectories, impedance parameters, and reference motions

9. Minimal Experiment: Does Human Video Actually Improve Cross-Embodiment Policies? ​

Fix the robot policy architecture and the number of robot trajectories, and compare five training settings: robot data only, raw human video added, object-centric motion representations added, latent actions added, and behavior tokens added. Use the same amount of human-side data in each applicable setting, and hold out novel viewpoints, objects, operators, and robot embodiments.

Minimum reporting requirements: robot closed-loop success rate, few-shot adaptation curves, linear probes for objects and phases, transfer to unseen embodiments, slip rate and peak force in contact tasks, a temporal-shuffling ablation on human videos, an appearance-replacement ablation, and the net gain under an equal amount of robot data.

A human video shows “a hand moving a cup to a rack,” and the robot uses a manipulator to perform the same function.

BaselineSupervisionEvaluation target
Robot actions onlyDirect behavior cloningBaseline success rate
Human visual pretraining addedTemporal video representationsNovel objects and viewpoints
Object-relative trajectories addedHand–object interaction intermediate variablesTransfer across cameras and scales
Latent actions addedState-change conditioningAdaptation with limited robot data
Event/skill labels addedHigh-level phase supervisionLong-horizon composition and recovery

The amount of robot data must be held equal, and tasks must be held out, to avoid mistaking greater data scale or repeated objects for gains from cross-embodiment representations.

10. Major Failure Modes ​

  • Treating human pixel velocity directly as robot control velocity.
  • Mistaking a shared vocabulary for shared action semantics.
  • Latent actions encoding only video appearance rather than controllable changes.
  • Correct object detection but missing contact phases and mechanical states.
  • Forcing frame-by-frame alignment even though human and robot tasks proceed at different timescales.
  • Reporting only representation similarity rather than closed-loop robot gains.

11. Credible Evaluation ​

  1. Linear probes of representations: whether objects, phases, and relative motion are predictable.
  2. Task holdouts across viewpoints, scales, and embodiments.
  3. Few-shot robot adaptation efficiency.
  4. Functional success rate of generated actions rather than reconstruction error alone.
  5. Recovery rate, peak force, and slip rate in contact tasks.
  6. Ablations that remove human video, shuffle its temporal order, or replace objects.

12. Exercises ​

  1. List the observable variables, unobservable variables, and learnable intermediate variables in human video.
  2. Derive the difference between the conditional probability used for inverse dynamics and that used for video prediction.
  3. Design an experiment that distinguishes “gains in pixel prediction” from “gains in control.”
  4. Define training objectives and failure modes for a cross-robot Behavior Tokenizer.
  5. Represent a grasping phase using the relative hand–object pose, and explain why it is more robust than absolute pixels.
  6. Explain the intersections among WAM-TTT, human video, and VLAs.

13. Paper Facts, Authors’ Interpretations, and Course Assessments ​

WorkPaper factsAuthors’ interpretationCourse assessment
R3MPretrains visual representations on large-scale egocentric human video and evaluates them on multiple categories of downstream robot tasksInteraction structure in human video can produce transferable visual representations for roboticsThis demonstrates the benefits of representation pretraining; it does not mean that video directly provides robot actions or contact forces
MimicPlayUses human play videos to provide high-level guidance, then learns low-level manipulation from a small number of robot demonstrationsHuman behavioral priors can support long-horizon robot imitation learningThe key bridge is the combination of high-level motion or goal conditioning with low-level robot data; the gains cannot be attributed entirely to cross-embodiment action alignment
XSkillStudies the discovery of cross-embodiment skills from human and robot videos and applies them to downstream imitationShared skill representations can connect different acting bodiesIt is necessary to verify that skills preserve task function, event boundaries, and object state rather than merely being close in embedding space
Open X-EmbodimentAggregates multi-robot, multi-task data and trains cross-embodiment policy modelsStandardized data formats and large-scale heterogeneous co-training can improve generalizationShared semantics do not require unified numerical action values; robot identity, action space, and control frequency must still be explicitly preserved

E1|Video Motion and Object-Centric Representations

E2|Latent Actions and Inverse Dynamics

E3|Behavior Tokenizer

E4|Cross-Embodiment Alignment and Heterogeneous Co-Training

Article text is licensed under the Apache License 2.0