Original Feishu Document · Source Revision 34
🎥
Core assessment: Ordinary human videos do not directly provide hand-control velocities or robot action labels. Models can estimate image motion, relative trajectories, and action phases from temporal sequences, but real robot velocities, forces, and control constraints must still be grounded through robot data and controllers.
This article connects two questions: the mathematical example of a continuous action probability density, and how human behavior data can be transferred into executable robot control.
1. What Data Is Actually Available in Human Videos
The most basic form of ordinary human video data is:
Reading: Video data V consists of T image frames and their corresponding timestamps.
Derivation: A camera directly records optical observations that change over time. Each frame I_t must be paired with its timestamp tau_t to calculate inter-frame intervals and apparent motion; the data still contains no human control variables, robot actions, forces, or contact labels.
Here, is the -th image frame, and is its corresponding timestamp. Video directly records optical observations, not action commands.
| Potentially available in video | Typically unavailable in video |
|---|---|
| Image frames and temporal order | Actual human joint-control commands |
| Visual positions of hands and objects | Robot joint velocities and torques |
| Object states and visible motion | Actual contact forces and friction |
| Action rhythm and task phases | Precise 3D scale and depth |
| Occlusion, co-motion, and outcomes | Actuator constraints and controller states |
Measurements closer to actual 3D human motion can be obtained only when the capture system additionally provides RGB-D, multiview, IMU, motion-capture, or glove signals.
📌
Key boundary: A hand trajectory is a behavioral observation in a video, not an action that can be sent directly to a robot.
2. How to Estimate Motion and Velocity from Image Sequences
Extended Topic: From Hand Keypoints to Hand–Object Interaction Representations
Hand velocity is not a raw label. A palm coordinate frame, grasp type, and relative motion can first be recovered from a 21-keypoint skeleton, then combined with object pose, contact probability, and changes in object state to construct a transferable representation.
E1|Video Motion and Object-Centric Representations
E1.5|Spatial State Estimation and 3D Geometry
2D Pixel Velocity
If a video tracker obtains the pixel position of a hand or object, finite differences can be used:
Reading: The estimated pixel velocity equals the difference between pixel positions in adjacent frames divided by the difference between their timestamps.
Derivation: Finite differences approximate a time derivative using discrete positions. The numerator is measured in pixels and the denominator in seconds, so the result is measured in pixels per second. Keypoint noise is amplified by differencing, so tracking, camera-motion compensation, and smoothing are required before practical use.
Its unit is pixel/s, and it represents only apparent motion in the image plane.
3D Velocity
If depth, camera calibration, or multiview reconstruction is available, the 3D position can be estimated:
Reading: The estimated 3D velocity equals the difference between 3D positions at adjacent time steps divided by the time interval.
Derivation: A positional difference has physical units such as meters only after depth, camera calibration, or multiview reconstruction provides a shared 3D coordinate system. If X_t still has scale ambiguity, the velocity can likewise represent only velocity at a relative scale.
Why Frames Cannot Simply Be Differenced Directly
Keypoint estimates contain discontinuities and occlusions, and direct differencing amplifies noise. A practical pipeline is typically:

A Kalman Filter, Savitzky–Golay filter, spline curve, or temporal neural network can be used to smooth the trajectory. Position and velocity confidence should both be retained in the output.
3. Fundamental Ambiguities in Monocular Video
Inferring 3D motion from 2D images in monocular video is an underdetermined problem: multiple real motions may produce similar images.
| Ambiguity | Consequence |
|---|---|
| Depth and scale | The number of centimeters corresponding to a pixel displacement is unknown |
| Motion toward the camera | The actual displacement may be large even when the 2D displacement is small |
| Camera motion | Head or camera shake may be mistaken for hand motion |
| Occlusion | Fingers and contact points may be invisible at critical moments |
| Motion blur | Fast actions make keypoint estimation unstable |
| Inconsistent frame rates | Velocities cannot be compared without reliable timestamps |
Monocular video can therefore provide the following with reasonable reliability:
- motion direction;
- relative speed;
- changes in hand–object distance;
- whether an object moves together with the hand;
- action phases and event order.
However, the estimates cannot be directly interpreted as precise velocities in meters per second, contact forces, or robot control commands.
4. How Video Models Learn Motion Without Velocity Labels
A model does not necessarily require explicit velocity labels. As long as the input includes temporal order, it can learn motion representations from changes between frames.
Video Prediction
Reading: Given past images and language conditions, the video model represents the joint distribution of the next H frames.
Derivation: To predict future pixels, the model must use object persistence, motion direction, velocity rhythm, and event phase. However, pixel likelihood constrains only the visual future; it does not automatically recover actual forces, robot joint commands, or executability.
To predict future frames, the model must implicitly estimate where an object is moving, how quickly it is moving, and which phase of the action is underway.
Contrastive and Temporal-Order Learning
A model can learn that adjacent frames are more similar than temporally distant frames, or determine the order of frames, thereby forming a representation of temporal progress.
Trajectory Prediction
Methods such as ATM directly predict the future trajectories of visual points:
Reading: Given historical images, current visual points, and a language task, the trajectory model predicts the positions of those points over the next H steps.
Derivation: Compared with generating an entire image, point trajectories retain only task-relevant motion. They can provide supervision for direction and relative path, but object identity, occlusion confidence, and robot data are still needed to map visual trajectories to executable actions.
Boundaries of Implicit Learning
A video model can learn “what motion looks like,” but it does not automatically recover:
- human muscle or joint-control variables;
- actual contact forces;
- the velocities and torques required by robot actuators;
- whether an action satisfies the target robot’s dynamics.
WAM-TTT follows this same approach: human videos write visual dynamics into memory through video prediction, while the robot Action Expert is still grounded using paired robot action data.
5. Why Human Hand Velocity Cannot Be Used Directly as a Robot Action
Even if human hand velocity can be estimated accurately:
Even if the velocity of a human hand in world coordinates or normalized coordinates can be estimated from video, this still does not provide the target robot’s action command.
One still cannot directly set:
Reading: Robot velocity cannot be directly equated with human hand velocity.
Derivation: Their coordinate systems, scales, degrees of freedom, control frequencies, dynamics, and safety constraints differ. What can be shared is task function, relative direction, and event timing; the specific velocity must be solved again using the robot state, controller, and robot action data.
This is because humans and robots are separated by an embodiment gap:
- arm lengths and joint topologies differ;
- workspaces and singular configurations differ;
- grasp types differ among human hands, two-finger grippers, and dexterous hands;
- maximum velocities, accelerations, and torques differ;
- control frequencies and latencies differ;
- payload, compliance, and collision constraints differ.

Human data is better suited to supervising “what to do, which relative path to follow, and when to make contact”; robot data supervises “the specific velocity, position, and force that the current embodiment should output.”
6. Recommended Object-Centric and Normalized Representations
For cross-embodiment sharing, object-relative quantities should be represented in preference to absolute velocities in human world coordinates.
Object-Relative Displacement
Reading: The position of the hand relative to the object equals the hand position minus the object position.
Derivation: Translating both the hand and object by the same amount does not change this difference, so relative position is better suited than absolute scene coordinates for describing approach, grasping, and placement. For cross-view stability, the difference vector should also be transformed into the object coordinate frame.
Scale-Normalized Velocity
Reading: The scale-normalized relative velocity equals the change in the hand–object relative position divided by the time interval and the object scale.
Derivation: Differencing the relative position first removes the object’s own motion, dividing by time produces relative velocity, and finally dividing by object scale removes size differences. It expresses “how many object scales are traversed per second,” rather than meters per second.
Here, is the object scale. This quantity represents how far the hand moves relative to the object’s size per unit time.
Direction and Speed Categories
When data quality is limited, exact numerical regression can be avoided in favor of:
[APPROACH direction=left speed=slow]
[CONTACT speed=decelerating]
[LIFT speed=steady]
[PLACE speed=slowing_down]Task Phase
Each behavior segment can also be mapped to normalized progress:
Reading: The task-phase variable phi_t lies between zero and one, where zero denotes the start of the phase and one denotes its end.
Derivation: Normalizing demonstrations of different durations to a common interval based on event boundaries makes it possible to compare action rhythm and phase progress. However, the same phase may correspond to different paths and velocities, so it cannot be treated as a control variable.
However, phase can represent only task progress and cannot replace actual velocity. A more robust Prompt includes the object trajectory, event boundaries, and relative rhythm together.
7. What Additional Signals Are Required for Contact-Rich Tasks?
For free-space approach and ordinary transport, RGB trajectories may already provide strong supervision. Video alone is generally insufficient, however, for the following tasks:
- insertion and assembly;
- twisting bottle caps and screws;
- wiping and scraping;
- deformable-object manipulation;
- grasping fragile objects;
- tool use that requires normal-force control.
| Additional signal | Problem addressed |
|---|---|
| RGB-D / multiview | 3D scale, depth, and occlusion |
| Camera or head pose | Compensation for egocentric camera motion |
| Glove / IMU / Mocap | High-frequency finger and arm motion |
| Object 6D Pose | Object-relative motion and rotation |
| Robot force/torque | Contact-force and dynamics grounding |
| Tactile sensing | Slip, contact distribution, and grasp stability |
A pragmatic approach is to use large volumes of low-cost RGB human video to learn behavioral structure, then use a small amount of high-quality multimodal human data and real robot interactions for calibration.
8. Recommended Data and Training Pipeline
Recommended Division of Training Responsibilities
- Pretrain motion, object-relation, and task-phase encoders on human videos.
- Use a small amount of paired human–robot data to align identical tasks and events.
- Train the actual Action Head under robot action supervision.
- Use robot force and tactile data to train contact phases.
- During deployment, use the human Prompt to specify behavioral structure and the robot’s current state to determine the specific velocity.
Label Confidence
Velocities derived from human videos should be stored as estimates:
{
"motion_direction": "left",
"normalized_speed": 0.42,
"speed_confidence": 0.68,
"source": "monocular_track_v2",
"metric_scale_available": false
}Monocular estimates must not be treated as true metric velocity.
9. How the One-Dimensional Gaussian Example Should Be Rewritten
The original one-dimensional Gaussian velocity example is suitable for explaining continuous probability density, but in the context of human video it may lead readers to mistakenly assume that the data contains actual hand velocities.
More Rigorous Formulation 1: Robot End-Effector Velocity
When discussing robot action data, it can be explicitly written as:
Reading: Under condition c, the robot’s one-dimensional velocity is modeled as a Gaussian random variable with mean mu and variance sigma squared.
Derivation: Robot control logs provide velocity samples with physical units. Describing their conditional distribution using the sample mean and variance is useful for teaching continuous probability density and uncertainty. A Gaussian is only a simplified model; multimodal actions should use mixture distributions or generative policies.
Here, the velocity comes from robot control logs and has actual units and action semantics.
More Rigorous Formulation 2: Normalized Relative Velocity from Human Video
When discussing human video, it should be written as:
Reading: Under condition c, the normalized relative velocity derived from human video is approximated by a Gaussian distribution.
Derivation: Here, the samples are computed from tracking, timestamps, and object scale, and contain estimation error rather than coming from actual control logs. The distribution describes relative motion rhythm, and both confidence and the availability of metric scale should be stored.
And define:
Here, the scale-normalized relative velocity definition from Section 6 is used: first compute the change in the position of the hand relative to the object, then divide by the time interval and object scale.
It expresses the rhythm of hand motion relative to the object’s scale, not the absolute velocity that the robot needs to execute.
A Simpler Teaching Example
If the goal is only to explain probability density, a “normalized one-dimensional action variable” can be used directly:
Reading: Under condition c, the normalized one-dimensional behavior variable z is modeled as a Gaussian distribution.
Derivation: Writing the teaching variable as a unitless z keeps the focus on probability density, mean, and variance without implying that human videos contain robot velocity labels. z may represent a normalized displacement, trajectory coefficient, or latent action.
Here, can be a normalized position increment, trajectory coefficient, or latent action, avoiding any implication of a specific data modality.
10. Final Conclusions
| Question | Answer |
|---|---|
| Do human videos contain hand-velocity labels? | Ordinary RGB videos contain no direct labels; velocity can only be estimated from temporal sequences |
| Can a model learn motion? | It can learn image motion, relative trajectories, rhythm, and task phases |
| Can actual 3D velocity be recovered? | Usually not from monocular video; depth, calibration, multiview data, or wearable devices are required |
| Can human velocity be sent directly to a robot? | No; the embodiment, actuators, and dynamics differ |
| Where does the robot’s actual velocity come from? | Robot demonstrations, the Action Head, the controller, and closed-loop feedback |
| What should be shared across embodiments? | Object relations, path directions, event order, and normalized rhythm |
✅
Final conclusion: Human videos provide “what motion looks like and how the task unfolds”; robot data provides “how the current embodiment should execute it with the correct velocity, force, and control frequency.” They are connected through object-centric behavior representations, not by directly copying human velocity.
11. Minimal Experiment: Does Video Motion Estimation Actually Help the Robot?
Collect a set of manipulation tasks for which monocular RGB, RGB-D or multiview 3D trajectories, and robot reproduction data are all available. Provide only monocular RGB to the training pipeline, and reserve the 3D trajectories as evaluation ground truth rather than training labels.
- Compare raw frame differencing, tracking plus smoothing, camera-motion compensation, object-centric relative trajectories, and trajectory-prediction models.
- Predict pixel velocity, scale-normalized relative velocity, action phase, and future point trajectories separately.
- Use these representations as pretraining or conditioning inputs for the same robot policy while holding the amount of robot demonstration data and the controller fixed.
- Test under novel viewpoints, novel object scales, moving cameras, and object self-motion.
- Add force or tactile controls for contact tasks to verify the boundaries of RGB representations.
Minimum reporting requirements: 2D tracking error, 3D direction error, scale-normalized velocity error, phase recognition, confidence calibration, robot closed-loop success rate, slip rate, peak force, and net improvement under the same amount of robot data.
12. Major Failure Modes
| Failure | Symptom | Diagnosis and correction |
|---|---|---|
| Treating pixel velocity as physical velocity | Values become invalid after changing depth, focal length, or cropping | Report units, camera calibration, and whether metric scale is available |
| Camera-motion leakage | Stationary objects are predicted to move together | Apply camera-motion compensation, use background-point controls, and perform fixed-camera ablations |
| Noise amplified by differencing | Slight keypoint jitter produces enormous velocity spikes | Apply trajectory smoothing, outlier rejection, and velocity confidence intervals |
| Ignoring object self-motion | Large relative velocity is still estimated when the hand and object move together | Difference the hand–object relative position rather than only the hand position |
| Occlusion and identity switching | The tracker switches hands or objects, or incorrectly connects trajectories across occlusions | Enforce object-identity consistency, segment trajectories, and apply uncertainty gating |
| Direct copying across embodiments | Mapping human trajectories to robot velocities causes collisions, limit violations, or instability | Condition on robot state and enforce reachability, inverse-kinematics, and control constraints |
| Evaluating only video metrics | Trajectory prediction improves, but robot task performance does not | Hold robot data and control budgets fixed and report actual closed-loop gains |
| Missing contact information | Visual trajectories are correct for insertion, wiping, or fragile-object grasping, but force control fails | Evaluate force sensing, tactile sensing, contact events, and slip rate |
13. Paper Facts, Author Interpretations, and Course Assessments
| Work | Paper facts | Author interpretation | Course assessment |
|---|---|---|---|
| R3M | Learns visual representations from large-scale egocentric human videos and transfers them to robot manipulation tasks | Human videos contain reusable interaction and object information | This demonstrates representation transfer, not the recovery of actual robot velocities or forces from video |
| MimicPlay | Uses human play videos to form high-level guidance and combines them with a small number of robot demonstrations to learn long-horizon manipulation | Human videos can provide task structure and motion priors | Robot action grounding still comes from robot data; the amount of robot demonstration data should be held fixed to verify the net benefit |
| ATM | Predicts the future trajectories of arbitrary visual points conditioned on video and uses those trajectories for policy learning | Point trajectories express motion intent more compactly than full videos | Trajectories provide observation-space supervision; object identity, occlusion, contact, and robot reachability still require additional handling |
| WAM-TTT | Uses auxiliary human-video signals to update fast memory at test time, while robot action task supervision updates the learning rule | Human demonstrations in the current environment can help robots adapt quickly | The critical evidence consists of strict holdouts, incorrect-demonstration controls, human–robot phase alignment, and improved robot closed-loop performance |
14. Exercises and Cross-Reading
- Explain why the same world-space velocity produces different pixel velocities at different depths.
- Derive the hand–object relative velocity and explain why the object’s own motion must be subtracted.
- Design an experiment that distinguishes camera motion from object motion.
- Explain what normalized phase, relative velocity, and robot action each represent.
- Design an ablation that demonstrates that video pretraining improves robot closed-loop performance rather than only video prediction.
E0|Data, Representations, and Cross-Embodiment Learning
E1|Video Motion and Object-Centric Representations
E1.5|Spatial State Estimation and 3D Geometry

