Skip to content

Original Feishu Document · Source Revision 15

💡

Mechanisms lesson: Starting from video observation models, this lesson explains optical flow, keypoints, 3D motion, object trajectories, and hand-object interaction graphs, and identifies which quantities can transfer across viewpoints and embodiments.

Learning Objectives ​

After completing this lesson, you should be able to distinguish pixel motion from physical motion; derive keypoint velocities and error propagation; construct object-centric and hand-object relative representations; understand monocular scale, occlusion, and camera motion; and design evaluations of the benefits of video representations for robot control.

1. Video Observation Model ​

A 3D point is projected through a camera:

Interpretation: The 3D point is first transformed by the camera extrinsics and intrinsics into homogeneous image coordinates, then converted to pixel coordinates through perspective division.

Derivation: R_t and T_t transform the world point into camera coordinates, and K maps it to homogeneous pixel coordinates; Pi then divides by the depth component. Because the final step depends on depth, the same 3D displacement does not correspond to a fixed proportional pixel displacement.

Pixel changes are simultaneously affected by object motion, camera motion, depth, and intrinsics. Therefore:

Interpretation: Pixel displacement cannot, in general, be represented as 3D displacement multiplied by a fixed constant c.

Derivation: The local scale of perspective projection varies with depth, focal length, and camera pose; the camera's own motion also contributes to pixel differences. A local linear approximation is therefore possible only when local depth is approximately constant and the camera has been calibrated and its motion compensated for.

Unless the camera parameters and depth are known, pixel velocity cannot be treated directly as physical velocity.

2. Optical Flow ​

Brightness constancy assumption:

Interpretation: After an image point moves by u and v, its brightness remains approximately constant over a short interval Delta t.

Derivation: If the illumination of the same surface point does not change significantly between adjacent frames, its new location should have a similar intensity. Reflections, shadows, occlusions, and contact deformation violate this assumption, so optical flow must be combined with robust losses or learned priors.

A first-order expansion yields the optical flow constraint:

Interpretation: The image brightness gradients along x, y, and time satisfy a first-order optical flow constraint with the 2D displacements u and v.

Derivation: Applying a first-order Taylor expansion to the brightness constancy equation at the current point and subtracting I(x,y,t) yields a sum of I_xu, I_yv, and I_t Delta t that is approximately zero. One equation contains two unknown motion components, so neighborhood smoothness, multiple pixels, or a learned model is also required.

One equation has two unknowns, requiring local smoothness or a deep model. Reflections, occlusions, and contact deformation violate the assumption.

3. Keypoints and Velocity ​

Given keypoint position estimates , the finite-difference velocity is:

Interpretation: The keypoint velocity estimate equals the difference between two consecutive position estimates divided by the inter-frame interval.

Derivation: This is the forward finite-difference approximation of the time derivative of position. It is simple, but it introduces the position errors from both frames into the numerator. Consequently, if the frame rate increases while localization accuracy remains unchanged, velocity noise may actually increase.

If the independent position noise has variance :

Interpretation: If the position noise in two adjacent frames is independent and has variance sigma squared in each frame, the finite-difference velocity has variance equal to twice sigma squared divided by the square of the time interval.

Derivation: The velocity error is epsilon_{t+1} minus epsilon_t, divided by Delta t. The variance of the difference between independent noise terms equals the sum of their variances, yielding 2 sigma squared; if the noise is temporally correlated, a negative twice-covariance term must also be included.

The smaller the inter-frame interval, the more direct differencing amplifies noise, making smoothing, state-space filtering, or trajectory fitting necessary.

Course whiteboard

4. Hand Skeleton ​

The set of hand keypoints is . It can be used to construct:

  • A palm-centered coordinate frame.
  • Finger-joint flexion angles.
  • Distances between fingertips and objects.
  • Grasp types and contact candidates.

A skeleton is more structured than pixels, but it does not encode actual forces, skin contact area, or object weight.

5. Object-Centric Representations ​

Relative position:

Interpretation: The hand's position vector relative to the object equals the hand position minus the object position.

Derivation: A common scene translation cancels under subtraction, so r_t provides a more stable description of approach, retreat, and grasping relationships. For transfer across viewpoints, this vector should also be expressed in an object-centered or camera-calibrated coordinate frame.

Relative orientation:

Interpretation: The rotation of the hand relative to the object equals the inverse of the object rotation multiplied by the hand rotation.

Derivation: A rotation matrix is orthogonal, so its inverse equals its transpose. First, the inverse object rotation transforms world directions into the object coordinate frame; multiplying by the hand rotation then yields the hand's relative orientation in the object coordinate frame.

Relative representations are better suited to transfer across cameras and scenes than absolute image coordinates. A hand-object interaction graph can also be constructed, with nodes representing the hand, objects, and environment, and edges representing distance, contact, and motion relationships.

6. Events and Phases ​

PhaseVisual cuesMissing information
ApproachDistance decreasesGrasp intent
ContactAbrupt change in relative motionContact force
Stable graspHand and object move togetherInternal slip
TransportObject trajectory follows the handLoad and impedance
ReleaseHand and object separateObject stability

Event-based representations are often better suited than fixed temporal windows for defining skill and Tokenizer boundaries.

7. 3D Reconstruction ​

Multiple views, depth cameras, human body models, and object models can reduce ambiguity. Monocular 3D reconstruction still depends on scale, priors, and camera assumptions. Evaluations should report 2D accuracy, 3D accuracy, and control utility separately rather than directly equating reconstruction accuracy with usability for robotics.

8. Minimal Experiment ​

Use the same set of videos with RGB, camera poses, and either depth or multi-view 3D ground truth, while providing only RGB as model input. Compare optical flow, keypoint differencing, smoothed trajectories, object-centric relative representations, and 3D reconstruction, then connect each representation to the same few-shot robot policy.

Minimum reporting requirements: optical flow endpoint error, keypoint tracking error, velocity direction and variance calibration, relative-pose error, event-boundary F1, recovery from occlusion, held-out novel viewpoints and moving-camera sequences, and closed-loop success rate under a fixed number of robot demonstrations.

  1. Extract hand and object keypoints.
  2. Compare direct differencing, smoothing, and Kalman-filtered velocities.
  3. Compare absolute coordinates with object-relative coordinates.
  4. Train a contact-phase classifier.
  5. Use the representation in a few-shot robot policy.
  6. Hold out camera motion, scale, and occlusion conditions.

9. Exercises ​

  1. Derive the noise variance of finite-difference velocity.
  2. Explain the optical flow aperture problem.
  3. Construct a hand-object interaction graph for grasping.
  4. Design a counterexample showing that “moving together does not imply a stable grasp.”
  5. Compare 2D, 3D, and object-centric representations.
  6. Explain how video representations connect to VLA models, world models, and hierarchical skills.

10. Major Failure Modes ​

FailureManifestationDiagnosis and correction
Brightness constancy violationReflections, shadows, or rapid exposure changes produce erroneous optical flowPhotometric augmentation, robust losses, and learned matching
Aperture problemOnly normal motion can be observed along a long edge, leaving tangential velocity uncertainBroader spatial context, multiple corners, or global matching
Keypoint noise amplificationSmall position errors cause severe jitter in finite-difference velocityReport noise covariance; use filtering and trajectory fitting
Identity switchingTracking switches to another hand or object after occlusionLong-term identity consistency, re-identification, and segment-level confidence
Camera motion mixed with object motionThe entire scene exhibits optical flow in a similar directionCamera-pose estimation, background-point compensation, and fixed-camera controls
Monocular scale driftThe 3D trajectory has a plausible shape, but its physical scale changes over timeKnown object scale, depth, multiple views, or scale-invariant metrics
Contact missing from the object-centric graphThe hand approaches the object, but the representation cannot determine whether the grasp is stable or slippingContact events, force sensing, tactile sensing, and object-state changes
Representation metrics decoupled from control2D/3D errors decrease, but robot success rate remains unchangedFreeze the controller and data budget; report downstream closed-loop ablations

11. Paper Facts, Author Interpretations, and Course Assessments ​

WorkPaper factsAuthor interpretationCourse assessment
RAFTEstimates dense optical flow using all-pairs correlation volumes and iterative updatesGlobal matching and repeated refinement can improve the estimation of complex motionHigh optical flow accuracy remains pixel-space evidence; it does not imply recovery of 3D velocity, contact, or actions
TAPIRPerforms long-term point tracking for arbitrary query points and explicitly predicts occlusion and uncertaintyPoint trajectories and visibility are well suited to describing object motion in long videosLong-term identity and occlusion modeling are useful for robotics, but they must also be grounded in object semantics, events, and control
R3MPretrains representations from egocentric human videos and transfers them to robotic manipulationInteraction structure in video can produce general-purpose visual representationsBenefits must be demonstrated on closed-loop tasks with a fixed amount of robot data, rather than only through representation probes
XSkillLearns cross-embodiment skill representations from human and robot videosA shared skill space can connect different bodiesEmbedding alignment does not guarantee action executability; event boundaries, object states, and few-shot reuse on robots must be validated

Human Videos Have No Velocity Labels

12. Cross-Reading ​

Human Videos Have No Velocity Labels

E1.5|Spatial State Estimation and 3D Geometry

E2|Latent Actions and Inverse Dynamics

E3|Behavior Tokenizer

Palm Skeleton Lines and Hand-Object Interaction Representations

Article text is licensed under the Apache License 2.0