Skip to content

Original Feishu Document · Source Revision 8

💡

Mechanisms lesson: A VLA can output actions directly from images, but a physical closed loop still implicitly requires position, orientation, velocity, occlusion, and uncertainty. This lesson establishes a unified language spanning Bayes filtering, SE(3), depth and point clouds, object-centric state, and affordances.

Learning Objectives ​

After completing this lesson, you should be able to distinguish observations from states; derive the prediction and correction steps of a Bayes filter; interpret poses, coordinate transformations, depth back-projection, and object-centric representations; explain why end-to-end policies still require state identifiability; and design experiments involving occlusion, camera changes, and calibration errors.

1. An Image Is Not a State ​

The state a robot actually needs may include its own pose, joint velocities, the six-degree-of-freedom poses of objects, contact modes, and obstacles. A camera provides only observations of these variables after projection, occlusion, and noise:

Interpretation: The current observation is generated stochastically from the current true state through the observation model.

Derivation: The same state produces different images under different lighting conditions and cameras; the same image may also correspond to different depths and occlusion relationships. Uncertainty must therefore be represented with probability distributions.

2. Bayes Filter: How History Becomes a Belief State ​

Prediction step:

Interpretation: Propagate all possible states from the previous step to the current time according to the action-conditioned dynamics, then aggregate them to obtain the predicted belief before incorporating the new observation.

Derivation: By the law of total probability, integrate over the unknown previous state; the dynamics provide the conditional probability of transitioning from the previous state to the current state.

Correction step:

Interpretation: Reweight the predicted belief by the likelihood of the current observation under each candidate state, then use a normalization constant to ensure that the probabilities sum to one.

Derivation: Apply Bayes' rule directly: the prior is the predicted belief, the likelihood is the observation model, and the posterior is the corrected belief state.

3. The Kalman Filter Is the Linear-Gaussian Special Case ​

Interpretation: The state evolves according to linear dynamics and control inputs, with additive process noise; the observation is a linear projection of the state, with additive observation noise.

Derivation: Approximate the nonlinear dynamics near the operating point with matrices A and B, and approximate the sensor with matrix C. If both process noise and observation noise are Gaussian, linear transformations and products of Gaussians remain Gaussian, so the belief need only maintain a mean and covariance.

Under linear-Gaussian assumptions, the belief always remains a Gaussian distribution, so only its mean and covariance must be maintained. Nonlinear systems use EKF/UKF, while strongly multimodal distributions and discrete contact modes often require particle filters or learned filters.

4. 3D Poses and SE(3) ​

A rigid-body pose is represented by a homogeneous transformation:

Interpretation: Matrix T consists of a 3D rotation R and a 3D translation t and is used to transform a point from one coordinate frame to another.

Derivation: A 3D point is first rotated and then translated. With homogeneous coordinates, rotation and translation can be combined into a single matrix multiplication, while successive coordinate transformations are composed through matrix multiplication.

5. Back-Projecting Pixels and Depth into 3D ​

In a pinhole camera, pixel and depth correspond to the following camera coordinates:

Interpretation: Multiply the pixel offset from the principal point by the depth, then divide by the focal length to obtain the horizontal and vertical positions in the camera coordinate frame.

Derivation: By similar triangles, the normalized image coordinates equal the 3D coordinates divided by depth; rearranging gives the back-projection equations. Depth error directly amplifies 3D position error as distance increases.

Formula Visualization|From Pixel and Depth to World Coordinates ​

Course Whiteboard

6. Point Clouds, Voxels, Occupancy, and Neural Fields ​

RepresentationAdvantagesLimitations
Point cloudPreserves direct 3D measurementsSparse, irregular, and strongly affected by occlusion
Voxel/occupancy gridConvenient for collision checking and spatial queriesHigh resolution is computationally expensive
TSDFWell suited to surface fusion and reconstructionDifficult to use in dynamic scenes
NeRF/Gaussian SplattingView synthesis and continuous scene representationGeometry for control and real-time updates require additional validation
Object-centric stateFacilitates task composition and relational reasoningDepends on stable detection, tracking, and identity preservation

7. Object-Centric State and Affordances ​

A policy does not necessarily require complete scene reconstruction; it may need only task-relevant state:

Interpretation: The latent state consists of N objects; each object contains a category or semantics, a 3D pose, velocity, and task-relevant relationships or attributes.

Derivation: Decomposing the entire scene into a set of objects allows the state to scale with the number of objects and associates grasping, occlusion, and inter-object relationships with specific entities. Category c, pose T, velocity v, and relationship r form a minimal example; a real system must also attach a confidence score or distribution to each item.

An affordance is not a fixed label of an object but an actionable relationship jointly determined by the object, the acting agent, and the task. For example, whether a cup handle is “graspable” depends on gripper size, approach direction, and current occlusion.

8. Does an End-to-End VLA Still Need State Estimation? ​

An end-to-end model can encode state estimation implicitly in a Transformer context, but this does not eliminate observability problems. A camera cannot see the back side of an object, velocity cannot be determined from a single frame, and occlusion makes object identity uncertain. Explicit state estimation is valuable because it can be debugged, calibrated, and used for safety constraints; implicit state is valuable because it reduces manual modeling and can leverage large-scale data.

9. The Boundary Between Visual and Geometric Representations ​

Visual representations such as DINO, R3M, and VC-1 can provide semantics and correspondences; depth, SLAM, and DUSt3R-like geometric models provide camera pose and 3D structure. Semantic similarity does not guarantee correct metric distance, while geometric accuracy does not guarantee an understanding of how an object is used. Physical AI usually requires both.

Course Whiteboard

10. Minimal Experiments ​

  1. Implement a Bayes/Kalman filter using a one-dimensional position and noisy observations, and plot the prior, likelihood, and posterior.
  2. Briefly occlude an object and compare a single-frame policy with a policy that uses historical state estimation.
  3. Add small perturbations to the camera extrinsics and measure the degradation in end-effector positioning and grasp success rate.
  4. Perform target grasping using a purely visual embedding, explicit 3D coordinates, and a fusion of the two.
  5. Design an object-swapping experiment to verify whether the model preserves object identity rather than merely memorizing pixel locations.

11. Lesson Exercises ​

  1. Derive the prediction and correction steps from Bayes' rule.
  2. Explain why absolute depth cannot be uniquely recovered from a single monocular frame.
  3. Write the transformation chain among the base, camera, tool, and object coordinate frames.
  4. Construct an example in which correct semantics but incorrect geometry causes a grasp to fail.
  5. Design an uncertainty evaluation for the reappearance of an object after occlusion.

12. Major Failure Modes ​

FailureSymptomsDiagnosis and Correction
Conflating observations with statesThe object is immediately forgotten after occlusion, and single-frame ambiguity causes abrupt action changesHistorical belief, object persistence, and uncertainty
Overconfident filterThe mean appears stable even though the true state has moved outside the covariance rangeInnovation residuals, coverage, and noise-parameter calibration
Linearization failureThe EKF diverges during large rotations, collisions, or discrete contact transitionsUKF, particle filtering, hybrid modes, or reinitialization
Incorrect coordinate-frame directionVisual localization is correct, but the robot arm moves in the opposite directionExplicitly define the from/to semantics of each T and perform closed-loop calibration tests
Depth-scale or extrinsic driftGrasp points exhibit systematic offsets with distance and over timeGround-truth scale, reprojection error, and online calibration
Object identity swappingAfter similar objects cross paths, their states and task histories are exchangedIdentity consistency, data association, and multi-hypothesis tracking
Correct semantics, incorrect geometryThe target category is recognized, but its pose is insufficiently accurate for grasping or insertionMetric pose, reachability, and actual control error
Accurate geometry, task-irrelevant representationThe scene is reconstructed comprehensively, but the policy does not improveTask-relevant object state, policy ablations, and compute budget

13. Paper Facts, Author Interpretations, and Course Assessments ​

WorkPaper FactsAuthor InterpretationCourse Assessment
Kalman FilterRecursively computes the posterior state mean and covariance in linear-Gaussian systemsHistorical dynamics and current observations can be fused optimally according to their uncertaintyIt is foundational for understanding the belief state, but contact, multimodal occlusion, and incorrect data association exceed its basic assumptions
ORB-SLAM3Unifies visual, visual-inertial, and multi-map SLAM to estimate camera trajectories and mapsGeometric tracking and relocalization can provide a stable spatial referenceCamera localization is not equivalent to object state and task affordances; the policy still requires object-level semantics and dynamic relationships
DUSt3RDirectly regresses pointmaps from image pairs and performs 3D reconstruction and correspondence estimationLearned geometry can reduce reliance on conventional camera calibration and matching pipelinesReconstruction accuracy and scale stability must be revalidated in robot coordinates and control tasks
DINOv2 / VC-1Obtains transferable visual-semantic representations through large-scale pretrainingGeneral-purpose visual features can support various embodied downstream tasksSemantic correspondence is not metric geometry; Physical AI usually requires the fusion of semantics, geometry, and proprioceptive state

14. Cross-Reading for This Course ​

E1|Video Motion and Object-Centric Representations

E2|Latent Actions and Inverse Dynamics

B0|World Models and Model-Based Planning

F0|Dynamics, Control, and Physical Interaction

We recommend reading classic SLAM/Kalman textbooks alongside R3M, VC-1, DINOv2, DUSt3R, object-centric learning, and 3D vision-language models. The course focuses on what state information these methods provide to a policy, rather than treating perception benchmark performance itself as robotic success.

Article text is licensed under the Apache License 2.0