Original Feishu Document · Source Revision 8
💡
Mechanisms lesson: A VLA can output actions directly from images, but a physical closed loop still implicitly requires position, orientation, velocity, occlusion, and uncertainty. This lesson establishes a unified language spanning Bayes filtering, SE(3), depth and point clouds, object-centric state, and affordances.
Learning Objectives
After completing this lesson, you should be able to distinguish observations from states; derive the prediction and correction steps of a Bayes filter; interpret poses, coordinate transformations, depth back-projection, and object-centric representations; explain why end-to-end policies still require state identifiability; and design experiments involving occlusion, camera changes, and calibration errors.
1. An Image Is Not a State
The state a robot actually needs may include its own pose, joint velocities, the six-degree-of-freedom poses of objects, contact modes, and obstacles. A camera provides only observations of these variables after projection, occlusion, and noise:
Interpretation: The current observation is generated stochastically from the current true state through the observation model.
Derivation: The same state produces different images under different lighting conditions and cameras; the same image may also correspond to different depths and occlusion relationships. Uncertainty must therefore be represented with probability distributions.
2. Bayes Filter: How History Becomes a Belief State
Prediction step:
Interpretation: Propagate all possible states from the previous step to the current time according to the action-conditioned dynamics, then aggregate them to obtain the predicted belief before incorporating the new observation.
Derivation: By the law of total probability, integrate over the unknown previous state; the dynamics provide the conditional probability of transitioning from the previous state to the current state.
Correction step:
Interpretation: Reweight the predicted belief by the likelihood of the current observation under each candidate state, then use a normalization constant to ensure that the probabilities sum to one.
Derivation: Apply Bayes' rule directly: the prior is the predicted belief, the likelihood is the observation model, and the posterior is the corrected belief state.
3. The Kalman Filter Is the Linear-Gaussian Special Case
Interpretation: The state evolves according to linear dynamics and control inputs, with additive process noise; the observation is a linear projection of the state, with additive observation noise.
Derivation: Approximate the nonlinear dynamics near the operating point with matrices A and B, and approximate the sensor with matrix C. If both process noise and observation noise are Gaussian, linear transformations and products of Gaussians remain Gaussian, so the belief need only maintain a mean and covariance.
Under linear-Gaussian assumptions, the belief always remains a Gaussian distribution, so only its mean and covariance must be maintained. Nonlinear systems use EKF/UKF, while strongly multimodal distributions and discrete contact modes often require particle filters or learned filters.
4. 3D Poses and SE(3)
A rigid-body pose is represented by a homogeneous transformation:
Interpretation: Matrix T consists of a 3D rotation R and a 3D translation t and is used to transform a point from one coordinate frame to another.
Derivation: A 3D point is first rotated and then translated. With homogeneous coordinates, rotation and translation can be combined into a single matrix multiplication, while successive coordinate transformations are composed through matrix multiplication.
5. Back-Projecting Pixels and Depth into 3D
In a pinhole camera, pixel and depth correspond to the following camera coordinates:
Interpretation: Multiply the pixel offset from the principal point by the depth, then divide by the focal length to obtain the horizontal and vertical positions in the camera coordinate frame.
Derivation: By similar triangles, the normalized image coordinates equal the 3D coordinates divided by depth; rearranging gives the back-projection equations. Depth error directly amplifies 3D position error as distance increases.
Formula Visualization|From Pixel and Depth to World Coordinates

6. Point Clouds, Voxels, Occupancy, and Neural Fields
| Representation | Advantages | Limitations |
|---|---|---|
| Point cloud | Preserves direct 3D measurements | Sparse, irregular, and strongly affected by occlusion |
| Voxel/occupancy grid | Convenient for collision checking and spatial queries | High resolution is computationally expensive |
| TSDF | Well suited to surface fusion and reconstruction | Difficult to use in dynamic scenes |
| NeRF/Gaussian Splatting | View synthesis and continuous scene representation | Geometry for control and real-time updates require additional validation |
| Object-centric state | Facilitates task composition and relational reasoning | Depends on stable detection, tracking, and identity preservation |
7. Object-Centric State and Affordances
A policy does not necessarily require complete scene reconstruction; it may need only task-relevant state:
Interpretation: The latent state consists of N objects; each object contains a category or semantics, a 3D pose, velocity, and task-relevant relationships or attributes.
Derivation: Decomposing the entire scene into a set of objects allows the state to scale with the number of objects and associates grasping, occlusion, and inter-object relationships with specific entities. Category c, pose T, velocity v, and relationship r form a minimal example; a real system must also attach a confidence score or distribution to each item.
An affordance is not a fixed label of an object but an actionable relationship jointly determined by the object, the acting agent, and the task. For example, whether a cup handle is “graspable” depends on gripper size, approach direction, and current occlusion.
8. Does an End-to-End VLA Still Need State Estimation?
An end-to-end model can encode state estimation implicitly in a Transformer context, but this does not eliminate observability problems. A camera cannot see the back side of an object, velocity cannot be determined from a single frame, and occlusion makes object identity uncertain. Explicit state estimation is valuable because it can be debugged, calibrated, and used for safety constraints; implicit state is valuable because it reduces manual modeling and can leverage large-scale data.
9. The Boundary Between Visual and Geometric Representations
Visual representations such as DINO, R3M, and VC-1 can provide semantics and correspondences; depth, SLAM, and DUSt3R-like geometric models provide camera pose and 3D structure. Semantic similarity does not guarantee correct metric distance, while geometric accuracy does not guarantee an understanding of how an object is used. Physical AI usually requires both.

10. Minimal Experiments
- Implement a Bayes/Kalman filter using a one-dimensional position and noisy observations, and plot the prior, likelihood, and posterior.
- Briefly occlude an object and compare a single-frame policy with a policy that uses historical state estimation.
- Add small perturbations to the camera extrinsics and measure the degradation in end-effector positioning and grasp success rate.
- Perform target grasping using a purely visual embedding, explicit 3D coordinates, and a fusion of the two.
- Design an object-swapping experiment to verify whether the model preserves object identity rather than merely memorizing pixel locations.
11. Lesson Exercises
- Derive the prediction and correction steps from Bayes' rule.
- Explain why absolute depth cannot be uniquely recovered from a single monocular frame.
- Write the transformation chain among the base, camera, tool, and object coordinate frames.
- Construct an example in which correct semantics but incorrect geometry causes a grasp to fail.
- Design an uncertainty evaluation for the reappearance of an object after occlusion.
12. Major Failure Modes
| Failure | Symptoms | Diagnosis and Correction |
|---|---|---|
| Conflating observations with states | The object is immediately forgotten after occlusion, and single-frame ambiguity causes abrupt action changes | Historical belief, object persistence, and uncertainty |
| Overconfident filter | The mean appears stable even though the true state has moved outside the covariance range | Innovation residuals, coverage, and noise-parameter calibration |
| Linearization failure | The EKF diverges during large rotations, collisions, or discrete contact transitions | UKF, particle filtering, hybrid modes, or reinitialization |
| Incorrect coordinate-frame direction | Visual localization is correct, but the robot arm moves in the opposite direction | Explicitly define the from/to semantics of each T and perform closed-loop calibration tests |
| Depth-scale or extrinsic drift | Grasp points exhibit systematic offsets with distance and over time | Ground-truth scale, reprojection error, and online calibration |
| Object identity swapping | After similar objects cross paths, their states and task histories are exchanged | Identity consistency, data association, and multi-hypothesis tracking |
| Correct semantics, incorrect geometry | The target category is recognized, but its pose is insufficiently accurate for grasping or insertion | Metric pose, reachability, and actual control error |
| Accurate geometry, task-irrelevant representation | The scene is reconstructed comprehensively, but the policy does not improve | Task-relevant object state, policy ablations, and compute budget |
13. Paper Facts, Author Interpretations, and Course Assessments
| Work | Paper Facts | Author Interpretation | Course Assessment |
|---|---|---|---|
| Kalman Filter | Recursively computes the posterior state mean and covariance in linear-Gaussian systems | Historical dynamics and current observations can be fused optimally according to their uncertainty | It is foundational for understanding the belief state, but contact, multimodal occlusion, and incorrect data association exceed its basic assumptions |
| ORB-SLAM3 | Unifies visual, visual-inertial, and multi-map SLAM to estimate camera trajectories and maps | Geometric tracking and relocalization can provide a stable spatial reference | Camera localization is not equivalent to object state and task affordances; the policy still requires object-level semantics and dynamic relationships |
| DUSt3R | Directly regresses pointmaps from image pairs and performs 3D reconstruction and correspondence estimation | Learned geometry can reduce reliance on conventional camera calibration and matching pipelines | Reconstruction accuracy and scale stability must be revalidated in robot coordinates and control tasks |
| DINOv2 / VC-1 | Obtains transferable visual-semantic representations through large-scale pretraining | General-purpose visual features can support various embodied downstream tasks | Semantic correspondence is not metric geometry; Physical AI usually requires the fusion of semantics, geometry, and proprioceptive state |
14. Cross-Reading for This Course
E1|Video Motion and Object-Centric Representations
E2|Latent Actions and Inverse Dynamics
B0|World Models and Model-Based Planning
F0|Dynamics, Control, and Physical Interaction
We recommend reading classic SLAM/Kalman textbooks alongside R3M, VC-1, DINOv2, DUSt3R, object-centric learning, and 3D vision-language models. The course focuses on what state information these methods provide to a policy, rather than treating perception benchmark performance itself as robotic success.

