Original Feishu Document · Source Revision 12
💡
Mechanisms lesson: Robots need to remember current task progress, past failures, object locations, and new demonstrations. This lesson compares in-context memory, external memory, parameter adaptation, and fast weights, with an emphasis on test-time training and WAM-TTT.
Learning Objectives
After completing this lesson, you should be able to distinguish working memory, episodic memory, external retrieval, and fast weights; formulate the inner and outer loops of test-time adaptation; analyze memory contamination, forgetting, and human-robot alignment; and design experiments involving incorrect demonstrations, irrelevant demonstrations, and safety gating.
1. Why Memory Is Necessary
A single-frame observation cannot answer:
- Which objects have already been handled.
- Whether the gripper has just slipped.
- Which grasping approach was previously attempted.
- What the user has just demonstrated.
- What new differences exist between the current environment and the training environment.
Memory compresses historical information into the conditioning context for the current decision.
2. Four Types of Memory
| Memory | Medium | Update | Risk |
|---|---|---|---|
| In-context memory | Historical tokens / hidden state | Every step | Context length and forgetting |
| External memory | Trajectory repository, object graph, database | Explicit writes | Retrieval errors and staleness |
| Parameter memory | Model weights | During training | Slow updates |
| Fast weights | Plastic parameters or adapters | At test time | Contamination and catastrophic forgetting |
3. In-Context Memory
Policy:
Interpretation: The current action is sampled from a policy distribution conditioned on observations from the most recent L steps, actions from the previous L steps, and the task goal.
Derivation: In a partially observable environment, a single frame cannot determine occluded objects, the task stage, or the consequences of actions. Adding a finite history window to the conditioning context can recover some state information, but critical events outside the window are truncated, and computational cost increases with L.
A finite window is simple, but early task states may be truncated. An RNN/Transformer hidden state can continuously aggregate information, but it is difficult to verify exactly what it has retained.
4. External Episodic Memory
Retrieve experiences similar to the current state:
Interpretation: From memory store M, select the memory item m_i whose key k_i is most similar to the encoding of the current observation and task.
Derivation: Encoder E maps the current situation to a query vector, and the retriever ranks historical experiences by similarity. If the key encodes appearance alone, similar images may correspond to different mass properties, contact states, or task stages. Reliable retrieval should therefore also include object identity, embodiment state, and failure cause.
Use the retrieved trajectories, failure causes, or recovery skills as policy conditioning. Visual similarity does not necessarily imply the same physical state; object, task-stage, and embodiment metadata are required.

5. Bilevel Optimization for Test-Time Adaptation
The inner loop uses data from the current environment to update fast parameters:
Interpretation: Update rule U_eta uses the original fast parameters phi and the current context data to produce the adapted parameters phi prime.
Derivation: The inner loop compresses demonstrations, unlabeled videos, or interaction feedback into fast weights. U_eta may consist of one or more gradient-descent steps, a learned optimizer, or an explicit memory writer. The key is to keep the update scope controlled and make the update reversible.
The outer loop uses robot task outcomes to optimize the update rule:
Interpretation: Select update rule eta such that, after it reads context data D_c, the expected loss on independent robot task data D_r is minimized.
Derivation: The outer loop must evaluate updates using robot outcomes that are separate from the write signal. Otherwise, the updater may merely reduce the auxiliary loss while degrading control. Taking the expectation over a distribution of tasks and environments trains a generalizable method for determining “how to update” across situations.
The key criterion is not whether the inner loss itself decreases, but whether robot task performance improves after the update.
Formula Visualization|Inner and Outer Loops for Test-Time Adaptation

6. Test-Time Training Signals
| Signal | Advantage | Risk |
|---|---|---|
| Video prediction | Requires no action labels | Appearance adaptation does not imply behavioral adaptation |
| Contrastive consistency | Uses multiple views / temporal information | Weak task relevance |
| Pseudo-actions | Connects state transitions | Pseudo-label bias |
| Human demonstrations | Specific to the task and environment | Cross-embodiment alignment |
| Robot failures | Matches the deployment distribution | Safety risks and erroneous writes |
7. The Mechanism of WAM-TTT
The World-Action Model learns shared structure among vision, latent actions, and robot actions. At test time, human videos write to fast-weight memory, and the robot policy reads the updated memory.
Meta-training requires:
- The inner loop to observe only human-side signals.
- The outer loop to use robot action tasks to determine whether an update is useful.
- Training environments to contain correspondences between human and robot tasks.
If human and robot tasks are not synchronized or stage-aligned, the memory may encode incorrect behavior.
8. Memory Gating
Not every experience should be written to memory. The gating probability is:
Interpretation: Based on candidate experience x, uncertainty u, and current task g, the gating model estimates the probability that the experience should be written to memory.
Derivation: The write decision can be represented as a binary variable w, and the gate can be supervised using correctness, task relevance, safety, and the benefit of rollback. High uncertainty does not necessarily mean that a write should be rejected, but it should trigger more conservative validation or sandboxed adaptation.
The following should be considered:
- Whether the demonstration is relevant to the current task.
- Whether the model understands the object and task stage.
- Whether a small-scale post-update evaluation shows improvement.
- Whether the update can be rolled back.
- Whether safety boundaries are preserved.
9. Memory Contamination and Forgetting
| Problem | Manifestation | Mitigation |
|---|---|---|
| Incorrect demonstration | Capability declines after rapid adaptation | Confidence gating and controlled validation |
| Irrelevant demonstration | Memory interferes with the current task | Task routing and isolation |
| Sequential users | Preferences overwrite one another | User-specific adapters |
| Long-horizon task | Early information is forgotten | External state graph |
| Excessive writes | Higher inference cost and retrieval noise | Compression, eviction, and summarization |
10. Evaluating Memory
- No-demonstration baseline.
- Correct demonstration.
- Irrelevant demonstration.
- Incorrect demonstration.
- Shuffled task stages.
- Same object but different task.
- Retention of old tasks after adaptation.
- Update rollback and safety gating.
11. Minimal Experiment: Does Writing to Memory Actually Improve the Robot Task?
Select a manipulation task involving a new object location, a new background, or new operator habits. Copy the same base policy into four systems: no adaptation, context window, external retrieval, and fast weights. Hold the base model, robot training data, number of frames visible at test time, number of update steps, and total inference time fixed.
- Provide correct demonstrations, irrelevant demonstrations, incorrect demonstrations, stage-shuffled demonstrations, and demonstrations involving the same object but a different task.
- Strictly separate the context used for writing from the robot rollouts used for evaluation.
- Compare full-parameter updates, restricted-layer updates, explicit memory tokens, and read-only retrieval.
- Re-evaluate old tasks after adaptation to measure forgetting and cross-task contamination.
- Enable shadow evaluation, gate-based rejection, and one-click rollback for high-risk actions.
Minimum reporting requirements: Success rates before and after adaptation, amount of context required to achieve the same improvement, performance degradation caused by incorrect demonstrations, sensitivity to irrelevant demonstrations, old-task retention rate, gating precision and recall, degree of recovery after rollback, update compute, and latency.
12. Lesson Exercises
- Design working-memory fields for a long-horizon organization task.
- Compare in-context memory with parameter adaptation.
- Specify the data boundary between the inner and outer loops.
- Design a safe experiment involving incorrect demonstrations.
- Explain why retrieval similarity must be conditioned on the task and embodiment.
- Explain WAM-TTT’s human-robot synchronization assumption.
13. Major Failure Modes
| Failure | Manifestation | Diagnosis and Correction |
|---|---|---|
| Memory truncation | Early information in a long-horizon task is forgotten after leaving the context window | Use explicit task state, summarized memory, or an external event log |
| Retrieval mismatch | Retrieved experience has a similar appearance but a different physical state or task stage | Add object, task, embodiment, and failure cause to the key, and report retrieval hit quality |
| Contamination from incorrect demonstrations | A single incorrect or adversarial demonstration continuously degrades subsequent control | Use write gating, an independent validation set, shadow updates, and rollback-capable snapshots |
| Auxiliary-objective shortcutting | Test-time loss decreases while robot success rate declines | Use independent robot outcomes in the outer loop and report the correlation between loss and task improvement |
| Catastrophic forgetting | Old tasks or safety skills degrade after adaptation to a new environment | Restrict updated parameters, apply regularization, replay old tasks, and evaluate retention |
| Human-robot stage misalignment | The stage shown in the human video differs from the robot’s current need | Use event alignment, task-predicate validation, and stage-shuffling controls |
| Update drift | After writes across multiple consecutive episodes, parameters gradually leave the safe region | Track update norms, maintain a version chain, reset periodically, and impose a maximum adaptation budget |
| Privacy and data persistence | External memory retains human or environmental information that should not be stored long term | Minimize storage, define expiration policies, enforce access control, and support deletion |
14. Paper Facts, Author Interpretations, and Course Assessments
| Work | Paper Fact | Author Interpretation | Course Assessment |
|---|---|---|---|
| MAML | Through meta-optimization across tasks, learns an initialization that can adapt rapidly using a small number of gradient steps | The training objective can directly shape a model’s adaptability | It provides the foundation for bilevel optimization, but real robots must also address update safety, data relevance, and online computational cost |
| Test-Time Training | Updates the model on test samples using a self-supervised auxiliary task before making predictions for the main task | The test distribution itself can provide adaptation signals | A decrease in auxiliary loss is not sufficient evidence of improved control; robots must validate benefits and safety on independent rollouts |
| RoboCat | Unifies multi-robot and multi-task data, and demonstrates subsequent training and data-driven self-improvement using demonstrations of new tasks | A generalist robotic agent can continuously expand its capabilities through new experience | This is closer to offline or staged self-improvement and should not be conflated with test-time fast weights within a single episode |
| WAM-TTT | Uses human-side context to write to fast memory at test time and supervises the update mechanism through robot action tasks | Human demonstrations can help robots adapt to the current environment or task | Key audit items include human-robot task correspondence, strictly held-out environments, controls using incorrect demonstrations, update stability, and rollback capability |
15. Cross-Reading for This Course
D0|Hierarchical Planning, Skills, and Memory
D2|Subgoal Planning and Embodied Reasoning
How Robots Learn Motion When Human Videos Have No Velocity Labels

