Original Feishu document · Source revision 18
💡
Mechanisms lesson: An offline checkpoint is not a robotics product. A deployment system requires safety envelopes, runtime monitoring, anomaly detection, fallback modes, human takeover, shadow evaluation, regression gates, and traceable rollback. This lesson examines how to translate model uncertainty into system actions.
Learning Objectives
After completing this lesson, you should be able to map hazards to constraints and monitors; distinguish model confidence, out-of-distribution detection, and physical safety metrics; design runtime assurance, shadow deployment, canary releases, and rollback; and make release decisions using risk budgets rather than average success rates.
1. Safety Is Not a Loss Term
Robot risk can be approximated as:
Interpretation: Expected risk equals the probability of each type of hazard multiplied by its consequence cost, summed over all hazard categories.
Derivation: Express total loss as a random variable partitioned by hazardous events. Applying the law of total expectation yields a weighted sum of each event’s probability multiplied by its conditional loss. This formulation requires sufficiently comprehensive hazard categories and comparable costs; correlated accident chains, unacceptable harm, and hard regulatory constraints cannot be diluted through averaging expected values.
Low-probability, high-consequence events cannot be diluted by large numbers of routine successes. Engineering systems must simultaneously reduce occurrence probability, limit consequences, increase detection rates, and shorten response times. Unacceptable harm should be treated directly as a hard constraint rather than something that task utility may offset.
2. From Hazard Analysis to Runtime Constraints
| Hazard | Observable signals | Runtime action |
|---|---|---|
| Collision/pinch injury | Distance, force, velocity, human-occupied regions | Limit speed, stop, retreat |
| Dropped object | Tactile slip, gripper width, vision | Regrasp, place in a safe zone |
| Policy loses track of the objective | Confidence, progress, cyclic trajectories | Re-perceive, request human assistance |
| System timeout | Inference latency, message age | Hold position, switch to a local controller |
| Sensor anomaly | Dropped frames, drift, timestamps | Degrade operation or terminate |
3. Safety Envelope
Let the admissible state set be . A simple runtime filter finds the safe action closest to the policy action:
Interpretation: Within the action bounds, find the command closest to the policy’s recommendation while requiring that the model-predicted next state remain within the safe set.
Derivation: Treat the policy action as the desired command, express safety requirements—such as joint limits, velocity, collision distance, and contact force—as constraints on the next state, and then project onto those constraints with the minimum possible modification. If the model or state estimate is incorrect, or constraints are missing, the projection itself does not guarantee safety; independent monitoring and hardware-enforced limits are therefore required.
Constraints may come from joint limits, velocity limits, collision distance, contact force, power, or barrier functions. The safety filter should record how much it modifies each policy action. Large corrections sustained over time indicate that the division of responsibilities between the policy and the safety layer has become misaligned.
4. Control Barrier Function Intuition
Use to denote the safe set. For the continuous system , we require:
Interpretation: The sum of the safety function’s rate of change under the uncontrolled dynamics, the component affected by the control input, and a boundary-recovery term must be nonnegative.
Derivation: Taking the time derivative of h(x) along the system x_dot=f(x)+g(x)u gives the gradient of h multiplied by f and g u, respectively—that is, the two Lie derivatives. After adding a monotone function alpha, the constraint requires the system, when approaching the boundary, to remain on or return to the safe side at a sufficient rate, thereby making the safe set forward invariant.
CBF guarantees depend on state observability, bounded model error, and the optimization completing on time. Real systems also require physical emergency stops, independent hardware-enforced limits, and conservative margins. Mathematical feasibility must not be equated with the safety of the product as a whole.

💡
Interactive Validation | Action Age, Canary Releases, Rollback, and Fault-Injection Lab
Vary action age, the safety response budget, canary traffic, and fault injection to observe when the system should degrade operation or roll back.
Action Age, Canary Releases, Rollback, and Fault-Injection Lab
:::5. Model Confidence Is Not Safety
Action-distribution entropy, ensemble variance, and VLM token probabilities can serve as uncertainty signals, but danger also depends on speed, contact, proximity to humans, and recoverability. The same low confidence may be acceptable during slow movement through open space but unacceptable when a cutting tool is close to a person.
Triggering rules should therefore be conditioned on the risk context:
Interpretation: The system must intervene when the estimated probability of harm in the current context multiplied by the consequence cost exceeds the risk threshold.
Derivation: The same uncertainty has different consequences in a low-speed open environment and near a person, so probability predictions are combined with contextual costs to produce conditional risk. The threshold should be determined jointly by the safety budget, calibration error, and response capability. High-consequence scenarios should also have hard exclusion zones that do not depend on probability estimates.
6. Four Layers of Online Monitoring
- System health: Latency, dropped frames, temperature, GPU status, network status, and message age.
- Physical health: Force, velocity, collision distance, remaining joint margin, and vibration.
- Task health: Progress, repeated actions, timeouts, and whether the target still exists.
- Distribution health: Drift in input embeddings, objects, scenes, actions, and failure slices.
Monitoring metrics must be tied to executable responses; otherwise, they are merely dashboard decorations.
6.1 Safety Responses Must Outrun the Time to Harm
It is not enough for monitoring to detect an issue eventually. For rapid collisions or dropped objects, the total time required for detection, decision-making, and braking must be shorter than the remaining time between hazard onset and irreversible harm:
Interpretation: The total duration of hazard detection, response selection, and braking must be shorter than the time required for the hazard to develop into harm.
Derivation: Starting from hazard onset, monitoring first incurs detection latency, the safety logic then incurs decision latency, and the actuator finally incurs braking or retreat time. A software response is physically meaningful only if the sum of these three intervals still precedes the moment of harm. Otherwise, the system must reduce speed in advance, increase the safety distance, or use faster local protection.
This inequality should be validated under worst-case rather than average latency and must account for sensor periods, queue jitter, network timeouts, braking distance, and load variations.
6.2 Drift Detection Requires Sequential Statistics
For latency, action-boundary, or embedding anomaly scores ell_t, a one-sided CUSUM can accumulate persistent shifts:
Interpretation: The current cumulative anomaly score equals the previous cumulative value plus the new score minus the tolerated drift, truncated at zero. An alert is triggered when the score exceeds threshold H.
Derivation: If the mean normal score is below the reference value nu, the cumulative value will frequently return to zero. If the distribution shifts upward persistently, each step contributes positive drift, and the cumulative value gradually crosses the threshold. H controls the trade-off between false alarms and detection latency and must be calibrated using both normal-operation and fault-injection data.
7. Shadow, Canary, and Staged Releases
| Stage | Model authority | Primary evidence |
|---|---|---|
| Offline replay | No control authority | Consistency, latency, action bounds |
| Shadow | Predicts but does not execute | Disagreement with the production policy, resource usage |
| Controlled experiment | Safe test site with immediate human takeover available | Real closed-loop behavior and fault response |
| Canary | Small subset of robots/tasks | Online risk and rollback capability |
| Expanded rollout | Gradually increased coverage | Continued compliance with stratified metrics |
8. Regression Testing Is Not Merely Rerunning an Old Benchmark
The regression suite should include core capabilities, severe historical failures, boundary conditions, degraded sensors, injected latency, hardware versions, and cross-task combinations. Every fix should add a test that reproduces the corresponding failure, creating a cumulative “incident immunity memory.”
9. Release Gates
Release gates should include all of the following:
- A lower bound on primary-task performance, not merely a point estimate.
- No degradation in any critical safety metric.
- All regression tests for severe historical failures must pass.
- Inference latency, resource usage, and timeouts must remain within budget.
- Data, models, code, firmware, and configurations must be traceable.
- Drills for stopping, degraded operation, human takeover, and rollback must pass.
10. Incidents and Near Misses
Near misses that cause no actual loss should still be recorded. At minimum, each event record should include the timeline, model and hardware versions, observations and actions, whether the monitor triggered, human interventions, direct causes, systemic causes, remediation items, and regression tests. Root-cause analysis must not stop at non-actionable labels such as “model hallucination.”
11. Minimum Deployment Drill
Drill system. Build a replay-plus-hardware-in-the-loop environment in which the production policy, new policy, safety filter, independent monitor, and low-level controller all run at their actual frequencies. Every fault injection must record the actual onset time, detection time, response command, and physical stop time.
| Fault injection | Expected response | Required measurements |
|---|---|---|
| Camera freeze, timestamp regression, force-sensor drift | Degrade perception, hold, or stop safely | Detection rate, false-positive rate, detection latency, and message age |
| Inference timeout, GPU overheating, network packet loss | Switch to a local controller and prohibit stale actions | Worst-case end-to-end latency, switchover gap, and control continuity |
| Human entry, reduced collision distance, sudden increase in contact force | Limit speed, retreat, or trigger an emergency stop | Response-time budget, stopping distance, peak force, and minimum safety margin |
| Target disappearance, cyclic actions, task timeout | Re-perceive, recover, or request human assistance | Task-monitor recall, duration of ineffective loops, and takeover success rate |
| Slow drift in input embeddings and task distributions | Alert, reduce coverage, or halt rollout expansion | CUSUM detection latency, interval between false alarms, and performance on drift slices |
| Degraded canary metrics and incompatible configurations | Atomically roll back the model, configuration, tokenizer, and control parameters | Rollback time, state recovery, version consistency, and post-rollback regression results |
Release sequence. Offline replay, shadow deployment, controlled experiments, canary releases, and expanded rollout must use progressively increasing authority. Pass, stop, and rollback criteria must be defined in advance for every stage. Shadow disagreement only indicates that the candidate behaves differently; it cannot replace real closed-loop safety testing.
Minimum outputs. A fault-injection coverage matrix, detection-latency distribution, safety response-time budget, risk-slice dashboard, shadow disagreement report, canary decision log, and one-click rollback drill record.
12. Failure Modes
| Failure mode | Systemic cause | Decisive check | Remediation |
|---|---|---|---|
| Safety logic runs in the same process as the policy | A model hang, memory exhaustion, or GPU failure disables protection at the same time | Kill the policy process and observe whether the protection system can still brake | Use an independent process, independent watchdog, and hardware emergency-stop chain |
| Thresholds were tuned only on the development set | Real-world noise, latency, and high-risk subgroups were excluded from calibration | Run fault injection and stratified reliability/risk-coverage tests | Use context-conditioned thresholds, conservative margins, and periodic recalibration |
| The emergency-stop action itself causes an object to be dropped | “Stop” does not account for the held object, gravity, or current contact state | Drill shutdowns across different poses, loads, and task stages | Design state-dependent fallback modes such as holding, placing in a safe zone, and controlled retreat |
| Average metrics pass while high-risk slices regress | Canary traffic is not stratified by hazard context | Slice results by human proximity, tool, speed, and task stage | Use risk-weighted sampling and gates that prohibit regression in any critical slice |
| Rollback switches only the checkpoint | The configuration, tokenizer, firmware, and control parameters remain on the new versions | Compare the complete deployment manifest after rollback | Use atomic release packages and transactional rollback |
| Monitoring raises alerts, but no one is responsible | There is no owner, response deadline, or automated action | Drill the complete timeline from an overnight alert to recovery | Bind every monitor to a runbook, owner, SLO, and automated fallback action |
13. Lesson Exercises
- List five hazard categories and corresponding monitors for a collaborative robotic-arm pick-and-place task.
- Construct an example in which the model has high confidence but the physical risk is high.
- Specify the release gates for moving from shadow deployment to a canary release.
- Design independent safety responses for camera freezes and inference timeouts.
- Convert an object-drop incident into an automatically replayable regression test.
Relationship to Other Tracks
G3 does not replace the algorithm tracks; it constrains their deployment authority. Track F provides physical limits and low-level protection; G1 provides statistical thresholds; G2 ensures version traceability; and human takeovers and failure data from Track C can become learning signals for the next iteration.
14. Evidence Matrix for Safety and Runtime Assurance Papers
| Work | Facts from the paper | Authors’ interpretation | Course assessment |
|---|---|---|---|
| Sha et al.|Simplex Architecture | The Simplex architecture allows a high-performance but not fully verified controller to coexist with a verified safe controller, with a decision module switching between them when necessary. | The authors use an independent safety baseline to limit the risks of online upgrades. | The core of runtime assurance is fault isolation and a switchable safe fallback—not adding another loss term to the policy. |
| Ames et al.|Control Barrier Functions | CBFs transform the forward-invariance condition for a safe set into constraints that can be added to online control optimization. | The authors use barrier inequalities to combine performance control with safety constraints. | CBFs cover only hazards that the model and state estimate can express; they must be combined with hardware protection, latency budgets, and fault monitoring. |
| García & Fernández|Safe RL Survey | This survey organizes safe reinforcement learning methods into two broad categories: modifying optimization criteria and restricting the exploration process. | The authors distinguish safety during training from performance guarantees during deployment. | Robot deployment safety cannot be justified solely by training rewards; evidence for exploration constraints, runtime protection, and rollback must be evaluated separately. |
| Lakshminarayanan et al.|Deep Ensembles | Deep Ensembles use multiple independently trained models to improve predictive uncertainty and calibration. | The authors use disagreement among models as a practical approximation of predictive uncertainty. | Ensemble variance is a signal, not a safety conclusion; it must be calibrated in conjunction with consequences, physical state, and distribution drift. |
| Geifman & El-Yaniv|Selective Classification | Selective classification studies the risk–coverage trade-off when a model rejects a subset of samples. | The authors reduce the error risk among accepted samples by ranking predictions by confidence. | Robotics systems must map rejection to re-perception, degraded operation, or human takeover while accounting for coverage and response costs. |
| Mitchell et al.|Model Cards | Model Cards proposes documenting intended uses, evaluation conditions, limitations, and performance across groups. | The authors use structured documentation to improve deployment transparency and clarify boundaries of responsibility. | Physical AI model cards should also be linked to control frequency, hardware, firmware, safety layers, rollback packages, and the history of real-world incidents. |
15. Cross-Reading
Read together with F1–F3. Runtime safety depends on stability, contact limits, latency, and system identification; it cannot be determined independently from high-level confidence.
Read together with G1. Thresholds, canary releases, and staged rollouts require calibration, risk coverage, and stratified confidence intervals.
Read together with G2. A complete rollback must restore the model, configuration, tokenizer, firmware, and control parameters, while preserving traceability for every event.

