A deployed learned dynamics model can identify a consequential actuator change, ask one bounded question, compile and verify a local repair, retain the old regime and recover multi-step competence—while a pre-question provenance gate prevents uncertain perception from becoming a false law change.
From process control to an embodied mission
Earlier phases repaired learned dynamics in controlled resource processes, nonlinear four-tank control and CSTR. Phase 7 moved the same evidence contract to the official Gymnasium MuJoCo Reacher-v5 task: a two-joint arm must reach a target within a fixed 12-step mission.
The learned state-action model contained 137,482 parameters. A hidden deployment change reduced only the first actuator channel to 40% authority. Public action names, target and topology stayed fixed; the model became wrong because the world changed underneath it.
MuJoCo provides real rendered observations of the simulated experiment, which makes the behavior inspectable. It is still physics simulation—not footage from a physical robot—and the distinction is kept visible in every public asset.
Two failed triggers were preserved before the benchmark passed
The first embodied candidate, FetchPush, passed competence and consequence gates but failed its trigger. A passive detector produced two nominal false positives and only five timely shifted detections. A preregistered free-space probe removed false positives but still detected only 18 of 32 changes against a 28-of-32 gate.
Those failures were not tuned away. FetchPush was closed under its stop rule. A fresh Reacher audit then froze a direct causal action-response probe on new partitions and passed all 10 gates.
| Measure | Nominal / matched | Shifted frozen | Exact input reference |
|---|---|---|---|
| H25 normalized error | 0.18158 | 0.61906 | 0.18251 |
| 12-step mission success | 32/32 | 12/32 | 26/32 |
| Detector | 0/32 false gaps | 32/32 changes | Evaluation only |
Repair was not evaluated until the learned model was competent, the hidden change was consequential, the trigger was clean and a local correction was demonstrably feasible.
One answer, one verified input change
Each shifted world received the same 16 public transitions. The causal probe identified public action joint_torque_0. Only then did the system ask one bounded question: what fraction of the commanded torque now reaches the system?
The answer—40%—compiled into a typed action-input transformation before the unchanged learned dynamics model. A visible H8 replay had to reduce causal error, preserve neutral action, remain finite and pass before an immutable shifted version could commit. The nominal route retained the exact original predictor object.

| Strategy | Mean H25 error | 12-step missions | Median update |
|---|---|---|---|
| Frozen | 0.55419 | 9/48 | 0 |
| Full fine-tuning | 0.55527 | 19/48 | 0.44781 s |
| Rank-two adapter | 0.42124 | 33/48 | 0.00032 s |
| Verified hybrid repair | 0.17071 | 39/48 | 0.00882 s |
| Matched nominal / exact reference | 0.17158 / 0.17071 | — / 39/48 | Evaluation only |
The hybrid update was 50.8× faster than full fine-tuning and updated zero learned weights. The closed-form adapter was faster than the hybrid; the result does not claim otherwise.
The system also had to earn the right not to ask
Across 12 shifted confirmation worlds, the system made 12 exact one-answer verified commits. Twelve fresh nominal controls spent zero answers, and four shifted worlds with deliberately missing provenance abstained before the question. Incorrect commits and invariant violations were zero.
This asymmetry is central to the project. A useful repair loop is not a detector followed by automatic mutation; it needs a typed evidence boundary, a counted external answer, a verifier and an explicit commit authority.
Perception uncertainty comes before dynamics repair
Phase 7C added three learned visual estimators—1.424 million parameters in total—as an independent camera view over public robot telemetry. This is not a vision-only world model. Public proprioception still drives dynamics detection and verification; the camera ensemble validates whether the observation provenance is trustworthy enough to open a law-changing question.
The ensemble trained only on nominal images. A fresh calibration froze its uncertainty and camera-to-proprioception disagreement boundaries before clean shifted or occluded images were opened.
| Condition | Trusted | Underlying dynamics gaps | Answers | Commits |
|---|---|---|---|---|
| Clean nominal | 32/32 | 0/32 | 0 | 0 |
| Clean actuator change | 32/32 | 32/32 | 32 | 32 exact |
| Occluded nominal camera | 0/32 | 0/32 | 0 | 0 |
| Occluded camera + actuator change | 0/32 | 32/32 | 0 | 0 |
Every registered central 32×32 occlusion abstained before the question. It is a controlled corruption result, not a claim of robustness to arbitrary visual failures.
A replay that keeps inconvenient outcomes
The visual demonstrator chose one acquisition and four mission seeds deterministically from the sealed confirmation hash before opening their outcomes. One mission isolated the repair benefit; one was already solved by every strategy; two failed even for the exact reference.
All four remain in the replay. Correcting the world model cannot make every finite-horizon task feasible, and a useful demonstration should show that boundary rather than curate it away.

What Phase 7 establishes—and what remains open
The result supports a bounded embodied claim: after a known class of hidden actuator change, one active probe and one exact typed correction restored the deployed learned model's matched competence and recovered aggregate multi-step planning faster than full fine-tuning, with exact old-regime routing.
It does not establish physical-robot transfer, arbitrary fault discovery, unrestricted language induction, robustness to arbitrary image corruption or universal task-level dominance over adapters. The visual layer is a cross-modal provenance gate, not a vision-only world model.
The next justified milestone is bounded physical transfer: reuse the frozen question, compiler, verifier and router while changing only the hardware adapter and instrumentation.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.