The same local action-effect repair works in a recognized nonlinear continuous-control environment without reading simulator coefficients. It restores continuous prediction and plan quality, but not yet task-level reliability dominance.
Why four-tank changes the evidence
The coupled four-tank benchmark introduces nonlinear continuous dynamics, two causal action channels and longer plan-once horizons. The learned backbone has 17,924 parameters and is trained for recursive competence rather than exact symbolic execution.
A hidden actuator degradation changes one action channel. Acquisition may read only visible trajectories and public affected-state footprints—not simulator equations, hidden coefficients or evaluation tasks.
Failures that shaped the protocol
The first nonlinear workload failed because the chosen shift barely changed consequential behavior. A causal-sensitivity protocol then failed because the shift signal was smaller than nominal model error. A stronger rollout-trained backbone fixed competence, but reactive control masked model differences.
Plan-once evaluation exposed the causal gap. A prospective 1% non-inferiority margin replaced an over-precise cost ordering that had failed by 0.0036%. R4 development then passed on both channels: 8/8 exact repairs, 31/32 hybrid successes versus 0/32 frozen, and 32/32 paired non-inferiority.
Two independently frozen evaluation partitions
Across 24 hidden changes, 96 held-out prediction episodes and 72 plan-once tasks, every change triggered exactly one correction and exact local commit. Every case recovered the registered H25 prediction gap and mean planning gap.
| Evidence | Result |
|---|---|
| Actionable trigger + exact commit | 24/24 |
| H25 prediction-gap recovery | 24/24 |
| Mean planning-gap recovery | 24/24 |
| Within 1% of better learned update | 72/72 tasks |
| Incorrect / cross-channel commits | 0 |
| Invariant violations / leaks | 0 |
Prediction, planning and update cost
On the final replication, hybrid final H25 error was 0.000180 versus 0.001046 for full fine-tuning. Mean planning cost fell from 0.002989 frozen to 0.000465 hybrid. The 178–179 byte patch updated in 2.51 ms median versus 100.27 ms for full fine-tuning.
| Metric | Frozen | Full FT | Hybrid |
|---|---|---|---|
| Final H25 error | — | 0.001046 | 0.000180 |
| H25 AUC | — | 0.002734 | 0.001264 |
| Mean planning cost | 0.002989 | — | 0.000465 |
| Median update latency | 0 | 100.27 ms | 2.51 ms |
The one gate that failed twice
B2-R4 and the feasibility-aware B2-R5 replication each passed nine of ten gates. The final replication covered 17 of 20 tasks that the shifted reference could reach—85% versus the frozen 90% requirement—and achieved 17 total successes versus 18 for the adapter.
This blocks any claim of task-level dominance or industrial reliability. The defensible claim is restoration of causal predictions and mean plan quality, materially faster than parameter updating, with locality and integrity preserved.
Further threshold, target, tolerance, seed, planner or backbone tuning was explicitly closed as a rabbit hole.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.