A public pretrained visual world model may be competent and sensitive to a hidden physical change without exposing a deployable, task-level repair signal. Repair must begin with a trustworthy contradiction—not merely a larger average prediction error.
Can a public visual world model become the repair substrate?
Phase 7 repaired a compact state-action model in MuJoCo. Phase 8 deliberately raised the bar: RGB observations became part of learned dynamics in contact-rich manipulation, and the candidate backbones were public pretrained world models rather than models designed around A08's repair loop.
The hypothesis was practical and falsifiable. If a pretrained model was already competent in the nominal world, a hidden contact change should create a public prediction contradiction; one bounded correction could then be compiled, verified and compared with adaptation baselines.
Repair never received an automatic pass. Before any question or model edit, the deployed trigger had to distinguish a real dynamics change from normal model error on unseen nominal episodes.
A reproducible GPU route on Scaleway
The large-model work ran on a disposable Scaleway L4-1-24G instance with an NVIDIA L4 GPU and a separately managed 500 GB persistent block volume. The VM could be destroyed and rebuilt while checkpoints, manifests and raw results remained versioned on persistent storage.
Ansible and pinned container images reconstructed the software environment. Every public checkpoint was identified by repository revision and artifact hash before an evidence run. Phase 8 used both a V-JEPA 2 action-conditioned route in ManiSkill PushCube and the public DINO-WM route in PushT.
This is an infrastructure disclosure, not a partnership claim. It records where the experiments ran and makes the compute boundary inspectable for future reproduction or technical collaboration.
The V-JEPA route used a frozen large visual encoder with a 102,343-parameter action-conditioned interface. The DINO-WM audit used the exact public 52,319,499-parameter checkpoint.
ManiSkill proved the changed world was real
The first ManiSkill audit used PushCube with a hidden high-resistance contact change. Its initial protocol failed because bit-exact CPU replay was unrealistic and its endpoint-only detector missed too many changes. A preregistered remediation changed only those two measurement assumptions.
On fresh episodes, the remediated audit passed all nine platform gates: 64 of 64 nominal plans succeeded, only 7 of 64 survived the hidden change, and the exact unchanged-model input reference recovered 48 of 57 paired losses with no regression. This established a consequential and locally correctable workload—but not learned repair.
The pretrained V-JEPA route then learned useful nominal visual dynamics. A 64-parameter typed patch improved long-horizon prediction and aggregate planning, yet the repair conjunction failed twice: nominal false triggers, incomplete H8 recovery and unsafe plans remained. Phase 8D therefore never opened.
| Observed positive | Measured result | Why no repair claim |
|---|---|---|
| Typed local update | 64 parameters | Development only |
| H8 competence recovery | 40.05% | Below the frozen boundary |
| Planning missions | 115/192 → 156/192 | Aggregate gain; unsafe plans remained |
| Update latency | 3.57× faster than full FT | Conjunction failed 5/8 |
DINO-WM was competent; its trigger was not
The second route changed both benchmark and model family. On PushT, the public DINO-WM checkpoint met the preregistered nominal H25 competence floor. A hidden contact intervention then collapsed task success and multiplied visual prediction error, while an evaluation-only unchanged-action reference recovered most lost missions.
That is exactly the setting in which repair could matter. But the deployed detector saw only public images and predictions. Whole-image surprise detected 4 of 24 changed worlds. The single permitted remediation measured the same frozen error only inside the manipulated block's public RGB footprint; detection improved to 14 of 24, still below the frozen 20-of-24 minimum.
The local audit also contained three no-contact plans, which failed the integrity requirement. Confirmation remained unopened. No answer, patch, fine-tuning comparison or repair claim was evaluated on this branch.

| Measure | Whole-image audit | Local-footprint remediation |
|---|---|---|
| Nominal H25 mission success | 15/24 | 17/24 |
| Changed-world success | 2/24 | 2/24 |
| Paired mission losses | 13 | 15 |
| Changed / matched visual error | 3.02× | 5.32× |
| Unchanged-action reference | 13/24 | 15/24 |
| Nominal false triggers | 0/24 | 0/24 |
| Changed worlds detected | 4/24 | 14/24 |
| Frozen gate result | 8/9 fail | 7/9 fail |
Why a larger average error was not enough
The 5.32-fold figure is a descriptive aggregate. It says that the affected visual region became harder to predict on average; it does not prove that every changed episode can be separated reliably from the natural error distribution of a large pretrained model.
This distinction is operational. A repair loop that asks or edits whenever aggregate error rises will eventually modify a correct model because of observation noise, unusual but valid states or ordinary approximation error. The trigger is therefore part of the scientific claim, not a dashboard metric.
The correct Phase 8 decision is a failure: the world changed, the model was affected and a privileged reference showed that local correction was feasible, but the deployed public evidence was insufficiently reliable to authorize repair.
Phase 8 established a reproducible visual-manipulation testbed, nominal pretrained competence, a consequential hidden contact change and bounded positive development evidence. It did not establish reliable self-diagnosis or repair of a public pretrained world model.
The architecture now has to meet the mechanism
Searching for another almost-compatible checkpoint would repeat the same integration problem. Public models were built primarily for prediction or control, not for identifying which local law changed, representing a bounded correction and proving that the rest of the world model remained intact.
A08 Labs is therefore beginning work on its own world model, designed from the start to integrate verified adaptive—and eventually self-repairing—behavior. We will disclose its architecture and evaluation protocol only when the next evidence boundary is ready.
The standard does not change: competence first, then a consequential hidden change, a public trigger, a bounded attributable repair, version retention, planning recovery and comparison with learning-based adaptation on fresh data.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.