Thesis tested

A public pretrained visual world model may be competent and sensitive to a hidden physical change without exposing a deployable, task-level repair signal. Repair must begin with a trustworthy contradiction—not merely a larger average prediction error.

Can a public visual world model become the repair substrate?

Phase 7 repaired a compact state-action model in MuJoCo. Phase 8 deliberately raised the bar: RGB observations became part of learned dynamics in contact-rich manipulation, and the candidate backbones were public pretrained world models rather than models designed around A08's repair loop.

The hypothesis was practical and falsifiable. If a pretrained model was already competent in the nominal world, a hidden contact change should create a public prediction contradiction; one bounded correction could then be compiled, verified and compared with adaptation baselines.

Repair never received an automatic pass. Before any question or model edit, the deployed trigger had to distinguish a real dynamics change from normal model error on unseen nominal episodes.

A reproducible GPU route on Scaleway

The large-model work ran on a disposable Scaleway L4-1-24G instance with an NVIDIA L4 GPU and a separately managed 500 GB persistent block volume. The VM could be destroyed and rebuilt while checkpoints, manifests and raw results remained versioned on persistent storage.

Ansible and pinned container images reconstructed the software environment. Every public checkpoint was identified by repository revision and artifact hash before an evidence run. Phase 8 used both a V-JEPA 2 action-conditioned route in ManiSkill PushCube and the public DINO-WM route in PushT.

This is an infrastructure disclosure, not a partnership claim. It records where the experiments ran and makes the compute boundary inspectable for future reproduction or technical collaboration.

Measured model scales

The V-JEPA route used a frozen large visual encoder with a 102,343-parameter action-conditioned interface. The DINO-WM audit used the exact public 52,319,499-parameter checkpoint.

ManiSkill proved the changed world was real

The first ManiSkill audit used PushCube with a hidden high-resistance contact change. Its initial protocol failed because bit-exact CPU replay was unrealistic and its endpoint-only detector missed too many changes. A preregistered remediation changed only those two measurement assumptions.

On fresh episodes, the remediated audit passed all nine platform gates: 64 of 64 nominal plans succeeded, only 7 of 64 survived the hidden change, and the exact unchanged-model input reference recovered 48 of 57 paired losses with no regression. This established a consequential and locally correctable workload—but not learned repair.

The pretrained V-JEPA route then learned useful nominal visual dynamics. A 64-parameter typed patch improved long-horizon prediction and aggregate planning, yet the repair conjunction failed twice: nominal false triggers, incomplete H8 recovery and unsafe plans remained. Phase 8D therefore never opened.

V-JEPA repair development boundary
Observed positiveMeasured resultWhy no repair claim
Typed local update64 parametersDevelopment only
H8 competence recovery40.05%Below the frozen boundary
Planning missions115/192 → 156/192Aggregate gain; unsafe plans remained
Update latency3.57× faster than full FTConjunction failed 5/8

DINO-WM was competent; its trigger was not

The second route changed both benchmark and model family. On PushT, the public DINO-WM checkpoint met the preregistered nominal H25 competence floor. A hidden contact intervention then collapsed task success and multiplied visual prediction error, while an evaluation-only unchanged-action reference recovered most lost missions.

That is exactly the setting in which repair could matter. But the deployed detector saw only public images and predictions. Whole-image surprise detected 4 of 24 changed worlds. The single permitted remediation measured the same frozen error only inside the manipulated block's public RGB footprint; detection improved to 14 of 24, still below the frozen 20-of-24 minimum.

The local audit also contained three no-contact plans, which failed the integrity requirement. Confirmation remained unopened. No answer, patch, fine-tuning comparison or repair claim was evaluated on this branch.

Two paired charts showing PushT mission success before and after the hidden contact change, and changed-world detections against the frozen threshold.
The change was consequential and visually measurable. The missing link was task-level detection reliability: the best registered detector reached 14 of 24 against a frozen 20-of-24 gate.
Once-only DINO-WM audits on fresh partitions
MeasureWhole-image auditLocal-footprint remediation
Nominal H25 mission success15/2417/24
Changed-world success2/242/24
Paired mission losses1315
Changed / matched visual error3.02×5.32×
Unchanged-action reference13/2415/24
Nominal false triggers0/240/24
Changed worlds detected4/2414/24
Frozen gate result8/9 fail7/9 fail

Why a larger average error was not enough

The 5.32-fold figure is a descriptive aggregate. It says that the affected visual region became harder to predict on average; it does not prove that every changed episode can be separated reliably from the natural error distribution of a large pretrained model.

This distinction is operational. A repair loop that asks or edits whenever aggregate error rises will eventually modify a correct model because of observation noise, unusual but valid states or ordinary approximation error. The trigger is therefore part of the scientific claim, not a dashboard metric.

The correct Phase 8 decision is a failure: the world changed, the model was affected and a privileged reference showed that local correction was feasible, but the deployed public evidence was insufficiently reliable to authorize repair.

What remains defensible

Phase 8 established a reproducible visual-manipulation testbed, nominal pretrained competence, a consequential hidden contact change and bounded positive development evidence. It did not establish reliable self-diagnosis or repair of a public pretrained world model.

The architecture now has to meet the mechanism

Searching for another almost-compatible checkpoint would repeat the same integration problem. Public models were built primarily for prediction or control, not for identifying which local law changed, representing a bounded correction and proving that the rest of the world model remained intact.

A08 Labs is therefore beginning work on its own world model, designed from the start to integrate verified adaptive—and eventually self-repairing—behavior. We will disclose its architecture and evaluation protocol only when the next evidence boundary is ready.

The standard does not change: competence first, then a consequential hidden change, a public trigger, a bounded attributable repair, version retention, planning recovery and comparison with learning-based adaptation on fresh data.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.