Thesis tested

Exact local repair now has positive evidence on two nonlinear process families. That closes one question and exposes the next: how should a world model act when the correction itself is approximate?

A bounded cross-process audit

The CSTR branch reused the same generic local action-effect patch, verifier and baselines. It did not read reactor equations or hidden coefficients. The 17,282-parameter backbone had to pass nominal competence before any shifted repair was evaluated.

The first audit stopped pre-shift because its initial-temperature distribution crossed thermal runaway. The only remediation narrowed temperature to the safe operating envelope on fresh seeds.

Prediction repair transferred

Four fresh shifted worlds triggered exactly once and committed the exact local change. Hybrid final H25 error fell from 0.060255 frozen to 0.000594, beating both full fine-tuning and a rank-2 adapter.

CSTR repair audit
StrategyFinal H25 errorUpdate latency
Frozen0.0602550
Full fine-tuning0.002054103.14 ms
Rank-2 adapter0.001755
Verified hybrid repair0.0005942.80 ms

Why the audit did not pass

The registered planning-consequence gate required frozen planning cost to be at least 1.50× the shifted reference. The observed ratio was only 1.272×. The full audit therefore failed six of seven gates, and downstream development and confirmation partitions remained unopened.

This is not a repair failure: prediction repair was evaluated and positive. It is an audit failure because the selected planning tasks did not make the model error consequential enough under the frozen rule.

Why exact answers are now the wrong easy case

With an exact expert correction, the hybrid succeeded on every reference-reachable audit task. There were no failed plans left for an uncertainty selector or abstention policy to identify.

The next milestone therefore changes one mechanism: teaching becomes approximate or interval-valued. The compiler must preserve that uncertainty, robust planning must propagate it, and the system must choose between acting, asking and abstaining on fresh partitions.

Next falsifiable milestone

Compile a bounded expert correction into an interval-valued dynamics patch, then test whether uncertainty-aware planning avoids unsafe confidence without surrendering useful recovery.

The project’s position now

The evidence does not yet support industrial readiness. It supports a research wedge: exact local edits can repair learned nonlinear dynamics across two external process families, faster and more transparently than full parameter updating.

The most valuable next proof is selective competence under imperfect knowledge—the point where interactive acquisition, language compilation and world-model planning become one problem.

Reading the evidence correctly

Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.