Interactive repair is only economical if the system can identify the smallest correction with the largest downstream effect. These phases establish that acquisition policy before any learned world model is introduced.
The scientific question
A model can be incomplete in thousands of places, but not every unknown matters to the task at hand. The first hypothesis was that question selection should follow the behavioral frontier: reveal the law that blocks the most currently reachable scenarios.
The protocol strictly separated acquisition trajectories from held-out and stress trajectories. Question selection, verification and stopping could not inspect evaluation outcomes.
The first run missed its answer-reduction gate by 0.16 percentage point. The threshold was not moved. A pre-registered replication passed, and the pooled rule opened the next phase.
From exact contexts to structural questions
Exact state/action identity fragmented one reusable law into many sparse questions. A domain-neutral structural identity compressed those duplicates using public action schemas, known dependencies and qualitative relations to thresholds.
The gain grew with topology size: larger worlds created more places where choosing the right question mattered.
| Regime | Frontier compression | AUC vs random | Answer reduction | Gate |
|---|---|---|---|---|
| Small | 52.4% | +0.0403 | 13.75% | Pass |
| Medium | 45.1% | +0.0533 | 21.80% | Pass |
| Large | 39.1% | +0.0884 | 26.13% | Pass |
One correction, multiple grounded instances
A confirmed concrete correction was then compiled into a typed transition motif and instantiated across compatible grounded components. Every instance still passed the existing verifier; ambiguous bindings abstained.
This is not open-ended schema induction. Public action schemas and typed role bindings were supplied. The result is narrower and more useful: validated knowledge can be reused without adding domain branches to the runtime.
| Regime | Concrete answers | Motif answers | Reduction | Exact final worlds |
|---|---|---|---|---|
| Small | 6.00 | 5.87 | 2.2% | 30/30 |
| Medium | 10.00 | 8.23 | 17.7% | 30/30 |
| Large | 16.00 | 9.83 | 38.5% | 30/30 |
What this established—and what it did not
The evidence supports a bounded acquisition claim: structural question identity and verified motif transfer reduce teaching cost on fresh controlled worlds, with the benefit increasing across the tested topology range.
It does not yet say that a deployed neural dynamics model can be repaired. It establishes the acquisition and compilation machinery that later phases must integrate with learned prediction and planning.
Combined totals are descriptive unless the article explicitly says they were a preregistered pooled gate. Machine timings describe the measured local implementation. A failed conjunction remains failed even when several sub-results are positive.