A08 Labs research · 01–08

The claim was built one failure boundary at a time.

This is the complete research path from question selection to cross-process nonlinear repair. Each chapter separates what passed, what failed, what changed next and which claim became defensible.

PassedMixed / failed conjunctionDevelopment only
01
Phases 1–316 July 20268 min

Learning what to ask

The project starts by separating a useful teacher from a talkative one: ask about the missing law that unlocks the largest frontier of behavior.

Open chapter
Passed after registered replication29/30world-level AUC wins
02
Phase 417 July 202610 min

From executable laws to learned dynamics

A learned action-conditioned predictor becomes wrong after a hidden regime shift. The repair must recover the changed behavior without turning the engine into a hand-coded simulator.

Open chapter
Mixed: core repair positive, key failures preserved100%affected accuracy
04
Phase 6A21 July 202610 min

The closed repair loop

Detection, one counted correction, verified commit, recursive prediction and planning recovery finally run inside one frozen experiment.

Open chapter
Passed54/60hybrid target recovery
05
Phases 6B–6C-A22 July 202611 min

Ask only when the deployed model is actually wrong

The trigger learns an important restraint: observing a public contradiction is not enough. If the learned model already predicts correctly, spend no answer and make no edit.

Open chapter
Trigger passed; model-scale conjunction failed 18/1958/58exact actionable commits
06
Phase 6C-B22 July 202612 min

External benchmark: nonlinear four-tank control

On PC-Gym’s coupled four-tank process, one verified actuator correction restores H25 prediction and mean H80 plan quality—about 40× faster than full fine-tuning.

Open chapter
Continuous recovery passed; formal confirmations failed 9/1024/24detect + exact commit
07
Phase 6D23 July 20269 min

A second process—and the uncertainty frontier

The generic patch transfers to a nonlinear CSTR reactor. Exact repair works; the next falsifiable question is whether approximate teaching can support robust planning and safe abstention.

Open chapter
Cross-process evidence positive; audit failed 6/74/4exact CSTR repairs
08
Phase 6E23 July 20268 min

When uncertainty is not yet the bottleneck

Approximate repair recovered prediction and every selective execution succeeded—but the benchmark produced too few naïve failures to establish calibrated ask-or-abstain behavior.

Open chapter
Development failed 5/9; confirmation unopened95.5%H25 competence recovered

Before reading the numbers

See what every experiment was forbidden to read.

Fresh partitions, counted answers, frozen gates and explicit leak audits define the evidence. The methodology page explains how a result earns—or fails to earn—a claim.

Read the evidence contract