Method / Evidence contract

A good repair story is easy. A falsifiable one is harder.

Every core milestone writes its gate before tuning, separates visible evidence from evaluation, compares against learned-update baselines and preserves failed decisions. This page is the compact contract behind the claims.

01

Fresh before frozen

Development may diagnose blockers. Confirmation seeds, thresholds and stop rules are registered before results are inspected.

02

One common budget

Frozen, fine-tuned, adapter and repaired systems receive the same visible post-shift interaction prefix.

03

Failure stays failed

A 9/10 or 18/19 conjunction is not rounded up. Positive sub-results remain reportable only inside their exact claim boundary.

The frozen comparison

Four strategies, one shifted world.

AFrozen learned model

No update after shift.

BFull fine-tuning

Weights updated on visible trajectories.

CVerified symbolic repair

Counted answer becomes a bounded patch.

DHybrid composition

Patch handles changed slice; learned model handles the rest.

Information boundary

What the repair loop may see—and what remains sealed.

The teacher can answer only after visible mismatch evidence creates a valid typed gap. Evaluation outcomes cannot affect question selection, compilation, verification or commit.

VISIBLE TO ACQUISITION + REPAIR
Pre-shift training data
Public model + groundings
Visible post-shift prefix
Counted expert answer
REPAIR LOOPdetect → ask → compile → verify → commit
EVALUATION ONLY
Held-out + stress trajectories
Hidden law identifiers
Reference candidates
Task outcomes + gates

Every gate reports

Competence, cost, locality and integrity.

PredictionOne-step error and affected-state accuracy.
Recursive rolloutH5, H8, H16 or H25 recovery—not one-step alone.
PlanningPlan quality, task success, false plans and safe abstention.
Teaching costCounted answers and visible trajectories.
ComputeWall-clock update latency and inference overhead.
LocalityUnaffected behavior and explicit old-version routing.
SafetyInvariant violations, incorrect commits and provenance.
IntegritySeed separation, leak audits and replayable decision traces.

How to read status

Evidence labels with teeth.

Passed

Every registered gate passed on its frozen fresh partition.

Mixed / failed conjunction

Useful sub-results exist, but at least one required gate failed.

Development only

Promising diagnostic or prototype evidence that cannot confirm a claim.

Anti-rabbit-hole rule

Three questions before adding any mechanism.

  1. Does it directly enable or measure one arrow in the end-to-end repair loop?
  2. Is that arrow currently blocked, rather than merely interesting to improve?
  3. Can the result be expressed as a falsifiable comparison on fresh data?

If any answer is no, the idea is deferred. This rule closed more prompt tuning, more four-tank threshold tuning and production-demo cosmetics once they stopped addressing the active scientific blocker.

Trace the evidence

Start at the latest result—or rebuild the argument from chapter one.