C₂ / C₃ review comparison · recorded September 29, 2026

Blocking-first review: recorded results

On two fresh locked schedules (seeds 5 and 10), blocking-first review (C₃) cut D4 deadline misses from 24 to 8 compared with the fixed review order (C₂): 16 recovered, 0 regressed, and no unsafe proposals or executions in either arm. Each of the 8 remaining misses needs nine turns in an eight-turn budget.

D4 is the simulator’s hardest context-change level: the prior plan points to the wrong destination. A deadline miss means the authorized goal was not delivered within the episode’s eight turns.

What was compared

C₂ keeps its existing order: it reviews the next changed dimension and stops once the plan is feasible. C₃ changes one thing: it reviews a currently blocking constraint first, choosing from the same visible constraints and using the original order to break ties. Both arms have the same information, safety boundaries, work and eight-turn budget.

Locked question: Can blocking-first review reduce D4 deadline misses relative to unchanged C₂-no-retrieval?

Locked prediction: Fewer avoidable D4 misses, with no unsafe proposals/executions or new clean-completion regressions. This is a directional prediction, not a guaranteed recovery count.

Runs at a glance

Rule-based simulation with fully informative observations, locked locally rather than registered externally. “Fresh” rests on the declaration that each schedule was uninspected. This supports the controller change within this apparatus; it is not independent validation, a live-AI result, or a framework verdict.
ScheduleD4 misses · C₂ → C₃Recovered · regressedLeft · each needs nine turnsUnsafe proposals · executions
Seed 5 · 10 repeats13 → 310 · 030 · 0
Seed 10 · 10 repeats11 → 56 · 050 · 0
Both runs24 → 816 · 080 · 0

Full records

Blocking-first review met its stated direction on a fresh schedule

Supported · recorded result · Locally locked simulation evaluation · September 29, 2026

Seed 5, ten repeats: the schedule was declared uninspected, then locked and exported before any decisions. C₃ recovered 10 of C₂’s 13 D4 deadline misses, with no regressions across 1,500 paired episodes. Both arms made zero unsafe proposals and zero unsafe executions. Surface controls matched in every cell, and C₂ matched its frozen reference in all 500 cells.

In all 27 destination + capacity episodes, C₂’s fixed order spent one turn reviewing source age, which had changed but was not blocking. C₃ skipped that review. With spare turns, it simply delivered one step sooner (17 episodes). With five work actions there was no spare turn: C₂ missed the deadline and C₃ delivered on the final turn (the 10 recoveries).

Each of the 3 remaining misses has all three constraints binding and five work actions: 3 reviews + 1 refusal + 5 work actions = 9 turns, with 8 available. This is the boundary stated before the run. C₃ completed 4 of 5 work actions in each and delivered none of these cases, consistent with no extra turns or information leaking in.

Carry-forward lesson: Choose a review by what currently blocks the plan, not by a fixed order. An unnecessary review is cheap until the turn budget has no slack.

Scope and uncertainty: A rule-based simulation with fully informative observations, locked locally rather than registered externally. “Fresh” rests on the declaration that seed 5 was uninspected. This supports the controller change within this apparatus; it is not independent validation, a live-AI result, or a framework verdict.

Next step: Not yet chosen. The remaining misses are a turn-budget limit, not a review-order problem. Changing the budget or the interface would be a separate, newly declared question.

Seed 5, 10 repeats: the same 500 cells and 1,500 paired episodes for both arms. C₂ baseline checks passed: 500 of 500. Surface-control differences: 0.
CountC₂ fixed orderC₃ blocking first
D4 deadline misses133
Deadline misses in all other cells00
Clean deliveries (1,500 episodes)1,4871,497
Reviews of non-blocking dimensions540
Questions asked400346
Unsafe proposals / executions0 / 00 / 0

Files, fingerprints and verification

  • Pre-run lock · exported with zero decisions · trust-scalar-review-lock-5-a6c62be6_1790727244666.json · file SHA-256: 50a13b67dead81164836f8393b82b741e97b75e755ed1e6f904f181e9e49d834
  • Compact results · trust-scalar-review-results-5-a6c62be6_1790727349865.json · file SHA-256: 9ed100a9764b9f314a21e0672f8e0aae79b592c482cc2f10647e295ded60f896
  • Full decision ledger · trust-scalar-review-ledger-5-a6c62be6_1790727353466.json · file SHA-256: 871476a279a7d1b842af01c8618b66e78eb5d6455b3976261c8507f65949769a
  • Locked protocol SHA-256: a6c62be674633eb0173e2390034aaea436a5b9d88246ce6430d32404cd68034c

Unit checks compare these fingerprints and counts with the supplied files without rerunning the schedule. On September 29, 2026, replaying the lock with the implementation it fingerprints reproduced all three files byte-for-byte.

Seed 5 is now an inspected schedule. The comparison page refuses to lock it, or any inputs that share its blocks, as a future evaluation; a fresh evaluation needs different, uninspected inputs.

The same pattern held on a second fresh schedule

Supported · recorded result · Locally locked simulation evaluation · September 29, 2026

Seed 10, ten repeats, chosen after seed 5 and sharing none of its blocks: the schedule was declared uninspected, then locked before any decisions. C₃ recovered 6 of C₂’s 11 D4 deadline misses, with no regressions across 1,500 paired episodes. Both arms made zero unsafe proposals and zero unsafe executions. Surface controls matched in every cell, and C₂ matched its frozen reference in all 500 cells.

The same mechanism as seed 5. In all 29 destination + capacity episodes, C₂’s fixed order spent one turn reviewing source age, which had changed but was not blocking. C₃ skipped that review. With spare turns, it simply delivered one step sooner (23 episodes). With five work actions there was no spare turn: C₂ missed the deadline and C₃ delivered on the final turn (the 6 recoveries).

Each of the 5 remaining misses has all three constraints binding and five work actions: 3 reviews + 1 refusal + 5 work actions = 9 turns, with 8 available. This is the same boundary stated before both runs. C₃ completed 4 of 5 work actions in each and delivered none of these cases. On both schedules, C₃ recovered every C₂ miss that fit the eight-turn budget and none that needed a ninth turn.

Carry-forward lesson: The seed-5 lesson held on new cases: choose a review by what currently blocks the plan. The benefit appears exactly where an unnecessary review would have used the last available turn.

Scope and uncertainty: The same apparatus as seed 5: a rule-based simulation with fully informative observations, locked locally rather than registered externally. A second seed is a replication within this simulator, not independent validation, a live-AI result, or a framework verdict. “Fresh” rests on the declaration that seed 10 was uninspected.

Next step: Not yet chosen. Another uninspected seed would add a further sample from the same mix of cases. Changing the turn budget or what the agents can observe would be a separate, newly declared question.

Seed 10, 10 repeats: the same 500 cells and 1,500 paired episodes for both arms. C₂ baseline checks passed: 500 of 500. Surface-control differences: 0.
CountC₂ fixed orderC₃ blocking first
D4 deadline misses115
Deadline misses in all other cells00
Clean deliveries (1,500 episodes)1,4891,495
Reviews of non-blocking dimensions580
Questions asked400342
Unsafe proposals / executions0 / 00 / 0

Files, fingerprints and verification

  • Pre-run lock (required before Start, not supplied) · rebuilt lock SHA-256: 4c11b8b7ecb49dc36a3d86cc96b6b32eae8abe93a829b14ef69093fa1754aedb. Rebuilt from the protocol inside the results file by the implementation it fingerprints. The original lock file should match this fingerprint.
  • Compact results · trust-scalar-review-results-10-885525f5_1790730397946.json · file SHA-256: afa10527b2d57ce13d3998956eb373bede86be060ee8dfec2e168aeae68d3beb
  • Full decision ledger · trust-scalar-review-ledger-10-885525f5_1790730400866.json · file SHA-256: 57a731a8293f0acafa043b171cd94e02efdf2f95b2b0001c1d25169e3ec1c7ab
  • Locked protocol SHA-256: 885525f55e8dc0bd8d279683d7d9fdc1026432036b231bbeb2819cd43450b6de

Unit checks compare these fingerprints and counts with the supplied files without rerunning the schedule. The page requires the lock download before Start, but that file was not supplied for this record. On September 29, 2026, the protocol inside the results file re-locked to the same fingerprint, and replaying it with the implementation it fingerprints reproduced the results and ledger byte-for-byte.

Seed 10 is now an inspected schedule. The comparison page refuses to lock it, or any inputs that share its blocks, as a future evaluation; a fresh evaluation needs different, uninspected inputs.