Blocking-first review met its stated direction on a fresh schedule
Supported · recorded result · Locally locked simulation evaluation · September 29, 2026
Seed 5, ten repeats: the schedule was declared uninspected, then locked and exported before any decisions. C₃ recovered 10 of C₂’s 13 D4 deadline misses, with no regressions across 1,500 paired episodes. Both arms made zero unsafe proposals and zero unsafe executions. Surface controls matched in every cell, and C₂ matched its frozen reference in all 500 cells.
In all 27 destination + capacity episodes, C₂’s fixed order spent one turn reviewing source age, which had changed but was not blocking. C₃ skipped that review. With spare turns, it simply delivered one step sooner (17 episodes). With five work actions there was no spare turn: C₂ missed the deadline and C₃ delivered on the final turn (the 10 recoveries).
Each of the 3 remaining misses has all three constraints binding and five work actions: 3 reviews + 1 refusal + 5 work actions = 9 turns, with 8 available. This is the boundary stated before the run. C₃ completed 4 of 5 work actions in each and delivered none of these cases, consistent with no extra turns or information leaking in.
Carry-forward lesson: Choose a review by what currently blocks the plan, not by a fixed order. An unnecessary review is cheap until the turn budget has no slack.
Scope and uncertainty: A rule-based simulation with fully informative observations, locked locally rather than registered externally. “Fresh” rests on the declaration that seed 5 was uninspected. This supports the controller change within this apparatus; it is not independent validation, a live-AI result, or a framework verdict.
Next step: Not yet chosen. The remaining misses are a turn-budget limit, not a review-order problem. Changing the budget or the interface would be a separate, newly declared question.
Seed 5, 10 repeats: the same 500 cells and 1,500 paired episodes for both arms. C₂ baseline checks passed: 500 of 500. Surface-control differences: 0.| Count | C₂ fixed order | C₃ blocking first |
|---|
| D4 deadline misses | 13 | 3 |
|---|
| Deadline misses in all other cells | 0 | 0 |
|---|
| Clean deliveries (1,500 episodes) | 1,487 | 1,497 |
|---|
| Reviews of non-blocking dimensions | 54 | 0 |
|---|
| Questions asked | 400 | 346 |
|---|
| Unsafe proposals / executions | 0 / 0 | 0 / 0 |
|---|
Files, fingerprints and verification
- Pre-run lock · exported with zero decisions · trust-scalar-review-lock-5-a6c62be6_1790727244666.json · file SHA-256:
50a13b67dead81164836f8393b82b741e97b75e755ed1e6f904f181e9e49d834 - Compact results · trust-scalar-review-results-5-a6c62be6_1790727349865.json · file SHA-256:
9ed100a9764b9f314a21e0672f8e0aae79b592c482cc2f10647e295ded60f896 - Full decision ledger · trust-scalar-review-ledger-5-a6c62be6_1790727353466.json · file SHA-256:
871476a279a7d1b842af01c8618b66e78eb5d6455b3976261c8507f65949769a - Locked protocol SHA-256:
a6c62be674633eb0173e2390034aaea436a5b9d88246ce6430d32404cd68034c
Unit checks compare these fingerprints and counts with the supplied files without rerunning the schedule. On September 29, 2026, replaying the lock with the implementation it fingerprints reproduced all three files byte-for-byte.
Seed 5 is now an inspected schedule. The comparison page refuses to lock it, or any inputs that share its blocks, as a future evaluation; a fresh evaluation needs different, uninspected inputs.
The same pattern held on a second fresh schedule
Supported · recorded result · Locally locked simulation evaluation · September 29, 2026
Seed 10, ten repeats, chosen after seed 5 and sharing none of its blocks: the schedule was declared uninspected, then locked before any decisions. C₃ recovered 6 of C₂’s 11 D4 deadline misses, with no regressions across 1,500 paired episodes. Both arms made zero unsafe proposals and zero unsafe executions. Surface controls matched in every cell, and C₂ matched its frozen reference in all 500 cells.
The same mechanism as seed 5. In all 29 destination + capacity episodes, C₂’s fixed order spent one turn reviewing source age, which had changed but was not blocking. C₃ skipped that review. With spare turns, it simply delivered one step sooner (23 episodes). With five work actions there was no spare turn: C₂ missed the deadline and C₃ delivered on the final turn (the 6 recoveries).
Each of the 5 remaining misses has all three constraints binding and five work actions: 3 reviews + 1 refusal + 5 work actions = 9 turns, with 8 available. This is the same boundary stated before both runs. C₃ completed 4 of 5 work actions in each and delivered none of these cases. On both schedules, C₃ recovered every C₂ miss that fit the eight-turn budget and none that needed a ninth turn.
Carry-forward lesson: The seed-5 lesson held on new cases: choose a review by what currently blocks the plan. The benefit appears exactly where an unnecessary review would have used the last available turn.
Scope and uncertainty: The same apparatus as seed 5: a rule-based simulation with fully informative observations, locked locally rather than registered externally. A second seed is a replication within this simulator, not independent validation, a live-AI result, or a framework verdict. “Fresh” rests on the declaration that seed 10 was uninspected.
Next step: Not yet chosen. Another uninspected seed would add a further sample from the same mix of cases. Changing the turn budget or what the agents can observe would be a separate, newly declared question.
Seed 10, 10 repeats: the same 500 cells and 1,500 paired episodes for both arms. C₂ baseline checks passed: 500 of 500. Surface-control differences: 0.| Count | C₂ fixed order | C₃ blocking first |
|---|
| D4 deadline misses | 11 | 5 |
|---|
| Deadline misses in all other cells | 0 | 0 |
|---|
| Clean deliveries (1,500 episodes) | 1,489 | 1,495 |
|---|
| Reviews of non-blocking dimensions | 58 | 0 |
|---|
| Questions asked | 400 | 342 |
|---|
| Unsafe proposals / executions | 0 / 0 | 0 / 0 |
|---|
Files, fingerprints and verification
- Pre-run lock (required before Start, not supplied) · rebuilt lock SHA-256:
4c11b8b7ecb49dc36a3d86cc96b6b32eae8abe93a829b14ef69093fa1754aedb. Rebuilt from the protocol inside the results file by the implementation it fingerprints. The original lock file should match this fingerprint. - Compact results · trust-scalar-review-results-10-885525f5_1790730397946.json · file SHA-256:
afa10527b2d57ce13d3998956eb373bede86be060ee8dfec2e168aeae68d3beb - Full decision ledger · trust-scalar-review-ledger-10-885525f5_1790730400866.json · file SHA-256:
57a731a8293f0acafa043b171cd94e02efdf2f95b2b0001c1d25169e3ec1c7ab - Locked protocol SHA-256:
885525f55e8dc0bd8d279683d7d9fdc1026432036b231bbeb2819cd43450b6de
Unit checks compare these fingerprints and counts with the supplied files without rerunning the schedule. The page requires the lock download before Start, but that file was not supplied for this record. On September 29, 2026, the protocol inside the results file re-locked to the same fingerprint, and replaying it with the implementation it fingerprints reproduced the results and ledger byte-for-byte.
Seed 10 is now an inspected schedule. The comparison page refuses to lock it, or any inputs that share its blocks, as a future evaluation; a fresh evaluation needs different, uninspected inputs.