Corpus v2.1 (1,628 scenarios, 16 families - PR #90) is fully verified on the structural and engine axes: spans/clean-twins/labels 1,628/1,628 and an EXACT-match engine pass of 1,526/1,526 (verify_each.py ledger).
The LLM behavior re-baseline (regime A + regime B, double-graded with Opus arbiter) completed 100/1,628 scenarios before the subscription OAuth token rotated during the long background run; the remaining 1,528 are pending.
Everything needed is already in-tree:
- full_bench.py --tier full --resume --out full_bench.json reruns only missing scenarios (auth resolves subscription OAuth from the keychain first; never uses a raw API key silently)
- the run-until-verified driver loops rounds until verify_each.py reports every scenario individually verified (nonzero exit otherwise - CI-gateable)
Practical notes for whoever runs it: keep a Claude Code session alive so the keychain token stays fresh, expect the shared subscription rate window to pace the run (hours, zero marginal dollars), and post the final per-family/per-sub tables with Wilson CIs to #89 when the ledger closes. Until then the citable LLM numbers are the completed v1 study (results/archive_v1_full_bench_390.json) documented in poc/README.md.
Corpus v2.1 (1,628 scenarios, 16 families - PR #90) is fully verified on the structural and engine axes: spans/clean-twins/labels 1,628/1,628 and an EXACT-match engine pass of 1,526/1,526 (verify_each.py ledger).
The LLM behavior re-baseline (regime A + regime B, double-graded with Opus arbiter) completed 100/1,628 scenarios before the subscription OAuth token rotated during the long background run; the remaining 1,528 are pending.
Everything needed is already in-tree:
Practical notes for whoever runs it: keep a Claude Code session alive so the keychain token stays fresh, expect the shared subscription rate window to pace the run (hours, zero marginal dollars), and post the final per-family/per-sub tables with Wilson CIs to #89 when the ledger closes. Until then the citable LLM numbers are the completed v1 study (results/archive_v1_full_bench_390.json) documented in poc/README.md.