FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Benchmark: finish the v2.1 LLM re-baseline (1,528 scenarios remaining) on subscription auth · Issue #95 · inputlayer/inputlayer · GitHub

Benchmark: finish the v2.1 LLM re-baseline (1,528 scenarios remaining) on subscription auth #95

Description

Corpus v2.1 (1,628 scenarios, 16 families - PR #90) is fully verified on the structural and engine axes: spans/clean-twins/labels 1,628/1,628 and an EXACT-match engine pass of 1,526/1,526 (verify_each.py ledger).

The LLM behavior re-baseline (regime A + regime B, double-graded with Opus arbiter) completed 100/1,628 scenarios before the subscription OAuth token rotated during the long background run; the remaining 1,528 are pending.

Everything needed is already in-tree:

  • full_bench.py --tier full --resume --out full_bench.json reruns only missing scenarios (auth resolves subscription OAuth from the keychain first; never uses a raw API key silently)
  • the run-until-verified driver loops rounds until verify_each.py reports every scenario individually verified (nonzero exit otherwise - CI-gateable)

Practical notes for whoever runs it: keep a Claude Code session alive so the keychain token stays fresh, expect the shared subscription rate window to pace the run (hours, zero marginal dollars), and post the final per-family/per-sub tables with Wilson CIs to #89 when the ledger closes. Until then the citable LLM numbers are the completed v1 study (results/archive_v1_full_bench_390.json) documented in poc/README.md.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestverified-completionsVerified Completions: OpenAI-compatible endpoint with consistency checking

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions


    Back | FazBrowse Home | New Git URL