FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

screenpipe/evals/coding-agent at main · screenpipe/screenpipe · GitHub

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Coding-agent regression evals

This is an agent eval suite, not a unit-test suite. Every case contains:

  • a sanitized task derived from an escaped product failure;
  • the historical broken repository revision;
  • an isolated trial workspace with no future git history;
  • a hidden deterministic outcome grader materialized only after the agent stops;
  • saved prompt, transcript, candidate patch, grader output, runtime fingerprint, and result;
  • repeated-trial reporting with success rate, pass@k, and pass^k.

The current app corpus contains 20 git-mined regressions. See DESIGN.md for the Anthropic guidance, source contract, and history-mining workflow. The companion website manifest contributes 20 more.

The regression inventory has an explicit owner and advisory/blocking policy. New cases may declare trigger paths for later change-aware selection. Validation prints the dataset fingerprint, and scored reports retain both dataset and runtime fingerprints for exact-run comparison.

The product tests referenced by the manifest are graders. Passing them directly is not the eval; the evaluated object is an agent trajectory and resulting patch from the historical broken state.

Validate the corpus

node evals/coding-agent/run.mjs --validate
node evals/coding-agent/run.mjs --verify

--verify proves each historical base fails its hidden grader and its known fix passes. A case is invalid if either side of that contrast is missing.

Run agent trials

node evals/coding-agent/run.mjs \
  --case app-chat-concurrent-save \
  --trials 3 \
  --agent-command 'codex exec --ephemeral --approve-for-me --json -C {workspace} - < {prompt_file}'

The agent receives only the task and archived broken tree. It does not receive the oracle commit, grader definition, or future repository history.

Results are written under evals/coding-agent/results/ unless --results-dir is supplied. Use --keep only for debugging a failed trial.

Agent-process and harness failures are reported as error and excluded from the success denominator. They are never silently converted into model failures.

The same runner can score another checkout and manifest with --repo and --manifest; the website corpus uses this so both repositories share exactly one harness implementation.

Modes

  • agent: run the configured coding agent and grade its patch.
  • baseline: make no change and grade the broken revision.
  • oracle: apply the historical fix and grade it.
  • regrade: apply a saved candidate.patch and run the current grader without another model call.
  • --verify: require baseline failure plus oracle success.

Do not turn capability scores into a release gate after one run. Establish matched-environment repeated baselines first. Regression cases intended to block should target reliable pass^k, not a lucky pass@k.


Back | FazBrowse Home | New Git URL