| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
ContractBench: evaluating observation contract failures (validity + integrity) in LLM agents. 33 harbor-runnable API-contract tasks with deterministic programmatic evaluation.
Inspect AI eval for ContractBench — 33 observation-contract tasks (validity + integrity). Register-ready for inspect_evals.
Everything needed to run the KORA benchmark for AI Child Safety.
PromtFuzz is an automated tool that generates high-quality fuzz drivers for libraries via a fuzz loop constructed on mutating LLMs' prompts.
OSS-Fuzz - continuous fuzzing for open source software.
A high-throughput and memory-efficient inference and serving engine for LLMs
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |