| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Autonomous orchestration framework for Claude Code with MemPalace-inspired memory (4-layer stack, 818-token wake-up), parallel-first Agent Teams (6 teammates), Aristotle First Principles methodology, and 4-stage quality gates. 925+ tests, 22 active hooks, automatic learning pipeline.
Ship evals before you ship features.
Eval framework. Define correct, test against it, get results.
A guard-railed, closed-loop workflow for AI coding agents: live state bus + execution-level hard intercepts for Claude Code and Codex (GitHub PR / GitLab MR). From step-level to requirement-level; eval-driven, spec-driven, human-in-the-loop.
AI-augmented QA platform for spec-driven development and testing, RAG-grounded analysis, eval-driven development and contract validation across Python, Go, Rust and Solidity.
Autonomous skill improvement loop for Claude Code plugins — inspired by Karpathy's autoresearch. Modify → evaluate → keep/discard → repeat until convergence. Zero-touch quality iteration at scale.
A hands-on learning repository exploring Spec-Driven Development (SDD) for building deterministic AI systems. Covers specs, evaluation loops, patterns, experiments, and failures to bridge theory with real-world AI engineering practices.
Production harness for a multi-agent BI system — eval-gated, guardrailed, cross-source-validated. LangGraph + hybrid RAG + FastAPI, live on AWS.
Multi-agent inspection pipeline for solar cell EL images: EfficientNet-B0 severity classifier + Qwen3-VL (Ollama) reasoning, served via FastAPI. 75.3% on a 20-criteria eval suite.
Companion code for the talk "Managing Production Agents at Scale — from Chaos to Reliability". One Google ADK agent, three production failure modes: eval-driven development, resilience, and zero-trust on Vertex AI.
Eval-driven development for LLM accounting skills. 50 test cases · 66% → 100% in 6 iterations · results reproducible with the included grader
Most AI plugins hope they work. These prove it. Eval-driven Claude plugins for product teams.
Ковенантный мониторинг корпоративных кредитов: 200 PDF + банковский леджер → статус ковенанта, значение метрики и транзакция-улика. Решение кейса Halyk AI Challenge с полным post-mortem.
Eval-first plugin builder for Claude Code — the eval suite is the contract; green is the definition of done. Primitive decision records, generated eval suites with arming gates, sha256-frozen contracts, goal-loop builds in isolated worktrees, verified ships.
Eval-first AI product engineering portfolio: 10 connected n8n Workflow-as-Code projects (eval harness, bilingual RAG with citation integrity, drift monitor, signed gateway, autonomous agent, Feishu adapter) - every subsystem gated by offline deterministic verification
Founds the AI-assisted development environment of a repository: spec-driven and eval-driven methodology, a domain-sized agent team, a living knowledge base in git, and explicit token economics.
🗂 Coordinate multiple AI agents using shared plans and structured tasks for efficient team-based coding with Claude Code Agent Teams.
The Ultimate Eval-Driven Claude Plugin Suite for Product Teams 2026 - Verified Toolkit
Catch LLM quality regressions before they reach production — eval-driven CI/CD with LLM-as-Judge scoring, Wilson 95% CI diffing, and automatic PR alerts.
Modular self-referencing Markdown grounding system for agentic AI software engineering and architecture
Add a description, image, and links to the eval-driven-development topic page so that developers can more easily learn about it.
To associate your repository with the eval-driven-development topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |