FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

evaluation-framework · GitHub Topics · GitHub

#

evaluation-framework

Here are 564 public repositories matching this topic...

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

  • Updated Aug 19, 2026
  • TypeScript

A framework for few-shot evaluation of language models.

  • Updated Aug 14, 2026
  • Python

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

  • Updated Aug 20, 2026
  • Python

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

  • Updated Aug 11, 2026
  • Python

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Aug 19, 2026
  • Python

This is the repository of our article published in RecSys 2019 "Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches" and of several follow-up studies.

  • Updated May 25, 2023
  • Python

Repo for AI Agents The Definitive Guide

  • Updated Aug 1, 2026
  • Jupyter Notebook

AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.

  • Updated Jul 17, 2026
  • Python

Data-Driven Evaluation for LLM-Powered Applications

  • Updated Aug 10, 2026
  • Python

Moonshot - A simple and modular tool to evaluate and red-team any LLM application.

  • Updated Jun 10, 2026
  • Python

Metrics to evaluate the quality of responses of your Retrieval Augmented Generation (RAG) applications.

  • Updated Jul 10, 2025
  • Python

Python SDK for running evaluations on LLM generated responses

  • Updated Jun 6, 2025
  • Python

MedEvalKit: A Unified Medical Evaluation Framework

  • Updated Feb 24, 2026
  • Python

build and benchmark deep research

  • Updated Mar 28, 2026
  • Python

A research library for automating experiments on Deep Graph Networks

  • Updated Dec 16, 2025
  • Python

AI Data Management & Evaluation Platform

  • Updated Oct 5, 2023
  • Svelte

Improve this page

Add a description, image, and links to the evaluation-framework topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the evaluation-framework topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL