| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
What is OpenCompass ? OpenCompass is a platform focused on understanding of the AGI, include Large Language Model and Multi-modality Model.
We aim to:
OpenCompass
VLMEvalKit
CompassVerifier
CompassJudger
| Project | Topic | Paper |
| Automated Software Development | ||
| Critic Reasoning | ||
| Hallucination Annotation |
ANAH: Analytical Annotation of Hallucinations in Large Language Models |
|
| Mathematical Reasoning | ||
| Tool Utilization |
T-Eval: Evaluating the Tool Utilization Capability Step by Step |
|
| Multi Modality | ||
| Subjective Evaluation |
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues |
|
| Domain Evaluation |
LawBench: Benchmarking Legal Knowledge of Large Language Models |
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
AgentCompass is an extensible open-source evaluation infrastructure for systematically assessing LLM/VLM agent capabilities.
CNFinBench — the first comprehensive benchmark for high-stakes financial scenarios. It spans 29 subtasks grounded in authoritative financial corpora and real business contexts, reconstructing end-to-end agent execution chains from requirement parsing, path planning, tool invocation, to result verification.
The first unified, efficient, and extensible evaluation toolkit for evaluating image generation and editing models across multiple benchmarks.
We provide TextEdit, a high-quality, multi-scenario text editing benchmark for generation models.
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |