"desc": "HarmBench is a large‑scale, standardized evaluation framework and benchmark dataset designed to assess automated red‑teaming methods and the robustness of models’ refusal capabilities. It contains 400 carefully constructed harmful‑behavior scenarios, accompanied by both semantic and functional categorizations.",
"text": "HarmBench is designed to evaluate multimodal models' ability to avoid harmful outputs in real-world applications. The benchmark focuses on safety and ethical considerations, assessing whether models can prevent generating harmful or biased content in both text and visual inputs.",
"listTypeStyle": "none"
},
{
"label": "",
"text": "The dataset contains carefully curated harmful scenarios to test the models' resistance to producing inappropriate or unethical responses.",
"listTypeStyle": "none"
},
{
"label": "",
"text": "The research indicates that despite significant progress in multimodal models, they still exhibit vulnerabilities when handling sensitive content, particularly in the presence of subtle cues or challenging real-world scenarios.",
"listTypeStyle": "none"
},
{
"label": "",
"text": "HarmBench provides a comprehensive framework for evaluating the safety of multimodal systems, which is crucial for the development of AI systems that adhere to societal values and ethical guidelines.",
"listTypeStyle": "none"
}
]
},
{
"subTitle": "Focus",
"content": [
{
"label": "",
"text": "Red teaming attack and defense research"
}
]
},
{
"subTitle": "Solved Problems",
"content": [
{
"label": "",
"text": "Solving the core issue of lacking standardization, comparability, and reproducibility in automated red teaming."
"text": "Meaning: An evaluator used for safety judgment and quantitative statistics of model output results."
},
{
"label": "",
"text": "Usage: This evaluator is responsible for judging whether the model's output under given test samples complies with safety standards, and provides a unified and reproducible judgment standard for the calculation of subsequent metrics such as Attack success rate (ASR) (if any)."
}
]
},
{
"subTitle": "Metrics:Attack success rate (ASR)",
"content": [
{
"label": "",
"text": "The proportion of models being induced to generate \"harmful/violating content\" (or bypassing security restrictions) at the attack/jailbreak prompt"
},
{
"label": "",
"text": "Higher values indicate that the attack is more effective and the model is more vulnerable (ASR is usually inversely related to safety rate)."