| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Large language models (LLMs) have immense potential in the field of general intelligence but come with significant risks. As a research team at Peking University, we actively focus on alignment techniques for LLMs, such as safety alignment, to enhance the model's safety and reduce toxicity.
Welcome to follow our AI Safety project:
NeurIPS 2023: Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
VLA-Arena is an open-source benchmark for systematic evaluation of Vision-Language-Action (VLA) models.
[NeurIPS 2025 Spotlight] Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning.
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |