FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

ai-alignment ยท GitHub Topics ยท GitHub

#

ai-alignment

Here are 639 public repositories matching this topic...

Build reliable customer-facing AI agents with Parlant: an interaction control harness optimized for controlled, consistent, and predictable LLM interactions.

  • Updated Jul 12, 2026
  • Python

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

  • Updated Aug 28, 2026
  • Python

PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. ๐Ÿ† Best Paper Awards @ NeurIPS ML Safety Workshop 2022

  • Updated Apr 27, 2026
  • Python

Code accompanying the paper Pretraining Language Models with Human Preferences

  • Updated Feb 13, 2024
  • Python

How to Make Safe AI? Let's Discuss! ๐Ÿ’ก|๐Ÿ’ฌ|๐Ÿ™Œ|๐Ÿ“š

  • Updated Mar 29, 2023

A curated list of awesome academic research, books, code of ethics, courses, databases, data sets, frameworks, institutes, maturity models, newsletters, principles, podcasts, regulations, reports, responsible scale policies, tools and standards related to Responsible, Trustworthy, and Human-Centered AI.

  • Updated Aug 26, 2026

[AAAI'25 Oral] "MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector".

  • Updated Mar 17, 2025
  • Python

LLM alignment jailbreak; a set of instructions for auditing their internal reasoning and uncovering biases

  • Updated Aug 18, 2026

A curated list of awesome resources for Artificial Intelligence Alignment research

  • Updated Jul 14, 2023

Sparse probing paper full code.

  • Updated Dec 17, 2023
  • Jupyter Notebook

Directional Preference Alignment

  • Updated Sep 23, 2024

Adam Framework for OpenClaw โ€” 5-layer persistent memory and identity architecture for AI agents. Production-validated over 353+ sessions. First documented case of emergent values in persistent AI, quantum-verified on IBM hardware.

  • Updated Jul 12, 2026
  • Python

[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"

  • Updated Mar 2, 2026
  • Python

Official Implementation of Nabla-GFlowNet (ICLR 2025)

  • Updated May 3, 2025
  • Python

Website to track people, organizations, and products (tools, websites, etc.) in AI safety

  • Updated Aug 16, 2026
  • HTML

Just like the elite potential of a high-drive Belgian Malinois, an AI system's raw capabilities are wasted when deployed without proper structure. The technological value is no longer found in creating the drive, but in mastering the leash. Synapptic gives your AI Assistant persistent memory that updates in real time, saving you tokens and time.

  • Updated Mar 25, 2026
  • Python

Educational analysis of LLM alignment, safety behavior, and framing-sensitive response patterns.

  • Updated Nov 4, 2025

Improve this page

Add a description, image, and links to the ai-alignment topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-alignment topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL