| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Forked from huggingface/trl
[Downstream Fork DO NOT EDIT MAIN] Train transformer language models with reinforcement learning.
Python
[Fork] An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)
Python
Forked from numirias/pytest-json-report
🗒️ A pytest plugin to report test results as JSON
Python
Forked from HKUNLP/critic-rl
Code for Paper: Teaching Language Models to Critique via Reinforcement Learning
Python
Forked from hiyouga/LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Python
Harbor is a framework for running agent evaluations and creating and using RL environments.
OpenSandbox is a general-purpose sandbox platform for AI applications, offering multi-language SDKs, unified sandbox APIs, and Docker/Kubernetes runtimes for scenarios like Coding Agents, GUI Agents, Agent Evaluation, AI Code Execution, and RL Training.
Convert GitHub PRs into Harbor tasks. Now with OBS cloning and Voyager^tm uploading capabilities!
DSM-AE: diagnostic engine for agentic ill-behaviours (indicator protocols + multi-model matrix)
New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment
A Claude Code plugin that reverse-engineers clean behavioral specs, test vectors, and acceptance criteria from any codebase, producing a provenance trail so a fresh team can reimplement without inheriting the original's internal structure.
SlopCodeBench: Measuring Code Erosion Under Iterative Specification Refinement
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
Loading…
| Back | FazBrowse Home | New Git URL |