FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

reward-modeling · GitHub Topics · GitHub

#

reward-modeling

Here are 67 public repositories matching this topic...

[ICLR 2025] IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation

  • Updated Feb 19, 2025
  • Python

[CVPR 2026] Official Code for "ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning"

  • Updated Feb 13, 2026
  • Python

Local outcome-loop + reward layer for coding agents — remember what worked, ship the next move, measure how it landed. Docs at nimroboai.com

  • Updated Jul 10, 2026
  • TypeScript

A comrephensive collection of learning from rewards in the post-training and test-time scaling of LLMs, with a focus on both reward models and learning strategies across training, inference, and post-inference stages.

  • Updated Jun 13, 2025

official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives

  • Updated Apr 2, 2025
  • Python

Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CISPO, VAPO, Reward Modeling, MoE RL stability, Verifier-Free RL, Training-Free RL, Agentic RL, DeepSeek-R1 reproduction.

  • Updated Aug 26, 2026

This training offers an intensive exploration into the frontier of reinforcement learning techniques with large language models (LLMs). We will explore advanced topics such as Reinforcement Learning with Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), Reasoning LLMs, and demonstrate practical applications such as fine-tuning

  • Updated Mar 9, 2026
  • Jupyter Notebook

An easy python package to run quick basic QA evaluations. This package includes standardized QA evaluation metrics and semantic evaluation metrics: Black-box and Open-Source large language model prompting and evaluation, exact match, F1 Score, PEDANT semantic match, transformer match. Our package also supports prompting OPENAI and Anthropic API.

  • Updated Jul 18, 2025
  • Python

[CVPR 2025] Science-T2I: Addressing Scientific Illusions in Image Synthesis

  • Updated Mar 31, 2026
  • Python

(🔥ICML2026) Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios

  • Updated Jan 24, 2026
  • Python

Learning to route instances for Human vs AI Feedback (ACL Main '25)

  • Updated Jul 23, 2025
  • Python

Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.

  • Updated Aug 25, 2026
  • Python

A modular, production-grade framework for Reinforcement Learning from Human Feedback (RLHF) with Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO).

  • Updated Jun 6, 2026
  • Python

Revealing and unlocking the context boundary of reward models

  • Updated May 10, 2026
  • Python

[ACL2024 Findings]DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling

  • Updated Jun 6, 2024
  • Python

Reward model engineering harness for evolutionary rubric search, deployable RM artifacts, online scoring, and RL experiment lineage.

  • Updated May 1, 2026
  • Python

ToolRM: Towards Agentic Tool-Use Reward Modeling

  • Updated Jan 14, 2026
  • Python

Code for SFT and RL

  • Updated Jun 22, 2025
  • Python

A novel Group Relative Reward Model (GRRM) framework enhances machine translation quality and reasoning capabilities by improving intra-group ranking through comparative analysis rather than isolated metric evaluation.

  • Updated Mar 27, 2026
  • Python

Improve this page

Add a description, image, and links to the reward-modeling topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the reward-modeling topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL