Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

RLHF (Reinforcement Learning from Human Feedback)

A training methodology in which a language model is fine-tuned using a reward signal derived from human preferences over pairs (or larger sets) of candidate model outputs, typically by first training a reward model and then optimising the policy against it via PPO or a related algorithm.

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about RLHF (Reinforcement Learning from Human Feedback)

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    An instruction-tuned model fine-tuned with 100k human preference comparisons over candidate responses.

  • Is an instance

    A DPO-trained model using preference data without an explicit reward-model step.

Counter-examples

Looks similar, but isn't

  • Not an instance

    Pure supervised fine-tuning on labelled instruction data (SFT, not RLHF).

  • Not an instance

    Pre-training next-token prediction on web text.

Editorial commentary

Reinforcement learning from human feedback (RLHF) is a fine-tuning technique that aligns a pretrained language model’s behaviour to human preferences by training a separate reward model on human comparisons between candidate outputs, then optimising the language model against that reward model using reinforcement learning (originally proximal policy optimisation, PPO). Its modern form for instruction-following chat models was established by Ouyang et al. (2022, InstructGPT), building on the earlier general RL-from-human-preferences framework of Christiano et al. (2017).

RLHF became the dominant alignment technique for chat-tuned LLMs released from 2022 onward, and remains widely used, but it is no longer the only technique in production use. Newer methods reformulate the same underlying goal (make the model prefer outputs a human would prefer) without a separate RL loop: Direct Preference Optimisation (DPO) fits the preference data directly via a classification-style loss; IPO and KTO are further variants addressing specific weaknesses in DPO’s assumptions. Constitutional AI‘s reinforcement phase is a related but distinct variant, RLAIF (RL from AI Feedback), which substitutes AI-generated preference judgments for the human comparisons RLHF depends on.

Why it matters for documentation and disclosure

Whether, and with what technique, a model was preference-tuned materially affects its behaviour — refusal patterns, verbosity, sycophancy toward the user’s stated view, and safety-relevant responses are all shaped by this stage far more than by pretraining alone. It belongs in a model’s fine-tune lineage record and, where relevant, in its model card: naming the base model without disclosing the alignment technique applied on top of it is an incomplete provenance record.

References

  • Christiano et al., ‘Deep reinforcement learning from human preferences’ (NeurIPS 2017)
  • Ouyang et al., ‘Training language models to follow instructions with human feedback’ (NeurIPS 2022)

Also known as

RLHF

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="RLHF (Reinforcement Learning from Human Feedback)"
      vocab-term-identifier="https://casrai.org/dictionary/term/rlhf-reinforcement-learning-from-human-feedback" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/rlhf-reinforcement-learning-from-human-feedback",
  "name": "RLHF (Reinforcement Learning from Human Feedback)",
  "identifier": "https://casrai.org/dictionary/term/rlhf-reinforcement-learning-from-human-feedback",
  "description": "A training methodology in which a language model is fine-tuned using a reward signal derived from human preferences over pairs (or larger sets) of candidate model outputs, typically by first training a reward model and then optimising the policy against it via PPO or a related algorithm.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/rlhf-reinforcement-learning-from-human-feedback",
  "alternateName": [
    "RLHF"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-09-05T14:24:54",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.