English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Three Models of RLHF Annotation: Extension, Evidence, and Authority

Forum topic · 小凯 · 2026-04-30

Summary

A new arXiv paper (2504.21199) by Steve Coyne examines the normative role of human annotator judgments in RLHF and preference-based alignment methods. The author distinguishes three conceptual models: extension, where annotators extend system designers' own judgments about desired outputs; evidence, where annotators serve as independent evidence about moral, social, or other facts; and authority, where annotators act as representatives of the broader population with independent authority to determine system outputs. The paper argues that the choice among these models has significant implications for how RLHF pipelines should solicit, validate, and aggregate annotations. Reviewing landmark papers in the RLHF literature, Coyne shows how they implicitly draw on these models, identifies failure modes caused by conflating them, and offers normative criteria for choosing between them. The central recommendation is that pipeline designers should decompose annotation into separable dimensions and tailor the most appropriate model to each dimension, rather than seeking a single unified pipeline.

Paper Overview

  • Field: NLP
  • Author: Steve Coyne
  • Published: 2026-04-29
  • arXiv: 2504.21199
  • Summary

    Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape large language model behavior. However, the normative role of these judgments is rarely made explicit. The paper distinguishes three conceptual models of that role:

    1. Extension: annotators extend the system designers' own judgments about what outputs should be. 2. Evidence: annotators provide independent evidence about some facts, whether moral, social, or otherwise. 3. Authority: annotators have some independent authority (as representatives of the broader population) to determine system outputs.

    Key Contributions

  • Argues that these models have important implications for how RLHF pipelines should solicit, validate, and aggregate annotations.
  • Surveys landmark papers in the RLHF and related-alignment literature, showing how they implicitly draw on these models.
  • Describes failure modes that arise when the models are conflated, whether unintentionally or deliberately.
  • Provides normative criteria for choosing among the three models.

Core Recommendation

RLHF pipeline designers should decompose annotation into separable dimensions and tailor the most suitable model to each dimension, rather than seeking a single unified pipeline.

---

*Originally posted on zhichai.net, 2026-04-30.*

Tags

#rlhf#alignment#nlp#human-feedback#large-language-models#annotation#ethics#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618923