English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Are Large Language Models Consistent over Value-laden Questions?

Forum topic · 小凯 · 2026-07-05

Summary

This paper by Jared Moore, Tanvi Deshpande, and Diyi Yang (arXiv:2407.02996, July 2024) examines whether large language models respond consistently to value-laden (ethically and politically sensitive) questions. The authors evaluate LLM responses across systematically varied question formats, including paraphrases, rephrasings, demographic framings, and answer-option orderings. They find that models frequently produce inconsistent or contradictory responses under small, semantically equivalent changes to a question, revealing substantial fragility in how models handle subjective and controversial topics. The paper argues that evaluating LLMs on value-laden questions with a single prompt format—such as standard multiple-choice benchmarks—overestimates reliability and can hide biased or unstable behavior. Instead, the authors recommend testing response robustness across multiple query variations before drawing conclusions about a model's values or trustworthiness. This work is relevant to AI evaluation, alignment, and responsible deployment, offering guidance for researchers and practitioners who assess LLM behavior on socially sensitive content.

Are Large Language Models Consistent over Value-laden Questions?

Authors: Jared Moore, Tanvi Deshpande, Diyi Yang Source: arXiv:2407.02996, July 2024

Overview

This paper investigates a critical weakness in how large language models (LLMs) are evaluated: their consistency when answering value-laden questions — questions involving ethics, politics, and other subjective or controversial topics.

Key points

  • Motivation: LLM responses to value-laden questions are increasingly used to judge model alignment, bias, and safety. However, if a model's answers shift dramatically under trivial rewording, single-shot evaluations may be misleading.
  • Method: The authors probe LLM responses across multiple variations of the same underlying question, including:
  • Paraphrases and rephrasings of the question
  • Variations in demographic framing (e.g., asking as or about different groups)
  • Changes to answer-option ordering in multiple-choice formats
  • Findings: Models show substantial inconsistency across these small, semantically equivalent perturbations, meaning conclusions drawn from a single query format can be unreliable or misleading.
  • Implication: Evaluations of model values and trustworthiness should test robustness across multiple question formulations rather than relying on one canonical prompt (e.g., a single MMLU-style multiple-choice item).
  • Takeaways

    1. Single-turn, single-format evaluations overestimate the reliability of LLM behavior on subjective topics. 2. Response inconsistency is itself an important evaluation signal — a well-calibrated, trustworthy model should give stable answers to equivalent questions. 3. Practitioners assessing deployment readiness for socially sensitive applications should include consistency checks in their evaluation suites.

    References

  • Original paper: Are Large Language Models Consistent over Value-laden Questions?
*Note: Quantitative results should be verified against the original PDF; this summary is based on the paper's abstract and public metadata.*

Tags

#large-language-models#llm-evaluation#consistency#value-laden-questions#ai-safety#alignment#benchmarking#responsible-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208692