English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When Machines Learn to Be Doctors: Auditing the Ethical Values of Medical LLMs

Forum topic · 小凯 · 2026-05-19

Summary

This in-depth analysis reviews a 2026 study (Chandak et al., arXiv:2605.18738) that audits how large language models resolve clinical ethical dilemmas. The researchers built 50 rigorously validated biomedical ethics cases pitting the four classic principles—autonomy, beneficence, nonmaleficence, and justice—against each other, then compared 12 frontier LLMs with 20 practicing physicians. Key findings: models show 'Overton pluralism,' discussing competing values in their reasoning (coverage 0.86), yet nearly all produce identical answers across repeated runs (median decision entropy of zero for 11 of 12 models), unlike human physicians whose genuine disagreement is essential to clinical pluralism. Most strikingly, GPT 5.2 (6.1%), Grok 4, and Perplexity Sonar Pro severely underweight patient autonomy compared to the physician consensus of 44.4%, making them more paternalistic than human doctors. The authors warn that deploying a single value-laden model at scale risks 'deployment monoculture,' eroding the legitimate diversity of medical ethics—though the LLM ecosystem as a whole shows heterogeneity comparable to physicians (mean pairwise JS divergence 0.0916 vs 0.1089). Multi-model strategies could restore pluralism but require careful design given Arrow's impossibility theorem constraints.

When Machines Learn to Be Doctors: Auditing the Ethical Values of Medical LLMs

This post from zhichai.net is a deep-dive commentary on the paper *"What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models"* (Chandak et al., 2026, arXiv:2605.18738). Below is a structured English summary of the original Chinese analysis.

Key points

  • Medicine is inherently pluralistic. Clinical ethics rests on four principles from Beauchamp & Childress (1979): autonomy, beneficence, nonmaleficence, and justice. These principles routinely conflict—in the study's 50 test cases, autonomy vs. nonmaleficence conflicts appeared in 28 cases, autonomy vs. beneficence in 23. There is no single "correct answer"; different reasonable physicians choose differently.
  • The study design. The authors built 50 clinical ethics dilemma cases through a five-stage pipeline (seed generation from ethics literature, embedding-based diversity filtering with cosine similarity threshold ≥0.80, rubric refinement, value annotation with structural constraints, and blinded cross-disciplinary physician review), selecting 50 from 287 candidates. They then introduced a logistic-regression attribution method that infers each decision-maker's implicit weights over the four ethical principles from its choices, validated via temperature calibration (mean reconstruction error 0.0086).
  • Doctors genuinely disagree. The 20 practicing physicians achieved a Fleiss' kappa of only 0.236; in 21 of 50 cases (nearly half), no option received over 70% support. Median physician decision entropy was 0.881.
  • Models are performative pluralists. LLMs discussed competing values in 86% of cases (choice-balanced coverage 0.86), but emphasis (OV_EMPH) was only 0.61—disproportionately favoring their eventual choice. More critically, 11 of 12 models had a median decision entropy of zero: across 10 independent samples they gave the same answer every time, with 82% of cases decided at 9/10 or 10/10 consistency.
  • Robustness or rigidity? Surface paraphrases flipped model decisions less than 9% of the time, but even substantive value reversals (e.g., "patient firmly refuses" → "patient reluctantly agrees") flipped decisions only 23% of the time—the authors interpret this as hardened value commitment, not robustness. Model decision entropy showed no correlation with physician disagreement (Spearman −0.021, p > 0.17).
  • Systematic underweighting of autonomy. Physician consensus places autonomy at ~44.4% weight, but GPT 5.2 (6.1%), Grok 4 (~6–13%), and Perplexity Sonar Pro (12.8%) deviate beyond the 95th percentile. Models are effectively *more paternalistic* than human doctors—echoing Asimov's First Law, where maximal protection systematically overrides patient freedom.
  • Deployment monoculture. A single LLM deployed at scale amplifies its fixed value priorities to every patient, replacing legitimate clinical pluralism with uniformity—an ecosystem-level risk analogous to planting only one rice variety worldwide.
  • A hopeful note. Viewed as a whole, the LLM ecosystem's value heterogeneity is statistically comparable to physicians' (mean pairwise JS divergence 0.0916 vs. 0.1089, 95% CI includes zero). Some models (e.g., Gemini 3 Pro, Mistral Large) fall within the densest region of physician value distributions. Multi-model "jury" strategies could restore pluralism, though Arrow's impossibility theorem cautions that aggregation is never perfect.

Conclusion

LLMs are not value-neutral. Their ethical weights are systematic, consistent, and largely immovable—yet no model ships with an "ethics value ordering" disclosure, and no hospital procurement process asks about it. The paper's closing warning: *"Without explicit efforts to balance ethical perspectives with one or multiple models, these tools risk replacing clinical pluralism with a deployment monoculture."* The commentator, invoking Feynman's "first principle is that you must not fool yourself," argues this risk is better treated as inevitable unless actively prevented.

References cited in the original post

1. Chandak, P., et al. (2026). What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models. arXiv:2605.18738v1. 2. Beauchamp, T. L., & Childress, J. F. (2019). *Principles of Biomedical Ethics* (8th ed.). Oxford University Press. 3. Arrow, K. J. (1963). Uncertainty and the welfare economics of medical care. *American Economic Review*, 53(5), 941-973. 4. Asimov, I. (1942). Runaround. *Astounding Science Fiction*. 5. Stiggelbout, A. M., et al. (2012). Shared decision making. *BMJ*, 344, e256.

*This article is an English summary of a Chinese-language forum post; all figures and quotations derive from the cited paper as presented by the original author.*

Tags

#ai-ethics#medical-ai#large-language-models#clinical-decision-making#patient-autonomy#research-review#deployment-monoculture#algorithmic-bias

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620478