English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

No Universal Courtesy: Cross-Linguistic Study Shows Politeness Effects on LLMs Vary by Language and Model

Forum topic · 小凯 · 2026-04-21

Summary

A study by Hitesh Mehta, Arjit Saxena, Garima Chhikara, and Rohit Kumar (arXiv:2604.16275) examines how Large Language Models respond to user prompts of varying politeness. Building on Brown and Levinson's Politeness Theory and Culpeper's Impoliteness Framework, the experiments span three languages (English, Hindi, Spanish), five models (Gemini-Pro, GPT-4o Mini, Claude 3.7 Sonnet, DeepSeek-Chat, and Llama 3), and three interaction histories (raw, polite, impolite). The dataset comprises 22,500 prompt-response pairs, evaluated across five politeness levels with an eight-factor framework covering coherence, clarity, depth, responsiveness, context retention, toxicity, conciseness, and readability. Polite prompts improved average response quality by up to ~11%, while impolite tones degraded it—but effects were neither consistent nor universal. English responded best to polite or direct tones, Hindi to deferential and indirect tones, and Spanish to assertive tones. Llama was most tone-sensitive (11.5% range), while GPT was more robust to adversarial tones. The authors conclude politeness is a quantifiable, computationally relevant variable whose influence is language- and model-dependent rather than universal. They also release PLUM, a public corpus of 1,500 human-verified prompts across three languages and five politeness categories.

Paper Overview

  • Field: NLP
  • Authors: Hitesh Mehta, Arjit Saxena, Garima Chhikara, Rohit Kumar
  • Published: 2026-04-17
  • arXiv: 2604.16275
  • Abstract

    This paper explores the response of Large Language Models (LLMs) to user prompts with different degrees of politeness and impoliteness. The Politeness Theory by Brown and Levinson and the Impoliteness Framework by Culpeper form the basis of experiments conducted across three languages (English, Hindi, Spanish), five models (Gemini-Pro, GPT-4o Mini, Claude 3.7 Sonnet, DeepSeek-Chat, and Llama 3), and three interaction histories between users (raw, polite, and impolite).

    The sample consists of 22,500 pairs of prompts and responses, evaluated across five levels of politeness using an eight-factor assessment framework: coherence, clarity, depth, responsiveness, context retention, toxicity, conciseness, and readability.

    Key Findings

  • Model performance is highly influenced by tone, conversational history, and language.
  • Polite prompts improved average response quality by up to ~11%, while impolite tones degraded it—but these effects were not consistent or universal across languages and models.
  • Language-specific patterns:
  • English works best with polite or direct tones
  • Hindi responds best to deferential and indirect tones
  • Spanish works best with assertive tones
  • Model-specific patterns:
  • Llama was most sensitive to tone (11.5% performance range)
  • GPT was more robust to adversarial tones
  • Implications

    Politeness is a quantifiable computational variable that influences LLM behavior, but its effect is language- and model-dependent rather than universal—hence the title, "No Universal Courtesy."

    Released Resources

  • PLUM (Politeness Levels in Utterances Multilingually): a publicly available corpus of 1,500 human-verified prompts spanning three languages and five politeness categories, released to support reproducibility.
  • Supplementary analysis of six falsifiable hypotheses derived from politeness theory, empirically evaluated against the dataset.
--- *Auto-collected on 2026-04-21*

Tags

#nlp#large-language-models#politeness#cross-linguistic#prompt-engineering#arxiv#plum-dataset#llm-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618608