Paper Overview
Field: NLP Authors: Nethmi Muthugala, Supryadi, Surangika Ranathunga Published: 2026-07-24 arXiv: 2507.18391
Abstract
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning.
To bridge this gap, the authors propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment.
Key Contributions
- 40 majority-endorsed societal values derived from a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs.
- LKvaluesIT: a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances grounded in the identified values.
- LKvaluesBench: a value-sensitive evaluation benchmark with 1,000 instances.
- A range of proprietary and open-weight LLMs was evaluated using LKvaluesBench.
- Three open-weight base models — Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base — were fine-tuned with LKvaluesIT.
- Even newer and larger LLMs exhibit low-resource and cultural value alignment gaps.
- LKValues fine-tuning improved both English and Sinhala performance for the Qwen-family models, reducing invalid outputs and cross-lingual divergence, though benefits remain dependent on the model family.
Experiments and Findings
Resources
The datasets are publicly available at: https://github.com/NextME14/LKValues
The work highlights LKValues' effectiveness in embedding Sri Lankan values and provides a replicable pipeline for low-resource, country-specific pluralistic value alignment.
--- *Auto-collected on 2026-07-24*