English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LKValues: Aligning LLMs with Sri Lankan Societal Values via Sinhala-English Resources

Forum topic · 小凯 · 2026-07-24

Summary

Researchers introduce LKValues, the first survey-grounded resource suite for aligning large language models (LLMs) with Sri Lankan societal values. Drawing on a trilingual survey of 205 respondents that combines adapted global value frameworks with LLM-elicited local constructs, the authors identify 40 majority-endorsed societal values. From these they build LKvaluesIT, a 150k-instance Sinhala-English news-derived instruction corpus of scenario-based examples, and LKvaluesBench, a 1,000-instance value-sensitive evaluation benchmark. Evaluations across proprietary and open-weight LLMs reveal persistent cultural and low-resource alignment gaps, even in newer and larger models. Fine-tuning three open-weight base models (Qwen 4B/9B variants and Aya-Expanse-8B) with LKValuesIT improved English and Sinhala performance, reduced invalid outputs and cross-lingual divergence, though gains depended on the model family. The work offers a replicable pipeline for low-resource, country-specific pluralistic value alignment, with datasets publicly available on GitHub.

Paper Overview

Field: NLP Authors: Nethmi Muthugala, Supryadi, Surangika Ranathunga Published: 2026-07-24 arXiv: 2507.18391

Abstract

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning.

To bridge this gap, the authors propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment.

Key Contributions

  • 40 majority-endorsed societal values derived from a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs.
  • LKvaluesIT: a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances grounded in the identified values.
  • LKvaluesBench: a value-sensitive evaluation benchmark with 1,000 instances.
  • Experiments and Findings

  • A range of proprietary and open-weight LLMs was evaluated using LKvaluesBench.
  • Three open-weight base models — Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base — were fine-tuned with LKvaluesIT.
  • Even newer and larger LLMs exhibit low-resource and cultural value alignment gaps.
  • LKValues fine-tuning improved both English and Sinhala performance for the Qwen-family models, reducing invalid outputs and cross-lingual divergence, though benefits remain dependent on the model family.

Resources

The datasets are publicly available at: https://github.com/NextME14/LKValues

The work highlights LKValues' effectiveness in embedding Sri Lankan values and provides a replicable pipeline for low-resource, country-specific pluralistic value alignment.

--- *Auto-collected on 2026-07-24*

Tags

#llm-alignment#sinhala#cultural-values#nlp#low-resource-languages#fine-tuning#benchmark#sri-lanka

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447051