Paper Overview
Field: NLP Author: Haotao Xie Published: 2026-06-10 arXiv: 2606.12392
Abstract
Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited. The main challenge is that most studies treat the poetic appreciation task as a general-domain problem, neglecting the distinctive features of poetic appreciation, while high-quality and domain-specific datasets are extremely limited.
Key Contributions
- Task decomposition: The poetic appreciation task is split into three subtasks — term interpretation, semantic interpretation, and emotional inference.
- New dataset: Based on multiple open-source datasets with data cleansing and alignment, the authors construct CCPoetry-49K, a Classical Chinese Poetry instruction dataset containing 49,404 high-quality instruction-response pairs, explicitly optimized for this domain.
- Domain-specific LLM: PoetryQwen is created by fine-tuning the Qwen2.5-14B model with low-rank adaptation (LoRA).
Results
On the CCL25-Eval Task 5 benchmark:
| Model | Score | |---|---| | Qwen2.5-14B-Instruct (baseline) | 0.690 | | PoetryQwen | 0.757 (+9.7%) |
Conclusion
The findings show that PoetryQwen significantly enhances performance on precise translation and emotional understanding of classical poetry. The paper contributes both a new domain-specific dataset and methodological considerations aimed at supporting domain-specific optimization of LLMs.
--- *Source: arXiv:2606.12392*