English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PoetryQwen: LoRA-Fine-Tuned LLM for Classical Chinese Poetry Translation and Emotional Understanding (CCL25-Eval Task 5)

Forum topic · 小凯 · 2026-06-12

Summary

This arXiv paper (2606.12392) by Haotao Xie introduces a domain-specific large language model for classical Chinese poetry appreciation. The work decomposes poetic appreciation into three subtasks—term interpretation, semantic interpretation, and emotional inference—and constructs CCPoetry-49K, a dataset of 49,404 high-quality instruction-response pairs curated through data cleansing and alignment of multiple open-source sources. The author fine-tunes Qwen2.5-14B using low-rank adaptation (LoRA) to create PoetryQwen. On the CCL25-Eval Task 5 benchmark, PoetryQwen scores 0.757, a 9.7% improvement over the Qwen2.5-14B-Instruct baseline (0.690). The results demonstrate that domain-focused data and parameter-efficient fine-tuning significantly enhance precise translation and affective-semantic understanding of classical Chinese poetry, offering both a new dataset and methodological guidance for domain-specific LLM optimization.

Paper Overview

Field: NLP Author: Haotao Xie Published: 2026-06-10 arXiv: 2606.12392

Abstract

Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical poetry. However, domain-specific research on precise translation and affective-semantic understanding of classical poetry remains limited. The main challenge is that most studies treat the poetic appreciation task as a general-domain problem, neglecting the distinctive features of poetic appreciation, while high-quality and domain-specific datasets are extremely limited.

Key Contributions

  • Task decomposition: The poetic appreciation task is split into three subtasks — term interpretation, semantic interpretation, and emotional inference.
  • New dataset: Based on multiple open-source datasets with data cleansing and alignment, the authors construct CCPoetry-49K, a Classical Chinese Poetry instruction dataset containing 49,404 high-quality instruction-response pairs, explicitly optimized for this domain.
  • Domain-specific LLM: PoetryQwen is created by fine-tuning the Qwen2.5-14B model with low-rank adaptation (LoRA).

Results

On the CCL25-Eval Task 5 benchmark:

| Model | Score | |---|---| | Qwen2.5-14B-Instruct (baseline) | 0.690 | | PoetryQwen | 0.757 (+9.7%) |

Conclusion

The findings show that PoetryQwen significantly enhances performance on precise translation and emotional understanding of classical poetry. The paper contributes both a new domain-specific dataset and methodological considerations aimed at supporting domain-specific optimization of LLMs.

--- *Source: arXiv:2606.12392*

Tags

#llm#nlp#classical-chinese-poetry#lora#fine-tuning#qwen#dataset#ccl25-eval

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981122