[论文] Retrieving Biblical Intertextual References in Karen Blixen's Seven Go...

研究领域: NLP 作者: András Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen 发布时间: 2026-09-28 arXiv: 2609.35765

论文概要

研究领域: NLP 作者: András Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen 发布时间: 2026-09-28 arXiv: 2609.35765

中文摘要

识别互文引用是文学研究的核心,但在源材料经过改写、典故、历史语言和翻译转换时,计算上非常困难。我们通过卡伦·布利克森《七个哥特式故事》中的圣经互文性来研究这个问题。我们构建了一个包含 189 条标注引用的基准,针对全部 31,170 节历史上可能的丹麦语旧约和新约译本进行评估。我们比较了 TF-IDF 和 BM25 与多语言及丹麦语句向量编码器,考察语言规范化的影响,并使用难负例和五折交叉验证对丹麦语编码器进行微调。语言规范化后的 BM25 提供了强大的零样本基线,总体 R@10 达 0.365,在排名前十的节中检索到了每一条引文。微调使 DFM-large 的总体 R@10 从 0.265 提升至 0.508,典故性能从 0.138 翻倍以上至 0.339。然而,仅凭编辑评注评估会低估模型的学术实用性:一位文学学者判定 30 个被选为假阳性的排名第一预测中有 7 个是有意义的额外引用。我们提出将检索模型视为启发式共同阅读者——既恢复已记录的引用,也生成供专家细读的候选。

原文摘要

Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap ...


*自动采集于 2026-09-30*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens