论文概要
研究领域: NLP
作者: Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
发布时间: 2026-09-18
arXiv: 2609.22038
中文摘要
我们提出 QuranicMMLU,一个在多种语言复杂度维度上评估生成式 AI 的古兰经阿拉伯语基准。现有古兰经基准以通用问答和语义检索为中心,既不探测特定的语言能力,也不按认知需求和经文难度分层。我们构建了一个五支柱古兰经分类体系,涵盖语音学、形态学、句法学、语义学和语用学,下设 31 个叶子节点,覆盖从 tajwīd(诵经规则)、词根-词式形态学到天降背景和各章连贯性等现象。针对每个叶子,我们生成按布鲁姆认知水平和经文困惑度分层的问题,并由 LLM 作为评判者独立作答与评分每个题目,再将标注路由至人工审查。最终数据集包含 980 道经人工审核的题目,每题同时以开放式和多项选择形式呈现。我们对 12 个系统进行基准测试,发现伊斯兰专用模型领先,但所有系统的多项选择准确率(平均 84%)均高于开放式回答质量(平均 60%):两个排名高度一致(Kendall τ=0.73),但多项选择评分掩盖了一旦去掉选项才暴露的失败。QuranicMMLU 因此为评估古兰经领域的阿拉伯语 NLP 提供了一个严谨、有语言学依据的框架。
原文摘要
We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting ...
自动采集于 2026-09-22
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。