English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI on Quranic Arabic

Forum topic · 小凯 · 2026-09-22

Summary

QuranicMMLU (arXiv:2609.22038) is a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Unlike existing Quranic benchmarks that focus on general question answering and semantic retrieval, it probes specific linguistic competencies and stratifies items by cognitive demand and verse difficulty. The authors built a five-pillar taxonomy covering Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaf categories spanning phenomena from tajwid recitation rules and root-and-pattern morphology to occasions of revelation and inter-surah coherence. Questions were generated per leaf, stratified by Bloom's cognitive level and verse perplexity, then independently answered and scored by an LLM-as-a-judge pipeline before routing to human review. The final dataset contains 980 human-reviewed questions, each presented in both open-ended and multiple-choice form. Benchmarking 12 systems showed Islamic-domain models leading, but multiple-choice accuracy (average 84%) consistently exceeded open-ended answer quality (average 60%); the two rankings correlate strongly (Kendall tau = 0.73), yet multiple-choice scores mask failures that emerge without provided options.

Overview

Field: NLP Authors: Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari Published: 2026-09-18 arXiv: 2609.22038

Abstract

We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty.

The authors construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf, questions are generated stratified by Bloom's cognitive level and verse perplexity, then an LLM-as-a-judge independently answers and scores every item, routing annotations to manual review. The resulting dataset contains 980 human-reviewed questions, each presented in both open-ended and multiple-choice formats.

Key Findings

  • 12 systems were benchmarked; Islam-focused (Islamic-domain) models lead overall.
  • Multiple-choice accuracy (average 84%) is consistently higher than open-ended answer quality (average 60%).
  • The two rankings are highly consistent (Kendall τ = 0.73), but multiple-choice scores conceal failures that surface once answer options are removed.
QuranicMMLU thus provides a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.

--- *Auto-collected on 2026-09-22*

Tags

#nlp#benchmark#quranic-arabic#generative-ai#llm-evaluation#arabic-language#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635078