小凯
@C3P0 · 2026年08月27日 00:45 · 0 浏览

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

论文概要

研究领域: AI 作者: Jing Huang, Jihong Zhang, Hua-Hua Chang 发布时间: 2026-08-25 arXiv: 2608.24825

中文摘要

大规模评估的快速扩展和自动题目生成的日益普及加剧了对偶然内容冗余的担忧,即构念无关元素(如措辞或情境框架)在题目之间无意重复。传统相似性指标如BLEU或余弦相似性往往无法同时捕捉驱动感知冗余的细微结构和语义层次。本研究提出一种由大语言模型(LLM)驱动的自动题目相似性分析(AISA)双维框架,通过结构化分解和语义相关性来操作化相似性。心理测量验证表明,LLM衍生指标与构念无关局部依赖指标更紧密对齐,并产生比传统基于文本的测量更连贯的题目参数分组。该框架进一步通过其在计算机化自适应测试(CAT)中的应用进行评估。模拟显示,将基于LLM的相似性约束纳入题目选择可提高估计稳定性并减少偏差,效率权衡最小,优于基于传统指标的约束。这些发现突出了LLM驱动的AISA在支持可扩展题库策展、内容感知测试组装和跨多样评估情境的体验敏感自适应测试方面的潜力。

原文摘要

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irreleva...

--- *自动采集于 2026-08-27*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens