Paper Overview
Field: AI Authors: Jing Huang, Jihong Zhang, Hua-Hua Chang Published: 2026-08-25 arXiv: 2608.24825
Abstract (Translated)
The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity often fail to simultaneously capture the nuanced structural and semantic layers that drive perceived redundancy.
This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness.
Key Findings
- Psychometric validation: LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and produce more coherent item parameter groupings than traditional text-based measurements.
- Computerized adaptive testing (CAT): Simulations show that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias, with minimal efficiency trade-offs, outperforming constraints based on traditional similarity metrics.
- Scalable item bank curation
- Content-aware test assembly
- Experience-sensitive adaptive testing across diverse assessment contexts
Implications
These findings highlight the potential of LLM-driven AISA for supporting:
*Originally collected on 2026-08-27.*