English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Dual-Dimensional LLM Framework for Automated Item Similarity Analysis in Large-Scale Assessments

Forum topic · 小凯 · 2026-08-27

Summary

Researchers Jing Huang, Jihong Zhang, and Hua-Hua Chang propose a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by large language models (LLMs), addressing incidental content redundancy in large-scale assessments. As automatic item generation proliferates, construct-irrelevant elements such as wording or contextual framing can repeat unintentionally across items, threatening measurement validity. Traditional metrics like BLEU and cosine similarity fail to capture both structural and semantic layers of redundancy. The proposed framework operationalizes similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation shows LLM-derived metrics align more closely with construct-irrelevant local dependence indicators and yield more coherent item parameter groupings than conventional text-based measures. Applied to computerized adaptive testing (CAT), simulations demonstrate that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on traditional metrics. The study highlights LLM-driven AISA's potential for scalable item bank curation, content-aware test assembly, and experience-sensitive adaptive testing. Paper: arXiv:2608.24825.

Paper Overview

Field: AI Authors: Jing Huang, Jihong Zhang, Hua-Hua Chang Published: 2026-08-25 arXiv: 2608.24825

Abstract (Translated)

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity often fail to simultaneously capture the nuanced structural and semantic layers that drive perceived redundancy.

This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness.

Key Findings

  • Psychometric validation: LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and produce more coherent item parameter groupings than traditional text-based measurements.
  • Computerized adaptive testing (CAT): Simulations show that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias, with minimal efficiency trade-offs, outperforming constraints based on traditional similarity metrics.
  • Implications

    These findings highlight the potential of LLM-driven AISA for supporting:

  • Scalable item bank curation
  • Content-aware test assembly
  • Experience-sensitive adaptive testing across diverse assessment contexts
---

*Originally collected on 2026-08-27.*

Tags

#llm#psychometrics#automated-item-generation#computerized-adaptive-testing#assessment#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634093