?
Toward Interpretable Automated Item Generation: Evaluating Q-Matrix Representations of Item Difficulty Using LLTM
Automated Item Generation (AIG) has expanded the scale and efficiency of test development yet reliably predicting and controlling item difficulty remains a central challenge. This study investigates how different Q-matrix representations influence the explanatory power of the Linear Logistic Test Model (LLTM) in accounting for difficulty in a mathematics literacy item bank. We compared expert-defined, framework-based, LLM-derived, and LLM-assisted Q-matrix configurations (a hybrid version combining LLM-derived structural features with thematic domain descriptors). Results show that broad domain-level and expert-defined matrices, while pedagogically meaningful, insufficiently captured variation in empirical difficulty. In contrast, LLM-assisted derived and LLM-assisted (a hybrid version combining LLM-derived structural features with thematic domain descriptors) hybrid configurations provided clearer and more consistent mappings between item features and difficulty patterns. These findings highlight the importance of systematically specifying item features and demonstrate that combining computational refinement with domain-informed attributes enhances interpretability. The study provides an empirical foundation for developing structured feature sets that can support difficulty-aware item generation, enabling the design of tasks with predictable cognitive demands and more stable psychometric properties.