?
Presentation modality and comprehension of a factually dense educational text: The role of cognitive load and metacognitive calibration
Introduction. The article addresses the scientific problem of whether a more vivid and engaging presentation modality necessarily supports more accurate comprehension of the same educational content. This problem is especially important for factually dense educational texts that require learners to retain proper names, chronological order, cultural references and causal relations. The aim of the study is to determine how written text, audio and video presentation of one and the same educational stimulus are associated with objective comprehension, productive recall, subjective cognitive load, evaluative perception and metacognitive calibration.
Materials and Methods. The study was conducted as an original between-subjects psycholinguistic experiment. A total of 456 participants worked with the stimulus “The History of Ice Cream” in one of three presentation conditions: written text (n = 146), audio (n = 167) or video (n = 143). The methodological framework combined the cognitive theory of multimedia learning, cognitive load theory, discourse-comprehension models and research on metacognitive calibration. The instruments included 14 objective comprehension items, an open summary scored with a 0–10 rubric, a
semantic differential, and subjective ratings of material difficulty, task difficulty, confidence, self-rated comprehension, memory accessibility and interest.
Subjective perception was measured with a 12-scale semantic differential and six additional ratings: difficulty of the material, difficulty of the tasks, confidence, self-rated comprehension, memory accessibility and interest. The data were analyzed using Kruskal–Wallis tests, Mann–Whitney tests with Holm correction, effect sizes, Spearman correlations, regression and exploratory mediation analysis.
Results. Written text produced the highest objective comprehension score, followed by video and audio. The written-text advantage was most pronounced for exact factual attribution, proper names and chronological reconstruction. Video was evaluated as more dynamic, interesting and visually rich, but this affective advantage did not result in the highest factual comprehension. Cognitive load was the strongest negative predictor of objective performance, whereas imagery/detail contributed positively and engagement did not show an independent direct effect after cognitive load was taken into account.
Conclusions. The main result of the study is an experimental confirmation of the dissociation between subjective engagement and precise comprehension in multimodal educational input. For factually dense material, the effectiveness of a modality should be evaluated not by its attractiveness alone, but by the balance between external cognitive support, experienced cognitive load, productive recall and metacognitive calibration.