multimodal-learning
Filtering by topic multimodal-learning(2)Clear all filters
- PaperComputers and Education: Artificial Intelligence21 Jun 2026
Students’ multimodal prompting practices as epistemic work in AI literacy development
Sylvana Sofkova Hashemi
Investigates prompting strategies of 28 postgraduate students using a generative AI tool in collaborative multimodal tasks, finding strategies ranging from basic input-output to strategic, iterative, and dialogic practices. Prompting emerges as an epistemic practice for AI literacy, fostering critical interpretation and awareness of system limitations, while ethical dimensions remain underdeveloped. The study highlights the value of iterative, reflective, and multimodal learning designs for fostering critical and strategic engagement with AI.
Original abstract
: As generative artificial intelligence (GenAI) rapidly transforms higher education, critical questions arise about how students engage with these open-ended tools and the implications for learning. This study provides empirical insight into this research gap investigating (1) the prompting strategies students develop when interacting with a university-provided GenAI tool and (2) how engagement in prompt engineering activities shapes their understanding of GenAI and AI literacy. Data were collected in an exploratory workshop with 28 postgraduate students engaged in collaborative multimodal prompting tasks, including the creation of short stories or poems and corresponding images. Students’ self-documented prompting histories and reflections were analysed qualitatively using reflexive thematic analysis, guided by frameworks for prompting methods and AI literacy. The findings show that students’ prompting strategies vary along a continuum from basic input-output use to strategic, iterative, and dialogic practices. Prompting emerges as a central epistemic practice through which students critically interpret, refine, and negotiate AI-generated outputs. Multimodal engagement exposes challenges in translating abstract meaning into machine-readable prompts, fostering awareness of system limitations, bias, and the need to actively construct coherence across modalities. While students demonstrate developing competence in evaluation and creation, ethical dimensions of AI literacy remain underdeveloped. The findings provide empirical insight into how AI literacy develops through hands-on engagement with GenAI, positioning prompting as an epistemic practice through which students learn to interpret, negotiate, and guide AI-generated outputs, while highlighting the value of iterative, reflective, and multimodal learning designs that foster critical, strategic, and responsible engagement with AI.
- PaperarXiv — Language & NLP (cs.CL)27 May 2026
VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading
Jinzhou Wu, Zhengwu Ma, Jixing Li, Baoping Tang et al.
Comparing matched large language models (LLMs) and vision-language models (VLMs) under text-only conditions reveals that multimodal pretraining does not provide a universal advantage in aligning with human neural and eye-tracking data during natural reading. However, VLMs show selective improvement for sentences with strong visual semantic content, indicating that language-internal representations remain the primary driver of human-like text processing.
Original abstract
Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-language learning makes text representations more human-like during natural reading. Here, we address this question by comparing tightly matched LLM and vision-language model (VLM) pairs under a strictly text-only setting, allowing us to isolate the effect of multimodal training history from online visual input or cross-modal fusion. We evaluate model alignment with a human natural-reading dataset that includes whole-cortex fMRI responses and synchronized eye-tracking saccades. Our findings demonstrate that multimodal pretraining may not confer a uniform, global advantage in human alignment during natural reading, indicating that language-internal representations remain the key factor for modeling human text processing. However, the VLM advantage could emerge more selectively when sentences contain stronger visual semantic content, with converging evidence from both fMRI and eye-movement alignments. Together, our findings provide a controlled in silico framework for testing how visual learning history shapes model-human alignment of language processing, suggesting that multimodal pretraining contributes selectively rather than globally to human-like language representations during natural reading.