Student Engagement Measurement Through Micro-Action Detection with Multi-Modal Foundation Model

International Conference · 2025

Student Engagement Measurement Through Micro-Action Detection with Multi-Modal Foundation Model

Masato Matsuura , Tatsuya Amano , Hirozumi Yamaguchi

2025 Fifteenth International Conference on Mobile Computing and Ubiquitous Networking (ICMU), 2025, pp. 1-6

DOI: 10.23919/ICMU65253.2025.11219132

Abstract

Accurately measuring student engagement in collaborative learning environments is essential yet challenging, especially in dynamic group settings where traditional methods like analyzing facial expressions and speech patterns are often inadequate. This paper introduces a novel approach that focuses on detecting micro-actions, which offer a more consistent and reliable gauge of engagement across such settings. Instead of creating individual detectors for each micro-action, our method utilizes ImageBind, a pre-trained multimodal foundation model. This model efficiently encodes text, video, and audio into a unified latent space. We generate textual descriptions of these microactions and convert them, along with real-time classroom video and audio data, into comparable latent space representations. The resulting similarities between these representations are used as features for an engagement-level classifier. Our method was validated in a simulated university setting with eight groups of three students each performing a 20 -minute programming task, achieving an accuracy of 80.1 % in classifying engagement levels. It demonstrated real-time processing capabilities, confirming its practicality for classroom use. This enhances the deployment of AI-driven tools to improve educational outcomes by providing reliable engagement measurement.

Research Note

Measuring whether students are engaged in collaborative learning is usually attempted through facial expression or speech analysis, but in a group that moves around, faces are often out of view and it is not always clear who is speaking. This work looks instead at micro-actions. Small movements such as nodding, reaching for a document or turning the body remain observable in these settings and track engagement fairly reliably.

Overview of the proposed method for student engagement measurement

Building a separate detector for each micro-action would require labelled data for every one of them. The alternative used here rests on multimodal foundation models. ImageBind is pretrained to embed text, video and audio into a single shared latent space, which makes data of different kinds directly comparable.

That property turns detection into a similarity computation. Written descriptions of the micro-actions are embedded alongside live classroom video and audio, and the resulting similarities indicate which actions are occurring. No per-action detector is needed, and covering a new action amounts to writing another description. These similarities are then used as features for an engagement-level classifier.

Validation in a simulated university setting, with eight groups of three students each working on a twenty-minute programming task, gave 80.1% accuracy in classifying engagement levels, and the system ran in real time, confirming that it is practical for classroom use.

Environment-Aware Distributed Scheduling for Emergency LoRa Networks

Yuto Inaba, Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), SPT-IoT 2026, pp. 1366–1371

DOI 10.1109/PerComWorkshops68308.2026.11585469

Disaster CommunicationLoRa +4

A Lightweight Vision-Language Model for Disaster Image Summarization

Hibiki Yoshizaki, Akira Uchiyama, Akihito Hiromori, Mineo Takai, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerconAI 2026, pp. 1203–1208

DOI 10.1109/PerComWorkshops68308.2026.11585419

Semantic CommunicationDisaster Response +4

Physics-Integrated Deep Learning for Urban Landslide Prediction

Ren Ozeki, Hamada Rizk, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), URBSENSE 2026, pp. 1094–1099

DOI 10.1109/PerComWorkshops68308.2026.11585337

Landslide PredictionPhysics-Integrated Learning +3

A Simulation Framework for Precision Formation Flying of Massive Satellite Swarms

Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi, Sumio Morioka

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerVehicle , pp. 230–235

DOI 10.1109/PerComWorkshops68308.2026.11585321

Satellite Formation FlyingDistributed Simulation +4

Ray-Tracing-Driven Pattern-Based Vehicle Recognition in ISAC Radar

Heetae Jin, Akira Uchiyama

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerRad 2026, pp. 328–333

DOI 10.1109/PerComWorkshops68308.2026.11585327

ISACBeyond 5G +4

A Questionnaire-Only Counterfactual Machine Learning Approach to Assess the Spatial Impact of Green Mobility Vehicles in Urban Parks

Rami Naeem, Srikant Manas, Tatsuya Amano, Hirozumi Yamaguchi

ICDCN 2026 Workshop: IWNDSC2026

DOI 10.1145/3737611.3776620