Student Engagement Measurement Through Micro-Action Detection with Multi-Modal Foundation Model

国際会議 · 2025

Student Engagement Measurement Through Micro-Action Detection with Multi-Modal Foundation Model

Masato Matsuura , Tatsuya Amano , Hirozumi Yamaguchi

2025 Fifteenth International Conference on Mobile Computing and Ubiquitous Networking (ICMU), 2025, pp. 1-6

DOI: 10.23919/ICMU65253.2025.11219132

Abstract

Accurately measuring student engagement in collaborative learning environments is essential yet challenging, especially in dynamic group settings where traditional methods like analyzing facial expressions and speech patterns are often inadequate. This paper introduces a novel approach that focuses on detecting micro-actions, which offer a more consistent and reliable gauge of engagement across such settings. Instead of creating individual detectors for each micro-action, our method utilizes ImageBind, a pre-trained multimodal foundation model. This model efficiently encodes text, video, and audio into a unified latent space. We generate textual descriptions of these microactions and convert them, along with real-time classroom video and audio data, into comparable latent space representations. The resulting similarities between these representations are used as features for an engagement-level classifier. Our method was validated in a simulated university setting with eight groups of three students each performing a 20 -minute programming task, achieving an accuracy of 80.1 % in classifying engagement levels. It demonstrated real-time processing capabilities, confirming its practicality for classroom use. This enhances the deployment of AI-driven tools to improve educational outcomes by providing reliable engagement measurement.

解説

グループ学習が身についているかどうかを測ろうとすると、表情や発話の分析がよく使われますが、複数人が動き回る場では顔が映らない時間も長く、誰が話しているのかもはっきりしません。そこで本研究が着目したのが微細動作です。うなずく、手元の資料に手を伸ばす、身体の向きを変えるといった小さな動きは、そうした場でも比較的安定して観測でき、関与の度合いをよく表します。

Overview of the proposed method for student engagement measurement

ただ、微細動作ごとに検出器を作っていくと種類の数だけ学習データが必要になります。ここで使われているのがマルチモーダル基盤モデルという考え方です。ImageBind は、テキストと映像と音声を一つの共通した潜在空間に埋め込むよう事前学習されたモデルで、種類の違うデータどうしを同じ土俵で比較できるようにします。

この性質を使うと、微細動作の検出を「その動作を説明する文章と、いま撮れている映像と音声が、どれくらい近いか」という類似度の計算に置き換えられます。動作ごとの検出器を用意せずに済み、新しい動作を足したいときは説明文を書き足すだけで対応できます。本研究はこの類似度を特徴量として関与度の分類器に入力しています。

大学を模した環境で 3 人ずつ 8 グループが 20 分のプログラミング課題に取り組む様子で検証したところ、関与度の分類精度は 80.1% でした。実時間で処理できることも確認されており、教室での利用に耐えることを示しています。

災害時LoRaネットワークのための環境認識型分散スケジューリング

Yuto Inaba, Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), SPT-IoT 2026, pp. 1366–1371

DOI 10.1109/PerComWorkshops68308.2026.11585469

災害通信LoRa +4

災害現場画像要約のための軽量Vision-Language Model

Hibiki Yoshizaki, Akira Uchiyama, Akihito Hiromori, Mineo Takai, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerconAI 2026, pp. 1203–1208

DOI 10.1109/PerComWorkshops68308.2026.11585419

セマンティック通信災害対応 +4

物理モデル統合型深層学習による都市の土砂災害予測

Ren Ozeki, Hamada Rizk, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), URBSENSE 2026, pp. 1094–1099

DOI 10.1109/PerComWorkshops68308.2026.11585337

土砂災害予測物理モデル統合学習 +3

超大規模衛星群の精密編隊飛行に向けたシミュレーションフレームワーク

Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi, Sumio Morioka

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerVehicle , pp. 230–235

DOI 10.1109/PerComWorkshops68308.2026.11585321

衛星編隊飛行分散シミュレーション +4

レイトレーシング駆動型ISACレーダによるパターンベース車両認識

Heetae Jin, Akira Uchiyama

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerRad 2026, pp. 328–333

DOI 10.1109/PerComWorkshops68308.2026.11585327

ISACBeyond 5G +4

A Questionnaire-Only Counterfactual Machine Learning Approach to Assess the Spatial Impact of Green Mobility Vehicles in Urban Parks

Rami Naeem, Srikant Manas, Tatsuya Amano, Hirozumi Yamaguchi

ICDCN 2026 Workshop: IWNDSC2026

DOI 10.1145/3737611.3776620