LLM-Driven Adaptive Autonomous Robot Navigation via Multimodal Fusion for Diverse Environments

International Conference · 2025

LLM-Driven Adaptive Autonomous Robot Navigation via Multimodal Fusion for Diverse Environments

Keywords

Autonomous NavigationMultimodal FusionFPGA AccelerationLLM (Large Language Model)Socially Compliant Path PlanningDynamic Environment Adaptation

Xuqing Liu , Ahmed Farid , Riki Ukyoh , Tatsuya Amano , Hamada Rizk , Hirozumi Yamaguchi

2025 IEEE Intelligent Vehicles Symposium (IV), Cluj-Napoca, Romania, 2025, pp. 2361-2368

DOI: 10.1109/IV64158.2025.11097694

Abstract

This paper presents a novel autonomous navigation framework that integrates Large Language Models (LLMs) with multimodal sensor fusion to enable dynamic obstacle avoidance and human-aware path planning in diverse environments. The proposed system leverages an FPGA-accelerated fusion pipeline, combining LiDAR and vision data for real-time perception. A Hungarian algorithm-based object matching technique ensures robust tracking, while a bird's-eye view (BEV) representation enhances spatial reasoning and occlusion handling. The fused sensory inputs are processed by a fine-tuned LLM, which contextualizes pedestrian behavior and environmental constraints to generate adaptive, human-centric navigation strategies. Unlike traditional rule-based methods, LLMs provide generalization capabilities to novel scenarios, significantly improving interaction with vulnerable pedestrians such as children, elderly individuals, and wheelchair users. Extensive evaluations in both simulated and real-world scenarios confirm the system's ability to reduce collisions and enhance navigation efficiency in high-density environments. By bridging semantic reasoning and robotic control, this work lays the foundation for next-generation intelligent navigation systems that are both safety-aware and scalable across autonomous platforms.

Research Note

This research addresses the challenges of autonomous robot navigation in dynamic, high-density environments (e.g., train stations and shopping malls) by proposing a novel framework that integrates multimodal sensor fusion (LiDAR and vision) with a Large Language Model (LLM). To overcome the limitations of rule-based methods in handling unpredictable human behavior and dynamic obstacles, our system combines FPGA-accelerated real-time data processing and LLM-driven socially compliant path planning. Specifically, LiDAR point clouds and Triple-RGB camera data are fused on an FPGA using the Hungarian algorithm, while the LLM analyzes pedestrian attributes (age, wheelchair usage) to dynamically adjust navigation priorities. Experimental results demonstrate a 40% reduction in pedestrian prediction error compared to baseline models, with FPGA processing achieving sub-10ms latency. Future work includes enhancing inference accuracy via Q-LoRA and independent FPGA module verification.


LLM-Driven Urban Transportation Simulation Platform

LLM-Driven Urban Transportation Simulation Platform

Japan Science and Technology Agency (JST) JST Strategic Basic Research Programs (PRESTO)
[Social Transformation Platform] Co-Creation of the Transformation Platform Technology for Human and Society by Integration of the Humanities and Sciences

Environment-Aware Distributed Scheduling for Emergency LoRa Networks

Yuto Inaba, Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), SPT-IoT 2026, pp. 1366–1371

DOI 10.1109/PerComWorkshops68308.2026.11585469

Disaster CommunicationLoRa +4

A Lightweight Vision-Language Model for Disaster Image Summarization

Hibiki Yoshizaki, Akira Uchiyama, Akihito Hiromori, Mineo Takai, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerconAI 2026, pp. 1203–1208

DOI 10.1109/PerComWorkshops68308.2026.11585419

Semantic CommunicationDisaster Response +4

Physics-Integrated Deep Learning for Urban Landslide Prediction

Ren Ozeki, Hamada Rizk, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), URBSENSE 2026, pp. 1094–1099

DOI 10.1109/PerComWorkshops68308.2026.11585337

Landslide PredictionPhysics-Integrated Learning +3

A Simulation Framework for Precision Formation Flying of Massive Satellite Swarms

Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi, Sumio Morioka

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerVehicle , pp. 230–235

DOI 10.1109/PerComWorkshops68308.2026.11585321

Satellite Formation FlyingDistributed Simulation +4

Ray-Tracing-Driven Pattern-Based Vehicle Recognition in ISAC Radar

Heetae Jin, Akira Uchiyama

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerRad 2026, pp. 328–333

DOI 10.1109/PerComWorkshops68308.2026.11585327

ISACBeyond 5G +4

A Questionnaire-Only Counterfactual Machine Learning Approach to Assess the Spatial Impact of Green Mobility Vehicles in Urban Parks

Rami Naeem, Srikant Manas, Tatsuya Amano, Hirozumi Yamaguchi

ICDCN 2026 Workshop: IWNDSC2026

DOI 10.1145/3737611.3776620