When Sensors Talk: LLM-Guided Object-Centric Collaboration for Distributed 3D Scene Understanding

International Conference · 2026

When Sensors Talk: LLM-Guided Object-Centric Collaboration for Distributed 3D Scene Understanding

Ziheng Xu , Tatsuya Amano , Hamada Rizk , Hirozumi Yamaguchi

ICDCN 2026 Poster/Demo

DOI: 10.1145/3737611.3776964

Abstract

Collaborative 3D scene understanding enables multiple LiDAR or depth sensors to jointly perceive and reason about shared spaces, but transmitting full point clouds across devices is impractical under real-world bandwidth and latency constraints. Existing large language model (LLM)–based frameworks also typically assume access to a single, holistic 3D view. We propose an object-centric collaborative reasoning framework in which each device encodes local objects, shares lightweight summaries for multi-view fusion, and performs high-level reasoning through an LLM interface. Experiments on dense indoor scenes show that our method achieves accuracy close to raw-data fusion while greatly reducing communication bandwidth.

Research Note

A single LiDAR or depth sensor cannot see what is hidden behind an object. Placing several sensors so that they cover each other's blind spots solves the geometry, but sharing raw point clouds between them does not fit within realistic bandwidth and latency budgets, since a point cloud amounts to hundreds of thousands of points per second.

Framework overview

The useful observation is that what needs to be shared is not the points but the meaning behind them. An object-centric representation describes a space as a set of objects, one desk and two chairs and a person, rather than as a cloud of measurements. It discards most of the data while keeping what later reasoning actually depends on.

In the proposed framework each device encodes the objects it can see locally and shares only a lightweight summary. Summaries arriving from several viewpoints are fused into a single account of the scene, and higher-level reasoning over that account is carried out through a large language model interface. This differs from earlier LLM-based frameworks, which generally assume access to one holistic 3D view, by taking distributed viewpoints as the starting point.

Experiments on dense indoor scenes achieved accuracy close to fusing the raw data while greatly reducing the communication required.

セマンティック通信による多端末連携型の状況理解と消防システムへの適用

セマンティック通信による多端末連携型の状況理解と消防システムへの適用

Ministry of Internal Affairs and Communications (MIC) FORWARD
デジタルインフラ構築部門

Environment-Aware Distributed Scheduling for Emergency LoRa Networks

Yuto Inaba, Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), SPT-IoT 2026, pp. 1366–1371

DOI 10.1109/PerComWorkshops68308.2026.11585469

Disaster CommunicationLoRa +4

A Lightweight Vision-Language Model for Disaster Image Summarization

Hibiki Yoshizaki, Akira Uchiyama, Akihito Hiromori, Mineo Takai, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerconAI 2026, pp. 1203–1208

DOI 10.1109/PerComWorkshops68308.2026.11585419

Semantic CommunicationDisaster Response +4

Physics-Integrated Deep Learning for Urban Landslide Prediction

Ren Ozeki, Hamada Rizk, Hirozumi Yamaguchi

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), URBSENSE 2026, pp. 1094–1099

DOI 10.1109/PerComWorkshops68308.2026.11585337

Landslide PredictionPhysics-Integrated Learning +3

A Simulation Framework for Precision Formation Flying of Massive Satellite Swarms

Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi, Sumio Morioka

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerVehicle , pp. 230–235

DOI 10.1109/PerComWorkshops68308.2026.11585321

Satellite Formation FlyingDistributed Simulation +4

Ray-Tracing-Driven Pattern-Based Vehicle Recognition in ISAC Radar

Heetae Jin, Akira Uchiyama

2026 IEEE International Conference on Pervasive Computing and Communications Workshops and other Affiliated Events (PerCom Workshops), PerRad 2026, pp. 328–333

DOI 10.1109/PerComWorkshops68308.2026.11585327

ISACBeyond 5G +4

A Questionnaire-Only Counterfactual Machine Learning Approach to Assess the Spatial Impact of Green Mobility Vehicles in Urban Parks

Rami Naeem, Srikant Manas, Tatsuya Amano, Hirozumi Yamaguchi

ICDCN 2026 Workshop: IWNDSC2026

DOI 10.1145/3737611.3776620