Co-located with the IEEE International Conference on Multimedia and Expo 2026
5 July - 9 July 2026, Bangkok, Thailand
Embodied Intelligence is redefining AI by shifting focus from abstract reasoning to learning through physical interaction. In embodied systems, multimedia serves as the essential medium for perception, communication, and action, enabling intelligent agents to see, hear, sense, and act within dynamic environments. Because embodied AI creates closed-loop “Perception-Decision-Action” structures, the behavior of such systems emerges from continuous environmental engagement rather than static dataset processing.
Our workshop brings together researchers from signal processing, computer vision, robotics, and network communications to address critical challenges in building next-generation intelligent systems. We propose a Five-Layer Unified Architecture encompassing:
With spatial computing devices and humanoid robots becoming reality, the time is ripe for establishing Embodied Multimedia as a coherent research agenda.
We invite submissions on topics including but not limited to:
Papers must adhere to the standard ICME 2026 format (IEEE conference style, double-column, up to 6 pages including references) and be submitted via:
The template can be found via:
All Submissions will undergo a double-blind peer review process. Accepted papers will be presented by the authors and be included in the IEEE Xplore.
Build and Leverage World Model in the Latent Representation Space
World model is becoming a core technological paradigm connecting the digital and physical worlds. Their value lies in endowing intelligent agents with the ability to understand, reason, and interact in space, enabling machines to truly "enter" the physical world. This report will systematically outline the complete technological chain of spatial intelligence from 3D perception and world model construction to embodied task execution, and highlight the team's representative work in this field: the UniScene unified generation framework for autonomous driving and robotic scenarios; the DreamVLA model, which introduces world models into the VLA architecture to enhance long-range reasoning and action prediction capabilities through "world representation embedding"; benchmark datasets such as InterVLA & SceneScribe-1M, filling the gaps in embodied interaction data and spatial geometric video scene data from the first-person perspective; the DeFI framework, which solves the forward-backward dynamics model learning process; and VLA-JEPA, which endows VLA models with efficient reasoning capabilities from the latent space level, etc.
Xin Jin received his PhD degree from the University of Science and Technology of China (USTC). He is currently an Assistant Professor at the Eastern Institute of Technology, Ningbo, China. He was a visiting scholar at the Learning and Vision Lab of the National University of Singapore (NUS) from 2021 to 2022.
He has received several academic awards, including 2024 VSPC Rising Star, 2026 MSA TC Best Paper Award, the ACM SIGAI China Doctoral Dissertation Award, the Chinese Academy of Sciences President's Scholarship—Special Award, and the “Rising Star of Microsoft Research Asia” title. He has published over 40 papers in top-tier conferences and journals, including CVPR, ICCV, ECCV, NeurIPS, ACM MM, ICIP, ICME, ISCAS, VCIP, TIP, TMM, TCSVT, and Pattern Recognition, with nearly 9,000 Google Scholar citations. He was recognized as one of the top 2% of scientists worldwide by Stanford University in 2024&2025.
He has actively contributed to international and domestic standardization efforts in image/video compression, such as MPEG VCM and China DCM (Data Coding for Machine). He has organized multiple tutorials and special sessions at major conferences, including CVPR 2024, ECCV 2024, VCIP 2024, and ICME 2025, focusing on generative AI, diffusion models, and visual disentanglement. He also served as a session chair at ICIP 2024. In CVPR 2024, ECCV 2024, ICCV 2025, he organized tutorial sessions and workshop related to “Visual Representation Learning of Disentanglement and Compositionality”. In VCIP 2024, he also built up a special session about “Generative AI for Image/Video Coding”.
Should We Cross This Frontier? Embodied AI and Human Sustainability
The world is aging. As demographic shifts reshape our societies, maintaining health, autonomy, and quality of life has become one of the defining challenges of the coming decades. Artificial intelligence is increasingly presented as a technological response, with embodied AI expected to assist, collaborate with, and augment human capabilities in the physical world.
Recent advances are transforming embodied AI. Beyond classical imitation learning, new paradigms allow robots to learn through interaction and to autonomously acquire increasingly sophisticated behaviors. These developments promise more adaptive and resilient machines, capable of operating in complex, real-world environments. But they also invite a more fundamental question: what kind of embodied intelligence are we actually building, and for what purpose?
True embodiment cannot be reduced to placing a foundation model inside a robot. Following Gibson’s ecological theory of perception, intelligence emerges from the continuous coupling between an organism and its environment through action. Similarly, the Free Energy Principle describes cognition as an active process of self-maintenance, shaped by the perception–action loop and by the finitude of living systems, their limited energy, vulnerability, and need to survive. These principles suggest that intelligence is not simply computed but enacted through the dynamic relationship between body, environment, and self.
Yet today’s embodied AI remains, by design, an imitation of these processes. The more convincing this imitation becomes, the greater the temptation to attribute to machines properties such as agency, understanding, consciousness, or a genuine self. This may be the real frontier of embodied AI. The challenge is not only to develop increasingly capable systems, but also to preserve the conceptual distinction between simulating the dynamics of life and reproducing life itself. If embodied AI is to support human sustainability, we must ask not only what we can build, but also which frontier we should choose not to cross.
Prof Patrick Le Callet (Fellow, IEEE) received the M.Sc. and Ph.D. degrees in image processing from the Ecole Polytechnique de l’Université de Nantes. He was an Assistant Professor and a full-time Lecturer with the Department of Electrical Engineering, Technical Institute of the University of Nantes, from 1997 to 1999 and from 1999 to 2003, respectively. He led the Image and Video Communication Laboratory, CNRS IRCCyN, from 2006 to 2016. Since 2015, he has been the Scientific Director of the cluster Ouest Industries Creatives, a five-year program gathering over ten institutions (including three universities). He was one of the five members of the Steering Board of CNRS from 2013 to 2016. Since 2017, he has been one of the seven members of the Steering Board of the CNRS LS2N Laboratory (450 researchers) as a Representative of the Polytech Nantes. He is mostly involved in research dealing with the application of human vision modeling in image and video processing.
Tongji University, China
Dr. Yang Liu is an Assistant Professor at Tongji University, China. Dr. Liu received his Ph.D. from Fudan University (2025) and was a visiting scholar at the University of Toronto. His research focuses on anomaly detection in embodied systems. He has published in ACM CSUR, IEEE TIP, TII, and TIE, and served as Guest Editor for IEEE TCSS, Area Chair for BMVC 2025 and IEEE ICIP 2025, and Workshop/Special Session Co-Chair for IEEE ICASSP, ICIP, and DSP.
University of British Columbia, Canada
Dr. Jing Liu is a Postdoctoral Research Fellow at The University of British Columbia, Canada. Dr. Liu received his Ph.D. from Fudan University (2023). His research interests include edge-cloud collaboration, anomaly detection, and multimodal learning. He has published multiple first-author papers in top tier venues and serves as a reviewer for IEEE TII, TCSVT, IoTJ, TITS, TMC, TNNLS, ACM CSUR, and as a TPC member for CVPR, ECCV, ICCV, NeurIPS, ICLR, IJCAI, and ACM MM.
Cardiff University, UK
Dr. Wei Zhou is an Assistant Professor (UK Lecturer) at Cardiff University. Wei’s research interests mainly focus on perceptual image processing, multimodality, and visual computing for healthcare. Dr Zhou has published over 70 papers in recent years, including publications in top-tier venues, e.g., IEEE TIP, IEEE TMM, IEEE TCSVT, IEEE TMI, CVPR, ACM MM, MICCAI, etc. Wei serves as General Chair for the 1st Cardiff Image & Vision Computing Workshop and Chair for the Elections Committee of IEEE UK & Ireland SPS Chapter. Wei is now an Associate Editor of IEEE Transactions on Neural Networks and Learning Systems (TNNLS), Pattern Recognition, Neurocomputing, Springer Signal, Image and Video Processing, and Human-centric Computing and Information Sciences. Wei has also served as the Topic Editor and Guest Editor for many journals, such as Elsevier Displays; the Area Chair for ACM MM 2024, ICME 2025, and IJCNN 2025; the Lead Special Session Chair for IEEE ICME 2025, IEEE QoMEX 2025, IEEE MMSP 2023, and the Special Session Co-Chair for IEEE ICIP 2025 and 2024.
Tongji University, China
Dr. Lulu Guo is currently a Research Professor at Tongji University, Shanghai, China. He received the B.S. degree in vehicle engineering and the Ph.D. degree in control engineering from Jilin University, Changchun, China, in 2014 and 2019, respectively. Before joining Tongji University, he was a Postdoctoral Research Associate with the University of Georgia, Athens, GA, USA. His current research interests include advanced vehicle control, energy management, and vehicle cybersecurity.
Cardiff University, UK
Dr. Minghao Zou is a postdoctoral researcher at Cardiff University. His research interests span computer vision and process mining, with a particular focus on action recognition, object detection, and human-object interaction detection. He serves as a reviewer for several journals and conferences, such as IEEE TCSVT, IEEE TNNLS, ACM MM, IEEE TCSS, ACM TOMM, and Pattern Recognition.
Oral Jul 9 2:00 pm - Jul 9 2:15 pm
We study example-level private supervised speech classification under a practical release constraint: training may access privileged side information, but the released model must be audio-only. This setting is important because speech systems can often exploit richer side...
Oral Jul 9 2:15 pm - Jul 9 2:30 pm
Wearable AI and voice-first companions often require continuous location access, increasing point and trajectory privacy risks. We present Twinkle GPS, a practical GI-based framework that pairs an always-on low-precision baseline with intent-triggered high-utility bursts, adds a...
Oral Jul 9 2:30 pm - Jul 9 2:45 pm
Human-robot collaboration (HRC) in shared environments remains challenging due to partial observability, dynamic changes, and unpredictable human behavior, which often lead to inconsistent task understanding and redundant or conflicting actions among agents. Effective...
Oral Jul 9 2:45 pm - Jul 9 3:00 pm
Abstract—Predicting a user’s interaction target early and accurately is critical for responsive AR telepresence systems, where participants are located in heterogeneous physical spaces. Earlier approaches mainly rely on rules, geometric cues, or motion patterns, and often struggle to...
Oral Jul 9 3:00 pm - Jul 9 3:15 pm
Emotion recognition from psychological signals including electroencephalogram (EEG) has garnered significant attention due to its potential to enable intelligent and adaptative human-computer interaction. However, the non-Euclidean structure, high dimensionality, and noise...
Oral Jul 9 3:15 pm - Jul 9 3:30 pm
As wireless networks carry more multimedia traffic, conventional bit-centric transmission limits task-oriented efficiency. Semantic communication sends task-relevant meaning instead of raw bits, but end-to-end designs are often task-specific and costly to retrain. We build semantic...
Oral Jul 9 3:30 pm - Jul 9 3:45 pm
Learning chunk-based visuomotor policies for long-horizon robot manipulation remains challenging. Recent action-chunking methods have shown promising performance by predicting temporally extended action sequences. However, their failures are often dominated...
Poster Jul 9 3:45 pm - Jul 9 4:00 pm
Vision-and-Language Navigation (VLN) requires an agent to both interpret linguistic instructions and visual scenes, and to analyze their cross-modal semantic relationships for accurate decision-making. Most existing methods rely on heterogeneous multimodal feature...
Poster Jul 9 4:00 pm - Jul 9 4:15 pm
Traditional multimedia technology has been built around optimizing content delivery for human observers, from perceptually driven compression standards to human-centric quality metrics. With the rapid rise of embodied intelligence, autonomous agents must perceive,...