Systems | Information | Learning | Optimization
 

SILO: Distribution Transport with Identifiability: A Signal Processing Perspective on Multimodal Generative AI

Abstract

A central objective of modern generative AI is to enable learning and inference across multiple modalities, supporting tasks such as cross-modal generation, multimodal translation, domain transfer, and data fusion. Despite impressive empirical progress, many multimodal AI frameworks still lack solid theoretical foundations, raising concerns about reliability and robustness. In this talk, we present a signal processing perspective on AI, where classical tools such as factor analysis-based modeling and identifiability analysis are used to understand and advance modern learning systems. Within this perspective, we revisit a core problem in multimodal learning: unsupervised cross-domain translation (CDT), which aims to learn mappings between two domains using only unpaired samples. CDT underlies many key applications in multimodal AI, including text-to-image generation, cross-platform medical image translation, and language model inference-time intervention.

Many CDT methods rely on distribution transport, a widely used paradigm for aligning data distributions across domains. However, such approaches often produce content-misaligned or hallucinated translations. We show that these failures stem from a fundamental identifiability issue caused by measure-preserving automorphisms (MPAs) in widely used CDT models. Building on this insight, we develop a framework that resolves the MPA ambiguity and establishes identifiability guarantees for distribution transport-based CDT. The resulting theory enables principled solutions to long-standing challenges across multiple scientific and engineering disciplines, including cross-platform optical coherence tomography (OCT) super-resolution and unregistered satellite spectral image fusion. These developments illustrate how signal processing perspectives can offer fundamental insights into modern generative AI systems and help bridge classical analytical principles with emerging AI paradigms.

Bio

Xiao Fu is an Associate Professor in the Department of Electrical and Computer Engineering at the University of Minnesota Twin Cities. Before joining the University of Minnesota, he was a faculty member in the School of Electrical Engineering and Computer Science at Oregon State University from 2017 to 2026. He received his Ph.D. in Electronic Engineering from The Chinese University of Hong Kong in 2014 and was a Postdoctoral Associate at the University of Minnesota from 2014 to 2017. His research lies at the intersection of machine learning and signal processing, with an emphasis on factor analysis, identifiability, optimization, multimodal learning, and generative AI. His work seeks to bring rigorous modeling and analytical principles to modern AI systems.

Dr. Fu received the IEEE Signal Processing Society Best Paper Award and Donald G. Fink Overview Paper Award in 2022. His other honors include the NSF CAREER Award, Oregon State University’s Engelbrecht Early Career Faculty Award and Promising Scholar Award, the ICASSP Best Student Paper Award, and the University of Minnesota Outstanding Postdoctoral Scholar Award.

He is a Senior Area Editor for IEEE Transactions on Signal Processing and served as a Senior Area Chair for NeurIPS 2026. He has also served as an Area Chair for NeurIPS, AAAI, and ICLR; Technical Program Co-Chair of IEEE SAM 2024; and Chair of the IEEE Signal Processing Society Oregon Chapter. He is an elected member of the IEEE SPS Signal Processing Theory and Methods and Sensor Array and Multichannel Technical Committees.

September 23, 2026
12:30 pm (1h)

Orchard View Room

Xiao Fu, University of Minnesota Twin Cities