CORTEXA
← Browse
arxiveess.IVeess.SP2026-07-03

AirTF: Over-the-Air Token Fusion for Task-Oriented Multi-Modal Token Communications

Bole Liu, Li Qiao, Minghui Wu, Yulin Shao, Zhen Gao

In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriented multi-modal token communications. Unlike existing schemes for segmentation that rely on convolutional neural networks (CNNs) with limited local receptive fields, AirTF leverages vision transformer (ViT) encoders to extract globally contextualized semantic tokens from distributed heterogeneous sensors. By concurrently transmitting these spatially aligned tokens over a shared wireless channel, our framework exploits the superposition property of the multiple access channel to inherently fuse complementary multi-modal semantics (e.g., RGB and infrared) directly over the air. This mechanism significantly enhances spectral efficiency compared to orthogonal transmission. Furthermore, the integration of a pre-trained foundation model provides critical visual priors, effectively addressing the data-hungry nature of ViTs on limited, scenario-specific semantic segmentation datasets. Experiments demonstrate that AirTF consistently outperforms orthogonal transmission and CNN-based fusion baselines across AWGN and fading channels. Additional evaluations under a three-user setting, residual synchronization errors, and imperfect channel state information estimation further confirm its robustness. The source code will be made publicly available upon acceptance.

View free PDFSource page

Related papers

arxiveess.SYeess.IVeess.SP2026-07-13

WULPUS PRO: Multi-mode Ultra-Low-Power Wearable Ultrasound and Array Imaging with CMUT Support

Sergei Vostrikov, Federico Villani, Cedric Hirschi, Jinhao Lu, Jonas Welsch, Martin Angerer, et al.

Wearable ultrasound enables continuous monitoring of physiological processes such as muscle dynamics, bladder volume, and cardiovascular activity. Existing fully wearable ultra-low-power platforms are limited to shallow, low-channel A-mode sensing, while larger multi-mode systems…

View free PDFSource page
arxiveess.IVcs.CVeess.SP2026-07-15

Video to All-in-focus Image Reconstruction Algorithm for Automated Microscopic Urinalysis

Chinmay Nema, Hari Om Aggrawal, Dipam Goswami, Rajiv Gupta, Vinti Agarwal

Microscopic urinalysis is a routine diagnostic test at hospitals. Recent studies have demonstrated the effectiveness of deep learning methods to automate microscopic urinalysis. These methods rely on high-quality images of the urine samples in which each cell is clearly identifia…

View free PDFSource page
arxiveess.SPcs.CVcs.LGeess.IV2026-07-15

ECG-LLM: Foundation Model for ECG-Based Cardiac Reasoning

Alexander Selivanov, Friederike Jungmann, Jan Kehrer, Karl-Ludwig Laugwitz, Eimo Martens, Daniel Rueckert

Electrocardiography (ECG) is an inexpensive, standard-of-care test for cardiac symptoms, but front-line triage often lacks immediate access to definitive imaging such as echocardiography (ECHO) or cardiac magnetic resonance (CMR). Furthermore, most existing ECGAI systems are limi…

View free PDFSource page
arxivcs.GRcs.ETcs.MMeess.IVeess.SP2026-07-22

Fast Wave-optics Rendering of Multiplane Images for 3D Holographic Displays

Brian Chao, Dario Seyb, Nathan Matsuda, Oliver Cossairt, Yang Zhou, Douglas Lanman, et al.

Recent advances in neural rendering have unlocked unprecedented capabilities in 3D reconstruction and novel view synthesis, giving rise to applications such as virtual fly-throughs of a 3D scene reconstructed from a set of sparse, casually captured images. However, these renderin…

View free PDFSource page
arxivcs.ETeess.IVeess.SPphysics.med-ph2026-07-23

Quantum Adaptive Sensing for Accelerated MRI

Asmit Ganguly, Suprajit Dewanji, Chenyang Zhao, Danny J. J. Wang

Compressed sensing accelerates MRI by reconstructing images from undersampled k-space, but performance depends strongly on sampling distribution. We propose an adaptive framework that selects Cartesian phase-encode lines sequentially using a fixed-cardinality quadratic unconstrai…

View free PDFSource page