CORTEXA
← Browse
arxiveess.SPcs.AI2026-07-07

Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction

Adrien Schneider, Kacper Zabkowski, Anderson Augusma, Frédérique Letué, Maria Camila Pinzon, Dominique Vaufreydaz

The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization and a HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddings of the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.

View free PDFSource page

Related papers

arxivcs.SDcs.AIeess.SP2026-06-29

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

Pranav Tushar, Xiao Xiao Miao, Rong Tong

Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are developed on adult speech, leading to degraded performance when applied to child speech. This paper investigates child-centric an…

View free PDFSource page
arxivq-bio.NCcs.AIcs.LGeess.SPmath.AT2026-07-10

PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis

Ren Takahashi, Emre Yusuf, Jayabrata Bhaduri

Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achieving a state-of-the-art area under the receiver operating characteristic curve (AUC) of approximately 0.70 on the DREAM database (Wong et al., 2…

View free PDFSource page
arxivcs.CVcs.AIeess.SP2026-07-10

Towards Objective Dysgraphia Detection: A Multi-Branch Deep Learning Approach for Online Handwriting Analysis

Lydia Ouhib, Yassine Ouzar, Zoé Pinseel, Stéphane Bouilland, Mehdi Ammi

Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluency, and legibility, often hindering academic achievement and early learning development. This motor coordination disorder is typically diagnos…

View free PDFSource page
arxivphysics.med-phcs.AIcs.CVeess.IVeess.SP2026-07-03

Harmonic-Aware Transformer for Real-Time Catheter Localization in Interventional Procedures of Magnetic Particle Imaging

Abuobaida M. Khair, Wenjing Jiang, Xiaoli Yang, Moritz Wildgruber, Xiaopeng Ma

Magnetic particle imaging (MPI) enables real-time, radiation-free tracking of magnetic nanoparticle-coated instruments, making it highly suitable for interventional procedures. This study proposes a harmonic-aware transformer framework that directly predicts catheter tip position…

View free PDFSource page
arxiveess.SPcs.AIphysics.ao-ph2026-07-07

Physics-Informed Feature Engineering 1D-CNN for Multilayer Cloud Detection from Geostationary Satellites

Fu Wang, Chi Yang, Qi-Feng Lu, Rui-Xia Liu, Xiao-Fei Yang, Xiao-Fang Liu, et al.

Multilayer cloud detection from active--passive observation is vital for numerical weather prediction. In this study, channel selections derived from threshold-based algorithms are embedded as feature-engineering priors into a 1D-CNN, and machine learning (ML) is used to learn la…

View free PDFSource page
arxiveess.AScs.AIeess.SP2026-07-04

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

Héctor Martel, Joe Hennessy-Priest, Taemin Cho

Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of th…

View free PDFSource page