CORTEXA
← Browse
arxivq-bio.QMcs.LG2026-06-30

Refnd: Preventing Data Leakage in Relational Datasets

Anthony Lavertu, Jacob Cote, Jacques Corbeil, Sophie Gobeil, Pascal Germain

Machine learning models trained on biochemical data are routinely evaluated using splits that fail to account for relational structure, causing information leakage and over-optimistic performance estimates. Existing splitting methods lack theoretical grounding and scale at best quadratically. We introduce the Relational Generative Process (RGP), a mathematical formalization explaining why relational structure arises in biochemical datasets, and Refnd, a splitting algorithm that leverages a proximity graph computed in loglinear time using Hierarchical Navigable Small World (HNSW). We validate on an antimicrobial peptide dataset, showing that Refnd splits yield lower but more realistic evaluation performance than traditional splits. Refnd is applicable to any dataset arising from an RGP such as protein sequences and structures, small molecules, and nucleotide sequences, and is openly available as a Rust accelerated Python package: pip install refnd.

View free PDFSource page

Related papers

arxivcs.LGcs.DB2026-07-31

Curriculum Matters: Data-Efficient Relational PFN Pretraining with Synthetic Data

Mohammad Sadeq Abolhasani, Viswanath Ganapathy

Relational Prior-Data Fitted Networks (PFNs) such as RDB-PFN approximate Bayesian inference over multi-table relational databases by pretraining on millions of synthetic tasks. We investigate three intertwined questions about this paradigm. First, can a structurally different syn…

View free PDFSource page
arxivcs.CVcs.AIcs.LG2026-07-31

Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

Paweł Borsukiewicz, Daniele Lunghi, Wendkûuni C. Ouédraogo, Jacques Klein, Tegawendé F. Bissyandé

Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk t…

View free PDFSource page
arxivcs.LGcs.AI2026-07-24

Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression

Hirotaka Sugawara, Yujin Taguchi, Kei Minagawa, Yusuke Hiki, Takashi Morikura, Akira Funahashi

Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables by balancing exploration and exploitation. BO is often used under a limited evaluation budget, such as hyperparameter tuning of…

View free PDFSource page
arxivphysics.opticscs.LGphysics.app-ph2026-07-23

Reliability-Aware Bayesian Optimization of 1310 nm PCSELs with FDTD Verification

Jinglin Yu, Feiyang Wu, Longying Wen, Chongxian Yuan, Renjie Li, Zhaoyu Zhang

Near 1310 nm photonic-crystal surface-emitting lasers (PCSELs) are attractive narrow-beam sources for optical communication and sensing, but their final design refinement is costly. Small geometry changes simultaneously shift the band-edge resonance, cavity leakage, far-field div…

View free PDFSource page