CORTEXA
← Browse
arxivcs.LG2026-07-13

Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers

Gurdeep Singh Virdee

Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-registered, condition-stratified robustness analysis comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1--C4). Four hypothesis groups are evaluated: discrimination deltas with Holm-corrected multiplicity control (H1), Brier score differences (H2), calibration slope outcomes (H3), and AUROC differences under best-condition setups (H4). TEMP-minus-ISO discrimination deltas remain small across all conditions (-0.0155 to 0.0139), with Holm-adjusted p-values of 0.9895 everywhere. TEMP Brier differences are consistently negative (C1: -0.0002 through C4: -0.0074), while ISO shows sign reversals. TEMP calibration slopes stay closer to unity in every condition (range 0.7597--0.9493) than ISO slopes (0.1364--0.2726). AUROC differences shift from near zero in C1 (-0.0004) to positive in C4 (0.0264). These results establish that in-dataset robustness is condition-dependent and metric-specific. No claim of external transportability is made.

View free PDFSource page

Related papers

arxivcs.CVcs.LG2026-07-02

Rethinking Post-Hoc Calibration in Semantic Segmentation

Tristan Kirscher, Kim-Celine Kahl, Balint Kovacs, Maximilian R. Rokuss, Klaus Maier-Hein, Xavier Coubez, et al.

Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead downstream decisions. Yet modern segmentation models often remain miscalibrated. Post-hoc calibration offers a practical way to corr…

View free PDFSource page
arxivcs.LGcs.AI2026-07-14

Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty

Sarah Al-Shareeda, Gulcihan Ozdemir, Heung Seok Jeon

Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs. Missing features must then be reconstructed, and their errors can propagate through the model. If this input unce…

View free PDFSource page
arxivcs.CRcs.LG2026-06-30

Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed

Zhichao Fan, Zexin Zhuang, Yanhang Li

We audit a fixed prefix-window mean-NLL memorization probe (K=20) on a Qwen2.5-VL-7B canary testbed and report three post-hoc cases where it disagrees with full-span secret NLL or greedy exact-recall. C3 (false negative, window truncation): damage lands on hex tokens outside K=20…

View free PDFSource page
arxivcs.LG2026-07-01

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, Stefan Heimersheim

Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural tran…

View free PDFSource page