Skip to main navigation Skip to search Skip to main content

Calibration granularity, not contamination: diagnosing a TCN anomaly detector’s false positive advantage in cross-dataset IoT traffic

    Research output: Contribution to journalArticlepeer-review

    4 Downloads (Pure)

    Abstract

    We set out to fix a “contamination” problem in reconstruction-based Temporal Convolutional Network VAEs (TCN-VAEs) for cross-dataset IoT flow anomaly detection: when attack flows share an encoder window with benign flows, the shared latent code is allegedly distorted, inflating benign reconstruction error and producing false positive rates (FPRs) of 22–65% despite an ROC-AUC above 0.93. Our proposed fix, TCN-Pred, excludes the target flow from the encoder and scores it by next-flow prediction error, reducing FPR to 0.65–13%. We subjected this causal explanation to a battery of controlled ablations, holding architecture, decoder, loss, and thresholding fixed while varying one factor at a time. Each one falsified the original hypothesis: target inclusion/masking changes FPR by at most 0.001; context shuffling/reversing/zeroing changes it by at most 0.003; a context-blind constant-output predictor matches TCN-Pred’s FPR and F1 to three decimal places on all three datasets. The actual cause, confirmed on the original trained models with no retraining, is a scoring-granularity mismatch: the TCN-VAE threshold is calibrated from per-window errors averaged over 20 flows but applied to per-flow errors at evaluation (standard deviation ≈20−−√× higher, measured ratio 4.46 against a predicted 4.47). Recalibrating the identical model at matching granularity drops FPR from 22.7/47.6/64.6% to 0.65/5.0/12.5% on BoT-IoT, IoT-23 and ToN-IoT, closing 89–97% of the reported FPR gap without changing a single model weight. We report this diagnostic chain, together with an attack-prevalence sensitivity analysis, sample-disjoint calibration, normality diagnostics, and label-free and redundancy-aware (mRMR) feature-selection benchmarks, as a methodology other work should apply before attributing fixed-threshold performance to architecture. The pipeline is supervised source-domain feature selection followed by benign-only detector training, not fully unsupervised, a distinction we quantify later in the paper. Investigating dataset representativeness, we found that all three provided files reduce to only ≈6000 genuinely distinct flows via an undocumented row-duplication procedure, causing 97.8% BoT-IoT train/test near-duplicate overlap; a leakage-free re-evaluation changes FPR by only 0.23 percentage points. We also found that the TLS-metadata columns are already transformed upstream of every available artefact, so the proportion of genuinely TLS-encrypted flows cannot be recovered, and we soften the paper’s encrypted-traffic framing accordingly.
    Original languageEnglish
    Article number447
    Number of pages45
    JournalFuture Internet
    Volume18
    Issue number9
    DOIs
    Publication statusPublished - 24 Aug 2026

    Keywords

    • IoT network traffic
    • anomaly detection
    • temporal convolutional network
    • threshold calibration
    • ablation study
    • posterior collapse
    • variational autoencoder
    • corss-dataset generalisation
    • network intrusion detection

    Fingerprint

    Dive into the research topics of 'Calibration granularity, not contamination: diagnosing a TCN anomaly detector’s false positive advantage in cross-dataset IoT traffic'. Together they form a unique fingerprint.

    Cite this