Pathologists at Severance Hospital in Seoul reviewed 12,104 clinical sequencing datasets from 1,454 batches run over four years and found that gene-level fold changes, the quantity behind every tumor-only copy-number call, drift with the batch that produced them. The drift can carry a gene across the line into a reported copy gain or deletion. Their remedy computes a median and median absolute deviation of fold change inside each batch and reads individual results against those figures.
- Data: 12,104 clinical tumor-only sequencing datasets from 1,454 batches, roughly eight samples per batch, over four years.
- Lab: Department of Pathology, Severance Hospital, Yonsei University College of Medicine, Seoul.
- Method examined: panel of normals comparison for copy number, the common workaround where matched normal samples are unavailable.
- Suspected driver: experimental variation the panel of normals does not absorb, including probe efficiency differences across reagent lots.
- Correction: per-batch median and median absolute deviation of gene-level fold change, carried into interpretation as reference metrics.
- Published: The Journal of Molecular Diagnostics, online ahead of print August 21, 2026.
Four Years of Batches, Read Back
Clinical laboratories running tumor-only panels rarely have a matched normal from the same patient, so copy number gets called by comparing the case against a fixed panel of normals assembled once and reused. Chung Lee, Sejoon Lee and colleagues started from a stated hypothesis: that a panel of normals and its caller, once established, may not fully compensate for all experimental variation, with differences in probe efficiency across reagent lots given as the example. Testing it meant going back over every run the laboratory had done, 12,104 datasets across 1,454 batches, an average near eight samples per batch and roughly 30 batches a month across the four-year span.
The Fluctuation Sat at the Batch Level
Gene-level fold changes fluctuated in patterns tied to the sequencing batch. A gene's fold change in a given sample therefore reflects both the tumor and the run, and the authors point to the interpretive consequence: incorrect classification of gene copy deletions or gains. Copy-number calls in comprehensive genomic profiling drive gene dosage findings that reach therapy selection.
The abstract names no panel, no CNV caller and no specific genes, so which loci proved most exposed is not visible from outside the paywall. Only subscribers can open the article, and no PubMed Central deposit exists. That left the published abstract and the affiliation data in PubMed and Crossref.
A Median and a Deviation for Every Run
The correction takes all samples in a sequencing batch, computes the median and the median absolute deviation of gene-level fold change across them, and carries both into result interpretation as batch-level reference metrics. A gene that moves in one sample against a stable batch background reads differently from a gene that moves in concert with everything else on the run. The authors report that putative batch-driven artifacts can be identified this way and that false-positive CNV calls fall, without a quantified effect size in the abstract.
Why This Matters to the APO|APE Reader
Corresponding authors Inho Park and Hyo Sup Shim place the correction inside interpretation rather than inside the caller, which is where a laboratory can adopt it without revalidating a pipeline or refiling a validation package. Batch-level medians only exist once batches do, and this group needed 1,454 runs before the pattern was legible, so a panel going live has no history to compare its first cases against. Reagent lots also turn over on the manufacturer's schedule, not the laboratory's, which argues for a running metric recalculated as each lot enters service rather than a threshold fixed at validation. Severance is a Korean academic hospital laboratory, and the paper offers no CLIA or CAP framing, so translating any of this into a US validation protocol is work still to be done.


