Skip to content
LB2

Methods and limitations

How candidates are selected, what each number means, and where the approach can mislead. The code for every step is in the LB2 repository.

Overview

A tumor-derived RNA is only useful as a blood test if it rises above two backgrounds: the RNA the same organ makes when healthy, and the RNA shed by blood cells, which dominates plasma. For each cancer type, LB2 compares TCGA tumors with normal tissue from the same organ and with GTEx whole blood, removes genes that sorted blood immune cells express, and records whether published plasma cell-free RNA (cfRNA) data detect each remaining gene.

The output is a ranked list of candidates for measurement in plasma. Nothing here has been validated in patient blood by LB2 itself.

Data

SourceWhat LB2 uses
UCSC Xena Toil recompute1,3TCGA and GTEx RNA-seq processed with one pipeline (RSEM, GENCODE v23, 60,498 genes), as TcgaTargetGtex_rsem_gene_tpm, stored as log2(TPM + 0.001), with its phenotype table.
TCGA9,059 primary tumors in 30 cancer types, and tumor-adjacent normal tissue.
GTEx2,4Normal tissue matched by organ, and 337 whole-blood samples.
TCGA-CDR5Curated overall survival, progression-free interval and disease-specific survival.
GENCODE v237Gene biotypes for every gene, collapsed to protein-coding, lncRNA, pseudogene, small RNA or other.
Human Protein Atlas6Normalized expression (nTPM) in 18 sorted blood immune cell types; the total PBMC mixture is not used.
Plasma cfRNA
  • Established circulating markers: 16 tumor markers measured clinically in serum or plasma, as proteins (positive control)
  • exoRBase 3.0: protein-coding genes detected in at least 10% of 175 healthy blood cfRNA samples8
  • GSE174302: mRNA detected in at least 20% of 230 plasma cfRNA-seq samples from five cancer types and healthy donors9

Sample groups

Tumor: TCGA samples labelled primary tumor for each cancer type. Normal tissue: TCGA tumor-adjacent normal samples from the same project, plus GTEx normal samples from the matching organ. Where GTEx lacks the organ, a related tissue stands in (lung for mesothelioma, skeletal muscle for sarcoma, liver for cholangiocarcinoma, salivary gland for head and neck), and cervical cancer is currently compared with uterus; each cancer page lists its own comparator. Whole blood: the same 337 GTEx samples for every cancer type. GTEx’s EBV-transformed lymphocyte cell lines, which a simple “blood” filter would include, are excluded.

Comparisons with fewer than 10 samples in a group are flagged as underpowered16. Thymoma has only 2 normal samples and yields no candidates. Uveal melanoma and diffuse large B-cell lymphoma have no normal tissue in either study and were not analysed. Sample sizes for every cancer type.

Statistics

Each comparison (tumor vs normal tissue, tumor vs whole blood) tests every gene with a two-sided Wilcoxon rank-sum test, using the tie-corrected normal approximation. Rank tests keep false discoveries closer to their nominal rate than parametric methods at population-scale sample sizes11. p-values are adjusted across all genes with the Benjamini–Hochberg procedure12.

Effect size is the AUC, the Mann–Whitney U divided by the product of the group sizes, with DeLong 95% confidence intervals13,14. The log2 fold change is the difference in mean log2(TPM + 0.001). A gene passes a comparison when all three criteria hold:

CriterionThresholdWhy
q-value (FDR)< 0.05controls false discoveries across 60,498 genes
log2 fold change≥ 1at least twofold higher in the tumor
Lower 95% bound of the AUC≥ 0.70small or noisy groups fail instead of ranking high

A candidate must pass both comparisons. With thousands of samples almost every p-value is tiny, so candidates are ranked by discrimination rather than by significance.

Study effects

The Toil recompute removes differences in processing, but not those between surgical TCGA samples and post-mortem GTEx samples. When a cancer type has normal samples from both studies, the tissue comparison removes the study effect with a limma-style linear model that protects the tumor/normal difference15.

That correction is impossible when every normal sample comes from GTEx, and in the blood comparison, where every tumor is from TCGA and every blood sample from GTEx: study and biology are completely confounded. Those comparisons are reported as they are, and the immune-cell filter below provides an independent check.

Immune-cell filter

Genes expressed above 1 nTPM in any of the 18 sorted immune cell types of the Human Protein Atlas are removed, even when they pass the blood comparison. The filter only removes genes. Genes absent from the reference, mostly non-coding RNAs, cannot be checked and are kept; the candidate table marks them with a dash.

Plasma evidence and ranking

Each candidate is matched against three sources (by Ensembl ID or symbol, as each source provides). Two are direct detection in cell-free RNA: exoRBase 3.0 (blood cfRNA from 175 healthy people, genes detected in at least 10% of samples) and GSE174302 (230 plasma samples from patients with five cancer types and healthy donors, genes detected in at least 20%). The third is a short list of 16 tumor markers measured clinically in serum or plasma as proteins. Evidence is recorded but never removes a gene. The Vorperian et al. 2022 deconvolution reference was a fourth source until September 2026; it was dropped because membership in a reference matrix is not detection in plasma.

These sources are permissive and cover protein-coding genes only. exoRBase lists 18,917 genes and GSE174302 17,780, so of the 21,471 protein-coding candidates across the atlas, all but 381 have evidence. The remaining 26,892 candidates are non-coding genes (long non-coding RNAs, pseudogenes, small RNAs), which the sources cannot list at all; the site reports them as “cannot be checked” rather than “not detected”, using each gene’s GENCODE v23 biotype.

The priority score is the mean of the two AUCs plus 0.05 for each supporting source (a maximum bonus of 0.15), so discrimination still dominates the ranking.

Tissue panels

For each cancer type, an elastic-net logistic regression17 separates tumor from normal tissue using the expression of up to 300 top-ranked candidates. Penalty strength (C = 0.05 or 0.5) and mixing (L1 ratio 0.5 or 0.9) are tuned in an inner five-fold cross-validation, inside an outer five-fold cross-validation whose held-out predictions give the reported AUC. The final model is refit on all samples, and its non-zero coefficients form the panel.

Before training, the TCGA-versus-GTEx offset is removed from the expression matrix with the same limma-style model used for the tissue comparison, with the tumor/normal contrast protected. That is possible for the 23 cancer types whose normal samples come from both studies; for the 6 whose normals are all GTEx (LGG, OV, UCS, MESO, TGCT, ACC) the offset cannot be separated from the tumor/normal contrast, so their panels are trained on uncorrected data and may partly learn study differences. The correction is fitted once on all samples before cross-validation, so even corrected held-out AUCs are not fully independent of it.

Tumor and normal tissue are easy to tell apart, so held-out AUCs close to 1 are expected. They say nothing about how well the same genes would work in plasma.

Survival

For the 150 top-ranked candidates of each cancer type, a Cox proportional-hazards model relates tumor expression, scaled to one standard deviation, to a TCGA-CDR endpoint5. LB2 uses whichever of overall survival, progression-free interval and disease-specific survival has the most events, requiring at least 10. p-values are FDR-adjusted within each cancer type; 95% confidence intervals for the hazard ratio come from the Wald standard error. Expression stays continuous: searching for an optimal high/low cutpoint inflates false positives18.

NSCLC and colorectal cancer

TCGA profiles non-small cell lung cancer as two projects, lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC), and colorectal cancer as colon (COAD) and rectal (READ) adenocarcinoma. For each pair, a shared view lists the genes, matched by Ensembl ID, that are candidates in both separate analyses: 826 for NSCLC (1,364 LUAD and 2,156 LUSC candidates), and 939 for CRC (1,356 COAD and 1,355 READ candidates).

Nothing is pooled or re-tested. Each gene keeps the statistics of each project, and shared genes are ranked by the lower of their two priority scores. Plasma evidence and immune-cell expression belong to the gene, so they are the same in both projects. The overlap is conservative: a gene that is strong in one project but just misses a threshold in the other is left out.

Rarer NSCLC types such as large-cell carcinoma are not in TCGA, and TCGA’s colorectal samples are primary tumors, so neither view describes metastatic disease. Search accepts common clinical terms (NSCLC, CRC, HCC, ccRCC and others) and leads to the TCGA projects that hold the data; where a term and the data do not match, as for mCRC, TNBC or ESCC, search says so.

Limitations

  • Tissue is not plasma. High expression in a tumor does not mean the RNA is released, survives in circulation or reaches measurable levels, especially in early disease, when the tumor’s share of cfRNA is small10.
  • Bulk tissue mixes cell types. A candidate can come from the tumor microenvironment rather than cancer cells; collagen genes made by cancer-associated fibroblasts are examples.
  • The blood comparison crosses studies. Some differences reflect TCGA versus GTEx handling rather than biology. The immune-cell filter reduces, but cannot remove, this risk.
  • Normal tissue is imperfect. Adjacent tissue can carry tumor-related changes, and stand-in organs (lung for mesothelioma, muscle for sarcoma) only approximate the tissue of origin.
  • Small groups give unstable ranks. Cholangiocarcinoma has 36 tumors and bladder cancer 28 normal samples. When the two groups separate completely (AUC = 1), the DeLong interval has zero width and overstates certainty.
  • Plasma evidence is a low bar. The two cfRNA datasets list most protein-coding genes at the thresholds used, and GSE174302 includes plasma from cancer patients. Evidence here means the transcript is plausibly present in plasma, not that it is measurable at useful levels.
  • Non-coding candidates are unchecked, not cleared. More than half of all candidates are lncRNAs, pseudogenes or small RNAs. No plasma source used here lists them, so they carry no plasma evidence either way and no plasma bonus in the ranking.
  • Multiple testing is controlled within, not across, cancer types. A gene reported in many cancer types has been tested many times.
  • Annotation is from 2015. GENCODE v23 symbols predate many renamings, and the immune-cell check covers protein-coding genes only.
  • Survival results are associations in tissue, for the top 150 candidates, with an endpoint chosen by event count.
  • No independent validation yet. Candidates need confirmation in a separate plasma cfRNA cohort with cancer cases and matched controls.

Reproduce

The pipeline runs on Python 3.14. Downloading the Toil matrices takes about 2.6 GB. These commands rebuild every number on this site:

python -m lb2.data.download          # Toil, phenotype, TCGA-CDR, HPA, plasma sources
python -m lb2.pipeline.run prepare   # parse TPM matrix once
python -m lb2.pipeline.run run       # both comparisons, all cancer types
python scripts/annotate_results.py   # symbols, immune filter, plasma evidence
python scripts/run_panels.py         # nested-CV tissue panels
python scripts/run_survival.py       # Cox models
python scripts/export_web_data.py    # data files for this site

How to cite

Cite LB2 and the datasets it builds on (references 1–9 below, GENCODE included). A citation file is in the repository.

Sanchez H. LB2: dual-axis discovery of tumor-enriched cell-free RNA biomarker candidates. Version 0.1.0. 2026. github.com/zenitmaster/LB2

@software{lb2,
  author  = {Sanchez, Hector},
  title   = {{LB2}: dual-axis discovery of tumor-enriched cell-free {RNA} biomarker candidates},
  version = {0.1.0},
  year    = {2026},
  url     = {https://github.com/zenitmaster/LB2},
  license = {MIT}
}

References

  1. Vivian J, et al. Toil enables reproducible, open source, big biomedical data analyses. Nat Biotechnol. 2017;35:314–316. doi:10.1038/nbt.3772
  2. Wang Q, et al. Unifying cancer and normal RNA sequencing data from different sources. Sci Data. 2018;5:180061. doi:10.1038/sdata.2018.61
  3. Goldman MJ, et al. Visualizing and interpreting cancer genomics data via the Xena platform. Nat Biotechnol. 2020;38:675–678. doi:10.1038/s41587-020-0546-8
  4. GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science. 2020;369:1318–1330. doi:10.1126/science.aaz1776
  5. Liu J, et al. An integrated TCGA pan-cancer clinical data resource to drive high-quality survival outcome analytics. Cell. 2018;173:400–416.e11. doi:10.1016/j.cell.2018.02.052
  6. Uhlen M, et al. A genome-wide transcriptomic analysis of protein-coding genes in human blood cells. Science. 2019;366:eaax9198. doi:10.1126/science.aax9198
  7. Harrow J, et al. GENCODE: the reference human genome annotation for The ENCODE Project. Genome Res. 2012;22:1760–1774. doi:10.1101/gr.135350.111
  8. Yu W, et al. exoRBase 3.0: an updated resource for extracellular RNAs from human biofluids. Nucleic Acids Res. 2026;54:D135–D146. doi:10.1093/nar/gkaf1004
  9. Chen S, et al. Cancer type classification using plasma cell-free RNAs derived from human and microbes. eLife. 2022;11:e75181. doi:10.7554/eLife.75181
  10. Larson MH, et al. A comprehensive characterization of the cell-free transcriptome reveals tissue- and subtype-specific biomarkers for cancer detection. Nat Commun. 2021;12:2357. doi:10.1038/s41467-021-22444-1
  11. Li Y, et al. Exaggerated false positives by popular differential expression methods when analyzing human population samples. Genome Biol. 2022;23:79. doi:10.1186/s13059-022-02648-4
  12. Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc B. 1995;57:289–300. doi:10.1111/j.2517-6161.1995.tb02031.x
  13. DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics. 1988;44:837–845. doi:10.2307/2531595
  14. Sun X, Xu W. Fast implementation of DeLong’s algorithm for comparing the areas under correlated receiver operating characteristic curves. IEEE Signal Process Lett. 2014;21:1389–1393. doi:10.1109/LSP.2014.2337313
  15. Ritchie ME, et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43:e47. doi:10.1093/nar/gkv007
  16. Schurch NJ, et al. How many biological replicates are needed in an RNA-seq experiment and which differential expression tool should you use? RNA. 2016;22:839–851. doi:10.1261/rna.053959.115
  17. Zou H, Hastie T. Regularization and variable selection via the elastic net. J R Stat Soc B. 2005;67:301–320. doi:10.1111/j.1467-9868.2005.00503.x
  18. Altman DG, et al. Dangers of using “optimal” cutpoints in the evaluation of prognostic factors. J Natl Cancer Inst. 1994;86:829–835. doi:10.1093/jnci/86.11.829