RESEARCH
Publications & Preprints
Explore our published research, preprints, inference models, and benchmarking results
that demonstrate clinical and translational relevance.
A Biologically Grounded Structural Causal Model Enables cfRNA Specific In-Context Learning
Publication
July 2026
Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typi
Authors: Ryan Kim, Beomsoo Kim, Hyunjin Kim, Sang Lee
Preprint · December 2025
Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typically trained on synthetic datasets that assume generic statistical properties, and as a result, they fail to capture the distinctive characteristics of cfRNA. These limitations become even more pronounced in settings where labeled data are scarce. We introduce cfRNA-ICL , a cfRNA-specific in-context learning model trained entirely on tasks generated from a biologically grounded structural causal model (SCM) . The SCM produces realistic cfRNA-like scenarios by incorporating empirical measurements of gene-level dropout, overdispersion, tissue-mixture–driven latent factors, compositional variability, and sequencing noise. This synthetic task universe enables cfRNA-ICL to acquire inductive biases that closely reflect the geometry of real cfRNA data. Across multiple cancer classification benchmarks, cfRNA-ICL demonstrates consistently higher performance than tabular ICL models trained on generic synthetic data. The gains are most sub-stantial in few-shot settings, where the model benefits from its exposure to cfRNA-specific statistical regimes during meta-training. Representation-level analyses further show that cfRNA-ICL organizes samples into biologically coherent manifolds, preserving cancer-type identity without the use of supervised constraints. This finding indicates strong alignment between the synthetic prior and real cfRNA structure. Taken together, these results show that domain-aware generative priors can meaningfully enhance in-context learning for biological tabular data. cfRNA-ICL provides a generalizable framework for cfRNA modeling and establishes a practical path toward foundation-scale models that are intrinsically adapted to the statistical landscape of plasma cfRNA.
View full publication →
cfRNA Codex: Integrated Cell-Free RNA Expression Atlas Across Diseases and States
Publication
July 2026
The cfRNA Codex is the largest curated aggregation of publicly available cell-free RNA (cfRNA) datasets to date, consolidating 127 cohorts and 14,791 samples from plasma and serum RNA-seq studies across oncology, maternal–fetal health, immunology, cardiometabolic diseases, and rare conditions. This
Authors: Ryan Kim, Beomsoo Kim, Hyunjin Kim, Sang Lee
Preprint · December 2025
The cfRNA Codex is the largest curated aggregation of publicly available cell-free RNA (cfRNA) datasets to date, consolidating 127 cohorts and 14,791 samples from plasma and serum RNA-seq studies across oncology, maternal–fetal health, immunology, cardiometabolic diseases, and rare conditions. This resource systematically collects raw sequencing reads, processed expression matrices, and harmonized metadata, while deliberately preserving each study's original format without cross-dataset normalization. As a result, the Codex enables rigorous within-cohort analyses, rapid replication of published findings, baseline machine-learning experiments on individual studies, and qualitative cross-cohort pattern discovery. By providing a unified, high-confidence catalog of true cfRNA datasets, the cfRNA Codex dramatically lowers the barrier to exploration of circulating transcriptomics and offers a strategic foundation for future harmonization, meta-analysis, and multi-cohort cfRNA modeling efforts. It serves as both an immediate research accelerator and a roadmap for next-generation liquid-biopsy analytics.
View full publication →
Beyond Annotation: Leveraging Raw RNA-seq Reads via Foundation models for Multi-Cancer Early Detection
Publication
July 2026
This study introduces an annotation-free, alignment-free diagnostic framework that learns directly from raw cell-free RNA (cfRNA) sequencing reads to detect early-stage cancer with high accuracy and interpretability. Instead of restricting analysis to annotated genes or relying on traditional alignm
Authors: Ryan Kim, Beomsoo Kim, Sang Lee
Preprint · December 2025
This study introduces an annotation-free, alignment-free diagnostic framework that learns directly from raw cell-free RNA (cfRNA) sequencing reads to detect early-stage cancer with high accuracy and interpretability. Instead of restricting analysis to annotated genes or relying on traditional alignment-based RNA-seq pipelines, we train a large-scale RNA foundation model (~2.5B parameters) on 10 billion 150 bp raw RNA reads to capture biological signals across the entire transcriptome—including intergenic regions, introns, transposable elements, orphan transcripts, and novel RNAs. Using publicly available cfRNA datasets (GSE174302 and GSE183635), we evaluate this approach across colorectal, liver, stomach, lung, and ovarian cancers (N=653). The model produces sample-level embeddings directly from raw FASTQ files through masked nucleotide modeling and contrastive learning. These embeddings are fine-tuned with a lightweight classification head for early-stage cancer detection. Across multiple cohorts, the model achieves high early-stage performance (AUROC up to 0.975) and demonstrates strong specificity (98.5%) and overall accuracy (93.2%). Importantly, attention-weight analysis shows that 30% of the most informative features arise from unannotated or poorly characterized regions, confirming the value of capturing the "dark transcriptome" beyond canonical gene annotations. This work establishes a scalable, explainable, and annotation-independent foundation model for liquid biopsy, offering a powerful alternative to gene-centric workflows and enabling sensitive detection of minimal disease burden in early cancer.
View full publication →