The cfRNA Codex: Building the World's Largest RNA Atlas
Machine learning models are only as good as the data they're trained on. Recognizing this, Eigen Bio undertook an ambitious effort: build the world's largest harmonized repository of cell-free RNA data.
The result is the cfRNA Codex — a curated collection of 253 million transcript-sample measurements across 15,207 plasma cfRNA samples from 127 independent cohorts.
What's in the Codex?
The Codex aggregates publicly available cfRNA datasets spanning oncology, maternal-fetal health, immunology, cardiometabolic diseases, and rare conditions. Each study is preserved in its original format, with raw sequencing reads, processed expression matrices, and harmonized metadata.
Why Harmonization Matters
By bringing together diverse datasets in a unified catalog, the Codex enables researchers to:
- Rapidly replicate published findings
- Run baseline machine learning experiments on individual studies
- Discover qualitative cross-cohort patterns
- Build a foundation for future multi-cohort cfRNA modeling
Open Science
The cfRNA Codex is freely available to the research community. We believe open access to high-quality data accelerates discovery — and the Codex is our contribution to that mission.