Most human genomics data is locked behind controlled access. To analyze a tumor genome from dbGaP or EGA, you apply to a data access committee, sign an agreement, wait weeks or months, and then work only in the places the agreement allows. These rules exist for good reasons. But they slow everyone down, and they apply even when the person who donated the data would have happily given it away.
Some people have given it away. A few patients, and a few institutions working with them, have released individual-level sequencing data that anyone can download and use for any purpose.
I think the field would benefit from a truly open database of voluntarily contributed genomics data with no usage restrictions. The idea is not new, and it is not impossible. It has already been done several times. What is missing is organization.
This post collects the examples I know about, so I can reason about what exists, and so you can see it too. Not everything here is truly open, but everything is related.
Would you like to share an example I missed? Let me know!
What counts as truly open? #
I am looking for data that meets four criteria:
- Individual-level data. Sequencing reads, or at least per-person variant calls, not summary statistics.
- Anyone can download it. No account, no application, no data access committee.
- No usage restrictions. A public domain dedication like CC0, or no conditions at all.
- The donor chose this. Explicit consent to public release, with the understanding that a genome identifies its owner.
Few datasets meet all four. Here is how the main examples compare:
| Dataset | Donors | Raw reads | Anyone can download | Usage restrictions |
|---|---|---|---|---|
| Sid Sijbrandij | 1 | Yes | Yes | None (CC0) |
| NIST HG008 and HG009 | 2 | Yes | Yes | None (public domain) |
| Texas Cancer Research Biobank | 7 | Yes | Yes, via SRA | Original terms barred resale |
| Personal Genome Project | Thousands | Mostly variant files | Yes | None (CC0) |
| Genome in a Bottle HG002 to HG007 | 6 | Yes | Yes | None |
| 1000 Genomes | 3,202 | Yes | Yes | None |
| Prostate cancer case | 1 | Yes | No, EGA | Data access agreement |
| Count Me In | Hundreds per project | No, dbGaP | Processed data only | Varies |
| Cancer Gene Trust | 18 | No, panels | Was yes, now offline | None stated |
| openSNP | Thousands | No, genotyping arrays | Deleted in 2025 | Was CC0 |
Truly open cancer genomes #
Sid Sijbrandij #
Sid Sijbrandij, the co-founder of GitLab, has osteosarcoma. He has released the data from his treatment in a public AWS bucket under CC0. As far as I know, no other cancer patient has released this much data: 37 TB in about 400,000 files, including:
- tumor and normal WGS and WES at several time points, plus an organoid
- PacBio long reads
- bulk RNA-seq, and single-cell RNA-seq of tumor and blood (Illumina and Oxford Nanopore)
- spatial data (Xenium, Visium HD, PhenoCycler, Orion)
- H&E and IHC images, and de-identified CT scans
- HLA typing, flow cytometry, MRD, and lab results
Anyone can list the bucket without an AWS account:
aws s3 ls --no-sign-request s3://sid-sijbrandij-osteosarc-dataset/
The dataset is listed on the AWS Registry of Open Data and managed by the Rare Cancer Research Foundation. Of the 48 registry entries tagged “cancer”, it is the only one donated by an individual patient.
NIST Cancer Genome in a Bottle: HG008 and HG009 #
https://www.nist.gov/programs-projects/cancer-genome-bottle
Genome in a Bottle is the NIST program that makes reference genomes for benchmarking variant calls. HG008 is its first tumor-normal pair. The donor was a 61-year-old woman with pancreatic ductal adenocarcinoma who gave explicit consent for public release of her genomic data. Her tumor was resected at Massachusetts General Hospital in 2020, and the Liss lab grew the HG008-T cell line from it. The matched normals are pancreatic and duodenal tissue.
Fourteen labs have sequenced HG008 with 17 technologies, including deep short- and long-read WGS, single-cell WGS, and Hi-C. HG009 is a second pair: a tumor cell line from a pancreatic cancer liver metastasis, with CD4+ T cell lines as the matched normal. NIST releases all datasets “publicly and without embargo as they are collected” on the GIAB FTP site, and its data policy puts them in the public domain.
The paper explains why this was needed:
While some tumor cell lines have been characterized as benchmarks by other groups, these data are from legacy cell lines with no consent or consent before whole genome sequencing was routine.
- McDaniel JH et al. (2025). Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair. Scientific Data 12:1195.
Texas Cancer Research Biobank open access pilot #
In 2016, the Texas Cancer Research Biobank released tumor and matched normal sequencing from 7 patients who consented to open access: 4 with pancreatic ductal adenocarcinoma, 2 with pancreatic neuroendocrine carcinoma, and 1 with follicular lymphoma. All 7 have whole-exome sequencing, and 2 also have whole-genome sequencing. The authors wrote:
To our knowledge, the TCRB OA dataset is the first individual-level, open-access genomic data release targeted to human cancers.
The release was open, but not unrestricted. The original portal required an account and a click-through agreement that prohibited re-identification and resale. That portal is gone, and its old domain now hosts an unrelated site. The reads are still downloadable from SRA and ENA under PRJNA285925.
- Becnel LB et al. (2016). An open access pilot freely sharing cancer genomic data from participants in Texas. Scientific Data 3:160010.
Truly open germline genomes #
None of these are cancer genomes. But they show that open consent works at scale, and that open genomes become shared infrastructure for the whole field.
The Personal Genome Project #
https://www.personalgenomes.org
George Church started the Personal Genome Project at Harvard in 2005. Participants pass an exam about the risks, then share their genomes, health records, and traits publicly and non-anonymously. Harvard releases the data under CC0. There are about 6,000 participant profiles, including a few hundred public whole genomes, and Harvard is still enrolling in 2026.
The project is built on open consent:
Open consent means that volunteers consent to unrestricted redisclosure of data originating from a confidential relationship, namely their health records, and to unrestricted disclosure of information that emerges from any future research on their genotype–phenotype data set… No promises of anonymity, privacy or confidentiality are made.
The PGP spread to other countries, with mixed results:
- PGP-UK (2013) released WGS, methylation, and RNA-seq data in ENA (PRJEB17529) with no access restrictions. Its last news post was in 2020.
- PGP-Canada (2012) is closed to new participants and is working to publish its data to a public repository.
- Genom Austria (2014) was suspended in 2018 for lack of funding. Its 20 pilot genomes are still online.
- PGP-China (2017) has not released data.
Citations:
- Lunshof JE, Chadwick R, Vorhaus DB, Church GM (2008). From genetic privacy to open consent. Nature Reviews Genetics 9:406–411.
- Chervova O et al. (2019). The Personal Genome Project-UK, an open access resource of human multi-omics data. Scientific Data 6:257.
Genome in a Bottle HG002 to HG007 #
https://www.nist.gov/programs-projects/genome-bottle
Some of the most widely used benchmark genomes come from PGP volunteers. NIST chose an Ashkenazi Jewish trio (HG002, HG003, HG004) and a Han Chinese trio (HG005, HG006, HG007) from the PGP because, in the words of the GIAB FAQ, “they are more broadly consented, including for commercial redistribution.” The original pilot genome (HG001, NA12878) did not have that consent.
1000 Genomes and the Human Pangenome Reference Consortium #
https://www.internationalgenome.org
Participants in the 1000 Genomes Project consented to open release, and the NIH cites it as its example of unrestricted-access data. The 2022 high-coverage resequencing of 3,202 samples describes itself as “the largest fully open resource of whole-genome sequencing (WGS) data consented for public distribution without access or use restrictions.” The Human Pangenome Reference Consortium builds on these samples and puts its data in the public domain.
- Byrska-Bishop M et al. (2022). High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185:3426–3440.
People who published their own genomes #
- Craig Venter (2007). The HuRef donor “gave full consent… to disclose publicly his genomic data in totality.” (Levy et al. 2007)
- James Watson (2008). Released with his APOE gene redacted at his request (Wheeler et al. 2008). A later study showed that his APOE status could be inferred from nearby markers (Nyholt et al. 2009).
- Genomes Unzipped (2010). Twelve members published their 23andMe data and waived all rights to it under CC0. https://genomesunzipped.org/data
- The Corpasome (2013). Manuel Corpas published his family’s genetic data openly on figshare. (Corpas 2013)
Open Humans #
Open Humans lets people upload their own data, including genome and exome VCFs, and decide who can use it. It is still running. Steven Keating (below) served on its board. I did not find any cancer genomics projects there.
Consented, but still behind a gate #
These projects involved patients who wanted their data to be used. But some or all of the data ended up behind access controls, or only processed results were released.
A man with advanced prostate cancer #
An 84-year-old man with SPOP-mutant prostate cancer worked with a team led by Andrew Armstrong at Duke and Mark Gerstein at Yale to sequence his tumor and germline genomes. The paper says:
Our subject agreed to release all of his genomic sequencing data for public use, which has been deposited into the European Genome-Phenome Archive
But EGA only accepts controlled-access data. His BAM files (EGAS00001004648) are behind a data access agreement, including a promise not to identify him. He wanted his data to be public, and the archive had no way to make it public.
- Armstrong AJ et al. (2021). Molecular medicine tumor board: whole-genome sequencing to inform on personalized medicine for a man with advanced prostate cancer. Prostate Cancer and Prostatic Diseases 24:786–793.
Count Me In #
https://www.broadinstitute.org/count-me-in
Count Me In, from the Broad Institute and Dana-Farber, let patients consent online and share their medical records and archived tumor tissue. Over 12,000 people enrolled in projects for metastatic breast cancer, angiosarcoma, prostate cancer, and other cancers. Mutation calls, copy number, expression, and patient-reported data are open on cBioPortal, but the raw sequencing is controlled access in dbGaP. Co-lead Corrie Painter is a scientist who became an advocate after her own angiosarcoma diagnosis in 2010. Most projects stopped enrolling in late 2025, when Count Me In joined Broad Clinical Labs.
- Painter CA et al. (2020). The Angiosarcoma Project: enabling genomic and clinical discoveries in a rare cancer through patient-partnered research. Nature Medicine 26:181–187.
Cancer Gene Trust #
Researchers at UCSC and UCSF built the Cancer Gene Trust to share somatic mutations, imaging, and clinical data from consented patients. A pilot shared data from 18 patients: panel sequencing results (Foundation Medicine and UCSF500), imaging, and health records, stored on IPFS with Ethereum for authentication. The site is now offline. In 2020, David Haussler told the National Academies:
The Cancer Gene Trust went precisely nowhere. […] But the idea of actually freely exchanging information about cancer genetics and cancer outcomes has never taken hold, and I think that’s a travesty.
- Glicksberg BS et al. (2020). Blockchain-authenticated sharing of genomic and clinical outcomes data of patients with cancer: a prospective cohort study. Journal of Medical Internet Research 22:e16810.
UCSC Treehouse Childhood Cancer Initiative #
https://treehousegenomics.ucsc.edu/public-data/
Treehouse compares pediatric tumor RNA-seq against large compendia to find treatment leads. Its public compendium mixes reprocessed TCGA, TARGET, and other public data with samples from Treehouse’s own clinical sites (IDs starting with “TH”). Anyone can download it, but it is gene expression (RSEM TPM), not reads.
Rare Cancer Research Foundation and the Pattern Data Commons #
https://rarecancer.org/data-commons
The Rare Cancer Research Foundation (Pattern.org) helps people with rare cancers donate tissue, records, and sequencing data, and it manages Sid’s dataset on AWS. It plans to make data in the Pattern Data Commons “publicly available on a quarterly schedule.” In 2026, it started a two-year pilot, funded by the Chan Zuckerberg Initiative’s Biohub, to generate single-cell datasets from 158 patients across 14 rare tumor types. I could not find a public portal or license yet. For now, access is by request.
Patients who opened their records #
These patients shared their medical data publicly. Most of it is imaging and clinical records, not sequencing reads.
Salvatore Iaconesi, La Cura #
https://opensourcecureforcancer.com
When artist Salvatore Iaconesi was diagnosed with a brain tumor in 2012, he converted his medical records and his CT and MRI scans into open formats. He published them under a CC BY 3.0 license and invited the public to respond with their own ideas for a cure. He ran the project with his partner, Oriana Persico. He died in 2022.
Steven Keating #
https://news.mit.edu/2019/celebrating-curious-mind-steven-keating-0722
Steven Keating, a PhD student at MIT, collected more than 200 GB of his own medical data after his glioma diagnosis in 2014. It included MRIs, pathology, bloodwork, video of his awake brain surgery, and 3D-printed tumor models. He posted it online under CC0. But he could not get the sequence of his own tumor. In a 2016 interview, he said:
I was not allowed, due to the data being generated on a non-CLIA certified machine.
He served on the board of Open Humans, and he died in 2019 at age 31.
Liz Salmi #
Liz Salmi has lived with a grade 2 astrocytoma since 2008. She works on patients’ access to their own records at OpenNotes (Beth Israel Deaconess Medical Center), and she appears with Keating in the 2016 documentary The Open Patient. She has donated tumor tissue for research-grade genomic sequencing and linked her records to All of Us (STAT, 2026), but I found no public release of her sequencing data.
Russ Read-Barrow #
Russ Read-Barrow has had stage 4 colorectal cancer since 2021. His site publishes his blood tests, CT reports, daily chemotherapy logs, and a summary of his tumor panel results, including CSV and JSON files. He credits Paul Conyngham (below) as his inspiration. So far, the genomics are transcribed panel results, not sequencing data, and the site’s terms reserve all rights. People like Russ show the demand: detailed public records, but no simple path to sequence a tumor and release the reads.
N-of-1 treatment stories #
These people used sequencing to guide treatment, but did not release the data.
Lukas Wartman #
When Lukas Wartman, a leukemia researcher at Washington University, relapsed with acute lymphoblastic leukemia, his colleagues sequenced his tumor genome, exome, and transcriptome. They found FLT3 overexpression, and he responded to sunitinib. His paper has no data deposition statement.
- Wartman LD (2015). A case of me: clinical cancer sequencing and the future of precision medicine. Cold Spring Harbor Molecular Case Studies 1:a000349.
Paul Conyngham and his dog Rosie #
Paul Conyngham had his dog Rosie’s mast cell tumor and blood sequenced at the UNSW Ramaciotti Centre for Genomics. He used ChatGPT and AlphaFold to help design a personalized mRNA vaccine, which the UNSW RNA Institute made (The Scientist). The case is widely cited alongside Sid’s. I found no public release of the sequencing data.
Beata Halassy #
Beata Halassy, a virologist at the University of Zagreb, treated her own recurrent breast cancer with oncolytic viruses (measles and vesicular stomatitis virus) and published the case. The paper includes histology and imaging, but no genomic data.
- Forčić D et al. (2024). An unconventional case study of neoadjuvant oncolytic virotherapy for recurrent breast cancer. Vaccines 12:958.
RareCure #
RareCure is an open-source pipeline that ranks treatments for rare solid tumors. The paper describes Sid and Paul Conyngham, then says:
These are existence proofs, but neither approach scales to the broader patient population.
It was tested on TCGA sarcoma data, not new patients. The code is on GitHub.
- Arogyasami D (2026). RareCure: an open-source artificial intelligence pipeline for context-adaptive treatment discovery in rare solid tumors. Cureus 18:e109744.
Patient communities and tools #
These groups help patients use their own data. Their members have data, and their tools need it.
- Cancer Patient Lab is a patient-led nonprofit founded in 2022 by Brad Power, Rick Stanton, and Brian McCloskey, for people with advanced prostate, brain, and pancreatic cancer. It runs weekly webinars and an online community.
- Cancer HackerLab is an accelerator for early-stage startups in cancer diagnostics and navigation, co-founded by Brad Power, who was diagnosed with lymphoma in 2018, and Ari Akerstein.
- Cancer Commons was founded in 2011 by Marty Tenenbaum, a metastatic melanoma survivor. Its site-less N-of-1 study (NCT07343024) uses genomic and drug-sensitivity tests to match people with advanced cancer to FDA-approved drugs. It does not plan to share individual participant data.
- OpenCancer.ai, from Ari Akerstein and Brad Power, is an AI assistant that reads a patient’s records, explains them in plain language, and matches trials, with review by scientists.
- AI Tumor Board is a research and education demo by oncologist Roupen Odabashian. Up to nine specialist AI agents research a case in PubMed, ClinicalTrials.gov, CIViC, and other sources, then debate a recommendation.
Rules and tools that already allow this #
- NIH Genomic Data Sharing Policy. The policy “permits unrestricted access to de-identified data, but only if participants have explicitly consented to sharing their data through unrestricted-access mechanisms.” A draft revision from December 2025 would keep this path for data with “informed consent explicitly stating data are to be shared openly without controls,” after an institutional risk review.
- NCBI SRA. The submission guide says: “Before depositing human data into the public SRA database make sure that you have consent from the donating individual to make this data available in an unprotected database.” EGA only accepts controlled-access data, but ENA hosts open human data like PGP-UK and COLO829.
- GA4GH Data Use Ontology. The code DUO:0000004, “no restriction”, means: “This data use permission indicates there is no restriction on use.”
- Somatic Tumor Twins. A method that removes germline variants from tumor-normal reads. On 47 PCAWG pilot samples, it removed all detectable germline variants and kept a median of 98% of validated somatic SNVs. This could give patients a middle option between releasing everything and releasing nothing. (Gaitán et al. 2026)
- Portable Legal Consent. John Wilbanks started “Consent to Research” in 2011 and continued it at Sage Bionetworks as Portable Legal Consent: people could donate data under one reusable consent. It was not unrestricted, since data went only to qualified researchers under contract, and enrollment has closed. Its ideas carried into Apple’s ResearchKit.
- EU data altruism. The Data Governance Act (2022) created a legal framework for people to donate data for general-interest purposes, like health research. It is restricted by design: recognized organizations may not use the data for other purposes.
Cautionary tales #
HeLa #
In March 2013, EMBL researchers published the genome of the HeLa cell line. Henrietta Lacks never consented to the use of her cells, and her family had not been asked. The data were withdrawn. In August 2013, the NIH and the Lacks family agreed to put HeLa genome data under controlled access in dbGaP, with family members on the committee that reviews requests.
Legacy cancer cell lines #
Common somatic benchmarks, like COLO829 and the SEQC2 breast cancer cell line HCC1395, are openly available. But as the HG008 paper notes, they come from “legacy cell lines with no consent or consent before whole genome sequencing was routine.” HG008 exists to fix this.
Re-identification #
Homer et al. (2008) showed that you can tell whether a person is in a study from its summary statistics, and the NIH moved aggregate GWAS data behind controlled access in response (it reversed this in 2018). Gymrek et al. (2013) inferred the surnames of anonymous research participants from their Y chromosomes and public genealogy databases. De-identification is not a promise anyone can keep. Open consent says so up front.
- Homer N et al. (2008). Resolving individuals contributing trace amounts of DNA to highly complex mixtures using high-density SNP genotyping microarrays. PLoS Genetics 4:e1000167.
- Gymrek M et al. (2013). Identifying personal genomes by surname inference. Science 339:321–324.
Insurance and relatives #
In the US, the Genetic Information Nondiscrimination Act (GINA, 2008) bans genetic discrimination in health insurance and employment, but not in life, disability, or long-term care insurance. Florida extended protection to life and long-term care insurance in 2020, but only “in the absence of a diagnosis”, which may not help someone who already has cancer. A germline genome also reveals information about the donor’s relatives, who did not consent.
openSNP #
openSNP let people publish their consumer genotyping data under CC0 from 2011 until 2025. After 23andMe went bankrupt, the founders shut it down and deleted all of the data on April 30, 2025. Co-founder Bastian Greshake Tzovaras wrote:
The risk/benefit calculus of providing free & open access to individual genetic data in 2025 is very different compared to 14 years ago.
A 2017 data dump survives at the Internet Archive.
What is missing #
Looking at this list, a few things stand out to me:
- There is no shared home. Sid’s data is in its own AWS bucket, HG008 is on NIST’s FTP site, and the Texas reads are in SRA. The PGP-UK authors wrote that “there is currently no single public repository for open access multi-omics data.”
- Open data disappears. The Texas portal, the Cancer Gene Trust, Steven Keating’s website, and openSNP are all gone.
- Consent is not enough. The prostate cancer patient agreed to public release, and his data still ended up behind a data access agreement.
- Patients have data, but no path to release it. People like Russ Read-Barrow and Liz Salmi keep detailed records, and some have donated tissue for sequencing. But there is no simple path from “I want my tumor genome to be public” to reads in a public archive.
- The pieces already exist. Open consent, the NIH policy, public SRA deposits, the DUO “no restriction” code, and NIST’s published consent language for HG008 are all in place. They need to be put together.