Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Whole-exome Sequencing of Intracranial Germ Cell Tumors performed at Human Genome Sequencing Center, Baylor College of Medicine
Intracranial germ cell tumors (IGCTs) are rare and biologically diverse tumors affecting mainly male adolescents with the highest incidence in Japan and other Asian countries. They are divided into two main groups, pure germinoma and nongerminomatous germ cell tumors (NGGCTs). Germinoma is the most common subtype. NGGCTs include teratoma, embryonal carcinoma, yolk sac tumor and choriocarcinoma. About 10% of germinomas and most NGGCTs remain refractory to multimodality therapy. Little is currently known about IGCTs except for KIT mutation or overexpression, observed in ~25% of pure germinomas and rarely seen in NGGCTs. As yet, there are no clues for the puzzle of onset during puberty, geographic and gender discrepancy in the incidence of IGCTs. With the collaboration of Texas Children's Hospital, Saitama Medical University Hospital, Kumamoto University Hospital, Nagoya University Hospital, Hokkaido University Hospital and Chinese University of Hong Kong, the Human Genome Sequencing Center at Baylor College of Medicine had access to 62 tumor specimens and 52 matched normal blood samples from 68 IGCT patients. We performed whole-exome sequencing, targeted deep sequencing and high-resolution SNP arrays to characterize the profile of somatic mutations, germline variants and DNA copy number alterations. The deposited BAM files record the sequence alignments used to generate the mutation data.
Genome-wide hepatic miRNAome profiling in chronic hepatitis C patients using next-generation sequencing [miRNA-seq]
GEO Series GSE87843. Homo sapiens. 10 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Whole genome bisulfite sequencing from primate brains (orthologous region to human BA46)
GEO Series GSE151768. Pan troglodytes; Macaca mulatta. 50 samples. Type: Methylation profiling by high throughput sequencing.
Whole genome bisulfite sequencing of arp6-5 and pie1-7 Arabidopsis mutants
GEO Series GSE122373. Arabidopsis thaliana. 2 samples. Type: Methylation profiling by high throughput sequencing.
RNA deep sequencing to compare genome-wide differences between PRMT5/knockdown and control AML cells
GEO Series GSE75215. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Genomic sequences and annotations for Solanum lycopersicum, Solanum pennellii and Solanum habrochaites
<p><strong>=== Genome sequences === </strong></p> <p> </p> <p>These are the different genome references (fasta formats) available for:</p> <ul> <li><em>Solanum lycopersicum</em>: <ul> <li><a href="https://zenodo.org/api/files/c5778399-5188-4959-93b7-1738584c0f72/S_lycopersicum_chromosomes.2.50.fa.gz">S_lycopersicum_chromosomes.2.50.fa.gz</a></li> <li><a href="https://zenodo.org/api/files/94d37cf6-6e8b-47c8-96c7-b82ae363ccd7/S_lycopersicum_chromosomes.3.00.fa.tar.gz?versionId=04729276-a75e-4652-8cfe-07359b13bd8c">S_lycopersicum_chromosomes.3.00.fa.tar.gz</a></li> <li><a href="https://zenodo.org/api/files/57840fa9-db90-4795-af03-8dbd2f711b69/S_lycopersicum_chromosomes.4.00.fa.tar.gz?versionId=a006a20e-5a34-47c9-8f81-feb782f9e68d">S_lycopersicum_chromosomes.4.00.fa.tar.gz</a></li> </ul> </li> <li><em>Solanum pennellii </em>(one version only from Bolger et al., 2014) : <ul> <li><a href="https://zenodo.org/api/files/94d37cf6-6e8b-47c8-96c7-b82ae363ccd7/Spenn.fasta.tar.gz?versionId=18b34053-cb37-4e47-9025-5213b0455347">Spenn.fasta.tar.gz</a></li> </ul> </li> <li><em>Solanum habrochaites</em> LA1777 (technology hotel project 2018): <ul> <li><a href="https://zenodo.org/api/files/94d37cf6-6e8b-47c8-96c7-b82ae363ccd7/LA1777.final.fasta">LA1777.final.fasta</a> </li> </ul> </li> <li><em>Solanum habrochaites</em> PI127826: <ul> <li>2018 Hotel Project: <a href="https://zenodo.org/api/files/aa27d9ec-6ab7-4582-a17c-be7dfb0952a4/PI127826.final.fasta?versionId=d9748404-5d79-4a7b-af63-8dfd792436a0">PI127826.final.fasta </a></li> <li>2021 Dovetails assembly: <a href="https://zenodo.org/api/files/2b9490ff-031c-49d7-85e3-c57ba37ff21a/PI127826_hirise_assembly.fasta.gz">PI127826_hirise_assembly.fasta.gz</a><a href="https://zenodo.org/api/files/aa27d9ec-6ab7-4582-a17c-be7dfb0952a4/PI127826.final.fasta?versionId=d9748404-5d79-4a7b-af63-8dfd792436a0"> </a></li> </ul> </li> <li><em>Solanum</em> <em>habrochaites</em> LYC4 (from the paper of <a href="https://pubmed.ncbi.nlm.nih.gov/25039268/">Aflitos et al. 2014</a>. 3rd assembly version): <ul> <li><a href="https://zenodo.org/api/files/77a91023-8321-4f8e-86a6-1ae3b8197edf/S_habrochaites_LYC4_genome_assembly_v3.0_scaffold_scarpa.fasta?versionId=71501782-d9a5-49c8-83ed-511ba7deb6fa">S_habrochaites_LYC4...</a></li> </ul> </li> <li><em>Solanum arcanum</em> LA2172 (from the paper of <a href="https://pubmed.ncbi.nlm.nih.gov/25039268/">Aflitos et al. 2014</a>. 3rd assembly version): <ul> <li><a href="https://zenodo.org/api/files/179d217c-2a83-47ca-8180-3e4821b2481e/LA2172.fasta.tar.gz">LA2172.fasta.tar.gz</a></li> </ul> </li> <li><em>Solanum chilense</em> LA3111 (from the paper of <a href="https://www.g3journal.org/content/9/12/3933">Stam et al. 2019</a>, NCBI assembly ASM601370v1): <ul> <li><a href="https://zenodo.org/api/files/179d217c-2a83-47ca-8180-3e4821b2481e/LA3111.fasta.tar.gz">LA3111.fasta.tar.gz</a></li> </ul> </li> <li><em>Solanum lycopersicoides</em> LA2951 (from the work of The Boyce Thompson Institute and RWTH Aachen University: <a href="https://solgenomics.net/organism/Solanum_lycopersicoides/genome">link</a>): <ul> <li><a href="https://zenodo.org/api/files/77a91023-8321-4f8e-86a6-1ae3b8197edf/S_lycopersicoides_LA2951_v1.0_chromosomes_contigs.fasta.tar.gz">S_lycopersicoides_LA2951_v1.0_chromosomes.fasta.tar.gz</a></li> </ul> </li> </ul> <p>The two genome assemblies of S. habrochaites LA1777 and PI127826 were obtained through a combination of 10X Linked-Reads and BioNano Optical Mapping. This sequencing has been funded by the DTL Technology Hotel 2018 funding scheme.</p> <p> </p> <p><strong>=== Transcriptomes and proteomes ===</strong></p> <ul> <li><em><strong>Solanum lycopersicum</strong> (assembly</em> 4.0): <ul> <li>Transcriptome: <a href="https://zenodo.org/api/files/2a1ca78d-e799-4034-9abd-b760e4ea3694/ITAG4.0_cDNA.fasta?versionId=76df390e-af94-4ebb-b6a7-05cf2dca5010">ITAG4.0_cDNA.fasta</a> </li> <li>Proteome: <a href="https://zenodo.org/api/files/2a1ca78d-e799-4034-9abd-b760e4ea3694/ITAG4.0_proteins.fasta?versionId=cdc2eea1-f5be-4153-92d8-18a4cd77931c">ITAG4.0_proteins.fasta</a></li> </ul> </li> <li><em><strong>Solanum pennellii</strong> </em>(one version only from Bolger et al., 2014): <ul> <li>Transcriptome: <a href="https://zenodo.org/record/3885088/files/Spenn-v2-cds-annot.fa?download=1">Spenn-v2-cds-annot.fa</a></li> <li>Proteome: <a href="https://zenodo.org/api/files/2a1ca78d-e799-4034-9abd-b760e4ea3694/Spenn-v2-aa-annot.fa">Spenn-v2-aa-annot.fa</a></li> </ul> </li> <li><strong><em>Solanum lycopersicoides</em></strong> (version 1.0) <ul> <li>Transcriptome: <a href="https://zenodo.org/api/files/77a91023-8321-4f8e-86a6-1ae3b8197edf/S_lycopersicoides_LA2951_v1.0_cds.fasta">S_lycopersicoides_LA2951_v1.0_cds.fasta </a></li> <li>Proteome: <a href="https://zenodo.org/api/files/77a91023-8321-4f8e-86a6-1ae3b8197edf/S_lycopersicoides_LA2951_v1.0_proteins.fasta">S_lycopersicoides_LA2951_v1.0_proteins.fasta </a></li> </ul> </li> <li><strong><em>Solanum habrochaites </em>PI127826 </strong> <ul> <li><strong>Transcriptome: </strong><a href="https://zenodo.org/api/files/7ba7db3d-f8dd-43f9-b990-c1a5d3487b8f/Solanum_habrochaites_PI127826_mRNAs.fasta">Solanum_habrochaites_PI12826_mRNAs.fasta</a> (2018 Hotel Project assembly)</li> <li><strong>Transcriptome (2021 Dovetails): </strong><a href="https://zenodo.org/api/files/13145d97-7394-4b89-83e6-039680fb8844/Solanum_habrochaites_PI127826_CDS_Dovetails_2021.fasta">Solanum_habrochaites_PI127826_CDS_Dovetails_2021.fasta </a></li> <li><strong>Proteome (2021 Dovetails): </strong><a href="https://zenodo.org/api/files/13145d97-7394-4b89-83e6-039680fb8844/Solanum_habrochaites_PI127826_protein_Dovetails_2021.fasta">Solanum_habrochaites_PI127826_protein_Dovetails_2021.fasta</a><strong> </strong></li> </ul> </li> </ul> <p> </p> <p><strong>=== Genome annotations files ===</strong></p> <p><strong><em>Solanum lycopersicum </em>Heinz1706</strong></p> <ul> <li><strong>ITAG2.4</strong> <ul> <li>Gene File Format (GFF): <a href="https://zenodo.org/api/files/c5778399-5188-4959-93b7-1738584c0f72/ITAG2.4_gene_models.gff3">ITAG2.4_gene_models.gff </a></li> <li>Gene Transfer Format (GTF): <a href="https://zenodo.org/api/files/3e34c90f-9fce-4947-9ad3-0573572d942b/ITAG2.4_gene_models.gtf">ITAG2.4_gene_models.gtf</a></li> </ul> </li> <li><strong>ITAG4.0</strong> <ul> <li>Gene File Format (GFF): <a href="https://zenodo.org/api/files/5d1c61b1-e0b9-4351-8fd4-127edb9b8e08/ITAG4.0_gene_models.gff?versionId=f22ed8c3-6629-4d88-b6b7-472f3cd7c975">ITAG4.0_gene_models.gff</a></li> <li>General Transfer Format (GTF): <a href="https://zenodo.org/api/files/57840fa9-db90-4795-af03-8dbd2f711b69/ITAG4.0_gene_models.gtf">ITAG4.0_gene_models.gtf</a></li> <li>MapMan annotation: <a href="https://zenodo.org/api/files/5d1c61b1-e0b9-4351-8fd4-127edb9b8e08/S_lycopersicum_ITAG4.0_mapping_Mercator_v.3.6.tsv?versionId=bfa0b3d1-352a-470a-9504-c0f7611045f1">S_lycopersicum_ITAG4.0_mapping_Mercator_v.3.6.tsv</a> was obtained with Mercator 3.6 using the ITAG4.0_proteins.fasta file.</li> </ul> </li> </ul> <p><strong><em>Solanum lycopersicoides </em>LA2951</strong></p> <ul> <li>Gene File Format: <a href="https://zenodo.org/api/files/77a91023-8321-4f8e-86a6-1ae3b8197edf/S_lycopersicoides_LA2951_v1.0_gene_models_all.gff3">S_lycopersicoides_LA2951_v1.0_gene_models_all.gff3 </a></li> </ul> <p><strong><em>Solanum habrochaites </em>PI127826</strong></p> <ul> <li>(Based on the 2018 Hotel Project assembly): a GFF file was produced using RepeatMasker and funannotate and is named<a href="https://zenodo.org/api/files/aa27d9ec-6ab7-4582-a17c-be7dfb0952a4/Solanum_habrochaites_PI127826.gff3?versionId=4bc7492e-2de1-4b67-a21c-bb04d2bc6c10"> Solanum_habrochaites_PI127826.gff3</a>. The companion script with the performed steps is available in this data record as well and is called <a href="https://zenodo.org/api/files/aa27d9ec-6ab7-4582-a17c-be7dfb0952a4/S_habrochaites_PI127826_funannotate_steps.sh">S_habrochaites_PI127826_funannotate_steps.sh</a></li> <li>(Based on the 2021 Dovetails Genomic project): <a href="https://zenodo.org/api/files/13145d97-7394-4b89-83e6-039680fb8844/Solanum_habrochaites_PI127826_gene_models.gff">Solanum_habrochaites_PI127826_gene_models.gff</a></li> </ul> <p><strong>Additional information:</strong></p> <ul> <li>2021 Dovetails Genomics complete <strong>assembly</strong> project report<strong>: </strong><a href="https://zenodo.org/api/files/2b9490ff-031c-49d7-85e3-c57ba37ff21a/dovetails_genomics_2021.tar.gz">dovetails_genomics_2021.tar.gz </a></li> <li>2021 Dovetails Genomics complete <strong>annotation </strong>project report: <a href="https://zenodo.org/api/files/13145d97-7394-4b89-83e6-039680fb8844/dovetails_genomics_annotation_report_2021.tar.gz">dovetails_genomics_annotation_report_2021.tar.gz</a></li> </ul> <p><strong>Reference:</strong></p> <p>Tomato Genome Sequencing Consortium. 2012. The tomato genome sequence provides insights into fleshy fruit evolution. Nature volume 485, pages 635–641.</p> <p>Bolger et al. 2014. The genome of the stress-tolerant wild tomato species Solanum pennellii http://www.nature.com/ng/journal/v46/n9/full/ng.3046.html </p> <p>Hosmani et al. 2019. An improved de novo assembly and annotation of the tomato reference genome using single-molecule sequencing, Hi-C proximity ligation and optical maps. <a href="https://www.biorxiv.org/content/10.1101/767764v1">https://www.biorxiv.org/content/10.1101/767764v1</a></p> <p>Aflitos et al. 2014. Exploring genetic variation in the tomato (<em>Solanum</em> section <em>Lycopersicon</em>) clade by whole‐genome sequencing. <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/tpj.12616">https://onlinelibrary.wiley.com/doi/full/10.1111/tpj.12616</a></p> <p>Stam et al. 2019. The <em>de Novo</em> Reference Genome and Transcriptome Assemblies of the Wild Tomato Species <em>Solanum chilense</em> Highlights Birth and Death of NLR Genes Between Tomato Species. G3: Genes, Genomes, Genetics December 1, 2019 vol. 9 no. 12 3933-3941; <a href="https://doi.org/10.1534/g3.119.400529">https://doi.org/10.1534/g3.119.400529 </a></p> <p> </p> <p> </p>
Aggregated variant data from whole-genome sequenced tinnitus patients (TIGER)
<p>Aggregated variant data obtained from tinnitus patients from Sweden.</p> <p>Uploaded datasets are storage in annotated csv files. Annotation was performed using VEP (v106), including population frequencies for each variant from gnomAD, non-finnish Europeans from gnomAD, and swedish population from SweGen project. Pathogenicity scores from CADD are also annotated for each variant. Variants from genes found to be enriched in a gene burden analysis can be found in this aggregated dataset.</p> <p><strong>agg.tiger.csv</strong> - TIGER cohort is composed by 97 swedish whole-genome sequenced constant tinnitus patients.</p> <p><strong>agg.jaguar.csv</strong> - JAGUAR cohort is composed by 147 swedish whole-exome sequenced tinnitus patients .</p> <p><strong>agg.sevtin.csv</strong> - SEVTIN cohort is a subcohort from TIGER, with 34 WGS patients seggregating severe tinnitus phenotype.</p> <p><strong>agg.controls.csv</strong> - Controls is a swedish population cohort composed by 151 whole-exome sequenced swedish individuals.</p>
Whole-genome bisulfite sequencing (WGBS) for LKB1 overexpressing and NNMT knockdown A549 cells
GEO Series GSE307897. Homo sapiens. 8 samples. Type: Methylation profiling by high throughput sequencing.
Genome level identification of transcription start sites by nanoCAGE sequence in soybean (Glycine max).
GEO Series GSE302313. Glycine max. 4 samples. Type: Expression profiling by high throughput sequencing.
Whole genome bisulfite sequencing data of Russian wheat aphid biotypes US-RWA1 (co-fed), US-RWA2 (co-fed) and US-RWA2 (isolated)
GEO Series GSE185975. Diuraphis noxia. 9 samples. Type: Methylation profiling by high throughput sequencing.
Genome-wide transcriptome sequencing at different stages of Drosophila development, RNA-seq
GEO Series GSE18068. Drosophila melanogaster. 12 samples. Type: Expression profiling by high throughput sequencing.
Whole genome bisulfite Sequencing data of four types mouse spermatogenesis
GEO Series GSE137743. Mus musculus. 4 samples. Type: Methylation profiling by high throughput sequencing.
single-nucleotide resolution sequencing of N6-methyladenine in E. coli genomes and mammalian mitochondrial DNA
GEO Series GSE194212. Homo sapiens; Escherichia coli. 9 samples. Type: Methylation profiling by high throughput sequencing.
Integrative Analysis of Tamoxifen-resistant Cell Line Models Based on Sequencing Genomes, Transcriptomes and Epigenomes [aCGH]
GEO Series GSE55380. Homo sapiens. 4 samples. Type: Genome variation profiling by array.
Simultaneous Single-cell Genome and Transcriptome Sequencing with Digital Microfluidics Identifying Essential Driving Genes
GEO Series GSE199010. Homo sapiens. 91 samples. Type: Expression profiling by high throughput sequencing; Other.
Genome-wide total RNA and MBD-sequencing in HCT116 and DKO cells as a global re-expression model [MBD-Seq]
GEO Series GSE45333. Homo sapiens. 8 samples. Type: Methylation profiling by high throughput sequencing.
Genome-wide total RNA and MBD-sequencing in HCT116 and DKO cells as a global re-expression model
GEO Series GSE45334. Homo sapiens. 25 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing.
Whole genome sequencing of T-47D cell colonies overexpressed with APOBEC3B and hUGI
GEO Series GSE193226. Homo sapiens. 12 samples. Type: Other.
Genome-wide mapping cis-regulatory sequence elements in the human genome using Nuclease-Hypersensitive-CGH technique
GEO Series GSE12650. Homo sapiens. 4 samples. Type: Expression profiling by array; Genome binding/occupancy profiling by genome tiling array.
Draft Genome sequences of Rhodotorula mucilaginosa isolated from the International space station
Draft Genome sequences of Rhodotorula mucilaginosa isolated from the International space station.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.