Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

199

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

199 results for “reference genome”

Learn how ShareScore rates datasets ↗
zenodo36/100

Revised transcript annotations for GRCh38 reference genome and Ensembl v87.

<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh38<br> Ensembl version: 87</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Revised transcript annotations for GRCh37 (hg19) reference genome and Ensembl v90.

<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh37<br> Ensembl version: 90</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>

opencc-by-4.0Sep 2017View details →
dryad36/100

Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024

<p>Expression quantitative trait loci (eQTLs) provide a key bridge between noncoding DNA sequence variants and organismal traits. The effects of eQTLs can differ among tissues, cell types, and cellular states, but these differences are obscured by gene expression measurements in bulk populations. We developed a one-pot approach to map eQTLs in <em>Saccharomyces cerevisiae</em> by single-cell RNA sequencing (scRNA-seq) and applied it to over 100,000 single cells from three crosses. We used scRNA-seq data to genotype each cell, measure gene expression, and classify the cells by cell-cycle stage. We mapped thousands of local and distant eQTLs and identified interactions between eQTL effects and cell-cycle stages. We took advantage of single-cell expression information to identify hundreds of genes with allele-specific effects on expression noise. We used cell-cycle stage classification to map 20 loci that influence cell-cycle progression. One of these loci influenced the expression of genes involved in the mating response. We showed that the effects of this locus arise from a common variant (W82R) in the gene <em>GPA1</em>, which encodes a signaling protein that negatively regulates the mating pathway. The 82R allele increases mating efficiency at the cost of slower cell-cycle progression and is associated with a higher rate of outcrossing in nature. Our results provide a more granular picture of the effects of genetic variants on gene expression and downstream traits.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Reference genome data used for benchmarking SingleM

<p>For use as part of https://github.com/wwood/singlem-benchmarking/</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Supplemental files for the manuscript "Chromosome assembly of large and complex genomes using multiple references"

<p>This archive contains supplemental&nbsp;files for the manuscipt: &quot;Chromosome assembly of large and complex genomes using multiple references&quot;.</p> <p>It contains assemblies generated by Ragout and RACA as well as evaluation scripts that were used in our analysis.</p> <p>Each subdirectory contains an additional README file with details.</p> <p>Please note that some intermediate files were deleted&nbsp;in the interest&nbsp;of saving space. If you need access to&nbsp;those files or having&nbsp;issues with reproducing our results, don&#39;t hesitate to contact Mihkail Kolmogorov: fenderglass@gmail.com</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

Updated annotation for Aedes aegypti reference genome AaegL5 with extended 3' UTRs

<p>Updated annotation file for the AaegL5 genome generated and used in the Adavi et al. 2024 <em>bioRxiv </em>preprint: https://doi.org/10.1101/2024.08.21.608847</p> <p>Key updates (to VectorBase-55_AaegyptiLVP_AGWG.gff) include:</p> <ul> <li>Addition of several chemoreceptors that were annotated in previous work</li> <li>Automated extension of 3' UTRs (by up to 750bp) for all genes where supported by antennal neuron snRNAseq data</li> <li>Further manual extension of 3'UTRs for some chemoreceptors where supported by antennal neuron snRNAseq data</li> </ul> <p>For more information on this annotation and the way it was generated, please see the Methods section of the above preprint.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Reference genome of Digitaria radicosa

<p>The reference genome and annotation data of&nbsp;<em>Digitaria radicosa</em>.</p> <p>Publication: A chromosome-scale genome assembly of Timorese crabgrass (<em>Digitaria radicosa</em>): a useful genomic resource for the Poaceae. doi: https://doi.org/10.1101/2024.05.14.594087</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

PkA1HT reference genome assembly (version 1.0)

<p>&nbsp;</p> <p><strong><em><span>PkA1HT reference genome assembly</span></em></strong></p> <p><span><span>Long-read PacBio HiFi sequencing was performed on two distinct <em>P. knowlesi piggyBac </em>clones (PkBc38 and PkBc44) to fully investigate potential structural variation (including gene duplication or deletion) amongst mutants that may have been overlooked by short-read sequencing. The <em>piggyBac</em> insertion was cut out of the PkBc38 clone assembly to generate a reference genome for our parental PkA1-H.1 line which we&rsquo;ve named PkA1HT. The final PkA1HT reference genome was annotated with Companion (v2.2.8) using the <em>P. knowlesi</em> H strain as reference, specifying the assembly option with no contiguation</span><span>. The new PkA1HT reference genome has 18 sequences (comprised of 14 chromosomes, the mitochondrial genome, the apicoplast genome, and two unordered contigs). It has no sequencing gaps and is 25.29MB long, compared to 142 gaps and a smaller 24.32Mb for PkA1-H.1. Sequences for ten chromosomes reach into the telomeric heptamer repeats on both ends, and sequences for the four remaining chromosomes reach into the telomeric repeats for just one end (the two unordered contigs contain the remaining two unassembled chromosome ends).</span></span></p> <p>&nbsp;</p> <p><em><span>Additional methods</span></em></p> <p><span>Long-read PacBio HiFi sequencing was performed on two distinct&nbsp;<em>P. knowlesi piggyBac </em>clones (PkBc38 and PkBc44) to fully investigate potential structural variation (including gene duplication or deletion) amongst mutants that may have been overlooked by short-read sequencing. High molecular weight (HMW) DNA was extracted from parasite-infected human RBCs using the MagAttract HMW DNA Kit (Qiagen #67563), following the manufacturer&rsquo;s protocol for the manual purification of genomic DNA from whole blood. DNA quality and quantity were assessed using a Qubit fluorometer and agarose gel electrophoresis, with samples meeting the criteria of a Qubit concentration &gt;50 ng/&micro;L and intact bands on the gel. Two &micro;g of HMW genomic DNA was then sheared to an average size of ~15-20 kb, and SMRTbell libraries were prepared using the PacBio SMRTbell Prep Kit 3.0 (PacBio #102-141-700), which included DNA end-repair, adapter ligation, and nuclease treatment to remove incomplete molecules. An additional gel-based size selection step on the PippinHT was performed to remove fragments smaller than 10 kb. The final library was purified, quantified, and sequenced on the PacBio Revio platform (1x Revio Cell) to generate long-read data. Samples were generated at the University of South Florida and sequenced at the Wellcome Sanger Institute.</span></p> <p><span>Sequencing data were processed for quality control using standard PacBio workflows and a genome assembly was generated for each clone. We performed our long-read assemblies using the Canu assembler (v2.2) with default parameters, followed by polishing with ILRA (v1.5.1)</span><span>. The ILRA workflow included running ABACAS against the current PkA1-H.1 reference genome (v. 55) followed by two iterations of short-read correction using Pilon with Illumina reads of the parental line</span><span>. We performed manual finishing with the Artemis Comparison Tool (ACT)</span><span>, informed by long reads mapped back against the draft assemblies. We obtained &gt;500x coverage from our PacBio sequencing with a median read length of 15kbp, and we had telomere-to-telomere completion on several chromosomes.</span></p> <p><span>Assemblies were then compared against each other using ACT. We found no significant structural variation between the two clonal lines, with near complete co-linearity save for each transposon insertion, indicating as expected that the <em>piggyBac </em>transposon insertion introduces no wider genomic changes. To compare differences between our parental line and the PkA1-H.1 reference, we used ACT to identify possible regions of recombination, followed by a mapping approach and manual analysis in Artemis for validation</span><span>. We found a duplication of 14 genes on chromosome 7 in both clones. We further found several synteny breaks between our <em>de novo</em> assemblies and the current PkA1-H.1 reference. Analyzing those &ldquo;breakpoints&rdquo; more closely in ACT, we found that they are actually misassemblies in the current reference&nbsp; This finding of misassemblies motivated us to generate a more complete reference genome (see next section). We otherwise found no evidence of recombination or large indels vs. the reference for either clone. <span>It should be noted that our PkA1HT assembly also supersedes both the PkH1 and PkA1 assemblies in terms of stats (details to be reported elsewhere).</span></span></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

hg19/GRCh37 and hg38 reference genome

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo36/100

Resolving Organoid Brain Region Identities by Mapping Single-Cell Genomic Data to Reference Atlases

<p>Data underlying the figures in the publication &ldquo;Resolving organoid brain region identities by mapping single-cell genomic data to reference atlases&rdquo;, published in <em>Cell Stem Cell, </em><strong>2021</strong><em>, </em>28, 1148&ndash;1159.</p> <p><a href="https://www.sciencedirect.com/science/article/pii/S1934590921000655">https://www.sciencedirect.com/science/article/pii/S1934590921000655</a></p> <p>Table of contents:</p> <p><strong>1. patscreen_srt.rds</strong>; Numerical data for <em>Figure 7</em>: RNA-seq data of a patterning screen in organoids with an array of small molecules. The dataset is in the rds data format, which can be opened in the R programming language using the function `readRDS()`. Once opened, the dataset is a Seurat object (https://satijalab.org/seurat/) and contains both the transcript counts and the metadata for all samples in the screen. The raw data used in figure 7 was also deposited in ArrayExpress (<a href="https://www.ebi.ac.uk/arrayexpress/experiments/E-MTAB-10037/">https://www.ebi.ac.uk/arrayexpress/experiments/E-MTAB-10037/</a>)</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Multi-reference genome and K-mer based association mapping in Zymoseptoria tritici

<p>Data tables for a study of multi-reference genome and K-mer based association mapping of the fungal wheat pathogen <em>Zymoseptoria tritici</em></p>

opencc-by-4.0Aug 2021View details →
dryad36/100

Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data

<p class="MsoNormal">Whole genome sequencing enables us to ask fundamental questions about the genetic basis of adaptation, population structure, and epigenetic mechanisms, but usually requires a suitable reference genome for making sense of the sequence data. While the availability of reference genomes has significantly improvement in both taxonomic coverage and overall quality, this poses a challenge for researchers in determining which reference genome best suits their data. Here we compare the use of two different reference genomes for the three-spined stickleback (<em>Gasterosteus aculeatus</em>), one novel genome from a European individual and the published reference genome of a North American individual. Specifically, we investigate the impact of using a local reference versus one generated from a differentiated population on several commonly used metrics in population genomics. Through mapping genome resequencing data of 60 sticklebacks from across Europe and North America, we confirmed genome quality is an important factor in choosing a reference genome. A local reference genome did offer increased mapping efficiency and genotyping accuracy, likely stemming from the higher similarity in genome sequence and synteny. Despite comparable distributions of the metrics generated across the genome using SNP data (i.e., π, Tajima's D, and FST), window-based statistics using different references resulted in different outlier genes and enriched gene functions. In contrast, the marker-based analysis utilising DNA methylation distributions had a considerably higher overlap in outlier genes and functions when using different reference genomes. Overall, our results highlight how using a local reference genome can increase the resolution of genome scans when multiple similar-quality reference genomes are available. Such results have implications in the detection of signatures of selection.</p>

opencc-zeroJan 2023View details →
dryad36/100

Viral reference genomes to disentangle the recombinant phylogenetic history of the potyviruses

<p>Potyviruses are a large genus of plant-infecting RNA viruses in the family Potyviridae. Due to their rapid diversification and frequent recombination, reconstructing the phylogenetic history of the potyviruses has proven difficult. Phylogenies reconstructed from different protein-coding regions of the viral genome often reveal conflicing or discordant relationships. But the extent to which discordance is due to interspecific recombination versus phylogenetic noise or errors in reconstruction is unclear.    </p> <p>To explore the recombinant history of the potyviruses, we assembled a dataset containing referece genomes for 131 species of potyviruses. High-quality, full-length reference genomes for all species were obtained form NCBI GenBank. Viral genomes were carefully aligned at the codon-level and screened for recombination. The full alignment was then partitioned into several sub-alignments between each detected recombination event, such that each sub-aligment corresponds to a non-recombinant block (NRB) free of detected recombination events. Local phylogenetic trees for each NRB were then reconstructed to explore how phylogenetic relationships varied across different regions of the potyvirus genome.  </p> <p>We then used our program Espalier to disentangle the phylogenetic history of the potyviruses. Espalier reconciles and removes discordances between phylogenetic trees that are likely attributable to phylogenetic error while retaining recombination events that are strongly supported by the sequence data. Applying Espalier to the potyviruses revealed that most phylogenetic discordace between local trees is likely attributable to phylogenetic noise. Removing the discordance attributable to phylogenetic error allows us to much more clearly visualize the phylogenetic history of the potyviruses.</p>

opencc-zeroFeb 2023View details →
zenodo36/100

SARS-COV-2 genomic sequences used as references for the NASCarD method

<p>A set of 9 SARS-CoV-2 genome sequences used as reference in the NASCarD process.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

SARS-CoV2 reference genome UCN20

<p>This genome is a&nbsp;majority consensus obtained from sequencing respiratory samples and&nbsp;first passage of&nbsp;dozens&nbsp;isolates&nbsp;(UCN1 to&nbsp;UCN20, University Hospital of Caen, data mapped on&nbsp;reference&nbsp;genome&nbsp;from WIV) &nbsp;from the early beginning of the epidemic (march - april 2020).</p>

opencc-by-4.0May 2023View details →
zenodo36/100

European Reference Genome Atlas Community - Phase 1 Members - 2020-2023

<p>This dataset presents a comprehensive list of members registered as part of the European Reference Genome Atlas (ERGA, erga-biodiversity.eu) Community during ERGA Phase 1, which spanned from 2020 to 2023. The dataset includes information on the various roles undertaken by these members, particularly those who played a key role in establishing ERGA. Founding members are individuals who joined ERGA prior to the first leadership elections in February 2021, and they were instrumental in building the initial structure of ERGA. Some of these founding members were also involved in the establishment of different ERGA Committees.</p> <p>The dataset further includes details about the current and former core members and chairs of the ERGA committees. Additionally, it provides a list of the current (2023) and former Council members, along with the countries or regions they represent. The dataset is organised in alphabetical order for ease of reference.</p> <p>The co-authors of this dataset encompass both current and former ERGA Council members, listed in alphabetical order, and the current ERGA chair, as the last author.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Massive Chinese domestic pigs provide missing sequences in reference genome and reveal non-coding sequence variations regulating gene expression across Eurasian boars

<p>This dataset contains novel sequences in Chinese domestic pigs but is absent in Sscrofa 11.1 reference genome. The detailed information for each file is recorded in the README file.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

SARS-CoV-2 Viral Samples and Reference Genome for Galaxy Training Network SARS with Galaxy on AnVIL Tutorial

<p>In the lab activity, we&#39;ll see if there are genomic differences in the collected sample compared to the original SARS-CoV-2 genome (the reference). We need three files to do this:</p> <ul> <li><strong>SARS-CoV-2_reference_genome.fasta</strong>&nbsp;: the reference genome</li> <li><strong>VA_sample_forward_reads.fastq.gz</strong>: 1 of 2 raw read data files</li> <li><strong>VA_sample_reverse_reads.fastq.gz</strong>: 2 of 2 raw read data files</li> </ul> <p>The sample for this activity was derived from data collected at Virginia Commonwealth University in Richmond, VA. The researchers collected the sample with the goal of being able to track the spread and evolution of this virus state-wide, nationally, and internationally. You can download the original data&nbsp;<a href="https://www.ncbi.nlm.nih.gov/sra/?term=XGTK449087">here</a>.</p> <p>You can download the reference genome&nbsp;<a href="https://www.ncbi.nlm.nih.gov/nuccore/1798174254">here</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

The Genomic Reference Resource for African Cattle: genome sequences and high-density array variants.

<p><em>The diversity in genome resources is fundamental to designing genomic strategies for local breed improvement and utilisation. These resources also support gene discovery and enhance our understanding of the mechanisms of resilience with applications beyond local breeds. We report here the genome sequences of 573 samples (198 new genomes) and high-density (HD) array genotyping of 1,082 samples (537 new samples) from indigenous African cattle populations. The new sequences have an average genome coverage of ~30X, three times higher than the average (~10X) of the over 300 sequences already in the public domain. Following variant quality checks, we identified approximately 32.4 million sequence variants and 661,943 HD autosomal variants mapped to the Bos taurus reference genome (ARS-UCD1.2). &nbsp;The new datasets were generated as part of the Centre for Tropical Livestock Genetic and Health (CTLGH) Genomic Reference Resource for African Cattle (GRRFAC) initiative, which aspires to facilitate the generation of this livestock resource. We hope this resource will be utilised by the global scientific community and breeders for sustainable global livestock improvement.</em></p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Dataset for "Comprehensive Identification of NUMTs in the Human Reference Genome through Pan-Mitogenome"

<p><strong>存放&quot;Comprehensive Identification of NUMTs in the Human Reference Genome through Pan-Mitogenome&quot;文章中的相关数据。</strong></p> <p>包括blastn出来的原始output文件;mtDNA-like short segments fastq文件;ATAC-seq的fastq文件;</p> <p>以及文章中提及的supplementary 表格和bed文件</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record