Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
199
datasets available to search
ShareScore release 0.7.1
Dataset results
199 results for “reference genome”
Reference genome resources associated with the project: Functional genetic diversity is correlated with intensity of genetic drift in populations of an endangered rattlesnake
<p class="MsoNormal">Theory predicts that genetic erosion in small, isolated populations of endangered species can be assessed using estimates of neutral genetic variation reflecting long-term impacts of genetic drift, yet this widely used approach has been questioned in the genomics era. Here we leverage a chromosome-level assembly and whole genome resequencing data (N=110 individuals) from an endangered rattlesnake (<em>Sistrurus catenatus</em>) to evaluate the relationship between genome-wide neutral and functional diversity over long- and short-term timescales. As predicted for populations at long-term equilibrium, we found a positive correlation between population-level estimates of neutral genetic diversity (π) and the mean number of highly detrimental loss-of-function mutations, and a negative relationship between neutral genetic diversity and an estimate of genetic load. In contrast, we found only a weak, non-significant positive correlation between levels of neutral and adaptive variation. Additional analyses using estimates of drift at more recent time scales (> 100 generations) show expected correlations between both measures of genetic load, but a lack of a significant correlation with levels of adaptive variation. Individual-based demographic metrics that capture drift impacts over recent time scales confirm these results. Broadly, our results confirm that estimates of diversity and demography based on neutral genetic variation provide an accurate measure of a key component of genetic erosion – genetic load – in populations of a threatened vertebrate. Our findings also provide nuance to the neutral-functional diversity controversy by demonstrating that neutral genetic diversity is useful in predicting some, but not all, components of functional genetic diversity.</p>
Gene prediction for: A reference genome for ecological restoration of the sunflower sea star, Pycnopodia helianthoides
<div> <div> <div> <div>Wildlife diseases, such as the sea star wasting (SSW) epizootic that outbroke in the mid-2010s, appear to be associated with acute and/or chronic abiotic environmental change; dissociating the effects of different drivers can be difficult. The sunflower sea star,<em> Pycnopodia helianthoides</em>, was the species most severely impacted during the SSW outbreak, which overlapped with periods of anomalous atmospheric and oceanographic conditions, and there is not yet a consensus on the cause(s). Genomic data may reveal underlying molecular signatures that implicate a subset of factors and, thus, clarify past events while also setting the scene for effective restoration efforts. To advance this goal, we used Pacific Biosciences HiFi long sequencing reads and Dovetail Omni-C proximity reads to generate a highly contiguous genome assembly that was then annotated using RNA-seq-informed gene prediction. The genome assembly is 484 Mb long, with contig N50 of 1.9 Mb, scaffold N50 of 21.8 Mb, BUSCO completeness score 96.1%, and 22 major scaffolds consistent with prior evidence that sea star genomes comprise 22 autosomes. These statistics generally fall between those of other recently assembled chromosome-scale assemblies for two species in the distantly related asteroid genus <em>Pisaster</em>. These novel genomic resources for <em>Pycnopodia helianthoides</em> will underwrite population genomic, comparative genomic, and phylogenomic analyses — as well as their integration across scales — of SSW and environmental stressors. This data resource contains the files associated with gene prediction.</div> </div> </div> </div>
A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants
<p>Underrepresentation of non-European populations hinders growth of global precision medicine. Resources such as imputation reference panels that match the study population are necessary to find low-frequency variants with substantial effects. We created a reference panel consisting of 14,393 whole-genome sequences including more than 11,000 Asian individuals. Genome-wide association studies were conducted using the reference panel and a population-specific genotype array of 72K subjects for eight phenotypes. This panel yields improved imputation accuracy of rare and low-frequency variants within East Asian populations compared with the largest reference panel. Thirty-nine previously unidentified associations were found, and more than half of the variants were East-Asian-specific. We discovered genes with rare protein-altering variants, including LTBP1 for height and GPR75 for body mass index, as well as putative regulatory mechanisms for rare noncoding variants with cell-type-specific effects. We suggest this data set will add to the potential value of Asian precision medicine.</p>
Genome and Transcriptome references based on hg19 from UCSC, 2015
<p>rsem.transcripts.nant2015.fa.gz - bgzipped FASTA reference of transcriptomes</p><p>genome.nant2015.fa.gz - bgzipped FASTA human genome reference, with several viral sequences added.</p><p>refseq.txt.gz - Exact sequence accessions and mapping coordinates for a RefSeq transcriptome based off the UCSC genome browser for hg19.</p><p>Coordinates are BED-style, with one row per transcript, and 1+ transcript per gene.</p><p>Column annotation</p><p>1. RefSeq Accession</p><p>2. Chromosome</p><p>3. Strand</p><p>4. thinStart (gene boundary, including UTR)</p><p>5. thinEnd (gene boundary, including UTR)</p><p>6. thickStart (CDS boundary)</p><p>7. thinStart (CDS boundary)</p><p>8. number of exons</p><p>9. comma separated exon starts</p><p>10. comma separate exon ends</p><p>11. common gene name</p><p>12. refseq gene id</p><p>13. 0 if non-primary transcript, 1 if primary transcript</p>
Data from: A latitudinal gradient of reference genomes
Open the record for dataset details and reuse information.
Data From: TERRA-REF, An open reference data set from high resolution genomics, phenomics, and imaging sensors
Open the record for dataset details and reuse information.
Gene prediction for: A reference genome for ecological restoration of the sunflower sea star, Pycnopodia helianthoides
Open the record for dataset details and reuse information.
Landscape connectivity and genetic structure in a mainstem and a tributary stonefly (Plecoptera) species using a novel reference genome
Open the record for dataset details and reuse information.
Transitioning from environmental genetics to genomics using mitogenome reference databases
Open the record for dataset details and reuse information.
Viral reference genomes to disentangle the recombinant phylogenetic history of the potyviruses
Open the record for dataset details and reuse information.
A chromosomal-scale reference genome of the New World Screwworm, Cochliomyia hominivorax
Open the record for dataset details and reuse information.
Novel Megaptera novaeangliae (Humpback whale) haplotype reference genome
Open the record for dataset details and reuse information.
Reference genome and annotation for Teleopsis dalmanni
Open the record for dataset details and reuse information.
A chromosome-scale reference genome and genome-wide genetic variations elucidate adaptation in yak
Open the record for dataset details and reuse information.
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
Open the record for dataset details and reuse information.
A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants
Open the record for dataset details and reuse information.
Reference genome resources associated with the project: Functional genetic diversity is correlated with intensity of genetic drift in populations of an endangered rattlesnake
Open the record for dataset details and reuse information.
Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024
Open the record for dataset details and reuse information.
Reference genome of an irruptive migrant, the pine siskin (<em>Spinus pinus</em>)
Open the record for dataset details and reuse information.
A highly contiguous reference genome for the Steller's jay (Cyanocitta stelleri)
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.