Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

200

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

200 results for “Genomic Resources”

Learn how ShareScore rates datasets ↗
zenodo44/100

SARS-CoV-2 genomics resources for Galaxy

<p>Reference and custom annotation data expected as input by Galaxy SARS-CoV-2 variation analysis workflows developed by covid19.galaxyproject.org</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

Aphidinae comparative genomics resource

<p>Here we provide early access to 18 new genome assemblies, including 8 assembled to chromosome-scale, for aphids from the subfamily Aphidinae.&nbsp;For consistency and to aid comparative analysis, all genomes have been annotated using the same repeat masking and RNA-seq-based gene prediction pipeline.&nbsp;Using this pipeline we also provide new annotations for three previously published genome assemblies.</p> <p>The genome assemblies and annotations are made freely available without restriction, we only request that this Zenodo resource is cited when using the data. Raw sequence data upload to NCBI is underway and full details of all accessions will be given in an updated version of this resource. Manuscripts are in preparation describing the individual genome assemblies in detail and larger comparative genome analyses and we will update this resource with additional citation information as papers are published.</p> <p>Full details of all genome assemblies and annotations included in this release are given in the attached &quot;Data_Description.pdf&quot; document.&nbsp;</p> <p><strong>Aphid species included in this release (bold type = chromosome-scale assembly):</strong></p> <p><em><strong>Aphis fabae</strong><br> Aphis glycines </em>(updated annotation)<br> <em><strong>Aphis gossypii</strong><br> Aphis thalictri<br> Aphis rumicis<br> Brachycaudus cardui<br> Brachycaudus helichrysi<br> Brachycaudus klugkisti<br> <strong>Brevicoryne brassicae</strong><br> Diuraphis noxia<br> <strong>Macrosiphum albifrons</strong><br> Metopolophium dirhodum<br> Myzus cerasi&nbsp;</em>(updated annotation)<br> <em>Myzus ligustri<br> Myzus lythri<br> Myzus varians<br> Pentalonia nigronervosa&nbsp;</em>(updated annotation)<br> <em><strong>Phorodon humuli</strong><br> <strong>Rhopalosiphum padi<br> Sitobion avenae<br> Sitobion miscanthi</strong></em></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix

<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation:&nbsp;https://doi.org/10.1101/2022.07.29.502101</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate FGOB10Ptm-1.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction

<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the&nbsp;<em>P. teres&nbsp;</em>f.<em>&nbsp;maculata&nbsp;</em>isolate P-A14.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

BacSPaD: A robust bacterial strains' pathogenicity resource based on integrated and curated genomic metadata

<p>The vast array of omics data in microbiology presents significant opportunities for studying bacterial pathogenesis and creating computational tools for predicting pathogenic potential. However, the field lacks a comprehensive, curated resource that catalogs bacterial strains and their ability to cause human infections. Current methods for identifying pathogenicity determinants often introduce biases and miss critical aspects of bacterial pathogenesis.<br>In response to this gap, we introduce BacSPaD (Bacterial Strains&rsquo; Pathogenicity Database), a thoroughly curated database focusing on pathogenicity annotations for a wide range of high-quality, complete bacterial genomes. Our rule-based annotation workflow combines metadata from trusted sources with automated keyword matching, extensive manual curation, and detailed literature review. Our analysis classified 5,502 genomes as pathogenic to humans (HP) and 490 as non-pathogenic to humans (NHP), encompassing 532 species, 193 genera, and 96 families. Statistical analysis demonstrated a significant but moderate correlation between virulence factors and HP classification, highlighting the complexity of bacterial pathogenicity and the need for ongoing research. This resource is poised to enhance our understanding of bacterial pathogenicity mechanisms and aid in the development of predictive models. To improve accessibility and provide key visualization statistics, we developed a user-friendly web interface, accessible at<a href="https://bacspad.altrabio.com/"> </a><a href="https://bacspad.altrabio.com/"><u>https://bacspad.altrabio.com</u></a>.</p>

opencc-by-nc-sa-4.0Aug 2024View details →
zenodo44/100

Genome data and resources on the recombination landscape and population history of the harlequin fly

<p>This dataset contains phased vcf files of <em>Chironomus riparius,&nbsp;</em>ouput files of RepeatMasker, MELT, RepeatOBserver, MSMC2, iSMC and bedtools, such as supporting files.&nbsp;</p> <p>For further details also check the GitHub page: <a href="https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024" target="_blank" rel="noopener">https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024</a></p> <ul> <li><strong>phased-vcfs: </strong>Artificially phased vcf-files of five populations with four individuals each. Needed to generate multihetsep files. Input files for iSMC.<br> <ul> <li>Hesse in Germany =&nbsp; MG</li> <li>Rh&ocirc;ne-Alpes in France = MF</li> <li>Lorraine in France = NMF</li> <li>Piemont in Italy = SI</li> <li>Andalusia&nbsp;in Spain = SS</li> </ul> </li> <li><strong>multihetsep-files:&nbsp;</strong>Created with msmc-tools. Input files for MSMC2.&nbsp;</li> <li><strong>RepeatMasker:&nbsp;</strong>Raw output of RepeatMasker run. Summary file and file with filtered <em>Cla</em>-element (a transposable element) included.<strong><br></strong></li> <li><strong>MELT: </strong>MELT ouput with added info on population and numbered insertions reflecting all 441 detected <em>Cla </em>insertions.<strong><br></strong></li> <li><strong>RepeatOBserver: </strong>Summary files on centromere predictions based on histograms and Shannon Diversity from RepeatOBserver. Genome-wise Shannon Diversity per chromosome included. <strong><br></strong></li> <li><strong>MSMC2: </strong>Raw ouput of combined cross-coalescence and mean values if MSMC2 per populations. <strong><br></strong></li> <li><strong>iSMC: </strong>Recombination rate rho in 10 kb windows and 100 kb windows along the genome. <strong><br></strong></li> <li><strong>bedtools closest ismc 10 kb: </strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 10 kb windows.<strong><br></strong></li> <li><strong>bedtools closest ismc 100 kb:&nbsp;</strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 100 kb windows.</li> <li><strong>input-files figures: </strong>Supporting files needed to create figures.<strong><br></strong></li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Mitochondrial genome sequencing and analysis of the invasive Microstegium vimineum: a resource for systematics, invasion history, and management

<p>Table S1: Accession data for Microstegium samples included in this study.</p> <p>File S1: Alignment of Mitochondrial CDS for Poales mitochondrial sequences.</p> <p>File S2: SNP data for Microstegium vimineum mitochondrial variants.</p> <p>Figure S1: Transposable element content in the Microstegium vimineum mitogenome.</p> <p>Figure S2: Summary of Kraken2 output.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Allliance of Genome Resources Orthology

<p>Tab separated formatted spreadsheet of orthology annotations from the Alliance of Genome Resources.</p> <p>The Alliance provides the results of all methods that have been benchmarked by the <a href="https://questfororthologs.org/">Quest for Orthologs Consortium (QfO)</a>, as well as curated ortholog inferences from HGNC (for human and mouse genes), Xenbase (for frog genes), and ZFIN (relating zebrafish genes to orthologs in human, mouse, and fly).</p> <p>The ortholog inferences from the different methods have been integrated using the DRSC Integrative Ortholog Prediction Tool (DIOPT). DIOPT integrates a number of existing methods including those used by the Alliance: Ensembl Compara, HGNC, Hieranoid, InParanoid, OMA, OrthoFinder, OrthoInspector, PANTHER, PhylomeDB, SonicParanoid, Xenbase, and ZFIN. See the <a href="https://fgr.hms.harvard.edu/diopt-documentation">DIOPT documentation</a> for additional information and references related to the included methods. DIOPT assigns a score/count based on the number of methods that call a specific ortholog. For noncoding RNA genes, currently only HGNC and ZFIN curated orthologs are included.</p> <p>File includes orthology relationships among genes from the following organisms:</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Alliance of Genome Resources Genetic Interactions

<p>These files provide a set of annotations of genetic interactions for genes for human, rat, mouse, zebrafish, fruit fly, nematode, African clawed frog,and yeast). The files are in the <a href="https://github.com/HUPO-PSI/miTab/blob/master/PSI-MITAB27Format.md">PSI-MI TAB 2.7 format</a>, a tab-delimited format established by the <a href="http://www.psidev.info/">HUPO Proteomics Standards Initiative</a> Molecular Interactions (PSI-MI) working group. The interaction data are sourced from Alliance members WormBase and FlyBase, as well as the <a href="https://thebiogrid.org/">BioGRID database</a>. Identities or types of genetic perturbations for each interactor (if available) are provided in columns 26 and 27 and relevant phenotypes or traits (if available) are provided in column 28.</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo44/100

New Soil Metagenome-Assembled Genomes Catalogue Boosts Genetic Resources

<p><strong>Soil harbors a vast expanse of unidentified microbes, termed as microbial dark matter, presenting an untapped reservoir of microbial biodiversity and genetic resources, but has yet to be fully explored. In this study, we conducted the first large-scale excavation of soil microbial dark matter by reconstructing 40,039 metagenome-assembled genome bins (the SMAG catalog) from 3,304 soil metagenomes. We identified 16,530 of 21,077 species-level genome bins (SGBs) as unknown SGBs (uSGBs), which greatly expand archaeal and bacterial diversity across the tree of life. We also illustrate the pivotal role of uSGBs in augmenting soil microbiome&#39;s functional landscape and intra-species genome diversity, providing large proportions of the 43,169 biosynthetic gene clusters and 8,545 CRISPR-Cas genes. Additionally, we determined that uSGBs contributed 84.6% of novel viral-host associations identified from the SMAG catalog. Our results propose the SMAG catalog, a novel and expansive genomic resource that brings the soil microbial biodiversity and novel genetic resources to light.</strong></p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Genomic resources of the Podospora anserina species complex

<p>The filamentous fungus <em>Podospora anserina</em> is a model organism used extensively in the study of molecular biology, senescence, prion biology, meiotic drive, mating-type chromosome evolution, and plant biomass degradation. It has recently been established that <em>P. anserina</em> is a member of a complex of seven, closely related species. In addition to <em>P. anserina</em>, high-quality genomic resources are available for two of these taxa. Here we provide chromosome-level annotated assemblies of the four remaining species of the complex, as well as a comprehensive dataset of annotated assemblies from a total of 28 <em>Podospora</em> genomes.</p>

opencc-zeroOct 2023View details →
dryad40/100

High-density genetic linkage mapping in Sitka spruce advances the integration of genomic resources in conifers

<p><span>In species with large and complex genomes such as conifers, dense linkage maps are a useful for supporting genome assembly and laying the genomic groundwork at the structural, populational and functional levels. However, most of the 600+ extant conifer species still lack extensive genotyping resources, which hampers the development of high-density linkage maps. In this study, </span><span><span>we developed a linkage map relying on 21,570 SNP makers in </span></span><span>Sitka spruce (<em>Picea sitchensis</em> [Bong.] Carr.)</span><span><em><span>, </span></em></span><span><span>a long-lived conifer from western North America that is widely planted for productive forestry in the British Isles. </span></span><span>We used a single-step mapping approach to efficiently combine RAD-Seq and genotyping array SNP data for 528 individuals from two full-sib families. As expected for spruce taxa, the saturated map contained 12 linkages groups with a total length of 2,142 cM. The positioning of 5,414 unique gene coding sequences allowed us to compare our map with that of other Pinaceae species, which provided evidence for high levels of synteny and gene order conservation in this family. We then developed an integrated map for <em>P. sitchensis</em> and <em>P. glauca</em> based on 27,052 makers and 11,609 gene sequences. Altogether, these two linkage maps, the accompanying catalog of 286,159 SNPs and the genotyping chip developed herein opens new perspectives for a variety of fundamental and more applied research objectives, such as for the improvement of spruce genome assemblies, or for marker-assisted sustainable management of genetic resources in Sitka spruce and related species.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome

<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the&nbsp;GRCh38 (hg38) assembly of the human genome&nbsp;using&nbsp;the&nbsp;<a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing &amp; Analysis Pipeline</a>&nbsp;(details in the documentation on GitHub).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Taking advantage from phenotype variability in a local animal genetic resource: identification of genomic regions associated with the hairless phenotype in Casertana pigs

<p>Ped and Map files for 96 Casertana breed pigs genotyped with Illumina BeadChip 60K Porcine.<br> The first field of the ped file contains the id of the farm (1az-6az).<br> The hairless phenotype, in the ped phenotype field, is codified&nbsp;as 1, the hairy phenotype is codified as 2.</p>

opencc-by-4.0Feb 2018View details →
zenodo40/100

Annotation Data: Decoding the chromosome-scale genome of the nutrient-rich Agaricus subrufescens: A Resource for fungal biology and biotechnology

<p><strong>Decoding the chromosome-scale genome of the nutrient-rich Agaricus subrufescens: A Resource for fungal biology and biotechnology</strong></p> <p>Genome annotation data</p> <p><strong>Genome Browser:</strong>&nbsp;<a href="https://plantgenomics.ncc.unesp.br/gen.php?id=Asub">https://plantgenomics.ncc.unesp.br/gen.php?id=Asub</a></p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Alliance of Genome Resources Gene Expression Data

<p>Tab separated formatted spreadsheets of gene expression annotations from the Alliance of Genome Resources. Gene expression data include temporal and/or spatial localization of transcripts and proteins in a wild-type background.</p> <p>File includes annotations for the following organisms:</p> <ul> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Alliance of Genome Resources Alleles

<p>Tab separated formatted spreadsheets of allele annotations from the Alliance of Genome Resources.</p> <p>Files include annotations for</p> <ul> <li>Caenorhabditis elegans (nematode; NCBITaxon 6239)</li> <li>Danio rerio (zebrafish;NCBITaxon 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBITaxon 7227)</li> <li>Mus musculus (mouse; NCBITaxon 10090)</li> <li>Rattus norvegicus (rat; NCBITaxon 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBITaxon&nbsp;559292 )</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Alliance of Genome Resources Molecular Interactions

<p>This file provides a set of annotations of molecular interactions for genes and gene products for human, rat, mouse, zebrafish, fruit fly, nematode, yeast, African and Western clawed frogs, and SARS-CoV-2. The file is in the <a href="https://github.com/HUPO-PSI/miTab/blob/master/PSI-MITAB27Format.md">PSI-MI TAB 2.7 format</a>, a tab-delimited format established by the <a href="http://www.psidev.info">HUPO Proteomics Standards Initiative</a> Molecular Interactions (PSI-MI) working group. The interaction data is sourced from Alliance members WormBase and FlyBase, as well as the <a href="http://www.imexconsortium.org">IMEx consortium</a> and the <a href="https://thebiogrid.org">BioGRID database</a>.</p> <p>The file contains molecular interaction annotations for the following organisms:</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> <li>Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2; NCBI:txid2697049)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Alliance of Genome Resources Disease Annotations

<p>Tab separated formatted spreadsheets of disease annotations from the Alliance of Genome Resources. Annotations are to terms in the Disease Ontology (DO), evidence codes from the Evidence Code Ontology (ECO).&nbsp; PubMed ids (PMID) are provided for the source of the annotation assertions.</p> <p>File includes annotations for the following organisms:</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> </ul>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record