Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
zenodo36/100

MACIE scores for human genome assembly GRCh37 Part 3 (Chr8 - Chr13)

<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of &ldquo;not evolutionarily conserved and regulatory functional&rdquo; (MACIE01); &ldquo;evolutionarily conserved and not regulatory functional&rdquo; (MACIE10); &ldquo;not evolutionarily conserved and not regulatory functional&rdquo; (MACIE00); &ldquo;both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo; or &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01, MACIE10, and MACIE11.</p>

opencc-by-4.0Dec 2021View details →
dryad36/100

Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C

<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>

opencc-zeroDec 2021View details →
zenodo36/100

Metagenome-assembled genomes(MAGs) generated from soil dataset.

<p>MAGs generated from soil dataset with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(pretrain).tar.gz.&nbsp;</p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

De novo assembly of 20 chicken genomes reveals the undetectable phenomenon for thousands of core genes on micro-chromosomes and sub-telomeric regions

<p>The gene numbers and evolutionary rates of birds were assumed to be much lower than those&nbsp;of mammals, which is&nbsp;in sharp contrast to the huge species number and morphological diversity of birds. It is therefore&nbsp;necessary to construct a complete avian genome and analyze its evolution. We constructed a chicken pan-genome from 20 <em>de novo</em>&nbsp;assembled&nbsp;genomes&nbsp;with high sequencing depth, and&nbsp;identified 1,335 protein-coding genes and 3,011 long noncoding RNAs not found in GRCg6a. The majority of these novel genes were detected across most individuals of the examined transcriptomes but were seldomly&nbsp;measured in each of the DNA sequencing data regardless of Illumina or PacBio technology. Furthermore, different from previous pan-genome models, most of these novel genes were overrepresented on chromosomal sub-telomeric regions&nbsp;and micro-chromosomes, surrounded by&nbsp;extremely high proportions of tandem repeats, which&nbsp;strongly blocks&nbsp;DNA sequencing. These hidden genes were proved to be shared by all chicken genomes, included many housekeeping genes, and enriched in immune pathways. Comparative genomics revealed the novel genes had three-fold elevated substitution rates than known ones, updating the knowledge about&nbsp;evolutionary rates in&nbsp;birds. Our study provides a framework for constructing a better chicken genome, which will contribute towards the understanding of avian evolution and improvement of poultry breeding.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Twenty-five metagenome assembled genomes recovered from the gut microbiome of the domestic ferret, Mustela putorius

<p>This dataset is composed of 25 unique metagenome assembled genomes (MAGs) recovered from the gut microbiome of three domestic ferrets (<em>Mustela putorius</em>). Details on both MAG and host ferret metadata, as well as information on sample collection, DNA sequencing, and bioinformatic processing can be found in the American Society for Microbiology Resource Announcement by Amundson et al. (in prep).&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

HiFi Metagenomic Sequencing Enables Assembly of Accurate and Complete Genomes from Human Gut Microbiota.

<p>We reported 102 complete metagenome assembled genomes (cMAGs) from five human fecal HiFi sequencing samples.</p> <p>102_cMAGs_fna.tar.gz: Fasta sequence files of 102 cMAGs.</p> <p>gc_skew_figures.tar.gz: GC-skew pattern figures of 102 cMAGs. (SVG format)</p> <p>coverage_plots.tar.gz: Genome coverage plot of 102 cMAGs.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Dereplicated Metagenome assembled genomes (MAGs) from Columbia River hyporheic sediments

<p>Fasta file containing 55&nbsp;metagenome assembled genomes (MAGs) from&nbsp;publication to be submitted titled&nbsp;&quot;<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments&quot;.&nbsp;</strong></p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments

<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2)&nbsp;A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p>&nbsp;</p> <p>Manuscript title&nbsp;Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

ATAC-seq processing resources for the GRCm38 (mm10) assembly of the mouse genome

<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the&nbsp;GRCm38 (mm10) assembly of the mouse genome&nbsp;using&nbsp;the&nbsp;<a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing &amp; Analysis Pipeline</a>&nbsp;(details in the documentation on GitHub).</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Genomic, transcriptomic and proteomic comparison of MRSA CC398 isolates collected from human and wild animal samples (Genome assembly and annotation dataset)

<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following methicillin-resistant <em>Staphylococcus aureus</em> (MRSA) strains: MRSA CC398 isolates recovered from humans, namely C5621 and C9017, and from a wild boar, namely OR418.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB35102).</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

A high-quality genome assembly and annotation of the dark-eyed junco Junco hyemalis, a recently diversified songbird

<p>The dark-eyed junco (<i>Junco hyemalis</i>) is one of the most common passerines of North America, and has served as a model organism in studies related to ecophysiology, behavior and evolutionary biology for over a century. It is composed by at least six distinct, geographically structured forms of recent evolutionary origin presenting remarkable variation in phenotypic traits, migratory behavior and habitat. Here we report a high-quality genome assembly and annotation of the dark-eyed junco generated using a combination of shotgun libraries and proximity ligation Chicago<sup>TM</sup> and Dovetail HiC<sup>TM</sup> libraries. The final assembly is 1,031,523,571 bp long, with 98.3% of the sequence located in 30 full or nearly full chromosome scaffolds, and with a N50/L50 of 71,3 Mb/5 scaffolds. We identified 19,026 functional genes combining gene prediction and similarity approaches, of which 15,967 were associated to GO terms. Genome assembly and annotated set of genes yielded 95.4% and 96.2% completeness scores, respectively, when compared with the BUSCO avian dataset. This new assembly for <i>J. hyemalis </i>provides a valuable resource for genome evolution analysis, as well as for identifying functional genes involved in adaptive processes and speciation.</p>

opencc-zeroApr 2022View details →
zenodo36/100

Metagenome-assembled genomes obtained from fecal and salivary microbiomes of pancreatic cancer patients and controls

<p>7,546 MAGs obtained from fecal and salivary metagenomes of pancreatic cancer patients and controls</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Hydractinia strain 236-21 genome assembly and Alr domain predictions

<p>This dataset is related to the preprint &quot;A family of unusual A family of unusual immunoglobulin superfamily genes in an invertebrate histocompatibility complex&quot; (<a href="https://www.biorxiv.org/content/10.1101/2022.03.04.482883v2">https://www.biorxiv.org/content/10.1101/2022.03.04.482883v2</a>).</p> <p><strong>Preprint Abstract:</strong></p> <p>Most colonial marine invertebrates are capable of allorecognition, the ability to distinguish between themselves and conspecifics. One long-standing question is whether invertebrate allorecognition genes are homologous to vertebrate histocompatibility genes. In the cnidarian <em>Hydractinia symbiolongicarpus, </em>allorecognition is controlled by at least two genes, <em>Allorecognition 1</em> (<em>Alr1</em>) and <em>Allorecognition 2 </em>(<em>Alr2</em>), which encode highly polymorphic cell surface proteins that serve as markers of self. Here, we show that <em>Alr1</em> and <em>Alr2</em> are part of a family of 41 <em>Alr </em>genes, all of which reside a single genomic interval called the Allorecognition Complex (ARC). Using sensitive homology searches and highly accurate structural predictions, we demonstrate that the Alr proteins are members of the immunoglobulin superfamily (IgSF) with V-set and I-set Ig domains unlike any previously identified in animals. Specifically, their primary amino acid sequences lack many of the motifs considered diagnostic for V-set and I-set domains, yet they adopt secondary and tertiary structures nearly identical to canonical Ig domains. Thus, the V-set domain, which played a central role in the evolution of vertebrate adaptive immunity, was present in the last common ancestor of cnidarians and bilaterians. Unexpectedly, several Alr proteins also have immunoreceptor tyrosine-based activation motifs (ITAMs) and immunoreceptor tyrosine-based inhibitory motifs (ITIMs) in their cytoplasmic tails, suggesting they could participate in pathways homologous to those that regulate immunity in humans and flies. This work expands our definition of the IgSF with the addition of a family of unusual members, several of which play a role in invertebrate histocompatibility.</p> <p><strong>This dataset contains:</strong></p> <ol> <li><strong>Hsym-236-21-genome-assembly.fa.gz</strong>: A gzip-compressed FASTA-formatted file of the genome assembly generated in the paper.&nbsp;</li> <li><strong>Alr-domain-structure-predictions.zip:</strong> a zip-compressed file with structural predictions produced with Colabfold for all domains of the Alr proteins described in that manuscript.</li> </ol>

opencc-by-4.0May 2022View details →
zenodo36/100

Genome assemblies of 118 Pisum accessions used for Pisum pan-genome analysis

<p>This repository stores genome assemblies of 118&nbsp;<em>Pisum</em> accessions used for <em>Pisum</em> pan-genome anaylsis.</p> <p>Please refer to the supplementary data in publication and NCBI BioSamples for details.</p> <p>Associated NCBI BioProject :&nbsp;<strong>PRJNA730094</strong></p> <p>Correspondance&nbsp;: gaoshh@im.ac.cn</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Hydractinia symbiolongicarpus genome and transcriptome assemblies

<p>Scaffold-level genome assemblies for Hydractinia symbiolongicarpus histoincompatible siblings BC-3 and BC-15 from a back-cross population derived from wildtype individuals&nbsp;. Assemblies generated with ABySS short read assembler using&nbsp;Illumina 200-bp insert&nbsp;paired-end libaries and 3-Kbp insert mate pair &#39;long jumping distance&#39; libraries. The trimmed read&nbsp;depths were ~36X and ~49X for BC-3 and BC-15, respectively.</p> <p>Transcriptome assemblies for Hydractinia symbiolongicarpus wildtype individuals HWB-103 and HWB-29. Assemblies generated with Trinity using Illumina mRNA libraries.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Acanthamoeba castellanii genome assembly and infection by Legionella pneumophila

<p>Data associated with the publication &quot;<em>Regulation of the Acanthamoeba castellanii genome upon infection by Legionella pneumophila</em>&quot;. The record contains 4 archives, each associated with a github repository, and a &quot;shared assets&quot; archive, which contains processed files used by some repositories. The code from github repositories is embedded in each tarball, along with input and output data. Analyses are organized as independent snakemake pipelines for each part.</p> <p>&nbsp;</p> <p>For convenient reanalysis, genomes, annotations and merged contact maps used in the publication can be found in the `shared_assets.tar.gz` archive. The infection analysis results are located in the `data/output` folder of Acastellanii_legionella_infection.tar.gz.</p> <p>All archives can be downloaded at the bottom of the page.</p> <p>&nbsp;</p> <p><strong>Hybrid genome assembly:</strong></p> <p>Genome assembly pipeline code and output data used for the assembly of 2 <em>A. castellanii</em> strains (Neff and C3) through a hybrid pipeline combining Illumina shotgun, Hi-C and Oxford Nanopore long reads.</p> <p>Github: <a href="https://github.com/cmdoret/Acastellanii_hybrid_assembly">https://github.com/cmdoret/Acastellanii_hybrid_assembly</a></p> <p>Archive: Acastellanii_hybrid_assembly.tar.gz</p> <p>&nbsp;</p> <p><strong>Genome annotation:</strong></p> <p>Genome annotation pipeline used for functional annotation of <em>A. castellanii</em> strains C3 and Neff, and associated output files.</p> <p>Github: <a href="https://github.com/cmdoret/Acastellanii_genome_annotation">https://github.com/cmdoret/Acastellanii_genome_annotation</a></p> <p>Archive: Acastellanii_genome_annotation.tar.gz</p> <p>&nbsp;</p> <p><strong>Genome analyses:</strong></p> <p>Code and data related to general analyses of genomic properties of <em>A. castellanii</em> strains C3 and Neff.</p> <p>Github: <a href="https://github.com/cmdoret/Acastellanii_genome_analysis">https://github.com/cmdoret/Acastellanii_genome_analysis</a></p> <p>Archive: Acastellanii_genome_analysis.tar.gz</p> <p>&nbsp;</p> <p><strong>Infection analyses:</strong></p> <p>Code and data related to the analysis of structural changes in the <em>A. castellanii</em> C3 genome during infection by <em>L. pneumophila</em>.</p> <p>Github: <a href="https://github.com/cmdoret/Acastellanii_legionella_infection">https://github.com/cmdoret/Acastellanii_legionella_infection</a></p> <p>Archive: Acastellanii_legionella_infection.tar.gz<br> &nbsp;</p> <p><strong>Shared assets:</strong></p> <p>This archive contains processed files (genomes, annotations, Hi-C matrices, differential expression results) which can be useful for reanalysis, and are automatically pulled when executing the pipeline of some repositories.</p> <p>Archive: shared_assets.tar.gz</p> <p>&nbsp;</p> <p><strong>Supp. analyses:</strong></p> <p>Code and data related to short ad-hoc analyses on the genomic location of specific sequences in the genomes of C3 and Neff. The archive contains two subfolders: `telomere_repeats` where we analyse the distribution of TTAGGG subtelomeric repeats throughout the A. castellanii assemblies, and `C3_exclusive_regions` where we visualize the genomic distribution of C3-specific sequences (i.e. absent from Neff) along the C3 assembly.</p> <p>&nbsp;</p> <p>Archive: supp_analyses.tar.gz<br> &nbsp;</p>

opencc-by-4.0Sep 2021View details →
dryad36/100

A chromosome-scale genome assembly of the okapi (Okapia johnstoni)

<p><span>The okapi (<em>Okapia johnstoni</em>), or forest giraffe, is the only species in its genus and the only extant sister group of the giraffe within the family Giraffidae. The species is one of the remaining large vertebrates surrounded by mystery because of its elusive behavior as well as the armed conflicts in the region where it occurs, making it difficult to study. Deforestation puts the okapi under constant anthropogenic pressure, and it is currently listed as "Endangered" on the IUCN Red List. Here, we present the first annotated de novo okapi genome assembly based on PacBio continuous long reads, polished with short reads, and anchored into chromosome-scale scaffolds using Hi-C proximity ligation sequencing. The final assembly (TBG_Okapi_asm_v1) has a length of 2.39 Gbp, of which 98% are represented by 28 scaffolds &gt;3.9 Mbp. The contig N50 of 61 Mbp and scaffold N50 of 102 Mbp, together with a BUSCO score of 94.7%, and 23,412 annotated genes, underline the high quality of the assembly. This chromosome-scale genome assembly is a valuable resource for future conservation of the species and comparative genomic studies among the giraffids and other ruminants.</span></p>

opencc-zeroJul 2022View details →
zenodo36/100

Early-life human gut metagenome-assembled genomes and proteins catalogs

<p>The description of the files:</p> <p>(1) The 32,277 genomes include&nbsp;six parts:&nbsp;ELGG_part_1.zip,&nbsp;ELGG_part_2.zip,&nbsp;ELGG_part_3.zip,&nbsp;ELGG_part_4.zip,&nbsp;ELGG_part_5.zip,&nbsp;ELGG_part_6.zip.</p> <p>(2) The 2,172 representative&nbsp;species: ELGG_representatives_2172.zip.</p> <p>(3) The&nbsp;ELGP&nbsp;catalog&nbsp;clustered at 95% amino acid identity:&nbsp;ELGP_95.faa.gz.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Enterobacter cloacae complex (E. bugandensis species, ST599) strain associated with a catheter-related bloodstream infection (CRBSI) (genome assembly and annotation dataset)

<p>This dataset includes the&nbsp;assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a&nbsp;<em><strong>Enterobacter&nbsp;cloacae </strong></em><strong>complex</strong><em><strong>&nbsp;(E. bugandensis </strong></em><strong>species</strong><em><strong>, </strong></em><strong>ST599</strong><em><strong>) </strong></em>strain associated with a catheter-related bloodstream infection&nbsp;(CRBSI) (genome anotation was performed using&nbsp;Prokka v1.14.5; https://github.com/tseemann/prokka)</p> <p>The raw sequence reads were&nbsp;deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB45360; Run&nbsp;Accession:&nbsp;ERR10044433).</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Metagenome-assembled genomes (MAGs), colorectal cancer (CRC)

<p>This archive contains (i) Metagenome assemblies of short-term enrichment cultures of CRC mucosal tissue microbiota, and (ii) Reconstructed metagenome-assembled genomes (MAGs) generated through binning of metagenome contigs.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record