Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
zenodo36/100

Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)

<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Alternative assemblies for Pteropus medius genome

<p>This dataset correspond to the alternative assemblies for P. medius genome.</p> <p>The final assembly has been deposited on ENA: https://www.ebi.ac.uk/ena/data/view/GCA_902729225</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Assemblies and annotations from the paper "Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria"

<p>These files are the assemblies and annotations produced and analyzed in the revised version of the manuscript entitled &quot;Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria&quot;.</p>

opencc-by-4.0Dec 2019View details →
dryad36/100

A high-quality genome assembly and annotation of the gray mangrove, Avicennia marina

<p class="CxSpFirst">The gray mangrove [<i>Avicennia marina</i> (Forsk.) Vierh.] is the most widely distributed mangrove species, ranging throughout the Indo-West Pacific. It presents remarkable levels of geographic variation both in phenotypic traits and habitat, often occupying extreme environments at the edges of its distribution. However, subspecific evolutionary relationships and adaptive mechanisms remain understudied, especially across populations of the West Indian Ocean. High-quality genomic resources accounting for such variability are also sparse. Here we report the first chromosome-level assembly of the genome of <i>A. marina</i>. We used a previously release draft assembly and proximity ligation libraries Chicago and Dovetail HiC for scaffolding, producing a 456,526,188 bp long genome. The largest 32 scaffolds (22.4 Mb to 10.5 Mb) accounted for 98 % of the genome assembly, with the remaining 2% distributed among much shorter 3,759 scaffolds (62.4 Kb to 1 Kb). We annotated 45,032 protein-coding genes using tissue-specific RNA-seq data in combination with <i>de novo</i> gene prediction, from which 34,442 were associated to GO terms. Genome assembly and annotated set of genes yield a 96.7% and 95.1% completeness score, respectively, when compared with the eudicots BUSCO dataset. Furthermore, an F<sub>ST</sub> survey based on resequencing data successfully identified a set of candidate genes potentially involved in local adaptation, and revealed patterns of adaptive variability correlating with a temperature gradient in Arabian mangrove populations. Our <i>A. marina </i>genomic<i> </i>assembly provides a highly valuable resource for genome evolution analysis, as well as for identifying functional genes involved in adaptive processes and speciation.</p>

opencc-zeroMay 2020View details →
dryad36/100

Data from: Genome assembly of the ragweed leaf beetle, a step forward to better predict rapid evolution of a weed biocontrol agent to environmental novelties

<p><span>Rapid evolution of weed biological control agents (BCAs) to new biotic and abiotic conditions is poorly understood and so far, only little considered both in pre-release and post-release studies, despite potential major negative or positive implications for risks of non-targeted attacks or for colonizing yet unsuitable habitats, respectively. Provision of genetic resources, such as assembled and annotated genomes, is essential to assess potential adaptive processes by identifying underlying genetic mechanisms. Here, we provide the first sequenced genome of a phytophagous insect used as a BCA, <i>i.e.</i> the leaf beetle <i>Ophraella communa</i>, a promising BCA of common ragweed, recently and accidentally introduced into Europe. A total 33.98 Gb of raw DNA sequences, representing c. 43-fold coverage, were obtained using the PacBio SMRT-Cell sequencing approach. Among the five different assemblers tested, the SMARTdenovo assembly displaying the best scores was then corrected with Illumina short reads. A final genome of 774 Mb containing 7,003 scaffolds was obtained. The reliability of the final assembly was then assessed by benchmarking universal single-copy orthologous genes (&gt; 96.0% of the 1,658 expected insect genes) and by remapping tests of Illumina short reads (average of 98.6% ± 0.7% without filtering). The number of protein-coding genes of 75,642, representing 82% of the published antennal transcriptome, and the phylogenetic analyses based on 825 orthologous genes placing <i>O. communa </i>in the monophyletic group of Chrysomelidae, confirm the relevance of our genome assembly. Overall, the genome provides a valuable resource for studying potential risks and benefits of this BCA facing environmental novelties.</span></p>

opencc-zeroMay 2020View details →
dryad36/100

Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways

Background: Teak, a member of the Lamiaceae family, produces one of the most expensive hardwoods in the world. High demand coupled with deforestation have caused a decrease in natural teak forests, and future supplies will be reliant on teak plantations. Hence, selection of teak tree varieties for clonal propagation with superior growth performance is of great importance, and access to high-quality genetic and genomic resources can accelerate the selection process by identifying genes underlying desired traits. Findings: To facilitate teak research and variety improvement, we generated a highly contiguous, chromosomal-scale genome assembly using high-coverage PacBio long reads coupled with high-throughput chromatin conformation capture. Of the 18 teak chromosomes, we generated 17 near-complete pseudomolecules with one chromosome present as two chromosome arm scaffolds. Genome annotation yielded 31,168 genes encoding 46,826 gene models, of which, 39,930 and 41,155 had Pfam domain and expression evidence, respectively. We identified 14 clusters of tandem-duplicated terpene synthases (TPSs), genes central to the biosynthesis of terpenes which are involved in plant defense and pollinator attraction. Transcriptome analysis revealed 10 TPSs highly expressed in woody tissues, of which, 8 were in tandem, revealing the importance of resolving tandemly duplicated genes and the quality of the assembly and annotation. We also validated the enzymatic activity of four TPSs to demonstrate the function of key TPSs. Conclusions: In summary, this high-quality chromosomal-scale assembly and functional annotation of the teak genome will facilitate the discovery of candidate genes related to traits critical for sustainable production of teak and for anti-insecticidal natural products.

opencc-zeroDec 2018View details →
dryad36/100

Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny

<p class="BodyAA">The pace of the sequencing and computational assembly of novel reference genomes is accelerating. Though DNA sequencing technologies and assembly software tools continue to improve, biological features of genomes such as repetitive sequence as well as molecular artifacts that often accompany sequencing library preparation can lead to fragmented or chimeric assemblies. If left uncorrected, defects like these trammel progress on understanding genome structure and function, or worse, positively mislead this research. Fortunately, integration of additional, independent streams of information, such as a marker-dense genetic map and conserved orthologous gene order from related taxa, can be used to scaffold together unlinked, disordered fragments and to restructure a reference genome where it is incorrectly joined. We present a tool set for automating these processes, one that additionally tracks any changes to the assembly and to the genetic map, and which allows the user to scrutinize these changes with the help of web-based, graphical visualizations. Chromonomer takes a user-defined reference genome, a map of genetic markers, and, optionally, conserved synteny information to construct an improved reference genome of chromosome models: a "chromonome". We demonstrate Chromonomer's performance on genome assemblies and genetic maps that have disparate characteristics and levels of quality.</p>

opencc-zeroAug 2020View details →
dryad36/100

Data from: NOVOWrap: an automated solution for plastid genome assembly and structure standardization

<p>Plastid genomes play an important role in genomics and evolutionary biology. Next-generation sequencing has revolutionized plastid genomic data acquisition to the point that genome assembly has become a bottlenecks for widespread utilization of plastid genome data. To solve this problem, we developed an open-source, cross-platform tool known as, NOVOWrap, which includes both command-line and graphical interfaces for automatically assembling plastid genomes on personal computers. With minimal inputs, settings, and user intervention, NOVOWrap can automatically assemble plastid genomes, validate results and standardize the structure using affordable computer resources. The performance of this software has been successfully benchmarked against the plastid genomes of 11 species belonging to lycopods, gymnosperms, and angiosperms. This program is expected to liberate researchers from laborious and cumbersome computer manipulations and create reliable and standardized genomic data.</p>

opencc-zeroSep 2020View details →
dryad36/100

Chromosome-level genome assembly of Poropuntius huangchuchieni

<p><i>Poropuntius huangchuchieni</i> is a diploid species in the family cyprinid, widely distributed in Mekong and Red River basins. Previous study suggested that it is one of the most closely related diploid ancestral species to common carp, which has allotetraploidized genome generated by merging two diploid genomes during evolution. Therefore, <i>P. huangchuchieni</i> is an ideal diploid model for polyploid evolution study in Cyprinidae. Here, we report a high-quality chromosome-level genome assembly of <i>P. huangchuchieni</i> by the integrating of the Oxford Nanopore Technology and Hi-C technology. The assembled genome size was 1021.38 Mb with a scaffold N50 of 32.93 Mb. More than 47.61% of the genome was identified as repetitive elements, and 895.66 Mb sequences were anchored onto 25 chromosomes. Of the 24,099 predicted protein-coding genes, 97.57% were functional annotated. Approximately 95.9% of complete BUSCOs were detected in the genome.The high-quality genomic data of <i>P. huangchuchieni</i> provides an ancestral diploid reference for the evolution and adaptation of allotetraploid carps.</p>

opencc-zeroDec 2019View details →
zenodo36/100

Eigen scores for human genome assembly GRCh37 Part 1 (Chr1 - Chr3)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

Eigen scores for human genome assembly GRCh37 Part 4 (Chr17 - Chr22)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

Antarctic endolithic bacterial metagenome-assembled genomes

<p>Bacterial&nbsp;assembled genomes and annotation data&nbsp;from the Antarctic cryptoendolithic communities&nbsp;collected during the XXXI (2015-16) Italian Antarctic Expedition.</p> <p>The dataset consists of 4&nbsp;zip&nbsp;archives and 3 files (comma-separated values). Here is a brief summary of their contents:</p> <ul> <li><strong>MAGs: </strong>high quality (HQ) and medium quality (MQ) bacterial&nbsp;metagenome assembled genomes.</li> <li><strong>MAGs_metadata: </strong>completeness, contamination, length, N50, GTDB classification for each MAG.</li> <li><strong>MAGs_HQ_CDS:</strong>&nbsp;&nbsp;translated coding sequences for each high quality MAG.</li> <li><strong>MAGs_HQ_Annotation: </strong>EggNOG annotation files. For each high quality&nbsp;MAG, the following files are included: <ul> <li>eggnog.emapper.annotations: the final EggNOG annotation;</li> <li>eggnog.emapper.hmm_hits:&nbsp;list of significant hits to eggNOG Orthologous Groups</li> <li>eggnog.emapper.seed_orthologs:&nbsp;best match of each query within the best Orthologous Group (OG) reported in the eggnog.emapper.hmm_hits file<strong>.</strong></li> </ul> </li> <li><strong>Jiangella_Antarctica:&nbsp;</strong><em>Candidatus Jiangella antarctica</em>&nbsp;representative genome (UniValnordMG_2_bin.36.fa) and the extracted ribosomal RNA genes (rRNA.fasta).</li> <li><strong>Order_MSA:&nbsp;</strong>protein multiple sequence alignments using the 120 GTDB bacterial marker genes. These alignments were used to estimate divergence times&nbsp;on orders containing at least 4 CBS, for a total of 19 orders.</li> <li><strong>Samples_accession</strong>: table that relates to the NCBI deposition of the shotgun metagenomes, the following info are included: <ul> <li>NCBI Sequence Read Archive (SRA)</li> <li>BioProject accession numbers</li> <li>JGI Integrated Microbial Genomes &amp; Microbiomes site IDs</li> <li>N50 values</li> <li>Metadata</li> </ul> </li> <li><strong>Samples_metadata: </strong>geographic coordinates, temperature, relative humidity and sampling date are reported.</li> </ul>

opencc-by-4.0Feb 2020View details →
zenodo36/100

De novo genome assembly of the Tobacco Hornworm moth (Manduca sexta)

<p><strong>We present the new reference genome for M sexta, JHU_Msex_v1.0, applying a combination of modern technologies in a de novo assembly to increase continuity, accuracy, and completeness. The assembly is 470 Mb and is ~25x more continuous than the original assembly, with scaffold N50 &gt;14 Mb. We annotated the assembly by lifting over existing annotations and supplementing with additional supporting RNA-based data for a total of 25,256 genes. The new reference assembly is accessible in annotated form for public use.</strong></p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

1263 Salmonella enterica draft genomes assembled from Bioproject PRJEB31846

<p>We assembled 1263&nbsp;Salmonella enterica draft genomes (raw data available from PRJEB31846).</p> <p>&nbsp;</p> <ul> <li>The dataset comprises diverse Salmonella enterica serovars collected between the years 1999 and 2019 and sequenced by the National Reference Laboratory for Salmonella on Illumina MiSeq and NextSeq technology. The data was described in more detail in &nbsp;10.1128/AEM.02265-19.</li> </ul> <ul> <li>Data were trimmed (with fastp, version 0.19.5) and assembled (with shovil-spades, version 1.1.0) using the AQUAMIS pipeline (https://gitlab.com/bfr_bioinformatics/AQUAMIS, version v1.2.0). All samples passed basic quality checks, such as sufficient base quality, coverage depth, genome length and contig number. Furthermore, no evidence for sample contamination was detected.</li> <li>The assemblies are input to a validation of chewieSnake (https://gitlab.com/bfr_bioinformatics/chewieSnake).</li> <li>The cgMLST analysis is available in https://bfr_bioinformatics.gitlab.io/chewiesnake_publicationdata/</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes

<p>This project contains anvi'o profiles and contigs databases that is used and/or referenced from the Lee STM and Khan SA, <em>et al.</em> study titled "<strong>High-resolution tracking of microbial colonization in Fecal Microbiota Transplantation experiments via metagenome-assembled genomes</strong>". The pre-print of this study is available via http://dx.doi.org/10.1101/090993.</p> <p>To be able to work with the data files you will need anvi'o <strong>v2.1.0</strong> to be installed on your system. For installation instructions, or to have access to a Docker image for anvi'o, please visit this URL: http://merenlab.org/software/anvio</p> <p>Public data:</p> <ul> <li><strong>ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION.tar.gz</strong>: Data files for a quick visualization of the 97 MAGs and their distribution across the two FMT recipients. A run script in the archive explains how to use this data.<br>  </li> <li><strong>ANVIO-FMT-D-R01-R02-MERGED-PROFILE.tar.gz</strong>: The merged anvi'o profile for the entire data, which also contains a collection of 97 MAGs identified in the donor. The profile database contains no hierarchical clustering of contigs, however, individual MAGs can be displayed via the following notation since the collection 'MAGs' describe the organization of contigs in each MAG referenced from the dataset `ANVIO-FMT-D-R01-R02-QUICK-VISUALIZATION`, as well as from the paper: "anvi-refine -c CONTIGS.db -p PROFILE.db -C MAGs -b <em>FMT-Donor_MAG_00054</em>". All MAG names are in the supplementary tables in our paper.<br>  </li> <li><strong>ANVIO-FMT-D-R01-R02-MAGs-SUMMARY.tar.gz</strong>: A static HTML website that contains FASTA files for each MAG, and TAB-delimited matrices for coverage and detection values, and others. After unpacking, you can double-click the index.html file.  </li> </ul>

opencc-by-4.0Nov 2016View details →
zenodo36/100

de novo genome assembly of the LNCaP human prostate cancer cell line

<p>Whole-genome sequencing reads from the LNCaP human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p>

opencc-by-4.0Jan 2017View details →
zenodo36/100

de novo genome assembly of the PC3 human prostate cancer cell line

<p>Whole-genome sequencing reads from the PC3 human prostate cancer cell line were used to generate a <em>de novo </em>assembly with SGA v0.10.15. Please see https://github.com/sciseim/PCaWGS for associated scripts. Library preparation was performed using a TruSeq Nano DNA kit (Illumina) with a target insert size of 350bp. Paired-end libraries (150bp) were sequenced using a HiSeqX sequencer (Illumina).</p> <p> </p> <p> </p> <p> </p>

opencc-by-4.0Jan 2017View details →
dryad36/100

A chromosome-scale assembly of the quinoa genome provides insights into the structure and dynamics of its subgenomes

<p>Quinoa (<em>Chenopodium</em> <em>quinoa</em> Willd.) is an allotetraploid seed crop with the potential to help address global food security concerns. Genomes have been assembled for three accessions of quinoa; however, all assemblies are fragmented and do not reflect known chromosome biology. Here, we used in vitro and in vivo Hi-C data to produce a chromosome-scale assembly of the Chilean quinoa accession PI 614886 (QQ74). The final assembly spanned 1.326 Gb, of which 90.5% was assembled into 18 chromosome-scale scaffolds. The genome was annotated with 54,499 protein-coding genes, 97% of which were located on the 18 largest scaffolds. We also produced an updated genome assembly for the B-genome diploid <em>C. suecicum</em> and used it, together with the A-genome diploid<em> C. pallidicaule</em>, to identify genomic rearrangements within the quinoa genome, including a large pericentromeric inversion representing 71.7% of chromosome Cq3B. Repetitive sequences comprise 65.20%, 48.61%, and 57.91% of the quinoa, <em>C. pallidicaule</em>, and <em>C. suecicum</em> genomes, respectively. Evidence suggests that the B subgenome is more dynamic and has expanded more than the A subgenome. These genomic resources will enable more accurate assessments of genome evolution within the Amaranthaceae and will facilitate future efforts to identify variation in genes underlying important agronomic traits in quinoa.</p>

opencc-zeroOct 2023View details →
zenodo36/100

Uncoupling of programmed DNA cleavage and repair jeopardizes the assembly of the Paramecium somatic genome

<p>In the ciliate <i>Paramecium</i>, the precise excision of numerous Internal Eliminated Sequences (IESs) from the somatic genome is essential at each sexual cycle. DNA double strands breaks (DSBs) are introduced by the PiggyMac endonuclease, and repaired in a highly concerted manner by the Non-Homologous End Joining pathway (NHEJ), as illustrated by the complete inhibition of DNA cleavage when Ku70/80 proteins are missing. We show here that expression of a DNA binding-deficient Ku70 mutant (Ku70-6E) permits DNA cleavage but not DSB repair, leading to accumulation of unrepaired DSBs. When wildtype and mutant Ku are co-expressed, the DSBs induced by Ku70-6E can be repaired by wildtype Ku, which uncouples DNA repair from the cleavage step. High-throughput sequencing of the developing MAC genome in these conditions reveals the presence of extremities healed by <i>de novo</i> telomere addition and numerous translocations between IES-flanking sequences.&nbsp;We conclude that coupling the two steps of IES excision ensures that both extremities are maintained together throughout the process, and propose that Ku assists PiggyMac during assembly of the synaptic pre-cleavage complex.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects

<p>This dataset represents all results files described in the paper 'MarkerScan: Separation and assembly of cobionts sequenced alongside target species in biodiversity genomics projects'.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record