Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
124
datasets available to search
ShareScore release 0.9.0
Dataset results
124 results for “de novo genome”
Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
Sex determination systems are diverse, especially among fish, and include genetic and/or environmental components. Unexpectedly for such a basic aspect of development, sex determination systems change rapidly during evolution and gonadal fate is not ultimate, being actively maintained lifelong. Here, sequences of expressed genes involved in maintenance of gonad identity and reproduction processes were obtained through transcriptome assembly of the brain-gonadal axis tissues of a freshwater fish inhabiting highly variable environments, the gonochoristic Iberian fish Squalius pyrenaicus. Through Illumina total RNA-sequencing, male and female transcriptomes of brain and gonad tissues were assembled with Trans-ABySS software and merged to produce a more comprehensive S. pyrenaicus transcriptome. Coding sequences (CDS) predicted by TransDecoder were annotated using blastx. By means of read mapping against the reference transcriptome and CDS datasets, using Bowtie2, the accuracy of read mapping was assessed. This first endemic Iberian cyprinid transcriptome of organs involved in reproduction processes may serve as a valuable genomic resource for studying sexual mechanisms and other aspects of evolution, such as speciation and responses to environmental changes, and may be a useful tool for conservation studies since S. pyrenaicus is an endangered species.
Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies
Short-read sequencing technologies have in principle made it feasible to draw detailed inferences about the recent history of any organism. In practice, however, this remains challenging due to the difficulty of genome assembly in most organisms and the lack of statistical methods powerful enough to discriminate among recent, non-equilibrium histories. We address both the assembly and inference challenges. We develop a bioinformatic pipeline for generating outgroup-rooted alignments of orthologous sequence blocks from de novo low-coverage short-read data for a small number of genomes, and show how such sequence blocks can be used to fit explicit models of population divergence and admixture in a likelihood framework. To illustrate our approach, we reconstruct the Pleistocene history of an oak-feeding insect (the oak gallwasp Biorhiza pallida) which, in common with many other taxa, was restricted during Pleistocene ice ages to a longitudinal series of southern refugia spanning theWestern Palaearctic. Our analysis of sequence blocks sampled from a single genome from each of three major glacial refugia reveals support for an unexpected history dominated by recent admixture. Despite the fact that 80% of the genome is affected by admixture during the last glacial cycle, we are able to infer the deeper divergence history of these populations. These inferences are robust to variation in block length, mutation model, and the sampling location of individual genomes within refugia. This combination of de novo assembly and numerical likelihood calculation provides a powerful framework for estimating recent population history that can be applied to any organism without the need for prior genetic resources.
De novo genome assembly of Meloidogyne chitwoodi
<p>A whole genome of <em>M. chitwoodi</em> was <em>de novo</em> assembled by empirically optimizing k-mer sizes </p>
The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications
<div> <div> <p><a href="https://onlinelibrary.wiley.com/doi/10.1111/age.13181">https://onlinelibrary.wiley.com/doi/10.1111/age.13181</a></p> <h1>The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications</h1> <div> </div> <div> <div> <div> <div><a href="https://onlinelibrary.wiley.com/authored-by/Chen/Jianhai">Jianhai Chen</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Zhong/Jie">Jie Zhong</a>, <a href="https://onlinelibrary.wiley.com/authored-by/He/Xuefei">Xuefei He</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Li/Xiaoyu">Xiaoyu Li</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Ni/Pan">Pan Ni</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Safner/Toni">Toni Safner</a>, <a href="https://onlinelibrary.wiley.com/authored-by/%C5%A0prem/Nikica">Nikica Šprem</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Han/Jianlin">Jianlin Han</a></div> </div> </div> </div> <p>The rapid progress of sequencing technology has greatly facilitated the de novo genome assembly of pig breeds. However, the assembly of the wild boar genome is still lacking, hampering our understanding of chromosomal and genomic evolution during domestication from wild boars into domestic pigs. Here, we sequenced and de novo assembled a European wild boar genome (ASM2165605v1) using the long-range information provided by 10× Linked-Reads sequencing. We achieved a high-quality assembly with contig N50 of 26.09 Mb. Additionally, 1.64% of the contigs (222) with lengths from 107.65 kb to 75.36 Mb covered 90.3% of the total genome size of ASM2165605v1 (~2.5 Gb). Mapping analysis revealed that the contigs can fill 24.73% (93/376) of the gaps present in the orthologous regions of the updated pig reference genome (Sscrofa11.1). We further improved the contigs into chromosome level with a reference-assistant scaffolding method. Using the ‘assembly-to-assembly’ approach, we identified intra-chromosomal large structural variations (SVs, length >1 kb) between ASM2165605v1 and Sscrofa11.1 assemblies. Interestingly, we found that the number of SV events on the X chromosome deviated significantly from the linear models fitting autosomes (<em>R</em><sup>2</sup> > 0.64, <em>p</em> < 0.001). Specifically, deletions and insertions were deficient on the X chromosome by 66.14 and 58.41% respectively, whereas duplications and inversions were excessive on the X chromosome by 71.96 and 107.61% respectively. We further used the large segmental duplications (SDs, >1 kb) events as a proxy to understand the large-scale inter-chromosomal evolution, by resolving parental-derived relationships for SD pairs. We revealed a significant excess of SD movements from the X chromosome to autosomes (<em>p</em> < 0.001), consistent with the expectation of meiotic sex chromosome inactivation. Enrichment analyses indicated that the genes within derived SD copies on autosomes were significantly related to biological processes involving nervous system, lipid biosynthesis and sperm motility (<em>p</em> < 0.01). Together, our analyses of the de novo assembly of ASM2165605v1 provides insight into the SVs between European wild boar and domestic pig, in addition to the ongoing process of meiotic sex chromosome inactivation in driving inter-chromosomal interaction between the sex chromosome and autosomes.</p> </div> </div> <div>The work has been pulished here: https://onlinelibrary.wiley.com/doi/full/10.1111/age.13181</div> <div> </div> <div>The current dataset include the genome annotation files.</div> <div> </div> <div>For the whole-genomic assembly, please check NCBI: </div> <div>https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_021656055.1/</div> <div> <table> <tbody> <tr> <th> </th> <th>GenBank</th> </tr> </tbody> <tbody> <tr> <td>Genome size</td> <td>2.5 Gb</td> </tr> <tr> <td>Total ungapped length</td> <td>2.4 Gb</td> </tr> <tr> <td>Number of scaffolds</td> <td>12,642</td> </tr> <tr> <td>Scaffold N50</td> <td>28.3 Mb</td> </tr> <tr> <td>Scaffold L50</td> <td>25</td> </tr> <tr> <td>Number of contigs</td> <td>41,323</td> </tr> <tr> <td>Contig N50</td> <td>157.9 kb</td> </tr> <tr> <td>Contig L50</td> <td>4,562</td> </tr> <tr> <td>GC percent</td> <td>42</td> </tr> <tr> <td>Genome coverage</td> <td>56.0x</td> </tr> <tr> <td>Assembly level</td> <td>Scaffold</td> </tr> </tbody> </table> <p> </p> <h2>Assembly methods</h2> <div>Sequencing technology 10xgenomics Assembly method Supernova v. 2.1.1 <p> </p> <p>part_** are genome fasta for the GCA_021656055.1</p> <p>You could use the following to combine and uncompress.</p> </div> </div> <div> <div> <div><code><span>cat</span> part_* > archive_combined.zip </code></div> </div> <div> <div> </div> <div><code>unzip archive_combined.zip</code></div> </div> </div> <div> </div> <div> </div>
Data from: De Novo Genome assembly of the Caucasian dwarf goby Knipowitschia cf. caucasica, a new alien Gobiidae invading the River Rhine
<p><strong>Background (Abstract from Paper)</strong></p> <p>The Caucasian dwarf goby <em>Knipowitschia</em> cf. <em>caucasica</em> is a new invasive alien Gobiidae spreading in the<br>Lower Rhine since 2019. Little is known about the invasion biology of the species and further investiga-<br>tions to reconstruct the invasion history are lacking genomic resources. We assembled a high-quality<br>chromosome-scale reference genome of <em>Knipowitschia</em> cf. <em>caucasica</em> by combining PacBio, Omni-C and<br>Illumina technologies. The size of the assembled genome is 956.58 Mb with a N50 scaffold length of 43 Mb,<br>which includes 92.3 % complete vertebrate/Actinopterygii Benchmarking Universal Single-Copy Orthologs.<br>98.96 % of the assembly sequence was assigned to 23 chromosome-level scaffolds, with a GC-content of<br>42.83 %. Repetitive elements account for 53.08 % of the genome. The chromosome-level genome contained<br>49,622 transcripts with 42,926 multi-exons, of which 45,512 genes were functionally annotated. In summary,<br>the high-quality genome assembly provides a fundamental basis to understand the adaptive advantage of<br>the species.<br><br>The file provided here is the <strong>non-redundant repeat library</strong> (2,812 consensus sequences of repeat families). </p> <p><strong>Method:</strong> Repetitive elements were identified de novo with RepeatModeler version 2.0.1. Repetitive DNA and soft-masking was performed with RepeatMasker version 4.1.1 (Smit et al., 2013) using the repeat library previously identified via RepeatModeler and skipping the bacterial insertion element check (-no_is) and run with rmblastn version 2.10.0+ (Flynn et al., 2020). </p>
De novo assembly and annotation of parasitic trematode genomes
<p>Contained in this release are 19 genome assemblies and annotations of parasitic trematodes, encompassing 13 species. This included representatives of the <em>Schistosoma</em> (<em>n</em> = 13 assemblies), <em>Trichobilharzia</em> (<em>n</em> = 2 assemblies), <em>Heterobilharzia americana</em> (<em>n</em> = 2 assemblies) and <em>Dicrocoelium dendriticum </em>(<em>n </em>= 1 assembly). The <em>Schistosoma curassoni</em> assembly has been released previously (10.5281/zenodo.6594833) but a new annotation is included with the original assembly here. </p> <p>These genomes were assembled from a variety of sources including stored parasites from museum collections, established laboratory strains and wild-caught isolates sampled from natural hosts in endemic regions. Using a combination of DNA sequencing approaches, all genomes were assembled into chromosomal-scale scaffolds. This was followed by genome annotation based on short-read RNA sequencing (RNA-seq) and long-read isoform sequencing (Iso-seq) transcriptomic data.</p> <p>Included here are the primary assemblies (representing a non-redundant haploid genome) for each species (*.primary.fa), alternate loci (alternate representations of loci found in a largely haploid assembly; *.haplotypes.fa) and annotations (*.gff3). Metadata for each assembly can be found in the included spreadsheets (metadata.xlsx). </p> <p>This data is part of a pre-publication release. For information on the proper use of pre-publication data shared by the Wellcome Trust Sanger Institute (including details of any publication moratoria), please see https://www.sanger.ac.uk/about/research-policies/open-access-science/.</p> <p>This repository will be updated with a complete list of collaborators/authors prior to publication. Please contact Duncan Berger (db22@sanger.ac.uk) with questions regarding pre-publication use of this dataset. </p>
CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing
<p><strong>Supplementary dataset for the manuscript:</strong></p> <p><strong>CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing</strong></p> <p> Xionghui Zhou1,*, Haizi Zheng1,*, Hailu Fu1,*, Kelsey L. Dillehay McKillip2-3, Susan M. Pinney2,4, Yaping Liu1-2,5-7 #</p> <p>Affiliations:</p> <p>1 Division of Human Genetics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>2 University of Cincinnati Cancer Center, Cincinnati, OH 45229</p> <p>3 Department of Pathology & Laboratory Medicine, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>4 Department of Environmental and Public Health Sciences, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>5 Division of Biomedical Informatics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>6 Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>7 Department of Electrical Engineering and Computing Sciences, University of Cincinnati College of Engineering and Applied Science, Cincinnati, OH 45229</p> <p>* These authors contributed equally</p> <p># Email: lyping1986@gmail.com</p>
De novo genome assembly for Eulemur rufifrons
<p>As one of the most threatened mammalian taxa, lemurs of Madagascar are facing unprecedented anthropogenic pressures. To address conservation imperatives such as this, researchers have increasingly relied on conservation genomics to identify populations of particular concern. However, many of these genomic approaches necessitate high-quality genomes. While the advent of next generation sequencing technologies and the resulting reduction of associated costs have led to the proliferation of genomic data and high-quality reference genomes, global discrepancies in genomic sequencing capabilities often result in biological samples from biodiverse host countries being exported to facilities in the Global North, creating inequalities in access and training within genomic research. Here, we present the first reference genome for the endangered red-fronted brown lemur (Eulemur rufifrons) from sequencing efforts conducted entirely within the host country using portable Oxford Nanopore sequencing. Using an archived E. rufifrons specimen, we conducted long-read, nanopore sequencing at the Centre ValBio Research Station near Ranomafana National Park, in rural Madagascar, generating over 750 Gb of sequencing data from 10 MinION flow cells. Exclusively using this long-read data, we assembled 2.215 gigabase, 20,330-contig assembly with an N50 of 98.9 Mb and a 17,108 bp mitogenome. The nuclear assembly had 31x average coverage and was comparable in completeness to other primate reference genomes, with a 95.51% BUSCO completeness score for primate-specific genes. As the first reference genome for E. rufifrons and the only annotated genome available for the speciose Eulemur genus, this resource will prove vital for conservation genomic studies while our efforts exhibit the potential of this protocol to address research inequalities and build genomic capacity. </p>
Assemblies generated in the manuscript "Geometric deep learning framework for de novo genome assembly"
<p>Assemblies evaluated in the manuscript "Geometric deep learning framework for de novo genome assembly". All the assemblies were generated by us, except CHM13.ONT.Flye-2.9.fa.gz which was generated by <a href="https://www.nature.com/articles/s41587-019-0072-8">Kolmogorov et al. (2019)</a>.</p>
Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"
<p>Quast outputs for "When do longer reads matter? A benchmark of long read de novo assembly tools for eukaryotic genomes"</p>
Risk of Recurrence of de Novo Mutations: Research and Quantification of Paternal Germinal Mosaicism by the Combined Use of Genomic Tools
ClinicalTrials.gov study NCT04564235. IPD Sharing: NO. Countries: 1. Publications: 0.
Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
Data from: "De novo transcriptome assembly and polymorphism detection in ecological important widely distributed Neotropical toads from the Rhinella marina species complex (Anura: Bufonidade)" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Open the record for dataset details and reuse information.
Data from: "De novo assembly transcriptome for the rostrum dace (Leuciscus burdigalensis, Cyprinidae: fish) naturally infected by a copepod ectoparasite" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
Open the record for dataset details and reuse information.
Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
Open the record for dataset details and reuse information.
De novo genome assembly of Leptodactylus fuscus
Open the record for dataset details and reuse information.
De novo genome assembly for Eulemur rufifrons
Open the record for dataset details and reuse information.
De novo genome assembly of Tectona grandis (Teak) with 2993 scaffolds
Open the record for dataset details and reuse information.
Data from: Two low coverage bird genomes and a comparison of reference-guided versus de novo genome assemblies
Open the record for dataset details and reuse information.
Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.