Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

181

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

181 results for “de novo assembly”

Learn how ShareScore rates datasets ↗
dryad36/100

Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad36/100

De novo assembly of SNPs in VCF format for 112 individualss of Campylorhynchus in western Ecuador

Open the record for dataset details and reuse information.

publicJun 2024View details →
dryad36/100

De novo genome assembly of Kallima inachus

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad36/100

Accuracy of de novo assembly of DNA sequences from double‐digest libraries varies substantially among software

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad36/100

Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference

Open the record for dataset details and reuse information.

publicJun 2014View details →
dryad36/100

Error, noise and bias in de novo transcriptome assemblies

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad36/100

De novo genome assembly of human cell line CHM13 nanopore ultra-long reads using Shasta

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad32/100

De novo genome assembly of Tectona grandis (Teak) with 2993 scaffolds

<p>Teak (<em>Tectona grandis</em> L. f.) is one of the precious bench mark tropical hardwood having qualities of durability, strength and visual pleasantries. Natural teak populations harbour a variety of characteristics that determine their economic, ecological and environmental importance. Sequencing of whole nuclear genome of teak provides a platform for functional analyses and development of genomic tools in applied tree improvement. A draft genome of 317 Mb was assembled at 151× coverage and annotated 36, 172 protein-coding genes. Approximately about 11.18% of the genome was repetitive. Microsatellites or simple sequence repeats (SSRs) are undoubtedly the most informative markers in genotyping, genetics and applied breeding applications. We generated 182,712 SSRs at the whole genome level, of which, 170,574 perfect SSRs were found; 16,252 perfect SSRs showed <em>in silico</em> polymorphisms across six genotypes suggesting their promising use in genetic conservation and tree improvement programmes. Genomic SSR markers developed in this study have high potential in advancing conservation and management of teak genetic resources. Phylogenetic studies confirmed the taxonomic position of the genus <em>Tectona</em> within the family Lamiaceae. Interestingly, estimation of divergence time inferred that the Miocene origin of the <em>Tectona</em> genus to be around 21.4508 million years ago.</p>

opencc-zeroAug 2020View details →
dryad32/100

Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014

Drosophila (Drosophila) nigrosparsa is a habitat specialist restricted to the European montane/alpine zone (Bächli 2008). Mountain biodiversity is considered highly vulnerable to ongoing climate warming (IPCC 2013), and organisms at high altitudes have only limited possibility to shift to cooler habitats at elevations above (Pertoldi &amp; Bach 2007). For such species, rapid evolution may offer a solution for long-term survival. We are establishing D. nigrosparsa as a model system to test the extent and tempo of adaptive evolution under thermal stress in the laboratory. In this study, we used Illumina high-throughput sequencing to assemble the species' transcriptome using the pooled mRNA from 22 developmental and physiological stages.

opencc-zeroDec 2013View details →
dryad32/100

Data from: De novo assembly of the transcriptome of an invasive snail and its multiple ecological applications

Studying how invasive species respond to environmental stress at the molecular level can help us assess their impact and predict their range expansion. Development of markers of genetic polymorphism can help us reconstruct their invasive route. However, to conduct such studies requires the presence of substantial amount of genomic resources. This study aimed to generate and characterize genomic resources using high throughput transcriptome sequencing for Pomacea canaliculata, a non-model gastropod indigenous to Argentina that has invaded Asia, Hawaii and southern United States. De novo assembly of the transcriptome resulted in 128,436 unigenes with an average length of 419 bp (range: 150 to 8,556 bp). Many of the unigenes (2,439) contained transposable elements, showing the existence of a source of genetic variability in response to stressful conditions. A total of 3,196 microsatellites were detected in the transcriptome; among 20 of the randomly tested microsatellites, 10 were validated to exhibit polymorphism. A total of 15,412 single-nucleotide polymorphisms (SNPs) were detected in the ORFs. LC-MS/MS analysis of the proteome of juveniles revealed 878 proteins, of which many are stress-related. This study has demonstrated the great potential of high throughput DNA sequencing for rapid development of genomic resources for a non-model organism. Such resources can facilitate various molecular ecological studies, such as stress physiology and range expansion.

opencc-zeroDec 2011View details →
dryad32/100

Data from: "De novo assembly transcriptome for the rostrum dace (Leuciscus burdigalensis, Cyprinidae: fish) naturally infected by a copepod ectoparasite" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015

The emergence of pathogens represents substantial threats to public health, livestock, domesticated animals, and biodiversity. How wild populations respond to emerging pathogens has generated a lot of interest in the last two decades. With the recent advent of high-throughput sequencing technologies it is now possible to develop large transcriptomic resources for non-model organisms, hence allowing new research avenues on the immune responses of hosts from a large taxonomic spectra. We here focused on a wild population of the rostrum dace (Leuciscus burgiladensis) that is infected by Tracheliastes polycolpus, an emerging freshwater ectoparasite copepod. We used next generation Illumina sequencing technology to sequence the transcriptome of eight L. burdigalensis adult individuals collected in natura from the same sampling site. Four individuals were non-infected and four individuals were infected by T. polycolpus. We specifically focused on the spleen, the head kidney and epithelial cells and mucus from the fins, three tissues known to be involved in the immune response of fish. We used the Trinity methodology to reconstruct a de novo full-length transcriptome for L. burdigalensis. The resulting transcriptome will serve as an important broad-scale genomic resource for further studying the response of local population of L. burdigalensis to T. polycolpus pressures.

opencc-zeroDec 2014View details →
dryad32/100

Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015

Sex determination systems are diverse, especially among fish, and include genetic and/or environmental components. Unexpectedly for such a basic aspect of development, sex determination systems change rapidly during evolution and gonadal fate is not ultimate, being actively maintained lifelong. Here, sequences of expressed genes involved in maintenance of gonad identity and reproduction processes were obtained through transcriptome assembly of the brain-gonadal axis tissues of a freshwater fish inhabiting highly variable environments, the gonochoristic Iberian fish Squalius pyrenaicus. Through Illumina total RNA-sequencing, male and female transcriptomes of brain and gonad tissues were assembled with Trans-ABySS software and merged to produce a more comprehensive S. pyrenaicus transcriptome. Coding sequences (CDS) predicted by TransDecoder were annotated using blastx. By means of read mapping against the reference transcriptome and CDS datasets, using Bowtie2, the accuracy of read mapping was assessed. This first endemic Iberian cyprinid transcriptome of organs involved in reproduction processes may serve as a valuable genomic resource for studying sexual mechanisms and other aspects of evolution, such as speciation and responses to environmental changes, and may be a useful tool for conservation studies since S. pyrenaicus is an endangered species.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies

Short-read sequencing technologies have in principle made it feasible to draw detailed inferences about the recent history of any organism. In practice, however, this remains challenging due to the difficulty of genome assembly in most organisms and the lack of statistical methods powerful enough to discriminate among recent, non-equilibrium histories. We address both the assembly and inference challenges. We develop a bioinformatic pipeline for generating outgroup-rooted alignments of orthologous sequence blocks from de novo low-coverage short-read data for a small number of genomes, and show how such sequence blocks can be used to fit explicit models of population divergence and admixture in a likelihood framework. To illustrate our approach, we reconstruct the Pleistocene history of an oak-feeding insect (the oak gallwasp Biorhiza pallida) which, in common with many other taxa, was restricted during Pleistocene ice ages to a longitudinal series of southern refugia spanning theWestern Palaearctic. Our analysis of sequence blocks sampled from a single genome from each of three major glacial refugia reveals support for an unexpected history dominated by recent admixture. Despite the fact that 80% of the genome is affected by admixture during the last glacial cycle, we are able to infer the deeper divergence history of these populations. These inferences are robust to variation in block length, mutation model, and the sampling location of individual genomes within refugia. This combination of de novo assembly and numerical likelihood calculation provides a powerful framework for estimating recent population history that can be applied to any organism without the need for prior genetic resources.

opencc-zeroDec 2012View details →
dryad32/100

Data from: De novo transcriptome assembly for the lobster Homarus americanus and characterization of differential gene expression across nervous system tissues

Background: The American lobster, Homarus americanus, is an important species as an economically valuable fishery, a key member in marine ecosystems, and a well-studied model for central pattern generation, the neural networks that control rhythmic motor patterns. Despite multi-faceted scientific interest in this species, currently our genetic resources for the lobster are limited. In this study, we de novo assemble a transcriptome for Homarus americanus using central nervous system (CNS), muscle, and hybrid neurosecretory tissues and compare gene expression across these tissue types. In particular, we focus our analysis on genes relevant to central pattern generation and the identity of the neurons in a neural network, which is defined by combinations of genes distinguishing the neuronal behavior and phenotype, including ion channels, neurotransmitters, neuromodulators, receptors, transcription factors, and other gene products. Results: Using samples from the central nervous system (brain, abdominal ganglia), abdominal muscle, and heart (cardiac ganglia, pericardial organs, muscle), we used RNA-Seq to characterize gene expression patterns across tissues types. We also compared control tissues with those challenged with the neuropeptide proctolin in vivo. Our transcriptome generated 34,813 transcripts with known protein annotations. Of these, 5,000-10,000 of annotated transcripts were significantly differentially expressed (DE) across tissue types. We found 421 transcripts for ion channels and identified receptors and/or proteins for over 20 different neurotransmitters and neuromodulators. Results indicated tissue-specific expression of select neuromodulator (allostatin, myomodulin, octopamine, nitric oxide) and neurotransmitter (glutamate, acetylcholine) pathways. We also identify differential expression of ion channel families, including kainite family glutamate receptors, inward-rectifying K+ (IRK) channels, and transient receptor potential (TRP) A family channels, across central pattern generating tissues. Conclusions: Our transcriptome-wide profiles of the rhythmic pattern generating abdominal and cardiac nervous systems in Homarus americanus reveal candidates for neuronal features that drive the production of motor output in these systems.

opencc-zeroDec 2015View details →
zenodo32/100

De novo genome assembly of Meloidogyne chitwoodi

<p>A whole genome of <em>M. chitwoodi</em> was <em>de novo</em> assembled by empirically optimizing k-mer sizes&nbsp;</p>

openother-openMar 2016View details →
zenodo32/100

The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications

<div> <div> <p><a href="https://onlinelibrary.wiley.com/doi/10.1111/age.13181">https://onlinelibrary.wiley.com/doi/10.1111/age.13181</a></p> <h1>The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications</h1> <div>&nbsp;</div> <div> <div> <div> <div><a href="https://onlinelibrary.wiley.com/authored-by/Chen/Jianhai">Jianhai Chen</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/Zhong/Jie">Jie Zhong</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/He/Xuefei">Xuefei He</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/Li/Xiaoyu">Xiaoyu Li</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/Ni/Pan">Pan Ni</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/Safner/Toni">Toni Safner</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/%C5%A0prem/Nikica">Nikica &Scaron;prem</a>,&nbsp;<a href="https://onlinelibrary.wiley.com/authored-by/Han/Jianlin">Jianlin Han</a></div> </div> </div> </div> <p>The rapid progress of sequencing technology has greatly facilitated the de novo genome assembly of pig breeds. However, the assembly of the wild boar genome is still lacking, hampering our understanding of chromosomal and genomic evolution during domestication from wild boars into domestic pigs. Here, we sequenced and de novo assembled a European wild boar genome (ASM2165605v1) using the long-range information provided by 10&times; Linked-Reads sequencing. We achieved a high-quality assembly with contig N50 of 26.09 Mb. Additionally, 1.64% of the contigs (222) with lengths from 107.65 kb to 75.36 Mb covered 90.3% of the total genome size of ASM2165605v1 (~2.5 Gb). Mapping analysis revealed that the contigs can fill 24.73% (93/376) of the gaps present in the orthologous regions of the updated pig reference genome (Sscrofa11.1). We further improved the contigs into chromosome level with a reference-assistant scaffolding method. Using the &lsquo;assembly-to-assembly&rsquo; approach, we identified intra-chromosomal large structural variations (SVs, length &gt;1 kb) between ASM2165605v1 and Sscrofa11.1 assemblies. Interestingly, we found that the number of SV events on the X chromosome deviated significantly from the linear models fitting autosomes (<em>R</em><sup>2</sup>&nbsp;&gt;&nbsp;0.64,&nbsp;<em>p</em>&nbsp;&lt;&nbsp;0.001). Specifically, deletions and insertions were deficient on the X chromosome by 66.14 and 58.41% respectively, whereas duplications and inversions were excessive on the X chromosome by 71.96 and 107.61% respectively. We further used the large segmental duplications (SDs, &gt;1&nbsp;kb) events as a proxy to understand the large-scale inter-chromosomal evolution, by resolving parental-derived relationships for SD pairs. We revealed a significant excess of SD movements from the X chromosome to autosomes (<em>p</em>&nbsp;&lt;&nbsp;0.001), consistent with the expectation of meiotic sex chromosome inactivation. Enrichment analyses indicated that the genes within derived SD copies on autosomes were significantly related to biological processes involving nervous system, lipid biosynthesis and sperm motility (<em>p</em>&nbsp;&lt;&nbsp;0.01). Together, our analyses of the de novo assembly of ASM2165605v1 provides insight into the SVs between European wild boar and domestic pig, in addition to the ongoing process of meiotic sex chromosome inactivation in driving inter-chromosomal interaction between the sex chromosome and autosomes.</p> </div> </div> <div>The work has been pulished here: https://onlinelibrary.wiley.com/doi/full/10.1111/age.13181</div> <div>&nbsp;</div> <div>The current dataset include the genome annotation files.</div> <div>&nbsp;</div> <div>For the whole-genomic assembly, please check NCBI:&nbsp;</div> <div>https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_021656055.1/</div> <div> <table> <tbody> <tr> <th>&nbsp;</th> <th>GenBank</th> </tr> </tbody> <tbody> <tr> <td>Genome size</td> <td>2.5 Gb</td> </tr> <tr> <td>Total ungapped length</td> <td>2.4 Gb</td> </tr> <tr> <td>Number of scaffolds</td> <td>12,642</td> </tr> <tr> <td>Scaffold N50</td> <td>28.3 Mb</td> </tr> <tr> <td>Scaffold L50</td> <td>25</td> </tr> <tr> <td>Number of contigs</td> <td>41,323</td> </tr> <tr> <td>Contig N50</td> <td>157.9 kb</td> </tr> <tr> <td>Contig L50</td> <td>4,562</td> </tr> <tr> <td>GC percent</td> <td>42</td> </tr> <tr> <td>Genome coverage</td> <td>56.0x</td> </tr> <tr> <td>Assembly level</td> <td>Scaffold</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2>Assembly methods</h2> <div>Sequencing technology 10xgenomics Assembly method Supernova v. 2.1.1 <p>&nbsp;</p> <p>part_** are genome fasta for the GCA_021656055.1</p> <p>You could use the following to combine and uncompress.</p> </div> </div> <div> <div> <div><code><span>cat</span> part_* &gt; archive_combined.zip </code></div> </div> <div> <div>&nbsp;</div> <div><code>unzip archive_combined.zip</code></div> </div> </div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Data from: De Novo Genome assembly of the Caucasian dwarf goby Knipowitschia cf. caucasica, a new alien Gobiidae invading the River Rhine

<p><strong>Background (Abstract from Paper)</strong></p> <p>The Caucasian dwarf goby <em>Knipowitschia</em> cf. <em>caucasica</em> is a new invasive alien Gobiidae spreading in the<br>Lower Rhine since 2019. Little is known about the invasion biology of the species and further investiga-<br>tions to reconstruct the invasion history are lacking genomic resources. We assembled a high-quality<br>chromosome-scale reference genome of <em>Knipowitschia</em> cf. <em>caucasica</em> by combining PacBio, Omni-C and<br>Illumina technologies. The size of the assembled genome is 956.58 Mb with a N50 scaffold length of 43 Mb,<br>which includes 92.3 % complete vertebrate/Actinopterygii Benchmarking Universal Single-Copy Orthologs.<br>98.96 % of the assembly sequence was assigned to 23 chromosome-level scaffolds, with a GC-content of<br>42.83 %. Repetitive elements account for 53.08 % of the genome. The chromosome-level genome contained<br>49,622 transcripts with 42,926 multi-exons, of which 45,512 genes were functionally annotated. In summary,<br>the high-quality genome assembly provides a fundamental basis to understand the adaptive advantage of<br>the species.<br><br>The file provided here is the <strong>non-redundant repeat library</strong> (2,812 consensus sequences of repeat families).&nbsp;</p> <p><strong>Method:</strong> Repetitive elements were identified de novo with RepeatModeler version 2.0.1. Repetitive DNA and soft-masking was performed with RepeatMasker version 4.1.1 (Smit et al., 2013) using the repeat library previously identified via RepeatModeler and skipping the bacterial insertion element check (-no_is) and run with rmblastn version 2.10.0+ (Flynn et al., 2020).&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

GALA: a computational framework for de novo chromosome-by-chromosome assembly with long reads

<p>C.elegans, O.sativa&nbsp;and Human assemblies produced by GALA software</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Supplementary data for Orciraptor agilis de novo transcriptome assembly study

<p>Supplementary Data for study &quot;Comparative transcriptomics reveals the molecular toolkit used by an algivorous protist for cell wall perforation&quot;</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

De novo assembly and annotation of parasitic trematode genomes

<p>Contained in this release are 19&nbsp;genome assemblies and annotations of parasitic trematodes, encompassing 13 species. This included representatives of the <em>Schistosoma</em> (<em>n</em> = 13 assemblies), <em>Trichobilharzia</em> (<em>n</em> = 2 assemblies), <em>Heterobilharzia americana</em> (<em>n</em> = 2 assemblies) and <em>Dicrocoelium dendriticum </em>(<em>n </em>= 1 assembly). The&nbsp;<em>Schistosoma curassoni</em>&nbsp;assembly has been released previously (10.5281/zenodo.6594833) but a&nbsp;new annotation is included with the original assembly&nbsp;here.&nbsp;</p> <p>These genomes were assembled from a variety of sources including stored parasites from museum collections, established laboratory strains and wild-caught isolates sampled from natural hosts in endemic regions. Using a combination of DNA sequencing approaches, all genomes were assembled into chromosomal-scale scaffolds. This was followed by genome annotation based on short-read RNA sequencing (RNA-seq) and long-read isoform sequencing (Iso-seq) transcriptomic data.</p> <p>Included here are the primary&nbsp;assemblies (representing a non-redundant haploid genome) for each species (*.primary.fa), alternate loci (alternate representations of loci&nbsp;found in a largely haploid assembly; *.haplotypes.fa) and annotations (*.gff3). Metadata for each assembly can be found in the included spreadsheets (metadata.xlsx).&nbsp;</p> <p>This data is part of a pre-publication release. For information on the proper use of pre-publication data shared by the Wellcome Trust Sanger Institute (including details of any publication moratoria), please see https://www.sanger.ac.uk/about/research-policies/open-access-science/.</p> <p>This&nbsp;repository will be updated with a complete list of collaborators/authors prior to publication. Please contact Duncan Berger (db22@sanger.ac.uk) with questions regarding pre-publication use of this dataset.&nbsp;</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record