Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo24/100

Full genome viral sequences inform patterns of SARS-CoV-2 spread into and within Israel

<p>Beast2 outputs for the analysis in the submitted manuscript.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
dryad24/100

The whole protein sequences of the nuclear genome of Chrysosplenium sinicum

<p>We used PacBio and Illumina sequencing to de novo assemble the nuclear genome of Chrysosplenium sinicum, and further annotated the whole protein  sequences of the nuclear genome of Chrysosplenium sinicum.</p>

opencc-zeroAug 2020View details →
zenodo24/100

Simulated nucleotide sequences for testing alignment-free genome distance estimates

<p>This repository contains (12&times;500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in <a href="https://riojournal.com/article/36178/">Criscuolo (2019)</a>. Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+&Gamma; evolutionary model).</p> <p>For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> &nbsp; [1] &nbsp; &nbsp; seed value used during simulation,<br> &nbsp; [2] &nbsp; &nbsp; true evolutionary distance <em>d</em> between the two simulated sequences,<br> &nbsp; [3] &nbsp; &nbsp; total number of simulated characters,<br> &nbsp; [4] &nbsp; &nbsp; number of non-indel characters with nucleotide mismatch,<br> &nbsp; [5] &nbsp; &nbsp; number of non-indel characters,<br> &nbsp; [6-9] &nbsp; A, C, G, T frequencies used during simulation,<br> &nbsp; [10-15] &nbsp; GTR parameters used during simulation,<br> &nbsp; [16] &nbsp; &nbsp; &Gamma; distribution parameter used during simulation,<br> &nbsp; [17-18] &nbsp; two simulated sequences with indel events as gaps.</p> <p>Of note, each pair of aligned sequences without gaps can be regenerated using <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree:</p> <pre>(t1:d,t2:0.000);</pre> <p>where <em>d</em> is given in field [2].</p> <p>___</p> <p>Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:<a href="https://doi.org/10.3897/rio.5.e36178">10.3897/rio.5.e36178</a></p>

opencc-by-4.0Sep 2020View details →
zenodo24/100

Figure 3 from: Sun G, Zhao C, Xia T, Wei Q, Yang X, Feng S, Sha W, Zhang H (2020) Sequence and organisation of the mitochondrial genome of Japanese Grosbeak (Eophona personata), and the phylogenetic relationships of Fringillidae. ZooKeys 995: 67-80. https://doi.org/10.3897/zookeys.995.34432

Figure 3 Codon distribution in the mitochondrial genome of Eophona personata.

opencc-by-4.0Nov 2020View details →
zenodo24/100

Figure 2 from: Sun G, Zhao C, Xia T, Wei Q, Yang X, Feng S, Sha W, Zhang H (2020) Sequence and organisation of the mitochondrial genome of Japanese Grosbeak (Eophona personata), and the phylogenetic relationships of Fringillidae. ZooKeys 995: 67-80. https://doi.org/10.3897/zookeys.995.34432

Figure 2 Predicted secondary structures for the 22 tRNAs in Eophona personata.

opencc-by-4.0Nov 2020View details →
zenodo24/100

Figure 4 from: Sun G, Zhao C, Xia T, Wei Q, Yang X, Feng S, Sha W, Zhang H (2020) Sequence and organisation of the mitochondrial genome of Japanese Grosbeak (Eophona personata), and the phylogenetic relationships of Fringillidae. ZooKeys 995: 67-80. https://doi.org/10.3897/zookeys.995.34432

Figure 4 Mitochondrial gene order and arrangement in Eophona personata.

opencc-by-4.0Nov 2020View details →
zenodo24/100

Figure 3 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 3 Evolutionary rates of the Lepus yarkandensis mitogenome by Ka/Ks.

opencc-by-4.0Feb 2021View details →
zenodo24/100

Figure 2 from: Shan W, Tursun M, Zhou S, Zhang Y, Dai H (2021) Complete mitochondrial genome sequence of Lepus yarkandensis Günther, 1875 (Lagomorpha, Leporidae): characterization and phylogenetic analysis. ZooKeys 1012: 135-150. https://doi.org/10.3897/zookeys.1012.59035

Figure 2 GC and AT skews for mitochondrial PCGs in Lepus yarkandensis.

opencc-by-4.0Feb 2021View details →
dryad24/100

Data from: Whole genome sequence accuracy is improved by replication in a population of mutagenized sorghum.

The accurate detection of induced mutations is critical for both forward and reverse genetics studies. Experimental chemical mutagenesis induces relatively few single base changes per individual. In a complex eukaryotic genome, false positive detection of mutations can occur at or above this mutagenesis rate. We demonstrate here, using a population of ethyl methanesulfonate (EMS) treated Sorghum bicolor BTx623 individuals, that using replication to detect false positive induced variants in next-generation sequencing data permits higher throughput variant detection with greater accuracy. We used a lower sequence coverage depth (average of 7X) from 586 independently mutagenized individuals and detected 5,399,493 homozygous SNPs. Of these, 76% originated from only 57,872 genomic positions prone to false positive variant calling. These positions are characterized by high copy number paralogs where the error-prone SNP positions are at copies containing a variant at the SNP position. The ability of short stretches of homology to generate these error prone positions suggests that incompletely assembled or poorly mapped repeated sequences are one driver of these error prone positions.. Removal of these false positives left 1,275,872 homozygous and 477,531 heterozygous EMS-induced SNPs which, congruent with the mutagenic mechanism of EMS, were greater than 98% G:C to A:T transitions. Through this analysis we generated a database of sequence indexed mutants of Sorghum. This collection contains 4,035 high impact homozygous mutations in 3,637 genes and 56,514 homozygous missense mutations in 23,227 genes. Each line contains, on average, 2,177 annotated homozygous SNPs per genome, including seven likely gene knockouts and 96 missense mutations. The number of mutations in a transcript was linearly correlated with the transcript length and also the G+C count, but not with the GC/AT ratio. Analysis of the detected mutagenized positions identified CG-rich patches, and flanking sequences strongly influenced EMS-induced mutation rates. Our method for detecting false-positive induced mutations is generally applicable to any organism, is independent of the choice of in silico variant-calling algorithm, and is most valuable when the true mutation rate is likely to be low, such as in laboratory induced mutations or somatic mutation detection in medicine.

opencc-zeroDec 2017View details →
dryad24/100

Data from: Whole genome amplification and reduced-representation genome sequencing of Schistosoma japonicum miracidia

Background: In areas where schistosomiasis control programs have been implemented, morbidity and prevalence have been greatly reduced. However, to sustain these reductions and move towards interruption of transmission, new tools for disease surveillance are needed. Genomic methods have the potential to help trace the sources of new infections, and allow us to monitor drug resistance. Large-scale genotyping efforts for schistosome species have been hindered by cost, limited numbers of established target loci, and the small amount of DNA obtained from miracidia, the life stage most readily acquired from humans. Here, we present a method using next generation sequencing to provide high-resolution genomic data from S. japonicum for population-based studies. Methodology/Principal Findings: We applied whole genome amplification followed by double digest restriction site associated DNA sequencing (ddRADseq) to individual S. japonicum miracidia preserved on Whatman FTA cards. We found that we could effectively and consistently survey hundreds of thousands of variants from 10,000 to 30,000 loci from archived miracidia as old as six years. An analysis of variation from eight miracidia obtained from three hosts in two villages in Sichuan showed clear population structuring by village and host even within this limited sample. Conclusions/Significance: This high-resolution sequencing approach yields three orders of magnitude more information than microsatellite genotyping methods that have been employed over the last decade, creating the potential to answer detailed questions about the sources of human infections and to monitor drug resistance. Costs per sample range from $50-$200, depending on the amount of sequence information desired, and we expect these costs can be reduced further given continued reductions in sequencing costs, improvement of protocols, and parallelization. This approach provides new promise for using modern genome-scale sampling to S. japonicum surveillance, and could be applied to other schistosome species and other parasitic helminthes

opencc-zeroDec 2016View details →
dryad24/100

Data from: Genomics of introgression in the Chinese horseshoe bat (Rhinolophus sinicus) revealed by transcriptome sequencing

Recent genomic studies show that introgression can occur at a genome-wide scale among recently diverged lineages. However, introgression is difficult to distinguish from incomplete lineage sorting (ILS), and these processes are expected to occur together. Moreover, ncDNA introgression is less easily detected than mtDNA introgression, and as such its prevalence is less well understood. The Chinese horseshoe bat (Rhinolophus sinicus) occurs as three distinct forms on mainland China: the subspecies R. s. septentrionalis and two parapatric clades of R. s. sinicus (Central and East R. s. sinicus). Previous work suggested widespread mtDNA introgression between these subspecies; however, no ncDNA introgression was detected. In this study we sampled the coding genomes of all three forms of R. sinicus in order to perform a more sensitive test for ncDNA introgression against an expected background of ILS. We assembled 3548 nuclear protein-coding genes from these and three congeneric species, and built a high-confidence species tree using maximum likelihood and Bayesian concordance methods. Phylogenetic analysis suggested a mosaic genome for Central R. s. sinicus derived from R. s. septentrionalis and East R. s. sinicus. Nuclear DNA introgression between Central R. s. sinicus and R. s. septentrionalis was supported by three different tests, whereas ILS could not be ruled out completely. Our findings, in line with other recent results, indicate that recently diverged taxa undergo large-scale secondary introgression, and that this process likely operates alongside ILS to give rise to phylogenomic discordances or even mosaic genomes.

opencc-zeroDec 2016View details →
dryad24/100

Data from: Large-scale ruminant genome sequencing provides insights into their evolution and distinct traits

The ruminants are one of the most successful mammalian lineages, exhibiting morphological and habitat diversity and containing several key livestock species. To better understand their evolution, we generated and analyzed de novo assembled genomes of 44 ruminant species, representing all six Ruminantia families. We used these genomes to create a time-calibrated phylogeny to resolve topological controversies, overcoming the challenges of incomplete lineage sorting. Population dynamic analyses show that population declines commenced between 100,000 and 50,000 years ago, which is concomitant with expansion in human populations. We also reveal genes and regulatory elements that possibly contribute to the evolution of the digestive system, cranial appendages, immune system, metabolism, body size, cursorial locomotion, and dentition of the ruminants.

opencc-zeroDec 2018View details →
zenodo24/100

Figure 3 from: Duan Y-B, Wang Y-J, Zhu D-H, Zeng Y, Wang X-D (2024) Description and mitochondrial genome sequencing of a new species of inquiline gall wasp, Synergus nanlingensis (Hymenoptera, Cynipidae, Synergini), from China. Journal of Hymenoptera Research 97: 105-126. https://doi.org/10.3897/jhr.97.119433

Figure 3 Gall of Synergus nanlingensis Wang &amp; Zeng, 2023, sp. nov. on Castanopsis eyrei Tutch.

opencc-by-4.0Mar 2024View details →
zenodo24/100

Statistical data for "Recent advances and perspectives on whole genome sequencing of insects: A review"

<p>This is the statistical data for "Recent advances and perspectives on whole genome sequencing of insects: A review", and derived from NCBI-Genome database.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo24/100

Figure 1 from: Shaoli M, Hao Y, Chao L, Yafu Z, Fuming S, Yuchao W (2018) The complete mitochondrial genome of Xizicus (Haploxizicus) maculatus revealed by next-generation sequencing and phylogenetic implication (Orthoptera, Meconematinae). ZooKeys 773: 57-67. https://doi.org/10.3897/zookeys.773.24156

Figure 1 Circular visualization of the mitogenome of Xizicus (Haploxizicus) maculatus.

opencc-by-4.0Jul 2018View details →
dryad24/100

Data from: Genomic patterns of introgression in rainbow and westslope cutthroat trout illuminated by overlapping paired-end RAD sequencing

Rapid and inexpensive methods for genomewide single nucleotide polymorphism (SNP) discovery and genotyping are urgently needed for population management and conservation. In hybridized populations, genomic techniques that can identify and genotype thousands of species-diagnostic markers would allow precise estimates of population- and individual-level admixture as well as identification of 'super invasive' alleles, which show elevated rates of introgression above the genomewide background (likely due to natural selection). Techniques like restriction-site-associated DNA (RAD) sequencing can discover and genotype large numbers of SNPs, but they have been limited by the length of continuous sequence data they produce with Illumina short-read sequencing. We present a novel approach, overlapping paired-end RAD sequencing, to generate RAD contigs of &gt;300–400 bp. These contigs provide sufficient flanking sequence for design of high-throughput SNP genotyping arrays and strict filtering to identify duplicate paralogous loci. We applied this approach in five populations of native westslope cutthroat trout that previously showed varying (low) levels of admixture from introduced rainbow trout (RBT). We produced 77 141 RAD contigs and used these data to filter and genotype 3180 previously identified species-diagnostic SNP loci. Our population-level and individual-level estimates of admixture were generally consistent with previous microsatellite-based estimates from the same individuals. However, we observed slightly lower admixture estimates from genomewide markers, which might result from natural selection against certain genome regions, different genomic locations for microsatellites vs. RAD-derived SNPs and/or sampling error from the small number of microsatellite loci (n = 7). We also identified candidate adaptive super invasive alleles from RBT that had excessively high admixture proportions in hybridized cutthroat trout populations.

opencc-zeroDec 2012View details →
dryad24/100

Data from: Obtaining mtDNA genomes from next-generation transcriptome sequencing: a case study on the basal Passerida (Aves: Passeriformes) phylogeny.

Classically, the mitochondrial genome is sequenced by a series of amplicons using conserved PCR primers. Here we show how shot-gun transcriptome sequencing can be used to obtain the complete set of protein-coding genes from the mtDNA of four passerine bird species. With these sequences, we address the still unresolved basal Passerida relationships (Aves: Passeriformes). Our analysis suggests a new hypothesis for the basal relationships of Passerida, namely a clade grouping Sylvioidea and Passeroidea, with Paridae and Muscicapidae as successive sister groups to this clade. This study demonstrates the usefulness of next-generation sequencing transcriptome sequencing for obtaining new mtDNA genomes.

opencc-zeroDec 2012View details →
zenodo24/100

Figure 2 from: Wang P, Yang H, Zhou W, Hwang C, Zhang W, Qian Z (2014) The mitochondrial genome of the land snail Camaena cicatricosa (Müller, 1774) (Stylommatophora, Camaenidae): the first complete sequence in the family Camaenidae. ZooKeys 451: 33-48. https://doi.org/10.3897/zookeys.451.8537

Figure 2 - Gene arrangement of nine mt genomes in the order Stylommatophora.

opencc-by-4.0Nov 2014View details →
zenodo24/100

Genome Wide Association Study on Reproductive Traits Using Imputation-Based Whole-Genome Sequence Data in Yorkshire Pigs

<p>These are the supplementary files of the article &quot;Genome Wide Association Study on Reproductive Traits Using Imputation-Based Whole-Genome Sequence Data in Yorkshire Pigs&quot;.</p>

opencc-by-4.0Mar 2023View details →
zenodo24/100

Sequencing Genome in a Bottle samples

<p>We are pleased to announce the release of a new addition to the Oxford Nanopore Open Data project: sequencing of several Genome in a Bottle samples (including the Ashkenazi Trio). Sequencing was performed with the 5 kHz upgrade to the&nbsp;<a href="https://store.nanoporetech.com/uk/ligation-sequencing-kit-v14.html">Ligation Sequencing Kit V14</a>&nbsp;released in MinKNOW 23.04.05. As such the quality of data presented here should be representative of routine sequencing that can be performed by any lab using this latest release.</p> <p>These reference samples were sequenced with two PromethION flow cells each to yield around more than 200 Gbases sequencing data per sample.</p> <blockquote> <p>The following cell line samples were obtained from the NIGMS Human Genetic Cell Repository at the Coriell Institute for Medical Research: GM12878, GM24143, GM24149, GM24385</p> </blockquote> <p><em>Data location</em></p> <p>As with previous releases the new dataset is available for anonymous download from an Amazon Web Services S3 bucket. The bucket is part of the&nbsp;<a href="https://aws.amazon.com/opendata/">Open Data on AWS</a>&nbsp;project enabling sharing and analysis of a wide range of data.</p> <p>The data is located in the bucket at:</p> <blockquote> <pre>s3://ont-open-data/giab_2023.05/</pre> </blockquote> <p>See the&nbsp;<a href="https://labs.epi2me.io/tutorials/">tutorials</a>&nbsp;page for information on downloading the dataset.</p> <p><em>Sequencing Outputs</em></p> <p>Two flowcells were used to sequence each of the samples to high depth:</p> <table> <thead> <tr> <th>Genome</th> <th>Description</th> <th>Cell line</th> </tr> </thead> <tbody> <tr> <td>HG001</td> <td>CEPH/UTAH</td> <td>GM12878</td> </tr> <tr> <td>HG002</td> <td>PGP Ashkenazi Son</td> <td>GM24385</td> </tr> <tr> <td>HG003</td> <td>PGP Ashkenazi Father</td> <td>GM24149</td> </tr> <tr> <td>HG004</td> <td>PGP Ashkenazi Mother</td> <td>GM24143</td> </tr> </tbody> </table> <p>For each flowcell used in the sequencing the PromethION device outputs are available. All data is present as&nbsp;<code>.pod</code>&nbsp;files, along with associated summary files in a structured fashion.</p> <p><em>Further Information</em></p> <ul> <li><em>https://labs.epi2me.io/giab-2023.05</em></li> </ul>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record