Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo28/100

Transposase-Assisted Tagmentation: An Economical and Scalable Strategy for Single-Worm Whole-Genome Sequencing (other species)

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Biostatistical aspects of whole genome sequencing studies: pre-processing and quality control

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Sequencing of historical plastid genomes reveal exceptional genetic diversity in rye before the start of systematic breeding

<p>The files found in this repository are sequencing files of historical rye samples used for generating the analyses presented in the Komluski et al. manuscript. The metadata.csv contains sampling sites and taxon information for the respective sequence file.</p>

opencc-by-4.0Jul 2024View details →
zenodo28/100

Figure 1. – Sequence lengths for 7760 in The complete mitochondrial genome of Thymallus thymallus (Linnaeus, 1758) (Actinopterygii, Salmonidae) obtained by long range PCRs and double multiplexing

Figure 1. – Sequence lengths for 7760 nucleotide sequences allowing the assembling of the Thymallus thymallus mitogenome.

opencc-by-4.0Dec 2020View details →
zenodo28/100

supplemental files for Assembly and analysis of sequence from a spring and winter type Camelina sativa by whole genome PacBio HiFi technologies

<p><span>Supplemental files for Assembly and analysis of sequence from a spring and winter type <em>Camelina sativa</em> by whole genome PacBio HiFi technologies</span></p>

opencc-by-4.0Jan 2024View details →
zenodo28/100

Figure 3 from: Shaoli M, Hao Y, Chao L, Yafu Z, Fuming S, Yuchao W (2018) The complete mitochondrial genome of Xizicus (Haploxizicus) maculatus revealed by next-generation sequencing and phylogenetic implication (Orthoptera, Meconematinae). ZooKeys 773: 57-67. https://doi.org/10.3897/zookeys.773.24156

Figure 3 Phylogenetic reconstruction of Tettigoniidea using mitochondrial PCGs and rRNA concatenated dataset. A Bayesian result, applicable posterior probability values are shown B Maximum likelihood result with applicable bootstrap values shown.

opencc-by-4.0Jul 2018View details →
zenodo28/100

Figure 2 from: Shaoli M, Hao Y, Chao L, Yafu Z, Fuming S, Yuchao W (2018) The complete mitochondrial genome of Xizicus (Haploxizicus) maculatus revealed by next-generation sequencing and phylogenetic implication (Orthoptera, Meconematinae). ZooKeys 773: 57-67. https://doi.org/10.3897/zookeys.773.24156

Figure 2 Relative synonymous codon usage of X. (X.) fascipes, X. (E.) howardi, X. (H.) maculatus mitochondrial protein-coding genes. Condon families are provided on the x-axis.

opencc-by-4.0Jul 2018View details →
zenodo28/100

Sequencing data for "Rapid heuristic inference of antibiotic resistance and susceptibility by genomic neighbor typing"

<p>This repository contains sequencing data from the following paper:</p> <p>Karel Břinda, Alanna Callendrello, Kevin C. Ma, Derek R MacFadden, Themoula Charalampous, Robyn S Lee, Lauren Cowley, Crista B Wadsworth, Yonatan H Grad, Gregory Kucherov, Justin O&rsquo;Grady, Michael Baym, and William P Hanage.&nbsp;Rapid heuristic inference of antibiotic resistance and susceptibility by genomic neighbor typing<strong>.</strong>&nbsp;2019.</p>

opencc-by-nc-4.0Jul 2019View details →
dryad28/100

Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species

Today most population genomic studies of nonmodel organisms either sequence a subset of the genome deeply in each individual or sequence pools of unlabelled individuals. With a step-by-step workflow, we illustrate how low-coverage whole-genome sequencing of hundreds of individually barcoded samples is now a practical alternative strategy for obtaining genomewide data on a population scale. We used a highly efficient protocol to generate high-quality libraries for ~6.5 USD from each of 876 Atlantic silversides (a teleost fish with a genome size ~730 Mb) that we sequenced to 1–4× genome coverage. In the absence of a reference genome, we developed a bioinformatic pipeline for mapping the genomic reads to a de novo assembled reference transcriptome. This provides an 'in silico' method for exome capture that avoids the complexities and expenses of using wet chemistry for target isolation. Using novel tools for analysis of low-coverage data, we extracted population allele frequencies, individual genotype likelihoods and polymorphism data for 2 504 335 SNPs across the exome for the 876 fish. To illustrate the use of the resulting data, we present a preliminary analysis of geographical patterns in the exome data and a comparison of complete mitochondrial genome sequences for each individual (constructed from the low-coverage data) that show population colonization patterns along the US east coast. With a total cost per sample of less than 50 USD (including sequencing) and ability to prepare 96 libraries in only 5 h, our approach adds a viable new option to the population genomics toolbox.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Molecular phylogeny and SNP variation of polar bears (Ursus maritimus), brown bears (U. arctos) and black bears (U. americanus) derived from genome sequences

We assessed the relationships of polar bears (Ursus maritimus), brown bears (U. arctos), and black bears (U. americanus) with high throughput genomic sequencing data with an average coverage of 25X for each species. A total of 1.4 billion 100-bp paired-end reads was assembled using the polar bear and annotated giant panda (Ailuropoda melanoleuca) genome sequences as references. We identified 13.8 million single nucleotide polymorphisms (SNP) in the three species aligned to the polar bear genome. These data indicate that polar bears and brown bears share more SNP with each other than either does with black bears. Concatenation and coalescence-based analysis of consensus sequences of approximately one million base pairs of ultra-conserved elements (UCE) in the nuclear genome resulted in a phylogeny with black bears as the sister group to brown and polar bears, and all brown bears are in a separate clade from polar bears. Genotypes for 162 SNP loci of 336 bears from Alaska and Montana showed that the species are genetically differentiated and there is geographic population structure of brown and black bears but not polar bears.

opencc-zeroDec 2012View details →
dryad28/100

Data from: An evaluation of the hybrid speciation hypothesis for Xiphophorus clemenciae based on whole genome sequences

Once thought rare in animal taxa, hybridization has been increasingly recognized as an important and common force in animal evolution. In the past decade, a number of studies have suggested that hybridization has driven speciation in some animal groups. We investigate the signature of hybridization in the genome of a putative hybrid species, Xiphophorus clemenciae, through whole genome sequencing of this species and its hypothesized progenitors. Based on analysis of this data, we find that X. clemenciae is unlikely to have been derived from admixture between its proposed parental species. However, we find significant evidence for recent gene flow between Xiphophorus species. Though we detect genetic exchange in two pairs of species analyzed, the proportion of genomic regions that can be attributed to hybrid origin is small, suggesting that strong behavioral pre-mating isolation prevents frequent hybridization in Xiphophorus. The direction of gene flow between species supports a role for sexual selection in mediating hybridization.

opencc-zeroDec 2011View details →
dryad28/100

Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis

Top predators are disappearing worldwide, significantly changing ecosystems that depend on top-down regulation. Conflict with humans remains the primary roadblock for large carnivore conservation, but for the eastern wolf (Canis lycaon), disagreement over its evolutionary origins presents a significant barrier to conservation in Canada and has impeded protection for grey wolves (Canis lupus) in the USA. Here, we use 127 235 single-nucleotide polymorphisms (SNPs) identified from restriction-site associated DNA sequencing (RAD-seq) of wolves and coyotes, in combination with genomic simulations, to test hypotheses of hybrid origins of Canis types in eastern North America. A principal components analysis revealed no evidence to support eastern wolves, or any other Canis type, as the product of grey wolf × western coyote hybridization. In contrast, simulations that included eastern wolves as a distinct taxon clarified the hybrid origins of Great Lakes-boreal wolves and eastern coyotes. Our results support the eastern wolf as a distinct genomic cluster in North America and help resolve hybrid origins of Great Lakes wolves and eastern coyotes. The data provide timely information that will shed new light on the debate over wolf conservation in eastern North America.

opencc-zeroDec 2016View details →
dryad28/100

Data from: "Transcriptome sequence identity between Lyme disease tick vectors, Ixodes scapularis and Ixodes ricinus" in Genomic Resources Notes accepted 1 April 2014 to 31 May 2014

Ixodes scapularis and I. ricinus transmit the Lyme disease agent Borrelia burgdorferi in the U.S. and Europe, respectively. The only tick genome sequence available is that of I. scapularis, which constitutes a limitation for tick research. Recent evidences suggest that I. ricinus and I. scapularis transcriptomes share some degree of sequence identity. However, only the global transcriptome comparison reported here demonstrated that I. ricinus and I. scapularis share a 99.232±0.005 percent sequence identity with a very low frequency of INDELs. However, due to limitations of the current I. scapularis genome assembly, the number of aligned reads was only 26-27%. These results support the use of I. scapularis genome sequence as a reference for the analysis of I. ricinus transcriptomics and proteomics data, but addressing the limitations associated with the I. scapularis genome assembly.

opencc-zeroDec 2013View details →
dryad28/100

Neisseria cinerea 346T whole genome sequence

<p>Type VI Secretion Systems (T6SS) are widespread in bacteria and can dictate the development and organisation of polymicrobial ecosystems by mediating contact dependent killing. In Neisseria species, including Neisseria cinerea a commensal of the human respiratory tract, interbacterial contacts are mediated by Type four pili (Tfp) which promote formation of aggregates and govern the spatial dynamics of growing Neisseria microcolonies. Here we show that N. cinerea expresses a plasmid-encoded T6SS that is active and can limit growth of related pathogens. We explored the impact of Tfp expression on N. cinerea T6SS-dependent killing and show that expression of Tfp by prey strains enhances their susceptibility to T6SS, by keeping them in close proximity of T6SS wielding attacker strains. Our findings have important implications for understanding how spatial constraints during contact-dependent antagonism can shape the evolution of microbial communities.</p>

opencc-zeroJul 2021View details →
dryad28/100

Data from: Whole genome sequencing of elite rice cultivars as a comprehensive information resource for marker assisted selection

Current advances in sequencing technologies and bioinformatics revealed the genomic background of rice, a staple food for the poor people, and provided the basis to develop large genomic variation databases for thousands of cultivars. Proper analysis of this massive resource is expected to give novel insights into the structure, function, and evolution of the rice genome, and to aid the development of rice varieties through marker assisted selection or genomic selection. In this work we present sequencing and bioinformatics analyses of 104 rice varieties belonging to the major subspecies of Oryza sativa. We identified repetitive elements and recurrent copy number variation covering about 200 Mbp of the rice genome. Genotyping of over 18 million polymorphic locations within O. sativa allowed us to reconstruct the individual haplotype patterns shaping the genomic background of elite varieties used by farmers throughout the Americas. Based on a reconstruction of the alleles for the gene GBSSI, we could identify novel genetic markers for selection of varieties with high amylose content. We expect that both the analysis methods and the genomic information described here would be of great use for the rice research community and for other groups carrying on similar sequencing efforts in other crops.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Genome sequences reveal cryptic speciation in the human pathogen Histoplasma capsulatum

Histoplasma capsulatum is a pathogenic fungus that causes life-threatening lung infections. About 500,000 people are exposed to H. capsulatum each year in the United States, and over 60% of the U.S. population has been exposed to the fungus at some point in their life. We performed genome-wide population genetics and phylogenetic analyses with 30 Histoplasma isolates representing four recognized areas where histoplasmosis is endemic and show that the Histoplasma genus is composed of at least four species that are genetically isolated and rarely interbreed. Therefore, we propose a taxonomic rearrangement of the genus. IMPORTANCE: The evolutionary processes that give rise to new pathogen lineages are critical to our understanding of how they adapt to new environments and how frequently they exchange genes with each other. The fungal pathogen Histoplasma capsulatum provides opportunities to precisely test hypotheses about the origin of new genetic variation. We find that H. capsulatum is composed of at least four different cryptic species that differ genetically and also in virulence. These results have implications for the epidemiology of histoplasmosis because not all Histoplasma species are equivalent in their geographic range and ability to cause disease.

opencc-zeroDec 2016View details →
zenodo28/100

Figure 1 in May a hybridogenetic complex regenerate the nuclear genome of both sexes of a missing ancestor? First evidence on the occurrence of a nuclear non-hybrid Squalius alburnoides (Cyprinidae) female based on DNA sequencing

Figure 1. Minimum spanning network among cytb haplotypes. The majority of the haplotypes (represented by circles) are exclusive of Squalius alburnoides (in grey) and of S. pyrenaicus (in white) individuals, except for the central one which is a haplotype shared between one S. alburnoides and two S. pyrenaicus individuals. NH indicates the haplotype of the non-hybrid female. The number of mutations between haplotypes is represented by small black dots.

opencc-by-4.0Oct 2006View details →
zenodo28/100

Figure 5 from: Wang P, Yang H, Zhou W, Hwang C, Zhang W, Qian Z (2014) The mitochondrial genome of the land snail Camaena cicatricosa (Müller, 1774) (Stylommatophora, Camaenidae): the first complete sequence in the family Camaenidae. ZooKeys 451: 33-48. https://doi.org/10.3897/zookeys.451.8537

Figure 5 - Phylogenetic tree inferred by maximum likelihood (ML) and maximum parsimony (MP) methods based on 13 protein genes. The tree is rooted with Aplysia californica. Numbers on or under the nodes represent bootstrap values of MP and ML respectively.

opencc-by-4.0Nov 2014View details →
zenodo28/100

Figure 4 from: Wang P, Yang H, Zhou W, Hwang C, Zhang W, Qian Z (2014) The mitochondrial genome of the land snail Camaena cicatricosa (Müller, 1774) (Stylommatophora, Camaenidae): the first complete sequence in the family Camaenidae. ZooKeys 451: 33-48. https://doi.org/10.3897/zookeys.451.8537

Figure 4 - Relative synonymous codon usage (RSCU) in the Camaena cicatricosa mt genome. Codon families are provided on the x axis.

opencc-by-4.0Nov 2014View details →
zenodo28/100

Figure 3 from: Wang P, Yang H, Zhou W, Hwang C, Zhang W, Qian Z (2014) The mitochondrial genome of the land snail Camaena cicatricosa (Müller, 1774) (Stylommatophora, Camaenidae): the first complete sequence in the family Camaenidae. ZooKeys 451: 33-48. https://doi.org/10.3897/zookeys.451.8537

Figure 3 - Inferred secondary structures of 22 tRNA genes in Camaena cicatricosa. Dashed (-) indicates Watson-Crick base pairing and (•) indicates G-U base pairing.

opencc-by-4.0Nov 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record