Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Supplementary Material for Ph.D. thesis: "Development of a data-intensive centralized system for surveillance and outbreak investigation of bacterial pathogens using whole-genome sequencing""
<p>Supplementary material for Ph.D. thesis.</p>
Data from: Diagnostic yield in epileptic encephalopathies is improved by genome sequencing and re-analysis
<p><b>Objective:</b> To assess the benefits and limitations of whole genome sequencing (WGS) compared to exome sequencing (ES) or multigene panel (MGP) in the molecular diagnosis of developmental and epileptic encephalopathies (DEE).</p> <p><b>Methods: </b>We performed WGS of 30 comprehensively phenotyped DEE patient trios that were undiagnosed after first-tier testing, including chromosomal microarray (CMA), and either research ES (n=15) or diagnostic MGP (n=15).</p> <p><b>Results</b>: 8 diagnoses were made in the 15 individuals who received prior ES (53%): 3 individuals had complex structural variants; 5 had ES-detectable variants which now had additional evidence for pathogenicity. 11 diagnoses were made in the 15 MGP-negative individuals (68%); the majority (n=10) involved genes not included in the panel, particularly in individuals with post-neonatal onset of seizures and those with more complex presentations including movement disorders, dysmorphic features and/or multi-organ involvement. 42% of diagnoses were autosomal recessive or X-chromosome linked.</p> <p><span><span><b>Conclusion:</b> WGS was able to improve diagnostic yield over ES primarily through the detection of complex structural variants (n=3). The higher diagnostic yield was otherwise better attributed to the power of re-analysis rather than inherent advantages of the WGS platform. Additional research is required to assist in the assessment of pathogenicity of novel non-coding and complex structural variants and further improve diagnostic yield for patients with DEE and other neurogenetic disorders.</span></span></p>
DNA sequence data use in phylogenetic analysis of eastern North American stitchworts
<p>Generic delimitation in Caryophyllaceae has been a challenge, and has been informed most recently by use of molecular phylogenetic data. In this study, analysis of 29 samples from the small segregate <em>Mononeuria</em> using nuclear ITS and plastid <em>rps16</em> data revealed it to be polyphyletic. The type species <em>Mononeuria patula</em> as well as two others (<em>M. muscorum</em> and <em>M. paludicola</em>) were shown to belong to the <em>Sabulina</em> clade. The remaining species formed a clade that also included the previously monotypic <em>Geocarpon</em> and was sister to a heterogeneous group that included the Hawaiian <em>Schiedea</em> and three other monotypic genera, <em>Honckenya</em>, <em>Wilhelmsia</em>, and <em>Triplateia</em>. Although several nomenclatural options are available, we propose to place the species from this clade into a single genus, <em>Geocarpon</em>, which basically follows the most recent treatment after exclusion of <em>Sabulina</em> species, but with the necessary new genus placements. New combinations are proposed: <em>Sabulina muscorum</em>, <em>Sabulina paludicola</em>, <em>Geocarpon carolinianum, Geocarpon cumberlandensis</em>, <em>Geocarpon glabrum</em>, <em>Geocarpon groendlandicum</em>, <em>Geocarpon nuttallii</em>, and <em>Geocarpon uniflorum.</em> Analysis of the sequence data revealed remarkable variability among populations of <em>Sabulina</em> (formerly <em>Mononeuria</em>) <em>patula</em>, <em>Sabulina</em> (formerly <em>Mononeuria) paludicola</em>, and <em>Geocarpon </em>(formerly <em>Mononeuria</em>) <em>groenlandicum</em>, suggesting that cryptic species may be present. The data also suggested that broader sampling of <em>Sabulina</em> and <em>Geocarpon</em> could lead to increased understanding of the timing and origins of occupation of calcareous glades and rock outcrop habitats in eastern North America.</p>
DNA sequence data - Bicyclus
<p>Compared to other regions, the drivers of diversification in Africa are poorly understood. We studied a radiation of insects with over 100 species occurring in a wide range of habitats across the Afrotropics to investigate the fundamental evolutionary processes and geological events that generate and maintain patterns of species richness on the continent. By investigating the evolutionary history of <em>Bicyclus</em> butterflies within a phylogenetic framework, we inferred the group's origin at the Oligo-Miocene boundary from ancestors in the Congolian rainforests of central Africa. Abrupt climatic fluctuations during the Miocene (<em>ca.</em> 19–17 Ma) likely fragmented ancestral populations, resulting in at least eight early-divergent lineages. Only one of these lineages appears to have diversified during the drastic climate and biome changes of the early Miocene, radiating into the largest group of extant species. The other seven lineages diversified in forest ecosystems during the late Miocene and Pleistocene when climatic conditions were more favorable—warmer and wetter. Our results suggest changing Neogene climate, uplift of eastern African orogens, and biotic interactions have had different effects on the various subclades of <em>Bicyclus</em>, producing one of the most spectacular butterfly radiations in Africa.</p>
Data from: Genotyping-in-Thousands by sequencing of archival fish scales reveals maintenance of genetic variation following a severe demographic contraction in kokanee salmon
<p>Historical DNA analysis of archival samples has added new dimensions to population genetic studies, enabling spatiotemporal approaches for reconstructing population histories and informing conservation management. Here we tested the efficacy of Genotyping-in-Thousands by sequencing (GT-seq) for collecting targeted single nucleotide polymorphism (SNP) genotypic data from archival scale samples, and demonstrate its application to a study of kokanee salmon (<i>Oncorhynchus nerka</i>) in Kluane National Park and Reserve (KNPR; Yukon, Canada) that underwent a severe 12-year population decline followed by a rapid rebound. We genotyped archival scales sampled pre-crash and contemporary fin clips collected post-crash, revealing high coverage (>90% average genotyping across all individuals) and low genotyping error (<0.01% within-libraries, 0.60% among-libraries) despite the relatively poor quality of recovered DNA. We observed slight decreases in expected heterozygosity, allelic diversity, and effective population size post-crash, but none were significant, suggesting genetic diversity was retained despite the severe demographic contraction. Genotypic data also revealed the genetic distinctiveness of a now extirpated population just outside of KNPR, revealing biodiversity loss at the northern edge of the species distribution. More broadly, we demonstrate GT-seq as a valuable tool for genome-wide data collection from archival samples to address basic questions in ecology and evolution, and inform applied research in wildlife conservation and fisheries management.</p>
Data related to the manuscript "Sequence-specific aggregation of magnetic nanoparticles and single-stranded DNA amplification products for detection of antibiotic resistance gene sul1."
<p>Absorbance and AC susceptometry raw and processed excel files used.</p>
research data supporting "Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis"
<p>This repository contains all the data generated and analyzed including the starting structures, the input files, the trajectory files, the output data from cpptraj and network analyses, and in-house scripts used to prepare the network and module files shown in the paper <strong>"Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis"</strong> published in <strong>Journal of Chemical Information and Modeling</strong> (DOI: 10.1021/acs.jcim.2c00101). </p>
Supplementary material 2 from: Naro-Maciel E, Ingala MR, Werner IE, Reid BN, Fitzgerald AM (2022) COI amplicon sequence data of environmental DNA collected from the Bronx River Estuary, New York City. Metabarcoding and Metagenomics 6: e80139. https://doi.org/10.3897/mbmg.6.80139
Tables S1,S2, Figures S1–S3
Supplementary material 1 from: Naro-Maciel E, Ingala MR, Werner IE, Reid BN, Fitzgerald AM (2022) COI amplicon sequence data of environmental DNA collected from the Bronx River Estuary, New York City. Metabarcoding and Metagenomics 6: e80139. https://doi.org/10.3897/mbmg.6.80139
Supplementary Data Files 1, 2
Data from: Next-generation museum genomics: phylogenetic relationships among palpimanoid spiders using sequence capture techniques (Araneae: Palpimanoidea)
Historical museum specimens are invaluable for morphological and taxonomic research, but typically the DNA is degraded making traditional sequencing techniques difficult to impossible for many specimens. Recent advances in Next-Generation Sequencing, specifically target capture, makes use of short fragment sizes typical of degraded DNA, opening up the possibilities for gathering genomic data from museum specimens. This study uses museum specimens and recent target capture sequencing techniques to sequence both Ultra-Conserved Elements (UCE) and exonic regions for lineages that span the modern spiders, Araneomorphae, with a focus on Palpimanoidea. While many previous studies have used target capture techniques on dried museum specimens (for example, skins, pinned insects), this study includes specimens that were collected over the last two decades and stored in 70% ethanol at room temperature. Our findings support the utility of target capture methods for examining deep relationships within Araneomorphae: sequences from both UCE and exonic loci were important for resolving relationships; a monophyletic Palpimanoidea was recovered in many analyses and there was strong support for family and generic-level palpimanoid relationships. Ancestral character state reconstructions reveal that the highly modified carapace observed in mecysmaucheniids and archaeids has evolved independently.
Data from: DNA and RNA-sequence based GWAS highlights membrane-transport genes as key modulators of milk lactose content
Lactose provides an easily-digested energy source for neonate mammals, and is the primary carbohydrate in milk. Lactose is also a key component of many human food products, though compared to analyses of other milk components, the genetic control of lactose has been little studied. Here we present the first GWAS of milk lactose concentration and yield, investigated in a population of 12,000 taurine dairy cattle. We detail 27 QTL spanning these traits, and subsequently validate the effects of 26 of these loci in a separate population of 18,000 cows. We next present data implicating causative genes and variants for these QTL. Fine mapping of these regions using imputed, whole genome sequence-resolution genotypes reveals protein-coding candidate causative variants affecting the ABCG2, DGAT1, STAT5B, KCNH4, NPFFR2 and RNF214 genes. Eleven of the remaining QTL appear to be driven by regulatory effects, suggested by the presence of co-locating, co-segregating eQTL discovered using mammary RNA sequence data representing a population of 357 lactating cows. Pathway analysis of genes representing all lactose-associated loci shows significant enrichment of genes located to the endoplasmic reticulum, with functions related to ion channel activity mediated through the LRRC8C, P2RX4, KCNJ2 and ANKH genes. Together, these findings highlight novel candidate genes and variants involved in milk lactose regulation, whose impacts on facilitated and active membrane transport mechanisms reinforce the key osmo-regulatory roles of lactose in milk.
Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species
The ribosomal RNA encapsulates a wealth of evolutionary information, including genetic variation that can be used to discriminate between organisms at a wide range of taxonomic levels. For example, the prokaryotic 16S rDNA sequence is very widely used both in phylogenetic studies and as a marker in metagenomic surveys and the ITS region, frequently used in plant phylogenetics, is now recognised as a fungal DNA barcode. However, this widespread use does not escape criticism, principally due to issues such as difficulties in classification of paralogous versus orthologous rDNA units and intragenomic variation, both of which may be significant barriers to accurate phylogenetic inference. We recently analysed datasets from the Saccharomyces Genome Resequencing Project, characterising rDNA sequence variation within multiple strains of the baker's yeast <i>Saccharomyces cerevisiae</i> and its nearest wild relative <i>Saccharomyces paradoxus</i> in unprecedented detail. Notably, both species possess single locus rDNA systems. Here, we use these new variation datasets to assess whether a more detailed characterisation of the rDNA locus can alleviate the second of these phylogenetic issues, sequence heterogeneity, while controlling for the first. We demonstrate that a strong phylogenetic signal exists within both datasets and illustrate how they can be used, with existing methodology, to estimate intra-species phylogenies of yeast strains consistent with those derived from whole-genome approaches. We also describe the use of partial Single Nucleotide Polymorphisms, a type of sequence variation found only in repetitive genomic regions, in identifying key evolutionary features such as genome hybridisation events and show their consistency with whole-genome Structure analyses. We conclude that our approach can transform rDNA sequence heterogeneity from a problem to a useful source of evolutionary information, enabling the estimation of highly accurate phylogenies of closely related organisms, and discuss how it could be extended to future studies of multi-locus rDNA systems.
Data from: "Transcriptome sequences for Campanula gentilis" in Genomic Resources Notes accepted 1 April 2015 – 31 May 2015
In this report, we present the transcriptome of a single accession of Campanula gentilis Kovanda, obtained through the sequencing of both a normalized and a non-normalized cDNA library generated from stem and leaf tissue. The resources we provide include the raw sequence reads, the assembled contigs, the putative open reading frames, the contig/ORF annotations and the normalized as well as non-normalized expression levels.
Data from: "RAD Sequencing for SNP Discovery in Two Populations of Bighorn Sheep (Ovis canadensis)" in Genomic Resources Notes accepted 1 April 2013 - 31 May 2013
In this work we present the development of a large set of single nucleotide polymorphisms (SNPs) discovered in two populations of bighorn sheep (Ovis canadensis). To do so we used restriction-site associated DNA (RAD) sequencing of four individuals from each population. Through alignment of reads to the domestic sheep (Ovis aries) genome we discovered >83,000 SNPs, of which >38,000 are suitable for assays such as an Illumina SNP chip. These loci will allow for fine-mapping of loci associated with horn size, and examination of the consequences of the genetic rescue including mapping genes underling differences in life-history characteristics.
Data from: "Transcriptome sequencing of the Queensland fruit fly, Bactrocera tryoni (Diptera: Tephritidae)" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014
[No abstract filled]
Data from: Characterization of microsatellite loci for the Gulf Coast waterdog (Necturus beyeri) using paired-end Illumina shotgun sequencing and cross-amplification in other Necturus
[No abstract filled]
Data from: Sequence-based association analysis reveals an MGST1 eQTL with pleiotropic effects on bovine milk composition
[No abstract entered]
Data from: Genotyping-in-Thousands by Sequencing panel development and application for high-resolution monitoring of introgressive hybridization within sockeye salmon
<p>Stocking programs have been widely implemented to re-establish extirpated fish species to their historical ranges; when employed in species with complex life histories, such management activities should include careful consideration of resulting hybridization dynamics with resident stocks and corresponding outcomes on recovery initiatives. Genetic monitoring can be instrumental for quantifying the extent of introgression over time, however, conventional markers typically have limited power for the identification of advanced hybrid classes, especially at the intra-specific level. Here, we demonstrate a workflow for developing, evaluating, and deploying a Genotyping-in-Thousands by Sequencing (GT-seq) SNP panel with the power to detect advanced hybrid classes to assess the extent and trajectory of intra-specific hybridization, using the sockeye salmon (<em>Oncorhynchus nerka)</em> stocking program in Skaha Lake, British Columbia, as a case study. Previous analyses detected significant levels of hybridization between the anadromous (sockeye) and freshwater resident (kokanee) forms of <em>O. nerka</em>, but were restricted to assigning individuals to pure-stock or "hybrid". Simulation analyses indicated our GT-seq panel had high accuracy, efficiency and power (> 94.5%) of assignment to pure-stock sockeye salmon/kokanee, F<sub>1</sub>, F<sub>2</sub>, and B<sub>2</sub> backcross-sockeye/kokanee. Re-analysis of 2016/2017 spawners previously analyzed using TaqMan<span> </span>assays and otolith microchemistry revealed shifts in assignment of some hybrids to adjacent pure-stock or B<sub>2</sub>-backcross classes, while new assignment of 2019 spawners revealed hybrids comprised 31% of the population, ~74% of which were B<sub>2</sub>-backcross or F<sub>2</sub>. Overall, the GT-seq panel development workflow presented here could be applied to virtually any system where genetic stock identification and intra-specific hybridization are important management parameters.</p>
Genotype by sequencing data from CHPRRU2, a panel of 282 maize inbred lines.
<p>Genotype by sequencing data from CHPRRU2, a panel of 282 maize inbred lines.</p>
MALDI-TOF MS spectra and sequence data of collagen of modern and archaeological flatfish from European waters
<p>MALDI-TOF MS spectra, LC-MS/MS datafiles, and Mascot MZID files of modern bone collagen of 18 species of Pleuronectiformes as reference spectra that were used to develop peptide biomarkers for ZooMS (Zooarchaeology by Mass Spectrometry). Details on the samples used can be found in the file "Reference spectra information.csv". Further information on the method and results can be found in the manuscript. The file names contain the type of data file and the species name. </p> <p>MALDI-TOF MS of 202 archaeological samples for Zooarchaeology by Mass Spectrometry (ZooMS) from three case study sites from around the North Sea: Barreau Saint-George ferroviaire in northern France, and 16-22 Coppergate and Blue Bridge Lane from York in the United Kingdom. Details on the samples can be found in the supplementary information of the manuscript. Further information on the method can be found in the manuscript. The file names are labeled with the sample ID number and the triplicate number (out of 3).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.