Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
130
datasets available to search
ShareScore release 0.7.1
Dataset results
130 results for “whole-genome sequencing”
EGP Mitochondrial Genome Analysis on Gambian Genome Variation Project Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GGVP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Gambian Genome Variation Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>d21e1e91e8b4c00627171fae79a1f54d</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>b359d1068d4f84f7746d1ebde82df29a</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>ee2b93aa93d2177d92ec0f8f308b43ed</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Copy Number</td> <td>fda509ba1d2bf33fd2d6b77b92e76c03</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>
EGP Mitochondrial Genome Analysis on GIAB Whole-Genome Sequencing Data
<div> <p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GIAB. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on GIAB</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://raw.githubusercontent.com/genome-in-a-bottle/giab_data_indexes/refs/heads/master/AshkenazimTrio/alignment.index.AJtrio_Illumina300X_wgs_novoalign_GRCh37_GRCh38_NHGRI_07282015</code></p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>5eac6ec7d36307aa401fd5441b38a506</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>1151ae74c8e515f4f39bef816bb55d6a</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Variant Tables</td> <td>b3e342fe9827df2e399f5685f84cd4dc</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Copy Number</td> <td>3c45c19f76f71b3ccaad155565ead5e4</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div> <p> </p> <p> </p> </div> <h2> </h2>
Supplementary dataset to publication: Whole-genome sequencing of Streptococcus uberis isolated from cows with mastitis in Thuringia
<p><span><strong>Introduction</strong>.</span> <em><span>Streptococcus uberis</span></em> is a common cause of mastitis in cattle, leading to significant economic losses. The widespread use of antimicrobials has contributed to the emergence of resistance, which poses a severe challenge in controlling <em><span>S. uberis</span></em> infection.</p> <p><span><strong>Aim</strong>.</span> The objective of this study was to gain insights into the antimicrobial resistance (AMR) and epidemiological typing of <em><span>S. uberis</span></em> isolated from milk collected from bovine mastitis on dairy farms in Thuringia.</p> <p><span><strong>Methodology</strong>.</span> In this study, 84 <em><span>S. uberis</span></em> isolates were obtained from cattle with clinical mastitis in Thuringia, their phenotypic and genotypic AMR were analyzed and their phylogenetic relationship was explored using whole-genome sequencing.</p> <p><span><strong>Results</strong>.</span> Genetically heterogeneous strains were found on the farms, but clusters of highly similar strains also circulated within the same farms. All isolates were sensitive to ampicillin, penicillin, ceftiofur, and vancomycin. However, 42.9%, 42.9%, 22.6%, 19.0%, and 13.0% were resistant to tetracycline, doxycycline, clindamycin, pirlimycin, and erythromycin, respectively. Thirty-nine strains were phenotypically resistant to two or more tested antibiotics. We identified a plasmid associated with macrolide and lincosamide resistance in 12% of the strains.</p> <p><span><strong>Conclusion</strong>.</span> The emergence of <em><span>S. uberis</span></em> strains resistant to multiple antibiotics highlights the importance of <em><span>S. uberis</span></em> surveillance and the prudent use of antimicrobials.</p>
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>
Whole-genome capture and sequencing of Francisella tularensis directly from clinical samples
<p>This dataset comprises:</p> <p><strong>1. The design of RNA oligonucleotide baits for Agilent Technologies’ SureSelect target enrichment - <strong>54756 RNA oligonucleotide "baits" (120 bp each) </strong></strong>designed to perform <strong>whole-genome capture and sequencing of </strong><strong>Francisella tularensis<strong> directly from clinical samples</strong></strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol.</p> <p>RNA oligonucleotide “baits” were designed to span the <em>Francisella tularensis </em>chromosome and plasmid, accounting for the genetic variability among publicly available genome sequences. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 54756 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies.</p> <p><strong>2. Genome assemblies of 17 Francisella tularensis samples generated in the context of validation and application of SureSelect target enrichment for Whole-genome capture and sequencing of Francisella tularensis directly from clinical samples </strong></p> <p>File “<strong>Ft_assembly_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers.</p> <p>The archive “<strong>Ft_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>More details can be found in the following publication: (available soon)</p>
Whole-genome sequencing confirms multiple species of Galapagos giant tortoises
Open the record for dataset details and reuse information.
PacBio whole-genome sequencing and draft assembly of Stentor coeruleus
Open the record for dataset details and reuse information.
A comparison of phylogenomic inference pipelines for low-coverage whole-genome sequencing in Formica ants
Open the record for dataset details and reuse information.
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
Open the record for dataset details and reuse information.
Rapid and Inexpensive Whole-Genome Sequencing of SARS-CoV2 using 1200 bp Tiled Amplicons and Oxford Nanopore Rapid Barcoding
<p>Description of 1200bp amplicon primer sets and .bed and .tsv files for SARS-CoV-2 assembly using the ARTIC bioinformatics pipeline.</p>
Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies
Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.
Data from: Development of diagnostic microsatellite markers from whole-genome sequences of Ammodramus sparrows for assessing admixture in a hybrid zone
Studies of hybridization and introgression and, in particular, the identification of admixed individuals in natural populations benefit from the use of diagnostic genetic markers that reliably differentiate pure species from each other and their hybrid forms. Such diagnostic markers are often infrequent in the genomes of closely related species, and genomewide data facilitate their discovery. We used whole-genome data from Illumina HiSeqS2000 sequencing of two recently diverged (600,000 years) and hybridizing, avian, sister species, the Saltmarsh (Ammodramus caudacutus) and Nelson's (A. nelsoni) Sparrow, to develop a suite of diagnostic markers for high-resolution identification of pure and admixed individuals. We compared the microsatellite repeat regions identified in the genomes of the two species and selected a subset of 37 loci that differed between the species in repeat number. We screened these loci on 12 pure individuals of each species and report on the 34 that successfully amplified. From these, we developed a panel of the 12 most diagnostic loci, which we evaluated on 96 individuals, including individuals from both allopatric populations and sympatric individuals from the hybrid zone. Using simulations, we evaluated the power of the marker panel for accurate assignments of individuals to their appropriate pure species and hybrid genotypic classes (F1, F2, and backcrosses). The markers proved highly informative for species discrimination and had high accuracy for classifying admixed individuals into their genotypic classes. These markers will aid future investigations of introgressive hybridization in this system and aid conservation efforts aimed at monitoring and preserving pure species. Our approach is transferable to other study systems consisting of closely related and incipient species.
Genotype likelihoods for low-coverage whole-genome sequencing data of yellow warblers
<p>The following datasets include the required input files used to empirically test population assignment in WGSassign on Yellow Warbler data. The file "yewa.known.ind105.ds_2x.beagle.gz" includes the filtered variants of 105 Yellow Warbler individuals output as genotype likelihoods and stored in a Beagle-formatted file. The ID file, "yewa.known.ind105.reference.IDs.txt", is a tab-delimited file with 2 columns, the first being the sample ID, and the second being the known reference population. The sample order in the ID file should match that of the input beagle file. To measure the assignment accuracy of WGSassign, we used leave-one-out cross validation using the input beagle file and our ID file.</p>
Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age
We investigated how population changes and fluctuations in the pink-footed goose might have been affected by climatic and anthropogenic factors. First, genomic data confirmed the existence of two separate populations: western (Iceland) and eastern (Svalbard/Denmark). Second, emographic inference suggests that the species survived the last glacial period as a single ancestral population with a low population size (100-1,000 individuals) that split into the current populations at the end of the Last Glacial Maximum with Iceland being the most plausible glacial refuge. While population changes during the last glaciation were clearly environmental, we hypothesize that more recent demographic changes are human-related: (1) the inferred population increase in the Neolithic is due to deforestation to establish new lands for agriculture, increasing available habitat for pink-footed geese (2) the decline inferred during the Middle Ages is due to human persecution and (3) improved protection explains the increasing demographic trends during the 20th century. Our results suggest both environmental (during glacial cycles) and anthropogenic effects (more recent) can be a threat to species survival.
Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".
<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>
BactPrep: A user-friendly whole-genome sequencing analysis platform for the detection of homologous recombination and horizontal gene transfer in bacteria - Sample Dataset
<p>This is the dataset used as the sample dataset for the pipeline BactPrep. This dataset consists of 218 <em>Streptococcus pneumoniae</em> PMEN1 WGS assemblies collected from the year 1984 - 2008 from 22 unique countries globally. The raw sequencing data was originally published in the work: Rapid pneumococcal evolution in response to clinical interventions (doi: 10.1371/journal.ppat.1002745) under the bioproject PRJEB2085.</p> <p>We have assembled the raw sequences records with the following steps: 1) raw reads were first quality checked using fastQC 0.11.9; 2) adapters and low quality reads were removed using Trimmomatic 0.39 with parameter “ILLUMINACLIP:TruSeq2-PE.fa:2:30:10:2:keepBothReads LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36”; 3) trimmed reads were error-corrected and assembled into WGS assemblies using SPAdes 3.15.0 with parameters "--careful --mismatch-correction”.</p>
Low-coverage whole-genome sequencing reveals molecular markers for spawning season and sex identification in Gulf of Maine Atlantic cod (Gadus morhua, Linnaeus 1758)
<p class="CxSpFirst">Atlantic cod (<i>Gadus morhua</i>,<i> </i>Linnaeus 1758) in the western Gulf of Maine are managed as a single stock despite several lines of evidence supporting two spawning groups (spring and winter) that overlap spatially, while exhibiting seasonal spawning isolation. Low-coverage whole genome sequencing was used to evaluate the genomic population structure of Atlantic cod spawning groups in the western Gulf of Maine and Georges Bank using 222 individuals collected over multiple years. Results indicated low total genomic differentiation, while also showing strong differentiation between spring and winter spawning groups at specific regions of the genome. Guided regularized random forest and ranked <i>F</i><sub>ST</sub> methods were used to select panels of single nucleotide polymorphisms (SNPs) that could reliably distinguish spring and winter-spawning Atlantic cod (88.5% assignment rate), as well as males and females (95.0% assignment rate) collected in the western Gulf of Maine. These SNP panels represent a valuable tool for fisheries research and management of Atlantic cod in the western Gulf of Maine that will aid investigations of stock production and support accuracy of future assessments.</p>
CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing
<p><strong>Supplementary dataset for the manuscript:</strong></p> <p><strong>CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing</strong></p> <p> Xionghui Zhou1,*, Haizi Zheng1,*, Hailu Fu1,*, Kelsey L. Dillehay McKillip2-3, Susan M. Pinney2,4, Yaping Liu1-2,5-7 #</p> <p>Affiliations:</p> <p>1 Division of Human Genetics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>2 University of Cincinnati Cancer Center, Cincinnati, OH 45229</p> <p>3 Department of Pathology & Laboratory Medicine, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>4 Department of Environmental and Public Health Sciences, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>5 Division of Biomedical Informatics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229</p> <p>6 Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>7 Department of Electrical Engineering and Computing Sciences, University of Cincinnati College of Engineering and Applied Science, Cincinnati, OH 45229</p> <p>* These authors contributed equally</p> <p># Email: lyping1986@gmail.com</p>
Transposase-Assisted Tagmentation: An Economical and Scalable Strategy for Single-Worm Whole-Genome Sequencing (elegans)
<p><span>AlphaMissense identifies 23 million human missense variants as likely pathogenic, but only 0.1% have been clinically classified. To experimentally validate these predictions, chemical mutagenesis presents a rapid, cost-effective method to produce billions of mutations in model organisms.</span><span> </span><span>However, the prohibitive costs and limitations in the throughput of whole-genome sequencing (WGS) technologies, crucial for variant identification, constrain its widespread application. Here, we introduce a Tn5 transposase-assisted tagmentation </span><span>technique</span><span> for conducting WGS in <em>C. elegans</em>, <em>E. coli</em>, <em>S. cervisiae</em>, and <em>C. reinhardtii</em>. This method, demands merely 2</span><span>0 minutes of hands-on time for a single worm or single-cell clones and incurs a cost below 10 US dollars. <span>It effectively pinpoints causal mutations in mutants defective in cilia or neurotransmitter secretion and in mutants synthetically sterile with a variant analogous to the oncogenic BRAF(V600E) mutation. Integrated with chemical mutagenesis, our approach can generate and identify missense variants economically and efficiently, facilitating experimental investigations of missense variants in diverse species.</span></span></p>
Fig 6 in Whole-Genome Optical Mapping and Finished Genome Sequence of Sphingobacterium deserti sp. nov., a New Species Isolated from the Western Desert of China
Fig 6. Venn diagram depicting orthologous groups of predicted proteins encoded in four sphingobacterial genomes. C1: Sphingobacterium spiritivorum ATCC 33300; C2: Sphingobacterium paucimobilis HER1398; C3: Sphingobacterium thalpophilum DSM11723; and C4: Sphingobacterium deserti ZWT. doi:10.1371/journal.pone.0122254.g006
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.