Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

155

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

155 results for “WGS”

Learn how ShareScore rates datasets ↗
zenodo44/100

Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis

<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>

opencc-zeroJan 2022View details →
zenodo44/100

QC and WGS data on compound heterozygous PRKN-mutant (R275W/dEx8) PD patient

<p>Primary/raw&nbsp;data for QC on the characterisation of an iPSC line from a PD patient (Stem Cell Research)</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

WGS pipeline supporting files

<p>These are supporting files used for our WGS analysis pipeline described in&nbsp;<a href="https://github.com/edg1983/WGS_pipeline">https://github.com/edg1983/WGS_pipeline</a></p> <p>Somalier files are derived from&nbsp;<a href="https://github.com/brentp/somalier">https://github.com/brentp/somalier</a></p> <p>Expansion Hunter files represent the variants catalog from Expansion Hunter v5.0.0 (<a href="https://github.com/Illumina/ExpansionHunter">https://github.com/Illumina/ExpansionHunter</a>)</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Dataset: GeneDx Holdings Corp. (WGS) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

AMP-PD release 4: snRNA-seq and WGS dataset

<p>snRNA-seq dataset from 100 Parkinson&rsquo;s disease cases and controls profiled in 5 postmortem brain regions.&nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Performance and agreement between WGS variant calling pipelines used for bovine tuberculosis control: towards international standardisation

<p>This repository contains the simulated genomes and FASTQ files used for the analyses reported in the publication: <strong>Performance and agreement between WGS variant calling pipelines used for bovine tuberculosis control: towards&nbsp; international standardisation</strong>.</p> <p>This dataset is based on previously published data:&nbsp;<a href="https://doi.org/10.1099/mgen.0.000388">https://doi.org/10.1099/mgen.0.000388</a></p> <p>Processing scripts can be found in: <a href="https://github.com/Viloleal/bTB-pipeline-comparison-data-and-tools">https://github.com/Viloleal/bTB-pipeline-comparison-data-and-tools</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Dataset for WGS of TiLV using Nanopore

<p>This zip file contains scripts, initial fastq files, assembled genomes (public and from this study) as well as bioinformatics intermediate files used for this study.</p> <p>File Structure and Descriptions</p> <p>├── 01.Filter.sh : Script to perform read filtering from raw fastq.gz<br> ├── 02.RefGenome_Assembly.sh : Script to perform reference-mapping genome assembly from the trimmed fastq file<br> ├── 03.Cleanup.sh : Script to reorganize folder and clean-up intermediate files<br> ├── 04.GapAnalysis.sh : Script to extract individual viral segment and perform QUAST analysis to calculate number of gaps<br> ├── 05.Phylogenetic.sh : Script to combine TiLV genome from public database (availabel in Phylogenetic folder) and generate Maximum likelihood tree<br> ├── backup : Raw Fastq (basecalled with Guppy super accuracy mode)<br> ├── Consensus : Assembled genome in fasta format<br> ├── Coverage: Contig coverage information<br> ├── Filtered_Fastq: Quality and length-filtered fastq used for generating the genome assembly<br> ├── Filtered_Segment.fasta: Viral segments from samples that have been filtered for high completeness ( &gt; 80% genome without gap)<br> ├── GapAnalysis.tsv: Table containing the gap information for each assembled viral segment of each sample<br> ├── Normalised_Bam: BAM alignment file used for variant calling<br> ├── Phylogenetic: Contains crucial whole genome sequences of publicly available virus downloaded from NCBI<br> ├── primer-schemes: Contains reference sequence used for reference-based mapping<br> ├── Raw_Bam: Raw BAM alignment file prior to normalisation. Used to estimate the read depth observed for each sample and each viral segment<br> ├── ReadDepth.tsv:&nbsp; Table file showing the read depth of each sample and its viral segment<br> ├── Sequencing_Stat.tsv: Sequencing statistics of samples before and after length/quality filter with NanoFilt<br> └── TILV.tre: Newick file containing the maximum likelihood tree generated from fasttree (-gtr -nt)</p> <p>8 directories, 10 files</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Chap 3 - Appendix 2 - WGS statistics

<p>WGS statistics from Chapter 3 &quot;&quot;</p> <p>Including: the number of scaffolds, assembly length, average scaffold length, largest scaffold, N50, N count, and the number of recovered genes per specimens; and&nbsp;Average Assembly Length, Min length, Max length, and Average Genes recovered.</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

gnomAD SQLite database WGS v4.0

<p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of &lt;100G &nbsp;and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p><p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p><p>gnomAD SQLite database WGS v4.0</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

16S And WGS Feature Tables: Tryptophan Metabolites And Their Predicted Microbial Sources In Fecal Samples Of Healthy Individuals

<p>The zip file contains 16S and WGS feature tables used in the following publication:</p> <p>Tryptophan Metabolites And Their Predicted Microbial Sources In Fecal Samples Of Healthy Individuals<br>Cynthia L. Chappell, Kristi L. Hoffman, Philip L. Lorenzi, Lin Tan, Joseph F. Petrosino, Richard A. Gibbs, Donna M. Muzny, Harsha Doddapaneni, Matthew C. Ross, Vipin K. Menon, Anil Surathu, Sara J. Javornik Cregeen, Anaid G. Reyes, Pablo C. Okhuysen&nbsp;bioRxiv 2023.12.20.572622; doi: https://doi.org/10.1101/2023.12.20.572622</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

Whole genome sequencing (WGS) data from invasive pine sawfly Diprion similis

<p>Biological introductions are unintended "natural experiments" that provide unique insights into evolutionary processes. Invasive phytophagous insects are of particular interest to evolutionary biologists studying adaptation, as introductions often require rapid adaptation to novel host plants. However, adaptive potential of invasive populations may be limited by reduced genetic diversity—a problem known as the "genetic paradox of invasions". One potential solution to this paradox is if there are multiple invasive waves that bolster genetic variation in invasive populations. Evaluating this hypothesis requires characterizing genetic variation and population structure in the invaded range. To this end, we assemble a reference genome and describe patterns of genetic variation in the introduced white pine sawfly, <em>Diprion</em> <em>similis</em>. This species was introduced to North America in 1914, where it has rapidly colonized the thin-needled eastern white pine (<em>Pinus</em> <em>strobus</em>), making it an ideal invasion system for studying adaptation to novel environments. To evaluate evidence of multiple introductions, we generated whole-genome resequencing data for 64 <em>D</em>. <em>similis</em> females sampled across the North American range. Both model-based and model-free clustering analyses supported a single population for North American <em>D</em>. <em>similis</em>. Within this population, we found evidence of isolation-by-distance and a pattern of declining heterozygosity with distance from the hypothesized introduction site. Together, these results support a single-introduction event. We consider implications of these findings for the genetic paradox of invasion and discuss priorities for future research in <em>D</em>. <em>similis</em>, a promising model system for invasion biology.</p>

opencc-zeroApr 2023View details →
zenodo36/100

gnomAD SQLite database WGS v4.1

<div> <div> <div>&nbsp;</div> </div> </div> <div> <p>This package scales the huge gnomAD files (on average ~120G/chrom) to a SQLite database with a size of &lt;100G &nbsp;and allows scientists to look for various variant annotations present in gnomAD (i.e. Allele Count, Depth, Minor Allele Frequency, etc.). (A query containing 300.000 variants takes ~40s.)</p> <p>Find more information on <a href="https://github.com/KalinNonchev/gnomAD_DB">here</a>.</p> <p>gnomAD SQLite database WGS v4.1</p> </div>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Fragment coordinates from shallow WGS of colorectal cancer from patients in Pakistan

<p>BED files indicating fragment coordinates from whole genome sequencing of FFPE tumors and adjacent normal tissue samples. Sequencing data was aligned to hg19 using bwa-mem, prior to conversion to BED format. Tumor samples are indicated with suffix "_T" and normal samples are indicated with suffix "_N"</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Summary statistics of the gene-based collapsing analyses for COVID-19 severity within the DeCOI WGS data set

<p>These are the summary statistics of the gene-based collapsing analyses for COVID-19 severity within the DeCOI Whole-Genome Sequencing data set from the European subcohort (n=1,017).</p> <p>Case - control definitions:<br>Ex (file: DeCOI_EUR_Ex_RVAS.tsv.gz): 272 cases (required mechanical ventilation or died because of COVID-19 = WHO score 6-10); 362 controls (not hospitalized = WHO score 1-3)<br>B1 (file: DeCOI_EUR_B1_RVAS.tsv.gz): 655 cases (at least hospitalized = WHO score 4-10); 362 controls (not hospitalized = WHO score 1-3)</p> <p>Column description:<br>CHROM: Chromosome of the respective gene<br>ID: Column of the format: [Ensemble Gene ID].[mask which was used to select variants].[upper alle frequency cut-off]<br>ALLELE0: The reference allele<br>ALLELE1: The alternative allele and the effect allele - here all variants which pass allele frequency and mask filters (see ID-column)<br>A1FREQ: Total frequency of Allele 1<br>A1FREQ_CASES: Frequency of Allele 1 in cases<br>A1FREQ_CONTROLS: Frequency of Allele 1 in controls<br>N: Count of individuals that were used in the analysis<br>N_CASES: Count of cases that were used in the analysis<br>N_CONTROLS: Count of controls that were used in the analysis<br>TEST: Test-mode of regenie - here an additive model was used<br>BETA: Estimated effect size given as beta<br>SE: Standard error of BETA<br>LOG10P: Negative decadic logarithm of the p-value</p> <p>Method:<br>Please refer to our accompanying manuscript for a detailed description. We conducted gene-based collapsing analysis using regenie 3.2.4 in logistic regression mode (without step 1). Covariates used were sex, age, age*age, age*sex and the first 10 principal components derived from common variants.</p> <p>Contact for further information:<br>dac_decoi_hostgenetics@listen.uni-bonn.de</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

QC and WGS around the breakpoints of deletions in the compound heterozygous PRKN-deficient PD iPSC line FINi006-A (FI.CS.PRKNDex2/Dex5-7.@40)

<p>BAM files from WGS on the regions of deletions in both <em>PRKN</em> gene alleles of the iPSC line (clone 18) derived using Sendai virus from PRKN 09/090 patient's fibroblasts&nbsp;&nbsp;</p>

opencc-by-4.0Aug 2024View details →
dryad36/100

Alignments for probes, raw WGS reads, and WGS assemblies

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad36/100

Whole genome sequencing (WGS) data from invasive pine sawfly Diprion similis

Open the record for dataset details and reuse information.

publicApr 2023View details →
dryad36/100

The consolidated and reconciled annotations from all of the WGS strains used in this study

Open the record for dataset details and reuse information.

publicSep 2024View details →
zenodo32/100

Neisseria gonorrhoeae clustering to reveal major European WGS-based genogroups in association with antimicrobial resistance (cgMLST and MScgMLST schemas, allelic profile matrices and GrapeTree input file)

<p>This dataset refers to the gene-by-gene analysis of 3791 <em>Neisseria gonorrhoeae</em>&nbsp;genomes from 21 European countries and&nbsp;includes the used cgMLST and MScgMLST loci schemas prepared for the chewBBACA core suite, as well as the associated allelic profile matrices for all genomes. Additionally a&nbsp;<em>.json</em> file is made available for direct input in the GrapTree vizualization software for data/metadata exploration.&nbsp;</p> <p>All novel raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB36482). Additional raw sequence read data used were retrieved from the following ENA BioProjects:&nbsp;PRJEB14933; PRJEB2124; PRJEB23008; PRJEB26560; PRJEB9227; PRJNA275092; PRJNA348107; PRJNA473385; PRJNA315363.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Data for bioinformatics practical course (TP) 1 - Whole Genome Sequencing (WGS)

<p>This data are fastq (.fq) files for the Whole Genome Sequencing (WGS) TP1 (09/12/2020) of the Master Infectiologie-Vaccinologie (University of Tours)</p>

opencc-by-4.0Nov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record