Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

484

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

484 results for “next-generation sequencing”

Learn how ShareScore rates datasets ↗
zenodo44/100

SNP and indel discovery and genotyping in next-generation sequencing data

<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels &gt;50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo44/100

Structural variant discovery and genotyping in next-generation sequencing data

<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels &gt;50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>

opencc-by-4.0Oct 2016View details →
zenodo44/100

Genotyping of European Toxoplasma gondii strains by a new high-resolution next-generation sequencing-based method

<p>The data set comprises 164&nbsp;FASTQ&nbsp;files generated with an Ion AmpliSeq-based genotyping method for <em>Toxoplasma gondii </em>and<em>&nbsp;</em>a BED file used for the design of the Ion AmpliSeq primer panel. The FASTA&nbsp;file named as "AmpliSeq-ME49-Reference" was used as a reference for mapping and data analysis of the FASTQ files. The GZ&nbsp;file named as "Tgondii_IonAmpliSeq_Results_SNPs_VCF" is a VCF file, which contains&nbsp;all SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference. The VCF&nbsp;file was converted into a FASTA file named as "Tgondii_IonAmpliSeq_Results_SNPs", which also contains the&nbsp;SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference.</p> <p>The work is published in the European Journal of Clinical Microbiology &amp; Infectious Diseases with the title "Genotyping of European <em>Toxoplasma gondii</em> strains by a new high‑resolution next‑generation sequencing‑based method";&nbsp;https://doi.org/10.1007/s10096-023-04721-7</p>

openAug 2023View details →
zenodo40/100

Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'

<p>This file contains  the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1  and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>

opencc-zeroJan 2016View details →
zenodo40/100

Graphing and tabulating next-generation sequencing and genotyping data

<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Read-mapping for next-generation sequencing data (Drosophila melanogaster)

<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Read-mapping for next-generation sequencing data (Wolbachia)

<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Genotype reproducibility testing in next-generation sequencing data

<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas

<p>Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas</p> <p>Pilocytic astrocytoma (PA) is a benign grade 1 glioma according to the World Health Organization (WHO), common in children but rare in adults, where it may have a worse prognosis. Pediatric PA is usually associated with dysregulation of the MAPK pathway, often involving BRAF alterations such as the KIAA1549::BRAF (K-B) fusion or the V600E mutation. This dataset contains molecular data of 28 cases of adult PA obtained by using gene-targeted next-generation sequencing (NGS).</p>

openNov 2024View details →
zenodo40/100

Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)

<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Next-generation sequencing of AAV.CAP-Mac from Chuapoco et al. (2023) Nature Nanotechnology

<p>Dataset of next-generation sequencing of enrichment of AAV.CAP-Mac in various tissues from the publication:</p> <p>Chuapoco, M.R., Flytzanis, N.C., Goeden, N.&nbsp;<em>et al.</em>&nbsp;Adeno-associated viral vectors for functional intravenous gene transfer throughout the non-human primate brain.&nbsp;<em>Nat. Nanotechnol.</em>&nbsp;(2023). https://doi.org/10.1038/s41565-023-01419-x</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Comparative assessment of line-probe assays and targeted next-generation sequencing in drug-resistant tuberculosis diagnosis

Open the record for dataset details and reuse information.

publicAug 2025View details →
zenodo36/100

Dataset used for "Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing data"

<p>Each simulated dataset was generated using the AbSim R package (version 0.2.6) in a B cell single-lineage fashion. Each B cell clone simulation begins with a random selection from sets of IGHV, IGHD, and IGHJ germline sequences to produce a unique V(D)J recombination event. Then, clones are made by introducing mutations using a local nucleotide context-dependent model (S5F model) along a phylogenetic tree in which branching events occur stochastically.&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Next-generation sequencing of newborn screening genes: The accuracy of short-read mapping

<p>We examine the effect of high homology genomic regions on the mapping of genes related to newborn screening while taking different read lengths and patient&#39;s ethnic background into consideration.</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

CusVarDB: A tool for building customized sample-specific variant protein database from Next-generation sequencing datasets

<p>CusVarDB is a windows based tool for creating a variant protein database from Next-generation sequencing datasets. The program supports variant calling for Genome, RNA-Seq and exome datasets.</p> <p>This repository will provide the resultant variant peptides identified in our study and its corresponding information. The detailed information of the table is given below.</p> <p>Supplementary Table 1. This table contains the resultant variant peptides along with its wild-type peptides from BT474, MDMAB157, MFM223, and HCC38 datasets. Along with mutant peptides, this section also provides additional information such as peptide-spectrum match (PSM), Protein accession, cross-correlation value from the search (Xcorr), and retention time (RT).</p> <p>Supplementary Table 2. This table provides the complete details of the resultant peptides. Here the mutant and corresponding wild-type peptides are mentioned in different sheets. For a given mutant peptide its wild-type peptide and corresponding information can be mapped using the VLOOKUP function in Excel by keeping column A (Sl.No) as lookup parameter.</p> <p>Supplementary Table 3. This table briefs about the variants which are already reported in other cancers.</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo36/100

Supplementary material Next-generation sequencing reveals that miR-16-5p, miR-19a-3p, miR-451a and miR-25-3p cargo in plasma extracellular vesicles differentiates sedentary young males from athletes.

<p>A sedentary lifestyle is a leading risk factor for global mortality. No objective molecular biomarker of sedentarism is available. Extracellular vesicles miRNAs have been described to respond to exercise. Our aim was to identify the extracellular vesicle miRNA profile of chronically trained young male athletes, endurance and resistance, compared to their sedentary counterparts.&nbsp;A descriptive case-control design with 16 sedentary young men, 16 Olympic male endurance athletes and 16 Olympic male resistance athletes.&nbsp;Next Generation Sequencing and RT-qPCR, external and internal validation, were performed in order to analysed extracellular vesicle miRNA profiles.&nbsp;miR-16-5p, miR-19a-3p and miR-451a were significantly upregulated in SED compared to END and RES. Besides, miR-25-3p was specifically down-regulated in END compared to SED. Extracellular vesicle miR-16-5p, miR-19a-3p, miR-451a provide an objective signature of sedentarism irrespective of the type of exercise and miR-25-3p as a specific responder to endurance training. Therefore, this study provides for the first time an objective measure to categorise individuals as sedentary or trained in young male population and moreover, it highlights a common epigenetic modulation between models of training.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".

<p>Original hepatitis C virus hypervariable region 1 NGS sequences&nbsp;in fastq format from patients analyzed in the study&nbsp; &quot;Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center&quot;.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
dryad36/100

Sequences of bacterial and fungal communities by Next-Generation Sequencing (NGS) associated to wall patinas

<p>These data are the DNA sequences obtained from samples of house plaster with chromatic alteration facies. Three areas (VA1, VA3 and VNA6) of the wall were sampled. the DNA was extracted using the Spin Kit For Soil MPBio. </p> <p>The DNA extracts were sequenced by NGS using Illumina MiSeq</p> <p>Two regions were selected V3V5 16S for Bacteria and ITS1 ITS for Fungi.  </p>

opencc-zeroFeb 2023View details →
dryad36/100

A next-generation sequencing study of arthropods in the diet of Laysan Teal (Anas laysanensis)

<p><span>The critically endangered Laysan Teal <em>Anas laysanensis</em> (known as koloa pōhaka in the Hawaiian language) in the Northwestern Hawaiian Islands has wild populations on Kamole (Laysan Island), Kuaihelani (Midway Atoll NWR), and Hōlanikū (Kure Atoll). The Laysan Teal faces a new risk on Sand Island, Kuaihelani: non-target poisoning via a pending House Mouse <em>Mus musculus</em> eradication. After mice were observed attacking and depredating Laysan Albatross <em>Phoebastria immutabilis</em> (mōlī) in 2015, plans to eradicate mice were developed to protect this seabird species. However, this approach risks poisoning the Laysan Teal. To reduce exposure, teal will be translocated during mouse eradication. Even so, there is a potential risk of secondary poisoning for teal by ingesting arthropods that feed on mouse bait. We therefore used next-generation sequencing (NGS) to identify which arthropods teal consume. From August 2019 to February 2020, we collected 71 fresh teal faecal samples on Sand Island, and successfully extracted DNA from 21 samples. Via NGS, we found that teal most frequently consume cockroaches (order: Blattodea), freshwater ostracods (Cyprididae), midges (Chironomidae), and isopods (Porcellionidae). To a lesser degree, teal also eat spiders (Araneae), moths (Lepidoptera), beetles (Coleoptera), springtails (Entomobryomorpha), thrips (Thysanoptera), and crabs (Decapoda). Notably, Sand Island's teal consume entirely different arthropods from teal on Kamole, which mainly eat flies (Diptera) and brine shrimp (Anostraca, <em>Artemia </em>sp.). Our study serves as a model for risk mitigation during invasive rodent eradications.</span></p>

opencc-zeroJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record