Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
484
datasets available to search
ShareScore release 0.7.1
Dataset results
484 results for “next-generation sequencing”
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Structural variant discovery and genotyping in next-generation sequencing data
<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels >50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Genotyping of European Toxoplasma gondii strains by a new high-resolution next-generation sequencing-based method
<p>The data set comprises 164 FASTQ files generated with an Ion AmpliSeq-based genotyping method for <em>Toxoplasma gondii </em>and<em> </em>a BED file used for the design of the Ion AmpliSeq primer panel. The FASTA file named as "AmpliSeq-ME49-Reference" was used as a reference for mapping and data analysis of the FASTQ files. The GZ file named as "Tgondii_IonAmpliSeq_Results_SNPs_VCF" is a VCF file, which contains all SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference. The VCF file was converted into a FASTA file named as "Tgondii_IonAmpliSeq_Results_SNPs", which also contains the SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference.</p> <p>The work is published in the European Journal of Clinical Microbiology & Infectious Diseases with the title "Genotyping of European <em>Toxoplasma gondii</em> strains by a new high‑resolution next‑generation sequencing‑based method"; https://doi.org/10.1007/s10096-023-04721-7</p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
Graphing and tabulating next-generation sequencing and genotyping data
<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Read-mapping for next-generation sequencing data (Drosophila melanogaster)
<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Read-mapping for next-generation sequencing data (Wolbachia)
<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Genotype reproducibility testing in next-generation sequencing data
<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas
<p>Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas</p> <p>Pilocytic astrocytoma (PA) is a benign grade 1 glioma according to the World Health Organization (WHO), common in children but rare in adults, where it may have a worse prognosis. Pediatric PA is usually associated with dysregulation of the MAPK pathway, often involving BRAF alterations such as the KIAA1549::BRAF (K-B) fusion or the V600E mutation. This dataset contains molecular data of 28 cases of adult PA obtained by using gene-targeted next-generation sequencing (NGS).</p>
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
Next-generation sequencing of AAV.CAP-Mac from Chuapoco et al. (2023) Nature Nanotechnology
<p>Dataset of next-generation sequencing of enrichment of AAV.CAP-Mac in various tissues from the publication:</p> <p>Chuapoco, M.R., Flytzanis, N.C., Goeden, N. <em>et al.</em> Adeno-associated viral vectors for functional intravenous gene transfer throughout the non-human primate brain. <em>Nat. Nanotechnol.</em> (2023). https://doi.org/10.1038/s41565-023-01419-x</p>
Comparative assessment of line-probe assays and targeted next-generation sequencing in drug-resistant tuberculosis diagnosis
Open the record for dataset details and reuse information.
Dataset used for "Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing data"
<p>Each simulated dataset was generated using the AbSim R package (version 0.2.6) in a B cell single-lineage fashion. Each B cell clone simulation begins with a random selection from sets of IGHV, IGHD, and IGHJ germline sequences to produce a unique V(D)J recombination event. Then, clones are made by introducing mutations using a local nucleotide context-dependent model (S5F model) along a phylogenetic tree in which branching events occur stochastically. </p>
Next-generation sequencing of newborn screening genes: The accuracy of short-read mapping
<p>We examine the effect of high homology genomic regions on the mapping of genes related to newborn screening while taking different read lengths and patient's ethnic background into consideration.</p>
CusVarDB: A tool for building customized sample-specific variant protein database from Next-generation sequencing datasets
<p>CusVarDB is a windows based tool for creating a variant protein database from Next-generation sequencing datasets. The program supports variant calling for Genome, RNA-Seq and exome datasets.</p> <p>This repository will provide the resultant variant peptides identified in our study and its corresponding information. The detailed information of the table is given below.</p> <p>Supplementary Table 1. This table contains the resultant variant peptides along with its wild-type peptides from BT474, MDMAB157, MFM223, and HCC38 datasets. Along with mutant peptides, this section also provides additional information such as peptide-spectrum match (PSM), Protein accession, cross-correlation value from the search (Xcorr), and retention time (RT).</p> <p>Supplementary Table 2. This table provides the complete details of the resultant peptides. Here the mutant and corresponding wild-type peptides are mentioned in different sheets. For a given mutant peptide its wild-type peptide and corresponding information can be mapped using the VLOOKUP function in Excel by keeping column A (Sl.No) as lookup parameter.</p> <p>Supplementary Table 3. This table briefs about the variants which are already reported in other cancers.</p>
Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"
Open the record for dataset details and reuse information.
Supplementary material Next-generation sequencing reveals that miR-16-5p, miR-19a-3p, miR-451a and miR-25-3p cargo in plasma extracellular vesicles differentiates sedentary young males from athletes.
<p>A sedentary lifestyle is a leading risk factor for global mortality. No objective molecular biomarker of sedentarism is available. Extracellular vesicles miRNAs have been described to respond to exercise. Our aim was to identify the extracellular vesicle miRNA profile of chronically trained young male athletes, endurance and resistance, compared to their sedentary counterparts. A descriptive case-control design with 16 sedentary young men, 16 Olympic male endurance athletes and 16 Olympic male resistance athletes. Next Generation Sequencing and RT-qPCR, external and internal validation, were performed in order to analysed extracellular vesicle miRNA profiles. miR-16-5p, miR-19a-3p and miR-451a were significantly upregulated in SED compared to END and RES. Besides, miR-25-3p was specifically down-regulated in END compared to SED. Extracellular vesicle miR-16-5p, miR-19a-3p, miR-451a provide an objective signature of sedentarism irrespective of the type of exercise and miR-25-3p as a specific responder to endurance training. Therefore, this study provides for the first time an objective measure to categorise individuals as sedentary or trained in young male population and moreover, it highlights a common epigenetic modulation between models of training.</p>
Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".
<p>Original hepatitis C virus hypervariable region 1 NGS sequences in fastq format from patients analyzed in the study "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center". </p> <p> </p>
Sequences of bacterial and fungal communities by Next-Generation Sequencing (NGS) associated to wall patinas
<p>These data are the DNA sequences obtained from samples of house plaster with chromatic alteration facies. Three areas (VA1, VA3 and VNA6) of the wall were sampled. the DNA was extracted using the Spin Kit For Soil MPBio. </p> <p>The DNA extracts were sequenced by NGS using Illumina MiSeq</p> <p>Two regions were selected V3V5 16S for Bacteria and ITS1 ITS for Fungi. </p>
A next-generation sequencing study of arthropods in the diet of Laysan Teal (Anas laysanensis)
<p><span>The critically endangered Laysan Teal <em>Anas laysanensis</em> (known as koloa pōhaka in the Hawaiian language) in the Northwestern Hawaiian Islands has wild populations on Kamole (Laysan Island), Kuaihelani (Midway Atoll NWR), and Hōlanikū (Kure Atoll). The Laysan Teal faces a new risk on Sand Island, Kuaihelani: non-target poisoning via a pending House Mouse <em>Mus musculus</em> eradication. After mice were observed attacking and depredating Laysan Albatross <em>Phoebastria immutabilis</em> (mōlī) in 2015, plans to eradicate mice were developed to protect this seabird species. However, this approach risks poisoning the Laysan Teal. To reduce exposure, teal will be translocated during mouse eradication. Even so, there is a potential risk of secondary poisoning for teal by ingesting arthropods that feed on mouse bait. We therefore used next-generation sequencing (NGS) to identify which arthropods teal consume. From August 2019 to February 2020, we collected 71 fresh teal faecal samples on Sand Island, and successfully extracted DNA from 21 samples. Via NGS, we found that teal most frequently consume cockroaches (order: Blattodea), freshwater ostracods (Cyprididae), midges (Chironomidae), and isopods (Porcellionidae). To a lesser degree, teal also eat spiders (Araneae), moths (Lepidoptera), beetles (Coleoptera), springtails (Entomobryomorpha), thrips (Thysanoptera), and crabs (Decapoda). Notably, Sand Island's teal consume entirely different arthropods from teal on Kamole, which mainly eat flies (Diptera) and brine shrimp (Anostraca, <em>Artemia </em>sp.). Our study serves as a model for risk mitigation during invasive rodent eradications.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.