Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,270
datasets available to search
ShareScore release 0.7.1
Dataset results
2,270 results for “next generation sequencing”
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Structural variant discovery and genotyping in next-generation sequencing data
<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels >50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Genotyping of European Toxoplasma gondii strains by a new high-resolution next-generation sequencing-based method
<p>The data set comprises 164 FASTQ files generated with an Ion AmpliSeq-based genotyping method for <em>Toxoplasma gondii </em>and<em> </em>a BED file used for the design of the Ion AmpliSeq primer panel. The FASTA file named as "AmpliSeq-ME49-Reference" was used as a reference for mapping and data analysis of the FASTQ files. The GZ file named as "Tgondii_IonAmpliSeq_Results_SNPs_VCF" is a VCF file, which contains all SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference. The VCF file was converted into a FASTA file named as "Tgondii_IonAmpliSeq_Results_SNPs", which also contains the SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference.</p> <p>The work is published in the European Journal of Clinical Microbiology & Infectious Diseases with the title "Genotyping of European <em>Toxoplasma gondii</em> strains by a new high‑resolution next‑generation sequencing‑based method"; https://doi.org/10.1007/s10096-023-04721-7</p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
Graphing and tabulating next-generation sequencing and genotyping data
<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Read-mapping for next-generation sequencing data (Drosophila melanogaster)
<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Read-mapping for next-generation sequencing data (Wolbachia)
<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Genotype reproducibility testing in next-generation sequencing data
<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas
<p>Next-Generation Sequencing Dataset of Adult Pilocytic Astrocytomas</p> <p>Pilocytic astrocytoma (PA) is a benign grade 1 glioma according to the World Health Organization (WHO), common in children but rare in adults, where it may have a worse prognosis. Pediatric PA is usually associated with dysregulation of the MAPK pathway, often involving BRAF alterations such as the KIAA1549::BRAF (K-B) fusion or the V600E mutation. This dataset contains molecular data of 28 cases of adult PA obtained by using gene-targeted next-generation sequencing (NGS).</p>
Fig. 4 in Marked genetic diversity within Blastocystis in Australian wildlife revealed using a next generation sequencing-phylogenetic approach
Fig. 4. Relative abundance of Blastocystis subtypes (STs) in marsupial and deer species. Marsupials are represented by eastern grey kangaroos and wallabies; deer are represented by red, fallow and sambar deer. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 3 in Marked genetic diversity within Blastocystis in Australian wildlife revealed using a next generation sequencing-phylogenetic approach
Fig. 3. Phylogenetic analysis of SSU-rRNA sequence data (aligned over 2035 positions) to infer the relationships of recognised Blastocystis subtypes (STs) as well as new STs discovered in the present study. The tree was constructed using Bayesian Inference method (MrBayes) and used Proteromonas lacertae as an outgroup. Posterior probabilities less than 0.95% are not displayed. The two novel subtypes and additional ST13 and ST24 sequences are indicated in bold. After the present analysis was completed, Santín et al. (2023) reported a subdivision of "ST10" into four STs (i.e. ST10, ST42, ST43 and ST44).
Fig. 2 in Marked genetic diversity within Blastocystis in Australian wildlife revealed using a next generation sequencing-phylogenetic approach
Fig. 2. Diagram of the method used to obtain sequence for a SSU-rRNA gene region (~1750 bp) of Blastocystis. Two primer sets were used to obtain overlapping sequences for this region.
Fig. 3 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 3. Phylogenetic relationships between Trypanosoma sp. ANU2 and other members of the Trypanosoma genus. Maximum likelihood tree is shown. Neighbour-joining and maximumlikelihood bootstrap support followed by Bayesian posterior probability is shown at nodes, respectively. Genbank accession numbers follow species/genotype description. Trypanosoma species/genotypes isolated in Australia are in bold. Scale bar represents substitution per site.
Fig. 6 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 6. Principle Coordinates Analysis (PCoA) plots demonstrating relationship between Trypanosoma spp. ZOTUs and host species. Dissimilarity matrices were generated using sqrt transformed Bray-Curtis distances to show distance between abundance of ZOTUs between host marsupial species; the woylie (blue squares) and brushtail possum (red circles). (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 2 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 2. Abundance and diversity map of Trypanosoma copemani genotype 1 (G1) and genotype 2 (G2) positive samples. ZOTUs were sorted into G1 or G2 based on phylogenetic inference shown in rows, while columns are individual marsupial blood samples from infected individuals. The map represents samples separated by host species, which were WOY = woylie, or BTP = brushtail possum. Grayscale indicates number of sequences obtained from that sample as shown in the scale of intensity on the right (log transformed abundance).
Fig. 5 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 5. Trypanosoma spp. polyparasitism in 70 blood samples taken from marsupials in the Upper Warren Region. Marsupial species include: woylie (WOY), brushtail possum (BTP) and chuditch (CHU). Trypanosoma spp. include; C = Trypanosoma copemani, V = T. vegrandis, N = T. noyesi, G = T. gilletti, A = T. sp. ANU2, I = T. irwini, AT = T. sp. AAT, U = unknown, AA = T. avium, and CR = Crithidia spp.
Fig. 4 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 4. Abundance and diversity map of Trypanosoma spp. ZOTUs in different marsupial blood samples assigned to species groups shown in rows, while columns are individual blood samples. Samples are separated by host species including; WOY = woylie (34), CHU = chuditch (3), and BTP = brushtail possum (33). The colour scale indicates increasing number of sequences obtained from that sample, in that species, as shown in the intensity bar on the right. Species include; C = Trypanosoma copemani, V = T. vegrandis, N = T. noyesi, G = T. gilletti, A = T. sp. ANU2, I = T. irwini, AT = T. sp. AAT, U = unknown, AA = T. avium, and CR = Crithidia spp. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 1 in Next generation sequencing reveals widespread trypanosome diversity and polyparasitism in marsupials from Western Australia
Fig. 1. Phylogenetic relationships between Trypanosoma spp. ZOTUs assigned to nine species groups compared to 24 representative reference strains downloaded from Genbank. Phytomonas serpens and Leptomonas sp. were used as outgroups. Bayesian analysis was used to produce tree topology and posterior probability is shown at nodes. Scale bar represents substitution per site.
Structural modelling results to accompany the paper "Uncommon mutational profiles of metastatic colorectal cancer detected during routine genotyping using next generation sequencing: an update"
<p>This repository contains the results of modelling missense mutants in KRAS, NRAS and BRAF observed in our study in the corresponding protein structures. Modelling was performed using FoldX.</p>
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.