Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
54
datasets available to search
ShareScore release 0.9.0
Dataset results
54 results for “NGS DATA”
Data set for publication: Determination of Virulence-Associated Genes and Antimicrobial Resistance Profiles in Brucella Isolates Recovered from Humans and Animals in Iran Using NGS Technology
<p>This dataset includes information on resistance profiling, as well as antimicrobial resistance (AMR) genes and virulence-related factors that were identified in <em>Brucella</em> isolates recovered from humans and animals in different regions of Iran using classical phenotyping and next-generation sequencing (NGS) technology.</p>
Example Dataset for npstat: Population genetics from Pooled NGS data NPStat v1: User guide
<p>Example Dataset for npstat to test the program and the different options.</p> <p>The example dataset contains a pileup file with sequences of of the 2L chromosome from fifteen pooled inbreed individuals of <em>Drosophila melanogaster </em>(<span>doi: 10.1038/nature10811</span>). The dataset also contains the sequence reference of the 2L chromosome in fasta format, an outgroup sequence in fasta format of <em>D. yakuba</em> (SRR26246471), a GFF3 annotation file and a file with a brief list of selected SNPs to be analyzed.</p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics"
<p>NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics", currently in review.</p>
NGS competence network - pipeline benchmark data
<p>VCF files generated with the megSAP pipeline for a pipeline benchmark performed by the NGS competence network.</p>
Virtual NGS data used in the paper "Optimal k-mer values for good-quality triploid genome assemblies" and its assembly results
<p>Despite technological advancements, whole-genome sequencing remains technically challenging for organisms with higher-ploidy genomes. Therefore, the use of short-read sequencing platforms for this purpose has been attempted, but the conditions that result in poor-quality genomes have not been elucidated. Therefore, in the present study, simulated sequences mimicking the accumulation of insertion/deletion mutations were created to clarify the permissible differences between homologous chromosomes and the k-mer sizes for good-quality genome assembly from short-read sequencing data for triploid species. The results illustrated that a narrow range of k-mers permits the generation of high-quality assemblies for any level of difference between homologous chromosomes. <br> This dataset consists of the virtual haploid genome (O.fasta), triploid genome data (reference_sequense), NGS read data (NGS_reads), and assembly results (all_contigs) created in this study.<br> </p>
Raw data NGS 12s
<p>Raw data from NGS sequencing targeting the 12S marker. The data were used in a study comparing the effectiveness of conventional methods (electrofishing) with an environmental DNA metabarcoding method, for studying the specific richness of fish communities in estuaries on Reunion Island (JANNEL et al. 2024).</p>
ModEst - Precise estimation of genome size from NGS data
<p>Accurate estimates of genome sizes are important parameters for both theoretical and practical biodiversity genomics. We present here a fast, easy-to-implement and precise method to estimate genome size from the number of bases sequenced and the mean sequencing depth. To estimate the latter, we take advantage of the fact that a precise estimation of the Poisson distribution parameter lambda is possible from truncated data, restricted to the part of the sequencing depth distribution representing the true underlying distribution. With simulations we could show that reasonable genome size estimates can be gained even from low-coverage (10X), highly discontinuous genome drafts. Comparison of estimates from a wide range of taxa and sequencing strategies with flow-cytometry estimates of the same individuals showed a very good fit and suggested that both methods yield comparable, interchangeable results.</p>
Mycobacterium tuberculosis complex NGS made easy: data analysis step-by-step (datasets)
<p>Datasets used in the workshop "<em>Mycobacterium tuberculosis</em> complex NGS made easy: data analysis step-by-step". A capacity development initiative by TB-CAPT, PanACEA and Seq&Treat.</p>
Raw sequencing data of NGS of stomach contents of sympatric species of weakly electric fish (genus: Campylomormyrus)
<p>This dataset contains the raw sequencing data of stomach contents of sympatric species of weakly electric fish (genus: <em>Campylomormyrus</em>) using next generation sequencing.</p> <p>Stomach content samples were collected from five <em>Campylomormyrus</em> species (<em>C. alces</em>, ; <em>C. compressirostris</em>, ; <em>C. curvirostris</em>, ; <em>C. numenius</em>, ; <em>C. tshokwe</em>) and samples of <em>Gnathonemus petersii</em> (<em>G. petersii, </em>), a sister genus of <em>Campylomormyrus.</em></p> <p>The fish specimens, from which these stomach content samples are extracted, were collected during an expedition to the Republic of the Congo in fall 2012.</p> <p>The dataset files are in FASTA format.</p>
Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"
<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication “Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins”. </p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of four homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>
NGS data produced in 'Rapid selection and identification of functional CD8+ T-cell epitopes from large peptide-coding libraries'; Nature Communications (2019)
<p>Sharma, G et al. Rapid selection and identification of functional CD8+ T-cell epitopes from large peptide-coding libraries. <em>Nature Communications</em>. Accepted (August 2019)</p> <p><strong>Abstract:</strong></p> <p>Cytotoxic CD8+ T-cells recognize and eliminate infected or malignant cells that present, at their cell surfaces, short peptide epitopes derived from intracellularly processed antigens. However, broadly searching for specific major histocompatibility complex (MHC)-bound peptide epitopes that are naturally processed and capable of eliciting a functional T-cell response has been challenging. Here, we report a method for deep and unbiased T-cell epitope profiling, which is done by using <em>in vitro</em> co-culture of CD8+ T-cells and target cells transduced with high-complexity epitope-encoding minigene libraries. Target cells that are subject to cytotoxic attack from T-cells in co-culture are isolated, before they are lost to apoptosis, by fluorescence-activated cell sorting and characterized by sequencing the minigenes encoded within. In the present study, we validate this highly parallelized method using known murine T-cell receptor/peptide-MHC pairs and diverse minigene-encoded epitope libraries to identify naturally processed and MHC-presented peptide epitopes unambiguously and with high sensitivity.</p>
NGSAP-VC : Genomic Variant Calling as an Installable GALAXY Workflow Using NGS data.
<p>Implementation of genomic variants calling as an installable GALAXY workflows using NGS data. Repository contains two separate sets of simulated ebola test data. One for SNPs and INDELs calling and another for Structural Variants calling.</p>
NGS data from: Detection of unintended on-target effects in CRISPR genome editing by DNA donors carrying diagnostic mutations
<p>We developed a method to detect copy number variation and loss of heterozygosity in genome-edited cell lines by only sequencing the edited site. We introduced different substitutions on both chromosomes of diploid H9 human embryonic stem cells using CRISPR/Cas9 and a mixture of homology-directed repair donor DNAs carrying different substitutions. Here, we deposit the associated sequencing data (BAM-files) generated by extracting genomic DNA from genome-edited H9 cells, PCR amplifying the region of interest and sequencing on a MiSeq platform. We sequenced the edited site of cell bulks to investigate the editing efficiencies of individual donors from the donor mixtures. For proof of concept, we isolated over 850 cellular clones from the edited cell bulks and sequenced the edited site and heterozygous positions flanking the edited site to determine the number of alleles at the edited site and detect copy-neutral loss of heterozygosity. Additionally, we sequenced the DNA donors to assess, the distribution of individual donors in the mixtures.</p>
N-TIMP2 CD loop extension NGS data and analysis script
<p><strong>NGS data and analysis script for the YSD sorts described in Bonadio et al. "Using designed loop extension and combinatorial screening to enhance specificity of a broad matrix metalloproteinase inhibitor"</strong></p>
NGS data from: Deploying synthetic coevolution and machine learning to engineer protein-protein interactions
<p>Fine-tuning of protein-protein interactions occurs naturally through coevolution, but this process is difficult to recapitulate in the laboratory. We describe a synthetic platform for protein-protein coevolution that can isolate matched pairs of interacting muteins from complex libraries. This large dataset of coevolved complexes<span class="Apple-converted-space"> </span>drove a systems-level analysis of molecular recognition between Z domain-affibody pairs spanning a wide range of structures, affinities, cross-reactivities, and orthogonalities, and captured a broad spectrum of coevolutionary networks. Furthermore, we harnessed pre-trained protein language models to expand, <em>in silico</em>, the amino acid diversity of our coevolution screen, predicting remodeled interfaces beyond the reach of the experimental library. The integration of these approaches provides a means of generating protein complexes with diverse molecular recognition properties as tools for biotechnology and synthetic biology.</p>
NGS data from: Detection of unintended on-target effects in CRISPR genome editing by DNA donors carrying diagnostic mutations
Open the record for dataset details and reuse information.
ModEst - Precise estimation of genome size from NGS data
Open the record for dataset details and reuse information.
NGS data from: Deploying synthetic coevolution and machine learning to engineer protein-protein interactions
Open the record for dataset details and reuse information.
Data from: A NGS approach to the encrusting Mediterranean sponge Crella elegans (Porifera, Demospongiae, Poecilosclerida): transcriptome sequencing, characterization and overview of the gene expression along three life cycle stages
Sponges can be dominant organisms in many marine and freshwater habitats where they play essential ecological roles. They also represent a key group to address important questions in early metazoan evolution. Recent approaches for improving knowledge on sponge biological and ecological functions as well as on animal evolution have focused on the genetic toolkits involved in ecological responses to environmental changes (biotic and abiotic), development and reproduction. These approaches are possible thanks to newly available, massive sequencing technologies–such as the Illumina platform, which facilitate genome and transcriptome sequencing in a cost-effective manner. Here we present the first NGS (next-generation sequencing) approach to understanding the life cycle of an encrusting marine sponge. For this we sequenced libraries of three different life cycle stages of the Mediterranean sponge Crella elegans and generated de novo transcriptome assemblies. Three assemblies were based on sponge tissue of a particular life cycle stage, including non-reproductive tissue, tissue with sperm cysts and tissue with larvae. The fourth assembly pooled the data from all three stages. By aggregating data from all the different life cycle stages we obtained a higher total number of contigs, contigs with blast hit and annotated contigs than from one stage-based assemblies. In that multi-stage assembly we obtained a larger number of the developmental regulatory genes known for metazoans than in any other assembly. We also advance the differential expression of selected genes in the three life cycle stages to explore the potential of RNA-seq for improving knowledge on functional processes along the sponge life cycle.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.