Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
140
datasets available to search
ShareScore release 0.9.0
Dataset results
140 results for “Illumina Sequencing”
Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets
<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>: BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em> to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016 <a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>
SARS-Cov-2 illumina sequencing training course
<p>Dataset with two samples of SARS-Cov-2 sequenced with Illumina using Artic v3 amplicon enrichment protocol.</p>
Illumina sequencing data of Agro-mediated gene-edited apple lines
<p>FastaQ pair-end Illumina sequencing data of the Dipm1/4, Hipm1, and Mlo19 genes from the different apple Agro-edited lines obtained in the project. GA = Gala; GD = Golden Delicious</p> <p>Lines included are:</p> <p>GA1, GA3, GA4, GA5, and GA WT</p> <p>GD1, 2, 6, 10, 11, 12, 15, 17, 18, 19, 20, 21, 23, 24, 26, 27, 29, 31, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and WT</p>
Illumina Sequencing Data for "Elucidating human gut microbiota interactions that robustly inhibit diverse Clostridioides difficile strains across different nutrient landscapes"
<p>Illumina Sequencing Data for Sulaiman et al., "Elucidating human gut microbiota interactions that robustly inhibit diverse Clostridioides difficile strains across different nutrient landscapes".</p>
Draft genome assemblies of killifish from the Fundulus genus with ONT and Illumina sequencing platforms
<p>Four species from the genus Fundulus were selected for genome sequencing to study the physiological and genetic mechanisms that diverge between euryhaline and stenohaline freshwater species within this cyprinodontiform order of ray-finned fishes.</p>
Development of Illumina oligos for fungal ITS sequencing
New PCR primers were designed with the following objectives: 1) targetting the ITS2 region to minimize bias due to introns at the 3' end of the SSU, 2) have the broadest possible coverage of Fungi, 3) minimize amplification of plant, protist and other Eukaryote sequences, 4) compatibility with Illumina sequencing adaptors and GOLAY barcodes. Primer regions were selected by eye from sequence alignments then further analyzed using PrimerProspector. At present, several mock communities and hundreds of soil DNA extracts have been successfully sequenced on the Illumina MiSeq platform using the oligos that were developed.
.bam alignment files of Illumina and ONT sequencing of pREF plasmid
<p>The expression of genes encompasses their transcription into mRNA followed by translation into protein. In recent years, next-generation sequencing and mass spectrometry methods have profiled DNA, RNA and protein abundance in cells. However, there are currently no reference standards that are compatible across these genomic, transcriptomic and proteomic methods, and provide an integrated measure of gene expression. Here, we use synthetic biology principles to engineer a multi-omics control, termed <em>pREF</em>, that can act as a universal molecular standard for next-generation sequencing and mass spectrometry methods. The <em>pREF</em> sequence encodes 21 synthetic genes that can be <em>in vitro</em> transcribed into spike-in mRNA controls, and <em>in vitro</em> translated to generate matched protein controls. The synthetic genes provide qualitative controls that can measure sensitivity and quantitative accuracy of DNA, RNA and peptide detection. We demonstrate the use of <em>pREF</em> in metagenome DNA sequencing and RNA sequencing experiments and evaluate the quantification of proteins using mass spectrometry. Unlike previous spike-in controls, <em>pREF</em> can be independently propagated and the synthetic mRNA and protein controls can be sustainably prepared by recipient laboratories using common molecular biology techniques. Together, this provides the first universal synthetic standard able to integrate genomic, transcriptomic and proteomic methods.</p>
Illumina Sequencing Data for "Phocaeicola vulgatus shapes the long-term growth dynamics and evolutionary adaptations of Clostridioides difficile"
<p>Illumina Sequencing Data for "<em>Phocaeicola vulgatus</em> shapes the long-term growth dynamics and evolutionary adaptations of <em>Clostridioides difficile</em>"</p>
Illumina TruSeq stranded mRNA sequences of Rhynchosporium commune isolate UK7 in plantae
<p>Transcription profiles were generated from the <em>Rhynchosporium commune</em> isolate UK7. RNA-seq experiments were conducted <em>in plantae</em> on the barley cultivar Beatrix (Viskosa 9 Pasadena, Saaten Union, breeders’ Reference NS01/2449). Leaves were collected at 9 and at 13 days post infection (dpi). All experiments were conducted in triplicates. Total RNA was extracted using TRIzol (Invitrogen Inc.) following the manufacturer’s recommendations. RNA integrity and quantity was assessed on a Bioanalyzer 2100 (Agilent) and a Qubit fluorometer (Life Technologies) and a Bioanalyzer 2100. Libraries were prepared using the TruSeq stranded mRNA sample prep kit (Illumina Inc.). Total RNA was ribosome-depleted by using polyA selection and reverse-transcribed into double-stranded cDNA.</p>
Illumina TruSeq stranded mRNA sequences of Rhynchosporium commune isolate UK7 in vitro
<p>Transcription profiles were generated from the <em>Rhynchosporium commune</em> isolate UK7. Cultures were grown either on Luria-Bertani (LBA) or Potato Dextrose Broth (PDB) media. Total RNA was extracted using TRIzol (Invitrogen Inc.) following the manufacturer’s recommendations. Experiments were conducted in triplicates. RNA integrity and quantity was assessed on a Bioanalyzer 2100 (Agilent) and a Qubit fluorometer (Life Technologies) and a Bioanalyzer 2100. Libraries were prepared using the TruSeq stranded mRNA sample prep kit (Illumina Inc.). Total RNA was ribosome-depleted by using polyA selection and reverse-transcribed into double-stranded cDNA.</p>
VADER TyrT Illumina sequencing data
<p>FASTQ files used in data analysis for "Virus-assisted directed evolution of enhanced suppressor tRNAs in mammalian cells."</p>
VADER PylT-PyOtR Illumina sequencing data
<p>FASTQ files used in data analysis for "PyOtR: a novel Pyrrolysyl tRNA evolved for enhanced unnatural amino acid incorporation."</p>
VADER TyrT-MARIO Illumina sequencing data
<p>FASTQ files used in data analysis for "Evolution of MARIO, an improved Tyrosyl tRNA for more efficient genetic code expansion."</p>
Illumina RNA-Sequencing fastq data from insecticide resistant Anopheles gambiae s.l
<p>This is a dataset of Illumina RNA sequencing reads, for a project investigating resistance to Pirimiphos-methyl in the major malaria vectors, Anopheles gambiae and Anopheles coluzzii. There are four biological replicates for the following conditions:</p> <p> </p> <p>Ngousso (susceptible)</p> <p>Kisumu (susceptible)</p> <p>Bouake gambiae unexposed</p> <p>Bouake gambiae PM survivors</p> <p>Bouake coluzzii unexposed </p> <p>Bouake coluzzii PM survivors </p> <p> </p> <p>SRA submission: SUB14596876</p> <p> </p>
Combined high-depth Illumina+PacBio Sequencing of several samples from FDA-ARGOS
<p>The (real) sequencing data is compiled from a concatenation of sequencing runs from Database for Reference Grade Microbial Sequences (FDA-ARGOS). Specifically, the following samples were sequenced with both Illumina and PacBio. The sample accessions are shown below.</p> <pre><code> BioSample Run Platform Organism bases source <chr> <chr> <chr> <chr> <dbl> <chr> 1 SAMN06173354 SRR5409204 ILLUMINA Bacillus anthracis 1778000000 Colorado Serum Co., Anthrax Spore Vaccine 2 SAMN06173354 SRR5409205 PACBIO_SMRT Bacillus anthracis 2420000000 Colorado Serum Co., Anthrax Spore Vaccine 3 SAMN06173356 SRR5448657 ILLUMINA Bacillus circulans 2311000000 swab with brown-gray powder 4 SAMN06173356 SRR5448656 PACBIO_SMRT Bacillus circulans 242000000 swab with brown-gray powder 5 SAMN04875535 SRR4123920 ILLUMINA Elizabethkingia anophelis 1357000000 blood 6 SAMN04875535 SRR4123919 PACBIO_SMRT Elizabethkingia anophelis 2173000000 blood 7 SAMN06173306 SRR5413253 ILLUMINA Escherichia coli O157 3051000000 clinical isolate 8 SAMN06173306 SRR5413252 PACBIO_SMRT Escherichia coli O157 915000000 clinical isolate 9 SAMN06173318 SRR5413272 ILLUMINA Mycobacterium avium subsp. paratuberculosis 1032000000 feces 10 SAMN06173318 SRR5413271 PACBIO_SMRT Mycobacterium avium subsp. paratuberculosis 462000000 feces 11 SAMN07312468 SRR5879398 ILLUMINA Mycobacterium tuberculosis 1054000000 human 12 SAMN07312468 SRR5879396 PACBIO_SMRT Mycobacterium tuberculosis 1854000000 human 13 SAMN04875542 SRR4123931 ILLUMINA Neisseria gonorrhoeae 1053000000 ATCC strain 14 SAMN04875542 SRR4123930 PACBIO_SMRT Neisseria gonorrhoeae 1117000000 ATCC strain </code></pre> <p>Samples were selected with the SRA Run selector. The SraRunTable.txt file was exported containing the metadata for each sample, and fastq-dump from the SRA toolkit was used to write out fastq files for each run, with paired Illumina data being split into separate _1.fastq.gz and _2.fastq.gz files. </p>
Illumina next generation ddRAD sequencing SNP data from: Contrasting genetic diversity and structure between endemic and widespread damselfishes are related to differing adaptive strategies
<p class="MsoNormal"><strong><u><span>Aim:</span></u></strong><span> Discerning when, where, and how processes of isolation lead to differing biogeography is especially complex for marine species with similar ecological niches and within the same geographic location. We assessed population genetics of congeneric and ecologically similar damselfishes within their overlapping distributions and across potential barriers to geneflow.</span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Taxon:</span></u></strong><span> <em>Dascyllus marginatus </em>(endemic) and <em>Dascyllus abudafur </em>(widespread)<em>.</em></span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Location:</span></u></strong><span> Coral reefs from the Red Sea, Djibouti, Yemen, Oman, and Madagascar. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Methods:</span></u></strong><span> We used RADseq derived SNPs to investigate key differences in population genetics between both species and discuss barriers shaping genetic differentiation (neutral vs. selective) and biogeography. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Results:</span></u></strong><strong><span> </span></strong><em><span>Dascyllus marginatus </span></em><span>inhabited the Red Sea, the coasts of Yemen (including Socotra), and the Gulf of Oman. <em>Dascyllus abudafur</em> species was present from the Red Sea to Madagascar but was absent from Yemen and Oman. Populations of <em>D. marginatus </em>had an order of magnitude higher genetic differentiation compared to <em>D. abudafur</em>, as well as several outlier loci (suggesting selective pressure), which were absent in <em>D. abudafur</em> despite equal sampling locations. In both species, specimens from the Red Sea and Djibouti formed one genetic cluster separated from all other locations. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Main conclusions:</span></u></strong><span> The stronger genetic structure at smaller geographic scale of the endemic species seems associated to faster adaptation to environmental differences; whereas the widespread species only experienced reduced geneflow and neutral differentiation at much larger geographic scales. Restrictive transitions (between the Gulf of Aqaba and the Red Sea or the Red Sea and the Gulf of Aden) did not affect the genetic architecture of either species, while the environmental shift within the Red Sea (at 22°N/20°N) affected the endemic but not the widespread species. Samples from continental Yemen revealed that a genetic break in the Gulf of Aden likely reflects historical colonization processes and not contemporary environmental regimes.</span></p>
Illumina next generation ddRAD sequencing SNP data from: Contrasting genetic diversity and structure between endemic and widespread damselfishes are related to differing adaptive strategies
Open the record for dataset details and reuse information.
.bam alignment files of Illumina and ONT sequencing of pREF plasmid
Open the record for dataset details and reuse information.
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Genetic information is a valuable component of biosystematics, especially specimen identification through the use of species-specific DNA barcodes. Although many genomics applications have shifted to High-Throughput Sequencing (HTS) or Next-Generation Sequencing (NGS) technologies, sample identification (e.g., via DNA barcoding) is still most often done with Sanger sequencing. Here, we present a scalable double dual-indexing approach using an Illumina Miseq platform to sequence DNA barcode markers. We achieved 97.3% success by using half of an Illumina Miseq flowcell to obtain 658 base pairs of the cytochrome c oxidase I DNA barcode in 1,010 specimens from eleven orders of arthropods. Our approach recovers a greater proportion of DNA barcode sequences from individuals than does conventional Sanger sequencing, while at the same time reducing both per specimen costs and labor time by nearly 80%. In addition, the use of HTS allows the recovery of multiple sequences per specimen, for deeper analysis of genetic variation in target gene regions.
Valenzuela phylogenomic dataset from: Illumina whole genome sequencing indicates ploidy level differences within the Valenzuela flavidus (Psocodea: Psocomorpha: Caeciliusidae) species complex
<p>This contains data for the manuscript: "Illumina Whole Genome Sequencing indicates Ploidy Level Differences within the <i>Valenzuela flavidus </i>(Psocodea: Psocomorpha: Caeciliusidae) Species Complex".</p> <p><i>Valenzuela flavidus</i> is a species of bark louse which is known to have asexual parthenogenetic populations in Europe but is believed to have sexual and asexual populations in North America as well. Historically, <i>Valenzuela aurantiacus</i> was the species epithet recognized for North American members until reports of asexual reproduction surfaced in certain North American populations. Cytogenetic studies have demonstrated European all-female populations are triploid. However, males are often reported in North America suggesting diploidy for sexual populations. With the use of Illumina whole genome sequencing, genetic diversity among North American and European populations was explored with phylogenomic methods. Ploidy level was estimated by examining allele frequencies of read-mapped homologous gene regions. Results indicate divergent populations between Europe and North America. North American populations containing males are estimated to be diploid suggesting a different mechanism of genomic reproduction. These results suggest divergent population structure among European asexual and North American sexual members of <i>V. flavidus</i> providing insight for future studies to understand patterns of asexuality reported within the complex.</p> <p>The following file contains all gene alignments, concatenated supermatrix, and mitochondrial alignment for this manuscript. In addition, the BAM files used to estimate allele frequencies. Also, gene trees for coalescent analysis, resultant treefiles from IQ-tree searches, and MCMCtree result.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.