Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.7.1
Dataset results
25 results for “MiSeq”
Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - MiSeq - U5-H1-M1-M2-M3-M4-M5 - FASTQ
<p>mtDNA mixture model of 2 mtDNA sequences belonging to haplogroups U5 and H1. Run on Illumina MiSeq with 3 different polymerases (Clontech, Herculase, NEB Taq), and different DNA extraction protocols - Paired-end Fastq files</p> <p>M1 = Mixture 1:2 i.e. 50%</p> <p>M2 = Mixture 1:10 i.e. 10%</p> <p>M3 = Mixture 1:50 i.e. 2%</p> <p>M4 = Mixture 1:100 i.e. 1%</p> <p>M5 = Mixture 1:200 i.e. 0.5%</p>
High-throughput poly(A) length measurement of HeLa and NIH 3T3 cells using TAIL-seq with MiSeq
<p>This dataset contains the full raw data directory from Illumina MiSeq generated for Chang et al. (2014, DOI: 10.1016/j.molcel.2014.02.007). Please refer to the original paper and its supplementary materials for further details.</p>
RSYD-BASIC results for AMR benchmarking dataset subset (MiSeq data)
<p><strong>Input data:</strong></p> <ul> <li>20240905_test_config.yaml: original config file used to run the pipeline</li> <li>20241004_rsyd_largeset_reads.zip: renamed Illumina MiSeq reads</li> <li>20241011-sample-overview.xlsx: overview of SRR accession numbers to internal sample numbers</li> <li>ILM_Run0001_Y20240904_kts_new.xlsx: runsheet </li> <li>input_en.yaml: column name configuration for the run</li> <li>lis_data.zip: LIS report and bacteria list used for LIS-specific results</li> </ul> <p><strong>Expected results:</strong></p> <ul> <li>20240910_test_illumina_largeset.zip: Results of the RSYD-BASIC pipeline, version 1.15.1, with the reads used</li> </ul>
Data from: A method to generate multi-locus barcodes of pinned insect specimens using MiSeq
Open the record for dataset details and reuse information.
MiSeq data from Henssen et al. Elife 2015 DOI: 10.7554/eLife.10565
<p>This data represents the output (FASTQ format) of sequencing FLEA-PCR products, as described in Henssen et al. "Genomic DNA transposition induced by human PGBD5" DOI: <a href="https://doi.org/10.7554/eLife.10565">10.7554/eLife.10565</a> </p> <p>All scripts that were used to analyse this data can be found <a href="https://zenodo.org/record/22206#.Xmd-IqhKiUk">https://zenodo.org/record/22206#.Xmd-IqhKiUk</a></p> <p>The annotation of the samples can be found in the attached excel file Summary_MiSeq.xlsx </p> <p>Any questions regarding this dataset should be addressed to kentsisresearchgroup@gmail.com or henssenlab@gmail.com </p>
Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - MiSeq - U5-H1-M1-M2-M3-M4-M5 - BAM
<p>mtDNA mixture model of 2 mtDNA sequences belonging to haplogroups U5 and H1. Run on Illumina MiSeq with 3 different polymerases (Clontech, Herculase, NEB Taq), and different DNA extraction protocols - <strong>BAM FILES </strong></p> <p>M1 = Mixture 1:2 i.e. 50%</p> <p>M2 = Mixture 1:10 i.e. 10%</p> <p>M3 = Mixture 1:50 i.e. 2%</p> <p>M4 = Mixture 1:100 i.e. 1%</p> <p>M5 = Mixture 1:200 i.e. 0.5%</p>
Mothur MiSeq SOP Galaxy Tutorial Data
<p>These are files for use with the Galaxy metagenomics tutorial "Mothur MiSeq SOP"</p>
Silene uralensis aggregate, circumpolar species : dataset, Miseq Illumina RAW READS
<p>This dataset includes 43 samples of circumpolar species included in the <em>Silene uralensis</em> aggregate, sensu http://panarcticflora.org/. Forty-eight low copy nuclear genes were enriched with <em>Silene-</em>specific probes. The samples were sequenced with the Miseq technology from the short read Illumina platform. An excel sheet with samples information is included.</p> <p>Two other datasets are associated to this one, called "Silene uralensis aggregate, circumpolar species : dataset, Novaseq Illumina RAW READS SET 1" on Zenodo 10.5281/zenodo.12699639 and "Silene uralensis aggregate, circumpolar species : dataset, Novaseq Illumina RAW READS SET 2" on Zenodo 10.5281/zenodo.12700012. These three datasets belong to a study about phylogenetics in the circumpolar <em>Silene uralensis</em> aggregate. </p> <div> <div> <div> </div> <div> <div> <div> </div> <div> <p> </p> <p> </p> </div> </div> </div> </div> </div>
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Genetic information is a valuable component of biosystematics, especially specimen identification through the use of species-specific DNA barcodes. Although many genomics applications have shifted to High-Throughput Sequencing (HTS) or Next-Generation Sequencing (NGS) technologies, sample identification (e.g., via DNA barcoding) is still most often done with Sanger sequencing. Here, we present a scalable double dual-indexing approach using an Illumina Miseq platform to sequence DNA barcode markers. We achieved 97.3% success by using half of an Illumina Miseq flowcell to obtain 658 base pairs of the cytochrome c oxidase I DNA barcode in 1,010 specimens from eleven orders of arthropods. Our approach recovers a greater proportion of DNA barcode sequences from individuals than does conventional Sanger sequencing, while at the same time reducing both per specimen costs and labor time by nearly 80%. In addition, the use of HTS allows the recovery of multiple sequences per specimen, for deeper analysis of genetic variation in target gene regions.
Raw cutadapt miseq output - minibarcode 18S-V7 of macroalgae
<p>Macroalgae are key primary producers in North Atlantic and Arctic coastal ecosystems, and tracing their fate and distribution is vital to improve our understanding of their ecological role and provision of ecosystem services. Recent advances from environmental DNA (eDNA) have added a new capacity to fingerprint and trace macroalgae. However, further development of resources for amplifying and identifying macroalgal eDNA are much needed. Here, we examined the performance in terms of resolution and specificity of two 18S primers (18S-V7 & 18S-V9) recently applied in identifying macroalgae from eDNA. We also built a local barcode database for primer 18S-V7 with 31 widespread Arctic and North Atlantic macroalgal species to complement the existing DNA databases. Furthermore, we applied metabarcoding of eDNA to identify macroalgae in Arctic marine sediments (Disko Bay, W. Greenland) and evaluated the contributions from our local barcode database. We identified macroalgal DNA from 19 families across 11 orders in surface (0-1 cm, with both primers) and sub-surface (5-10 cm, with 18S-V7 primer) sediments. The barcode database developed here with the 18S-V7 primer improved the identification of unique families, from 16 to 19 families, thereby strengthening the taxonomic assignment possible relative to pre-existing barcode reference sequences. Overall, this study demonstrates the feasibility of eDNA to resolve contributions of macroalgae in Arctic marine sediments, and enhances the fingerprinting resolution. We thereby document a novel pathway to answer key questions on the ecological role and fate of macroalgae in the Arctic.</p>
Raw cutadapt miseq output - minibarcode 18S-V7 of macroalgae
Open the record for dataset details and reuse information.
Data from: Characterisation of microsatellite and SNP markers from Miseq and genotyping-by-sequencing data among parapatric Urophora cardui (Tephritidae) populations
Open the record for dataset details and reuse information.
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Open the record for dataset details and reuse information.
Data supplementing the article "Aquatic biofilms as passive environmental DNA samplers: application to benthic macroinvertebrate communities in rivers" - raw MiSeq data inventories
<p>These data supplement the article “Aquatic biofilms as passive environmental DNA samplers: application to benthic macroinvertebrate communities in rivers” Sinziana F. Rivera, Valentin Vasselon, Nathalie Mary, Olivier Monnier, Fréderic Rimet & Agnès Bouchez submitted to “Molecular Ecology Resources” journal.</p> <p>The directory is composed of: “38_samples_fastq:files”: contains raw demultiplexed fastq files (R1. fastq and R2. fastq) for each of the 38 samples used in this study to produce OTUs and taxonomic inventories.</p> <p>“Samples id.xlsx”: contains the samples ID of the fastq files</p> <p>“Inventories.xlsx”: contains single and multi-habitat morphological inventories as well as molecular inventories resulting from the study</p>
Data from: Bacterial characterization of Beijing drinking water by flow cytometry and MiSeq sequencing of the 16S rRNA gene
Flow cytometry (FCM) and 16S rRNA gene sequencing data are commonly used to monitor and characterize microbial differences in drinking water distribution systems. In this study, to assess microbial differences in drinking water distribution systems, 12 water samples from different sources water (groundwater, GW; surface water, SW) were analyzed by FCM, heterotrophic plate count (HPC), and 16S rRNA gene sequencing. FCM intact cell concentrations varied from 2.2 × 103 cells/mL to 1.6 × 104 cells/mL in the network. Characteristics of each water sample were also observed by FCM fluorescence fingerprint analysis. 16S rRNA gene sequencing showed that Proteobacteria (76.9–42.3%) or Cyanobacteria (42.0–3.1%) was most abundant among samples. Proteobacteria were abundant in samples containing chlorine, indicating resistance to disinfection. Interestingly, Mycobacterium, Corynebacterium, and Pseudomonas, were detected in drinking water distribution systems. There was no evidence that these microorganisms represented a health concern through water consumption by the general population. However, they provided a health risk for special crowd, such as the elderly or infants, patients with burns and immune-compromised people exposed by drinking. The combined use of FCM to detect total bacteria concentrations and sequencing to determine the relative abundance of pathogenic bacteria resulted in the quantitative evaluation of drinking water distribution systems. Knowledge regarding the concentration of opportunistic pathogenic bacteria will be particularly useful for epidemiological studies.
Data from: Bacterial characterization of Beijing drinking water by flow cytometry and MiSeq sequencing of the 16S rRNA gene
Open the record for dataset details and reuse information.
SPLICER: A Highly Efficient Base Editing Toolbox That Enables in vivo Exon Skipping For Targeting Alzheimer’s Disease [miSeq]
GEO Series GSE246586. Homo sapiens; Mus musculus. 177 samples. Type: Other.
Evaluation of MiSeq for Microbial Identification in Specimens
ClinicalTrials.gov study NCT02578875. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Gut Microbiota Orchestrates Energy Homeostasis during Cold [16S rRNA gene V4-region miSeq]
GEO Series GSE74227. mouse gut metagenome. 48 samples. Type: Other.
Genome editing reveals a role for OCT4 in human preimplantation development [MiSeq]
GEO Series GSE100119. Homo sapiens; Mus musculus. 244 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.