Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
718
datasets available to search
ShareScore release 0.9.0
Dataset results
718 results for “illumina”
Crossreactive probes on Illumina DNA methylation arrays: a large study on ALS shows that a cautionary approach is warranted in interpreting epigenome-wide association studies
<p>Data corresponding to the paper "Crossreactive probes on Illumina DNA methylation arrays: a large study on ALS shows that a cautionary approach is warranted in interpreting epigenome-wide association studies."<br> <br> Corresponding scripts can be found at: <a href="https://github.com/pjhop/dnamarray_crossreactivity">https://github.com/pjhop/dnamarray_crossreactivity</a><br> All downstream analyses in <a href="https://github.com/pjhop/dnamarray_crossreactivity/blob/master/analysis/c9_analysis.Rmd">c9_analysis.Rmd</a> and in<a href="https://github.com/pjhop/dnamarray_crossreactivity/blob/master/analysis/supplementary_note.Rmd"> supplementary_note.Rmd</a> can be reproduced using the deposited data as follows:</p> <ul> <li>Clone the dnamarray_crossreactivity repository: < git clone https://github.com/pjhop/dnamarray_crossreactivity.git ></li> <li>Download the data ('data.zip') and place it in the 'dnamarray_crossreactivity' folder.</li> <li>Unzip the data.zip folder</li> </ul> <p>Scripts used to generate the data in each subdirectory can be found at:</p> <ul> <li>data/processed/c9_matches/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/c9_matches">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/c9_matches</a></li> <li>data/output/ewas/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/ewas">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/ewas</a></li> <li>data/output/figs/: empty folder, running 'c9_analysis.Rmd' will save figures here.</li> <li>data/misc/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/other">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/other</a></li> <li>data/extdata: <ul> <li>Zhou <em>et al.</em> annotations (EPIC.hg19.manifest.tsv.gz, HM450.hg19.manifest.pop.tsv.gz, HM450.hg19.manifest.tsv.gz) were downloaded from: <a href="https://zwdzwd.github.io/InfiniumAnnotation">https://zwdzwd.github.io/InfiniumAnnotation</a> (downloaded at 17/09/2020)</li> <li>Naeem <em>et al.</em><em> </em>data (12864_2013_7006_MOESM2_ESM.csv) was downloaded from: <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3943510/">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3943510/</a></li> <li>Chen <em>et al.</em> data (48639-non-specific-probes-Illumina450k.xlsx) was downloaded from <a href="https://github.com/Jfortin1/funnorm_repro/blob/master/bad_probes/48639-non-specific-probes-Illumina450k.xlsx">https://github.com/Jfortin1/funnorm_repro/blob/master/bad_probes/48639-non-specific-probes-Illumina450k.xlsx</a></li> <li>The anno_450k.txt.gz and anno_EPIC.txt.gz are subsets of the annotation files included in the following package respectively: <a href="https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylation450kanno.ilmn12.hg19.html">https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylation450kanno.ilmn12.hg19.html</a> and <a href="https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylationEPICanno.ilm10b2.hg19.html">https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylationEPICanno.ilm10b2.hg19.html</a></li> </ul> </li> <li> data/genome_bs: Scripts used to generate these data can be found at <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R</a> and <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R</a> .</li> <li> data/raw: Individual-level data is available upon access at: <a href="https://ega-archive.org/studies/EGAS00001004587">https://ega-archive.org/studies/EGAS00001004587</a></li> </ul>
Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis
<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>
Genome wide Illumina 450K array in AML patients with or without CEBPA mutation
<ul> <li>Num: Row Number</li> <li>Hybridization REF: CpG Reference ID</li> <li>Position: Position in hg19</li> <li>Gene: Entrez Gene Symbol</li> <li>Chrome: Chromosome Number</li> <li>p_vals: P value for t-test between patients with CEBPA and without CEBPA mutation</li> <li>mean_cebpa_mut: mean methylation score for patients with CEBPA mutation</li> <li>mean_no_cebpa_mut: mean methylation score for patients with CEBPA mutation </li> <li> <p>site_status: prediction status for CEBPA sites </p> </li> <li> <p>fdr: fdr p_value </p> </li> </ul>
Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets
<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>: BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em> to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016 <a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>
SARS-Cov-2 illumina sequencing training course
<p>Dataset with two samples of SARS-Cov-2 sequenced with Illumina using Artic v3 amplicon enrichment protocol.</p>
Improved genome annotation of Rhynchosporium commune isolate UK7 using Illumina short reads of in vitro and in plantae conditions
<p>Improved genome annotation of the <em>Rhynchosporium commune</em> isolate UK7 using Illumina short reads of <em>in vitro</em> and <em>in plantae</em> conditions. The short reads used for the annotation are available at <a href="https://doi.org/10.5281/zenodo.5729968">https://doi.org/10.5281/zenodo.5729968</a> and <a href="https://doi.org/10.5281/zenodo.5729863">https://doi.org/10.5281/zenodo.5729863</a>.To create the gene models, we used tophat v. 2.0.14 to align short reads to the UK7 reference genome (Trapnell et al., 2009). The Intron splice site hints were generated using bam2hints, included in the AUGUSTUS v. 3.2.1 software (Stanke et al., 2006). Due to the very high RNA-sequencing depth available, intron splice hints were filtered for a minimum coverage of 20 reads to avoid an impact of spurious splice signals on gene prediction. To produce <em>ab initio </em>gene models, the BRAKER v. 1.0 pipeline (Hoff et al., 2016) combining GeneMark-ET <em>ab initio </em>gene model predictions and AUGUSTUS v. 3.2.1. GeneMark-ET was trained using the RNA-seq-based splice information as hints. AUGUSTUS was automatically trained using <em>ab initio </em>gene models that were fully supported by splice information. Finally, AUGUSTUS was used to predict gene models using both RNA-seq splice information and coding sequence hints based on exonerate protein alignments as extrinsic evidence.</p>
Illumina sequencing data of Agro-mediated gene-edited apple lines
<p>FastaQ pair-end Illumina sequencing data of the Dipm1/4, Hipm1, and Mlo19 genes from the different apple Agro-edited lines obtained in the project. GA = Gala; GD = Golden Delicious</p> <p>Lines included are:</p> <p>GA1, GA3, GA4, GA5, and GA WT</p> <p>GD1, 2, 6, 10, 11, 12, 15, 17, 18, 19, 20, 21, 23, 24, 26, 27, 29, 31, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and WT</p>
The impact of low input DNA on the reliability of DNA methylation as measured by the Illumina Infinium MethylationEPIC BeadChip, supplementary table 3
<p>Supplementary table 3: Summary statistics from an EWAS assessing the relationship between variance in DNA methylation value and DNA input level.</p>
Dataset: Illumina, Inc. (ILMN) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Illumina Sequencing Data for "Elucidating human gut microbiota interactions that robustly inhibit diverse Clostridioides difficile strains across different nutrient landscapes"
<p>Illumina Sequencing Data for Sulaiman et al., "Elucidating human gut microbiota interactions that robustly inhibit diverse Clostridioides difficile strains across different nutrient landscapes".</p>
Reduced size Illumina NovaSeq runfolder
<p>This dataset is a reduced size version of an Illumina NovaSeq run of PhiX, where all but the first four cycles have been removed. All metadata associated with the run has been kept intact. This dataset is e.g. useful in testing workflows which require raw Illumina data (i.e. bcl files rather than fastq files).</p> <p>To demultiplex this data with Illumina bclfastq (https://support.illumina.com/sequencing/sequencing_software/bcl2fastq-conversion-software.html) it is recommended that the following options are used:</p> <pre><code>--ignore-missing-bcls --ignore-missing-filter --ignore-missing-positions --use-bases-mask y4n*,n*</code></pre> <p>This work was supported by the R&D group at the SNP&SEQ Technology Platform in Uppsala. This facility is part of the National Genomics Infrastructure (NGI) Sweden and Science for Life Laboratory. The SNP&SEQ Platform is also supported by the Swedish Research Council and the Knut and Alice Wallenberg Foundation.</p>
Draft genome assemblies of killifish from the Fundulus genus with ONT and Illumina sequencing platforms
<p>Four species from the genus Fundulus were selected for genome sequencing to study the physiological and genetic mechanisms that diverge between euryhaline and stenohaline freshwater species within this cyprinodontiform order of ray-finned fishes.</p>
Development of Illumina oligos for fungal ITS sequencing
New PCR primers were designed with the following objectives: 1) targetting the ITS2 region to minimize bias due to introns at the 3' end of the SSU, 2) have the broadest possible coverage of Fungi, 3) minimize amplification of plant, protist and other Eukaryote sequences, 4) compatibility with Illumina sequencing adaptors and GOLAY barcodes. Primer regions were selected by eye from sequence alignments then further analyzed using PrimerProspector. At present, several mock communities and hundreds of soil DNA extracts have been successfully sequenced on the Illumina MiSeq platform using the oligos that were developed.
Garfagnina goats with Illumina CaprineSNP50 BeadChip
<p>The objective of this study was to investigate the genetic diversity of the Garfagnina (GRF) goat, a breed that currently risks extinction. For this purpose, 48 goats were genotyped with the Illumina CaprineSNP50 BeadChip and analyzed together with 214 goats belonging to 9 other Italian breeds (~25 goats/breed) from the AdaptMap project [Argentata (ARG), Bionda dell'Adamello (BIO), Ciociara Grigia (CCG), Di Teramo (DIT), Garganica (GAR), Girgentana (GGT), Orobica (ORO), Valdostana (VAL) and Valpassiria (VSS)]. Comparative analyses were conducted on i) runs of homozygosity (ROH), ii) admixture ancestries and iii) the accuracy of breed traceability via discriminant analysis on principal components (DAPC) based on cross-validation. For GRF, an excess of ROH (more than 45% in GRF samples) was detected on CHR 12 at, roughly 50.25-50.94Mbp (ARS1 assembly), which spans the CENPJ (centromere protein) and IL17D (interleukin 17D) genes. The same area of excess ROH was also present in DIT, while a broader region (~49.25-51.94Mbp) was shared among the ARG, CCG, and GGT. Admixture analysis revealed a small region of common ancestry from GRFshared by BIO, VSS, ARG and CCG breeds. The DAPC model yielded 100% assignment success for GRF. Overall, our results support the identification of GRF as a distinct native Italian goat breed. This work can contribute to planning conservation programmes to save GRF from extinction and will improve the understanding of the socio-agro-economic factors related with the farming of GRF.</p>
.bam alignment files of Illumina and ONT sequencing of pREF plasmid
<p>The expression of genes encompasses their transcription into mRNA followed by translation into protein. In recent years, next-generation sequencing and mass spectrometry methods have profiled DNA, RNA and protein abundance in cells. However, there are currently no reference standards that are compatible across these genomic, transcriptomic and proteomic methods, and provide an integrated measure of gene expression. Here, we use synthetic biology principles to engineer a multi-omics control, termed <em>pREF</em>, that can act as a universal molecular standard for next-generation sequencing and mass spectrometry methods. The <em>pREF</em> sequence encodes 21 synthetic genes that can be <em>in vitro</em> transcribed into spike-in mRNA controls, and <em>in vitro</em> translated to generate matched protein controls. The synthetic genes provide qualitative controls that can measure sensitivity and quantitative accuracy of DNA, RNA and peptide detection. We demonstrate the use of <em>pREF</em> in metagenome DNA sequencing and RNA sequencing experiments and evaluate the quantification of proteins using mass spectrometry. Unlike previous spike-in controls, <em>pREF</em> can be independently propagated and the synthetic mRNA and protein controls can be sustainably prepared by recipient laboratories using common molecular biology techniques. Together, this provides the first universal synthetic standard able to integrate genomic, transcriptomic and proteomic methods.</p>
Illumina Sequencing Data for "Phocaeicola vulgatus shapes the long-term growth dynamics and evolutionary adaptations of Clostridioides difficile"
<p>Illumina Sequencing Data for "<em>Phocaeicola vulgatus</em> shapes the long-term growth dynamics and evolutionary adaptations of <em>Clostridioides difficile</em>"</p>
Illumina TruSeq stranded mRNA sequences of Rhynchosporium commune isolate UK7 in plantae
<p>Transcription profiles were generated from the <em>Rhynchosporium commune</em> isolate UK7. RNA-seq experiments were conducted <em>in plantae</em> on the barley cultivar Beatrix (Viskosa 9 Pasadena, Saaten Union, breeders’ Reference NS01/2449). Leaves were collected at 9 and at 13 days post infection (dpi). All experiments were conducted in triplicates. Total RNA was extracted using TRIzol (Invitrogen Inc.) following the manufacturer’s recommendations. RNA integrity and quantity was assessed on a Bioanalyzer 2100 (Agilent) and a Qubit fluorometer (Life Technologies) and a Bioanalyzer 2100. Libraries were prepared using the TruSeq stranded mRNA sample prep kit (Illumina Inc.). Total RNA was ribosome-depleted by using polyA selection and reverse-transcribed into double-stranded cDNA.</p>
Illumina TruSeq stranded mRNA sequences of Rhynchosporium commune isolate UK7 in vitro
<p>Transcription profiles were generated from the <em>Rhynchosporium commune</em> isolate UK7. Cultures were grown either on Luria-Bertani (LBA) or Potato Dextrose Broth (PDB) media. Total RNA was extracted using TRIzol (Invitrogen Inc.) following the manufacturer’s recommendations. Experiments were conducted in triplicates. RNA integrity and quantity was assessed on a Bioanalyzer 2100 (Agilent) and a Qubit fluorometer (Life Technologies) and a Bioanalyzer 2100. Libraries were prepared using the TruSeq stranded mRNA sample prep kit (Illumina Inc.). Total RNA was ribosome-depleted by using polyA selection and reverse-transcribed into double-stranded cDNA.</p>
Illumina HD genotypes for 3,092 cattle from Burkina Faso, Ghana, Nigeria and Tanzania for: "Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle"
<p>Raw HD data for Riggio et al. 2022: Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle</p> <p>This repository contains the raw Illumina HD genotypes (i.e., 777,962 SNPs) mapped to the bovine UMD3.1 genome assembly for 3,092 animals from four African countries (namely Burkina Faso, Ghana, Nigeria and Tanzania). </p>
VADER TyrT Illumina sequencing data
<p>FASTQ files used in data analysis for "Virus-assisted directed evolution of enhanced suppressor tRNAs in mammalian cells."</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.