Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.7.1
Dataset results
47 results for “Galaxy Training Material”
Training data for 'Upload data to ENA' (Galaxy Training Material)
<p>The data here is a subset of the data published in 10.5281/zenodo.3732359 to be used in GTN 'Upload data to ENA' tutorial.</p> <p>Human traces have been removed following <a href="https://training.galaxyproject.org/training-material/topics/sequence-analysis/tutorials/human-reads-removal/tutorial.html">https://training.galaxyproject.org/training-material/topics/sequence-analysis/tutorials/human-reads-removal/tutorial.html</a></p> <p>We produced consensus sequences (*.fasta) for the Illumina PE data following SARS-CoV-2-PE-Illumina-WGS-variant-calling (https://workflowhub.eu/workflows/113?version=4), SARS-CoV-2-variation-reporting (https://workflowhub.eu/workflows/109?version=5) and COVID-19-consensus-construction (https://workflowhub.eu/workflows/138?version=4) workflows.</p>
Dataset for Training Material - Galaxy Workflow - Analyse unaligned ncRNAs
<p>Input dataset for Galaxy Training Material for the Analyze unaligned ncRNAs workflow.</p> <p>See https://github.com/galaxyproject/training-material for more information.</p>
Training data for 'Preparing genomic data for phylogeny reconstruction' (Galaxy Training Material)
<p>This data is used for Galaxy Training Network training 'Preparing genomic data for phylogeny reconstruction'. There are four nucleotide sequences from chromosome 5 of four strains of S. cerevisiae. The GenBank annotated sequenced were produced using 'funannotate predict annotation' (Galaxy Version 1.8.9+galaxy2) on the nucleotide sequences sequences. References: DOI: 10.1126/science.274.5287.546; DOI: 10.1126/science.1189015; DOI: 10.1016/j.cell.2016.08.020</p>
Galaxy Training Material for Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)
<p>This dataset contains the training data for the <strong>Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)</strong> GTN tutorial. It includes 3 GC-[EI+]-HRMS files from seminal plasma samples, the RECETOX Metabolome HR-[EI+]-MS library collected from mostly endogoenous compounds from MetaSci Human Metabolite Library, reference alkanes, sample metadata table, and preprocessed XCMS object.</p>
Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)
<p>The data used in this tutorial are a subset of the data published previously in <a href="https://zenodo.org/record/3243160">Training material for the course "Exome analysis with GALAXY"</a>. Credit for uploading the original data goes to Paolo Uva and Gianmauro Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p> </p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you'll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load and, in fact, without having to install GEMINI's bundled annotation data.</p>
Training data for 'Genome annotation with Maker' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with Maker.</p> <p>It is based on data used in <a href="http://weatherby.genetics.utah.edu/MAKER/wiki/index.php/MAKER_Tutorial_for_WGS_Assembly_and_Annotation_Winter_School_2018">another Maker tutorial</a>.</p> <p>The full genome was <a href="https://www.ncbi.nlm.nih.gov/genome/?term=Schizosaccharomyces%20pombe[Organism]&cmd=DetailsSearch">downloaded from NCBI</a>, and mitochondria sequence removed from it for simplicity.</p>
Training material for small RNA-seq data analysis (Galaxy Training Network tutorial)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes small RNA-seq (sRNA-seq) data from a study published by Harrington et al. (DOI:10.1186/s12864-017-3692-8) to detect differential abundance of various classes of endogenous short interfering RNAs (esiRNAs). The goal of this study was to investigate "connections between differential retroTn and hp-derived esiRNA processing and cellular location, and to investigate the potential link between mRNA 3’ end cleavage and esiRNA biogenesis." To this end, sRNA-seq libraries were constructed from triplicate <em>Drosophila</em> tissue culture samples under conditions of either control RNAi or RNAi knockdown of a factor involved in mRNA 3’ end processing, <em>Symplekin</em>. This dataset (GEO Accession: GSE82128) consists of single-end, size-selected, non-rRNA-depleted sRNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to a subset of interesting transcript features including: (1) transposable elements, (2) <em>Drosophila</em> piRNA clusters, (3) <em>Symplekin</em>, and (4) genes encoding mass spectrometry-defined protein binding partners of <em>Symplekin</em> from Additional File 2 in the indicated paper by Harrington et al. More details on features 1 and 2 can be found here: https://github.com/bowhan/piPipes/blob/master/common/dm3/genomic_features (piRNA_Cluster, Trn). All features are from the <em>Drosophila</em> genome Apr. 2006 (BDGP R5/<em>dm3</em>) release.</p>
Training material for flye genome assembly (Galaxy Training Network tutorial)
<p>Datasets are subsets of 3 public datasets (Mucor mucedo Fresen. NRRL 3635 Standard Draft genome sequencing with PacBio technology)</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336965[accn]</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336964[accn]</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336963[accn]</p> <p> </p> <p> </p>
Long reads training material for 'Quality Control' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for reads Quality Control.</p> <p>PacBio HiFi reads were provided by PacBio - GIAB sample HG002 (https://www.pacb.com/smrt-science/smrt-resources/datasets/) and was downsampled using seqtk (https://github.com/lh3/seqtk)</p> <p>Nanopore reads were provided by Tim Kahlke as part of "Long-Read, long reach Bioinformatics Tutorials" (https://timkahlke.github.io/LongRead_tutorials/) and was basecalled using Guppy v5.0.2 (dna_r9.4.1_450bps_sup.cfg).</p>
URL list for downloading training data for 'Maximum Likelihood Phylogeny Reconstruction'' (Galaxy Training Material)
<p>This data is used for Galaxy Training Network (GTN) training 'Maximum Likelihood Phylogeny Reconstruction'. It is a list of Zenodo URL pointers to a dataset of 173 amino acid alignments of orthologs found in chromosome 5 of four strains of S. cerevisiae. Original sequence data (https://zenodo.org/record/6610704) was processed in Galaxy following GTN 'Preparing genomic data for phylogeny reconstruction' training (10.48546/workflowhub.workflow.359.1) to generate alignments of orthologs.</p>
Training data for 'Functional annotation of protein sequences' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for functional annotation of protein sequences.</p>
Galaxy Training Material - Peak detection with recetox-aplcms
<p>This dataset contains the training data for the recetox-aplcms GTN Tutorial. It includes 3 LC-ESI+-HRMS DDA files containing a mixture of pesticides and 3 GC-EI+-HRMS files from seminal plasma samples.</p>
Training data for 'Refining Manual Genome Annotations with Apollo (eukaryotes)' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for manual curation of eukaryotic genome annotation using Apollo.</p>
Training material for Genome assembly quality control (Galaxy Training Network tutorial)
<p>This Zenodo repository includes the required datasets for following the GTN: Genome assembly quality control.</p>
Training data for 'Repeat masking with RepeatMasker' tutorial (Galaxy Training Material)
<p>Data needed for the 'Repeat masking with RepeatMasker' tutorial (Galaxy Training Material).</p> <p>The assembly was generated following the 'Genome assembly using PacBio data' tutorial</p>
Training data for ChIP-seq data analysis (Galaxy Training Material): Identification of the binding sites of the Estrogen receptor
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes ChIP-seq data from a study published by Ross-Inness et al., 2012 (DOI:10.1038/nature10730) to identify the binding sites of the Estrogen receptor, a transcription factor known to be associated with different types of breast cancer.</p>
Training data for 'From peaks to gene' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes peaks from a study published by Li et al., 2012 (DOI:10.1016/j.stem.2012.04.023) to identify target genes</p>
Training data for 'Reference based RADSeq ' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes RAD-seq data from a study published by Hohelnlohe et al., 2010 (DOI:10.1371/journal.pgen.1000862) to identify and type single nucleotide polymorphisms (SNPs) in each of 100 individuals from two oceanic and three freshwater populations and thus estimate genetic diversity and differentiation among populations. </p>
Training data for 'Long non-coding RNAs (lncRNAs) annotation with FEELnc' tutorial (Galaxy Training Material)
<p>Data needed for the 'Long non-coding RNAs (lncRNAs) annotation with FEELnc' tutorial (Galaxy Training Material).<br>The assembly was generated following the 'Genome assembly using PacBio data' tutorial.<br>The annotation was generated following the 'Genome annotation with Funannotate ' tutorial.</p> <p>The bam file is RNASeq SRR8534859_1.fastq.gz and SRR8534859_2.fastq.gz mapping on the genome assembly.</p>
Data for Galaxy CLIP-Seq Training Material
<p>The eCLIP data provided here is a subset of the eCLIP data of RBFOX2 from a study published by Nostrand et al. (2016, http://dx.doi.org/10.1038/nmeth.3810). The dataset contains the first biological replicate of RBFOX2 CLIP-seq and the input control experiment (*fastq files). The data was changed and downsampled to reduce data processing time, thus the datasets does not correspond to the original data pulled from Nostrand et al. (2016, http://dx.doi.org/10.1038/nmeth.3810). Also included is a text file (.txt) encompassing the chromosome sizes of hg19 and hg38 obtained from UCSC (http://hgdownload.cse.ucsc.edu/goldenPath/hg19/bigZips/hg19.chrom.sizes, http://hgdownload.cse.ucsc.edu/goldenPath/hg38/bigZips/hg38.chrom.sizes) and a genome annotation for hg19 (.gtf) taken from Ensembl (http://ftp.ensemblorg.ebi.ac.uk/pub/release-74/gtf/homo_sapiens/) and for hg38 taken from the Galaxy libraries (https://usegalaxy.eu/library/list#folders/F30cab321d898d2fb/datasets/9ba790aa79c9cf23). The data is used for a galaxy training course about CLIP-Seq data analysis. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.