Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.7.1
Dataset results
47 results for “Galaxy Training Material”
Training data for 'Somatic variant calling' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates identification of somatic and germline variants from tumor and normal sample pairs.</p>
Training material for the course "Exome analysis with GALAXY"
<p>Galaxy is an open source, web-based platform for data intensive biomedical research. It makes accessible bioinformatics applications to users lacking programming skills, enabling them to easily build analysis workflows for NGS data.<br /> <br /> The course "<strong>Exome analysis using Galaxy</strong>" is aimed at PhD student, biologists, clinicians and researchers who are analysing, or need to analyse in the near future, high throughput exome sequencing data. The aim of the course is to make participants familiarise with the Galaxy platform and prepare them to work independently, using state-of-the art tools for the analysis of exome sequencing data.</p> <p>The course will be delivered using a mixture of lectures and computer based hands-on practical sessions. Lectures will provide an up-to-date overview of the strategies for the analysis of exome next-generation experiments, starting from the raw sequence data. Analyses include sequence quality control, alignment to a reference genome, refinement of aligned sequences, variant calling, annotation and interpretation, and tools for visual inspection of results. Participants will apply the knowledge gained during the course to the analysis of Illumina’s real exome datasets, and implement workflows to reproduce the complete analysis. After the course, participants will be able to create pipeline for their individual analyses.</p> <p>Those are the needed datasets for this course.</p>
Training data for 'Genome annotation with Apollo' tutorial (Galaxy Training Material)
<p>Published scaffolds from the Apis mellifera assembly Amel_4.5 and Official Gene Set 3.2.</p> <p>Source: <a href="http://hymenopteragenome.org/beebase/?q=download_sequences">http://hymenopteragenome.org/beebase/?q=download_sequences</a></p>
LFY ChIP-SEQ analysis Galaxy Training Material
<p>Datasets for Galaxy Training on ChIP-SEQ analysis. Raw files can be downloaded from SRA project SRP051214</p>
Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value < 0.0001).</p>
Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)
<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>
Alevin in commandline - Galaxy Training Material
<p>Input datasets for Generating a single cell matrix using Alevin (bash + R) tutorial on Galaxy Training Network. </p>
Training data for 'Mapping-by-sequencing' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates mapping-by-sequencing analysis and represent a subsample of the data used in Sun & Schneeberger, 2015 (DOI:10.1007/978-1-4939-2444-8_19).</p>
Galaxy Hi-C Training material dm3
<p>Hi-C data for Galaxy training, dm3 cells.</p>
Training data for 'Genetic map RADSeq ' tutorial (Galaxy Training Material)
<p>The data provided here are part of a study published by Amores<em> et al.</em> (2011) (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3176089/">doi 10.1534/genetics.111.127324</a>), exploiting massively parallel DNA sequencing to develop meiotic maps by genotyping F<sub>1</sub> offspring of a single female and a single male spotted gar (<em>Lepisosteus oculatus</em>).</p>
Training material for analysis small RNA-seq data (Galaxy Training Network tutorial)
<p>The data provided here is part of the Galaxy Training Network tutorial for analysis of small RNA-seq (sRNA-seq) data using mirdeep2 and miranda. This dataset is provided by INRA (Le Rheu, France).</p>
Nanopore sequence analysis - Galaxy Training Material
<p>Twelve MDR plasmids harboring samples were prepared according to the MinION library construction protocols, followed by library sequencing. After 8 hours of sequencing run, a total of 287 725 reads ranging from dozens to tens of thousands of bases in length were obtained, covering a total of 493 Mbp. The raw data were subjected to several stages of processing, including basecalling, de-multiplexing, fasta sequence extraction. For this tutorial one out of the twelve samples is chosen as example.</p> <p>This dataset is extracted of a project studying the Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data (<a href="https://doi.org/10.1093/gigascience/gix132">https://doi.org/10.1093/gigascience/gix132</a>)</p>
Training data for 'Beacon' tutorial (Galaxy Training Material)
<p>The data files are from the 1000 Genomes Project (1000HG) and GDC database. These datasets will be utilized in the Galaxy training session titled "Working with Beacon V2: A Comprehensive Guide to Creating, Uploading, and Searching for Variants with Beacons and Querying the University of Bradford GDC Beacon Database for Copy Number Variants (CNVs)." This training aims to equip participants with the skills necessary to construct Beacons, prepare and transform data into Beacon-compatible formats, seamlessly import data, and proficiently query Beacons for genetic variants. The provided data sets are integral for hands-on practice and will guide users through working with Beacon V2.</p>
Detection of SARS-CoV-2 variants by genomic analysis of wastewater ampliconic samples (Galaxy Training Material)
<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater ampliconic samples. (https://training.galaxyproject.org/training-material/)</p>
Detection of SARS-CoV-2 variants by genomic analysis of wastewater metatranscriptomic samples (Galaxy Training Material)
<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater metatranscriptomic samples. (https://training.galaxyproject.org/training-material/)</p>
Scanpy Parameter Iterator - Galaxy Training Material
<p>Input dataset for the Scanpy Parameter Iterator tutorial on Galaxy Training Network. It is an extension of <a href="https://training.galaxyproject.org/training-material/topics/single-cell/tutorials/scrna-case_basic-pipeline/tutorial.html">Filter, Plot and Explore Single-cell RNA-seq Data</a> tutorial.</p>
Downsampled AnnData input file - Galaxy Training Material
<p>Downsampled input dataset for the Single Cell Data Formats Conversion tutorial on Galaxy Training Network. </p>
CDS input for Monocle3 tutorial - Galaxy Training Material
<p>CDS input file for Monocle3 trajectory analysis tutorial. Created from AnnData object from the upstream pre-processing.</p>
Combining datasets after Alevin pre-processing - Galaxy Training Material
<p>These are the downsampled datasets for one of the single cell case study tutorials in Galaxy: "Combining single cell datasets after pre-processing". </p>
Training data for 'Maximum Likelihood Phylogeny Reconstruction'' (Galaxy Training Material)
<p>This data is used for Galaxy Training Network (GTN) training 'Maximum Likelihood Phylogeny Reconstruction'. It consists of 173 amino acid alignments of orthologs found in chromosome 5 of four strains of S. cerevisiae. Original sequence data (https://zenodo.org/record/6610704) was processed in Galaxy following GTN 'Preparing genomic data for phylogeny reconstruction' training (10.48546/workflowhub.workflow.359.1) to generate alignments of orthologs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.