Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

47

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

47 results for “Galaxy Training Material”

Learn how ShareScore rates datasets ↗
zenodo40/100

Training data for 'Somatic variant calling' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates identification of somatic and germline variants from tumor and normal sample&nbsp;pairs.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Training material for the course "Exome analysis with GALAXY"

<p>Galaxy is an open source, web-based platform for data intensive biomedical research. It makes accessible bioinformatics applications to users lacking programming skills, enabling them to easily build analysis workflows for NGS data.<br /> &nbsp;<br /> The course &quot;<strong>Exome analysis using Galaxy</strong>&quot; is aimed at PhD student, biologists, clinicians and researchers who are analysing, or need to analyse in the near future, high throughput exome sequencing data. The aim of the course is to make participants familiarise with the Galaxy platform and prepare them to work independently, using state-of-the art tools for the analysis of exome sequencing data.</p> <p>The course will be delivered using a mixture of lectures and computer based hands-on practical sessions. Lectures will provide an up-to-date overview of the strategies for the analysis of exome next-generation experiments, starting from the raw sequence data. Analyses include sequence quality control, alignment to a reference genome, refinement of aligned sequences, variant calling, annotation and interpretation, and tools for visual inspection of results. Participants will apply the knowledge gained during the course to the analysis of Illumina&rsquo;s real exome datasets, and implement workflows to reproduce the complete analysis. After the course, participants will be able to create pipeline for their individual analyses.</p> <p>Those are the needed datasets for this course.</p>

opencc-zeroSep 2016View details →
zenodo40/100

Training data for 'Genome annotation with Apollo' tutorial (Galaxy Training Material)

<p>Published scaffolds from the Apis mellifera assembly Amel_4.5 and Official Gene Set 3.2.</p> <p>Source:&nbsp;<a href="http://hymenopteragenome.org/beebase/?q=download_sequences">http://hymenopteragenome.org/beebase/?q=download_sequences</a></p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

LFY ChIP-SEQ analysis Galaxy Training Material

<p>Datasets for Galaxy Training on ChIP-SEQ analysis. Raw files can be downloaded from SRA project&nbsp;SRP051214</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value &lt; 0.0001).</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Training data for 'Unicycler assembly of SARS-CoV-2 genome with preprocessing to remove human genome reads' tutorial (Galaxy Training Material)

<p>The data here is a copy of the corresponding SRR records in the NCBI SRA. The duplication serves a dual purpose:</p> <ol> <li>as a backup should there be problems connecting to NCBI servers, e.g., during Galaxy user trainings.</li> <li>to illustrate how to obtain raw sequencing data from alternative sources, and to organize the data into the same collection structure in a Galaxy history that is generated by specialized Galaxy SRA download tools.</li> </ol>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Alevin in commandline - Galaxy Training Material

<p>Input datasets for Generating a single cell matrix using Alevin (bash + R) tutorial on Galaxy Training Network.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Training data for 'Mapping-by-sequencing' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates mapping-by-sequencing analysis and represent a subsample of the data used in Sun &amp; Schneeberger, 2015 (DOI:10.1007/978-1-4939-2444-8_19).</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

Galaxy Hi-C Training material dm3

<p>Hi-C data for Galaxy training, dm3 cells.</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Training data for 'Genetic map RADSeq ' tutorial (Galaxy Training Material)

<p>The data provided here are part of a study published by Amores<em> et al.</em> (2011) (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3176089/">doi 10.1534/genetics.111.127324</a>), exploiting massively parallel DNA sequencing to develop meiotic maps by genotyping F<sub>1</sub> offspring of a single female and a single male spotted gar (<em>Lepisosteus oculatus</em>).</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

Training material for analysis small RNA-seq data (Galaxy Training Network tutorial)

<p>The data provided here is part of the Galaxy Training Network tutorial for analysis of small RNA-seq (sRNA-seq) data using mirdeep2 and miranda. This dataset is provided by INRA (Le Rheu, France).</p>

opencc-by-4.0Apr 2019View details →
zenodo36/100

Nanopore sequence analysis - Galaxy Training Material

<p>Twelve MDR plasmids harboring samples were prepared according to the MinION library construction protocols, followed by library sequencing. After 8 hours of sequencing run, a total of 287 725 reads ranging from dozens to tens of thousands of bases in length were obtained, covering a total of 493 Mbp. The raw data were subjected to several stages of processing, including basecalling, de-multiplexing, fasta sequence extraction. For this tutorial one out of the twelve samples is chosen as example.</p> <p>This dataset is extracted of a&nbsp;project studying the Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data (<a href="https://doi.org/10.1093/gigascience/gix132">https://doi.org/10.1093/gigascience/gix132</a>)</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

Training data for 'Beacon' tutorial (Galaxy Training Material)

<p>The data files are from the 1000 Genomes Project (1000HG) and GDC database. These datasets will be utilized in the Galaxy training session titled "Working with Beacon V2: A Comprehensive Guide to Creating, Uploading, and Searching for Variants with Beacons and Querying the University of Bradford GDC Beacon Database for Copy Number Variants (CNVs)." This training aims to equip participants with the skills necessary to construct Beacons, prepare and transform data into Beacon-compatible formats, seamlessly import data, and proficiently query Beacons for genetic variants. The provided data sets are integral for hands-on practice and will guide users through working with Beacon V2.</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater ampliconic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater ampliconic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater metatranscriptomic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater metatranscriptomic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Scanpy Parameter Iterator - Galaxy Training Material

<p>Input dataset for the&nbsp;Scanpy Parameter Iterator tutorial on Galaxy Training Network. It is an extension of <a href="https://training.galaxyproject.org/training-material/topics/single-cell/tutorials/scrna-case_basic-pipeline/tutorial.html">Filter, Plot and Explore Single-cell RNA-seq Data</a> tutorial.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Downsampled AnnData input file - Galaxy Training Material

<p>Downsampled input dataset for the Single Cell Data Formats Conversion tutorial on Galaxy Training Network.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

CDS input for Monocle3 tutorial - Galaxy Training Material

<p>CDS input file for Monocle3 trajectory analysis tutorial. Created from AnnData object from the upstream pre-processing.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Combining datasets after Alevin pre-processing - Galaxy Training Material

<p>These are the downsampled datasets for one of the single cell case study tutorials in Galaxy: "Combining single cell datasets after pre-processing".&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Training data for 'Maximum Likelihood Phylogeny Reconstruction'' (Galaxy Training Material)

<p>This data is used for Galaxy Training Network (GTN) training &#39;Maximum Likelihood Phylogeny Reconstruction&#39;. It consists of 173 amino acid alignments of orthologs found in chromosome 5 of four strains of S. cerevisiae. Original sequence data (https://zenodo.org/record/6610704) was processed in Galaxy following GTN &#39;Preparing genomic data for phylogeny reconstruction&#39; training (10.48546/workflowhub.workflow.359.1) to generate alignments of orthologs.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record