Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
443
datasets available to search
ShareScore release 0.9.0
Dataset results
443 results for “galaxy”
Training data for 'Genetic map RADSeq ' tutorial (Galaxy Training Material)
<p>The data provided here are part of a study published by Amores<em> et al.</em> (2011) (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3176089/">doi 10.1534/genetics.111.127324</a>), exploiting massively parallel DNA sequencing to develop meiotic maps by genotyping F<sub>1</sub> offspring of a single female and a single male spotted gar (<em>Lepisosteus oculatus</em>).</p>
Dataset for RNA-seq basic tutorial in Galaxy Australia
<p>Dataset of synthetic RNA-seq data for Drosophila melanogaster: 2 conditions, 3 replicates in each, paired-end reads. Reference genome in GTF format. </p>
The Binary Fraction of Stars in Dwarf Galaxies: The Cases of Draco and Ursa Minor
<p>supplementary data products, including all sky-subtracted spectra from individual targets, as well as random draws from posterior PDFs for model parameters (see enclosed README file)</p>
Training material for analysis small RNA-seq data (Galaxy Training Network tutorial)
<p>The data provided here is part of the Galaxy Training Network tutorial for analysis of small RNA-seq (sRNA-seq) data using mirdeep2 and miranda. This dataset is provided by INRA (Le Rheu, France).</p>
Nanopore sequence analysis - Galaxy Training Material
<p>Twelve MDR plasmids harboring samples were prepared according to the MinION library construction protocols, followed by library sequencing. After 8 hours of sequencing run, a total of 287 725 reads ranging from dozens to tens of thousands of bases in length were obtained, covering a total of 493 Mbp. The raw data were subjected to several stages of processing, including basecalling, de-multiplexing, fasta sequence extraction. For this tutorial one out of the twelve samples is chosen as example.</p> <p>This dataset is extracted of a project studying the Efficient generation of complete sequences of MDR-encoding plasmids by rapid assembly of MinION barcoding sequencing data (<a href="https://doi.org/10.1093/gigascience/gix132">https://doi.org/10.1093/gigascience/gix132</a>)</p>
Trimmed RNASeq pair for the Galaxy Training Network tutorial - "Metatranscriptomics analysis using microbiome RNASeq data"
<p>Functional microbiome analysis which estimates the functional groups expressed by microbial community enables researchers to look beyond taxonomic composition and correlation with the condition under study. Using microbial community RNA-Seq data and subsequent metatranscriptomics workflows to elucidate the functional complement of the microbiome is gaining interest in the field. <br> This Galaxy training network tutorial will introduce researchers to the basic concepts and tools from the published ASaiM workflow (Batut et al, <em>GigaScience</em> (2018), 7 (6),<a href="http://dx.doi.org/10.1093/gigascience/giy057"> http://dx.doi.org/10.1093/gigascience/giy057</a>). </p> <p>The dataset is a trimmed version of one of the time points from a cellulose degradation biogas reactor dataset. The dataset has been trimmed to facilitate running the workflows for this tutorial. Any biological interpretation from the results would be incorrect, due to the trimmed version of the dataset.</p>
Stellar Density Profiles of Dwarf Spheroidal Galaxies: Posterior Distributions
<p>This archive contains samples from the posterior distributions of our 3-plummer, 1-plummer, 3-steeper and 1-steeper fits for 38 dwarf spheroidal galaxies.</p>
NGSAP-VC : Genomic Variant Calling as an Installable GALAXY Workflow Using NGS data.
<p>Implementation of genomic variants calling as an installable GALAXY workflows using NGS data. Repository contains two separate sets of simulated ebola test data. One for SNPs and INDELs calling and another for Structural Variants calling.</p>
Training data for 'Beacon' tutorial (Galaxy Training Material)
<p>The data files are from the 1000 Genomes Project (1000HG) and GDC database. These datasets will be utilized in the Galaxy training session titled "Working with Beacon V2: A Comprehensive Guide to Creating, Uploading, and Searching for Variants with Beacons and Querying the University of Bradford GDC Beacon Database for Copy Number Variants (CNVs)." This training aims to equip participants with the skills necessary to construct Beacons, prepare and transform data into Beacon-compatible formats, seamlessly import data, and proficiently query Beacons for genetic variants. The provided data sets are integral for hands-on practice and will guide users through working with Beacon V2.</p>
Isoform analysis Galaxy training with EWS-FLI1 example
<p>The fastqs are subsampled from the following SRA:</p> <p><br>SRR8561337<br>SRR8561334<br>SRR8561335<br>SRR8561332<br>SRR8561333<br>SRR8561347</p> <p>The gtf file was reduced from https://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz and only lines with 'chr10' in first column were kept.</p> <p>The fasta file comes from https://hgdownload.soe.ucsc.edu/goldenPath/hg19/chromosomes/chr10.fa.gz</p>
Test data for Galaxy IUC Seurat_v5 tools
Open the record for dataset details and reuse information.
Data for DIMet Galaxy (Bioprotocol)
<p>Dataset for the DIMet Galaxy step-by-step use case, in preparation for submission to the Bioprotocol journal.</p> <p>This dataset was originally published by <a href="https://www.embopress.org/doi/full/10.15252/emmm.202115343">Guyon J, <em>et al</em>., 2022</a>. The subset of it, with exclusively LDHAB KO and Control samples at 0h, 24h and 48h, is the main file of the present Zenodo record. Also, the Differentially expressed genes (DEG) when comparing LDHAB KO <em>vs</em>. Control at 0 h and 48 h are available and explained in the '<em>Instructions for users'</em> below.</p> <p>The present dataset serves for demonstrating the step-by-step usage of the Galaxy versions of <a href="https://github.com/cbib/TraceGroomer/wiki">TraceGroomer</a> and <a href="https://github.com/cbib/DIMet/wiki">DIMet</a> tools.</p> <p><em>Instructions for users:</em></p> <ul> <li>Please download and unzip the entire folder, in your machine.</li> <li>Take into account the files description below <ul> <li>'dimet_bioprotocol/' folder: <br> <table> <tbody> <tr> <td><strong>Type of file</strong></td> <td><strong>File name</strong></td> </tr> <tr> <td>Labeled metabolomics data as received from Metabolomics facility (main file)</td> <td>isocor_LDHABKO_Ctrl.tsv</td> </tr> <tr> <td>Samples metadata (the file that explains the experimental setup)</td> <td>metadata_LDHABKO_Ctrl.tsv</td> </tr> </tbody> </table> </li> <li>'dimet_bioprotocol/metabologram_data/' subfolder: <table> <tbody> <tr> <td><strong>Type of file</strong></td> <td><strong>File name</strong></td> </tr> <tr> <td>file of DEG at T0 (0 h)</td> <td>LDHABKO_Ctrl_0h_DEG.tsv</td> </tr> <tr> <td>file of DEG at T48 (48 h)</td> <td>LDHABKO_Ctrl_48h_DEG.tsv</td> </tr> <tr> <td>file with pathways* for metabolites</td> <td>pathways_metabolites_custom.tsv</td> </tr> <tr> <td>file with pathways* for transcripts (or genes)</td> <td>pathways_genes_list_custom.tsv</td> </tr> </tbody> </table> *the lists of pathways in the file are disposed in columns, for each column: the top cell is the name of the pathway, whereas the rest of the cells are the elements belonging to that pathway.</li> </ul> </li> </ul> <p>For any inquiries or help, contact <a href="mailto:deisy-johanna.galvis-rodriguez@u-bordeaux.fr">Johanna Galvis.</a></p> <p> </p>
Test data for Galaxy IUC Seurat Inspect & Manipulate Tool (Merge)
Open the record for dataset details and reuse information.
Lotus2 on Mycorrhizal Fungi in the Galaxy Training Network - Sample files for creating Mapping TSV
<p>Sample files for the last step on creating a Mapping TSV in the Galaxy Training Network tutorial on "Identifying Mycorrhizal Fungi from ITS2 sequencing using LotuS2"</p>
Appendix Figures B.1–B.7 from the article 'Core Prominence as a Signature of Restarted Jet Activity in the LOFAR Radio-Galaxy Population' (Accepted for publication in the journal Astronomy & Astrophysics on August 23, 2024)
<p><strong>Figure captions:</strong></p> <p> </p> <p><strong>Figs. B.1–B.5.</strong> Images of 69 candidate restarted galaxies selected based on high radio $\mathrm{CP_{1400}}$ combined with low SB of extended emission, a steep spectrum of the core, and USS extended emission coupled with a bright core and summarised in Tables A.1 and A.2. Radio contours from VLA FIRST maps (white, $5^{\prime\prime}$), LOFAR high-resolution maps (black, $6^{\prime\prime}$), and NVSS maps (purple, $45^{\prime\prime}$) are overlaid on the LOFAR low-resolution resolution maps (orange, $20^{\prime\prime}$). The contouring of all the maps is made at $\,\sigma_\mathrm{local}\times(-3,3,5,10,20,30,40,50,100,150,200)$ levels, with $\sigma_\mathrm{local}$ representing the local RMS noise of the corresponding maps. The host galaxy position is marked with a yellow cross.</p> <p> </p> <p><strong>Figs. B.6–B.7. </strong>Images of sources excluded from the sample of restarted candidates following the criteria discussed in Sect. 3.1, Sect. 3.2 and Sect. 3.3 and summarised in Tables A.3 and A.4. Radio contours from VLA FIRST maps (white, $5^{\prime\prime}$), LOFAR high-resolution maps (black, $6^{\prime\prime}$), and NVSS maps (purple, $45^{\prime\prime}$) are overlaid on the LOFAR low-resolution resolution maps (orange, $20^{\prime\prime}$). The contouring of all the maps is made at $\,\sigma_\mathrm{local}\times(-3,3,5,10,20,30,40,50,100,150,200)$ levels, with $\sigma_\mathrm{local}$ representing the local RMS noise of the corresponding maps. The host galaxy position is marked with a yellow cross.</p> <p> </p>
"Filaments in and between galaxy clusters at low and mid-frequency with the SKA telescope" (Figures 6 to 13, B.1 and B.2)
<p>This file contains Figures 6 to 13, B.1 and B.2, from the paper "Filaments in and between galaxy clusters at low and mid-frequency with the SKA telescope", Vacca et al., accepted for publication on A&A on 25 September 2024.</p>
CO emission line spectra of the detected IRAM 30m CO-CAVITY galaxies.
<p>Figures show the observed spectra of the CO(1-0) and CO(2-1) emission lines of the CO-CAVITY galaxies.</p>
JWST spectra and best-fit SED models of massive quiescent galaxies from the Blue Jay survey.
<ul> <li>The JWST/NIRSpec spectra are stored as FITS tables which include wavelength (in angstrom, in rest-frame), calibrated flux, uncertainty, and best-fit model (in 1e-19 erg/s/cm^2/AA).</li> </ul>
Datasets for Galaxy Collection Operations Tutorial
<p>This repository contains datasets using in Galaxy tutorial introducing users to dataset collection operations: https://training.galaxyproject.org/training-material/topics/galaxy-interface/tutorials/collections/tutorial.html</p>
Training data for MaxQuant and Msstats TMT analysis in Galaxy
<p>The files serve as input and intermediate results for a MaxQuant and MsstatsTMT training on lysine methyl transferase 9 knockdown and control cell proteomics (https://doi.org/10.1186/s12935-020-1141-2) in the Galaxy training network (https://training.galaxyproject.org).</p> <p>Input files: human FASTA protein database for Maxquant. MaxQuant experimental design template, MSstatsTMT annotation file</p> <p>Intermediate result files: MaxQuant protein groups and evidence</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.