Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
443
datasets available to search
ShareScore release 0.9.0
Dataset results
443 results for “galaxy”
First Light And Reionisation Epoch Simulations (FLARES) IV: The size evolution of galaxies at z≥5
<p>Galaxy size results from the FLARES simulations. This is the companion dataset to: https://arxiv.org/abs/2203.12627 containing the data plotted within. The codes used to produce the data are available on GitHub: https://github.com/WillJRoper/flares-sizes-obs</p>
There and back again: understanding the critical properties of backsplash galaxies (Data)
<p>Data from the paper "There and back again: understanding the critical properties of backsplash galaxies (Data)" and code used to reduce it from the publicly available TNG-300 dataset.</p>
Gaia EDR3 in nearby galaxies
<p>Gaia EDR3 sources in the vicinity of nearby galaxies, as described in Barmby (2022, MNRAS submitted).</p>
Galaxy Training Data for "Evaluating and ranking a set of pathways based on multiple metrics"
<pre>This dataset provides the inputs needed for the Galaxy Pathway Analysis workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow asseses the performance of predicted pathways by computing 4 criteria (target product flux, thermodynamic feasibility, pathway length, and enzyme availability). A score inform the user about the best candidate pathways to produce a compound of interest. The generated output is a collection of scored and ranked heterologous pathways. The content of the dataset is as follows: - A set of pathways provided in the SBML format (Systems Biology Markup Language) to be ranked, modeling heterologous pathways such as those outputted by the RetroSynthesis workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). - The GEM (Genome-scale metabolic models) which is a formalized representation of the metabolism of the host organism (the model is E. coli iML1515), provided in the SBML format.</pre>
NED and SIMBAD in nearby galaxies
<p>Data files produced to compare contents of the NED and SIMBAD databases in the vicinity of nearby galaxies.</p> <p>ned_simbad_pos_offset.csv: comparison of NED and SIMBAD positions for galaxies searched by name</p> <p>gdm_output.csv: main output containing the category matches for each galaxy</p> <p>CS?_NEDQuery.csv: NED listing of individual objects within galaxy for case studies.</p> <p>CS?_SIMBADQuery.csv: SIMBAD listing of individual objects within galaxy for case studies.</p> <p>CS1_UGCA086.csv: match details for individual objects in UGCA086</p> <p>CS2_NGC2903.csv: match details for individual objects in NGC2903</p> <p>Files with _5as.csv represent queries to later versions of NED and SIMBAD (as of June 2022), and match results with a 5 arcsec tolerance instead of 1 arcsec.</p>
URL list for downloading training data for 'Maximum Likelihood Phylogeny Reconstruction'' (Galaxy Training Material)
<p>This data is used for Galaxy Training Network (GTN) training 'Maximum Likelihood Phylogeny Reconstruction'. It is a list of Zenodo URL pointers to a dataset of 173 amino acid alignments of orthologs found in chromosome 5 of four strains of S. cerevisiae. Original sequence data (https://zenodo.org/record/6610704) was processed in Galaxy following GTN 'Preparing genomic data for phylogeny reconstruction' training (10.48546/workflowhub.workflow.359.1) to generate alignments of orthologs.</p>
Training data for 'Functional annotation of protein sequences' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for functional annotation of protein sequences.</p>
Galaxy Training Material - Peak detection with recetox-aplcms
<p>This dataset contains the training data for the recetox-aplcms GTN Tutorial. It includes 3 LC-ESI+-HRMS DDA files containing a mixture of pesticides and 3 GC-EI+-HRMS files from seminal plasma samples.</p>
Training data for 'Refining Manual Genome Annotations with Apollo (eukaryotes)' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for manual curation of eukaryotic genome annotation using Apollo.</p>
Training material for Genome assembly quality control (Galaxy Training Network tutorial)
<p>This Zenodo repository includes the required datasets for following the GTN: Genome assembly quality control.</p>
SDSS Galaxy Subset
<p>The <a href="https://www.sdss.org">Sloan Digital Sky Survey</a> (SDSS) is a comprehensive survey of the northern sky. This dataset contains a subset of this survey, of 100077 objects classified as galaxies, it includes a CSV file with a collection of information and a set of files for each object, namely JPG image files, FITS and spectra data. This dataset is used to train and explore the <a href="https://github.com/nunorc/astromlp-models">astromlp-models</a> collection of deep learning models for galaxies characterisation.</p> <p>The dataset includes a CSV data file where each row is an object from the SDSS database, and with the following columns (note that some data may not be available for all objects):</p> <ul> <li><strong>objid</strong>: unique SDSS object identifier </li> <li><strong>mjd</strong>: MJD of observation</li> <li><strong>plate</strong>: plate identifier</li> <li><strong>tile</strong>: tile identifier</li> <li><strong>fiberid</strong>: fiber identifier</li> <li><strong>run</strong>: run number</li> <li><strong>rerun</strong>: rerun number</li> <li><strong>camcol</strong>: camera column</li> <li><strong>field</strong>: field number</li> <li><strong>ra</strong>: right ascension</li> <li><strong>dec</strong>: declination</li> <li><strong>class</strong>: spectroscopic class (only objetcs with GALAXY are included)</li> <li><strong>subclass</strong>: spectroscopic subclass</li> <li><strong>modelMag_u</strong>: better of DeV/Exp magnitude fit for band u</li> <li><strong>modelMag_g</strong>: better of DeV/Exp magnitude fit for band g</li> <li><strong>modelMag_r</strong>: better of DeV/Exp magnitude fit for band r</li> <li><strong>modelMag_i</strong>: better of DeV/Exp magnitude fit for band i</li> <li><strong>modelMag_z</strong>: better of DeV/Exp magnitude fit for band z</li> <li><strong>redshift</strong>: final redshift from SDSS data z</li> <li><strong>stellarmass</strong>: stellar mass extracted from the <a href="https://www.sdss.org/dr16/spectro/eboss-firefly-value-added-catalog">eBOSS Firefly catalog</a></li> <li><strong>w1mag</strong>: WISE W1 "standard" aperture magnitude</li> <li><strong>w2mag</strong>: WISE W2 "standard" aperture magnitude</li> <li><strong>w3mag</strong>: WISE W3 "standard" aperture magnitude</li> <li><strong>w4mag</strong>: WISE W4 "standard" aperture magnitude</li> <li><strong>gz2c_f</strong>: Galaxy Zoo 2 classification from <a href="https://academic.oup.com/mnras/article/435/4/2835/1022913">Willett et al 2013</a></li> <li><strong>gz2c_s</strong>: simplified version of Galaxy Zoo 2 classification (<a href="https://github.com/nunorc/astromlp-models#galaxy-zoo-2-simplified-classes-gz2c">labels set</a>)</li> </ul> <p>Besides the CSV file a set of directories are included in the dataset, in each directory you'll find a list of files named after the <strong>objid </strong>column from the CSV file, with the corresponding data, the following directories tree is available:</p> <pre><code class="language-bash">sdss-gs/ ├── data.csv ├── fits ├── img ├── spectra └── ssel</code></pre> <p>Where, each directory contains:</p> <ul> <li><strong>img</strong>: RGB images from the object in JPEG format, 150x150 pixels, generated using the <a href="https://skyserver.sdss.org/dr16/en/help/docs/api.aspx">SkyServer DR16 API</a></li> <li><strong>fits</strong>: FITS data subsets around the object across the u, g, r, i, z bands; cut is done using the <a href="https://github.com/jhoar/ImageCutter">ImageCutter</a> library</li> <li><strong>spectra</strong>: full best fit spectra data from SDSS between 4000 and 9000 wavelengths</li> <li><strong>ssel</strong>: best fit spectra data from SDSS for specific selected intervals of wavelengths discussed by <a href="https://arxiv.org/abs/1003.3186">Sánchez Almeida 2010</a></li> </ul> <p><strong>Changelog</strong></p> <ul> <li>v0.0.4 - Increase number of objects to ~100k.</li> <li>v0.0.3 - Increase number of objects to ~80k.</li> <li>v0.0.2 - Increase number of objects to ~60k.</li> <li>v0.0.1 - Initial import.</li> </ul>
Training data for 'Repeat masking with RepeatMasker' tutorial (Galaxy Training Material)
<p>Data needed for the 'Repeat masking with RepeatMasker' tutorial (Galaxy Training Material).</p> <p>The assembly was generated following the 'Genome assembly using PacBio data' tutorial</p>
Multiwavelength classification of X-ray selected galaxy cluster candidates using convolutional neural networks
<p>Classification dataset used in Kosiba et al. 2020 (10.1093/mnras/staa1723) consisting of candidate clusters in the XCLASS survey</p> <p>Training and testing images and corresponding labels low-z (clusters, 0<z<0.3), hi-z (clusters, z>0.3), nearby galaxy, point source (point, double source, star/AGN), and other (artefact, edge)</p>
Training data for ChIP-seq data analysis (Galaxy Training Material): Identification of the binding sites of the Estrogen receptor
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes ChIP-seq data from a study published by Ross-Inness et al., 2012 (DOI:10.1038/nature10730) to identify the binding sites of the Estrogen receptor, a transcription factor known to be associated with different types of breast cancer.</p>
Training data for 'From peaks to gene' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes peaks from a study published by Li et al., 2012 (DOI:10.1016/j.stem.2012.04.023) to identify target genes</p>
Disseminating metaproteomic informatics capabilities and knowledge using the Galaxy-P framework
<p>Data for the "<strong>Disseminating metaproteomic informatics capabilities and knowledge using the Galaxy-P framework</strong>" paper and training.</p>
Training data for 'Reference based RADSeq ' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes RAD-seq data from a study published by Hohelnlohe et al., 2010 (DOI:10.1371/journal.pgen.1000862) to identify and type single nucleotide polymorphisms (SNPs) in each of 100 individuals from two oceanic and three freshwater populations and thus estimate genetic diversity and differentiation among populations. </p>
Photometric Galaxy Redshift Prediction
<h3>Sloan Digital Sky Survey (SDSS) Galaxy Redshift Dataset</h3> <p>This dataset comprises a curated collection of galaxy observations from the Sloan Digital Sky Survey (SDSS). It features photometric and spectroscopic data for 100 galaxies, specifically selected to cover a range of redshifts from 0 to 0.4. The dataset includes the following key parameters for each galaxy:</p> <ul> <li><strong>Photometric Data</strong>: Magnitudes in the SDSS 'u', 'g', 'r', 'i', and 'z' bands.</li> <li><strong>Spectroscopic Data</strong>: Measured redshift (<code>redshift</code>) and its error (<code>redshift_error</code>).</li> <li><strong>Additional Metadata</strong>: <ul> <li><code>objid</code>: Unique identifier for the photometric object.</li> <li><code>specObjID</code>: Unique identifier for the spectroscopic object.</li> <li><code>ra</code>: Right ascension in decimal degrees.</li> <li><code>dec</code>: Declination in decimal degrees.</li> <li><code>class</code>: Classification of the object, all marked as 'GALAXY'.</li> </ul> </li> </ul> <h3>Purpose and Use</h3> <p>This dataset is intended for use in astronomical research and education, particularly in studies involving galaxy properties and distribution, cosmology, and machine learning applications such as redshift prediction models. The data is well-suited for developing and testing predictive models that estimate redshifts from photometric data, aiding in the expansion of accessible astronomical analysis tools.</p> <h3>Data Collection Method</h3> <p>The data was extracted using SQL queries against the public SDSS DR16 database, ensuring accuracy and relevance in current astronomical research contexts.</p> <h3>Accessibility</h3> <p>The dataset is made available under a CC0 license to promote open scientific research and collaboration within the astronomical community and beyond.</p>
Best-fit SED templates of each CAMIRA member galaxy (Table 3)
<p>This upload includes individual SED templates as described in Table 3 of the related manuscript, "<em>Active Galactic Nucleus Properties of ~1 Million Member Galaxies of Galaxy Groups and Clusters at z < 1.4 Based on the Subaru Hyper Suprime-Cam Survey</em>", by Yoshiki Toba et al., published in the Astrophysical Journal (Accepted: 15-Feb-2024; published: 17-May-2024, doi:<a href="https://doi.org/10.3847/1538-4357/ad32c6">10.3847/1538-4357/ad32c6</a>).</p> <p>This archive is compressed with xz, part of <a href="https://github.com/tukaani-project/xz">XZ Utils</a>. </p> <p>A description of these files is given below.</p> <p>The "ID" field is captured in the file naming in the compressed archive:</p> <ul> <li> <p>1_SED.eps [EPS plot of SED, SED fit, for Object ID = 1]</p> </li> <li> <p>1_SED_data.fits [Binary Table FITS (HDU=1) where the contents of HDU=1 are given in the table below]</p> </li> </ul> <p> </p> <table> <tbody> <tr> <td> <div> <div> <div> <p>Column name</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>Format</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>Unit</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>Description</p> </div> </div> </div> </td> </tr> <tr> <td> <div> <div> <div> <p>Wavelength</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>DOUBLE</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>μm</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>Wavelength (observed frame)</p> </div> </div> </div> </td> </tr> <tr> <td> <div> <div> <div> <p>FNU</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>DOUBLE</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>mJy</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>Flux density at each wavelength</p> </div> </div> </div> </td> </tr> <tr> <td> <div> <div> <div> <p>log_L</p> </div> </div> </div> </td> <td> <div> <div> <div> <p>DOUBLE</p> </div> </div> </div> </td> <td>erg/s</td> <td> <div> <div> <div> <p>Luminosity at each wavelength</p> </div> </div> </div> </td> </tr> </tbody> </table> <p> </p> <p> </p>
Training data for 'Long non-coding RNAs (lncRNAs) annotation with FEELnc' tutorial (Galaxy Training Material)
<p>Data needed for the 'Long non-coding RNAs (lncRNAs) annotation with FEELnc' tutorial (Galaxy Training Material).<br>The assembly was generated following the 'Genome assembly using PacBio data' tutorial.<br>The annotation was generated following the 'Genome annotation with Funannotate ' tutorial.</p> <p>The bam file is RNASeq SRR8534859_1.fastq.gz and SRR8534859_2.fastq.gz mapping on the genome assembly.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.