Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
146
datasets available to search
ShareScore release 0.9.0
Dataset results
146 results for “data workflow”
Test data for Imputation Workflow
<p>Test data for Imputation Workflow</p>
DataverseNO curation workflow used by the Research data team at the University Library of Bergen
<p>DataverseNO curation workflow used by the Research data team at the University Library of Bergen.</p> <p>Original Google Slides are available here: https://docs.google.com/presentation/d/1fxpAnvca7jisqjb0oqD1L3N3Kxj4hrpn_ywxSEl6-RU/edit?usp=sharing</p>
Data and analysis pipelines used in Increasing workflow development speed and reproducibility with Vectools
<p>The methods, pipelines, and data used for the examples shown in the Vectools manuscript.</p>
Test data for running snakePipes : WGBS workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the WGBS workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for the mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Test data for running snakePipes : RNA-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the RNA-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Test data for running snakePipes : HiC workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the HiC workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for the mouse (mm9) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example command in <strong>README</strong>.</li> </ul>
Test data for running snakePipes : ATAC-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ATAC-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for fruit fly (<strong>dm6</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Test data for running snakePipes : ChIP-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ChIP-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for human (<strong>hg38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
pyKNEEr: An image analysis workflow for open and reproducible research on femoral knee cartilage - Validation data
<p>Image data used in the paper introducing pyKNEEr. Explanations about these data are in the <a href="https://github.com/sbonaretti/pyKNEEr/tree/master/publication">GitHub</a> repository</p> <p>Changes in version 0.2.0: </p> <p>- Added inHouse images </p> <p>- Segmented images casted to int16 for smaller file size</p>
A Unified Multi-Wavelength Data Analysis Workflow with gammapy.
<p>Repository containing multi-wavelength (MWL) data of the distant quasar OP 313 (z=0.997) taken during the nights of MJD 60373 (3-4 March 2024) and MJD 60384 (14-15 March 2024) using <em>Fermi</em>-LAT, <em>Swift</em>-XRT, <em>Swift</em>-UVOT, and the Liverpool IO:O photometric data.</p> <div> <h3>Directory structure:</h3> <a href="https://github.com/mireianievas/gammapy_mwl_workflow/tree/main#directory-structure"></a></div> <ul> <li>Notebooks/DatasetGenerator: Notebooks (one per instrument) that summarize the steps to generate gammapy-compliant 1D and 3D binned datasets.</li> <li>Notebooks/DatasetAnalysis: Notebooks (one per instrument) to analyse each dataset independently and one notebook to perform the MWL joint analysis.</li> <li>Helpers: Auxiliary functions to generate native multiplicative models for dust extinction (out of xspec's redden), neutral hydrogen (out of xspec's tbabs), EBL absorption, and utility functions for file handling and plotting.</li> <li>Models: multiplicative models in tabular format for easy 2d interpolation.</li> <li>Figures: collection of figures for the paper.</li> </ul> <p>Link: <a href="https://github.com/mireianievas/gammapy_mwl_workflow">https://github.com/mireianievas/gammapy_mwl_workflow/ </a></p> <p>Based on the original work from <a href="https://github.com/luca-giunti/gammapyXray">https://github.com/luca-giunti/gammapyXray</a></p>
Source data for "Benchmarking commonly used software suites and analysis workflows for DIA proteomics and phosphoproteomics"
<p>Source data of "Benchmarking commonly used software suites and analysis workflows for DIA proteomics and phosphoproteomics".</p> <p>For reproduction of main and supplementary figures.</p> <p>MS raw files, spectral libraries, and MS data search results are stored in iProX with identifier IPX0004576001.</p> <p> </p>
Raw data of workflow execution results used in tonkaz's experiments
<p>This constitutes the raw data of workflow execution results employed in <a href="http://github.com/sapporo-wes/tonkaz">Tonkaz</a>'s experiments.</p> <p>Further information regarding the generation methods and additional details can be found at <a href="https://github.com/sapporo-wes/tonkaz/tree/main/tests">https://github.com/sapporo-wes/tonkaz/tree/main/tests</a>."</p> <p>The contents of this deposit are basically licensed under <a href="https://spdx.org/licenses/CC0-1.0.html">the Creative Commons Zero v1.0 Universal</a>.<br> However, there are files that could be licensed under other licenses, such as the nf-core workflow and its dependencies.<br> Because Zenodo does not provide the capability to attach licenses to individual files, we have described the licenses for these workflows in license.txt and ro-crate-metadata.json.<br> Please check them.</p>
Fig. 5. A in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 5. A boxplot showing the ion intensities of MS/MS feature 10 (gallic acid) in extracts which were active (IC50 <30 μg/mL) and inactive against α-glucosidase.
Fig. 4 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 4. Discrimination of the analyzed Alnus extracts into chemogroups. The analyzed extracts can be discriminated into three chemogroups by visualizing the CSCS distance metric between samples as PCoA plot (A) and chemical dendrogram (B). On the other hand, conventional methods such as PCA score plot (C) or hierarchical clustering analysis (HCA) using the Euclidean distance (D; chemogroups 1–3 are visualized with same colors used in B to make it easy to be compared) could not discriminate the samples into the same chemotypes. By mapping the chemogrouping of samples on the molecular network, it could be visualized that the three chemogroups were rich in diarylheptanoid, flavonoid, and tannins, respectively (E). (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 3 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 3. MS2LDA-driven substructural annotation of diarylheptanoids of Alnus species. Integrated with GNPS library matching and NAP in silico annotation, diarylheptanoid-related Mass2Motifs 41, 49, 72, and 81 could be characterized and correlated with specific substructures of diarylheptanoid aglycones. Scaffold diversity within diarylheptanoid molecular families A, D, and I were revealed by mapping these Mass2Motifs on the molecular network with different colors. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 1 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 1. LC–MS base peak ion (BPI) chromatograms of 15 Alnus extracts. Gaps between chromatogram were added to visualize their difference, so y-axis values do not equal to the absolute intensities.
Fig. 2 in Assessing specialized metabolite diversity of Alnus species by a digitized LC-MS/MS data analysis workflow
Fig. 2. The MS/MS spectral network of specialized metabolites contained in the bark, twigs, leaves, and fruits of A. japonica, A. firma, A. hirsuta, and A. hirsuta var. sibirica. Spectral nodes are colored according to the mean precursor ion intensity per different plant parts: bark, twigs, leaves, and fruits. Molecular families A–I are highlighted.
Data from: Expanding the described metabolome of the marine cyanobacterium Moorea producens JHB through orthogonal natural products workflows
Open the record for dataset details and reuse information.
Data from: Specimens at the center: an informatics workflow and toolkit for specimen-level analysis of public DNA database data
Open the record for dataset details and reuse information.
Data from: A data-driven geospatial workflow to map species distributions for conservation assessments
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.