Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
109
datasets available to search
ShareScore release 0.9.0
Dataset results
109 results for “data pipelines”
MUSE HUDF survey I, Section 4: data and reproduction pipeline for photometry and astrometry
<p>Necessary data and <a href="http://akhlaghi.org/reproducible-science.html">Reproduction pipeline</a> for <a href="https://www.aanda.org/articles/aa/full_html/2017/12/aa30833-17/aa30833-17.html#S14">Section 4</a> of "<em>The MUSE Hubble Ultra Deep Field Survey: I. Survey description, data reduction and source detection</em>", Bacon et al. (2017), <a href="https://www.aanda.org/articles/aa/abs/2017/12/aa30833-17/aa30833-17.html">Astronomy & Astrophysics, 608, A1</a>. The purpose of this section in the paper is to show the photometric and astrometric precision of the processed <a href="http://muse-vlt.eu/science/">MUSE</a> 3D data cubes discussed in the paper (pseudo-broad-band images created from the cubes) in comparison with broad-band images of the Hubble Space Telescope (HST).</p> <p>This repository on Zenodo contains all the necessary input data, software and <a href="http://akhlaghi.org/reproducible-science.html">reproduction pipeline</a> (containing the scripts, configuration files and settings to exactly reproduce the results in Section 4 of the paper). Below is a description of the contents:</p> <ul> <li> <p><a href="https://zenodo.org/record/1163746/files/gnuastro-0.2.51-bc56.tar.gz"><code>gnuastro-0.2.51-bc56.tar.gz</code></a>: The version of <a href="https://www.gnu.org/software/gnuastro">GNU Astronomy Utilities</a> (Gnuastro) that is necessary for this pipeline. Gnuastro is a large collection of programs for astronomical data analysis on the command-line (and in scripts). Note that the reproduction pipeline <em>only</em> works with Gnuastro version 0.2.51, it will complain and abort if another version is installed.</p> <p>IMPORTANT NOTE: Since version 0.2.51 of Gnuastro was released, CFITSIO (one of Gnuastro's dependencies) has added a dependency for the cURL library (to read https URLs). Therefore, to install Gnuastro 0.2.51, please install <a href="https://heasarc.gsfc.nasa.gov/FTP/software/fitsio/c/cfitsio3410.tar.gz">CFITSIO version 3.41</a> or earlier.</p> </li> <li> <p><a href="https://zenodo.org/record/1163746/files/gnuastro-dependencies.tar.gz"><code>gnuastro-dependencies.tar.gz</code></a>: Software libraries necessary to build Gnuastro as it is used here. With these, a working C compiler is enough (currently only tested in a GNU/Linux environment) to exactly reproduce the results (tables).</p> </li> <li> <p><a href="https://zenodo.org/record/1163746/files/hst-acs-images.tar.gz"><code>hst-acs-images.tar.gz</code></a>: Necessary images from HST's <a href="https://archive.stsci.edu/prepds/xdf/">eXtreme Deep Field</a> survey <a href="https://archive.stsci.edu/pub/hlsp/xdf">archives</a>. These images are not necessary to run the reproduction pipeline (they will be downloaded from the HST archives if not present). They are stored here for the self-sufficiency of this repository and faster download: in this lossless compressed format, they are roughly 1/3rd the volume of the same files in HST archives.</p> </li> <li> <p><a href="https://zenodo.org/record/1163746/files/hst-acs-throughputs.tar.gz"><code>hst-acs-throughputs.tar.gz</code></a>: The throughputs of HST Advanced Camera for Surveys (ACS) filters necessary in this study. These are also available from the <a href="http://www.stsci.edu/hst/acs/analysis/throughputs/tables">HST archives</a> and are kept here with similar reasons to above.</p> </li> <li> <p><a href="https://zenodo.org/record/1163746/files/muse-pseudo-broadband-images.tar.gz"><code>muse-pseudo-broadband-images.tar.gz</code></a>: Pseudo-broad-band images generated from the MUSE 3D data cube. These images are only released in this repository. However, to run the reproduction pipeline, it isn't necessary to download them directly from here. The script will download them from Zenodo automatically.</p> </li> <li> <p><a href="https://zenodo.org/record/1163746/files/reproduce-v1-4-gaafdb04.tar.gz"><code>reproduce-v1-4-gaafdb04.tar.gz</code></a>: The <a href="http://akhlaghi.org/reproducible-science.html">reproduction pipeline</a> (version 1-4-gaafdb04) that produces the results (tables) plotted in the paper. The full Git version controlled history of this repository is available on <a href="https://git-cral.univ-lyon1.fr/mohammad.akhlaghi/muse-udf-photometry-astrometry">git-cral.univ-lyon1.fr</a> or <a href="https://gitlab.com/makhlaghi/muse-udf-photometry-astrometry">gitlab.com</a>. We recommend cloning from the Git repository if it is available. This tarball is kept here in case those servers don't work or Git is no longer in common use. Please see the <code>README</code> file in this repository for instructions on how to run the reproduction pipeline and exactly reproduce the results. This pipeline will download all the necessary data if they aren't already present on the system (it is probably just necessary to install the required version of Gnuastro).</p> </li> </ul> <p>The Creative Commons Attribution-NonCommercial 4.0 copyright mentioned in the Zenodo webpage is only applicable to files that don't have an explicit copyright within them. The copyright of other files (mainly scripts and software) is mentioned within them (all are <a href="https://www.gnu.org/licenses/licenses.en.html">free licenses</a>).</p> <p>For any issues with the pipeline/processing, please contact <a href="http://akhlaghi.org">Mohammad Akhlaghi</a>.</p>
IMC Segmentation Pipeline results of example IMC data
<p>If you are working with these files, please cite them as follows:<br><br>Windhager, J., Zanotelli, V.R.T., Schulz, D. et al. An end-to-end workflow for multiplexed image processing and analysis. Nat Protoc (2023). <a href="https://doi.org/10.1038/s41596-023-00881-0">https://doi.org/10.1038/s41596-023-00881-0</a></p><p>This repository hosts the results of processing example imaging mass cytometry (IMC) data hosted at <a href="http://10.5281/zenodo.5949116">10.5281/zenodo.5949116</a> using the IMC Segmentation Pipeline available at <a href="https://github.com/BodenmillerGroup/ImcSegmentationPipeline">https://github.com/BodenmillerGroup/ImcSegmentationPipeline</a> (DOI: <a href="http://10.5281/zenodo.6402666">10.5281/zenodo.6402666</a>) v3.6. Please refer to <a href="https://github.com/BodenmillerGroup/steinbock">https://github.com/BodenmillerGroup/steinbock</a> as alternative processing framework and <a href="http://10.5281/zenodo.6043600">10.5281/zenodo.6043600</a> for the data generated by <i>steinbock</i>.</p><p>The following files are part of the <strong>analysis.zip</strong> folder when running the IMC Segmentation Pipeline:</p><ul><li><strong>cpinp</strong>: contains input files for the segmentation pipeline</li><li><strong>cpout</strong>: contains all final output files of the pipeline: <i>cell.csv</i> containing the single-cell features; <i>Experiment.csv</i> containing CellProfiler metadata; <i>Image.csv</i> containing acquisition metadata; <i>Object relationships.csv</i> containing an edge list indicating interacting cells; <i>panel.csv</i> containing channel information; <i>var_cell.csv</i> containing cell feature information; <i>var_Image.csv</i> containing acquisition feature information; <i>images </i>containing the hot pixel filtered multi-channel images and the channel order; <i>masks</i> containing the segmentation masks; <i>probabilities </i>containing the pixel probabilities.</li><li><strong>histocat</strong>: contains single channel .tiff files per acquisition for upload to histoCAT (<a href="https://bodenmillergroup.github.io/histoCAT/">https://bodenmillergroup.github.io/histoCAT/</a>)</li><li><strong>crops</strong>: contains upscaled image crops in .h5 format for ilastik (<a href="https://www.ilastik.org/">https://www.ilastik.org/</a>) training</li><li><strong>ometiff</strong>: contains .ome.tiff files per acquisition, .png files per panorama and additional metadata files per slide</li><li><strong>ilastik</strong>: multi channel images for ilastik pixel classification (<i>_ilastik.full</i>) and their channel order (<i>_ilastik.csv</i>); upscaled multi channel images for ilastik pixel prediction (<i>_ilastik_s2.h5</i>); upscaled 3 channel images containing ilastik pixel probabilities (<i>_ilastik_s2_Probabilities.tiff</i>).</li></ul><p>The remaining files are part of the root directory:</p><ul><li><strong>docs.zip: </strong>Documentation of the pipeline in markdown format</li><li>I<strong>MCWorkflow.ilp: </strong>Ilastik pixel classifier pre-trained on the example data</li><li><strong>resources.zip: </strong>The CellProfiler pipelines and CellProfiler plugins used for the analysis</li><li><strong>scripts.zip: </strong>Python notebooks used for pre-processing and downloading the example data</li><li><strong>src.zip: </strong>Scripts for the imcsegpipe python package</li></ul>
Simulated BCRseq data for the nf-core/airrflow pipeline benchmark
<p>Simulated BCRseq data for the nf-core/airrflow pipeline benchmark.</p> <p>The original repertoires simulated with ImmuneSim (sim_repertoire_orig_repA/repB/repC.tsv) and the clonally expanded repertoires (sim_repertoire_clonally_expandedrepA/repB/repC.tsv), as well as the fasta file formats from the clonally expanded repertoires with and without UMIs are also shared. The fasta files were used to simulate the sequencing reads with various degrees of sequencing errors.</p>
Supplemental Data from the article "The SmARTR pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data"
<h1><strong>Please, refer to <a href="https://github.com/MeVisLab/SmARTR-Networks">this GitHub repository</a> for additional info, updates, issue reports, and discussion<br></strong></h1> <p><strong>A collection of configuration files (SmARTR networks) published in "<a href="https://doi.org/10.1016/j.isci.2024.111475">The SmARTR Pipeline: a modular workflow for the cinematic rendering of 3D scientific imaging data</a>", enabling the creation of cinematic (photorealistic) renderings of 3D data in the FREE software <a href="https://www.mevislab.de/download">MeVisLab</a><br></strong></p> <ul> <li>Each folder in the archive contains one or more SmARTR network files, the scan and mask files required for the practical examples detailed in the <a href="https://www.cell.com/cms/10.1016/j.isci.2024.111475/attachment/8d79036b-acb6-4cda-a5ff-f56317691ebc/mmc1.pdf">Supplemental Data</a> of the article, and an additional folder with LUT presets.</li> </ul>
Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline
<p>This upload contains the HZV029 Plasma and HZV029 Two-Phase dataset for reviewers of the "Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline" submission. </p> <p>Both datasets will be uploaded to metabolomics workbench and the upload completed before final publication of the manuscript. For the he HZV029 Plasma datasets only the final run is included for any sample (i.e., failed injections or other samples with data quality issues that were reran during acquisition were omitted).</p> <p>Also included in the upload is the source code for the MetDataModel and the pcpfm at the time of manuscript re-submission and the pcpfm itself. If you find this upload in the future, please check out the github repos for more updated versions:</p> <p>https://github.com/shuzhao-li-lab/PythonCentricPipelineForMetabolomics</p> <p>https://github.com/shuzhao-li-lab/metDataModel</p> <p>The github repo does not store the input the data for space reasons, they only have the notebooks. However, the .zip here has both the notebooks by themselves in the notebook subdirectory and a separate directory with the notebooks and the data used to generate all the figures and results in the manuscript.</p> <p><strong>Some information that is needed to rerun this analysis:</strong></p> <p>Sequence files are critical to the functioning of the pipeline. The sequence files for all analyses are provided under sequence_files.zip. These can be used to recapitulate the analysis by eitehr changing the filepath to each acquisition to where you put it on your sytem or by placing the sequence file in the same directory as the mzml or raw. In the latter case, the pipeline will search for filenames matching the sample names. The sequence files also store some sample metadata such as the type of sample a given acquisition is (unknown, pooled, qc, etc...)</p> <p>.raw to .mzML conversion works well on MacOS but may not work well on other systems. You will need to use the ability to specify your own conversion command or convert files outside of the pipeline. </p> <p>To replicate the results, you do need to have the annotation sources downloaded which can be done using the pipeline. MS2 annotation requires the files in the AcquireX directory which is MS2 acquisitions on pooled HZV029 plasma samples.</p> <p>For the comparison between MetaboAnalystR and the pcpfm, subsets of the datasets were used. These subsets and the sequence files are in Subsets_for_performance_testing.zip. The sequences are also in the sequence_files directory as well</p> <p>The notebooks reference data in the analysis folders. Copies of these files are located with the notebooks to ease reproduction of the exact results in the paper; however, to do so, you will need to change paths to this data in the notebook. This lets the notebooks be ran during a rerun without copying intermediates back and forth and it keeps the github repo clean.</p> <p><strong>Version History:</strong></p> <p>This version is after reviewer comments and is for resubmission.</p> <p> </p> <p><strong>Contributions:</strong></p> <p>Joshua M Mitchell implemented the pipeline and was first author on the manuscript. Shuzhao Li is the corresponding author on the manuscript. </p> <p>Maheshwor Thapa performed the experiments to collect the HZV029 data. Yuanye Chi helped with testing and documenting the pipeline. </p> <p>Jiangou (Jeff) Xia and Zhiqiang Pang provided the R portion of the analysis. </p>
Data for: Tang et al., Interpretable classification of Alzheimer's disease pathologies with a convolutional neural network pipeline. bioRxiv 2018.
<p>Datasets containing 63 whole slide images (WSIs) and their segmented 256x256 pixel tiles with approximately 80,000 tile-level amyloid-β pathology expert annotations.</p> <p><strong>Paper</strong>: "Interpretable classification of Alzheimer's disease pathologies with a convolutional neural network pipeline", bioRxiv 454793; DOI: <a href="https://doi.org/10.1101/454793">https://doi.org/10.1101/454793</a>.</p> <p><strong>Details:</strong> A total of 63 WSIs for 63 unique decedent cases spanning Alzheimer’s disease (AD) to non-AD and possessing a variety of CERAD scores. WSIs comprise three datasets as follows:</p> <ol> <li><em>Development (Phases I-II)</em>. 33 WSIs used for convolutional neural network (CNN) model development (29 training, 4 validation).</li> <li><em>Hold-out (Phase III)</em>. 10 WSIs selected by an expert neuropathologist as a held-out test set to assess the generalizability of the CNN model.</li> <li><em>CERAD-like hold-out</em>. 20 blinded WSIs collected solely for use in a CERAD-like scoring comparison study.</li> </ol> <p>Datasets 1 and 2 were color-normalized and segmented to 256x256 pixel image tiles for model training set (61,370 images), validation set (8,630 images), and hold-out test set (10,873 images). Dataset 3 was color-normalized but not segmented.</p> <p>Expert labels of plaques for Dataset 1 and 2 tiles are included in corresponding CSV files.</p> <p><strong>Slide source and preparation:</strong> All samples were retrieved from archives of the University of California, Davis Alzheimer’s Disease Center Brain Bank (<a href="https://www.ucdmc.ucdavis.edu/alzheimers/">https://www.ucdmc.ucdavis.edu/alzheimers/</a>). Archival samples analyzed in this study were 5 μm formalin fixed, paraffin embedded sections of the superior and middle temporal gyrus from human brain. The tissue had been previously stained with an amyloid-β antibody (4G8, recognizing residues 17-24, BioLegend, formerly Covance) that were first pretreated with formic acid to rid samples of endogenous protein. All slides were digitized using an Aperio AT2 up to 40x magnification.</p> <p><strong>Code:</strong> Please visit <a href="https://github.com/keiserlab/plaquebox-paper">https://github.com/keiserlab/plaquebox-paper</a></p> <p> </p>
Supplementary data to accompany "phyloFlash: A pipeline for rapid SSU rRNA-targeted profiling of metagenomes"
<p>Usage examples of the phyloFlash pipeline applied to shotgun metagenomic data sets.</p> <p>The phyloFlash software is available from https://github.com/HRGV/phyloFlash. Examples were generated with phyloFlash v3.3b.</p>
Building a data curation pipeline for complex diseases: the case of Major Depression - Supplementary Material
<p>This entry contains the data generated by the study "Building a data curation pipeline for complex diseases: the case of Major Depression".</p>
Raw data and analysis pipeline for producing figures in F.W. Carter, et. al., 2016
<p>This data release accompanies a manuscript submitted to IEEE Transactions on Applied Superconductivity on Sept. 6, 2016. If accepted, DOI of the publication will be linked to this.</p> <p>This data release (and accompanying JuPyter notebook) were used to generate all of the figures in the manuscript. The software package used to crunch the data is archived here: http://dx.doi.org/10.5281/zenodo.61512</p>
Identifying Easy Instances to Improve Efficiency of ML Pipelines for Algorithm-Selection - Code and Data
<p>This repository contains the code and data for reproducibility of the paper 'Identifying Easy Instances to Improve Efficiency of ML Pipelines for Algorithm-Selection'. </p> <p>The following files are included:</p> <ul> <li>best_algo.csv : labels for the classification and median performance of algorithms;</li> <li>ML_models.ipyb : jupyter notebook with the definition of the neural networks for both classifiers;</li> <li>pickle.zip : pickled models for the hardness classification and the algorithm selector;</li> <li>trajectories.zip : raw data files containing parts of the trajectories of each algorithm;</li> <li>Results.zip : results obtained using the approach on the stream of instances.</li> </ul>
Code to generate figures 3 and 4 of: "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."
<p>Code to generate figures 3 and 4 of the manuscript titled "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."</p> <p> </p>
Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).
<p>*******************************************************************<br> MetaDrugs workflow<br> *******************************************************************</p> <p>Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).</p> <p>For questions and requests, please contact:<br> Sofia K. Forslund (sofia.forslund@mdc-berlin.de)<br> and Till Birkner (till.birkner@mdc-berlin.de)</p> <p>*******************************************************************<br> Contents:<br> -------------------------------------------------------------------<br> Data files:<br> metadata.tar.gz - archived cohort metadata files*<br> input_features.tar.gz - archived preprocessed serum and urine metabolome and gut microbiome features<br> output_complete.tar.gz - archived example analysis output files for each of the input feature file<br> output_rerun.tar.gz - archived empty directory for generating test output files as described in this document<br> <br> *Please note: Due to conflicts with Danish Data Protection laws, metadata from the Danish subset of the cohort were removed in this repository. Please reach out for a potential case-by-case access request for access to the complete set of metadata.<br> -------------------------------------------------------------------<br> Text files:<br> archived in feature_names.tar.gz:<br> atcs_names - full names for atcs drug compounds<br> contrast_names - full names for disease comparison groups<br> file_names - brief description of the files in input_features folder<br> gmm_names - full names of GMM modules<br> kegg_names - full names of KEGG modules<br> ko_names - full names of KO modules<br> metadata_names - full names of metadata features<br> mOTU_names - species names for metagenomics data<br> taxon_names - taxon names for metagenomics data<br> -------------------------------------------------------------------<br> Scripts:<br> -------------------------------------------------------------------<br> runFrame.r - main wrapper script envoking the analysis pipeline<br> -------------------------------------------------------------------<br> runFrame_rel_comb.r - script calculating drug combination effects<br> runFrame_rel.r - script calculating dosage effects<br> testCombPresenceSeparate.r - testing of significant drug combination effects beyond single drug effects<br> testDosagePresenceSeparate.pl - testing of significant drug dosage effects beyond single drug effects<br> testDosagePresenceSeparateNegative.pl - testing of unique drug dosage effects beyond single drug effects<br> -------------------------------------------------------------------<br> prettifyResults_uncollapsed.pl - wrapper scripts to create and format a single analysis output file<br> makeTables.r - wrapper script to make excel tables with analysis results<br> -------------------------------------------------------------------<br> Example output file:<br> -------------------------------------------------------------------<br> output_all_formatted_noc_uncollapsed_complete.tsv - contains all disease-drug-host-microbiome feature analysis results in one place.<br> *******************************************************************</p>
Saccharomyces cerevisiae Bud-Annotation pipeline: Napari Example Data
<p>Example dataset for bud-annotation plugin Napari.</p> <p>Here we provide two sets of 3 images of single molecule mRNA FISH on Saccheromyces cerevisiae (BY4741) strains containing an mRNA bud localization reporter at the DOA1 locus. <br> The image set labeled EXPERIMENT contains images of the strain in which a localization element was present in the reporter mRNA. The reporter can be seen to localize to the bud. <br> The image set labeled CONTROL contains images of the control strain without localization element present in the reporter mRNA. The reporter can be seen to be randomly distributed throughout the cell. </p> <p>Images were acquired as 41 z-stacks per fluorescence channel CY5, CY3.5, CY3 and DAPI. For each fluorescence image a corresponding DIC image was acquired as well. </p> <p>The following smFISH probes were used: <br> - CY5: probes targeting the endogenous ASH1 and CLB2 mRNA to be used as bud marker (Quasar 670)<br> - CY3.5: probes targeting the DOA1 mRNA reporter (CAL Fluor Red 610)<br> - CY3: probes targeting the MS2v6 sequence inserted at the 3' end of the DOA1 mRNA reporter (Quasar 570) </p> <p>FISH-QUANT spot analysis results for the CY3.5 channel from these images are included. This data can be used to extract the spot information and assign spots to either mother cell or bud. Examples of nuclear, bud and cell masks for all images are provided as well.</p>
The structural basis for the self-inhibition of DNA binding by apo-σ70 - smFRET raw data and analyses pipeline
<p>This dataset includes all raw data of nsALEX smFRET measurements of doubly-labeled sigma70 reported in Joron et al. ("The structural basis for the self-inhibition of DNA binding by apo-σ70"), as well as Jupyter Notebooks documenting the analysis pipeline that takes us from the raw data to dual channel burst search and filtered bursts, and to the analyses of within-burst dynamics in the system</p>
Data for publication: A pipeline for in-depth analysis of DNA virus populations by profiling the low abundant virus variants and partial genomic components
<p>Raw and processed sequence data from Oxford Nanopore and BGI short read sequencing platforms used in the publication: "A pipeline for in-depth analysis of DNA virus populations by profiling the low abundant virus variants and partial genomic components".</p>
Example data for "Sunbeam: an extensible pipeline for analyzing metagenomic sequencing experiments" [Version 2]
<p>This repository contains the example datasets analyzed in the Sunbeam paper, version 2. Please see the current <a href="http://sunbeam.readthedocs.io/en/latest/quickstart.html">Sunbeam Quickstart Guide</a> for up-to-date instructions on installing and running Sunbeam.</p>
NGS competence network - pipeline benchmark data
<p>VCF files generated with the megSAP pipeline for a pipeline benchmark performed by the NGS competence network.</p>
RVFV data aligning only to M-Fragment to test the PARANOiD pipeline
<p>PARANOiD is a versatile software for fully automated analysis of iCLIP and iCLIP2 data. It contains all steps necessary for preprocessing, the determination of cross-link locations and several additional steps, which can be used to detect specific characteristics, e.g. definite distances between cross-link events or identify binding motifs. The cross-link sites are presented as WIG files that can be easily visualized e.g. using IGV, for which a config file is offered. Additionally, results are offered as statistical plots for a quick overview and as standardized bioinformatics file formats or TSV files, which can be used for further analysis steps.</p> <p>The data provided are used as a test case for PARANOiD.</p> <p>The data was extracted from RVFV MP-12 virions (virion-reads-M-fragment-only.fastq) and BHK cells infected with RVFV (BHK-reads-M-fragment-only.fastq) applying the iCLIP2 method for RVFV N iCLIP. Three independent biological replicates were performed for each sample. Sequencing was performed using the MiSeq Sequencer (Illumina) with MiSeq Reagent Kit v2 Micro (Illumina) for N-iCLIP from virus particles and MiSeq Reagent Kit v3 (Illumina) for N-iCLIP from infected BHK cells.</p> <p>The original reads have been aligned to the RVFV MP-12 reference genome and only reads aligning to the M-fragments were extracted. The whole dataset will be publish at a later date</p> <p>File description:</p> <p>virion-reads-M-fragment-only.fastq - Reads obtained from RVFV virions</p> <p>BHK-reads-M-fragment-only.fastq - Reads obtained from BHK cells infected with RVFV</p> <p>reference_RVFV.fasta - RVFV MP-12 reference genome</p> <p>barcodes-RVFV.tsv - Barcodes for virion-reads-M-fragment-only.fastq</p> <p>barcodes-RVFV-merge-all.tsv - Barcodes for merging all samples of virion-reads-M-fragment-only.fastq</p> <p>barcodes-BHK.tsv - Barcodes for BHK-reads-M-fragment-only.fastq</p> <p>barcodes-BHK-merge-all.tsv - Barcodes for merging all relevant samples of BHK-reads-M-fragment-only.fastq</p>
Data Supplement: GIRFReco.jl: An Open-Source Pipeline for Spiral Magnetic Resonance Image (MRI) Reconstruction in Julia
<p><strong>Dataset for GIRFReco.jl Paper</strong><br> <br> Please download this and extract to an appropriate location prior to running the demonstration code in GIRFReco.jl. The extracted folder will serve as the root directory in the demo code.</p>
Input data for Atollgen pipeline
<p>Input data for Atollgen pipeline</p> <p>Contains:</p> <ul> <li>Frozen island raw sources (atollgen database inputs)</li> <li>hmm database (integrase and mobility signatures coming from ConjScan and Pfam-A)</li> <li>categorisation metadata for each signature contained in the integrase and mobility database</li> <li>Frozen genomes sequences from the NCBI</li> <li>Frozen list of actinobacteria taxonomy ids</li> <li>Frozen defense-finder database</li> <li>Frozen cards</li> </ul> <p>Frozen data ensure reproducibility for the pipeline, but up-to-date data should give similar (yet not identical) results.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.