Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
Processed FDG-PET data from: A computational model of neurodegeneration in Alzheimer's disease
<p>Disruption of mental functions in Alzheimer's disease (AD) and related disorders is accompanied by selective degeneration of brain regions. These regions comprise large-scale ensembles of cells organized into systems for mental functioning, however the relationship between clinical symptoms of dementia, patterns of neurodegeneration, and functional systems is not clear. We developed a model of the association between dementia symptoms and degenerative brain anatomy using F18-fluorodeoxyglucose (FDG) PET and dimensionality reduction techniques patients with AD. This data and code package contains preprocessed FDG-PET images from 423 subjects across the Alzheimer's disease spectrum and the MATLAB code to produce eigenbrains from this data.</p>
Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"
<p>Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"</p>
Pathogen reduction data when applying high pressure processing to milk/colostrum and ready-to-eat foods
<p><strong>Introduction</strong></p> <p>This dataset was built by members of the EFSA Working Group on high pressure processing (HPP) of food during the preparatory work on the BIOHAZ Scientific Opinion on the efficacy and safety of high pressure processing of food (EFSA-Q-2020-00380) (<a href="https://doi.org/10.2903/j.efsa.2022.7128">https://doi.org/10.2903/j.efsa.2022.7128</a>).</p> <p>It was used to answer to the following Terms of Reference (ToR) of the mandate:</p> <p>ToR2. To assess the efficacy of HPP when applied to raw milk and raw colostrum from ruminants, and in particular:</p> <p>a. To recommend minimum requirements as regards time and pressure of the HPP, and other factors if relevant, for the control of <em>Mycobacterium</em> spp., <em>Brucella</em> spp., <em>Listeria monocytogenes</em>, <em>Salmonella</em> spp. and Shiga toxin-producing <em>Escherichia coli</em> (STEC), to achieve an equivalent efficacy to that of thermal pasteurisation.</p> <p> </p> <p>ToR3. To assess the efficacy of HPP when applied to foods known to cause human listeriosis and in particular:</p> <p>a. To recommend minimum requirements as regards time and pressure of the HPP, and other factors if relevant, to reduce significantly <em>L. monocytogenes</em> levels (e.g. by a certain log reduction), and assuming that the parameters influencing the growth of <em>L. monocytogenes</em> remain unchanged (e.g. shelf-life and storage conditions);</p> <p>b. To assess the efficacy on other relevant pathogens when applying the minimum requirements identified in a.</p> <p>A literature search was conducted to retrieve studies reporting on pathogen-specific parameters when treating milk/colostrum and ready-to-eat foods with HPP. A modelling approach was used to fit the derived log<sub>10</sub> reductions (logR) at specific pressure-holding time combinations. More information can be found in the Scientific Opinion.</p> <p><strong>Description</strong></p> <p>Two MS-Excel files contain the extracted data.</p> <p>One file consists of the data used for answering ToR2 and contains data on HPP inactivation of <em>Mycobacterium bovis</em>, <em>L. monocytogenes</em>, <em>Salmonella</em> spp., STEC, <em>Campylobacter</em> spp., and <em>Staphylococcus aureus</em> in milk and colostrum from ruminants.</p> <p>The other file was used for answering ToR3 and contains data on HPP inactivation of <em>L. monocytogenes</em>, <em>Salmonella</em> spp. and <em>E. coli</em> in three types of RTE food categories: category “cooked meat products”, category “smoked and gravad fish” and category “soft or semi-soft and fresh cheese”.</p> <p>Both files contain metadata related to product characteristics (e.g. the food matrix and composition), contamination characteristics (e.g. pathogen/strain used for inoculation, inoculation level, medium used for enumeration), HPP treatment characteristics (e.g. target pressure, time), and outcome (e.g. log<sub>10</sub> reduction).</p>
Data sets used in "Neural network processing of holographic images"
<p>Included are the training, validation, and testing data sets for synthetic holograms (netCDF), the HOLODEC data set containing the RF07 examples (netCDF), and the two splits of manually labeled HOLODEC image tiles (numpy arrays). The source code for using the data sets can be found at https://github.com/NCAR/holodec-ml </p>
Data for "Paranormal experiences, sensory-processing sensitivity, and the priming of pareidolia".
<p>Data from a study investigating paranormal priming effects, the detection of voices in ambiguous stimuli, and the perceptual advantage of sensory-processing sensitivity. </p>
Processed MinION genome sequencing data for strains ILHA G3AG5 and ILHA G3AA5
<p>Strain G3AA5 datasets include:</p> <p>Galaxy2750, Galaxy2754, Galaxy3522, Galaxy3523, Galaxy3524, Galaxy3541</p> <p> </p> <p>Strain G3AG4 datasets include:</p> <p>Galaxy2404, Galaxy2408, Galaxy3517, Galaxy3518, Galaxy3519, Galaxy3539</p>
Processed and additional data for our publication titled "Endoplasmic reticulum stress activates human IRE1α through reversible assembly of inactive dimers into small oligomers"
<p>This is an updated version of our original data archive (which can be found under the doi 10.5281/zenodo.5513025) that reflects changes we've made to the manuscript over the course of the review process and incorporates the new data we've collected since the time of the initial bioRxiv submission.</p> <p>This data archive contains all raw data EXCEPT for single-particle microscopy movies (which are deposited separately due to their size) for our paper titled "Endoplasmic reticulum stress activates human IRE1a through reversible assembly of inactive dimers into small oligomers". These raw data are stored in "non_SPT_data_final_v2.zip". Additionally, full plasmid sequences for all plasmids used in this paper are stored in GenBank format in the file "plasmid_sequences_v2.zip". Finally, this repository contains processed single-particle movies in the form of dual-color tracks from the TrackMate ImageJ plugin in XML format (file: "SPT_processed_data_and_settings_final_v2.zip").</p> <p>The processed XML tracks are organized in the same way as the raw data files in the separate repository. They are sorted into subfolders by date of acquisition first, followed by experimental conditions. To recreate the figures from the paper, follow instructions in the README.md file included with the source code repository and use the JSON settings files saved here under "analysis_settings".</p>
Processed ADCP and CTD Data for the Processes Driving Exchange at Cape Hatteras (PEACH) Program
<p>These are the hourly acoustic doppler current profiler (ADCP) and conductivity-temperature-depth (CTD files from the ten moored bottom frames and buoys deployed as part of the PEACH program. Level 2 data have been quality-controlled and gridded to an hourly time-base. A technical report describing the details of processing and moored aspects of the experiment can be found in: Haines, S., <em>et al.</em> (2022) Mooring Data Report for the Processes driving Exchange at Cape Hatteras (PEACH) Project. (doi:10.5281/zenodo.6380851). </p>
Processed Buoy and Bulk Heat Flux Data for the Processes Driving Exchange at Cape Hatteras (PEACH) Program
<p>These are the hourly buoy and bulk heat flux from National Data Buoy Center (NDBC) buoys and the two buoys deployed as part of the PEACH program. Level 2 data have been quality-controlled and gridded to an hourly time-base. A technical report describing the details of processing and the moored aspects of the experiment can be found in: Haines, S., et al. (2022) Mooring Data Report for the Processes driving Exchange at Cape Hatteras (PEACH) Project. (doi:10.5281/zenodo.6380851).</p>
Processed data for MethylBoostER: an XGBoost model to classify kidney cancer subtypes
<p>This is a repository containing processed data for MethylBoostER, an XGBoost model that classifies kidney cancer subtypes. The open-source code can be found here: https://github.com/ss-lab-cancerunit/MethylBoostER.</p>
Genome graphs detect human polymorphisms in active epigenomic states during influenza infection: code and processed data
<p>Manuscript, figure, and analysis code and processed data for the Groza et al (2022) preprint.</p>
Data for 'Using 40 years of spot measurements to assess stream temperature response and recovery for different harvesting systems in northern hardwood forests' by Jason Leach, Danielle Hudson and R. Dan Moore. Submitted to Hydrological Processes.
<p>This dataset contains spot stream temperature measurements taken at 5 headwater streams draining forested hillslopes (C31, C32, C33, C34, C35) in the Turkey Lakes Watershed, approximately 65 km northwest of Sault Ste. Marie, Ontario, Canada.</p>
Datasets and Code for "Hypothesis Tests with Functional Data for Surface Quality Change Detection in Surface Finishing Processes"
<p>This is the set of data and computer code used for reproducing the results in Jin, Tuo, Tiwari, Bukkapatnam, Aracne-Ruddle, Lighty, Hamza, and Ding, 2022, “Hypothesis tests with functional data for surface quality change detection in surface finishing processes,” <em>IISE Transactions</em>, in press.</p>
Data Availability - DRAFT - 'Process-based similarity' revealed by discharge-dependent relative submergence dynamics of thousands of large bed elements
<p>Datasets and R code related to confidential manuscript submission entitled "‘Process-based similarity’ revealed by discharge-dependent relative submergence dynamics of thousands of large bed elements". Restrictions apply to the availability of the 2014 DTM and 2D model results, which were used under contractual agreement from the project sponsor. These are available from the senior author with the permission of Yuba Water Agency. Restrictions apply to the availability of the 2014 DTM and 2D model results, which were used under contractual agreement from the project sponsor. These are available from the senior author with the permission of Yuba Water Agency.</p>
Processed ground observation and WRF-CAMQ data for Greater Bay Area, 2015-2021
<p>The pre-processed training and testing dataset for the paper <em>Development of an LSTM-Broadcasting deep-learning framework for regional air pollution forecast improvement. </em>The format of data is compatible with <a href="https://github.com/jvhs0706/regional-forecast-new">the official implementation</a>.</p>
Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"
<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores. <br> <br> For the details please see the original publication or the bioarxiv preprint: https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment. <br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>
Pre-processed data for "Does host plant drive variation in microbial gut communities in a recently shifted pest?"
<p>Pre-processed fastqs files generated by Illumina sequencing associated with the publication by Javal et al. entitled "Does host plant drive variation in microbial gut communities in a recently shifted pest?".</p>
Data and code for "Non-local parameterization of atmospheric subgrid processes with neural networks" (Wang et al. 2022 submit to JAMES)
<p>Data and code for "Non-local parameterization of atmospheric subgrid processes with neural networks" (Wang et al. 2022 submit to JAMES). Detailed description of the files in README.txt.</p>
Dataset and R script for Collection and Processing of Behavioural Data of the Olive Fruit Fly, Bactrocera oleae, when Exposed to Olive Twigs Treated with Different Commercial Products
<p>We provide raw data and R script for analysis of data published in:</p> <p>1) Daher, E.; Cinosi, N.; Chierici, E.; Rondoni, G.; Famiani, F.; Conti, E. Field and Laboratory Efficacy of Low-Impact Commercial<br> Products in Preventing Olive Fruit Fly, Bactrocera oleae, Infestation. Insects 2022, 13, 213. https://doi.org/10.3390/insects13020213</p> <p>2) Daher, E.; Chierici, E.; Cinosi, N.; Rondoni, G.; Famiani, F.; Conti, E. Collection and Processing of Behavioural Data of the Olive Fruit Fly, <em>Bactrocera oleae</em>, when Exposed to Olive Twigs Treated with Different Commercial Products. <em>Data </em><strong>2022</strong>, <em>7</em>,</p>
Dataset and data processing tools of EPL 139 (2022) 55002 — Harmonic calibration of quadrature phase interferometry
<p>Dataset for <em>EPL</em> <strong>139</strong> (2022) 55002 — Harmonic calibration of quadrature phase interferometry</p> <p><strong>Scripts: </strong>the main Matlab script to analyse all data is <strong>analyseall.m</strong>, and uses all other scripts and functions (<strong>*.m</strong>). See comments in this main script for details. The script <strong>DefineDatasets.m,</strong> and its output <strong>Dataset.mat</strong>, were run before to define the main parameters of the 20 datasets, and define the frequency range where to estimate the background noise to be subtracted from the spectra of the driven cantilever when computing the total harmonic distortion.</p> <p><strong>Data:</strong> the 100 data files (.mat) have the following name scheme: [driving-frequency]-[Amplitude]-(low frequency external phase)-001 to 005.mat. For example 1kHz-10nm-w1Hz150nm-001.mat correspond to a driving of the cantilever at 1kHz, with an oscillation amplitude close to 10nm, with a low frequency driving of the external optical phase at 1Hz and equivalent 150nm amplitude. 001 stand for the first of the five files corresponding to this measurement. Note that the low frequency driving is optional. The main driving frequency is either directly given in kHz, or corresponds to one of the 2 first resonance modes of the cantilever (mode1 at 13.4kHz, and mode2approx at 84.6 kHz). Each of these data files contains 5 variables: Cx and Cy are the 2 outputs of the quadrature phase interferometer (1s of acquisition, 2e6 samples), fs is the sampling frequency (2 MHz), Ix and Iy are the mean total intensities on the photodiodes corresponding to Cx and Cy (units: A, this information can be useful to estimate the expected shot noise floor on the measurement).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.