Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
77
datasets available to search
ShareScore release 0.9.0
Dataset results
77 results for “predictive processing”
pKaDatabase for Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge
<p>A curated a database of small molecules with experimentally measured pKa values. </p> <p>This pickle file can be loaded into memory using Pandas. In the code block below we will print out the columns of the DataFrame:</p> <pre><code class="language-python">import pandas as pd df = pd.load("pKaDatabase.pkl") print(df.keys()). # print the columns</code></pre> <blockquote> <p>['deprotonated microstate ID', 'protonated microstate ID', 'deprotonated microstate smiles', 'protonated microstate smiles', 'AM1BCC partial charge (prot. atom)', 'AM1BCC partial charge (deprot. atom)', 'AM1BCC partial charge (prot. atoms 1 bond away)', 'AM1BCC partial charge (deprot. atoms 1 bond away)', 'AM1BCC partial charge (prot. atoms 2 bond away)', 'AM1BCC partial charge (deprot. atoms 2 bond away)', 'Gasteiger partial charge (prot. atom)', 'Gasteiger partial charge (deprot. atom)', 'Gasteiger partial charge (prot. atoms 1 bond away)', 'Gasteiger partial charge (deprot. atoms 1 bond away)', 'Gasteiger partial charge (prot. atoms 2 bond away)', 'Gasteiger partial charge (deprot. atoms 2 bond away)', 'Extented Hückel partial charge (prot. atom)', 'Extented Hückel partial charge (deprot. atom)', 'Extented Hückel partial charge (prot. atoms 1 bond away)', 'Extented Hückel partial charge (deprot. atoms 1 bond away)', 'Extented Hückel partial charge (prot. atoms 2 bond away)', 'Extented Hückel partial charge (deprot. atoms 2 bond away)', '∆G_solv (kJ/mol) (prot-deprot)', 'SASA (Shrake)', 'SASA (Lee)', 'Bond Order', 'Change in Enthalpy (kJ/mol) (prot-deprot)', 'pKa','href', 'num ionizable groups', 'Weight', 'pKa source']</p> </blockquote> <p> </p> <p>For more information regarding feature calculations, please read the following paper:</p> <blockquote> <p>Raddi, Robert, and Vincent Voelz. "Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge." (2021). <a href="https://doi.org/10.26434/chemrxiv.14650302.v1">10.26434/chemrxiv.14650302.v1</a></p> </blockquote>
Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements
The Forest and Biodiversity (FAB2) experiment uses native tree species in varying levels of species richness, phylogenetic diversity, and functional diversity planted in 100 m2 and 400 m2 plots at 1 m spacing, appropriate for testing long-term ecosystem consequences. FAB2 was designed and established in conjunction with a prior experiment (FAB1) in which the same set of twelve species was planted in 16 m2 plots at 0.5 m spacing. This data package examines the connections between aboveground and belowground processes in FAB2. This data package includes information on tree diversity and community composition, forest structure, forest understories, soil microbes, net nitrogen mineralization, and canopy nitrogen. A wide variety of data types are included, such as data from hyperspectral and LiDAR remote sensing, percent cover analysis, soil microbial analyses, and soil assays including C:N, pH, and net nitrogen mineralization. This data package is included in the submission of the manuscript entitled “Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements.”
Dataset: Brain negativity as an indicator of predictive error processing: The contribution of visual action effect monitoring
<p>There are two files for each subject:</p> <p>1. sub##_error.dat -> Contains EEG Segments, that were recorded while the subject executed a clear target miss (minimal distance between the center of the ball and target > 12 cm) in the task (segment and electrode information can be found below).</p> <p>2. sub##_hit.dat -> Contains EEG Segments, that were recorded while the subject executed a clear target hit (minimal distance between the center of the ball and the target < 7 cm) in the task (segment and electrode information can be found below).</p> <p><br> The data in the *.dat-files are stored in a two dimensional matrix: n*1400 datapoints x 15 electrodes</p> <p>n represents the number of segments. 1400 datapoints per segment translate to a segment length of 2800 ms (from 600 ms before to 2200 ms after ball release). The ball´s release is located at the 301st datapoint and the feedback was presented at datapoint 726 (850 ms after ball release) in every segment.</p> <p>datapoints: The first dimension (rows) includes the measured neural activations in microvolts. The data is stored vectorized,<br> i.e. hit/error #1 -> row 1 to 1400, hit/error #2 -> row 1401 to 2800, ..., hit/error #n -> (n-1) * 1400 + 1 to n * 1400</p> <p>electrodes: The second dimension (columns) consists of the 15 different electrodes that were used during data recording in this exact order: [F3 Fz F4 C4 Cz C3 P3 Pz P4 VEOGu VEOGo HEOGre HEOGli FCz Mastre]</p>
Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models
<p>This data set contains the data, JMP scripts, and figures of the article titled "Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models" to be published in the journal Animal - Open Space.</p>
Ion Implantation Sensor and Process Target Data for Predicting Ion Beam Tuning in Semiconductor Manufacturing
<h2><strong>Dataset Description:</strong></h2> <p>This dataset is designed to predict ion beam tuning setup processes in semiconductor manufacturing, in terms of tuning success or failure, and tuning duration. It is split into <strong><code>X</code></strong> and <code><strong>y</strong></code> to allow for supervised learning approaches.</p> <ul> <li><code><strong>X</strong></code> represents the current equipment condition and the process targets of the currently processed and the upcoming lot, as defined within recipes.</li> <li><code><strong>y</strong></code> represents the ion beam tuning setup report, which informs about the tuning success ratio and tuning duration. These setups are necessary, when switching between recipes to prepare the equipment for processing the next lot. <strong><code>y</code></strong> contains three labels, enabling classification of (1) tuning success or fail, and (2) prolonged tuning, as well as (3) estimation of tuning duration as a regression task.</li> </ul> <p>About <strong><code>X</code></strong>:</p> <p>Each lot is processed with a specific recipe to achieve the process target. The tuning takes place before the first wafer of the to-be-tuned recipe is processed. Each row in <strong><code>X</code></strong> includes logistical information such as the equipment used for processing and parsed recipe / process target information for the current and upcoming lot. The majority of data consists out of aggregated metrics of equipment-internally tracked sensor traces, recording physical parameters such as gas flows, temperatures, voltages and currents. When analyzed in conjunction with the processed recipe, these sensors provide insights into the current equipment condition. </p> <p>About <code><strong>y</strong></code>:</p> <p>The <code>setup_result</code> column indicates the success or failure of tuning - with <code>setup_result=0</code> indicating tuning success, while <code>setup_result=1</code> signals tuning failure. If the first tuning attempt fails, there may be follow-up attempts, but these are not included in this dataset. The <code>duration</code> column represents the tuning duration in seconds, as used for regression analysis. The <code>duration_interval</code> column is a binary label for prolonged tunings, i.e. <code>duration_interval=1</code> for instances, which take more than 6 minutes to tune.</p> <p>For reproducibility of the corresponding paper's results:</p> <ol> <li>The dataset contains the same carefully curated subset of features.</li> <li>The train_test_split() has already been performed, thus we provide <code>x_train</code> and <code>x_valid</code> separately.</li> <li>To reduce the effect of outliers in the data, the sensor data has already been scaled, as derived from <code>x_train</code>.</li> </ol> <p>In summary, these datasets (<code><strong>X</strong></code>, <code><strong>y</strong></code>) provide comprehensive information for predicting ion beam tuning in semiconductor manufacturing, making it a valuable resource for researchers and practitioners in the field.</p> <h2><strong>Python Code for Reproducibility:</strong></h2> <p>Furthermore, we share a jupyter notebook <code>ionbeamtuning.ipynb</code> with Python code to train the best performing model on the provided data, as described in the paper. To execute the code, you may need to install any missing packages specified in the <code>requirements.txt</code>, as indicated within the notebook.</p>
Pathobionts in the tumour microbiota predict survival following resection for colorectal cancer - pre-processed data
<p>A multicentre, prospective observational study was conducted of colorectal cancer (CRC) patients undergoing primary surgical resection in the United Kingdom and Czech Republic. Analysis was performed using metataxonomics (microbiome) and ultra-performance liquid chromatography mass spectrometry (UPLC-MS, metabolomics). Both datasets were pre-processed as described in the methods section of the main article. The data here were used as the input to the data analysis workflows available from <a href="https://github.com/jmp111/CRC">Github</a>.</p>
Potential distribution of invasive boxwood blight pathogen (Calonectria pseudonaviculata) as predicted by process-based and correlative models
<p>R project, R scripts, and data files for reproducing most of the analyses presented in a climatic suitability study for boxwood blight. The README. md file describes how to run the scripts and provides details on data inputs.</p> <p><strong>Abstract: </strong>Boxwood blight caused by <em>Cps</em> is an emerging disease that has had devastating impacts on <em>Buxus</em> spp. in the horticultural sector, landscapes, and native ecosystems. In this study, we produced a process-based climatic suitability model in the CLIMEX program and combined outputs of four different correlative modeling algorithms to generate an ensemble correlative model. All models were fit and validated using a presence record dataset comprised of <em>Cps</em> detections across its entire known invaded range. Evaluations of model performance provided validation of good model fit for all models. A consensus map of CLIMEX and ensemble correlative model predictions indicated that not-yet-invaded areas in eastern and southern Europe and in the southeastern, midwestern, and Pacific coast regions of North America are climatically suitable for <em>Cps</em> establishment. Most regions of the world where<em> Buxus</em> and its congeners are native are also at risk of establishment. These findings provide the first insights into <em>Cps</em> global invasion threat, suggesting that this invasive pathogen has the potential to significantly expand its range.</p>
Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models
<p>This data set contains the data, JMP scripts, and figures of the article titled "Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models" to be published in the journal Animal - Open Space.</p>
Interception of virtual throws reveals predictive skills based on the visual processing of throwing kinematics - Dataset
<p>Dataset consists of a list of Matlab structures, one for each of the 21 participants. For each participant all the recorded trials are reported ("trials" field). For each trial, the dataset reports information about the associated experimental condition and the kinematics of the ball and the racket trajectories. Information about the experimental condition are specified in the "info" field, which provides the experimental phase (Training and Experimental), the visibility (AllVisible, ThrowerOnly, BallOnly), the target (1, 2, 3, 4), and the thrower ID (1, 2, 3, 4). The kinematics data, starting from the time of ball release, are given in the field "trajectories", which provides the time vector (in seconds), and the corresponding 3D positions of the racket and the ball (in meters) in the reference frame shown in Figure 1 of the paper.</p>
Figure 5. Sensory score and period of storage for processed cheese-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese
<p>R2 was found to be 96.5 percent of the total variation as explained by sensory scores. Period<br> of storage (days) for which the processed cheese has been in the shelf can be determined based on<br> sensory score (Fig. 5).</p>
Figure 4. Comparison of ASS and PSS for multilayer model R-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese
<p>TDNN models with single and multi layers were developed taking soluble nitrogen, pH,<br> standard plate count, yeast & mould count, spore count as input parameters, and sensory score as<br> output parameter for predicting the shelf life of processed cheese stored at 30o C. Mean Square<br> Error, Root Mean Square Error, Coefficient of Determination and Nash - Sutcliffo Coefficient were<br> used in order to compare the prediction ability of the developed TDNN models. Regression<br> equations were developed for predicting the shelf life of processed cheese, which came out as 28.25<br> days. Since, predicted value is close to the experimentally determined shelf life of 30 days, hence<br> from the study it can be concluded that TDNN artificial neural network models are quite efficient in<br> predicting shelf life of processed cheese.</p>
Figure 2. Training pattern of TDNN models-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese
<p>The Neural Network Toolbox under MATLAB software was used for developing the TDNN<br> models. Training pattern of TDNN models is presented in Fig.2.</p>
Figure 1. Inputs and output parameters for TDNN models-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese
<p>The data consisted of 36 samples, which were divided into two subsets, i.e., 30 used for<br> training the network and 6 for testing the TDNN models. Soluble nitrogen, pH, standard plate<br> count, yeast & mould count, and spore count were taken as input parameters, and sensory score as<br> output parameter for developing TDNN single and multilayer models (Fig.1).</p>
Figure 3. Comparison of ASS and PSS single layer model-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese
<p>TDNN models with single and multi layers were developed taking soluble nitrogen, pH,<br> standard plate count, yeast & mould count, spore count as input parameters, and sensory score as<br> output parameter for predicting the shelf life of processed cheese stored at 30o C. Mean Square<br> Error, Root Mean Square Error, Coefficient of Determination and Nash - Sutcliffo Coefficient were<br> used in order to compare the prediction ability of the developed TDNN models. Regression<br> equations were developed for predicting the shelf life of processed cheese, which came out as 28.25<br> days. Since, predicted value is close to the experimentally determined shelf life of 30 days, hence<br> from the study it can be concluded that TDNN artificial neural network models are quite efficient in<br> predicting shelf life of processed cheese.</p>
AbDb processed and pickled for use in deep learning CDR-H3 Structure prediction
<p>This is a pickle file, ready for training by the neural network described in "Improving CDR-H3 modelling in Antibodies" found at the following URL:</p> <p><a href="https://github.com/OniDaito/MRes">https://github.com/OniDaito/MRes</a></p> <p>The data is derived from the AbDb dataset found at:</p> <p><a href="http://www.bioinf.org.uk/abs/abdb/">http://www.bioinf.org.uk/abs/abdb/</a></p>
Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading"
<p>Title of dataset: Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading".</p> <p>Name/institution/contact information: Dr. Michal Bartošák, Czech Technical University in Prague - Faculty of Mechanical Engineering, email: michal.bartosak@fs.cvut.cz.</p> <p>Date of data collection: The data were collected between 2021 and 2024.</p> <p>File name structure: The data consists of two files: "316L_fatigue_and_defects.xls," which contains fatigue lifetime data and defect characteristics, and an associated description file, "read_me.txt."</p> <p>See "https://doi.org/10.1016/j.ijfatigue.2024.108608" for the associated article and a detailed description of the methods.</p>
Dataset and codes of the article "Neural correlates of hierarchical predictive processes in autistic adults"
<p>Data and code related to the article "Neural correlates of hierarchical predictive processes in autistic adults" by Laurie-Anne Sapey-Triomphe, Lauren Pattyn, Veith Weilnhammer, Philipp Sterzer and Johan Wagemans (Nature Communications):</p> <p>- Behavioral dataset of the 26 neurotypical participants (NT_behavioral_data.zip) and of the 26 autistic participants (ASD_behavioral_data.zip)</p> <p>- Source data of the graphics appearing in the article (Source data.xls)</p> <p>- Matlab codes used to run the experiment (Codes_to_run_experiment.zip)</p> <p>- Matlab codes to perform the main behavioral analyses (Codes_behavioral_analyses.zip) and to analyze the behavioral data with the HGF models (Codes_comput_model_analyses.zip)</p> <p>- Matlab codes to preprocess (Codes_fMRI_preprocessing.zip) and run the main fMRI analyses (Codes_fMRI_analyses.zip)</p>
Dataset for surface waves height prediction through the video and image processing
<p>Image-based study of surface waves is a long lasting topic in ocean science and remote sensing. We believe that modern computers and new programming techniques can make a break-through in this area.</p> <p> </p> <p>This dataset provides some video files of surface wind waves of two kinds. First is a video snapshot of a quite large area. Second one is a zoom-in video of a spar-buoy (a stick) located in this field. According to the zoom-in video we may see the actual height of the wave in this particular point. This should be treated as a reliable data and so it can be used to calibrate the brightness field. I.e. the users of this dataset are welcome to train their model to obtain the height of the wave out of its brightness on the zoom-out large-area videos.</p> <p> </p> <p>All video files are readable by a conventional software. Records were taken at mild wind conditions in a gulf (fjord or skerry) of the Ladoga Lake. See "readme.pdf" for the details</p>
Processed data and scripts supporting the manuscript "Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma"
<p>This submission contains the compiled count table, processed R objects and various scripts and output files accompanying our manuscript "Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma" (Nature Communications, 2023, https://doi.org/10.1038/s41467-023-38886-8)</p>
Modelling heterogeneity in the classification process in multi-species distribution models can improve predictive performance
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.