Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

77

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

77 results for “predictive processing”

Learn how ShareScore rates datasets ↗
zenodo48/100

pKaDatabase for Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge

<p>A curated a database of small molecules with experimentally measured pKa values.&nbsp;</p> <p>This pickle file can be loaded into memory using Pandas. In the code block below we will print out the columns of the DataFrame:</p> <pre><code class="language-python">import pandas as pd df = pd.load("pKaDatabase.pkl") print(df.keys()). # print the columns</code></pre> <blockquote> <p>[&#39;deprotonated microstate ID&#39;, &#39;protonated microstate ID&#39;, &#39;deprotonated microstate smiles&#39;, &#39;protonated microstate smiles&#39;, &#39;AM1BCC partial charge (prot. atom)&#39;, &#39;AM1BCC partial charge (deprot. atom)&#39;, &#39;AM1BCC partial charge (prot. atoms 1 bond away)&#39;, &#39;AM1BCC partial charge (deprot. atoms 1 bond away)&#39;, &#39;AM1BCC partial charge (prot. atoms 2 bond away)&#39;, &#39;AM1BCC partial charge (deprot. atoms 2 bond away)&#39;, &#39;Gasteiger partial charge (prot. atom)&#39;, &#39;Gasteiger partial charge (deprot. atom)&#39;, &#39;Gasteiger partial charge (prot. atoms 1 bond away)&#39;, &#39;Gasteiger partial charge (deprot. atoms 1 bond away)&#39;, &#39;Gasteiger partial charge (prot. atoms 2 bond away)&#39;, &#39;Gasteiger partial charge (deprot. atoms 2 bond away)&#39;, &#39;Extented H&uuml;ckel partial charge (prot. atom)&#39;, &#39;Extented H&uuml;ckel partial charge (deprot. atom)&#39;, &#39;Extented H&uuml;ckel partial charge (prot. atoms 1 bond away)&#39;, &#39;Extented H&uuml;ckel partial charge (deprot. atoms 1 bond away)&#39;, &#39;Extented H&uuml;ckel partial charge (prot. atoms 2 bond away)&#39;, &#39;Extented H&uuml;ckel partial charge (deprot. atoms 2 bond away)&#39;, &#39;∆G_solv (kJ/mol) (prot-deprot)&#39;, &#39;SASA (Shrake)&#39;, &#39;SASA (Lee)&#39;, &#39;Bond Order&#39;, &#39;Change in Enthalpy (kJ/mol) (prot-deprot)&#39;, &#39;pKa&#39;,&#39;href&#39;, &#39;num ionizable groups&#39;, &#39;Weight&#39;, &#39;pKa source&#39;]</p> </blockquote> <p>&nbsp;</p> <p>For more information regarding feature calculations, please read&nbsp;the following paper:</p> <blockquote> <p>Raddi, Robert, and Vincent Voelz. &quot;Stacking Gaussian Processes to Improve pKa Predictions in the SAMPL7 Challenge.&quot; (2021).&nbsp;<a href="https://doi.org/10.26434/chemrxiv.14650302.v1">10.26434/chemrxiv.14650302.v1</a></p> </blockquote>

opencc-by-4.0Jun 2021View details →
edi48/100

Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements

The Forest and Biodiversity (FAB2) experiment uses native tree species in varying levels of species richness, phylogenetic diversity, and functional diversity planted in 100 m2 and 400 m2 plots at 1 m spacing, appropriate for testing long-term ecosystem consequences. FAB2 was designed and established in conjunction with a prior experiment (FAB1) in which the same set of twelve species was planted in 16 m2 plots at 0.5 m spacing. This data package examines the connections between aboveground and belowground processes in FAB2. This data package includes information on tree diversity and community composition, forest structure, forest understories, soil microbes, net nitrogen mineralization, and canopy nitrogen. A wide variety of data types are included, such as data from hyperspectral and LiDAR remote sensing, percent cover analysis, soil microbial analyses, and soil assays including C:N, pH, and net nitrogen mineralization. This data package is included in the submission of the manuscript entitled “Predicting aboveground and belowground processes in diverse forest ecosystems using remote sensing and in-situ measurements.”

openCC0Jan 2026View details →
zenodo44/100

Dataset: Brain negativity as an indicator of predictive error processing: The contribution of visual action effect monitoring

<p>There are two files for each subject:</p> <p>1. sub##_error.dat -&gt; Contains EEG Segments, that were recorded while the subject executed a clear target miss (minimal distance between the center of the ball and target &gt; 12 cm) in the task (segment and electrode information can be found below).</p> <p>2. sub##_hit.dat -&gt; Contains EEG Segments, that were recorded while the subject executed a clear target hit (minimal distance between the center of the ball and the target &lt; 7 cm) in the task (segment and electrode information can be found below).</p> <p><br> The data in the *.dat-files are stored in a two dimensional matrix: n*1400 datapoints x 15 electrodes</p> <p>n represents the number of segments. 1400 datapoints per segment translate to a segment length of 2800 ms (from 600 ms before to 2200 ms after ball release). The ball´s release is located at the 301st datapoint and the feedback was presented at datapoint 726  (850 ms after ball release) in every segment.</p> <p>datapoints: The first dimension (rows) includes the measured neural activations in microvolts. The data is stored vectorized,<br> i.e. hit/error #1 -&gt; row 1 to 1400, hit/error #2 -&gt; row 1401 to 2800, ..., hit/error #n -&gt; (n-1) * 1400 + 1 to n * 1400</p> <p>electrodes: The second dimension (columns) consists of the 15 different electrodes that were used during data recording in this exact order: [F3 Fz F4 C4 Cz C3 P3 Pz P4 VEOGu VEOGo HEOGre HEOGli FCz Mastre]</p>

opencc-by-4.0May 2017View details →
zenodo44/100

Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models

<p>This data set contains the data, JMP scripts, and figures of the article titled &quot;Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models&quot; to be published in the journal Animal - Open Space.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Ion Implantation Sensor and Process Target Data for Predicting Ion Beam Tuning in Semiconductor Manufacturing

<h2><strong>Dataset Description:</strong></h2> <p>This dataset is designed to predict ion beam tuning setup processes in semiconductor manufacturing, in terms of tuning success or failure, and tuning duration. It is split into&nbsp;<strong><code>X</code></strong> and <code><strong>y</strong></code> to allow for supervised learning approaches.</p> <ul> <li><code><strong>X</strong></code> represents the current equipment condition and the process targets of the currently processed and the upcoming lot, as defined within recipes.</li> <li><code><strong>y</strong></code> represents the ion beam tuning setup report, which informs about the tuning success ratio and tuning duration. These setups are necessary, when switching between recipes to prepare the equipment for processing the next lot.&nbsp;<strong><code>y</code></strong> contains three labels, enabling classification of (1) tuning success or fail, and (2) prolonged tuning, as well as (3) estimation of tuning duration as a regression task.</li> </ul> <p>About <strong><code>X</code></strong>:</p> <p>Each lot is processed with a specific recipe to achieve the process target. The tuning takes place before the first wafer of the to-be-tuned recipe is processed. Each row in <strong><code>X</code></strong> includes logistical information such as the equipment used for processing and parsed recipe / process target information for the current and upcoming lot. The majority of data consists out of aggregated metrics of equipment-internally tracked sensor traces, recording physical parameters such as gas flows, temperatures, voltages and currents. When analyzed in conjunction with the processed recipe, these sensors provide insights into the current equipment condition.&nbsp;</p> <p>About <code><strong>y</strong></code>:</p> <p>The&nbsp;<code>setup_result</code> column indicates the success or failure of tuning - with <code>setup_result=0</code> indicating tuning success, while&nbsp;<code>setup_result=1</code> signals tuning failure. If the first tuning attempt fails, there may be follow-up attempts, but these are not included in this dataset. The&nbsp;<code>duration</code> column represents the tuning duration in seconds, as used for regression analysis. The&nbsp;<code>duration_interval</code> column is a binary label for prolonged tunings, i.e. <code>duration_interval=1</code> for instances, which take more than 6 minutes to tune.</p> <p>For reproducibility of the corresponding paper's results:</p> <ol> <li>The dataset contains the same carefully curated subset of features.</li> <li>The train_test_split() has already been performed, thus we provide&nbsp;<code>x_train</code> and <code>x_valid</code> separately.</li> <li>To reduce the effect of outliers in the data, the sensor data has already been scaled, as derived from&nbsp;<code>x_train</code>.</li> </ol> <p>In summary, these datasets (<code><strong>X</strong></code>, <code><strong>y</strong></code>) provide comprehensive information for predicting ion beam tuning in semiconductor manufacturing, making it a valuable resource for researchers and practitioners in the field.</p> <h2><strong>Python Code for Reproducibility:</strong></h2> <p>Furthermore, we share a jupyter notebook <code>ionbeamtuning.ipynb</code> with Python code to train the best performing model on the provided data, as described in the paper. To execute the code, you may need to install any missing packages specified in the <code>requirements.txt</code>, as indicated within the notebook.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Pathobionts in the tumour microbiota predict survival following resection for colorectal cancer - pre-processed data

<p>A multicentre, prospective observational study was conducted of colorectal cancer (CRC) patients undergoing primary surgical resection in the United Kingdom and Czech Republic. Analysis was performed using metataxonomics (microbiome) and ultra-performance liquid chromatography mass spectrometry (UPLC-MS, metabolomics). Both datasets were pre-processed as described in the methods section of the main article. The data here were used as the input to the data analysis workflows available from <a href="https://github.com/jmp111/CRC">Github</a>.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Potential distribution of invasive boxwood blight pathogen (Calonectria pseudonaviculata) as predicted by process-based and correlative models

<p>R project, R scripts, and data files for reproducing most of the analyses presented in a climatic suitability study for boxwood blight. The README. md file describes how to run the scripts and provides details on data inputs.</p> <p><strong>Abstract: </strong>Boxwood blight caused by <em>Cps</em> is an emerging disease that has had devastating impacts on <em>Buxus</em> spp. in the horticultural sector, landscapes, and native ecosystems. In this study, we produced a process-based climatic suitability model in the CLIMEX program and combined outputs of four different correlative modeling algorithms to generate an ensemble correlative model. All models were fit and validated using a presence record dataset comprised of <em>Cps</em> detections across its entire known invaded range. Evaluations of model performance provided validation of good model fit for all models. A consensus map of CLIMEX and ensemble correlative model predictions indicated that not-yet-invaded areas in eastern and southern Europe and in the southeastern, midwestern, and Pacific coast regions of North America are climatically suitable for <em>Cps</em> establishment. Most regions of the world where<em> Buxus</em> and its congeners are native are also at risk of establishment. These findings provide the first insights into <em>Cps</em> global invasion threat, suggesting that this invasive pathogen has the potential to significantly expand its range.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models

<p>This data set contains the data, JMP scripts, and figures of the article titled &quot;Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models&quot; to be published in the journal Animal - Open Space.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Interception of virtual throws reveals predictive skills based on the visual processing of throwing kinematics - Dataset

<p>Dataset consists of a list of Matlab structures, one for each of the 21 participants. For each participant all the recorded trials are reported (&quot;trials&quot; field). For each trial, the dataset reports information about the associated experimental condition and the kinematics of the ball and the racket trajectories. &nbsp;Information about the experimental condition are specified in the &quot;info&quot; field, which provides the experimental phase (Training and Experimental), the visibility &nbsp;(AllVisible, ThrowerOnly, BallOnly), the target (1, 2, 3, 4), and the thrower ID (1, 2, 3, 4). The kinematics data, starting from the time of ball release, are given in the field &quot;trajectories&quot;, which provides the time vector (in seconds), and the corresponding 3D positions of the racket and the ball (in meters) in the reference frame shown in Figure 1 of the paper.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Figure 5. Sensory score and period of storage for processed cheese-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese

<p>R2 was found to be 96.5 percent of the total variation as explained by sensory scores. Period<br> of storage (days) for which the processed cheese has been in the shelf can be determined based on<br> sensory score (Fig. 5).</p>

opencc-by-4.0Jan 2012View details →
zenodo40/100

Figure 4. Comparison of ASS and PSS for multilayer model R-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese

<p>TDNN models with single and multi layers were developed taking soluble nitrogen, pH,<br> standard plate count, yeast &amp; mould count, spore count as input parameters, and sensory score as<br> output parameter for predicting the shelf life of processed cheese stored at 30o C. Mean Square<br> Error, Root Mean Square Error, Coefficient of Determination and Nash - Sutcliffo Coefficient were<br> used in order to compare the prediction ability of the developed TDNN models. Regression<br> equations were developed for predicting the shelf life of processed cheese, which came out as 28.25<br> days. Since, predicted value is close to the experimentally determined shelf life of 30 days, hence<br> from the study it can be concluded that TDNN artificial neural network models are quite efficient in<br> predicting shelf life of processed cheese.</p>

opencc-by-4.0Jan 2012View details →
zenodo40/100

Figure 2. Training pattern of TDNN models-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese

<p>The Neural Network Toolbox under MATLAB software was used for developing the TDNN<br> models. Training pattern of TDNN models is presented in Fig.2.</p>

opencc-by-4.0Jan 2012View details →
zenodo40/100

Figure 1. Inputs and output parameters for TDNN models-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese

<p>The data consisted of 36 samples, which were divided into two subsets, i.e., 30 used for<br> training the network and 6 for testing the TDNN models. Soluble nitrogen, pH, standard plate<br> count, yeast &amp; mould count, and spore count were taken as input parameters, and sensory score as<br> output parameter for developing TDNN single and multilayer models (Fig.1).</p>

opencc-by-4.0Jan 2012View details →
zenodo40/100

Figure 3. Comparison of ASS and PSS single layer model-Time-Delay Artificial Neural Network Computing Models for Predicting Shelf Life of Processed Cheese

<p>TDNN models with single and multi layers were developed taking soluble nitrogen, pH,<br> standard plate count, yeast &amp; mould count, spore count as input parameters, and sensory score as<br> output parameter for predicting the shelf life of processed cheese stored at 30o C. Mean Square<br> Error, Root Mean Square Error, Coefficient of Determination and Nash - Sutcliffo Coefficient were<br> used in order to compare the prediction ability of the developed TDNN models. Regression<br> equations were developed for predicting the shelf life of processed cheese, which came out as 28.25<br> days. Since, predicted value is close to the experimentally determined shelf life of 30 days, hence<br> from the study it can be concluded that TDNN artificial neural network models are quite efficient in<br> predicting shelf life of processed cheese.</p>

opencc-by-4.0Jan 2012View details →
zenodo40/100

AbDb processed and pickled for use in deep learning CDR-H3 Structure prediction

<p>This is a pickle file, ready for training by the neural network described in &quot;Improving CDR-H3 modelling in Antibodies&quot; found at the following URL:</p> <p><a href="https://github.com/OniDaito/MRes">https://github.com/OniDaito/MRes</a></p> <p>The data is derived from the AbDb dataset found at:</p> <p><a href="http://www.bioinf.org.uk/abs/abdb/">http://www.bioinf.org.uk/abs/abdb/</a></p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading"

<p>Title of dataset: Data for "Using physics-informed neural networks to predict the lifetime of laser powder bed fusion processed 316L stainless steel under multiaxial low-cycle fatigue loading".</p> <p>Name/institution/contact information: Dr. Michal Barto&scaron;&aacute;k, Czech Technical University in Prague - Faculty of Mechanical Engineering, email: michal.bartosak@fs.cvut.cz.</p> <p>Date of data collection: The data were collected between 2021 and 2024.</p> <p>File name structure: The data consists of two files: "316L_fatigue_and_defects.xls," which contains fatigue lifetime data and defect characteristics, and an associated description file, "read_me.txt."</p> <p>See "https://doi.org/10.1016/j.ijfatigue.2024.108608" for the associated article and a detailed description of the methods.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Dataset and codes of the article "Neural correlates of hierarchical predictive processes in autistic adults"

<p>Data and code related to the article&nbsp;&quot;Neural correlates of hierarchical predictive processes in autistic adults&quot;&nbsp; by Laurie-Anne Sapey-Triomphe, Lauren Pattyn, Veith Weilnhammer, Philipp Sterzer and Johan Wagemans (Nature Communications):</p> <p>-&nbsp;Behavioral dataset&nbsp;of the 26 neurotypical participants (NT_behavioral_data.zip) and of the 26 autistic participants (ASD_behavioral_data.zip)</p> <p>- Source data of the graphics appearing in the article (Source data.xls)</p> <p>- Matlab codes used to run the experiment (Codes_to_run_experiment.zip)</p> <p>- Matlab codes to perform&nbsp;the main behavioral analyses (Codes_behavioral_analyses.zip) and to analyze the behavioral data with the HGF models (Codes_comput_model_analyses.zip)</p> <p>- Matlab codes to preprocess (Codes_fMRI_preprocessing.zip) and run the main fMRI analyses (Codes_fMRI_analyses.zip)</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Dataset for surface waves height prediction through the video and image processing

<p>Image-based study of surface waves is a long lasting topic in ocean science and remote sensing. We believe that modern computers and new programming techniques can make a break-through in this area.</p> <p>&nbsp;</p> <p>This dataset provides some video files of surface wind waves of two kinds. First is a video snapshot of a quite large area. Second one is a zoom-in video of a spar-buoy (a stick) located in this field. According to the zoom-in video we may see the actual height of the wave in this particular point. This should be treated as a reliable data and so it can be used to calibrate the brightness field. I.e. the users of this dataset are welcome to train their model to obtain the height of the wave out of its brightness on the zoom-out large-area videos.</p> <p>&nbsp;</p> <p>All video files are readable by a conventional software. Records were taken at mild wind conditions in a gulf (fjord or skerry) of the Ladoga Lake. See &quot;readme.pdf&quot; for the details</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Processed data and scripts supporting the manuscript "Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma"

<p>This submission contains the compiled count table,&nbsp;processed R objects and various scripts and output files&nbsp;accompanying our manuscript &quot;Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma&quot; (Nature Communications, 2023,&nbsp;https://doi.org/10.1038/s41467-023-38886-8)</p>

opencc-by-4.0May 2023View details →
dryad40/100

Modelling heterogeneity in the classification process in multi-species distribution models can improve predictive performance

Open the record for dataset details and reuse information.

publicMar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record