Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
159
datasets available to search
ShareScore release 0.9.0
Dataset results
159 results for “code prediction”
Data and code for Causality analysis and prediction of riverine algal blooms by combining empirical dynamic modeling and machine learning techniques
<p>Hydrological data (including daily water levels, flow velocities, and streamflow discharges) from two hydrological stations, the Hankou Station in the Yangtze River (YR) and the Hanchuan Station in the Han River (HR), were obtained from Hubei Province Hydrology and Water Resources Center.</p> <p>Water quality data (i.e., total nitrogen (TOTN), total phosphorus (TOTP), and water temperature in the Han River) and algae densities at three sections (Baihezui, Qinduankou and Zongguan) were acquired from the Yangtze River Basin Ecological and Environmental Supervision Authority. </p> <p><span>The R script(s) for machine learning models can also be found at <a href="../api/records/10901736/draft/files/Code%20for%20machine%20learning%20classification%20model.R/content" target="_blank" rel="noopener noreferrer">Code for machine learning classification model.R</a>.</span></p> <p> </p>
Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) - code and data sets
<p>This repository contains the scripts for the paper in revision to the AAPS J: Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) </p> <p>Authors: Niels Hendrickx, MSc, France Mentré, MD, PhD, Andreas Traschütz, MD, PhD, Cynthia Gagnon, PhD, Rebecca Schüle, MD, ARCA Study Group, EVIDENCE-RND consortium, Matthis Synofzik, MD, Emmanuelle Comets, PhD</p> <p>A simulated dataset (<strong>simulated_arsacs.csv</strong>) has been included in the repository to make the code executable as a standalone. Four main scripts have been provided in addition with the present Readme describing the files. The repository also includes 3 R objects and 2 folders which will be overwritten when the scripts are run, and are included as examples of the expected outputs. The main scripts are:</p> <p>- <strong>Script_imputation_selection.R</strong>: runs the covariate selection method. It uses a simulated dataset provided in the depot. The multiple imputation model is hardcoded as an input to the mice package to generate 10 imputed datasets, saved in current_directory/imputed_data_sets/df_arsacs_mi_i.csv. The script then runs the covariate selection method. The script prints out the list of selected covariates and returns a saemixObject containing the fit of the selected covariate model.<br> After the script executes, a list will be saved with the name of the selected covariates in the current directory (an example is included under the name "cov_matrix_model.RData" in the repository), the output of the selection, containing the whole history of runs will be saved under "final_covariate_model.RData", the list of selected covariate names will be saved under "list_covariates.RData".</p> <p>- <strong>source_mi.R</strong>: contains the functions used by Script_imputation_selection.R</p> <p>- <strong>script_bootstrap_indfit.R</strong>: This script loads "cov_matrix_model.RData" containing the matrix of covariate effects (used by saemix) and "list_covariates.RData", the list of covariates included, fits the model on the imputed data sets and computes its bootstrap distribution for each imputed data set (in the script, using only 20 samples for computation time, saved in current_directory/bootstrap/boot.arsacs.case.mi.i). It then computes the mean parameter and relative standard error of each parameter. It then computes the conditional distribution of each patient in each bootstrap samples and returns a data frame of individual predictions. The script will then plot 4 indivudal predictions. </p> <p>-<strong> source_bootstrap.R</strong>: contains the functions used by script_bootstrap_indfit.R</p> <p>Both scripts need the saemix package to run, which we haven’t included in the repository as it is freely available on the CRAN (https://cran.r-project.org/web/packages/saemix/index.html). Additional libraries we make use of in the code (MICE, tidyverse, ggplot2) also need to be installed prior to execution. <br>The R code provided can be further customised to be adapted to different scenarios.</p> <p>For the code to run, it is preferable to unzip the whole folder and set the working directory to the source file location as the script uses the "bootstrap" and "imputed_data_sets" sub-folders</p> <p>To execute this code, assuming the required libraries are available in the local R installation, please open an R session and run:<br>source("Script_imputation_selection.R") # for the covariate selection method (runtime: 3h on a i7-8565U laptop)<br>source("script_bootstrap_indfit.R") # to obtain individual trajectories (runtime: 1h on a i7-8565U laptop)</p>
Raw datasets and code generation for the strength prediction of TBC incorporating CCA and GOS
Open the record for dataset details and reuse information.
Raw strength datasets and code generation for predicting the strengths of TBC modified with SNA and OSP
Open the record for dataset details and reuse information.
Raw datasets and code generation for strength prediction of BCC incorporating CP and CCA
Open the record for dataset details and reuse information.
Original features and code of clinical prediction model
Open the record for dataset details and reuse information.
Data and code for: A sensory ecology of fear: Eye size predicts moonlight avoidance responses in Neotropical electric fishes
<p>Data in support of: Eye size predicts moonlight avoidance responses in Neotropical electric fishes</p>
Code and Training Data for "Cascaded Machine Learning of Soil Moisture and Salinity Prediction in Estuarine Wetlands based on In-situ Internet of Things Monitoring"
Open the record for dataset details and reuse information.
Code and Data for "Global Surface Eddy Mixing Ellipses: Spatio-temporal Variability and Machine Learning Prediction" By Jing et al. Submitted to Frontiers in Marine Science.
<p>This repository contains the code and data for the study of "Global Surface Eddy Mixing Ellipses: Spatio-temporal Variability and Machine Learning Prediction" By Jing et al. Submitted to Frontiers in Marine Science.</p> <p>Specifically, this repository contains the following items: </p> <p>(1) The codes needed for assessing the representation and prediction skills of Random Forest (RF), Convolutional Neural Network (CNN) and Spatial Transformer Networks (STN) models. </p> <p>(2) Original and normalized data to run these codes.</p> <p>(3) Code here is built on early work from our laboratory (Jaderberg et al., 2015; Guan et al., 2022; Zhang et al., 2023), though great modifications have been made tailored to our scientific question.</p> <p>[1] Jaderberg, M., Simonyan, K., Zisserman, A., et al. (2015). Spatial transformer networks. Advances in neural information processing systems, 28.</p> <p>[2] Guan, W., Chen, R., Zhang, H., Yang, Y., & Wei, H. (2022). Seasonal surface eddy mixing in the Kuroshio Extension: Estimation and machine learning prediction. Journal of Geophysical Research: Oceans, 127 (3), e2021JC017967.</p> <div>[3] Zhang, G., Chen, R., Li, X., Li, L., Wei, H., & Guan, W. (2023). Temporal variability of global surface eddy diffusivities: Estimates and machine learning prediction. Journal of Physical Oceanography, 53 (7), 1711–1730.</div>
Supplementary code for: "Historical glacier change on Svalbard predicts doubling of mass loss by 2100"
<p>Code to perform the analysis in:</p> <p>Geyman, E.C., van Pelt, W., Maloof, A.C., Faste Aas, H., and Kohler, J., 2021. "Historical glacier change on Svalbard predicts doubling of mass loss by 2100." Nature.</p> <p>Abstract:</p> <p>The melting of glaciers and ice caps accounts for about one third of current sea level rise, exceeding the mass loss from the more voluminous Greenland or Antarctic Ice Sheets. The Arctic archipelago of Svalbard, which hosts spatial climate gradients that are larger than the expected temporal shifts over the next century, is a natural laboratory to constrain the climate sensitivity of glaciers and predict their response to future warming. Leveraging an archive of historical aerial images from 1936 and 1938, we use structure-from-motion (SfM) photogrammetry to reconstruct the 3D geometry of 1,594 glaciers across Svalbard. We compare these reconstructions to modern ice elevation data to derive the spatial pattern of mass balance over a >70-year timespan, allowing us to see through the noise of annual and decadal variability to quantify how variables such as temperature and precipitation control ice loss. We find a robust temperature dependence of melt rates, whereby a 1°C rise in mean summer temperature corresponds to a decrease in area-normalized mass balance of -0.27 m yr<sup>-1</sup> of water equivalent. Finally, we design a space-for-time substitution to make first-order predictions of 21st century glacier change across Svalbard. Even in the most modest scenario (a ~1.4°C rise in mean summer temperature by 2100), we predict average glacier thinning rates in 2010-2100 of -0.67 m yr<sup>-1</sup>, approximately twice the 1936-2010 rates.</p>
Predictive utility of task-related functional connectivity vs. voxel activation - Data and code archive
<p>Functional connectivity, both in resting state and task performance, has steadily increased its share of neuroimaging research effort in the last 1.5 decades. In the current study, we investigated the predictive utility regarding behavioral performance and task information for 240 participants, aged 20-77, for both voxel activation and functional connectivity in 12 cognitive tasks, belonging to 4 cognitive reference domains (Episodic Memory, Fluid Reasoning, Perceptual Speed, and Vocabulary). We also added a model only comprising brain-structure information not specifically acquired during performance of a cognitive task. We used a simple brain-behavioral prediction technique based on Principal Component Analysis (PCA) and regression and studied the utility of both modalities in quasi out-of-sample predictions, using split-sample simulations (=5-fold Monte Carlo cross validation) with 1,000 iterations for which a regression model predicting a cognitive outcome was estimated in a training sample, with a subsequent assessment of prediction success in a non-overlapping test sample. The sample assignments were identical for functional connectivity, voxel activation, and brain structure, enabling apples-to-apples comparisons of predictive utility. All 3 models that were investigated included the demographic covariates age, gender, and years of education. A minimal reference model using simple linear regression with just these 3 covariates was included for comparison as well and was evaluated with the same resampling scheme as described above. Results of the comparison between voxel activation and functional connectivity were mixed and showed some dependency on cognitive outcome; however, mean differences in predictive utility between voxel activation and functional connectivity were rather small in terms of within-modality variability or predictive success. More notably, only in the case of Fluid Reasoning did concurrent functional neuroimaging provided compelling about cognitive performance beyond structural brain imaging or the minimal reference model.</p>
Build Prediction in Continuous Integration Using Textual Analysis of Source Code and Traditional Software Metrics
<p>Continuous Integration (CI) systems integrate code changes committed by software developers, tests the results of the integration, and feed developers with information about the outcome of the integration and testing. Predicting the outcome of the integration is important since it reduces the feedback time between the CI system and the developers. This data-set comprises of historical code changes extracted from the TravisTorrent data-set (found in the train-lines folder) and their corresponding feature vectors (found in the train-bag-of-words folder) for Java projects. It also includes a set of files that contains historical build records and a set of traditional software metrics.</p>
Code and data for "Seasonal Surface Eddy Mixing in the Kuroshio Extension: Estimation and Machine Learning Prediction" By Guan et al. Submitted to JGR Oceans.
<p>This repository contains the code and data for the machine learning analysis of "Seasonal Surface Eddy Mixing in the Kuroshio Extension: Estimation and Machine Learning Prediction”By Guan et al. Submitted to JGR Oceans.</p> <p>Specifically, this repository contains the following items:<br> (1) Codes for assessing the representation skill of the machine learning and linear regression (LR) methods. Three machine learning methods are considered: random forest (RF), back-propagation neural network (BP), and convolutional neural network (CNN).<br> (2) Codes for assessing the prediction skill of the machine learning and LR methods. <br> (3) Seasonal-mean and annual-mean input data to run these codes. <br> (4) The package needed to run the random forest code, i.e. the RF_MexStandalone-v0.02 program package from https://code.google.com/archive/p/randomforest-matlab/downloads .</p>
Code for: Comparison and interpretability of machine learning models to predict severity of chest injury
<p><span><span><span><span><span><span><span><span><span><span><span><b>Objective:</b> Trauma quality improvement programs and registries improve care and outcomes for injured patients. Designated trauma centers calculate injury scores using dedicated trauma registrars; however, many injuries arrive at non-trauma centers, leaving a substantial amount of data uncaptured. We propose automated methods to identify severe chest injury using machine learning (ML) and natural language processing (NLP) methods from the electronic health record (EHR) for quality reporting.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Materials and Methods:</b> A level I trauma center was queried for patients presenting after injury between 2014 and 2018. Prediction modeling was performed to classify severe chest injury using a reference dataset labeled by certified registrars. Clinical documents from trauma encounters were processed into concept unique identifiers for inputs to ML models: logistic regression with elastic net regularization (EN), extreme gradient boosted machines (XGB), and convolutional neural networks (CNN). The optimal model was identified by examining predictive and face validity metrics using global explanations.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Results:</b> Of 8,952 encounters, 542 (6.1%) had a severe chest injury. CNN and EN had the highest discrimination, with an area under the receiver operating characteristic curve of 0.93 and calibration slopes between 0.88 and 0.97. CNN had better performance across risk thresholds with fewer discordant cases. Examination of global explanations demonstrated the CNN model had better face validity, with top features including "contusion of lung" and "hemopneumothorax." </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Discussion: </b>The CNN model featured optimal discrimination, calibration, and clinically relevant features selected. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Conclusion:</b> NLP and ML methods to populate trauma registries for quality analyses are feasible.</span></span></span></span></span></span></span></span></span></span></span></p>
Deficits of hierarchical predictive coding in left spatial neglect
<p>Right brain-damaged patients with unilateral spatial neglect fail to explore the left side of space. Recent EEG and clinical evidence suggests that neglect patients might suffer deficits in predictive coding, i.e. in identifying and exploiting probabilistic associations among sensory stimuli in the environment. To gain direct insights on this issue, we focussed on the hierarchical components of predictive coding. We recorded EEG responses evoked by central, left-side or right-side tones that were presented at the end of sequences of four central tones. Left-side and right-side deviant tones produce a pre-attentive Mismatch Negativity that reflects a lower-order prediction error for the ‘Local’ deviation of the tone at the end of the sequence. Higher-order prediction errors for the frequency of these deviations in the acoustic environment, i.e. ‘Global’ deviation, are marked by the P3 response. We show that when neglect patients are immersed in an acoustic environment characterized by frequent left-side deviant tones, they display no pre-attentive Mismatch Negativity both for left-side deviant tones and infrequent omissions of the last tone, while they have Mismatch Negativity for infrequent right-side deviant tones. In the same condition, neglect patients show no P300 response to ‘Global’ prediction errors for deviant tones, including those in the non-neglected right-side, and omissions. In contrast to this, when right-side deviant tones are predominant in the acoustic environment, neglect patients have pre-attentive Mismatch Negativity both for right-side deviant tones and infrequent omissions, while they display no Mismatch Negativity for infrequent left-side deviant tones. Most importantly, in the same condition neglect patients show enhanced P300 response to infrequent left-side deviant tones, notwithstanding that these tones evoked no pre-attentive Mismatch Negativity. This latter finding indicates that ‘Global’ predictions are independent of ‘Local’ error signals provided by the Mismatch Negativity. These results qualify deficits of predictive coding in the spatial neglect syndrome and show that neglect patients base their predictive behaviour only on statistical regularities that are related to the frequent occurrence of sensory events on the right side of space.</p>
Code for Atmospheric Research publication - Height correction method based on the Monin–Obukhov similarity theory for better prediction of near-surface wind fields
<p>In this repository, we include the source codes for WRF namelist, figures, and height correction used in the Atmospheric Research publication "Height correction method based on the Monin–Obukhov similarity theory for better prediction of near-surface wind fields"</p> <p>The namelist.wps and namelist.input in WRF namelist are using for making input and running simulation, and Fig scripts in Figure scripts are using for plotting the figures in the paper.</p> <p>hgt_corr in Height correction method is a code to correct the disparity of the 10-m height definition between the model and observation by applying the developed the height correction algorithm based on the Monin-Obukhov similarity theory.</p>
Build Prediction in Continuous Integration Using Textual Analysis of Source Code
<p>The data-set comprises of software metrics that characterize build outcomes in CI using traditional and token frequency metrics.</p>
Supporting data and code for High Concentrations of Nanoparticles from Isoprene Nitrates Predicted in Convective Outflow Over the Amazon
<p>The Archive includes the CloudChem model code, which is used to simulate gas-phase chemical reactions and gas interactions with clouds. Additionally, it includes the simulated gas and particle concentration dataset used to produce figures for the paper "High Concentrations of Low Volatility Isoprene Nitrates Predicted in Convective Outflow Over the Amazon". For more information about the simulation setup and the results, please refer to the paper.</p>
Code Generation for Predicting the Radium Activity Equivalent
Open the record for dataset details and reuse information.
Data for the thesis "Exploring Heuristics for Predicting Microbenchmark Stability and Code Coverage using Static Code Analysis"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.