Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21
datasets available to search
ShareScore release 0.9.0
Dataset results
21 results for “Probabilistic learning”
EEG: Probabilistic Learning with Affective Feedback: Exp #2
Open the record for dataset details and reuse information.
EEG: Probabilistic Learning with Affective Feedback: Exp #1
Open the record for dataset details and reuse information.
Probabilistic volumetric speckle suppression in OCT using deep learning: Dataset
<p>This file contains a retinal OCT intensity volume as a demo dataset to generate volumetric speckle-suppressed training data using our non-local-means despeckling (TNode) script and four OCT intensity volumes and their corresponding TNode-processed intensity volumes of different tissue samples to train and test our deep learning framework used in "Probabilistic volumetric speckle suppression in OCT using deep learning" by Chintada et al. 2023. The TNode code for generating the training data and the source code for our deep learning framework are available at https://github.com/bhaskarachintada/DLTNode.git</p>
Data from: Learning of probabilistic punishment as a model of anxiety produces changes in action but not punisher encoding in the dmPFC and VTA
<p>Previously, we developed a novel model for anxiety during motivated behavior by training rats to perform a task where actions executed to obtain a reward were probabilistically punished and observed that after learning, neuronal activity in the ventral tegmental area (VTA) and dorsomedial prefrontal cortex (dmPFC) represent the relationship between action and punishment risk (Park & Moghaddam, 2017). Here we used male and female rats to expand on the previous work by focusing on neural changes in the dmPFC and VTA that were associated with the learning of probabilistic punishment, and anxiolytic treatment with diazepam after learning. We find that adaptive neural responses of dmPFC and VTA during the learning of anxiogenic contingencies are independent from the punisher experience and occur primarily during the peri-action and reward period. Our results also identify peri-action ramping of VTA neural calcium activity, and VTA-dmPFC correlated activity, as potential markers for the anxiolytic properties of diazepam.</p>
Data from: Learning of probabilistic punishment as a model of anxiety produces changes in action but not punisher encoding in the dmPFC and VTA
Open the record for dataset details and reuse information.
Data and code for paper "A gray-box model for a probabilistic estimate of regional ground magnetic perturbations: Enhancing the NOAA operational Geospace model with machine learning"
<p>Simulation results from the NOAA/SWPC Geospace model used in the paper Camporeale et al. (2020) "A gray-box model for a probabilistic estimate of regional ground magnetic perturbations: Enhancing the NOAA operational Geospace model with machine learning" published in J. Geophys. Res. (2020)</p> <p>MATLAB code is provided to train process the data and train the machine learning model and plot results.</p> <p>Manuscript available on <a href="https://arxiv.org/abs/1912.01038">https://arxiv.org/abs/1912.01038</a></p>
Interpretable Deep Learning for Probabilistic MJO Prediction: CNN Forecasts
<p>This repository contains data produced for the paper "Interpretable Deep Learning for Probabilistic MJO Prediction" by A. Delaunay and H. M. Christensen (2021).</p> <p>>> mu_ens_XX.pt<br> contains the mean forecasts from each ensemble member at a lead time of XX days</p> <p>>> cov_alea_XX.pt<br> contains the aleatoric predictions of each ensemble member at a lead time of XX days</p>
Dataset for MRes Thesis: Probabilistic Operator Learning for Climate Model Parameterisation
<p>This dataset was collated for use in experiments presented in the MRes Thesis "Probabilistic Operator Learning for Climate Model Parameterisation" submitted to the University of Cambridge.<br><br>All data included was generated by other researchers, this is simply a subset to allow easy reproduction of the experiments contained in the work above.</p> <p>The data for the Burgers' equation (<a href="../api/records/12529654/draft/files/burgers_data_R10.mat/content" target="_blank" rel="noopener noreferrer">burgers_data_R10.mat</a>) and the Darcy Flow (<a href="../api/records/12529654/draft/files/piececonst_r421_N1024_smooth1.mat/content" target="_blank" rel="noopener noreferrer">piececonst_r421_N1024_smooth1.mat</a>, <a href="../api/records/12529654/draft/files/piececonst_r421_N1024_smooth2.mat/content" target="_blank" rel="noopener noreferrer">piececonst_r421_N1024_smooth2.mat</a>) experiments were generated by Lu et al. (2022). Creative Commons Attribution Non Commercial Share Alike 4.0 International applies.</p> <p>The data for the Helmholtz (<a href="../api/records/12529654/draft/files/Helmholtz_inputs.npy/content" target="_blank" rel="noopener noreferrer">Helmholtz_inputs.npy</a>, <a href="../api/records/12529654/draft/files/Helmholtz_outputs.npy/content" target="_blank" rel="noopener noreferrer">Helmholtz_outputs.npy</a>) and Navier-Stokes (<a href="../api/records/12529654/draft/files/NavierStokes_inputs.npy/content" target="_blank" rel="noopener noreferrer">NavierStokes_inputs.npy</a>, <a href="../api/records/12529654/draft/files/NavierStokes_outputs.npy/content" target="_blank" rel="noopener noreferrer">NavierStokes_outputs.npy</a>) experiments were generated by de Hoop et al. (2022). Creative Commons Attribution 4.0 International applies.</p> <p>References:</p> <div> <div> <div>1. Lu L, Meng X, Cai S, Mao Z, Goswami S, Zhang Z, et al. A comprehensive and fair comparison of two neural operators (with practical extensions) based on FAIR data. Computer Methods in Applied Mechanics and Engineering. 2022 Apr 1;393:114778.</div> <div> </div> <div> <div> <div> <div>2. de Hoop MV, Huang DZ, Null EQ, Stuart AM. The Cost-Accuracy Trade-Off in Operator Learning with Neural Networks. JML. 2022 Jun;1(3):299–341.</div> </div> </div> </div> </div> </div> <p> </p>
Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations: Preprocessed Satellite and In-situ observation datasets
<p>This record includes all of the prepared data used in the manuscript, "Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations" (citation information forthcoming). As a part of this manuscript, we analyzed the ability for machine learning models to extract sea surface information (from salinity, temperature, sea height anomaly) to predict mixed layer depth. In this manuscript there are two experimental datasets: (1) info derived from CESM POP2 ocean model dataset (1989-1998), and (2) info derived from a combination of satellite sources and MLD from Argo profiles. More details below. </p> <p>All of these data files are preprocessed and organized to be used with the ml-ocean-bl github code found at https://github.com/NCAR/ml-ocean-bl/mloceanbl/.</p> <ul> <li><strong>CESM POP2 Ocean model dataset</strong></li> </ul> <p>Preprocessed sea surface salinity (SSS), temperature (SST), sea surface height anomalies (SSH), and ocean mixed layer depth (MLD, or HMXL) derived from the CESM POP2 Ocean model. Specifically, CESM POP2 model in a hindcast forced by JRA55do atmospheric reanalysis from 1958 to present and initialized with an oceanic climatology as in e.g. <a href="https://journals.ametsoc.org/view/journals/phoc/aop/JPO-D-20-0217.1/JPO-D-20-0217.1.xml">Deppenmeier et al. (2021)</a>. The model outputs include the ocean mixed layer depth (MLD), sea surface salinity (SSS), sea surface temperature (SST), and sea height anomaly (SSH) at a temporal frequency of 5-days and an approximate latitude and longitude resolution of 0.1 degrees.</p> <p>Relevant files:</p> <ol> <li>full_EPO.nc, full_SIO.nc <ul> <li>NetCDF4 containing SSS, SST, SSH, MLD for the equatorial Pacific Ocean (EPO) and southern Indian Ocean (SIO) (see manuscript for details). Data is regridded onto a 1/2 degree lat/lon 5 day grid to correspond with data used for Argo datasets (see below).</li> </ul> </li> <li>clim_EPO.nc, clim_SIO.nc, clim_std_EPO.nc, std_clim_EPO.nc, std_clim_SIO.nc <ul> <li>NetCDF4 containing mean and standard deviation climatologies of SSS, SST, SSH, and MLD for EPO and SIO.</li> </ul> </li> <li>std_anomalies_EPO.nc, std_anomalies_SIO.nc <ul> <li>NetCDF4 containing SSS, SST, SSH, and MLD standardized anomalies for EPO and SIO. This is the dataset directly used for training in aforementioned manuscript. Use with ml-ocean-bl/ml-ocean-test/data. </li> </ul> </li> </ol> <ul> <li><strong>Satellite and Argo datasets</strong></li> </ul> <p>Preprocessed satellite sea surface salinity (SSS), temperature (SST), and sea surface height anomalies (SSH) and Argo-based mixed layer depth (MLD) profiles. Original data can be found at:</p> <p>(SST): Remote Sensing Systems. 2017. MW optimum interpolated SST data set. Ver. 5.0. PO.DAAC, CA, USA. Further information available at at <a href="https://doi.org/10.5067/GHMWO-4FR05">https://doi.org/10.5067/GHMWO-4FR05</a>. Data can be accessed at https://podaac-tools.jpl.nasa.gov/drive/files/allData/ghrsst/data/GDS2/L4/GLOB/REMSS/mw_OI/v5.0/.</p> <p>(SSS): Oleg Melnichenko. 2018. Aquarius L4 Optimally Interpolated Sea Surface Salinity. Ver. 5.0. PO.DAAC, CA, USA. Further information at <a href="https://doi.org/10.5067/AQR50-4U7CS">https://doi.org/10.5067/AQR50-4U7CS</a>. Data can be accessed at https://podaac-tools.jpl.nasa.gov/drive/files/SalinityDensity/aquarius/L4/IPRC/v5/7day. </p> <p>(SSH): Zlotnicki, Victor; Qu, Zheng; Willis, Joshua. 2019. SEA_SURFACE_HEIGHT_ALT_GRIDS_L4_2SATS_5DAY_6THDEG_V_JPL1609. Ver. 1812. PO.DAAC, CA, USA. Information available at <a href="https://doi.org/10.5067/SLREF-CDRV2">https://doi.org/10.5067/SLREF-CDRV2</a>. Data can be accessed at https://podaac-tools.jpl.nasa.gov/drive/files/SeaSurfaceTopography/merged_alt/L4/cdr_grid</p> <p>(MLD) Argo-based ocean surface mixed layer depths using the buoyancy gradient definition of Whitt Nicholson and Carranza (2019) processed dataset available at https://doi.org/10.5281/zenodo.4291175.</p> <p>Relevant files:</p> <ol> <li>https://github.com/NCAR/ml-ocean-bl/mloceanbl/preprocess_mld.py and .../preprocess_sss_sst_ssh.py. <ul> <li>Preprocessing code</li> </ul> </li> <li>sss_sst_ssh_anomalies.nc. <ul> <li>Regridded and resampled SSS, SST, SSH onto a 1/2 degree lat/lon 7day grid. Contains preprocessed seasonal data along with anomalies.</li> </ul> </li> <li> mldb_climatology_climatologystd_binned.nc <ul> <li>Smoothed argo-based mixed layer depths are used to calculate climatologies and standardized climatologies. 4 degree lat/lon gridded climatologies.</li> </ul> </li> <li>mldb_full_anomalies_stdanomalies_climatology_stdclimatology.nc <ul> <li>Contains the Argo profile-derived MLD, anomalies, standard anomalies, climatologies, and standardized climatologies with corresponding argo locations, times, and corresponding weeks. </li> </ul> </li> <li>equatorial_pacific_model_oi_re.nc, southern_indian_model_oi_re.nc <ul> <li>Model outputs for the Equatorial Pacific Ocean and Southern Indian Ocean. These gridded files contain the model outputs (vlcnn, vlcnn variance, OI, OI variance, reanalysis, and reanalysis variance - see manuscript for nomenclature details) at each of the 200 weeks available. It should be noted that, in the equatorial Pacific Ocean, the lat/lon location of (-138.75, -9.75) is masked during the training and filled with a NaN in the .nc files. </li> </ul> </li> </ol> <p> </p> <p> </p> <p>Contact D. Foster with any questions.</p> <p> </p>
On the importance of additional behavioral observation in behavioral psychopharmacology research: a case study on agomelatine's effects on feedback sensitivity in probabilistic reversal learning test in rats (data&code)
<p>Dataset to the manuscript entitled: On the importance of additional behavioral observation in behavioral psychopharmacology research: a case study on agomelatine's effects on feedback sensitivity in probabilistic reversal learning test in rats. Code in R used for analysis data from probabilistic reversal learning (PRL) test performed in operant conditioning boxes. Script analyzes data files from MED-PC IV software.</p>
Data from: Stimulus discriminability may bias value-based probabilistic learning
Reinforcement learning tasks are often used to assess participants' tendency to learn more from the positive or more from the negative consequences of one's action. However, this assessment often requires comparison in learning performance across different task conditions, which may differ in the relative salience or discriminability of the stimuli associated with more and less rewarding outcomes, respectively. To address this issue, in a first set of studies, participants were subjected to two versions of a common probabilistic learning task. The two versions differed with respect to the stimulus (Hiragana) characters associated with reward probability. The assignment of character to reward probability was fixed within version but reversed between versions. We found that performance was highly influenced by task version, which could be explained by the relative perceptual discriminability of characters assigned to high or low reward probabilities, as assessed by a separate discrimination experiment. Participants were more reliable in selecting rewarding characters that were more discriminable, leading to differences in learning curves and their sensitivity to reward probability. This difference in experienced reinforcement history was accompanied by performance biases in a test phase assessing ability to learn from positive vs. negative outcomes. In a subsequent large-scale web-based experiment, this impact of task version on learning and test measures was replicated and extended. Collectively, these findings imply a key role for perceptual factors in guiding reward learning and underscore the need to control stimulus discriminability when making inferences about individual differences in reinforcement learning.
SixthSense: Debugging Convergence Problems in Probabilistic Programs via Program Representation Learning
<p>This is a dataset for our paper: "SixthSense: Debugging Convergence Problems in Probabilistic Programs via Program Representation Learning" published at FASE 2022. Find more details at https://github.com/uiuc-arc/sixthsense</p>
Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"
<p>Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"</p> <p>Yuhang Zhang1, Aizhong Ye1*, Phu Nguyen2, Bita Analui2, Soroosh Sorooshian2, Kuolin Hsu2</p> <p>1 State Key Laboratory of Earth Surface Processes and Resource Ecology, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China.</p> <p>2 Center for Hydrometeorology and Remote Sensing, Department of Civil and Environmental Engineering, University of California, Irvine, Irvine, California, CA 92697, USA.</p> <p>## Dataset </p> <p>Streamflow simulations from one observed precipitation (CMA) and three satellite precipitation products (PDIR, IMERG-F, and GSMaP) for 522 sub-basins.</p> <p>- Q-CMA (streamflow reference)<br> - Q-PDIR (uncorrected)<br> - Q-IMERGF (uncorrected)<br> - Q-GSMAP (uncorrected)</p> <p>### Data structure</p> <p>- Head section (row1-row5)<br> - SubNO: 522 <br> - BeginT: 2003-01-01 00:00 <br> - EndT: 2019-12-31 00:00 <br> - Interval: 1440s (daily)<br> - Revise: 10 (scaling factor to keep int datatype)<br> - Point1 Point2 ... (Subbasin No.)<br> - Data section<br> - 6209 rows, 522 cols</p> <p>## Results</p> <p>Two post-processing model results for test period (2015-1-1 to 2018-12-31).</p> <p>### Data structure</p> <p>- 1462 rows, every row denotes each day from 2015-1-1 to 2018-12-31</p> <p>- 100 columns, every column denotes each quantile from 0.005 to 0.995, total 100 quantiles.</p> <p>### qrf-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>### lstm-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p> </p>
Datasets for "Scalable interpolation of satellite altimetry data with probabilistic machine learning"
<p>Elevation (radar freeboard and sea-level anomaly) fields from CryoSat-2, Sentinel-3A, and Sentinel-3B, over the period December 1st 2018 - April 30th 2019. These data were processed for the Arctic domain using the European Space Agency's Grid Processing on Demand (GPOD) service. Processing follows the steps outlined in Lawrence et al., 2021 (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.asr.2019.10.011" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.asr.2019.10.011</a>). These data are provided at along-track, 5 km and 50 km resolution, where gridded data follow the EASE grid definition (<a href="https://doi.org/10.3390/ijgi1010032">https://doi.org/10.3390/ijgi1010032</a>).</p> <p>These data were used to develop the open-source Python programming library GPSat (https://github.com/CPOMUCL/GPSat), which uses local Gaussian Process models to perform scalable interpolation of non-stationary satellite altimetry data. The 'Source_data.xlsx' file contains the data corresponding to figures in the published Nature Communications article 'Scalable interpolation of satellite altimetry data with probabilistic machine learning'.</p>
Data from: Stimulus discriminability may bias value-based probabilistic learning
Open the record for dataset details and reuse information.
Data from: Neural structure mapping in human probabilistic reward learning
Humans can learn abstract concepts that describe invariances over relational patterns in data. One such concept, known as magnitude, allows stimuli to be compactly represented on a single dimension (i.e. on a mental line). Here, we measured representations of magnitude in humans by recording neural signals whilst they viewed symbolic numbers. During a subsequent reward-guided learning task, the neural patterns elicited by novel complex visual images reflected their payout probability in a way that suggested they were encoded onto the same mental number line, with 'bad' bandits sharing neural representation with 'small' numbers and 'good' bandits with 'large' numbers. Using neural network simulations, we provide a mechanistic model that explains our findings and shows how structural alignment can promote transfer learning. Our findings suggest that in humans, learning about reward probability is accompanied by structural alignment of value representations with neural codes for the abstract concept of magnitude.
Probabilistic Emulation of the Community Radiative Transfer Model Using Machine Learning
Open the record for dataset details and reuse information.
FMRI Study of Performance During a Probabilistic Reversal Learning Task in Depression
ClinicalTrials.gov study NCT00075296. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Neural structure mapping in human probabilistic reward learning
Open the record for dataset details and reuse information.
Mechanistic Exploration and Kinetic Modeling through In-Silico Data Generation and Probabilistic Machine Learning Analysis
<p>This zip file includes the dataset 'two_reactions_022624.csv,' which is used for training and testing ML/DL models in the paper 'Mechanistic Exploration and Kinetic Modeling through In-Silico Data Generation and Probabilistic Machine Learning Analysis,' as well as trained models and some files used for training the model. When running the model downloaded from GitHub, copy and paste the files downloaded from here into the subfolder with the same name and path as the one downloaded from GitHub.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.