Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

37 results for “Statistical learning”

Learn how ShareScore rates datasets ↗
dryad36/100

Data for: Intracranial entrainment reveals statistical learning across levels of abstraction

Open the record for dataset details and reuse information.

publicJul 2023View details →
zenodo32/100

Silicon-29 NMR Experimental Datasets used in Statistical Learning of NMR tensors from 2D Isotropic/Anisotropic Correlation Nuclear Magnetic Resonance Spectra

<p>Processed silicon-29 Magic-Angle Flipping and Magic-Angle Turning Nuclear Magnetic Resonance spectra used as input to the smooth-LASSO linear inversion algorithm along with their corresponding NMR tensor parameter distributions as described in the paper &quot;Statistical Learning of NMR tensors from 2D Isotropic/Anisotropic Correlation Nuclear Magnetic Resonance Spectra&quot;, by&nbsp;Srivastava and Grandinetti. &nbsp;</p> <p>Files with names containing &quot;MAF&quot; or &quot;MAT&quot; are the corresponding experimental Si-29 NMR MAF or MAT dataset on the composition given in the filename. &nbsp;Files with names containing &quot;inverse&quot;&nbsp;are the corresponding NMR tensor parameter distributions obtained from the inversion of the corresponding experimental MAF and MAT dataset with&nbsp;the composition given in the filename. &nbsp;</p> <p><br> &nbsp;</p> <p>Details of the csdf dataset format are given in&nbsp;<a href="https://doi.org/10.1371/journal.pone.0225953"><em>PLOS ONE,</em>&nbsp;15(1): e0225953 (2020)</a>, &quot;Core Scientific Dataset Model: A lightweight and portable model and file format for multi-dimensional scientific data,&quot;&nbsp;D. Srivastava, T. Vosegaard, D. Massiot, and P.J. Grandinetti. &nbsp;The data within&nbsp;csdf&nbsp;files can be accessed with the Python package&nbsp;<a href="https://csdmpy.readthedocs.io/en/stable">csdmpy</a>, or other CSDM-compliant software.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
dryad32/100

Calibration of probability predictions from machine-learning and statistical models

<p>This data set describes the occurrence (yes/no) of a bird, the Southern Whiteface (<i>Aphelocephala leucopsis)</i> in Australia. A suite of environmental variables is provided, which are used in the paper to illustrate a statistical problem. The data are meant to allow reproduction of the analysis in this paper. They are not intended for actual ecological analysis. The data come as .Rdata-file, i.e. as an R-dataset (described technically here: https://www.loc.gov/preservation/digital/formats/fdd/fdd000470.shtml).</p> <p>Here is the paper's abstract:</p> <p><span>Aim: Predictions from statistical models may be uncalibrated, meaning that the predicted values do not have the nominal coverage probability. This is easiest seen with probability predictions in machine-learning classification, including the common species occurrence probabilities. Here, a predicted probability of, say, 0.7 should indicate that out of 100 cases with these environmental conditions, and hence the same predicted probability, the species should be present in 70 and absent in 30.</span><br> <span>Innovation: A simple calibration plot shows that this is not necessarily the case, particularly not for over-fitted models or algorithms that use non-likelihood target functions. As a consequence, "raw" predictions from such model could easily be off by 0.2, are unsuitable for averaging across model types, and resulting maps hence be substantially distorted. The solution, a flexible calibration regression, is simple and can be applied whenever deviations are observed.</span><br> <span>Conclusion: "Raw", uncalibrated probability predictions should be calibrated before interpreting or averaging them in a probabilistic way.</span></p>

opencc-zeroJan 2021View details →
zenodo32/100

All data and code from Honeyman, et al. 2022: Statistical learning and uncommon soil microbiota explain biogeochemical responses after wildfire

<p>The zipped folder contains all data and code necessary to reproduce analyses from the manuscript. Note that raw sequencing reads are available at NCBI SRA under&nbsp;BioProject PRJNA767436.</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

DeepCellMap: a Deep Learning Approach Coupled to Spatial Statistics to unravel Microglial Spatial Organization in the Developing Human Brain

<p>DeepCellMap is a deep-learning-assisted tool that integrates multi-scale image processing with advanced spatial and clustering statistics. This pipeline is designed to map microglial organization during normal and pathological brain development but can be adapted to any cell type.} Using DeepCellMap, we can capture the morphological diversity of microglia,&nbsp; identify strong coupling between proliferative and phagocytic phenotypes, and show that distinct spatial clusters rarely overlap as human brain development progresses. Additionally, we uncover a novel association between microglia and blood vessels in fetal brains exposed to maternal SARS-CoV-2. These findings offer insights into whether various microglial phenotypes form networks in the developing brain to occupy space, and in conditions involving haemorrhages, whether microglia respond to, or influence changes in blood vessel integrity. DeepCellMap is available as open-source software and is a powerful tool for extracting spatial statistics and analyzing cellular organization in large tissue sections, accommodating various imaging modalities. </p>

opencc-by-4.0Oct 2024View details →
dryad32/100

Data from: Statistical context dictates the relationship between feedback-related EEG signals and learning

Learning should be adjusted according to the surprise associated with observed outcomes but calibrated according to statistical context. For example, when occasional changepoints are expected, surprising outcomes should be weighted heavily to speed learning. In contrast, when uninformative outliers are expected to occur occasionally, surprising outcomes should be less influential. Here we dissociate surprising outcomes from the degree to which they demand learning using a predictive inference task and computational modeling. We show that the P300, a stimulus-locked electrophysiological response previously associated with adjustments in learning behavior, does so conditionally on the source of surprise. Larger P300 signals predicted greater learning in a changing context, but less learning in a context where surprise indicated a one-off outlier (oddball). Our results suggest that the P300 provides a surprise signal that is interpreted by downstream learning processes differentially according to statistical context in order to appropriately calibrate learning across complex environments.

opencc-zeroAug 2019View details →
zenodo32/100

The machine learning based statistical emulators of GGCMI phase 2

<p>A statistical emulator with machine learning algorithm to reproduce the response of year-to-year variation of four crop yield to CO<sub>2</sub> (C), temperature (T), water (W) and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2 experiment.</p>

opencc-by-4.0Apr 2023View details →
dryad32/100

Data from: Statistical context dictates the relationship between feedback-related EEG signals and learning

Open the record for dataset details and reuse information.

publicNov 2019View details →
dryad32/100

Calibration of probability predictions from machine-learning and statistical models

Open the record for dataset details and reuse information.

publicMar 2020View details →
zenodo28/100

Data from: Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling

<p>As stated in the Read Me file:</p> <p>These data and resources are associated with the manuscript:</p> <p><em>Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling</em>, Mersmann et. al., as submitted to biorXiv in July 2020.</p> <p>The raw imaging data relates to Figure 5, S1, S2 and Table S1. The images are fluorescent micrographs displaying immobilised adenovirus particles bound to a monoclonal antibody 9C12.</p> <p>Each experiment folder is numbered, as in Table S1, and appended with the mixing proportion (Fl), as defined in the manuscript. Within each folder there are 6 subfolders, representing samples incubated with different concentrations of 9C12 antibody.</p> <p>Each image is a 3 channel 1024x1024 tif. Channel 1 = 9C12 Alexa Fluor 647. Channel 2 = 9C12 Biotin + QDot655. Channel 3 = Adenovirus Alexa Fluor 488. &nbsp;Samples were illuminated in TIRF mode using a 100X objective, images were captured on a Hamamatsu OCRA Flash 4 sCMOS camera. Further details are available in the header of each file.</p> <p>The control samples are labelled with 100% 9C12 Alexa Fluor 647 or 100% 9C12 Biotin, as described in the manuscript.</p> <p>The data analysis script is an imageJ macro. It runs on the FIJI version of ImageJ with the NanoJ package installed (https://github.com/HenriquesLab). It outputs fluorescent measurements for each identified AdV particle. Note that the script rearranges the channel order such that Channel 1 = Adenovirus Alexa Fluor 488, Channel 2 = 9C12 Alexa Fluor 647, Channel 3 = 9C12 Biotin + QDot655.&nbsp;</p> <p>The channels require registration due to chromatic aberration, this is achieved using the Realign Channels function in NanoJ, appropriate translation masks are provided along with the script.</p> <p>Any question about the data or script should be addressed in Joe Grove (j.grove@ucl.ac.uk)</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Predicting permeability via statistical learning on higher-order microstructural information

<p>Dataset and code used in M. R&ouml;ding, et al, &quot;Predicting permeability via statistical learning on higher-order microstructural information&quot;, published in Scientific Reports, 2020. In this work, we study permeability prediction in a large data set of 30,000 virtual, porous microstructures of different types, including both granular and continuous solid phases. The permeabilities are computed using the lattice Boltzmann method. The pore space geometries are charaterized using the following descriptors: one-point correlation functions (porosity, specific surface), two-point surface-surface, surface-void, and void-void correlation functions, and geodesic tortuosity. Linear regression with linear and quadratic terms as well articifical neural networks are used for prediction. As a reference, Kozeny-Carman regression with only lowest-order descriptors (porosity and specific surface) is also studied. Herein, the descriptors, the permeabilities, and the Matlab and Python/Tensorflow code used for prediction are supplied.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Large-scale statistical learning for mass transport prediction in porous materials using 90,000 artificially generated microstructures

<p>Dataset and code used in B Prifling, et al, &quot;Large-scale statistical learning for mass transport prediction in porous materials using 90,000 artificially generated microstructures&quot;, published in Frontiers in Materials. In this work, we investigate relationships between 3D microstructure and effective diffusivity and permeability, based on a dataset of 90,000 structures and using analytical formulas, artificial neural networks (ANNs), and convolutional neural networks (CNNs). Herein, the codes in Matlab and Python/Tensorflow necessary to investigate the prediction models and reproduce the results of the paper are supplied. Also, microstructures together with their computed geometrical descriptors and effective properties are included.</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Forecasting Cryptocurrency Markets: Predictive Modelling Using Statistical and Machine Learning Approaches

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
geo24/100

Statistical learning quantifies transposable element-mediated cis-regulation

GEO Series GSE208403. Homo sapiens. 9 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenOct 2023View details →
zenodo24/100

Results of experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks

<p>Results of the experiments&nbsp;in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks</p>

opencc-by-4.0Jun 2020View details →
ClinicalTrials.gov24/100

Statistical Learning as a Predictor of Attention Bias Modification Outcome

ClinicalTrials.gov study NCT03424967. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Statistical Learning as a Novel Intervention for Cortical Blindness

ClinicalTrials.gov study NCT06578117. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record