Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
37
datasets available to search
ShareScore release 0.7.1
Dataset results
37 results for “Statistical learning”
Data for: Intracranial entrainment reveals statistical learning across levels of abstraction
Open the record for dataset details and reuse information.
Silicon-29 NMR Experimental Datasets used in Statistical Learning of NMR tensors from 2D Isotropic/Anisotropic Correlation Nuclear Magnetic Resonance Spectra
<p>Processed silicon-29 Magic-Angle Flipping and Magic-Angle Turning Nuclear Magnetic Resonance spectra used as input to the smooth-LASSO linear inversion algorithm along with their corresponding NMR tensor parameter distributions as described in the paper "Statistical Learning of NMR tensors from 2D Isotropic/Anisotropic Correlation Nuclear Magnetic Resonance Spectra", by Srivastava and Grandinetti. </p> <p>Files with names containing "MAF" or "MAT" are the corresponding experimental Si-29 NMR MAF or MAT dataset on the composition given in the filename. Files with names containing "inverse" are the corresponding NMR tensor parameter distributions obtained from the inversion of the corresponding experimental MAF and MAT dataset with the composition given in the filename. </p> <p><br> </p> <p>Details of the csdf dataset format are given in <a href="https://doi.org/10.1371/journal.pone.0225953"><em>PLOS ONE,</em> 15(1): e0225953 (2020)</a>, "Core Scientific Dataset Model: A lightweight and portable model and file format for multi-dimensional scientific data," D. Srivastava, T. Vosegaard, D. Massiot, and P.J. Grandinetti. The data within csdf files can be accessed with the Python package <a href="https://csdmpy.readthedocs.io/en/stable">csdmpy</a>, or other CSDM-compliant software.</p> <p> </p>
Calibration of probability predictions from machine-learning and statistical models
<p>This data set describes the occurrence (yes/no) of a bird, the Southern Whiteface (<i>Aphelocephala leucopsis)</i> in Australia. A suite of environmental variables is provided, which are used in the paper to illustrate a statistical problem. The data are meant to allow reproduction of the analysis in this paper. They are not intended for actual ecological analysis. The data come as .Rdata-file, i.e. as an R-dataset (described technically here: https://www.loc.gov/preservation/digital/formats/fdd/fdd000470.shtml).</p> <p>Here is the paper's abstract:</p> <p><span>Aim: Predictions from statistical models may be uncalibrated, meaning that the predicted values do not have the nominal coverage probability. This is easiest seen with probability predictions in machine-learning classification, including the common species occurrence probabilities. Here, a predicted probability of, say, 0.7 should indicate that out of 100 cases with these environmental conditions, and hence the same predicted probability, the species should be present in 70 and absent in 30.</span><br> <span>Innovation: A simple calibration plot shows that this is not necessarily the case, particularly not for over-fitted models or algorithms that use non-likelihood target functions. As a consequence, "raw" predictions from such model could easily be off by 0.2, are unsuitable for averaging across model types, and resulting maps hence be substantially distorted. The solution, a flexible calibration regression, is simple and can be applied whenever deviations are observed.</span><br> <span>Conclusion: "Raw", uncalibrated probability predictions should be calibrated before interpreting or averaging them in a probabilistic way.</span></p>
All data and code from Honeyman, et al. 2022: Statistical learning and uncommon soil microbiota explain biogeochemical responses after wildfire
<p>The zipped folder contains all data and code necessary to reproduce analyses from the manuscript. Note that raw sequencing reads are available at NCBI SRA under BioProject PRJNA767436.</p>
DeepCellMap: a Deep Learning Approach Coupled to Spatial Statistics to unravel Microglial Spatial Organization in the Developing Human Brain
<p>DeepCellMap is a deep-learning-assisted tool that integrates multi-scale image processing with advanced spatial and clustering statistics. This pipeline is designed to map microglial organization during normal and pathological brain development but can be adapted to any cell type.} Using DeepCellMap, we can capture the morphological diversity of microglia, identify strong coupling between proliferative and phagocytic phenotypes, and show that distinct spatial clusters rarely overlap as human brain development progresses. Additionally, we uncover a novel association between microglia and blood vessels in fetal brains exposed to maternal SARS-CoV-2. These findings offer insights into whether various microglial phenotypes form networks in the developing brain to occupy space, and in conditions involving haemorrhages, whether microglia respond to, or influence changes in blood vessel integrity. DeepCellMap is available as open-source software and is a powerful tool for extracting spatial statistics and analyzing cellular organization in large tissue sections, accommodating various imaging modalities. </p>
Data from: Statistical context dictates the relationship between feedback-related EEG signals and learning
Learning should be adjusted according to the surprise associated with observed outcomes but calibrated according to statistical context. For example, when occasional changepoints are expected, surprising outcomes should be weighted heavily to speed learning. In contrast, when uninformative outliers are expected to occur occasionally, surprising outcomes should be less influential. Here we dissociate surprising outcomes from the degree to which they demand learning using a predictive inference task and computational modeling. We show that the P300, a stimulus-locked electrophysiological response previously associated with adjustments in learning behavior, does so conditionally on the source of surprise. Larger P300 signals predicted greater learning in a changing context, but less learning in a context where surprise indicated a one-off outlier (oddball). Our results suggest that the P300 provides a surprise signal that is interpreted by downstream learning processes differentially according to statistical context in order to appropriately calibrate learning across complex environments.
The machine learning based statistical emulators of GGCMI phase 2
<p>A statistical emulator with machine learning algorithm to reproduce the response of year-to-year variation of four crop yield to CO<sub>2</sub> (C), temperature (T), water (W) and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2 experiment.</p>
Data from: Statistical context dictates the relationship between feedback-related EEG signals and learning
Open the record for dataset details and reuse information.
Calibration of probability predictions from machine-learning and statistical models
Open the record for dataset details and reuse information.
Data from: Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling
<p>As stated in the Read Me file:</p> <p>These data and resources are associated with the manuscript:</p> <p><em>Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling</em>, Mersmann et. al., as submitted to biorXiv in July 2020.</p> <p>The raw imaging data relates to Figure 5, S1, S2 and Table S1. The images are fluorescent micrographs displaying immobilised adenovirus particles bound to a monoclonal antibody 9C12.</p> <p>Each experiment folder is numbered, as in Table S1, and appended with the mixing proportion (Fl), as defined in the manuscript. Within each folder there are 6 subfolders, representing samples incubated with different concentrations of 9C12 antibody.</p> <p>Each image is a 3 channel 1024x1024 tif. Channel 1 = 9C12 Alexa Fluor 647. Channel 2 = 9C12 Biotin + QDot655. Channel 3 = Adenovirus Alexa Fluor 488. Samples were illuminated in TIRF mode using a 100X objective, images were captured on a Hamamatsu OCRA Flash 4 sCMOS camera. Further details are available in the header of each file.</p> <p>The control samples are labelled with 100% 9C12 Alexa Fluor 647 or 100% 9C12 Biotin, as described in the manuscript.</p> <p>The data analysis script is an imageJ macro. It runs on the FIJI version of ImageJ with the NanoJ package installed (https://github.com/HenriquesLab). It outputs fluorescent measurements for each identified AdV particle. Note that the script rearranges the channel order such that Channel 1 = Adenovirus Alexa Fluor 488, Channel 2 = 9C12 Alexa Fluor 647, Channel 3 = 9C12 Biotin + QDot655. </p> <p>The channels require registration due to chromatic aberration, this is achieved using the Realign Channels function in NanoJ, appropriate translation masks are provided along with the script.</p> <p>Any question about the data or script should be addressed in Joe Grove (j.grove@ucl.ac.uk)</p>
Predicting permeability via statistical learning on higher-order microstructural information
<p>Dataset and code used in M. Röding, et al, "Predicting permeability via statistical learning on higher-order microstructural information", published in Scientific Reports, 2020. In this work, we study permeability prediction in a large data set of 30,000 virtual, porous microstructures of different types, including both granular and continuous solid phases. The permeabilities are computed using the lattice Boltzmann method. The pore space geometries are charaterized using the following descriptors: one-point correlation functions (porosity, specific surface), two-point surface-surface, surface-void, and void-void correlation functions, and geodesic tortuosity. Linear regression with linear and quadratic terms as well articifical neural networks are used for prediction. As a reference, Kozeny-Carman regression with only lowest-order descriptors (porosity and specific surface) is also studied. Herein, the descriptors, the permeabilities, and the Matlab and Python/Tensorflow code used for prediction are supplied.</p>
Large-scale statistical learning for mass transport prediction in porous materials using 90,000 artificially generated microstructures
<p>Dataset and code used in B Prifling, et al, "Large-scale statistical learning for mass transport prediction in porous materials using 90,000 artificially generated microstructures", published in Frontiers in Materials. In this work, we investigate relationships between 3D microstructure and effective diffusivity and permeability, based on a dataset of 90,000 structures and using analytical formulas, artificial neural networks (ANNs), and convolutional neural networks (CNNs). Herein, the codes in Matlab and Python/Tensorflow necessary to investigate the prediction models and reproduce the results of the paper are supplied. Also, microstructures together with their computed geometrical descriptors and effective properties are included.</p>
Forecasting Cryptocurrency Markets: Predictive Modelling Using Statistical and Machine Learning Approaches
Open the record for dataset details and reuse information.
Statistical learning quantifies transposable element-mediated cis-regulation
GEO Series GSE208403. Homo sapiens. 9 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Results of experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks
<p>Results of the experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks</p>
Statistical Learning as a Predictor of Attention Bias Modification Outcome
ClinicalTrials.gov study NCT03424967. IPD Sharing: NO. Countries: 1. Publications: 0.
Statistical Learning as a Novel Intervention for Cortical Blindness
ClinicalTrials.gov study NCT06578117. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.