Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

37 results for “Statistical learning”

Learn how ShareScore rates datasets ↗
OpenNeuro52/100

OLVSL_ Object-location visual statistical learning

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo52/100

SQLite database to accompany the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring"

<p>This dataset is a SQLite database that accompanies methods and analysis described in the paper, &quot;Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring&quot; (Balantic &amp; Donovan 2019, Bioacoustics, https://www.tandfonline.com/doi/full/10.1080/09524622.2019.1605309).&nbsp;</p> <p>A Github repository containing code for using the SQLite&nbsp;database also accompanies this paper at:&nbsp;<a href="https://github.com/cbalantic/false-positive-mitigation">http://github.com/cbalantic/false-positive-mitigation</a></p>

opencc-by-4.0May 2019View details →
zenodo44/100

MEG dataset nonlinguistic auditory statistical learning

<p>MEG data of 24&nbsp;healthy adults with an auditory nonlinguistic statistical learning paradigm plus data from two subsequent behavioral tasks. For closer description of data see data description file.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Code and data to "Statistical learning and topkriging improve spatio-temporal low-flow estimation"

<p>This data and software supports the manuscript "Statistical learning and topkriging improve spatio-temporal low-flow estimation" (https:://doi.org/<span>10.1029/2024WR038329</span>).</p> <p>The dataset consists of:</p> <ul> <li>all produced predictions of the models (data/predictions.RDS and data/predictions_csv/*)</li> <li>observational data (data/observations.csv)</li> <li>additional catchment data (data/catchment_data.csv) used for presenting the figures</li> <li>state boundaries of Austria as a shape file (data/boundaries.*)</li> <li>partial predictions of a model-based boosting approach (data/partial_predictions.csv)</li> <li>Example output of number of EOF, due to long computational time (data/number_eofs.RDS)</li> <li>IDs of near natural catchments (data/ids_low_flow.csv)</li> </ul> <p>Additionally, the code is provided to:</p> <ul> <li>Compute the number of EOFs (functions/number_eofs.R)</li> <li>Produce all the figures and tables in the paper (scripts/plotting_results.R)</li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Statistical analysis and dataset for: Invasive ant learning is not affected by seven potential neuroactive chemicals

<p>Linked to the journal article&nbsp;published in Current Zoology (<a href="https://doi.org/10.1093/cz/zoad001">https://doi.org/10.1093/cz/zoad001</a>).</p> <p><em><strong>Abstract</strong></em></p> <p>Argentine ants (<em>Linepithema humile</em>) are one of the most damaging invasive alien species worldwide. Enhancing or disrupting cognitive abilities, such as learning, has the potential to improve management efforts, for example by increasing preference for a bait, or improving ants&rsquo; ability to learn its characteristics or location. Nectar-feeding insects are often the victims of psychoactive manipulation, with plants lacing their nectar with secondary metabolites such as alkaloids and non-protein amino acids which often alter learning, foraging, or recruitment. However, the effect of neuroactive chemicals has seldomly been explored in ants. Here, we test the effects of seven potential neuroactive chemicals - two alkaloids: caffeine and nicotine; two biogenic amines: dopamine and octopamine, and three non-protein amino acids: &beta;-alanine, GABA and taurine - on the cognitive abilities of invasive&nbsp;<em>L. humile</em>&nbsp;using bifurcation mazes. Our results confirm that these ants are strong associative learners, requiring as little as one experience to develop an association. However, we show no short-term effect of any of the chemicals tested on spatial learning, and in addition no effect of caffeine on short-term olfactory learning. This lack of effect is surprising, given the extensive reports of the tested chemicals affecting learning and foraging in bees. This mismatch could be due to the heavy bias towards bees in the literature, a positive result&nbsp;publication bias, or differences in methodology.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Predictors and predictand for "Repeatable high-resolution statistical downscaling through deep learning"

<p>Predictors and predictand for &quot;Repeatable high-resolution statistical downscaling through deep learning&quot;. Predictors from the ERA5 reanalysis and predictand from ReKIS (https://rekis.hydro.tu-dresden.de). Data is saved in &quot;.rda&quot; format, to be read from R.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Hyperparameter tuning and performance assessment of statistical and machine-learning models using spatial data.

<p>This is a research compendium (RC) for the publication &quot;Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data&quot;.</p> <p>The code (including figures, appendices and the manuscript) is packed in <strong>pathogen-modeling-3.zip&nbsp;</strong>or can be found directly in the <a href="https://github.com/pat-s/pathogen-modeling">Github repository</a>.</p> <ul> <li><strong>Publication figures</strong>:&nbsp;analysis/paper/submission/3/latex-source-files/</li> <li><strong>Appendices</strong>: analysis/paper/submission/3/</li> </ul> <p>This RC represents a static snapshot at the time of submission. The Github repository will receive changes after the publication was published.</p> <p><strong>Data sources</strong></p> <ul> <li>Atlas Climatico:&nbsp;<a href="http://opengis.uab.es/wms/iberia/index.htm">http://opengis.uab.es/wms/iberia/index.htm</a></li> <li>DEM:&nbsp;ftp://ftp.geo.euskadi.eus/lidar/MDE_LIDAR_2016_ETRS89/</li> <li>Lithology:&nbsp;<a href="http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home">http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home</a></li> <li>pH:&nbsp;<a href="https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0">https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0</a></li> <li>soil:&nbsp;<a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a></li> </ul> <p><strong>Licenses</strong></p> <p>All files are shared via the given license with the exception of &quot;soil.tif&quot; which is shared via the&nbsp;<strong>ODbL </strong>license<strong>.</strong></p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

DIETxPOSOME - Summary statistics from papers obtained from literature mining and machine learning protocols

<p>DIETxPOSOME database concerning literature selection of potentially useful papers retrieved from PubMed search API, concerning contaminants quantification in food items of worldwide highest supply and using FoodMine code (text matching filter) and machine learning (ML) protocols.&nbsp;11,723 data points were collected from 254 papers from the last two decades&nbsp;in 72 foods to obtain relevant information on 96 contaminants, including heavy metals, polychlorinated biphenyls, dioxins, furans, polycyclic aromatic hydrocarbons (PAHs), pesticides, mycotoxins, and heterocyclic aromatic amines (HAAs).</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Statistics and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"

<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This publication contains aggregated data for the paper. It also contains the evaluation data of all model/hyper-parameter training and test runs.</p>

opencc-by-4.0Oct 2019View details →
dryad36/100

Data from: Statistically testing the role of individual learning and decision-making in trapline foraging

Trapline foraging, a behavior consisting of repeated visitation to spatially fixed resources in a predictable sequence, has been observed over diverse taxa and is important ecologically for efficient resource gathering. Despite this, few null models exist to test the significance of suspected traplines, particularly for studies interested in the role of individual decision-making in the formation of traplines versus the role of resource layouts and random movement patterns. Here we present a spatially explicit, individual-based null model, which may be used to test whether resource layout and realistic forager movement may account for sequence repeats in suspected traplines. In our model, we generate resource visitation sequences by modeling a forager without spatial memory using a random walk to discover and visit spatially-fixed resources. We quantify traplining using Determinism, a metric derived from recurrence quantification analysis. Using both simulated and empirical bee foraging data, we compared our model with two existing null models—a completely random model and a sample randomization model. The former creates null sequences by randomly selecting available resources, while the latter randomizes the order of visits in observed sequences. We found that our model has a higher propensity of being (correctly) rejected than a sample randomization model for trapliners, and a lower propensity of being (incorrectly) rejected for non-trapliners compared to a completely random model. The use of a spatially explicit individual-based null model to test the statistical significance of patterns in empirical data is a novel approach that may be useful for other spatial and individual-based processes.

opencc-zeroDec 2017View details →
zenodo36/100

Biological and environmental data for a study on transferability of statistical and machine learning models using North Sea Macrozoobenthos

<p>General</p> <p>Data documented here are not the product of our research but was scraped from various sources and processed - so no genuine reupload. This collection is a contribution to reproduceable reseach. All datasets are given in "RData" binary format</p> <p> </p> <p>Data description</p> <p>majornorthseabenthos </p> <p>This is macrozoobenthos data as data frame scraped from the GBIF repository (gbif.org). Species are  Corbula gibba, Tellina fabula, Turritella communis, Euspira pulchella, Corystes cassive- launus, Upogebia deltaura, Lanice conchilega, Nephtys hombergii, Echinocardium cordatum, and Amphiura filiformis. data was postprocessed to have only single occurrence fon the approxinatel 1x1 km grid used for this study. Also, occurrences closer than 5 km  close to shore were removed - including occurrences on land.</p> <p> </p> <p>Predictors</p> <p>A SpatialPixelsDataFrame in EPSG 4326 with five layers: Median grain size in micrometers, mud content in percent (both MUDAB database), water depth in meters above MSL (Weatherall et al, 2015), modelled average bottom shear stress from waves in N/sqrm (The Wamdi Group, 1988) and climatologival average winter bottom water temperature in deg. C (Stips et al, 2004).</p> <p> </p> <p> </p> <p>References</p> <p>Stips A, Bolding K, Pohlmann T, Burchard H (2004) Simulating the temporal and spatial dy- namics of the North Sea using the new model GETM (general estuarine transport model). Ocean Dynamics 54(2):266–283</p> <p>The Wamdi Group (1988) The WAM model-a third generation ocean wave prediction model. Journal of Physical Oceanography 18(12):1775–1810</p> <p>Weatherall P, Marks K, Jakobsson M, Schmitt T, Tani S, Arndt JE, Rovere M, Chayes D, Ferrini V, Wigley R (2015) A new digital bathymetric model of the world’s oceans. Earth and Space Science 2(8):331–345</p> <p> </p>

opencc-by-4.0Apr 2017View details →
zenodo36/100

Undergraduate Students' Perceptions and Experiences in Learning Statistics

<p>This dataset is an online survey participated voluntarily by undergraduate students in 2021 at the University of Hargeisa. The main variables included were sociodemographic variables, students' perceptions of statistics, challenges faced by students in learning statistics, engagement with statistics by students, and finally (but not least) students' performance in statistics.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Statistical and machine learning methods for evaluating trends in air quality under changing meteorological conditions

<p>This repo includes the GEOS-Chem simulations and R scripts that are needed to replicate and evaluate the conclusions from&nbsp;Qiu, Zigler, and Selin, ACP, 2022 &quot;Statistical and machine learning methods for evaluating trends in air quality under changing meteorological conditions&quot;.</p> <p><strong>The GEOS-Chem simulations</strong></p> <ul> <li>For the US (2011-2017): <ul> <li><em>observational_o3_pm_2011_2017_us.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the observational scenarios (<strong>changing</strong> meteorology, <strong>changing </strong>emissions).</li> <li><em>counterfactual_o3_pm_2011_2017_us.rds</em><em>:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the counterfactual scenarios (<strong>constant</strong>&nbsp;meteorology, <strong>changing</strong> emissions).</li> <li><em>constant_emis_o3_pm_2012_2017_us.rds:&nbsp;</em>the simulated daily PM2.5 and O3 concentrations in the constant-emission scenarios (<strong>constant</strong>&nbsp;meteorology, <strong>constant</strong>&nbsp;emissions).</li> <li><em>regional_features_2011_2017_4x5_us.rds:&nbsp;</em>the&nbsp;MERRA-2 meteorological features in the observational scenarios (aggregated to 4x5 degrees), inputs&nbsp;for the &quot;RF-regional&quot; model.</li> </ul> </li> <li>For China&nbsp;(2013-2017): <ul> <li><em>observational_o3_pm_2013_2017_china.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the observational scenarios (<strong>changing</strong> meteorology, <strong>changing </strong>emissions).</li> <li><em>counterfactual_o3_pm_2013_2017_china.rds:</em> the simulated daily PM2.5 and O3 concentrations, and MERRA-2 meteorological features in the counterfactual scenarios (<strong>constant</strong>&nbsp;meteorology, <strong>changing</strong> emissions).</li> <li><em>constant_emis_o3_pm_2014_2017_china.rds</em><em>:&nbsp;</em>the simulated daily PM2.5 and O3 concentrations in the constant-emission scenarios (<strong>constant</strong>&nbsp;meteorology, <strong>constant</strong>&nbsp;emissions).</li> <li><em>regional_features_2013_2017_4x5_china.rds:&nbsp;</em>the&nbsp;MERRA-2 meteorological features in the observational scenarios (aggregated to 4x5 degrees), inputs for the &quot;RF-regional&quot; model.</li> </ul> </li> </ul> <p><strong>R scripts:</strong></p> <ul> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/main.r">main.r</a>: the main script to perform statistical correction of meteorological variability.</li> <li>main.r uses functions from&nbsp;the other R script files (see below)&nbsp;which perform different&nbsp;statistical correction methods, respectively.&nbsp;&nbsp;</li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/parametric_regression_methods.r">parametric_regression_methods.r</a>: performs meteorological correction with parametric regression methods (MLR, polynomial, spline, GAM)</li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/tune_RF_regional.r">tune_RF_regional.r</a>&nbsp;and&nbsp;<a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/RF_regional.r">RF_regional.r</a>: perform&nbsp;the&nbsp;meteorological correction with the &quot;RF-regional&quot; model</li> <li><a href="https://zenodo.org/api/files/065be469-ef8d-4c9b-9bd8-a6a808275237/GEOS_Chem_constant_emis.r">GEOS_Chem_constant_emis.r</a>: performs the&nbsp;meteorological correction using the simulations from the constant emission scenarios from the GEOS-Chem model</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Statistical learning shapes pain perception and prediction independently of external cues

<h1>Dataset and code for the relevant analysis and results:</h1> <h3>"Statistical learning shapes pain perception and prediction independently of external cues"</h3> <p>Onysk, J., Whitefield, M., Gregory, N., Jain, M., Turner, G., Seymour, B., Mancini, F. (2024). eLife. <a href="https://doi.org/10.7554/eLife.90634.2">https://doi.org/10.7554/eLife.90634.2</a></p> <h2>1 - data_collection</h2> <p>Contains the code for the psychophysical experiment (PsychToolBox), including the sequence generations scripts.</p> <h2>2 - preprocessing</h2> <p>Contains code that preprocesses behavioural data from PsychToolBox. This includes linear transformation of inputs, exporting data to stan readable format and plotting Supplement figures.</p> <p>- The raw behavioural data can be found in <strong><em>preprocessing/data</em></strong>. Stan ready ready for each condition is found in <strong><em>preprocessing/stan_data</em></strong></p> <h2>3 - model_fit_analysis</h2> <p>Contains stan models used in the paper ('models/'), model fitting code ('fit_models_cs.R) (inlcuding HPC setup in 'hpc/'), initial analysis script for processing stan samples ('primary_analysis_cs.R'), as well as additional analyis scripts ('extra_analysis_cs.R', 'correlation_beh_model_cs.ipynb') that generate figures from the paper and supplement.</p> <p>- The posterior draws for parameters can be found in <em><strong>model_fit_analysis/output/cs_results</strong></em>&nbsp;</p> <h2>4 - model_recovery</h2> <p>Contains code that execute model and parameter recovery analysis ('mp_recovery.R'), including HPC setup ('hpc/'). The 'mp_rec_analyse.R' reproduces model and parameter recovery results from the supplement.</p> <h2>5 - Figures</h2> <p>Contains all the figures from the manuscript and the supplement.</p> <h2>6 - RDS_fits</h2> <p>Contains RStan fit objects for each condition for each model</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Does evaluative learning depend on the statistical relationship between stimuli? On the sensitivity of evaluative cue conditioning to ecological contingencies

<p>Evaluative conditioning (EC) is concerned with the transfer of valence from an unconditioned stimulus (US) to a conditioned stimulus (CS). Only recently the notion of EC was extended from individual CSs to categories of CSs that share certain cues. That research shows that a contingency implemented between a cue dimension and US valence has a direct effect on the evaluation of stimuli carrying values of this cue dimension. This phenomenon was coined &ldquo;evaluative cue conditioning&rdquo; (ECC). The present research tests whether ECC is sensitive to the contingency between a cue dimension and US valence. The present work thereby investigates the impact of ecological contingencies that define contingency by reference to all other CS-US pairings in the learning environment. Two experiments demonstrate that ECC effects are sensitive to the strength of the ecological contingency, suggesting that ECC is a relative phenomenon that depends on the valence of the CS-US pairings in the reference set. Moreover, these findings call for more research on the contingency sensitivity of standard EC effects using this revised definition of contingency.</p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

The Impact of Generative AI on Student Learning Outcomes: A Statistical Analytical Approach - Dataset

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
dryad36/100

Data for: Intracranial entrainment reveals statistical learning across levels of abstraction

<div class="t-landing__text-wall "> <p>The following submission contains the data reported from the manuscript "Intracranial entrainment reveals statistical learning across levels of abstraction". The dataset was obtained from 8 neurosurgical patients who had intracranially implanted electrodes for seizure monitoring. Intracranial EEG (iEEG) data were recorded while the patients viewed a rapid stream of scene images. In the Category-Level Structured condition, patients viewed a series of trial-unique scene images, in which the categories of scenes were paired across repetitions (e.g., images of beach always followed by images of canyons). In the Exemplar-level Structured condition, participants viewed a sequence of 6 repeating scene images (from non-overlapping scene categories), which were paired across repetitions (e.g., image A always followed by image B). In the baseline Random condition, participants again viewed a sequence of 6 repeating scene images, but the images were presented in a random temporal order. Each image was presented for 250 ms, followed by a 250 ms inter-stimulus-interval period.</p> <p>The data presented here contain the raw iEEG data from the 8 patients during these task conditions, as well as information about the anatomical placement of their electrodes.</p> </div>

opencc-zeroJun 2023View details →
ClinicalTrials.gov36/100

Pattern Recognition and Anomaly Detection in Fetal Morphology Using Deep Learning and Statistical Learning

ClinicalTrials.gov study NCT05738954. IPD Sharing: YES. Countries: 1. Publications: 4.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Predicting Language and Literacy Growth in Children With ASD Using Statistical Learning

ClinicalTrials.gov study NCT06332144. IPD Sharing: YES. Countries: 1. Publications: 4.

controlledIPD-YESFeb 2026View details →
dryad36/100

Data from: Statistically testing the role of individual learning and decision-making in trapline foraging

Open the record for dataset details and reuse information.

publicMar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record