Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

317

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

317 results for “model selection”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 5 in A Review Of Major Impact Factors Of Hostilities Influencing Biodiversity In The Eastern Ukraine (Modeled On Selected Animal Species)

Fig. 5. Distribution of two snake species, H. caspius and E. dione, in Ukrainian East (ATO zone is indicated by dotted line, burnt area marked inside zone).

opencc-by-4.0Mar 2015View details →
zenodo40/100

Fig. 3 in A Review Of Major Impact Factors Of Hostilities Influencing Biodiversity In The Eastern Ukraine (Modeled On Selected Animal Species)

Fig. 3. Spatial local distribution of ignitions in 2010–2014 in the outskirts of Slavyanoserbsk, Luhansk Region.

opencc-by-4.0Mar 2015View details →
zenodo40/100

Global and tropical band averages for a selection of CMIP5 and CMIP6 models: piControl and abrupt-4xCO2 experiments (compressed)

<p>This dataset provides post-processed spatial averages for a selection of 55 CMIP5 and CMIP6 models. The experiments contained in this dataset are only the pre-industrial controls (piControl) and the experiments with a four-fold increase in the atmospheric CO$_{2}$ concentration in relation to the pre-industrial level (abrupt-4xCO2). The spatial averages are&nbsp;global and tropical bands from x&deg;S to x&deg;N, where the x value is between 5&nbsp;and 40 in increments of 5&deg;. This dataset was created to study climate sensitivity in general and&nbsp;the effect of stratospheric circulation changes on&nbsp;the tropical&nbsp;equilibrium climate sensitivity. It contains the following variables:</p> <ul> <li>incoming (d) short-wave (SW, s) radiative flux (RF, r) at the top of the atmosphere (TOA, t): rsdt</li> <li>outgoing (u) SW&nbsp;RF&nbsp;at TOA: rsut</li> <li>outgoing long-wave (LW, l) RF&nbsp;at TOA: rlut</li> <li>net (n) RF at TOA: rnt*</li> <li>incoming SW RF at the surface (s): rsds</li> <li>outgoing SW RF at&nbsp;the surface: rsus</li> <li>incoming LW RF at the surface: rlds</li> <li>outgoing LW RF at&nbsp;the surface: rlus</li> <li>net RF at the surface: rns*</li> <li>sensible heat flux (hfs) at the surface: hfss</li> <li>latent heat flux (hfl) at the surface: hfls</li> <li>surface temperature (t): ts</li> <li>atmospheric temperature: ta</li> <li>specific humidity: hus</li> <li>zonal component of&nbsp;wind: ua</li> <li>meridional component of wind: va</li> <li>lagrangian tendency of pressure (vertical&nbsp;component of wind in pressure per time&nbsp;dimensions): wap</li> <li>surface pressure: ps</li> <li>geopotential height: zg</li> </ul> <p>*rnt and rns were calculated for the creation of this dataset using the model output&nbsp;rlut rsut, rsds, rsus, rlds, rlus.</p> <p>This version updates the previous version by providing the dataset in a compressed tarball. Use `tar -xzvf CMIP_data.tar.gz` to extract and decompress.</p>

opencc-by-4.0May 2022View details →
dryad40/100

Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models

<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model.  We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate.  We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>

opencc-zeroJun 2022View details →
zenodo40/100

Bulk and critical material demand for selected 'Starter Kit' energy system models - dataset

<p>This repository contains the data related to the Data in Brief article titled: <strong>Bulk and critical material demand for selected &lsquo;Starter Kit&rsquo; energy system models.</strong></p> <p>The data include the modeled mass of materials and their embodied emissions. A metadata file is also included to clarify the units, materials and scenario names.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria

<p>Cyanobacteria are the only prokaryotes to have evolved oxygenic photosynthesis paving the way for complex life. Studying the evolution and ecological niche of cyanobacteria and their ancestors is crucial for understanding the intricate dynamics of biosphere evolution. These organisms frequently deal with environmental stressors such as salinity and drought, and they employ compatible solutes as a mechanism to cope with these challenges. Compatible solutes are small molecules that help maintain cellular osmotic balance in high-salinity environments, such as marine waters. Their production plays a crucial role in salt tolerance, which, in turn, influences habitat preference. Among the five known compatible solutes produced by cyanobacteria (sucrose, trehalose, glucosylglycerol, glucosylglycerate, and glycine betaine), their synthesis varies between individual strains. In this study, we work in a Bayesian stochastic mapping framework, integrating multiple sources of information about compatible solute biosynthesis in order to predict the ancestral habitat preference of Cyanobacteria. Through extensive model selection analyses and statistical tests for correlation, we identify glucosylglycerol and glucosylglycerate as the most significantly correlated with habitat preference, while trehalose exhibits the weakest correlation. Additionally, glucosylglycerol, glucosylglycerate, and glycine betaine show high loss/gain rate ratios, indicating their potential role in adaptability, while sucrose and trehalose are less likely to be lost due to their additional cellular functions. Contrary to previous findings, our analyses predict that the last common ancestor of Cyanobacteria (living at around 3180 Ma) had a 97% probability of a high salinity habitat preference and was likely able to synthesize glucosylglycerol and glucosylglycerate. Nevertheless, cyanobacteria likely colonized low-salinity environments shortly after their origin, with an 89% probability of the first cyanobacterium with low-salinity habitat preference arising prior to the Great Oxygenation Event (2460 Ma). Stochastic mapping analyses provide evidence of cyanobacteria inhabiting early marine habitats, aiding in the interpretation of the geological record. Our age estimate of ~2590 Ma for the divergence of two major cyanobacterial clades (Macro- and Microcyanobacteria) suggests that these were likely significant contributors to primary productivity in marine habitats in the lead-up to the Great Oxygenation Event, and thus played a pivotal role in triggering the sudden increase in atmospheric oxygen.</p>

opencc-zeroMay 2024View details →
zenodo40/100

The WRF model output for selected foehn events at the northern foreland of the Moravian-Silesian Beskids, Czech Republic

<p>This dataset includes three output files from the Weather Research and Forecasting (WRF) model in the NetCDF format. The files are valid for three selected foehn events which affected the northern foreland of the Moravian-Silesian Beskids, Czech Republic and allows a deeper analysis of mechanism and impacts of these events.</p> <p>The dataset contains the model output for the following events:</p> <p>1) 14 January 2008, 10:00 UTC</p> <p>2) 31 October 2010, 06:00 UTC</p> <p>3) 30 October 2021, 06:00 UTC</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Data and code for FishMIP global marine ecosystem model ensemble projections summarised by countries and territories and other selected marine spatial regions.

<p>R code to extract and create data tables and summary plots of FishMIP mean ensemble projections provided are for percentage change in "exploitable fish biomass", which is a proxy for the biomass available to fisheries, consisting of marine animals spanning the size range 10 g to 100 kg: this is typically dominated by fish, but is also inclusive of other animals such as crustaceans and cephalopods.</p> <p>This release contains scripts and summary data for producing figures in Part A of the following report:</p> <p>Blanchard, J.L., Novaglio, C., eds. (2024). Climate change risks to marine ecosystems and fisheries: Future projections from the Fisheries and Marine Ecosystems Model Intercomparison Project. FAO Fisheries and Aquaculture Technical Paper No. 707. Rome, FAO.</p> <p>Please refer to the above report to cite and for more information.</p> <p>The summary data are here:</p> <p>https://github.com/Fish-MIP/FAO_Report/blob/main/data/table_stats_formatted_admin_full.csv</p> <p>Where the column 'spatial_scale' refers to the type of aggregation:</p> <p>FAO_area = High Sea areas grouped by FAO Major Fishing Areas</p> <p>countries = Exclusive Economic Zones</p> <p>countries_admin = Exclusive Economic Zones results aggregated into Administrative Countries</p> <p>Please note that these results can also be visualised and downloaded from our shiny app: https://rstudio.global-ecosystem-model.cloud.edu.au/shiny/FAO_report_shiny/</p> <p>&nbsp;</p>

openapache2.0Oct 2024View details →
zenodo40/100

Simulated pp collisions at 13 TeV with 2 leptons + 1 b jet final state and selected benchmark Beyond the Standard Model signals

<p>This data-set is comprised of simulated events of pp collisions at 13 TeV with 2 leptons + 1 bottom jet sinal state, with HT &gt; 500 GeV. It includes the following samples</p> <ul> <li>Standard-Model background (bkg), generated at leading order includes the sub-samples Z+Jets, ttbar, WW, WZ, and ZZ. <ul> <li>The processes were generated in kinematic regions to ensure good statistics across the whole phase space. The sampling was carried out using event generation filters at parton level as follows <ul> <li>ttbar: pT &lt;100 GeV; pT in [100, 250] GeV; pT &gt; 250 GeV</li> <li>The scalar sum of the pT of outgoing particles for Z+Jet: ST &lt; 250 Gev; ST in [250, 500] GeV; ST &gt;&nbsp;500 GeV</li> <li>W/Z pT for dibosons: pT &lt; 250 GeV; pT in [250, 500] GeV; pT &gt; 500 GeV</li> </ul> </li> </ul> </li> <li>Vector-like T-quarks with masses 1.0, 1.2, 1.4 TeV (hq1000, hq1200, hq14000) pair produced either through the Standard-Model gluon (wohg) or through a BSM 3TeV heavy gluon (hg3000)</li> <li>tZ production through a Flavour Changing Neutral Current (fcnc) vertex</li> </ul> <p>The samples are provided with both a full set of features, or with a sanitised set of features. The sanitised features remove some accumulation at zeros from non-reconstructed objects (i.e. missing values).&nbsp;All samples were generated using MadGraph5 2.6.5 and the detector was simulated using Delphes 3 with the default CMS card. For the Standard-Model background, both Pythia 8.2 (with CMS CUETP8M1 underlying event tune&nbsp;and NNPDF 2.3 parton distribution functions)&nbsp;(pythia) and Herwig 7 (herwig) hadronisations are provided to compare the background simulation. For the BSM signals only Pythia is provided.</p> <p>For the details of the generation and on the differences between the two feature sets please refer&nbsp;to&nbsp;<a href="https://link.springer.com/article/10.1140%2Fepjc%2Fs10052-020-08807-w">Finding new physics without learning about it: anomaly detection as a tool for searches at colliders</a>&nbsp;for more details. Each file provides a train:validation:split with the ratios 1:1:1 to ensure equal statistical description of the events at each step of the machine learning workflow.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Populations of local direction-selective cells encode global motion patterns generated by self-motion. Data, Code and Model.

<p>Directional tuning of the population of local motion detectors T4/T5 in the visual system of the fruit fly <em>Drosophila melanogaster</em>. Direction tuning and receptive field location was measured by recording responses to visual stimuli containing dark or bright edges/stripes moving into 8 directions. All provided MATLAB scripts were used to analyze and illustrate data show in the manuscript &#39;Populations of local direction-selective cells encode global motion patterns generated by self-motion.&#39;</p> <p>All data were obtained using <em>in vivo </em>two photon microscopy. Image time series were preprocessed using SIMA python software for motion alignment and further processed using custom written matlab or python code.</p> <p>Please find all relevant information to use the code in the README file.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Simulations for pre-industrial climate using EC-Earth3-LR model — selected data for a study on AMOC

<p>A long-term control simulation of pre-industrial period (1850 CE) climates were performed by the EC-Earth3-LR climate model with a horizontal resolution of ~1.125&deg;. The dataset contains selected output data from the simulations.</p> <p>In total, a 2000-year long control simulation was made, which has pre-industrial orbital boundary conditions, initialized by a pre-run steady restart file (the output of approximately 500-year pre-industrial control simulation). This dataset is used to investigate internal climate variability without external forcing changes under pre-industrial climate conditions.</p> <p>The dataset contains Earth system model results from EC-Earth3 presented in the study by Cao et al. (2022).</p> <p><strong>Model configuration</strong><br> Time periods: Pre-Industrial (2000-year time slice)<br> ESM configuration: EC-Earth3-LR<br> Horizontal resolution: ~1.125&deg; (~125 km)</p> <p><strong>Available data</strong><br> Annual mean data for standard oceanographic and meteorological variables.</p>

opencc-by-4.0Nov 2022View details →
dryad40/100

Supplementary tables for: Dependent variable selection in phylogenetic generalized least squares regression analysis under Pagel's lambda model

<p class="MsoNormal"><span>Phylogenetic generalized least squares (PGLS) regression is widely used to detect evolutionary correlations. In contrast to the equal treatment of analyzed traits in conventional correlation methods such as Pearson and Spearman's rank tests, we must designate one trait as the independent variable and the other as the dependent variable. However, in our PGLS regression analyses (using Pagel's <em>λ</em> model) of both empirical and simulated datasets, switching independent and dependent variables yielded many conflicting results. A serious problem with PGLS regression that has not been noticed before is that selecting an inappropriate trait as the dependent variable will often result in an error. To assess correlations in simulated data, we established a gold standard by analyzing changes in traits along phylogenetic branches. Next, we tested seven potential criteria for dependent variable selection: log-likelihood, Akaike information criterion, <em>R</em><sup>2</sup>, <em>p</em>-value, Pagel's <em>λ</em>, Blomberg et al.'s <em>K</em>, and the estimated <em>λ</em> in <a name="_Hlk136010442"></a>Pagel's <em>λ</em> model. We determined that the last three criteria performed equally well in selecting the dependent variable and were superior to the other four. For practicality, we suggest using the trait with a higher <em>λ</em></span><span> or <em>K</em> </span><span>value as the dependent variable in future PGLS regressions. In analyzing the evolutionary relationship between two traits, we should designate the trait with a stronger phylogenetic signal as the dependent variable even if it could logically assume the cause in the relationship.</span></p>

opencc-zeroJun 2023View details →
zenodo40/100

Data for 'Population density affects sexual selection in an insect model'

<p>Data set (.csv file), analysis code (.R file) and readme (.txt file giving details for&nbsp;dataset and code) accompanying&nbsp;the publication&nbsp;&#39;Population density affects sexual selection in an insect model&#39; (Winkler L, Eilhardt R, Janicke T,&nbsp;2023).</p>

opencc-by-4.0Jun 2023View details →
dryad40/100

Isotope mixing scenarios for: To what extent are the source mixing models accurate: evaluation of the model accuracy and guidelines for the site-specific model selection

<p><span>We selected 10 types of distinct isotope signatures that can be found in the samples of natural water. Every 3–10 types of hypothetical isotope signatures were conceptually grouped together. There would be 968 possible combinations based on combinatorics theory. However, we needed distinct mixing polygons to facilitate our determination of model capacity in dealing with uncertainties. Therefore, we </span><span>kept </span><span>only 240 such groups in </span><span>the </span><span>final</span><span> analysis</span><span>. Each group was designated with a </span><span>predefined</span><span> mixing ratio. After that, we ran all the examined models through these mixing scenarios to </span><span>obtain</span><span> the model estimation of the mixing ratios.</span></p>

opencc-zeroOct 2023View details →
dryad40/100

Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad40/100

Selected large model output files and Buffalo sounding data from: Lake Huron enhances snowfall downwind of Lake Erie: a modeling study of the 2010 near year’s Lake-effect snowfall event

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad40/100

Data from: Alternative forms of brook trout nest site selection alter modeled offspring thermal experience and emergence phenology in groundwater-influenced streambeds

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad40/100

A selective small-molecule agonist of G protein-gated inwardly-rectifying potassium channels reduces epileptiform activity in a mouse model of tumor associated epilepsy - Thy1-GCaMP Tumor Electrophysiology

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad40/100

Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Supplementary tables for: Dependent variable selection in phylogenetic generalized least squares regression analysis under Pagel’s lambda model

Open the record for dataset details and reuse information.

publicJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record