Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

34

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

34 results for “cross-validation”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data files belonging to the paper "Dealing with clustered samples for assessing map accuracy by cross-validation"

<p>Mapping of environmental variables often relies on map accuracy assessment through cross-validation with the data used for calibrating the underlying mapping model. When the data points are spatially clustered, conventional cross-validation leads to optimistically biased estimates of map accuracy. Several papers have promoted spatial cross-validation as a means to tackle this over-optimism. Many of these papers blame spatial autocorrelation as the cause of the bias and propagate the widespread misconception that spatial proximity of calibration points to validation points invalidates classical statistical validation of maps. In the paper related to these data, we present and evaluate alternative cross-validation approaches for assessing map accuracy from clustered sample data.&nbsp;</p> <p>&nbsp;</p> <p>The study area is western Europe, constrained in the north at 52&deg; latitude&nbsp;and at -10&deg; and 24&deg; longitude The projection is IGNF:ETRS89LAEA (Lambert azimuthal equal area projection).</p> <p>&nbsp;</p> <p><strong>Files:</strong></p> <p>agb.tif&nbsp; = above ground biomass (AGB) map from&nbsp;version 3 of the 2017 CCI-Biomass product (<a href="https://catalogue.ceda.ac.uk/uuid/5f331c418e9f4935b8eb1b836f8a91b8">https://catalogue.ceda.ac.uk/uuid/5f331c418e9f4935b8eb1b836f8a91b8</a>)<br> AGBstack.tif&nbsp; = covariates used for predicting AGB<br> aggArea.tif&nbsp; = coarse&nbsp;grid used for simulation in the model-based methods<br> ocs.tif&nbsp; = soil organic carbon stock (OCS) map (0-30 cm) from&nbsp;Soilgrids (<a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a>)<br> OCSstack.tif&nbsp; = covariates used for predicting OCS<br> strata.xxx&nbsp;= 100 compact geo-strata (ESRI shape) created with the spcosa package; used for generating clustered samples<br> TOTmask.tif&nbsp; = mask of the area covered by the covariates</p> <p>&nbsp;</p> <p><strong>Details and data sources of the covariates in AGBstack.tif and OCSstack.tif:</strong></p> <table> <tbody> <tr> <td> <p><strong>Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Note</strong></p> </td> </tr> <tr> <td> <p>ai</p> </td> <td> <p>Aridity Index</p> </td> <td> <p><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></p> </td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio1</p> </td> <td> <p>Mean annual air temperature [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio5</p> </td> <td> <p>Mean daily maximum air temperature of the warmest month [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio7</p> </td> <td> <p>Annual range of air temperature [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio12</p> </td> <td> <p>Annual precipitation [kg/m<sup>2</sup>]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio15</p> </td> <td> <p>Precipitation seasonality [kg/m<sup>2</sup>]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>gdd10</p> </td> <td> <p>Growing degree days heat sum above 10&deg;C</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>clay</p> </td> <td> <p>Clay content [g/kg] of the 0-5cm layer</p> </td> <td> <p><a href="https://soilgrids.org/">https://soilgrids.org/</a></p> <p>&nbsp;</p> </td> <td> <p>Only used for AGB</p> </td> </tr> <tr> <td> <p>sand</p> </td> <td> <p>Sand content [g/kg] of the 0-5cm layer</p> </td> <td><a href="https://soilgrids.org/">https://soilgrids.org/</a></td> <td>as above</td> </tr> <tr> <td> <p>pH</p> </td> <td> <p>Acidity (Ph(water)) of the 0-5cm layer</p> </td> <td><a href="https://soilgrids.org/">https://soilgrids.org/</a></td> <td>as above</td> </tr> <tr> <td> <p>glc2017</p> </td> <td> <p>Landcover 2017</p> </td> <td> <p><a href="https://land.copernicus.eu/global/products/lc">https://land.copernicus.eu/global/products/lc</a>, reclassified&nbsp; to: closed forest, open forest,&nbsp; natural non-forest veg., bare &amp; sparse veg. cropland, built-up, water</p> </td> <td> <p>Categorical variable</p> </td> </tr> <tr> <td> <p>dem</p> </td> <td> <p>Elevation</p> </td> <td> <p><a href="https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-eu-dem">https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-eu-dem</a></p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>cosasp</p> </td> <td> <p>Cosine of slope aspect</p> </td> <td> <p>Computed with the terra package from elevation</p> </td> <td>Computed @25m resolution; next aggregated to 0.5km</td> </tr> <tr> <td> <p>sinasp</p> </td> <td> <p>Sine of slope aspect</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>slope</p> </td> <td> <p>Slope</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TPI</p> </td> <td> <p>Topographic position index</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TRI</p> </td> <td> <p>Terrain ruggedness index</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TWI</p> </td> <td> <p>Topographic wetness index</p> </td> <td> <p>Computed with SAGA from 500m resolution (aggregated) dem</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>gedi</p> </td> <td> <p>Forest height</p> </td> <td> <p><a href="https://glad.umd.edu/dataset/gedi">https://glad.umd.edu/dataset/gedi</a></p> </td> <td> <p>Zone: NAFR</p> </td> </tr> <tr> <td> <p>xcoord</p> </td> <td> <p>X coordinate</p> </td> <td> <p>Using a mask created from the other covariates</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>ycoord</p> </td> <td> <p>Y coordinate</p> </td> <td>Using a mask created from the other covariates</td> <td>&nbsp;</td> </tr> <tr> <td> <p>Dcoast</p> </td> <td> <p>Distance from coast</p> </td> <td> <p>Using a land mask created from the other covariates</p> </td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

CROSS-VALIDATION OF FUNCTIONAL MRI and PARANOID-DEPRESSIVE SCALE: BRAIN SIGNATURES FROM MULTIVARIATE ANALYSIS

<p>Brain signatures identified by bottom-up unsupervised machine learning: three principal components based on activations yielded from the three kinds of diagnostically relevant stimuli are used in order to produce cross-validation markers which may effectively predict the variance on the level of clinical populations and eventually delineate diagnostic and classification groups.&nbsp; The stimuli represent items from a paranoid-depressive self-evaluation scale, administered simultaneously with functional magnetic resonance imaging (fMRI).</p> <p>We have been able to separate the two investigated clinical entities &ndash; schizophrenia and recurrent depression by use of multivariate linear model and principal component analysis. This is a confirmation of the possibility to achieve bottom-up classification of mental disorders, by use of the brain signatures relevant to clinical evaluation tests.</p>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper

<p>Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper (complement for a tutorial at&nbsp;https://github.com/datngu/LmTag)</p> <p>This repo includes chromosome 10 reference panel data constructed for 3 populations:</p> <p>- EAS</p> <p>- EUR</p> <p>- SAS</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data for the Article: Cross-validation of a semantic segmentation network for natural history collection specimens

<p>This deposit contains six datasets which were used for testing and validating a semantic segmentation network. The purpose was to evaluate the suitability of the segmentation network for use in the processing of images from Natural History Collections.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

DocTOR models and cross-validation dataset

<p>Dataset necessary for DocTOR utility.</p> <p>DocTOR (Direct fOreCast Target On Reaction), is a utility written in python3.9 (using the conda workframe) that allows the user to upload a list of Uniprot IDs and Adverse reactions (from the available models) in order to study the relationship between the two.</p> <p>On output the program will assign a positive or negative class to the protein, assessing its possible involvement in the selected ADRs onset.</p> <p>DocTOR exploits the data coming from T-ARDIS [https://doi.org/10.1093/database/baab068] to train different Machine Learning approaches (SVM, RF, NN) using network topological measurements as features.</p> <p>The prediction coming from the single trained models are combined in a meta-predictor exploiting three different voting systems.</p> <p>The results of the meta-predictor together with the ones from the single ML method will be available in the output log file (named &quot;predictions_community&quot; or &quot;predictions_curated&quot; based on the database type).</p> <p>The DocTOR utility is avaiable at&nbsp;https://github.com/cristian931/DocTOR</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

MDSINE2 Cross-Validation Analysis (Healthy Cohort)

<p>MDSINE2 Inference Analysis files. This archive contains all output files from the cross-validation inference (with comparator analysis included) for Healthy cohort. (Both healthy and dysbiotic cohorts are required to run the Jupyter Notebook (`fig4_semisynthetic_v2_cache.ipynb`) on our MDSINE2_Paper repo. For the Dysbiotic cohort files, refer to <a href="https://zenodo.org/records/16915340" target="_blank" rel="noopener">https://zenodo.org/records/16915340</a>.</p> <p>For the full project/source pipeline, refer to <a href="https://github.com/gerberlab/MDSINE2_Paper">https://github.com/gerberlab/MDSINE2_Paper</a>.</p> <p>The archive here was created using the command `tar --zstd -cvf`, and then split using the unix "split" command. To unpack these files, please use the following command:</p> <pre><code>cat cross_validation_healthy.tar.zst.part* &gt; cross_validation_healthy.tar.zst tar -I zstd -xvf cross_validation_healthy.tar.zst</code></pre> <p>These files should be unpacked and placed in the paper repository directory (wherever you did `git clone`), so that the `datasets` directory is directly inside `MDSINE2_Paper` repository directory.&nbsp;</p> <p>------------</p> <p>Related zenodo records:</p> <p><a href="https://doi.org/10.5281/zenodo.8208502">https://doi.org/10.5281/zenodo.8208502</a> -- Full MDSINE2 inference on Healthy cohort</p> <p><a href="https://doi.org/10.5281/zenodo.16915340">https://doi.org/10.5281/zenodo.16915340</a> -- MDSINE2 Cross-Validation run (Dysbiotic cohort)</p> <p><a href="https://doi.org/10.5281/zenodo.16915311">https://doi.org/10.5281/zenodo.16915311</a> -- MDSINE2 semisynthetic dataset</p>

opencc-by-4.0Jun 2023View details →
dryad36/100

Data reported in development and cross-validation of a veterans mental health risk factor screen

<p>Background. VA primary care patients are routinely screened for current symptoms of PTSD, depression, and alcohol disorders, but many who screen positive do not engage in care. In addition to stigma about mental disorders and a high value on autonomy, some veterans may not seek care because of uncertainty about whether they need treatment to recover. A screen for mental health risk could provide an alternative motivation for patients to engage in care.</p> <p>Results. Twelve items assessing dissociation, emotional lability, life stress, and moral injury correctly classified 86% of those who later had elevated PTSD and/or depression symptoms (sensitivity) and 75% of those whose later symptoms were not elevated (specificity). Performance was also very good for 110 veterans who identified as members of ethnic/racial minorities.</p> <p>Conclusions. Mental health status was prospectively predicted in VA primary care patients with high accuracy using a screen that is brief, easy to administer, score, and interpret, and fits well into VA's integrated primary care. When care is readily accessible, appealing to veterans, and not perceived as stigmatizing, information about mental health risk may result in higher rates of engagement than information about current mental disorder status.</p>

opencc-zeroOct 2022View details →
dryad36/100

Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation

<p>The mosquito <em>Aedes aegypti </em>is the primary vector of many human arboviruses such as dengue, yellow fever, chikungunya, and Zika, which affect millions of people world-wide. Population genetics studies on this mosquito have been important in understanding its invasion pathways and success as a vector of human disease. The Axiom aegypti1 SNP chip was developed from a sample of geographically diverse <em>Ae. aegypti </em>populations to facilitate genomic studies on this species. Here we evaluate the utility of the Axiom aegypti1 SNP chip for population genetics and compare it with a low-depth shot-gun sequencing approach using mosquitoes from the species' native (Africa) and invasive range (outside Africa). These analyses indicate that the results from the SNP chip are highly reproducible and have a higher sensitivity to capture alternative alleles than a low-coverage whole-genome sequencing approach. Although the SNP chip suffers from ascertainment bias, results from population structure, ancestry, demographic, and phylogenetic analyses using the SNP chip were congruent with those derived from low coverage whole genome sequencing, and consistent with previous reports on Africa and outside Africa populations using microsatellites. More importantly, we identified a subset of SNPs that can be reliably used to generate merged databases, opening the door to combined analyses. We conclude that the Axiom aegypti1 SNP chip is a convenient, more accurate, low-cost alternative to low-depth whole genome sequencing for population genetic studies of <em>Ae. aegypti</em> that do not rely on full allelic frequency spectra. Whole genome sequencing and SNP chip data can be easily merged, extending the usefulness of both approaches. </p>

opencc-zeroApr 2024View details →
zenodo36/100

Data for "Identifying and Fitting Eclipse Maps of Exoplanets with Cross-Validation" (Hammond et al. 2024)

<p>This archive contains the data and scripts needed to reproduce the analysis in "Identifying and Fitting Eclipse Maps of Exoplanets with Cross-Validation" (Hammond et al. 2024).</p> <p>Contents</p> <p>data/: Input data and posterior distributions of different model fits</p> <p>datasets/: Observational datasets</p> <p>figures/: Folder to save figures in</p> <p>fluxes/: Saved lightcurves for auxiliary plotting purposes</p> <p>archive_paper_plotter.ipynb: Example script to plot fitted eclipse maps</p> <p>eclipse_pixel_sampling.py: Script to fit eclipse map and test k-fold CV score</p> <p>paper_eclipse_suite.py: Script to use simulated or observational data to fit an eclipse map</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Dissimilarity-adaptive cross-validation experiments and datasets

<p>This data includes all datasets and codes for implementing dissimilarity-adaptive cross-validation experiments, datasets and code, Reademe.txt explains each file's meaning. Appendix includes the descriptions of datasets.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Python functions -- cross-validation methods from a data-driven perspective

<p>This is the organized python functions of proposed methods in Yanwen Wang PhD research. Researchers can directly use these functions to conduct spatial+ cross-validation (SP-CV), dissimilarity quantification by adversarial validation (AVD), and dissimilarity-adaptive cross-validation (DA-CV). The description of how to run codes is in Readme.txt. The descriptions of functions are in functions.docx.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Spatial Plus Cross-Validation experiments datasets and codes

<p>This zip file includes all materials of Spatial Plus Cross-Validation experiments.&nbsp;</p> <p>They are ordered by the first number of folder&#39;s name.</p> <p>In each folder, the order of running code scripts are labeled by the first number of code&#39;s name.</p>

opencc-by-4.0Jan 2023View details →
dryad36/100

Identifying the best approximating model in Bayesian phylogenetics: Bayes factors, cross-validation or wAIC?

<p>There is still no consensus as to how to select models in Bayesian phylogenetics, and more generally in applied Bayesian statistics. Bayes factors are often presented as the method of choice, yet other approaches have been proposed, such as cross-validation or information criteria. Each of these paradigms raises specific computational challenges, but they also differ in their statistical meaning, being motivated by different objectives: either testing hypotheses or finding the best-approximating model. These alternative goals entail different compromises, and as a result, Bayes factors, cross-validation and information criteria may be valid for addressing different questions. Here, the question of Bayesian model selection is revisited, with a focus on the problem of finding the best-approximating model. Several model selection approaches were re-implemented, numerically assessed and compared: Bayes factors, cross-validation (CV), in its different forms (k-fold or leave-one-out), and the widely applicable information criterion (wAIC), which is asymptotically equivalent to leave-one-out cross validation (LOO-CV). Using a combination of analytical results and empirical and simulation analyses, it is shown that Bayes factors are unduly conservative. In contrast, cross-validation represents a more adequate formalism for selecting the model returning the best approximation of the data-generating process and the most accurate estimates of the parameters of interest. Among alternative CV schemes, LOO-CV and its asymptotic equivalent represented by the wAIC, stand out as the best choices, conceptually and computationally, given that both can be simultaneously computed based on standard MCMC runs under the posterior distribution.</p>

opencc-zeroFeb 2023View details →
dryad36/100

Identifying the best approximating model in Bayesian phylogenetics: Bayes factors, cross-validation or wAIC?

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad36/100

Genotypes of Aedes aegypti mosquitoes derived from SNP chip and low-coverage whole genome sequencing for platform cross-validation

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Data from: cross-validation matters in species distribution models: a case study with goatfish species

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad36/100

Data reported in development and cross-validation of a veterans mental health risk factor screen

Open the record for dataset details and reuse information.

publicOct 2022View details →
zenodo32/100

GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank (5-fold cross-validation)

<p>This dataset contains GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank.</p> <p>The dataset is designed to enable systematic PRS analyses with 5-fold cross validation. For each phenotype and fold, we provide GWAS summary statistics for the training, validation, and test sets. The validation summary statistics can be used for model selection/tuning. The test summary statistics can be used to evaluate PRS models via pseudo-validation metrics. Association testing for all phenotypes and samples was done with <strong>plink2</strong>.</p> <p>&nbsp;</p> <p>The&nbsp;<strong>phenotypes</strong> included in this dataset are:</p> <ul> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=50">HEIGHT</a>: Standing height (Data-Field: 50)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=21001">BMI</a>: Body mass index (Data-Field: 21001)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=48">WC</a>: Waist circumference (Data-Field: 48)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=49">HC</a>: Hip circumference (Data-Field: 49)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=20022">BW</a>: Birth weight (Data-Field: 20022)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3062">FVC</a>: Forced vital capacity (Data-Field: 3062)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3063">FEV1</a>: Forced expiratory volume in 1-second (Data-Field: 3063)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30760">HDL</a>: HDL cholesterol (Data-Field: 30760)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30780">LDL</a>: LDL cholesterol (Data-Field: 30780)</li> </ul> <p>&nbsp;</p> <p>To allow users to assess PRS performance as a function of sample size, we also provide <strong>subsampled training GWAS summary statistics</strong>. This is done by taking the training samples and randomly selecting (without replacement) a subset of them for conducting association testing. The training sample sizes are:</p> <ul> <li>N = 5000</li> <li>N = 10000</li> <li>N = 20000</li> <li>N = 40000</li> <li>N = 80000</li> <li>N = 160000</li> <li>Full training set (sample size varies by phenotype).</li> </ul> <p><strong>NOTE</strong>: Due to the smaller overall sample size for the Birth weight phenotype, we do not include training data for the `N=160000` setting.<br><br></p> <p>The <strong>folder structure</strong> of the GWAS data for each phenotype is as follows:</p> <ul> <li><code>train</code> <ul> <li><code>N_5000</code> <ul> <li><code>&nbsp;fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>N_10000</code></li> <li><code>N_20000</code></li> <li><code>N_40000</code></li> <li><code>N_80000</code></li> <li><code>N_160000</code></li> <li><code>full</code></li> </ul> </li> <li><code>validation</code> <ul> <li><code>fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>test</code> <ul> <li><code>fold_1</code></li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> </ul> <p>For more details about the GWAS study, Quality Control (QC) criteria, or other information, please consult our publication:</p> <p>Zabad, S., Gravel, S., &amp; Li, Y. (2023).&nbsp;<strong>Fast and accurate Bayesian polygenic risk modeling with variational inference.</strong>&nbsp;The American Journal of Human Genetics, 110(5), 741&ndash;761.&nbsp;<a href="https://doi.org/10.1016/j.ajhg.2023.03.009" rel="nofollow">https://doi.org/10.1016/j.ajhg.2023.03.009</a></p> <p>If you use this data in your work, please cite the publication above.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Outputs from fitted models across the cross-validation scenarios for 'Space-time species distribution modeling with opportunistic presence-only data: a case study of passerines in a protected area'

<p>Three Zenodo repositories are linked to the preprint <em>Space-time Species Distribution Modeling for Opportunistic Presence-Only Data: A Case Study of Passerines in a Protected Area&nbsp; </em>(Lasgorceux et al., unpublished, <a href="https://hal.science/hal-04616332">https://hal.science/hal-04616332</a>):</p> <ul> <li>Data, scripts and, code (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12545052">https://doi.org/10.5281/zenodo.12545052</a>)</li> <li>Outputs from fitted models across the cross-validation scenarios (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12544212">https://doi.org/10.5281/zenodo.12544212</a>)</li> <li>Supplementary information at (Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12541412">https://doi.org/10.5281/zenodo.12541412</a>)</li> </ul> <p>This repository contains the outputs from fitted models across the cross-validation scenarios.</p> <p>In the folder <em>Ouputs_cross_validation</em>, each species is represented by a .RData file, numbered from 1 to 77 (excluding 7, which corresponds to <em>Bombycilla garrulus</em>; see the preprint for details).&nbsp;This dataset is specifically used to generate Figure 1, which shows the AUC of various cross-validation scenarios. To reproduce this figure in R, place all the files in the&nbsp;<em>Results/Fitted_models</em> folder and run the <em>Models_Outputs.R</em> script located in the <em>Results</em> folder of Lasgorceux et al., Zenodo, <a href="https://doi.org/10.5281/zenodo.12545052">https://doi.org/10.5281/zenodo.12545052.</a></p> <p>Note: These data have been separated due to memory requirements (23.14GB).</p>

opencc-by-4.0Jun 2024View details →
dryad28/100

Data from: Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure

Ecological data often show temporal, spatial, hierarchical (random effects), or phylogenetic structure. Modern statistical approaches are increasingly accounting for such dependencies. However, when performing cross-validation, these structures are regularly ignored, resulting in serious underestimation of predictive error. One cause for the poor performance of uncorrected (random) cross-validation, noted often by modellers, are dependence structures in the data that persist as dependence structures in model residuals, violating the assumption of independence. Even more concerning, because often overlooked, is that structured data also provides ample opportunity for overfitting with non-causal predictors. This problem can persist even if remedies such as autoregressive models, generalized least squares, or mixed models are used. Block cross-validation, where data are split strategically rather than randomly, can address these issues. However, the blocking strategy must be carefully considered. Blocking in space, time, random effects or phylogenetic distance, while accounting for dependencies in the data, may also unwittingly induce extrapolations by restricting the ranges or combinations of predictor variables available for model training, thus overestimating interpolation errors. On the other hand, deliberate blocking in predictor space may also improve error estimates when extrapolation is the modelling goal. Here, we review the ecological literature on non-random and blocked cross-validation approaches. We also provide a series of simulations and case studies, in which we show that, for all instances tested, block cross-validation is nearly universally more appropriate than random cross-validation if the goal is predicting to new data or predictor space, or for selecting causal predictors. We recommend that block cross-validation be used wherever dependence structures exist in a dataset, even if no correlation structure is visible in the fitted model residuals, or if the fitted models account for such correlations.

opencc-zeroDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record