Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41
datasets available to search
ShareScore release 0.7.1
Dataset results
41 results for “Error Estimation”
Supporting Dataset (Tables S1-S7, Figure S1) for "Heavy-mineral grain counting: Counting techniques, error estimation, and the number of grains to be counted"
<p>This Dataset comprises Tables S1-S7 for the article <strong>"Heavy-mineral grain counting: Counting techniques, error estimation, and the number of grains to be counted"</strong> authored by Jan Schönig and submitted to J<em>ournal of Geophysical Research: Earth Surface</em>.</p> <p>Figure S1: Error of different counting methods in comparison with theory and without considering the finite population correction</p> <p>Table S1: Heavy-mineral dataset from semi-automated Raman analysis</p> <p>Table S2: Summary of heavy-mineral composition for individual samples in percent</p> <p>Table S3: Computed ribbon compositions for individual samples and ribbon sizes given in numbers of counts</p> <p>Table S4: Relative errors at 95 % quantile for consecutive ribbon counting simulations</p> <p>Table S5: Relative errors at 95 % quantile for maximum distance ribbon counting simulations</p> <p>Table S6: Relative errors at 95 % quantile for cluster counting simulations</p> <p>Table S7: Numerical solution for determining the number of counts required to discriminate the content of two mineral species in an aliquot</p>
Рис. 1. Вероятность обнаружения меченых животных (среΑнее ± ошибка) при пяти- и Αесятиметровых интерваΛах межΑу прикормочными станциями в Αвух экспериментах. По второму эксперименту расчеты сΑеΛаны ΑΛя резуΛьтатов отΛова в течение первых трех и поΛных Αесяти Αней. Значение «p» отражает уровень статистической значимости разΛичий межΑу ΑоΛями животных с меткой при Αвух интерваΛах Fig. 1. Probability of finding marked animals (average±standard error) between feeding stations placed at intervals of five and ten meters in the two experiments. In the second experiment, calculations were made for the results of trapping during the first three days and during the whole period of ten days. The p value reflects the statistical significance of differences between the fractions of animals with a mark for two types of intervals in Verification of the bottle-based method for estimating abundance of small mammals using biomarkers
Рис. 1. Вероятность обнаружения меченых животных (среΑнее ± ошибка) при пяти- и Αесятиметровых интерваΛах межΑу прикормочными станциями в Αвух экспериментах. По второму эксперименту расчеты сΑеΛаны ΑΛя резуΛьтатов отΛова в течение первых трех и поΛных Αесяти Αней. Значение «p» отражает уровень статистической значимости разΛичий межΑу ΑоΛями животных с меткой при Αвух интерваΛах Fig. 1. Probability of finding marked animals (average±standard error) between feeding stations placed at intervals of five and ten meters in the two experiments. In the second experiment, calculations were made for the results of trapping during the first three days and during the whole period of ten days. The p value reflects the statistical significance of differences between the fractions of animals with a mark for two types of intervals
Relative Random Errors in the Convective Atmospheric Boundary Layer Estimated by the Relaxed Filtering Method from Large Eddy Simulations
<p>Data supporting the paper "How representative are uncrewed aircraft system measurements of the convective boundary layer?" by Brian R. Greene, Leia M. Otterstatter, and Scott T. Salesky, submitted to Geophysical Research Letters in 2024. Data are postprocessed from large-eddy simulations of the convective atmospheric boundary layer that are used to produce the figures within the paper. Details on the production of these files are included in the supplementary informatin of this paper.</p>
Using unoccupied aerial vehicles to estimate availability and group size error for aerial surveys of coastal dolphins
<p><span>Aerial surveys are frequently used to estimate the abundance of marine mammals, but their accuracy is dependent upon obtaining a measure of the availability of animals for visual detection. Existing methods for characterizing availability have limitations and do not necessarily reflect true availability. Here, we present a method of using small, vessel‐launched, multi‐rotor Unoccupied Aerial Vehicles (UAVs or drones) to collect video of dolphins to characterize availability and investigate errors surrounding group size estimates. We collected over 20 h of aerial video of dive‐surfacing behaviour across 32 encounters with the Australian humpback dolphin </span><em><span>Sousa sahulensis</span></em><span> off north‐western Australia. Mean surfacing and dive periods were 7.85 sec (</span><span>se</span><span> = 0.26) and 39.27 sec (</span><span>se</span><span> = 1.31) respectively. Dolphin encounters were split into 56 focal follows of consistent group composition to which example approaches to estimating availability were applied. Non‐instantaneous availability estimates, assuming a 7-sec observation window, ranged between 0.22 and 0.88, with a mean availability of 0.46 (CV = 0.34). Availability tended to increase with increasing group size. We found a downward bias in group size estimation, with true group size typically one individual more than would have been estimated by a human observer during a standard aerial survey. The variability of availability estimates between focal follows highlights the importance of sampling across a variety of group sizes, compositions, and environmental conditions. Through data re‐sampling exercises, we explored the influence of sample size on availability estimates and their precision, with results providing an indication of target sample sizes to minimize bias in future research. We show that UAVs can provide an effective and relatively inexpensive method of characterizing dolphin availability with several advantages over existing approaches. The example estimates obtained for humpback dolphins are within the range of values obtained for other shallow‐water, small cetaceans, and will directly inform a government‐run program of aerial surveys in the region.</span></p>
Standard Error Estimates from ARRI ensemble GAM model outputs
<p><strong>Summary (Purpose)</strong></p> <p>Basal area per acre (BAA) standard error estimate (SEE) for all trees, pine trees, and non-pine trees across three diameter at breast height (DBH) size classes, 2- to 10-inch, 10- to 14-inch, and 14+ inch. Models were informed by relative density rasters from 2018 Light Detection and Ranging (lidar) point clouds.</p> <p><strong>Description</strong></p> <p>The GAM_SEE_rasters are modelled basal area per acre standard error estimate single band rasters. The units for the rasters’ are square foot per acre. Rasters are divided into three tree species groups: 'All' trees, 'Pine' trees (defined as trees of the genus<em> Pinus)</em>, and 'No-Pine' trees, and three size classes: LT 10 for trees with DBH between 2- and 10-inches, 10-14 for trees with DBH between 10- and 14- inch DBH, and GT 14 for trees with DBH greater than 14-inches.</p> <p>Ensemble generalized additive models (GAM) of estimated tree basal area per DBH class were created from Restore field plots and relative density canopy cover rasters, or RDCC (St. Peter, et al., 2023). This ensemble GAM model was created using the R script detailed in Hogland, 2021. The parameters used were 0.75 for the percent of data used to train the model (selected using random sampling with replacement), 50 models, and using gaussian family. The estimated BAA for each of the 50 ensemble GAM models for each cell were averaged (mean) to produce the GAM estimate, additionally the variability between these estimates was used to create the standard error estimate (SEE) for each cell. </p> <p> </p> <p>The 246 Restore field plots used to train the GAM model of basal area include measurements of all trees within four non-overlapping 9m radius circular subplots within a 36m square plot. Tree diameter at breast height (DBH), species, count, and condition measurements were recorded. Measurements were summarized to the plot and DBH (square inches) was converted to basal area per acre (square feet per acre) using the formula 0.005454 * DBH^2. Restore field plots were measured in the Spring of 2018. The RDCC metrics are 5m resolution multiband rasters produced by applying a custom r software function that uses the r software’s ‘lidR’ package to produce forest metrics summarized from Light Imaging Detection and Ranging (LiDAR) point clouds. ARSA LiDAR is a combination of three collections, Block 2 and 3 were collected in early 2018 and has a NPS of 0.7-m using a Riegl VQ-1560i lidar system. Leon county LiDAR data has a nominal pulse spacing (NPS) of 0.35-m and was acquired between February 05, 2018 and April 25, 2018 using the Leica ALS80 HP SN8137 and SN8235 lidar systems. Choctawhatchee data was acquired in early 2017, using the Riegl LMS-Q1560 lidar system and has a NPS of 0.7-m. Rasters were generated in their vendor provided spatial projection before being reprojected to UTM Zone 16.</p> <p> </p> <p>The 5m resolution RDCC bands were summarized to 40x40m to correspond with the area of our plots and used as predictor variables in the models of all trees basal area. Each 5m pixel value represents the estimated standard error of the basal area per acre estimate as if it was the center of a 40x40m (8x8 cells) plot surrounding that pixel. </p> <p> </p> <p>Reference:</p> <p>Hogland, J. (2021). Ensemble Generalized Additive Models (EGAM). Retrieved from Jupyter Notebook: <a href="https://colab.research.google.com/drive/1GnRagruTUCoPJQZSkZ2vMKS9aAKgnhEw?usp=sharing">https://colab.research.google.com/drive/1GnRagruTUCoPJQZSkZ2vMKS9aAKgnhEw?usp=sharing</a></p> <p>St. Peter, Joseph, Drake, Jason, Medley, Paul, & Ibeanusi, Victor. (2023). Relative Density Canopy Cover Outputs for Leon Lidar data in the Florida Panhandle 2018 [Data set]. In Remote Sensing (Vol. 13, Number 23, p. 4763). Zenodo. https://doi.org/10.5281/zenodo.8222114</p> <p><strong>Credits</strong></p> <p>This dataset was built by Joseph St. Peter of FAMU’s Center for Spatial Ecology and Restoration using 2018 LiDAR data funded by Leon County, Northwest Florida Water Management District, US Geological Survey and the USDA Forest Service and processed using the r package lidR. Restore plots were funded by the Gulf Coast Ecosystem Restoration Council (RESTORE Council) through an interagency agreement with the USDA Forest Service (17-IA-11083150-001) for the Apalachicola Tate’s Hell Strategy 1 project.</p> <p><strong>Use Limitations</strong></p> <p>This spatial data is based on various data collection and processing techniques as well as on modeling or interpretation. While this data uses the most current and complete information available at the time of production, spatial data and derivative products may vary in accuracy. Spatial data are often developed from sources of differing accuracy which may be accurate only at certain scales. This data has been quality checked but may contain spurious errors or be incomplete or inappropriate for certain uses. Spatial data products used for purposes other than those for which they were created, may yield inaccurate or misleading results. The USDA Forest Service, Florida A&M University, and the Center for Spatial Ecology & Restoration (CSER) reserves the right to correct, update, modify, remove or replace GIS products at any time and without notification. This data may not be distributed without written permission from the Center for Spatial Ecology & Restoration (CSER) at Florida A&M University, and/or the USDA Forest Service.</p>
Data from: Robustness of divergence time estimation despite gene tree error: A case study of fireflies (Coleoptera: Lampyridae)
Open the record for dataset details and reuse information.
Data from: Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC
Open the record for dataset details and reuse information.
Using unoccupied aerial vehicles to estimate availability and group size error for aerial surveys of coastal dolphins
Open the record for dataset details and reuse information.
Results for paper "Energy dependent mesh adaptivity of discontinuous isogeometric discrete ordinate methods with dual weighted residual error estimators"
<p>This spreadsheet contains the results used to generate the plots in the paper "Energy dependent mesh adaptivity of discontinuous isogeometric discrete ordinate methods with dual weighted residual error estimators".</p>
East Cascades Error Adjusted Area Estimates
<p>The files in this folder include the error-adjusted area estimates of seven structural classes for five strata, as well as the annual area burned for each class at different levels of fire severity. Files with “_entire_” include the error-adjusted area estimates of each class for the complete extent of each strata, while files with “_unburned_” include the error-adjusted area estimates of each class for only the portion of each strata that was unburned during the study period. All area estimates are in hectares.</p>
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.
Observing system simulation experiments to evaluate transport model error on CO2 flux estimates
<p>Atmospheric CO2 inversion using coarse-resolution transport model can cause large errors on surface carbon flux estimates. The transport model errors on flux estimates are isolated using observing system simulation experiments presented here.</p>
Standard Error Estimates for ARRI Ensemble LM model outputs
<p>Summary</p> <p>Basal area per acre (BAA) standard error estimate (SEE) for all trees, pine trees, and non-pine trees across three diameter at breast height (DBH) size classes, 2- to 10-inch, 10- to 14-inch, and 14+ inch. Models were informed by relative density rasters from 2018 Light Detection and Ranging (lidar) point clouds.</p> <p>Description</p> <p>The LM_SEE_rasters are modelled basal area per acre standard error estimate single band rasters. The units for the rasters’ are square foot per acre. Rasters are divided into three tree species groups: 'All' trees, 'Pine' trees (defined as trees of the genus<em> Pinus)</em>, and 'No-Pine' trees, and three size classes: LT 10 for trees with DBH between 2- and 10-inches, 10-14 for trees with DBH between 10- and 14- inch DBH, and GT 14 for trees with DBH greater than 14-inches.</p> <p>Ensemble linear regression models (LM) of estimated tree basal area per DBH class were created from Restore field plots and relative density canopy cover rasters, or RDCC (St. Peter, et al., 2023). This ensemble LM model was created using a custom R script that was based off the work detailed in Hogland, 2021. The ensemble LM script was modified to use the lm() function in place of the GAM modelling functions. The parameters used were 0.75 for the percent of data used to train the model (selected using random sampling with replacement), 50 models, and using gaussian family. The estimated BAA for each of the 50 ensemble LM models for each cell were averaged (mean) to produce the LM estimate, additionally the variability between these estimates was used to create the standard error estimate (SEE) for each cell. </p> <p>The 246 Restore field plots used to train the LM model of basal area include measurements of all trees within four non-overlapping 9m radius circular subplots within a 36m square plot. Tree diameter at breast height (DBH), species, count, and condition measurements were recorded. Measurements were summarized to the plot and DBH (square inches) was converted to basal area per acre (square feet per acre) using the formula 0.005454 * DBH^2. Restore field plots were measured in the Spring of 2018. The RDCC metrics are 5m resolution multiband rasters produced by applying a custom r software function that uses the r software’s ‘lidR’ package to produce forest metrics summarized from Light Imaging Detection and Ranging (LiDAR) point clouds. ARSA LiDAR is a combination of three collections, Block 2 and 3 were collected in early 2018 and has a NPS of 0.7-m using a Riegl VQ-1560i lidar system. Leon county LiDAR data has a nominal pulse spacing (NPS) of 0.35-m and was acquired between February 05, 2018 and April 25, 2018 using the Leica ALS80 HP SN8137 and SN8235 lidar systems. Choctawhatchee data was acquired in early 2017, using the Riegl LMS-Q1560 lidar system and has a NPS of 0.7-m. Rasters were generated in their vendor provided spatial projection before being reprojected to UTM Zone 16.</p> <p>The 5m resolution RDCC bands were summarized to 40x40m to correspond with the area of our plots and used as predictor variables in the models of all trees basal area. Each 5m pixel value represents the estimated standard error of the basal area per acre estimate as if it was the center of a 40x40m (8x8 cells) plot surrounding that pixel. </p> <p>References:</p> <p>Hogland, J. (2021). Ensemble Generalized Additive Models (EGAM). Retrieved from Jupyter Notebook: <a href="https://colab.research.google.com/drive/1GnRagruTUCoPJQZSkZ2vMKS9aAKgnhEw?usp=sharing">https://colab.research.google.com/drive/1GnRagruTUCoPJQZSkZ2vMKS9aAKgnhEw?usp=sharing</a></p> <p>St. Peter, Joseph, Drake, Jason, Medley, Paul, & Ibeanusi, Victor. (2023). Relative Density Canopy Cover Outputs for Leon Lidar data in the Florida Panhandle 2018 [Data set]. In Remote Sensing (Vol. 13, Number 23, p. 4763). Zenodo. <a href="https://doi.org/10.5281/zenodo.8222114">https://doi.org/10.5281/zenodo.8222114</a></p> <p>Credits</p> <p>This dataset was built by Joseph St. Peter of FAMU’s Center for Spatial Ecology and Restoration using 2018 LiDAR data funded by Leon County, Northwest Florida Water Management District, US Geological Survey and the USDA Forest Service and processed using the r package lidR. Restore plots were funded by the Gulf Coast Ecosystem Restoration Council (RESTORE Council) through an interagency agreement with the USDA Forest Service (17-IA-11083150-001) for the Apalachicola Tate’s Hell Strategy 1 project.</p> <p>Use Limitations</p> <p>This spatial data is based on various data collection and processing techniques as well as on modeling or interpretation. While this data uses the most current and complete information available at the time of production, spatial data and derivative products may vary in accuracy. Spatial data are often developed from sources of differing accuracy which may be accurate only at certain scales. This data has been quality checked but may contain spurious errors or be incomplete or inappropriate for certain uses. Spatial data products used for purposes other than those for which they were created, may yield inaccurate or misleading results. The USDA Forest Service, Florida A&M University, and the Center for Spatial Ecology & Restoration (CSER) reserves the right to correct, update, modify, remove or replace GIS products at any time and without notification. This data may not be distributed without written permission from the Center for Spatial Ecology & Restoration (CSER) at Florida A&M University, and/or the USDA Forest Service.</p>
Spatial uncertainty in herbarium data: Simulated displacement but not error distance alters estimates of phenological sensitivity to climate in a widespread California wildflower
Open the record for dataset details and reuse information.
Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference
Open the record for dataset details and reuse information.
The perfect storm: Gene tree estimation error, incomplete lineage sorting, and ancient gene flow explain the most recalcitrant ancient angiosperm clade, Malpighiales
<p>The genomic revolution offers renewed hope of resolving rapid radiations in the Tree of Life. The development of the multispecies coalescent (MSC) model and improved gene tree estimation methods can better accommodate gene tree heterogeneity caused by incomplete lineage sorting (ILS) and gene tree estimation error stemming from the short internal branches. However, the relative influence of these factors in species tree inference is not well understood. Using anchored hybrid enrichment, we generated a data set including 423 single-copy loci from 64 taxa representing 39 families to infer the species tree of the flowering plant order Malpighiales. This order includes nine of the top ten most unstable nodes in angiosperms, which have been hypothesized to arise from the rapid radiation during the Cretaceous. Here, we show that coalescent-based methods do not resolve the backbone of Malpighiales and concatenation methods yield inconsistent estimations, providing evidence that gene tree heterogeneity is high in this clade. Despite high levels of ILS and gene tree estimation error, our simulations demonstrate that these two factors alone are insufficient to explain the lack of resolution in this order. To explore this further, we examined triplet frequencies among empirical gene trees and discovered some of them deviated significantly from those attributed to ILS and estimation error, suggesting gene flow as an additional and previously unappreciated phenomenon promoting gene tree variation in Malpighiales. Finally, we applied a novel method to quantify the relative contribution of these three primary sources of gene tree heterogeneity and demonstrated that ILS, gene tree estimation error, and gene flow contributed to 15%, 52%, and 32% of the variation, respectively. Together, our results suggest that a perfect storm of factors likely influence this lack of resolution, and further indicate that recalcitrant phylogenetic relationships like the backbone of Malpighiales may be better represented as phylogenetic networks. Thus, reducing such groups solely to existing models that adhere strictly to bifurcating trees greatly oversimplifies reality, and obscures our ability to more clearly discern the process of evolution.</p>
Data for: Gene tree estimation error with ultraconserved elements: An empirical study on Pseudapis bees
<p>Summarizing individual gene trees to species phylogenies using two-step coalescent methods is now a standard strategy in the field of phylogenomics. However, practical implementations of summary methods suffer from gene tree estimation error, which is caused by various biological and analytical factors. Greatly understudied is the choice of gene tree inference method and downstream effects on species tree estimation for empirical data sets. To better understand the impact of this method choice on gene and species tree accuracy, we compare gene trees estimated through four widely used programs under different model-selection criteria: PhyloBayes, MrBayes, IQ-Tree and RAxML. We study their performance in the phylogenomic framework of > 800 ultraconserved elements from the bee subfamily Nomiinae (Halictidae). Our taxon sampling focuses on the genus <i>Pseudapis</i>, a distinct lineage with diverse morphological features, but contentious morphology-based taxonomic classifications and no molecular phylogenetic guidance. We approximate topological accuracy of gene trees by assessing their ability to recover two uncontroversial, monophyletic groups, and compare branch lengths of individual trees using the stemminess metric (the relative length of internal branches). We further examine different strategies of removing uninformative loci and the collapsing of weakly supported nodes into polytomies. We then summarize gene trees with ASTRAL and compare resulting species phylogenies, including comparisons to concatenation-based estimates. Gene trees obtained with the reversible jump model search in MrBayes were most concordant on average and all Bayesian methods yielded gene trees with better stemminess values. The only gene tree estimation approach whose ASTRAL summary trees consistently produced the most likely correct topology, however, was IQ-Tree with automated model designation (MFP). We discuss these findings and provide practical advice on gene tree estimation for summary methods. Lastly, we establish the first phylogeny-informed classification for <i>Pseudapis</i> s. l. and map the distribution of distinct morphological features of the group.</p>
Data from: Pedigree error due to extra-pair reproduction substantially biases estimates of inbreeding depression
Understanding the evolutionary dynamics of inbreeding and inbreeding depression requires unbiased estimation of inbreeding depression across diverse mating systems. However, studies estimating inbreeding depression often measure inbreeding with error, for example, based on pedigree data derived from observed parental behavior that ignore paternity error stemming from multiple mating. Such paternity error causes error in estimated coefficients of inbreeding (f) and reproductive success and could bias estimates of inbreeding depression. We used complete "apparent" pedigree data compiled from observed parental behavior and analogous "actual" pedigree data comprising genetic parentage to quantify effects of paternity error stemming from extra-pair reproduction on estimates of f, reproductive success, and inbreeding depression in free-living song sparrows (Melospiza melodia). Paternity error caused widespread error in estimates of f and male reproductive success, causing inbreeding depression in male and female annual and lifetime reproductive success and juvenile male survival to be substantially underestimated. Conversely, inbreeding depression in adult male survival tended to be overestimated when paternity error was ignored. Pedigree error stemming from extra-pair reproduction therefore caused substantial and divergent bias in estimates of inbreeding depression that could bias tests of evolutionary theories regarding inbreeding and inbreeding depression and their links to variation in mating system.
Data from: Estimation of genotyping error rate from repeat genotyping, unintentional recaptures and known parent-offspring comparisons in 16 microsatellite loci for brown rockfish (Sebastes auriculatus)
Genotyping errors are present in almost all genetic data and can affect biological conclusions of a study, particularly for studies based on individual identification and parentage. Many statistical approaches can incorporate genotyping errors, but usually need accurate estimates of error rates. Here, we used a new microsatellite data set developed for brown rockfish (Sebastes auriculatus) to estimate genotyping error using three approaches: (i) repeat genotyping 5% of samples, (ii) comparing unintentionally recaptured individuals and (iii) Mendelian inheritance error checking for known parent–offspring pairs. In each data set, we quantified genotyping error rate per allele due to allele drop-out and false alleles. Genotyping error rate per locus revealed an average overall genotyping error rate by direct count of 0.3%, 1.5% and 1.7% (0.002, 0.007 and 0.008 per allele error rate) from replicate genotypes, known parent–offspring pairs and unintentionally recaptured individuals, respectively. By direct-count error estimates, the recapture and known parent–offspring data sets revealed an error rate four times greater than estimated using repeat genotypes. There was no evidence of correlation between error rates and locus variability for all three data sets, and errors appeared to occur randomly over loci in the repeat genotypes, but not in recaptures and parent–offspring comparisons. Furthermore, there was no correlation in locus-specific error rates between any two of the three data sets. Our data suggest that repeat genotyping may underestimate true error rates and may not estimate locus-specific error rates accurately. We therefore suggest using methods for error estimation that correspond to the overall aim of the study (e.g. known parent–offspring comparisons in parentage studies).
Final geometries and energies, statistical analysis and estimated errors of single metals and bimetallics for CO2 to methanol conversion
<p>The dataset accommodate all the extra data discussed in:<br>Pisal, P., Krejčí, O. & Rinke, P. Machine learning accelerated descriptor design for catalyst discovery in CO<sub>2</sub> to methanol conversion. <em>npj Comput Mater</em> <strong>11</strong>, 213 (2025). https://doi.org/10.1038/s41524-025-01664-9 </p> <p>The datased contains four types of data:</p> <ol> <li>All the final geometries and energies of adsorbated (*H, *O, *OCHO & *OCH3) and all the 158 single metals and bimetallic alloys on all the surfaces with Miller indices in {-2, -1, ... 2} optimized with Open Catalyst Project (OCP) 20 <em>equiformer_V2</em> machine-learned force-field model. These are in the <a href="https://zenodo.org/api/records/15587232/draft/files/geometries_and_energies.zip/content" target="_blank" rel="noopener noreferrer">geometries_and_energies.zip</a> file organized by the metal/alloys name, with the final geometries and enerigies in a json file, using a json ASE format.</li> <li>All the estimated mean absolute errors (MAE) of predicted adsorption energies for all the considered metals and bimetallic alloys in <a href="https://zenodo.org/api/records/15587232/draft/files/Estimated_MAEs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Estimated_MAEs_metals_bimetallics.csv</a> and xlsx file. The data content is identical, files differs only by a format.</li> <li>All the adsorption energy disctibutions (AEDs) for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> <li>All the statistical information of the adsorption energies for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/Statistics_AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Statistics_AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.