Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,064
datasets available to search
ShareScore release 0.9.0
Dataset results
13,064 results for “Prediction”
Dataset for "Machine learning predictions on an extensive geotechnical dataset of laboratory tests in Austria"
<p>This dataset comprises over 20 years of geotechnical laboratory testing data collected primarily from Vienna, Lower Austria, and Burgenland. It includes 24 features documenting critical soil properties derived from particle size distributions, Atterberg limits, Proctor tests, permeability tests, and direct shear tests. Locations for a subset of samples are provided, enabling spatial analysis.</p> <p>The dataset is a valuable resource for geotechnical research and education, allowing users to explore correlations among soil parameters and develop predictive models. Examples of such correlations include liquidity index with undrained shear strength, particle size distribution with friction angle, and liquid limit and plasticity index with residual friction angle.</p> <p>Python-based exploratory data analysis and machine learning applications have demonstrated the dataset's potential for predictive modeling, achieving moderate accuracy for parameters such as cohesion and friction angle. Its temporal and spatial breadth, combined with repeated testing, enhances its reliability and applicability for benchmarking and validating analytical and computational geotechnical methods.</p> <p>This dataset is intended for researchers, educators, and practitioners in geotechnical engineering. Potential use cases include refining empirical correlations, training machine learning models, and advancing soil mechanics understanding. Users should note that preprocessing steps, such as imputation for missing values and outlier detection, may be necessary for specific applications.</p> <p><strong>Key Features</strong>:</p> <ul> <li><strong>Temporal Coverage</strong>: Over 20 years of data.</li> <li><strong>Geographical Coverage</strong>: Vienna, Lower Austria, and Burgenland.</li> <li><strong>Tests Included</strong>: <ul> <li>Particle Size Distribution</li> <li>Atterberg Limits</li> <li>Proctor Tests</li> <li>Permeability Tests</li> <li>Direct Shear Tests</li> </ul> </li> <li><strong>Number of Variables</strong>: 24</li> <li><strong>Potential Applications</strong>: Correlation analysis, predictive modeling, and geotechnical design.</li> </ul> <p><strong>Technical Details</strong>:</p> <ul> <li>Missing values have been addressed using K-Nearest Neighbors (KNN) imputation, and anomalies identified using Local Outlier Factor (LOF) methods in previous studies.</li> <li>Data normalization and standardization steps are recommended for specific analyses.</li> </ul> <p><strong>Acknowledgments</strong>:<br>The dataset was compiled with support from the European Union's MSCA Staff Exchanges project 101182689 Geotechnical Resilience through Intelligent Design (GRID).</p>
Predicted and experimental chemical and ecotoxicological properties for the toxic unit based hazard assessment
<p><strong>Description</strong></p> <p>This dataset contains ecotoxicity data of 1585 chemicals of environmental concern (CECs) and chemical identifiers. The ecotoxicity data was retrieved from <a href="https://cfpub.epa.gov/ecotox">US EPA ECOTOX Knowlegdebase</a> in ASCII file format and was aggregated for the ecotoxicity groups algae, crustaceans, and fish based on the ideas of <a href="https://dx.doi.org/10.1002/etc.3460">Busch et al. 2016</a>. The dataset includes the 5-percentile, the mean and the geomean of all retrieved ecotoxicity for each compound. Missing ecotoxicity data was estimated with ECOSAR 1.0 algorithms for green algae, daphnids, and fish using <a href="https://www.ufz.de/index.php?en=34593">ChemProp 6.8</a>. The main purpose of this dataset is the <a href="http://doi.org/10.1016/0043-1354(70)90018-7">toxic unit</a> (TU) based hazard assessment of environmental water samples. Chemical properties were estimated using <a href="https://github.com/kmansouri/OPERA">OPERA 2.7</a>, <a href="https://chemaxon.com/products/instant-jchem">Instant JChem</a>, and ACD Percepta 2015 based on QSAR-ready SMILES derived from OPERA 2.7. All data aggregated from EcoTox Knowledgebase (e.g., raw values, species, etc.) is available in the dataset in the detailed sheets. REcoTox, the processing script written in R is available on <a href="https://github.com/tsufz/REcoTox/releases/latest">GitHub</a>.</p> <p><strong>CAUTION</strong></p> <p>It needs to be emphasized that quantitative-structure activity relationship data is just an estimate, which does not necessarily reflect the real property and behaviour of a modelled compound. The calculated data needs to be reviewed in deep. Especially for non-polar or very polar compounds, the QSAR predictions might fail. If a compound ranks high in the TU ranking, it is required to search for literature or regulative data evidences to underpin the finding to avoid false positive prioritizations.</p> <p><strong>RELEASE NOTE</strong></p> <p>Version 210714_v1 was created with <a href="https://github.com/tsufz/REcoTox/releases/tag/v0.1.0">REcoTox version v0.1.0</a>.</p>
Predictive maintenance of the Baghouse [CAO1]
<p> </p> <p>The baghouse filter consists of a collector which removes dust, mainly filler content in dry aggregates during drying process in drum. Baghouse performance is heavily depended on inlet and outlet gas temperature and flow speed as well as opacity climatic conditions and pressure drop, in the bag house (temperature and humidity, the recipe of asphalt). The final user of this asset is the plant operator (EIFFAGE), which is interested in the predictive maintenance of the baghouse.</p> <p>In the context of the developed Cognitive Solution (CS), two prediction models will be developed:</p> <ul> <li>The first model predicts if the baghouse is working properly by attempting to identify any abnormal behavior.</li> <li>The second model predict the remaining useful life of a component of baghouse (days or hours).</li> </ul> <p>The former is utilized for generating a set of alarms based on the process measurements while the latter is utilized for providing an estimation of the saturation level of the filters as well as a value of the evolution of the saturation level of the filters.</p>
Prediction and analysis of phenotypes in the Arabidopsis clock mutant prr7prr9 using the Framework Model v2 (FMv2)
<p>This upload contains or links to the biological data, FMv2 model and simulations for the Chew et al. 2017 paper (bioRxiv <a href="https://doi.org/10.1101/105437">https://doi.org/10.1101/105437</a> ), updated 2022 as bioRxiv <a href="https://doi.org/10.1101/105437v2">https://doi.org/10.1101/105437v2</a>, mostly testing and simulating the effect of a slow circadian clock in the <em>prr7prr9 </em>double mutant compared to the Col wild type plants, with controls in <em>lsf1 </em>and <em>prr7 </em>single mutants. This is one of the outputs from the EU TiMet project, <a href="https://fairdomhub.org/projects/92">https://fairdomhub.org/projects/92</a>.</p> <p>Several data files contain results generated in the same studies, but not covered by the publication. For example, additional time points (18 or 21 days of growth), many additional metabolites, and additional genotypes including <em>pgm</em>, <em>lhy cca1, </em>and in one case, <em>toc1 </em>and <em>gi</em>.</p> <p>This data archive was updated during submisson to the journal _in Silico _Plants in 2022, and is formatted as a Research Object, generated by the Snapshot function of FairdomHub, based on <a href="https://fairdomhub.org/investigations/123">Investigation https://fairdomhub.org/investigations/123.</a> The same Snapshot is shared on FairdomHub and will be from the University of Edinburgh Datashare.</p> <p>We request that users gives appropriate credit to the authors of any data released here, as a norm of academic practice, including data released under CC-0 licence on the FairdomHub.</p>
Data from: Land use, season, and parasitism predict metal concentrations in Australian flying fox fur
<p>There are two .csv files in this upload. The "Pteropus_metal_data_wide.csv" file contains metal concentrations (reported in ng/g) measured in fur samples collected from <em>Pteropus </em>flying foxes (<em>P. alecto</em>, <em>P. conspicillatus</em>, <em>P. poliocephalus</em>). Flying foxes were captured from 2015-2018 at multiple sites across Australia. The file also contains capture information (e.g. date, location) and biological information (e.g. species, sex, age class, parasitism) for each flying fox. The "Pteropus_metadata.csv" file provides further details on all column names in the primary data file, including the specific metals that were quantified. Detailed information on the study methods and results can be found in the associated Science of the Total Environment publication, "Land use, season, and parasitism predict metal concentrations in Australian flying fox fur" by Sánchez et al.</p>
Atmospheric angular momentum predictions
<p>Monthly mean predictions of atmospheric angular momentum for each year from 1960.</p> <p>All predictions start on 1 November.</p> <p>There is one value per month per latitude per year.</p> <p>Summing the values for each latitude gives the global mean.</p>
Ground ice content predictions for the Northern Hemisphere permafrost region at 1-km resolution, version 1.1
<p>Ground ice content is one of the least known characteristics of the permafrost-affected soils in the Northern Hemisphere. At the same time, ground ice content exerts a crucial effect on the thermal response of permafrost to changing climate and environmental conditions, and dictates the permafrost degradation-related geomorphic, hydrologic, and ecological processes, including thermokarst. This dataset presents numerical estimates of volumetric ice content over the permafrost region at a 1-km spatial resolution. The predictions are representative of pore and segregated ice contents in the topmost five meters of permafrost. We use compilations of field measurements of ground ice contents from across the permafrost region to train statistical models and to predict volumetric ice content with the aid of high-resolution geospatial data on climatic, soil and topography conditions. The dataset facilitates assessments of conditions of changing permafrost landscapes at an improved spatial and thematic resolution.</p>
Epidemiological and clinical characteristics predictive of ICU mortality of traumatic brain injury patients treated at a trauma reference hospital – A cohort study - Dataset
<p><strong>Dataset of a cohort whose summary is described below.</strong></p> <p><strong>ABSTRACT</strong></p> <p><strong>Background</strong>: Traumatic brain injury (TBI) has substantial physical, psychological, social and economic impacts, with high rates of morbidity and mortality. Considering its high incidence, the aim of this study was to identify epidemiological and clinical characteristics that predict mortality in patients hospitalized for TBI in intensive care units (ICUs). <strong>Methods</strong>: A retrospective cohort study was carried out with patients over 18 years old with TBI admitted to an ICU of a Brazilian trauma referral hospital between January 2012 and August 2019. TBI was compared with other traumas in terms of clinical characteristics of ICU admission and outcome. Univariate and multivariate analyses were used to estimate the odds ratio for mortality. <strong>Results</strong>: Of the 4816 patients included, 1114 had TBI, with a predominance of males (85.1%). Compared with patients with other traumas, patients with TBI had a lower mean age (45.3 ± 19.1 versus 57.1 ± 24.1 years, p < 0.001), higher median APACHE II (19 versus 15, p <0.001) and SOFA (6 versus 3, p < 0.001) scores, lower median Glasgow Coma Scale (GCS) score (10 versus 15, p < 0.001), higher median length of stay (7 days versus 4 days, p < 0.001) and higher mortality (27.6% versus 13.3%, p < 0.001). In the multivariate analysis, the predictors of mortality were older age (OR: 1.008 [1.002-1.015], p = 0.016), higher APACHE II score (OR: 1.180 [1.155-1.204], p < 0.001), lower GCS score for the first 24 hours (OR: 0.730 [0.700-0.760], p < 0.001), and greater number of brain injuries and presence of associated chest trauma (OR: 1.727 [1.192-2.501], p < 0.001). <strong>Conclusion</strong>: Patients admitted to the ICU for TBI were younger and had worse prognostic scores, longer hospital stays and higher mortality than those admitted to the ICU for other traumas. The independent predictors of mortality were advanced age, APACHE II score, first 24-hour GCS score, number of brain injuries and chest trauma.</p>
UPWARDS - Høgjaeren noise prediction benchmark
<p>The benchmark consists in a reduced layout of nine wind turbines located in the Høgjaeren wind farm in Norway for two wind conditions of similar wind speed amplitude but having opposite wind directions. The necessary information for the user to produce the noise footprint on an observer grid are detailed in the Upwards_D4_6_v1.pdf document.</p>
MHD Model of Ganymede's Magnetosphere: Predicted OCFB and magnetic footprint surface locations for Juno's flyby
<p>This dataset contains model results from a magnetohydrodynamic (MHD) model of Ganymede's magnetosphere adapted to Juno's PJ34 flyby in 2021. Here we publish coordinates for the predicted location of the open-closed-field line-boundary (OCFB) on Ganymede's surface. Additionally we provide coordinates of Juno's magnetic footprint, namely the surface locations that connect to Juno's trajectory through magnetic field lines.</p> <p>For the surface locations we use a western longitude planetographic coordinate system where 0° longitude is in direction of the y-axis and 90° in direction of the x-axis of the cartesian GPhiO system. The GPhiO system is defined by the primary direction<br> z parallel to Jupiter’s rotation axis, the secondary direction y is pointing towards Jupiter barycenter<br> and x completes the right-handed system approximately in direction of plasma flow.</p> <p><strong>Duling2022_JunoGanymede_modeled_surface_OCFB.txt</strong></p> <p>Columns:</p> <p>Longitude [°]<br> Northern OCFB latitude [°]<br> Southern OCFB latitude [°]</p> <p><strong>Duling2022_JunoGanymede_modeled_magnetic_footprint.txt</strong></p> <p>Columns:</p> <p>Spacecraft time [UTC]<br> Magnetic footprint longitude [°]<br> Magnetic footprint latitude [°]<br> Length of field line between Juno and surface [radii]<br> Length of field line between Juno and surface [km]<br> r coordinate of Juno [radii]<br> Latitude of Juno [°]<br> Longitude of Juno [°]<br> x of Juno in GPhiO [km]<br> y of Juno in GPhiO [km]<br> z of Juno in GPhiO [km]</p> <p><strong>Duling2022_JunoGanymede_surface_map.png</strong></p> <p>A plot that visualizes the data of this repository.</p>
Prediction of synaptic activity
<p>This folder contains supplementary material for the paper <a href="https://doi.org/10.1007/s12021-022-09609-z">An Algorithm Based on a Cable‑Nernst Planck Model Predicting Synaptic Activity throughout the Dendritic Arbor with Micron Specificity</a>:</p> <ul> <li>Experimental data: <ul> <li>fluorescence data</li> <li>morphometric data</li> </ul> </li> <li>A notebook to simulate calcium dynamics in the dentritic arbor using <a href="https://joss.theoj.org/papers/10.21105/joss.04012">sinaps</a> software, and the algorithm to predict synaptic activity in fluorescence data</li> <li>Simulation results</li> </ul> <p> </p>
Soil water content (volumetric %) for 33kPa and 1500kPa suctions predicted at 6 standard depths (0, 10, 30, 60, 100 and 200 cm) at 250 m resolution
<p>Soil water content (volumetric) in percent for 33 kPa and 1500 kPa suctions predicted at 6 standard depths (0, 10, 30, 60, 100 and 200 cm) at 250 m resolution. Training points are based on a global compilation of soil profiles (<a href="https://ncsslabdatamart.sc.egov.usda.gov/">USDA NCSS</a>, <a href="https://www.isric.org/projects/africa-soil-profiles-database-afsp">AfSPDB</a>, <a href="https://data.isric.org/geonetwork/srv/eng/catalog.search#/metadata/a351682c-330a-4995-a5a1-57ad160e621c">ISRIC WISE</a>, <a href="http://egrpr.esoil.ru/">EGRPR</a>, <a href="https://esdac.jrc.ec.europa.eu/content/soil-profile-analytical-database-2">SPADE</a>, <a href="https://open.canada.ca/data/en/dataset/6457fad6-b6f5-47a3-9bd1-ad14aea4b9e0">CanNPDB</a>, <a href="https://data.nal.usda.gov/dataset/unsoda-20-unsaturated-soil-hydraulic-database-database-and-program-indirect-methods-estimating-unsaturated-hydraulic-properties">UNSODA</a>, <a href="https://doi.pangaea.de/10.1594/PANGAEA.885492">SWIG</a>, <a href="http://www.cprm.gov.br/en/Hydrology/Research-and-Innovation/HYBRAS-4208.html">HYBRAS</a> and <a href="http://dx.doi.org/10.4228/ZALF.2003.273">HydroS</a>). Data import steps are available <a href="https://gitlab.com/openlandmap/compiled-ess-point-data-sets/-/tree/master/themes/sol/SoilHydroDB"><strong>here</strong></a>. Spatial prediction steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil/soil_water">here</a></strong>. Note: these are actually measured and mapped soil content values; no Pedo-Transfer-Functions have been used (except to fill-in the missing NCSS bulk densities). Available water capacity in mm (derived as a difference between field capacity and wilting point multiplied by layer thickness) per layer is available <strong><a href="https://doi.org/10.5281/zenodo.2629148">here</a></strong>. Antarctica is not included.</p> <p>To access and visualize some of the maps use: <a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>watercontent.33kPa = water content (volumetric percent) under field capacity (33 kPa suction),</li> <li>usda.4b1c = determination method: laboratory method code,</li> <li>m = mean value,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>b10..10cm = vertical reference: 10 cm depth below surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>
Short-term traffic flow prediction based on secondary hybrid decomposition and deep echo state networks
<p>The publication titled "Short-term traffic flow prediction based on secondary hybrid decomposition and deep echo state networks" is supported by the STRIDE K3 project. The dataset used in the publication is uploaded here.</p>
Supplemental Figures for "On the comparative utility of entropic learning versus deep learning for long-range ENSO prediction"
<p>Supplemental figures for the paper "On the comparative utility of entropic learning versus deep learning for long-range ENSO prediction".</p>
TCM: Benchmark Datasets for Predictive Maintenance in Steel Manufacturing
<h1>Anomaly-TCM</h1> <p>Predictive Maintenance (PdM) is a strategy that uses advanced data analytics to predict equipment failures and maintain industrial machinery in good condition. Its goals are to minimize downtime, reduce operational costs, and ensure product quality. PdM methods are applicable across various industries, including steel manufacturing.</p> <p>In steel production, cold rolling is a critical process that reduces the thickness of hot-rolled steel. Developing PdM methods for tandem cold mills (TCM) can significantly improve production efficiency. However, researchers often rely on real manufacturing data, which is typically unavailable, unlabeled, and noisy, making it difficult to validate and compare methods.</p> <p>To overcome this, we created synthetic datasets for the cold rolling process to identify anomalies based on physical principles. These datasets were generated using a mathematical model of a 5-stand TCM, calculating key process parameters like rolling force, torque, speed, tension, gap, thickness reduction, and motor power. We introduced anomalies related to specific failures in the process.</p> <p>We produced six diverse datasets, each with varying complexity, to enable benchmarking of machine learning-based PdM methods for the cold rolling process. Four different types of anomalies were introduced, which are related to a physics-based deviations in the process:</p> <ol> <li>Anomaly in reduction scheme</li> <li>Anomaly in work roll (increased work roll friction)</li> <li>Anomaly in bearing (increased motor torque)</li> <li>Anomaly in electric motor (decrease efficiency)</li> </ol> <p> The details of the datasets are provided below.</p> <table> <tbody> <tr> <td><strong>Dataset</strong></td> <td><strong>Observations</strong></td> <td><strong>Anomalies</strong></td> <td><strong>Share of Anomalies</strong></td> <td><strong>Features</strong></td> <td><strong>Anomaly Types</strong></td> <td><strong>Products</strong></td> <td><strong>Data Drift</strong></td> </tr> <tr> <td>tcm5_dataset_1</td> <td>20009</td> <td>1045</td> <td>5.2%</td> <td>51</td> <td>1</td> <td>4</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_2</td> <td>20001</td> <td>1035</td> <td>5.2%</td> <td>51</td> <td>1</td> <td>20</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_3</td> <td>20003</td> <td>981</td> <td>4.9%</td> <td>51</td> <td>4 (16)</td> <td>4</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_4</td> <td>20001</td> <td>925</td> <td>4.6%</td> <td>51</td> <td>4 (16)</td> <td>20</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_5</td> <td>20005</td> <td>1031</td> <td>5.2%</td> <td>51</td> <td>4 (16)</td> <td>5</td> <td>TRUE</td> </tr> <tr> <td>tcm5_dataset_6</td> <td>20008</td> <td>954</td> <td>4.8%</td> <td>51</td> <td>4 (16)</td> <td>25</td> <td>TRUE</td> </tr> </tbody> </table> <p> </p> <p>Each dataset is generated as a data stream, meaning the observations follow a chronological order, represented by increasing work roll mileage (which is reset after a predefined threshold). The table below provides details about the features and labels present in the datasets. Several features are recorded for each rolling stand, totaling 51 features. Apart from the anomaly related to reduction, the other anomalies are specific to individual stands, resulting in 16 anomaly labels in total.</p> <table> <tbody> <tr> <td><strong>Feature</strong></td> <td><strong>Suffixes</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>thickness_entry</td> <td>-</td> <td>mm</td> <td>steel entry thickness</td> </tr> <tr> <td>thickness_exit</td> <td>-</td> <td>mm</td> <td>steel exit thickness</td> </tr> <tr> <td>width</td> <td>-</td> <td>mm</td> <td>steel width</td> </tr> <tr> <td>ys_entry</td> <td>-</td> <td>MPa</td> <td>steel entry yield strength</td> </tr> <tr> <td>ys_exit</td> <td>-</td> <td>MPa</td> <td>steel exit yield strength</td> </tr> <tr> <td>work_roll_diam</td> <td>1 to 5</td> <td>mm</td> <td>work roll diamaeter (stands 1 to 5)</td> </tr> <tr> <td>work_roll_mileage</td> <td>1 to 5</td> <td>km</td> <td>work roll mileage (stands 1 to 5)</td> </tr> <tr> <td>reduction</td> <td>1 to 5</td> <td>-</td> <td>thickness reduction (stands 1 to 5)</td> </tr> <tr> <td>tension</td> <td>0 to 5</td> <td>N</td> <td>interstand tension (0 is tension before stand 1, 1-5 refer to tension after stands 1-5)</td> </tr> <tr> <td>roll_speed</td> <td>1 to 5</td> <td>NaN</td> <td>linear work roll speed (stands 1 to 5)</td> </tr> <tr> <td>force</td> <td>1 to 5</td> <td>N</td> <td>rolling force (stands 1 to 5)</td> </tr> <tr> <td>torque</td> <td>1 to 5</td> <td>Nm</td> <td>rolling torque (stands 1 to 5)</td> </tr> <tr> <td>gap</td> <td>1 to 5</td> <td>mm</td> <td>stand gap (stands 1 to 5)</td> </tr> <tr> <td>motor_power</td> <td>1 to 5</td> <td>kW</td> <td>electric motor power (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_Reduction</td> <td>-</td> <td>-</td> <td>(label) anomaly in reduction scheme</td> </tr> <tr> <td>Anomaly_Electric</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in electric motor (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_Bearing</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in stand bearing (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_WorkRoll</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in work roll friction (stands 1 to 5)</td> </tr> </tbody> </table>
Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory: Figures
<p>This entry contains the sources for the figures included in the body of the paper titled "Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory", by Yassine Bouchafra, Avijit Shee, Florent Réal, Valérie Vallet and André Severo Pereira Gomes, as well as those found in the supplementary information.</p> <p>It accompanies the dataset found at the DOI: 10.5281/zenodo.1477004</p> <p> </p> <p> </p>
Predicted USDA soil orders at 250 m (probabilities)
<p>Distribution of the USDA orders (12) based on machine learning predictions of great groups (<a href="https://doi.org/10.5281/zenodo.1476844">https://doi.org/10.5281/zenodo.1476844</a>) from global compilation of soil profiles. To learn more about soil orders and great groups please refer to the <a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil/tree/master/soil">here</a></strong>. Antartica is not included.</p> <p>To access and visualize maps use: <a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>order = variable: USDA order,</li> <li>usda.histosols = determination method: USDA soil taxonomy class Histosols,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>
Predicted USDA soil suborders at 250 m (probabilities)
<p>Distribution of the USDA suborders based on machine learning predictions of great groups (<a href="https://doi.org/10.5281/zenodo.1476844">https://doi.org/10.5281/zenodo.1476844</a>) from global compilation of soil profiles. To learn more about soil suborders and great groups please refer to the <a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil">here</a></strong>. Antartica is not included.</p> <p>To access and visualize maps use: <a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/-/issues">https://gitlab.com/openlandmap/global-layers/-/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>suborder = variable: USDA suborder,</li> <li>usda.ustolls = determination method: USDA soil taxonomy class Ustolls,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>
Predicted USDA soil great groups at 250 m (probabilities)
<p>Distribution of the USDA soil great groups based on machine learning predictions from global compilation of soil profiles (>350,000 training points). To learn more about soil great groups please refer to the <a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil">here</a></strong>. Antarctica is not included.</p> <p>To access and visualize maps use: <a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>A back-up copy of all predictions (>65GB) can be downloaded from: <a href="http://gofile.me/6J25n/mQ3cHOOMr">http://gofile.me/6J25n/mQ3cHOOMr</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>grtgroup = variable: USDA great group,</li> <li>usda.argiustolls = determination method: USDA soil taxonomy class Argiustolls,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.2 = version number: 0.2,</li> </ul>
Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis
<p>The attached two datasets are the optimized inputs used to analyze predictability limits in the paper Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis. Specifically, the datasets correspond to the inputs used to produce the blue (global) and green (regional) loss curves in Figure S2. They are NetCDF files of dimensions batch (1), time (2), latitude (181), longitude (360), pressure levels (13), and may be run as Graphcast model inputs to initiate a forecast at 00 UTC 20 June 2021. Both datasets have been systematically perturbed to reduce the Graphcast model's loss function, which minimizes forecast eror as described in the manuscript. The global input seeks to reduce the loss over the entire globe, while the regional input seeks only to minimize error within the Pacific Northwest (42N to 60N and 130W to 110W). The optimized inputs result in a reduction of the loss by approximately 85% (global) and 93% (regional) when compared to a control Graphcast forecast without perturbations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.