Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,773 results for “Prediction Models”

Learn how ShareScore rates datasets ↗
edi64/100

Spartina alterniflora above- and belowground biomass predictions and inundation intensity as estimated by the Belowground Ecosystem Resiliency Model for U.S. Georgia marshes from 2014 to 2023.

We applied the Belowground Ecosystem Resiliency Model (BERM) to estimate monthly aboveground biomass (AGB) and belowground biomass (BGB) in U.S. Georgia Spartina alterniflora marshes from 2014 to 2023 at 30 m scale. This application involved BERM version 2.0 (https://doi.org/10.5281/zenodo.13306821), which was built using data in the PLT-GCET-2308 dataset (https://dx.doi.org/10.6073/pasta/4a0b715104849d98320fcc34e7cd63a4). Data sources for BERM application included Landsat-8/9, NOAA CO-OPS Station ID: 8670870, Daymet, and USGS 3DEP 2018 DEM. Download and processing steps are described in the BERM code and in metadata methods section. Specific descriptions of data processing are available in model code: https://doi.org/10.5281/zenodo.13306821. Data provided here include model output of AGB estimates, BGB estimates, and calculated inundation intensity. See "Data reporting" method in the metadata for description of data files. For logisitical purposes here we present only select data from the model input and output. All model input data sources as listed in the abstract are publicly available. Model calibration data and code are published as well. Additional predictions not published here include foliar chlorophyll, foliar nitrogen, and leaf area index.

openCC (other)Dec 2024View details →
edi52/100

Code for Random Forest models that predict pharmaceutical and water chemistry measurements in Baltimore Ecosystem Study streams

This file contains code to model the relationship between the water chemistry measurements and discharge measured as part of BES routine sampling and the pharmaceuticals measured in WY 2018. We use Random Forest models to predict 1) total (i.e., summed) concentration of the pharmaceuticals for which we screened, 2) total nutrient concentrations (TN & TP), 3) whether or not the antibiotic trimethoprim was detected in a given sample, and 4) whether or not nitrate and TP were above or below environmentally-relevant threshold concentrations. We also use RF models to predict N and P concentrations over a longer period, in order to compare models for nutrients to pharma. Code and analyses here rely on data processed in the file "BESPharma_WY2018.Rmd", published on EDI (doi:10.6073/pasta/610cb67fcbc8982c2af8ed946dce8ea5) and BES water chemistry data published on EDI (doi:10.6073/pasta/ce7f30e6013e003bfe28c5fd7d4aed23 )

openCC0Feb 2025View details →
edi52/100

SBC LTER: Daily averages of modeled significant wave height (Hs) and peak wave period (Tp) in the Santa Barbara Coastal area from the Coastal Data Information Program - Monitoring and Prediction System (CDIP MOP)

From http://cdip.ucsb.edu: The Coastal Data Information Program (CDIP) is a research group at Scripps Institution of Oceanography that monitors coastal waves and nearshore sand levels on regional scales. CDIP maintains a network of optimally-placed, directional wave buoys from San Diego to Eureka. The buoy measurements are used to initialize a high spatial resolution (100m x 100m) linear spectral wave propagation model. The resulting hourly hindcasts and nowcasts of CA coastal wave conditions have a level of accuracy that is not possible with more traditional wind-wave generation models that are initialized with modeled wind fields.

openCC (other)Jun 2025View details →
zenodo48/100

Datasets from study: "Land surface observations boost temperature forecast skill: experiments using Long Short-Term Memory surrogate for physics-based models to assess potential predictability"

<p>This repository contains the datasets needed to reproduce the figures from manuscript: Land surface observations boost temperature forecast skill: experiments using Long Short-Term Memory surrogate for physics-based models to&nbsp;assess potential predictability.</p> <p>In this study, we examine the potential of land surface temperature and vegetation data, which are not routinely assimilated in NWP models, for enhancing temperature forecast skill. We build surrogate models for NWP using Long Short-Term Memory.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Prediction and analysis of phenotypes in the Arabidopsis clock mutant prr7prr9 using the Framework Model v2 (FMv2)

<p>This upload contains or links to the biological data, FMv2 model and simulations for the Chew et al. 2017 paper (bioRxiv <a href="https://doi.org/10.1101/105437">https://doi.org/10.1101/105437</a> ), updated 2022 as bioRxiv <a href="https://doi.org/10.1101/105437v2">https://doi.org/10.1101/105437v2</a>, mostly testing and simulating the effect of a slow circadian clock in the <em>prr7prr9 </em>double mutant compared to the Col wild type plants, with controls in <em>lsf1 </em>and <em>prr7 </em>single mutants. This is one of the outputs from the EU TiMet project, <a href="https://fairdomhub.org/projects/92">https://fairdomhub.org/projects/92</a>.</p> <p>Several data files contain results generated in the same studies, but not covered by the publication. For example, additional time points (18 or 21 days of growth), many additional metabolites, and additional genotypes including <em>pgm</em>, <em>lhy cca1, </em>and in one case, <em>toc1 </em>and <em>gi</em>.</p> <p>This data archive was updated during submisson to the journal _in Silico _Plants in 2022, and is formatted as a Research Object, generated by the Snapshot function of FairdomHub, based on&nbsp;<a href="https://fairdomhub.org/investigations/123">Investigation https://fairdomhub.org/investigations/123.</a> The same Snapshot is shared on FairdomHub and will be from the University of Edinburgh Datashare.</p> <p>We request that users gives appropriate credit to the authors of any data released here, as a norm of academic practice, including data released under CC-0 licence on the FairdomHub.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

MHD Model of Ganymede's Magnetosphere: Predicted OCFB and magnetic footprint surface locations for Juno's flyby

<p>This dataset contains model results from a magnetohydrodynamic (MHD) model of Ganymede&#39;s magnetosphere adapted to Juno&#39;s PJ34 flyby in 2021. Here we publish coordinates for the predicted location of the open-closed-field line-boundary (OCFB) on Ganymede&#39;s surface.&nbsp;Additionally we provide coordinates of Juno&#39;s magnetic footprint, namely the surface locations that connect to Juno&#39;s trajectory through magnetic field lines.</p> <p>For the surface locations we use a western longitude planetographic coordinate system where 0&deg; longitude is in direction of the y-axis and 90&deg; in direction of the x-axis of the cartesian GPhiO system.&nbsp;The GPhiO system is defined by the&nbsp;primary direction<br> z&nbsp;parallel to Jupiter&rsquo;s rotation axis, the secondary direction y is pointing towards Jupiter barycenter<br> and x completes the right-handed system approximately in direction of plasma flow.</p> <p><strong>Duling2022_JunoGanymede_modeled_surface_OCFB.txt</strong></p> <p>Columns:</p> <p>Longitude [&deg;]<br> Northern OCFB latitude [&deg;]<br> Southern OCFB latitude [&deg;]</p> <p><strong>Duling2022_JunoGanymede_modeled_magnetic_footprint.txt</strong></p> <p>Columns:</p> <p>Spacecraft time [UTC]<br> Magnetic footprint longitude [&deg;]<br> Magnetic footprint latitude [&deg;]<br> Length of field line between Juno and surface [radii]<br> Length of field line between Juno and surface [km]<br> r coordinate of Juno [radii]<br> Latitude of Juno [&deg;]<br> Longitude of Juno [&deg;]<br> x of Juno in GPhiO [km]<br> y of Juno in GPhiO [km]<br> z of Juno in GPhiO [km]</p> <p><strong>Duling2022_JunoGanymede_surface_map.png</strong></p> <p>A plot that visualizes the data of this repository.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Perturbative gravitational wave predictions for the real scalar extended Standard Model, dataset

<p>This deposit contains data from a perturbative study of cosmological phase transitions in the real singlet scalar extension of the Standard Model (xSM). The data relates to the paper "Perturbative gravitational wave predictions for the real scalar extended Standard Model". Everything is contained within the archive file <em>xsm_results.tar.gz</em>, a tarball compressed with Gzip.</p> <p>The data covers phase transition properties for a scan of 100,000 parameter points in the xSM. Further details on the contents of the dataset are explained in the <em>README.md</em> within the tarball.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data

<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed &ldquo;similarity-based merger models&rdquo; which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC&gt;0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

MHD Model of Ganymede's Magnetosphere: Predicted magnetic field on Juno's trajectory

<p>This dataset contains model results from a magnetohydrodynamic (MHD) model of Ganymede&#39;s magnetosphere adapted to Juno&#39;s PJ34 flyby in 2021. Here we publish predicted magnetic field components on Juno&#39;s trajectory that can be compared to MAG measurements and are displayed in Figure 3 of Duling et al. (2022).</p> <p>Each file contains data from one model. The dataset includes all models with parameter variations from Duling et al. (2022). These are summarized in Table 1 of Duling et al. (2022) and displayed in Figure 3 with the gray lines.</p> <p>If not varied, all models are run with the following parameters:</p> <p>Upstream Jovian background magnetic field B<sub>0&nbsp;</sub>= (&minus;15,24,&minus;75) nT<br> Upstream plasma velocity v<sub>0</sub>&nbsp;= 140 km/s<br> Upstream plasma mass density <span class="math-tex">\(\rho\)</span><sub>0</sub>&nbsp;=&nbsp;100 amu/cm<sup>3</sup><br> Upstream plasma thermal pressure p<sub>0</sub> = 2.8 nPa<br> Ionization frequency&nbsp;<span class="math-tex">\(\nu_{ion}\)</span>&nbsp;= 2.2e-8/s<br> Atmospheric surface mass density&nbsp;<span class="math-tex">\(n_{n,0}\)</span>&nbsp;=&nbsp;&nbsp;8e6/cm<sup>3</sup><br> Dipole Gauss coefficient&nbsp;<span class="math-tex">\(g_1^0\)</span>&nbsp;= &minus;716.8 nT</p> <p>&nbsp;</p> <p>The published data files correspond to the following models with each one parameter variation:</p> <table> <thead> <tr> <th scope="col">Parameter</th> <th scope="col">Value</th> <th scope="col">Filename Suffix</th> </tr> </thead> <tbody> <tr> <td>default model</td> <td>&nbsp;-&nbsp;</td> <td>default</td> </tr> <tr> <td>Upstream Jovian background magnetic field (measured before flyby)</td> <td>B<sub>0&nbsp;</sub>= (&minus;16,3,&minus;70) nT</td> <td>B0before</td> </tr> <tr> <td>Upstream Jovian background magnetic field (measured after flyby)</td> <td>B<sub>0&nbsp;</sub>= &nbsp;(&minus;14,43,&minus;80) nT</td> <td>B0after</td> </tr> <tr> <td>Upstream plasma velocity (min)</td> <td>v<sub>0</sub>&nbsp;= 120 km/s</td> <td>v-</td> </tr> <tr> <td>Upstream plasma velocity (max)</td> <td>v<sub>0</sub>&nbsp;= 160 km/s</td> <td>v+</td> </tr> <tr> <td>Upstream plasma mass density (min)</td> <td><span class="math-tex">\(\rho\)</span><sub>0</sub>&nbsp;=&nbsp;10 amu/cm<sup>3</sup></td> <td>rho-</td> </tr> <tr> <td>Upstream plasma mass density (max)</td> <td><span class="math-tex">\(\rho\)</span><sub>0</sub>&nbsp;=&nbsp;160 amu/cm<sup>3</sup></td> <td>rho+</td> </tr> <tr> <td>Upstream plasma thermal pressure (min)</td> <td>p<sub>0</sub> = 1.0 nPa</td> <td>p-</td> </tr> <tr> <td>Upstream plasma thermal pressure (max)</td> <td>p<sub>0</sub> = 5.0 nPa</td> <td>p+</td> </tr> <tr> <td>Ionization frequency (min)</td> <td>&nbsp;<span class="math-tex">\(\nu_{ion}\)</span>&nbsp;= 0.5e-8/s</td> <td>prod-</td> </tr> <tr> <td>Ionization frequency (max)</td> <td>&nbsp;<span class="math-tex">\(\nu_{ion}\)</span>&nbsp;= 10.0e-8/s</td> <td>prod+</td> </tr> <tr> <td>Atmospheric surface mass density (min)</td> <td>&nbsp;<span class="math-tex">\(n_{n,0}\)</span>&nbsp;=&nbsp; 1.6e6/cm<sup>3</sup></td> <td>nn-</td> </tr> <tr> <td>Atmospheric surface mass density (max)</td> <td>&nbsp;<span class="math-tex">\(n_{n,0}\)</span>&nbsp;=&nbsp; 40e6/cm<sup>3</sup></td> <td>nn+</td> </tr> <tr> <td>Dipole Gauss coefficient (min)</td> <td>&nbsp;<span class="math-tex">\(g_1^0\)</span>&nbsp;= &minus;702.5 nT</td> <td>dipole-</td> </tr> <tr> <td>Dipole Gauss coefficient (max)</td> <td>&nbsp;<span class="math-tex">\(g_1^0\)</span>&nbsp;= &minus;731.1 nT</td> <td>dipole+</td> </tr> </tbody> </table> <p>Magnetic Field components and Juno&#39;s position are in&nbsp;GPhiO system. GPhiO is defined by the&nbsp;primary direction z&nbsp;parallel to Jupiter&rsquo;s rotation axis, the secondary direction y is pointing from Ganymede&#39;s&nbsp;towards Jupiter&#39;s barycenter and x completes the right-handed system approximately in direction of plasma flow.</p> <p>Columns:</p> <p>Spacecraft time [UTC]<br> Bx modeled magnetic field in GPhiO [nT]<br> By&nbsp;modeled magnetic field in GPhiO [nT]<br> Bz&nbsp;modeled magnetic field in GPhiO [nT]<br> B&nbsp;modeled magnetic field magnitude&nbsp;[nT]<br> x of Juno in GPhiO [km]<br> y of Juno in GPhiO [km]<br> z of Juno in GPhiO [km]</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Dataset for "Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns"

<p>This archive contains the source code and data sets presented in the publication "Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns".</p> <p>Kabus, D., De Coster, T., de Vries, A. A., Pijnappels, D. A., &amp; Dierckx, H. (2024). Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns.&nbsp;<em>Computers in Biology and Medicine</em>, 107949. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.compbiomed.2024.107949" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.compbiomed.2024.107949</span></a></p>

opencc-by-4.0Jul 2023View details →
edi48/100

Lake chloride concentrations and model predictions for 49,432 lakes in the Midwest and Northeast United States.

Lakes in the Midwest and Northeast United States are at risk of anthropogenic chloride contamination, but we have little knowledge of the prevalence and spatial distribution of the problem. The majority of salt pollution in north temperate regions stems from road salt application but other chloride sources include water softeners, synthetic fertilizers, and livestock excretion. Although chloride contamination of lakes is well documented, it is unknown how many lakes are at risk of long-term salinization. We used a quantile regression forest to leverage information from 2,773 lakes to predict the chloride concentration of all 49,432 lakes greater than 4 ha in a 17-state area. The QRF used 22 predictor variables, which included lake morphometry characteristics, watershed land use, and distance to the nearest interstate and road. Model predictions had an r2 of 0.94 for all chloride observations, and 0.87 for predictions of the mean chloride concentration observed at each lake.

openCC (other)Apr 2020View details →
edi48/100

Arthropod biomass captured by sweepnet (weekly) and sweepnet biomass model predictions (daily) near Toolik Field Station, Alaska, summers 2012-2016

This data set contains information about the per sample sweepnet arthropod biomass captured (or modeled using GAM modelling approaches) near Toolik Field Station from 2012 to 2016 under National Science Foundation (NSF) Office of Polar Programs ARC 0908444 (to Laura Gough), ARC 0908602 (to Natalie Boelman), and ARC 0909133 (to John Wingfield). It is associated with publication DOI: 10.1111/jav.01712.

openCC (other)Jan 2020View details →
zenodo44/100

RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)

<p>RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

DWCox: A Density-Weighted Cox Model for Outlier-Robust Prediction of Prostate Cancer Survival

<p>This package, <strong>DWCox</strong>, implements a <strong>d</strong>ensity-<strong>w</strong>eighted <strong>Cox</strong> regression model that is more robust against outliers in the training data. DWCox gives more accurate predictions than the standard Cox regression on prostate cancer survival, especially in cases where the training data are expected to contain a lot of outliers. More details can be found in our paper (coming soon) and the README file inside this package.</p>

openmit-licenseNov 2016View details →
zenodo44/100

Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.

<p>This repository contains input and ouput files&nbsp;used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives:&nbsp;</p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;-&gt; DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex&nbsp; &nbsp;-&gt; DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures&nbsp;&nbsp;&nbsp; -&gt; DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i>&lt;index&gt;_&lt;pdbcode&gt;</i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. &nbsp;</p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -&gt; FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain &nbsp; &nbsp; &nbsp; &nbsp;-&gt; DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -&gt; DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 &nbsp; -&gt; DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -&gt; DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models &nbsp; -&gt; DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Stiffness Moduli Modelling and Prediction in Four-Point Bending of Asphalt Mixtures: A Machine Learning-Based Framework within Weave-UNISONO 2021 project, NCN project No 2021/03/Y/ST8/00079, and GACR project GA22-04047K

<div><strong>Summary:</strong></div> <div>Two selected mixtures were thoroughly investigated in an experimental trial carried out by means of a four-point bending test (4PBT) apparatus. The mixtures were prepared using spilite aggregate, a conventional 50/70 penetration grade bitumen, and limestone filler. Their stiffness moduli (SM) were determined while samples were exposed to 11 loading frequencies (from 0.1 to 50 Hz) and 4 testing temperatures (from 0 to 30 &deg;C). Observations were recorded and used to develop a machine learning (ML) model. The main scope was the prediction of the stiffness moduli based on the volumetric properties and testing conditions of the corresponding mixtures, which would provide the advantage of reducing the laboratory efforts required to determine them.</div> <div>&nbsp;</div> <div><strong>The dataset includes:</strong></div> <div>Characteristics of bituminous binder, CSV raw data</div> <div> <ul> <li>bituminous binder.csv</li> </ul> </div> <div>Grading curves of tested asphalt mixtures</div> <ul> <li>AML16 Grading curves.csv</li> <li>AMP22 Grading curves.csv</li> </ul> <div>Volumetric characterizations of AML16 and AMP22 mixtures</div> <ul> <li>AML16 Volumetric characterizations.csv</li> <li>AMP22 Volumetric characterizations.csv</li> </ul> <div>Outcomes of the 4PBT experimental trial carried out on AML16 and AMP22 mixtures</div> <ul> <li>AML16 Stiffness Modulus 4PB.csv</li> <li>AMP22 Stiffness Modulus 4PB.csv</li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Data to Support Predictive Models for Detrital Titanite Provenance with application to the Nanga Parbat syntaxial massif, western Himalaya."

<p>The files published here are metadata that are being used to support a manuscript currently (Mar, 2024) undergoing final reviews in Journal of Geophysical Research: Earth Surface.</p> <p>The intention of these data and code is to support a publication that is about generating a predictive categorisation scheme for the mineral titanite.</p> <p>The code to generate the titanite classification schemes was created in Python3, using Jupyter Notebook. The files also provide more motivation for why a predictive categorisation scheme for the mineral titanite is desirable, and other similar context. Chiefly, the dataset and random forest models published here will allow us to trace titanite in detritus.</p> <p>For info on running Jupyter Notebook, please visit (<a href="https://jupyter-notebook-beginner-guide.readthedocs.io/en/latest/execute.html">https://jupyter-notebook-beginner-guide.readthedocs.io/en/latest/execute.html</a>) to seek instructions. We also provide a readme file with some instructions. If you get really stuck, just email the authors.</p> <p>Our Model can be compared to similar previously published works (e.g.&nbsp;<a href="https://doi.org/10.1111/ter.12574">https://doi.org/10.1111/ter.12574</a>). Model was trained using skikit-learn v1.41.</p> <p>The supplementary file "Table_S4_Merged.csv" was used to train and generate the model.</p> <p>Your unknowns must contain the correct elements and labelling for the code to successfully run, these details are provided in the code (Titanite_Random_Forest_Model1_Mar24.ipynb). A template is also provided for you to paste your unknown data into (titanite_data_template.csv)</p> <p>Any new published data are titanite compositional or isotopic data collected by LA-ICP-MS. Description of how those data were collected is given in "OSullivan_et_al_Supp..." file.</p> <p>Some of the data, information and code in this submission has been subject to change after journal review, this is a second version of this content.</p> <p>References for the dataset compilation are provided in File S3.</p> <p>If you have any queries contact:<br>Gary O'Sullivan, Trinity College Dublin</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities

<p>Numerical data used to make Figures in the manuscript entitled "Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities". We provide the data to create Figures 1 to 4 from the main text and Figures S1 to S9 from the supplementary information. We also provide Python scripts to plot them.</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Datasets for Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes

<p>All used Datasets to pair with &quot;Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes&quot; by Nicholas Ouassil*, Rebecca L. Pinals*, Jackson Travis Del Bonis-O&#39;Donnell, Jeffrey W. Wang, and Markita P. Landry</p> <p>*Co-authors</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record