Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
53
datasets available to search
ShareScore release 0.9.0
Dataset results
53 results for “data driven Approach”
Global sea-surface DMS concentrations estimated through data-driven approaches
<p>DMS_ML_average.nc and DMS_SAT-OPT_average.nc represent the 10-year (2011-2020) average monthly seawater DMS concentrations (1°x1°) calculated using two different methods, ML and SAT-OPT, respectively.</p>
Phase field modelling combined with data-driven approach to unravel the orientation influenced growth of interfacial Cu6Sn5 intermetallics under electric current stressing
<p><strong>Description:</strong></p> <p>The datasets are constituted by two folders, namely, (A) data_features_and_metric.zip and (B) grain_area_prediction.zip. </p> <p><strong>(A) data_features_and_metric.zip:</strong></p> <p>The following are the contents of this folder</p> <p>(i) <em>grainTheta.csv file</em> : The "grainTheta.csv" file consists the datasets generated from multiple phase field simulations. Name of the columns in the csv file are:</p> <p> <strong>gnid </strong>= grain id number "n", <strong>ntheta = </strong>orientation angle of n<sup>th</sup> grain (<sup>o</sup>); <strong>nltheta </strong>= orientation angle of grain to the left of n<sup>th</sup> grain (<sup>o</sup>); <strong>nrtheta </strong>= orientation angle of grain to the right of n<sup>th</sup> grain (<sup>o</sup>); <strong>j </strong>= current density (A/m<sup>2</sup>);<strong> t =</strong> time (s); <strong>area</strong> = area of n<sup>th</sup> grain (m<sup>2</sup>);<strong> tl</strong> = horizontal length of the top edge of grain "n" (m) ; <strong>bl </strong>= horizontal length of the bottom edge of grain "n" (m) </p> <p>The features gnid, ntheta, nltheta and nrtheta for a given observation are determined during the design of initial conditions of the corresponding phase field simulation. The value of "j" for the observation is determined via the boundary condition in the same numerical simulation. The result from the finite element method based phase field simulation has provided the numerical quantities for t, area, tl and bl attributes. The multiple observations in the data file have been obtained from multiple phase field simulations. </p> <p>(ii) <em>imc_theta.ipynb, imc_theta.py and imc_theta.html files</em>: These files contain the code to build the Pearson's Correlation Coefficient (PCC) heatmap analysis of the data contained in grainTheta.csv file. </p> <p>(iii) <em>comparison_mse.csv</em>: This data file includes the information about mean square error for training data (tmse) and mean square error for validation data (vmse) at Epoch = 199 resulting from 10 different artificial neural network (ANN) models distinguished by 10 different values of learning rates (lr) . Thus, the name of the columns in this csv file are <strong>modelno</strong>, <strong>lr</strong>, <strong>tmse</strong> and <strong>vmse</strong>. </p> <p>(iv) <em>mse_comparison.gnu</em>: This file consists the codes required to output a png image from the data provided in comparison_mse.csv<em>. </em></p> <p>(v) train_loss.csv and val_loss.csv: These files consist of the data of tmse and vmse at all points of Epochs for the ANN model with lr = 2.5E-4 . Thus, the first column in train_loss.csv file corresponds to tmse whereas the second column is Epochs number. Similarly, vmse and Epochs represent the two columns in val_loss.csv file. </p> <p> </p> <p>(vi) <em>mse_lr2p5e-4.gnu</em> : This file consists the codes required to output a png image from the data provided in train_loss.csv and val_loss.csv<em>. </em></p> <p><strong>(B) grain_area_prediction.zip:</strong></p> <p>Inside this folder, there is a folder named "prediction_of_grain_area" consisting of the following files:</p> <p><em>initial_area.csv file</em>: This file consists the value of the initial grain area of grain 4. It is a constant at all orientation angle.</p> <p><em>predicted_result_00_5e4.csv</em>: This file consists of the prediction result of grain 4 area (at different orientation angles and t = 1250 s) for grain 3 and grain 5 at orientation angles of 0<sup>o</sup> and 0<sup>o </sup>respectively, and for applied current density of 5.0E+4 J/m<sup>2</sup> .</p> <p><em>predicted_result_00_5e5.csv</em>: This file consists of the prediction result of grain 4 area (at different orientation angles and t = 1250 s) for grain 3 and grain 5 at orientation angles of 0<sup>o</sup> and 0<sup>o </sup>respectively, and for applied current density of 5.0E+5 J/m<sup>2</sup> .</p> <p><em>predicted_result_9090_5e4.csv</em>: This file consists of the prediction result of grain 4 area (at different orientation angles and t = 1250 s) for grain 3 and grain 5 at orientation angles of 90<sup>o</sup> and 90<sup>o </sup>respectively, and for applied current density of 5.0E+4 J/m<sup>2</sup> .</p> <p><em>predicted_result_9090_5e5.csv</em>: This file consists of the prediction result of grain 4 area (at different orientation angles and t = 1250 s) for grain 3 and grain 5 at orientation angles of 90<sup>o</sup> and 90<sup>o </sup>respectively, and for applied current density of 5.0E+5 J/m<sup>2</sup> .</p> <p><em>area_00_adj.gnu</em> : This gnu file contains the code to produce the png image from the data contained in <em>predicted_result_00_5e4.csv </em>and<em> predicted_result_00_5e5.csv </em>. The information about the initial area of grain 4 is obtained from <em>initial_area.csv</em> file by the code.</p> <p><em>area_9090_adj.gnu</em> : This gnu file contains the code to produce the png image from the data contained in <em>predicted_result_9090_5e4.csv </em>and<em> predicted_result_9090_5e5.csv </em>. The information about the initial area of grain 4 is obtained from <em>initial_area.csv</em> file by the code.</p> <p> </p>
Dataset of A User-driven Hybrid Neuro-symbolic Approach for Knowledge Graph Creation from Relational Data
<p>This dataset contains the following:</p> <p>1. achieved percentage values of the generated RML rules using LXS and manually</p> <p>2. basic information about the example used and with which creation type users started</p> <p>3. all answers of users to the User Experience Questionnaire</p> <p>4. Answers to the structured part of the user interview</p>
Data from: A data-driven approach to establishing cell motility patterns as predictors of macrophage subtypes and their relation to cell morphology
Open the record for dataset details and reuse information.
A data driven network approach to rank countries production diversity and food specialization
<p>This file explains the database, that can be freely downloaded, used in the work A data driven network approach to rank food complexity and countries production diversity?(arXiv:1606.01270v1) by Chengyi Tu, Joel Carr and Samir Suweis. If you use the data, please cite the article. The file OrganizeData.xlsx includes bipartite adjacency matrix, both binary (b) and weighted (w) in tons of food, describing the food commodities produced (country-product), imported (country-import) and exported (country-export) by the considered countries from the year 1992 to the year 2011. GDP data are easily available from UN (http://data.un.org/Data.aspx?d=WDI&f=Indicator_Code%3ANY.GDP.MKTP.CD) or worldbank (http://data.worldbank.org/indicator/NY.GDP.MKTP.CD) website. From these data the results on the Minimum Spanning Forest and on the country fitness and food specialization can be reproduced.</p>
Delving into the Impact of Stress on Mental Health: A Data-Driven Approach
<p>This comprehensive dataset delves into the intricate relationship between stress and mental health, providing valuable insights for researchers, healthcare professionals, and individuals alike. The dataset encompasses a diverse range of variables, including self-reported stress levels, standardized mental health assessments, demographic information, and lifestyle factors. This rich data empowers researchers to conduct in-depth analyses and uncover the multifaceted connections between stress and mental Health.</p>
Figures for Developing an online data-driven approach for prognostics and health management of lithium-ion batteries
Open the record for dataset details and reuse information.
Data-driven approaches to bestow environmental management through linking wastewater data to source estimation of hazardous waste [Data]
<p>Data of article <em>Data-driven approaches to bestow environmental management through linking wastewater data to source estimation of hazardous waste</em></p>
Data-driven approaches to bestow environmental management through linking wastewater data to source estimation of hazardous waste
<p>Data of article <em>Data-driven approaches to bestow environmental management through linking wastewater data to source estimation of hazardous waste</em></p>
An Integrated Approach for enhanced SMAP Soil Moisture Retrieval: Multi-Source Data Fusion and Data-Driven Machine Learning
<p><span>Accurate satellite-based soil moisture (SM) retrieval is essential for hydrometeorological and agroecological applications, yet traditional physical models for L-band SM retrieval are hindered by uncertainties stemming from inaccuracies in prior parameters. This work combines multi-source data fusion and a physically-guided machine learning framework to develop a Soil Moisture Active Passive (SMAP) SM retrieval model (Fusion-LightGBM, F-LGB) that bypasses the need for static prior parameters, resulting in a new SM product. The retrieval benchmark is a new seamless SM data constructed by combining Triple Collection correlation coefficients (TC-R) and the Maximized-R method, which demonstrates superior temporal correlation on 20 International Soil Moisture Network (ISMN) <em>in-situ</em> networks compared to existing SM data, including ECMWF Reanalysis v5-Land (ERA5-Land), SMAP Level 4 (SMAP L4), and Global Land Data Assimilation System (GLDAS) Noah. The machine learning model incorporates input variables that represent the Tau-Omega model’s radiative transfer process, including brightness temperature, vegetation optical depth, soil temperature, and an external variable for precipitation. In the 2015-2020 validation set, F-LGB demonstrated the highest correlation (mean R = 0.72, significantly surpassing the second-best SMAP-INRAE-BORDEAUX (SMAP-IB) SM and deep neural network (DNN) SM at 0.67) and the lowest ubRMSE (mean value of 0.052 m<sup>3</sup>/m<sup>3</sup>, better than 0.055 m<sup>3</sup>/m<sup>3</sup> for both DNN and SMAP-IB). F-LGB performed well across diverse land covers, vegetation densities, and climates, with SHAP analysis showing H-polarized brightness temperature as crucial, especially in areas with low to moderate vegetation. This new machine learning-based SMAP SM product may improve global satellite-based SM estimation capabilities.</span></p>
Dataset for paper Pavel Perezhogin, Laure Zanna, Carlos Fernandez-Granda "Generative data-driven approaches for stochastic subgrid parameterizations in an idealized ocean model" submitted to JAMES.
<p>The dataset consists of the directory tree of .zarr archives. See <a href="https://github.com/m2lines/pyqg_generative/blob/master/Google-Colab/dataset.ipynb">Github repository</a> for the description of the dataset.</p> <p>The directory tree is:</p> <pre><code>├── eddy │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires ├── jet │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires</code></pre> <ul> <li>Every individual dataset is a <code>.zarr</code> <a href="https://zarr.readthedocs.io/en/stable/">archive</a></li> <li><code>eddy/jet</code> - configuration of the pyqg; eddy is default; See <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2022MS003258">Ross2022</a> for description</li> <li><code>hires.zarr</code> - high-resolution simulation at 256x256 grid</li> <li><code>48/64/96</code> - resolution of the coarse models</li> <li><code>lores.zarr</code> - low-resolution simulation</li> <li><code>gauss.zarr</code>, <code>sharp.zarr</code> - training datasets for prediction of subgrid forcing obtained with Gaussian or Sharp filters</li> <li><code>hires-gauss.zarr</code>, <code>hires-sharp.zarr</code> - high-resolution simulation projected onto coarse grid with Gaussian or Sharp filters</li> </ul> <p>The directory tree is split into small tar.gz files each representing a separate .zarr archive. Download any required parts of the dataset and unpack with:</p> <p><strong>tar -xf *.tar.gz </strong></p> <p><strong>The directory tree will be restored automatically!</strong></p>
Source data belonged to "Establishing structure-property linkages for wicking time predictions in porous polymeric membranes using a data-driven approach"
<p>This record contains all the necessary data to obtain the results of the study "Establishing structure-property linkages for wicking time predictions in porous polymeric membranes using a data-driven approach"</p>
A Data-Driven Approach for Finding Requirements Relevant Feedback from TikTok and YouTube
<p>This dataset includes the list of videos from TikTok and YouTube, regarding 20 different products, used in our study on utilizing videos to identify requirements relevant user feedback. We also provide the content and labeling for each video. In addition, we provide the search terms for each of the products that helped us find the videos. </p>
Ideal Solutions for the Evaluation of A User-driven Hybrid Neuro-symbolic Approach for Knowledge Graph Creation from Relational Data
<p>Those two files represent one ideal solution of RML rules for the provided ontology and database by the study conductors.</p> <p>Note: in those two Turtle files joins are not necessary since templates could have been defined, too. This would lead to the same result.</p>
Automating the interpretation of PM2.5 time-resolved measurements using a data-driven approach
Open the record for dataset details and reuse information.
Evaluation of the potential incidence of COVID-19 and effectiveness of contention measures in Spain: a data-driven approach
<p>Data used in the work "Evaluation of the potential incidence of COVID-19 and effectiveness of contention measures in Spain: a data-driven approach":</p> <p>- Population in each province in Spain in January 2019. Dataset adapted from the data available in the National Statistics Institute (Spanish: Instituto Nacional de Estadística, INE), www.ine.es</p> <p>- Average number of individuals going from province to province by main transportation mode used per day. Dataset adapted from the data provided by the Ministry of Development, https://observatoriotransporte.mitma.gob.es/estudio-experimental</p> <p> </p>
Data from: Temperature drives epidemics in a zooplankton-fungus disease system: a trait-driven approach points to transmission via host foraging
Climatic warming will likely have idiosyncratic impacts on infectious diseases, causing some to increase while others decrease or shift geographically. A mechanistic framework could better predict these different temperature-disease outcomes. However, such a framework remains challenging to develop, due to the non-linear and (sometimes) opposing thermal responses of different host and parasite traits, and due to the difficulty of validating model predictions with observations and experiments. We address these challenges in a zooplankton-fungus (Daphnia dentifera-Metschnikowia bicuspidata) system. We test the hypothesis that warmer temperatures promote disease spread and produce larger epidemics. In lakes, epidemics that start earlier and warmer in autumn grow much larger. In a mesocosm experiment, warmer temperatures produced larger epidemics. A mechanistic model parameterized with trait assays revealed that this pattern arose primarily from the temperature-dependence of transmission rate (β), governed by the increasing foraging (and hence parasite exposure) rate of hosts (f). In the trait assays, parasite production seemed sufficiently responsive to shape epidemics as well; however, this trait proved too thermally insensitive in the mesocosm experiment and lake survey to matter much. Thus, in warmer environments, increased foraging of hosts raised transmission rate, yielding bigger epidemics through a potentially general, exposure-based mechanism for ectotherms. This mechanistic approach highlights how a trait-based framework will enhance predictive insight into responses of infectious disease to a warmer world.
Supplementary material 3 from: Kendig AE, Canavan S, Anderson PJ, Flory SL, Gettys LA, Gordon DR, Iannone III BV, Kunzer JM, Petri T, Pfingsten IA, Lieurance D (2022) Scanning the horizon for invasive plant threats using a data-driven approach. NeoBiota 74: 129-154. https://doi.org/10.3897/neobiota.74.83312
Table S2
Supplementary material 2 from: Kendig AE, Canavan S, Anderson PJ, Flory SL, Gettys LA, Gordon DR, Iannone III BV, Kunzer JM, Petri T, Pfingsten IA, Lieurance D (2022) Scanning the horizon for invasive plant threats using a data-driven approach. NeoBiota 74: 129-154. https://doi.org/10.3897/neobiota.74.83312
Table S1
Supplementary material 4 from: Kendig AE, Canavan S, Anderson PJ, Flory SL, Gettys LA, Gordon DR, Iannone III BV, Kunzer JM, Petri T, Pfingsten IA, Lieurance D (2022) Scanning the horizon for invasive plant threats using a data-driven approach. NeoBiota 74: 129-154. https://doi.org/10.3897/neobiota.74.83312
Table S3
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.