Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

493

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

493 results for “Predictive factors”

Learn how ShareScore rates datasets ↗
dryad40/100

How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad40/100

Data and R code from: Spatiotemporal risk factors predict landscape-scale survivorship for a northern ungulate

Open the record for dataset details and reuse information.

publicAug 2022View details →
zenodo36/100

Exploratory analysis using machine learning of predictive factors for falls in persons with type 2 diabetes: A Longitudinal Study

<p>The risk of falls in elderly individuals with diabetes was reported to be 1.5 - 3 times higher than in those without diabetes. However, it is not clear what risk factors are strongly related to falls in those with diabetes. In this study, we aimed to investigate the status of falls and to identify important risk factors for falls in persons with type 2 diabetes (T2D) including the non-elderly. Participants were 316 persons with T2D who were admitted to the University of Tsukuba Hospital for treatment of diabetes. They were assessed for medical history, laboratory data and physical capabilities during the hospitalization and were given a questionnaire on falls one year after discharge. Two different statistical models, logistic regression and random forest classifier, were used to investigate important predictors of falls. The response rate to the survey was 72%; of the 226 respondents, there were 129 males and 97 females (median age 62 years). The fall rate during the first year after discharge was 19% and increased with age; fall rates were 17% for those &lt;60 years, 20% for those aged 60 &ndash; 69 years and 24% for those &ge;70 years. Logistic regression revealed that knee extension strength (&beta;= -0.698, P = 0.002), fasting C-peptide (F-CPR) level (&beta;= 0.492, P = 0.009) and dorsiflexion strength (&beta;= -0.432, P = 0.047) were independent predictors of falls. The random forest classifier placed knee extension strength (covariate importance = 0.304), grip strength (0.234), F-CPR level (0.232) and dorsiflexion strength (0.230) in the top 4 important variables for falls. The rate of falls in persons with T2D was high even in middle age. Lower extremity muscle weakness as well as elevated F-CPR levels and reduced grip strength were shown to be important risk factors for falls in T2D.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

DATASET RELATED TO ARTICLE "Brain Tumor Resection in Elderly Patients Potential Factors of Postoperative Worsening in a Predictive Outcome Model"

<p><span>clinical database including information on patients included in the study at title</span></p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

MeDeMo - a dependency model for DNA methylation-aware transcription factor binding predictions

<p>The uploaded <em>fasta </em>files contain extended reference genomes for three cell lines HepG2, GM12878, K562 (ENCODE) and two primary liver hepatocyte samples from the german epigenomics consortium (DEEP). The extended reference genomes&nbsp;contain information on DNA methylation in a CpG context. They can be used as input for <em>MeDeMo</em>, a tool to infer transcription factor binding sites incorporating not only sequence specificity but also DNA methylation. <em>MeDeMo&nbsp;</em>is available online at:&nbsp;<a href="http://www.jstacs.de/index.php/MeDeMo">http://www.jstacs.de/index.php/MeDeMo</a>.</p> <p>We considered the files ENCFF279HCL and ENCFF835NTC for GM12878, ENCFF867JRG and&nbsp;ENCFF721JMB for K562 as well as ENCFF064GJQ and ENCFF369YQW for HepG2. From DEEP, we considered samples&nbsp;41_Hf01 and&nbsp;41_Hf03 which are available through the International Human Epigenomics Consortium (IHEC).</p> <p>In addition, we provide all models trained using the mentioned data sets as well models for and motifs from genome wide predictions.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Supplementary File S2 for the publication 'Predicting Bacterial Virulence Factors - Evaluation of Machine Learning and Negative Data Strategies' by Rentzsch, R et al.

<p>Supplementary File S2 for the publication &#39;Predicting Bacterial Virulence Factors - Evaluation of Machine Learning and Negative Data Strategies&#39; by Robert Rentzsch, Carlus Deneke, Andreas Nitsche, and Bernhard Y. Renard</p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Fig. 1 in Predicting the risk of Alaria alata infestation in wild boar on the basis of environmental factors

Fig. 1. The trend in prevalence of A. alata in provinces with WETLANDS.

opencc-by-4.0Apr 2022View details →
zenodo36/100

Fig. 1 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 1. The Iguatemi River basin, showing the sampling sites in the Jogui and Iguatemi rivers.

opencc-by-4.0Mar 2007View details →
dryad36/100

Dataset for: Indirect nitrous oxide emission factors of fluvial networks can be predicted by dissolved organic carbon and nitrate from local to global scales

<p>Streams and rivers are important sources of nitrous oxide (N<sub>2</sub>O), a powerful greenhouse gas. Estimating global riverine N<sub>2</sub>O emissions is critical for the assessment of anthropogenic N<sub>2</sub>O emission inventories. The indirect N<sub>2</sub>O emission factor (EF<sub>5r</sub>) model, one of the bottom-up approaches, adopts a fixed EF<sub>5r</sub> value to estimate riverine N<sub>2</sub>O emissions based on IPCC methodology. However, the estimates have considerable uncertainty due to the large spatiotemporal variations in EF<sub>5r</sub> values. Factors regulating EF<sub>5r</sub> are poorly understood at the global scale. Here, we combine 4-year in situ observations across rivers of different land use types in China, with a global meta-analysis over six continents, to explore the spatiotemporal variations and controls on EF<sub>5r</sub> values. Our results show that the EF<sub>5r</sub> values in China and other regions with high N loads are lower than those for regions with lower N loads. Although the global mean EF<sub>5r</sub> value is comparable to the IPCC default value, the global EF<sub>5r</sub> values are highly skewed with large variations, indicating that adopting region-specific EF<sub>5r</sub> values rather than revising the fixed default value is more appropriate for the estimation of regional and global riverine N<sub>2</sub>O emissions. The ratio of dissolved organic carbon to nitrate (DOC/NO<sub>3</sub><sup>-</sup>) and NO<sub>3</sub><sup>-</sup> concentration are identified as the dominant predictors of region-specific EF<sub>5r</sub> values at both regional and global scales because stoichiometry and nutrients strictly regulate denitrification and N<sub>2</sub>O production efficiency in rivers. A multiple linear regression model using DOC/NO<sub>3</sub><sup>-</sup> and NO<sub>3</sub><sup>-</sup> is proposed to predict region-specific EF<sub>5r</sub> values. The good fit of the model associated with easily obtained water quality variables allows its widespread application. This study fills a key knowledge gap in predicting region-specific EF<sub>5r</sub> values at the global scale and provides a pathway to estimate global riverine N<sub>2</sub>O emissions more accurately based on IPCC methodology.</p> <p>This dataset is a global integrated N<sub>2</sub>O dataset including data from 4-year (2017-2020) in situ measurements of six large rivers in China, 3-year (2018-2020) in situ measurements of urban river networks in Beijing of China, and 825 measurements from 70 published papers over six continents. The data includes dissolved N<sub>2</sub>O concentration, biogeochemical (DOC, NO<sub>3</sub><sup>-</sup>, NH<sub>4</sub><sup>+</sup>, temperature, and DO), climatological (climate zones), and geographic (region, location, and land cover) information.</p>

opencc-zeroJan 2023View details →
dryad36/100

Data from: Predicting regional carbon price in China based on multi-factor HKELM by combining secondary decomposition and ensemble learning

<p class="MsoNormal"><span>Accurately predicting carbon price is crucial for risk avoidance in the carbon financial market. In light of the complex characteristics of the regional carbon price in China, this paper proposes a model to forecast carbon price based on the multi-factor hybrid kernel-based extreme learning machine (HKELM) by combining secondary decomposition and ensemble learning. Variational mode decomposition (VMD) is first used to decompose the carbon price into several modes, and range entropy is then used to reconstruct these modes. The multi-factor HKELM optimized by the sparrow search algorithm is used to forecast the reconstructed subsequences, where the main external factors innovatively selected by maximum information coefficient and historical time-series data on carbon prices are both considered as input variables to the forecasting model. Following this, the improved complete ensemble-based empirical mode decomposition with adaptive noise and range entropy are respectively used to decompose and reconstruct the residual term generated by VMD. Finally, the nonlinear ensemble learning method is introduced to determine the predictions of residual term and final carbon price. In the empirical analysis of Guangzhou market, the root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE) of the model are 0.1716, 0.1218 and 0.0026, respectively. The proposed model outperforms other comparative models in predicting accuracy. The work here extends the research on forecasting theory and methods of predicting the carbon price.</span></p>

opencc-zeroMay 2023View details →
zenodo36/100

Data sets and machine learning models for: Predicting critical properties and acentric factor of fluids using multi-task machine learning

<p>The experimental data sets, data splits, additional features, QM calculations, model predictions, and final machine learning models for the manuscript &quot;Predicting Critical Properties and Acentric Factor of Fluids Using Multi-Task Machine Learning&quot;.&nbsp;<strong>Citation should refer directly to the manuscript:</strong></p> <ul> <li> <p>Biswas, S.;&nbsp;Chung, Y.;&nbsp;Ramirez, J.;&nbsp;Wu, H.;&nbsp;Green, W. H.&nbsp;Predicting Critical Properties and Acentric Factors of Fluids Using Multitask Machine Learning. <em>Journal of Chemical Information and Modeling.</em>&nbsp;<strong>2023</strong>&nbsp;<em>63</em>&nbsp;(15), 4574-4588. DOI: <a href="https://doi.org/10.1021/acs.jcim.3c00546">10.1021/acs.jcim.3c00546</a></p> </li> </ul> <p>To use the machine learning&nbsp;models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a>.&nbsp;</p> <p>Detailed information&nbsp;can be found in README.md file.</p> <p>&nbsp;</p> <p><strong>Details on the properties considered</strong></p> <p>The data set includes the following 8 properties:</p> <ul> <li>Tc: critical temperature, in K</li> <li>Pc: critical pressure, in bar</li> <li>rhoc: critical density, in mol/L</li> <li>omega: acentric factor, unitless</li> <li>Tb: boiling point, in K</li> <li>Tm: melting point, in K</li> <li>dHvap: enthalpy of vaporization at boiling point, in kJ/mol</li> <li>dHfus: enthalpy of fusion at melting point, in kJ/mol</li> </ul> <p><strong>Details on the files</strong></p> <p>1. Data sets under CritProp_v1.1.0:</p> <ul> <li>all_data: includes the data sets used in this work. All data points are listed for each chemical compound as well as&nbsp;its corresponding data source. The details of the data sources can be found in the README.md file. The distribution of the data set is included in each folder. <ul> <li>estimated_data_for_pretraining:&nbsp;contains the estimated data from Yaws&#39; handbook that are used to pre-train our machine learning&nbsp;(ML) model.</li> <li>experimental_data: contains the experimental data (references 1 - 15) used to fine-tune our final ML model.</li> </ul> </li> <li>additional_features: includes the additional features tested for the ML model.&nbsp;The Abraham features are generated for all data (references 1 - 15)&nbsp;while the acsf, qm, and rdkit features are only generated for the data from references 1 - 9. <ul> <li>abraham: Abraham solute parameters (E, S, A, B, L). Molecular features.</li> <li>acsf: ACSF (atom-centered symmetry functions). Atomic features that are coverted from the 3D coordinates of the compound</li> <li>qm_atom: QM (quantum chemical) atomic feature.&nbsp;</li> <li>qm_mol: QM molecular feature.</li> <li>rdkit: Selected RDKit 2D molecular features.</li> </ul> </li> <li>data_splits_and_model_predictions: contains the training and test sets used to evaluate the model. It also&nbsp;contains the predicted values from our final ML model for each test set. <ul> <li>random and scaffold splits: training and test sets that include the data from references 1 - 9.</li> <li>external test set: a test set that includes the data from only references 10 - 15.</li> </ul> </li> </ul> <p>2. Machine learning (ML) model files:</p> <ul> <li>CritProp_ML_model_files_with_abraham_feat.zip: contains the Chemprop ML model files that are trained using Abraham features&nbsp;as additional molecular features. This gives the best results.</li> <li>CritProp_ML_model_files_without_additional_feat.zip: contains the Chemprop ML model files that are trained without&nbsp;any additional features. This gives the second best results.</li> </ul> <p>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a></p> <p>3. QM (quantum chemical) calculations:</p> <ul> <li>QM_calculations.zip: contains the results of the QM calculations that are performed to compute QM features.</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data files and taxonomic classifiers for Pseudomonas syringae classification and virulence factor prediction

<p>PSSC.tree : core-genome tree of 2,161 high quality <em>Pseudomonas syringae</em> genomes</p> <p>metadata.csv: A CSV file containing taxonomic data, type strain designations, phylogroups as assigned in this study, LIN clusters assigned for classification purposes, presence/absence of key virulence factors, and metadata found in each genome&rsquo;s Biosample record for all genomes found in PSSC.tree</p> <p>CLASSIFIER_xxx: QIIME 2 classifier artifacts trained on amplicons generated from primer sets indicated in file name</p> <p>xxx_VFOC.JSON: HMMER results for T3SS and effectors and WHOP genes, structured with both genome and gene product accession numbers as primary key, depending on file</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
ClinicalTrials.gov36/100

Risk Factors and Prediction Model of Cancer-associated Venous Thromboembolism

ClinicalTrials.gov study NCT05729464. IPD Sharing: UNDECIDED. Countries: 1. Publications: 11.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Identifying Factors That Predict Antidepressant Treatment Response

ClinicalTrials.gov study NCT00360399. IPD Sharing: Not stated. Countries: 1. Publications: 6.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Predicting Risk Factors of Postoperative Hypocalcemia After Total Thyroidectomy

ClinicalTrials.gov study NCT04372225. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
dryad36/100

Data from: Predicting regional carbon price in China based on multi-factor HKELM by combining secondary decomposition and ensemble learning

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad36/100

Dataset for: Indirect nitrous oxide emission factors of fluvial networks can be predicted by dissolved organic carbon and nitrate from local to global scales

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad32/100

Data from: Predicted effects of climate factors on mountain species are not uniform over different spatial scales

The selection of relevant factors and appropriate spatial scale(s) is fundamental when modelling species response to climate change. We evaluated whether the effects of climate factors on species distribution/occurrence are consistently modelled over different spatial scales in birds, and used a two-scale approach to identify species-climate correlations unlikely to represent causal effects. We used passerine birds inhabiting mountain grassland in the Apennines (Italy) as a model. We surveyed four grassland species at 400 sampling points, and built habitat selection models (territory scale) and distribution models (7 algorithms, landscape scale). We compared the effect of climatic predictors on occurrence/distribution highlighted by models over to the two spatial scales, and with the effects supposed a priori based on the climatic niche of each species. Models at the territory level included at least one climatic predictor for three species; the observed effect of climatic predictors was seldom consistent with supposed effects. At the broadest scale, distribution models for all species included climatic predictors, with varying consistence with supposed effects and findings at the finer scale. Despite the importance of climate for species distribution, occurrence could be more directly related to other factors, with important implications for understanding/predicting the impacts of climate/environmental changes. Our approach revealed key variables for grassland birds, and highlighted the scale-dependent perceived importance of climate. At the local scale, climate effects were weak or hard to interpret. We found a general lack of consistence between supposed and observed effects at the territory level, and between landscape and territory models. Our results show the importance of predicting the potential effect of climatic factors prior to the analyses, carefully selecting ecologically meaningful variables and scales, and evaluating the nature and scale of climate-species links. We call for caution when predicting under future climates, especially when mechanistic effects and consistency across scales lack.

opencc-zeroAug 2019View details →
dryad32/100

Data from: Detection error influences both temporal seroprevalence predictions and risk factors associations in wildlife disease models

Understanding the prevalence of pathogens in invasive species is essential to guide efforts to prevent transmission to agricultural animals, wildlife, and humans. Pathogen prevalence can be difficult to estimate for wild species due to imperfect sampling and testing (pathogens may not be detected in infected individuals and erroneously detected in individuals that are not infected). The invasive wild pig (Sus scrofa, also referred to as wild boar and feral swine) is one of the most widespread hosts of domestic animal and human pathogens in North America. We developed hierarchical Bayesian models that account for imperfect detection to estimate the seroprevalence of five pathogens (porcine reproductive and respiratory syndrome virus, pseudorabies virus, Influenza A virus in swine, Hepatitis E virus, and Brucella spp.) in wild pigs in the United States using a dataset of over 50,000 samples across nine years. To assess the effect of incorporating detection error in models, we also evaluated models that ignored detection error. Both sets of models included effects of demographic parameters on seroprevalence. We compared our predictions of seroprevalence to 40 published studies, only one of which accounted for imperfect detection. We found a range of seroprevalence among the pathogens with a high seroprevalence of pseudorabies virus, indicating significant risk to livestock and wildlife. Demographics had mostly weak effects, indicating that other variables may have greater effects in predicting seroprevalence. Models that ignored detection error led to different predictions of seroprevalence as well as different inferences on the effects of demographic parameters. Our results highlight the importance of incorporating detection error in models of seroprevalence and demonstrate that ignoring such error may lead to erroneous conclusions about the risk associated with pathogen transmission. When using opportunistic sampling data to model seroprevalence and evaluate risk factors, detection error should be included.

opencc-zeroAug 2019View details →
dryad32/100

Data from: SIDER: an R package for predicting trophic discrimination factors of consumers based on their ecology and phylogenetic relatedness

Stable isotope mixing models (SIMMs) are an important tool used to study species' trophic ecology. These models are dependent on, and sensitive to, the choice of trophic discrimination factors (TDF) representing the offset in stable isotope delta values between a consumer and their food source when they are at equilibrium. Ideally, controlled feeding trials should be conducted to determine the appropriate TDF for each consumer, tissue type, food source, and isotope combination used in a study. In reality however, this is often not feasible nor practical. In the absence of species-specific information, many researchers either default to an average TDF value for the major taxonomic group of their consumer, or they choose the nearest phylogenetic neighbour for which a TDF is available. Here, we present the SIDER package for R, which uses a phylogenetic regression model based on a compiled dataset to impute (estimate) a TDF of a consumer. We apply information on the tissue type and feeding ecology of the consumer, all of which are known to affect TDFs, using Bayesian inference. Presently, our approach can estimate TDFs for two commonly used isotopes (nitrogen and carbon), for species of mammals and birds with or without previous TDF information. The estimated posterior probability provides both a mean and variance, reflecting the uncertainty of the estimate, and can be subsequently used in the current suite of SIMM software. SIDER allows users to place a greater degree of confidence on their choice of TDF and its associated uncertainty, thereby leading to more robust predictions about trophic relationships in cases where study-specific data from feeding trials is unavailable. The underlying database can be updated readily to incorporate more stable isotope tracers, replicates and taxonomic groups to further increase the confidence in dietary estimates from stable isotope mixing models, as this information becomes available.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record