Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
493
datasets available to search
ShareScore release 0.9.0
Dataset results
493 results for “Predictive factors”
How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design
Open the record for dataset details and reuse information.
Data and R code from: Spatiotemporal risk factors predict landscape-scale survivorship for a northern ungulate
Open the record for dataset details and reuse information.
Exploratory analysis using machine learning of predictive factors for falls in persons with type 2 diabetes: A Longitudinal Study
<p>The risk of falls in elderly individuals with diabetes was reported to be 1.5 - 3 times higher than in those without diabetes. However, it is not clear what risk factors are strongly related to falls in those with diabetes. In this study, we aimed to investigate the status of falls and to identify important risk factors for falls in persons with type 2 diabetes (T2D) including the non-elderly. Participants were 316 persons with T2D who were admitted to the University of Tsukuba Hospital for treatment of diabetes. They were assessed for medical history, laboratory data and physical capabilities during the hospitalization and were given a questionnaire on falls one year after discharge. Two different statistical models, logistic regression and random forest classifier, were used to investigate important predictors of falls. The response rate to the survey was 72%; of the 226 respondents, there were 129 males and 97 females (median age 62 years). The fall rate during the first year after discharge was 19% and increased with age; fall rates were 17% for those <60 years, 20% for those aged 60 – 69 years and 24% for those ≥70 years. Logistic regression revealed that knee extension strength (β= -0.698, P = 0.002), fasting C-peptide (F-CPR) level (β= 0.492, P = 0.009) and dorsiflexion strength (β= -0.432, P = 0.047) were independent predictors of falls. The random forest classifier placed knee extension strength (covariate importance = 0.304), grip strength (0.234), F-CPR level (0.232) and dorsiflexion strength (0.230) in the top 4 important variables for falls. The rate of falls in persons with T2D was high even in middle age. Lower extremity muscle weakness as well as elevated F-CPR levels and reduced grip strength were shown to be important risk factors for falls in T2D.</p>
DATASET RELATED TO ARTICLE "Brain Tumor Resection in Elderly Patients Potential Factors of Postoperative Worsening in a Predictive Outcome Model"
<p><span>clinical database including information on patients included in the study at title</span></p>
MeDeMo - a dependency model for DNA methylation-aware transcription factor binding predictions
<p>The uploaded <em>fasta </em>files contain extended reference genomes for three cell lines HepG2, GM12878, K562 (ENCODE) and two primary liver hepatocyte samples from the german epigenomics consortium (DEEP). The extended reference genomes contain information on DNA methylation in a CpG context. They can be used as input for <em>MeDeMo</em>, a tool to infer transcription factor binding sites incorporating not only sequence specificity but also DNA methylation. <em>MeDeMo </em>is available online at: <a href="http://www.jstacs.de/index.php/MeDeMo">http://www.jstacs.de/index.php/MeDeMo</a>.</p> <p>We considered the files ENCFF279HCL and ENCFF835NTC for GM12878, ENCFF867JRG and ENCFF721JMB for K562 as well as ENCFF064GJQ and ENCFF369YQW for HepG2. From DEEP, we considered samples 41_Hf01 and 41_Hf03 which are available through the International Human Epigenomics Consortium (IHEC).</p> <p>In addition, we provide all models trained using the mentioned data sets as well models for and motifs from genome wide predictions.</p>
Supplementary File S2 for the publication 'Predicting Bacterial Virulence Factors - Evaluation of Machine Learning and Negative Data Strategies' by Rentzsch, R et al.
<p>Supplementary File S2 for the publication 'Predicting Bacterial Virulence Factors - Evaluation of Machine Learning and Negative Data Strategies' by Robert Rentzsch, Carlus Deneke, Andreas Nitsche, and Bernhard Y. Renard</p>
Fig. 1 in Predicting the risk of Alaria alata infestation in wild boar on the basis of environmental factors
Fig. 1. The trend in prevalence of A. alata in provinces with WETLANDS.
Fig. 1 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil
Fig. 1. The Iguatemi River basin, showing the sampling sites in the Jogui and Iguatemi rivers.
Dataset for: Indirect nitrous oxide emission factors of fluvial networks can be predicted by dissolved organic carbon and nitrate from local to global scales
<p>Streams and rivers are important sources of nitrous oxide (N<sub>2</sub>O), a powerful greenhouse gas. Estimating global riverine N<sub>2</sub>O emissions is critical for the assessment of anthropogenic N<sub>2</sub>O emission inventories. The indirect N<sub>2</sub>O emission factor (EF<sub>5r</sub>) model, one of the bottom-up approaches, adopts a fixed EF<sub>5r</sub> value to estimate riverine N<sub>2</sub>O emissions based on IPCC methodology. However, the estimates have considerable uncertainty due to the large spatiotemporal variations in EF<sub>5r</sub> values. Factors regulating EF<sub>5r</sub> are poorly understood at the global scale. Here, we combine 4-year in situ observations across rivers of different land use types in China, with a global meta-analysis over six continents, to explore the spatiotemporal variations and controls on EF<sub>5r</sub> values. Our results show that the EF<sub>5r</sub> values in China and other regions with high N loads are lower than those for regions with lower N loads. Although the global mean EF<sub>5r</sub> value is comparable to the IPCC default value, the global EF<sub>5r</sub> values are highly skewed with large variations, indicating that adopting region-specific EF<sub>5r</sub> values rather than revising the fixed default value is more appropriate for the estimation of regional and global riverine N<sub>2</sub>O emissions. The ratio of dissolved organic carbon to nitrate (DOC/NO<sub>3</sub><sup>-</sup>) and NO<sub>3</sub><sup>-</sup> concentration are identified as the dominant predictors of region-specific EF<sub>5r</sub> values at both regional and global scales because stoichiometry and nutrients strictly regulate denitrification and N<sub>2</sub>O production efficiency in rivers. A multiple linear regression model using DOC/NO<sub>3</sub><sup>-</sup> and NO<sub>3</sub><sup>-</sup> is proposed to predict region-specific EF<sub>5r</sub> values. The good fit of the model associated with easily obtained water quality variables allows its widespread application. This study fills a key knowledge gap in predicting region-specific EF<sub>5r</sub> values at the global scale and provides a pathway to estimate global riverine N<sub>2</sub>O emissions more accurately based on IPCC methodology.</p> <p>This dataset is a global integrated N<sub>2</sub>O dataset including data from 4-year (2017-2020) in situ measurements of six large rivers in China, 3-year (2018-2020) in situ measurements of urban river networks in Beijing of China, and 825 measurements from 70 published papers over six continents. The data includes dissolved N<sub>2</sub>O concentration, biogeochemical (DOC, NO<sub>3</sub><sup>-</sup>, NH<sub>4</sub><sup>+</sup>, temperature, and DO), climatological (climate zones), and geographic (region, location, and land cover) information.</p>
Data from: Predicting regional carbon price in China based on multi-factor HKELM by combining secondary decomposition and ensemble learning
<p class="MsoNormal"><span>Accurately predicting carbon price is crucial for risk avoidance in the carbon financial market. In light of the complex characteristics of the regional carbon price in China, this paper proposes a model to forecast carbon price based on the multi-factor hybrid kernel-based extreme learning machine (HKELM) by combining secondary decomposition and ensemble learning. Variational mode decomposition (VMD) is first used to decompose the carbon price into several modes, and range entropy is then used to reconstruct these modes. The multi-factor HKELM optimized by the sparrow search algorithm is used to forecast the reconstructed subsequences, where the main external factors innovatively selected by maximum information coefficient and historical time-series data on carbon prices are both considered as input variables to the forecasting model. Following this, the improved complete ensemble-based empirical mode decomposition with adaptive noise and range entropy are respectively used to decompose and reconstruct the residual term generated by VMD. Finally, the nonlinear ensemble learning method is introduced to determine the predictions of residual term and final carbon price. In the empirical analysis of Guangzhou market, the root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE) of the model are 0.1716, 0.1218 and 0.0026, respectively. The proposed model outperforms other comparative models in predicting accuracy. The work here extends the research on forecasting theory and methods of predicting the carbon price.</span></p>
Data sets and machine learning models for: Predicting critical properties and acentric factor of fluids using multi-task machine learning
<p>The experimental data sets, data splits, additional features, QM calculations, model predictions, and final machine learning models for the manuscript "Predicting Critical Properties and Acentric Factor of Fluids Using Multi-Task Machine Learning". <strong>Citation should refer directly to the manuscript:</strong></p> <ul> <li> <p>Biswas, S.; Chung, Y.; Ramirez, J.; Wu, H.; Green, W. H. Predicting Critical Properties and Acentric Factors of Fluids Using Multitask Machine Learning. <em>Journal of Chemical Information and Modeling.</em> <strong>2023</strong> <em>63</em> (15), 4574-4588. DOI: <a href="https://doi.org/10.1021/acs.jcim.3c00546">10.1021/acs.jcim.3c00546</a></p> </li> </ul> <p>To use the machine learning models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a>. </p> <p>Detailed information can be found in README.md file.</p> <p> </p> <p><strong>Details on the properties considered</strong></p> <p>The data set includes the following 8 properties:</p> <ul> <li>Tc: critical temperature, in K</li> <li>Pc: critical pressure, in bar</li> <li>rhoc: critical density, in mol/L</li> <li>omega: acentric factor, unitless</li> <li>Tb: boiling point, in K</li> <li>Tm: melting point, in K</li> <li>dHvap: enthalpy of vaporization at boiling point, in kJ/mol</li> <li>dHfus: enthalpy of fusion at melting point, in kJ/mol</li> </ul> <p><strong>Details on the files</strong></p> <p>1. Data sets under CritProp_v1.1.0:</p> <ul> <li>all_data: includes the data sets used in this work. All data points are listed for each chemical compound as well as its corresponding data source. The details of the data sources can be found in the README.md file. The distribution of the data set is included in each folder. <ul> <li>estimated_data_for_pretraining: contains the estimated data from Yaws' handbook that are used to pre-train our machine learning (ML) model.</li> <li>experimental_data: contains the experimental data (references 1 - 15) used to fine-tune our final ML model.</li> </ul> </li> <li>additional_features: includes the additional features tested for the ML model. The Abraham features are generated for all data (references 1 - 15) while the acsf, qm, and rdkit features are only generated for the data from references 1 - 9. <ul> <li>abraham: Abraham solute parameters (E, S, A, B, L). Molecular features.</li> <li>acsf: ACSF (atom-centered symmetry functions). Atomic features that are coverted from the 3D coordinates of the compound</li> <li>qm_atom: QM (quantum chemical) atomic feature. </li> <li>qm_mol: QM molecular feature.</li> <li>rdkit: Selected RDKit 2D molecular features.</li> </ul> </li> <li>data_splits_and_model_predictions: contains the training and test sets used to evaluate the model. It also contains the predicted values from our final ML model for each test set. <ul> <li>random and scaffold splits: training and test sets that include the data from references 1 - 9.</li> <li>external test set: a test set that includes the data from only references 10 - 15.</li> </ul> </li> </ul> <p>2. Machine learning (ML) model files:</p> <ul> <li>CritProp_ML_model_files_with_abraham_feat.zip: contains the Chemprop ML model files that are trained using Abraham features as additional molecular features. This gives the best results.</li> <li>CritProp_ML_model_files_without_additional_feat.zip: contains the Chemprop ML model files that are trained without any additional features. This gives the second best results.</li> </ul> <p>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a></p> <p>3. QM (quantum chemical) calculations:</p> <ul> <li>QM_calculations.zip: contains the results of the QM calculations that are performed to compute QM features.</li> </ul> <p> </p> <p> </p>
Data files and taxonomic classifiers for Pseudomonas syringae classification and virulence factor prediction
<p>PSSC.tree : core-genome tree of 2,161 high quality <em>Pseudomonas syringae</em> genomes</p> <p>metadata.csv: A CSV file containing taxonomic data, type strain designations, phylogroups as assigned in this study, LIN clusters assigned for classification purposes, presence/absence of key virulence factors, and metadata found in each genome’s Biosample record for all genomes found in PSSC.tree</p> <p>CLASSIFIER_xxx: QIIME 2 classifier artifacts trained on amplicons generated from primer sets indicated in file name</p> <p>xxx_VFOC.JSON: HMMER results for T3SS and effectors and WHOP genes, structured with both genome and gene product accession numbers as primary key, depending on file</p> <p> </p> <p> </p>
Risk Factors and Prediction Model of Cancer-associated Venous Thromboembolism
ClinicalTrials.gov study NCT05729464. IPD Sharing: UNDECIDED. Countries: 1. Publications: 11.
Identifying Factors That Predict Antidepressant Treatment Response
ClinicalTrials.gov study NCT00360399. IPD Sharing: Not stated. Countries: 1. Publications: 6.
Predicting Risk Factors of Postoperative Hypocalcemia After Total Thyroidectomy
ClinicalTrials.gov study NCT04372225. IPD Sharing: NO. Countries: 1. Publications: 1.
Data from: Predicting regional carbon price in China based on multi-factor HKELM by combining secondary decomposition and ensemble learning
Open the record for dataset details and reuse information.
Dataset for: Indirect nitrous oxide emission factors of fluvial networks can be predicted by dissolved organic carbon and nitrate from local to global scales
Open the record for dataset details and reuse information.
Data from: Predicted effects of climate factors on mountain species are not uniform over different spatial scales
The selection of relevant factors and appropriate spatial scale(s) is fundamental when modelling species response to climate change. We evaluated whether the effects of climate factors on species distribution/occurrence are consistently modelled over different spatial scales in birds, and used a two-scale approach to identify species-climate correlations unlikely to represent causal effects. We used passerine birds inhabiting mountain grassland in the Apennines (Italy) as a model. We surveyed four grassland species at 400 sampling points, and built habitat selection models (territory scale) and distribution models (7 algorithms, landscape scale). We compared the effect of climatic predictors on occurrence/distribution highlighted by models over to the two spatial scales, and with the effects supposed a priori based on the climatic niche of each species. Models at the territory level included at least one climatic predictor for three species; the observed effect of climatic predictors was seldom consistent with supposed effects. At the broadest scale, distribution models for all species included climatic predictors, with varying consistence with supposed effects and findings at the finer scale. Despite the importance of climate for species distribution, occurrence could be more directly related to other factors, with important implications for understanding/predicting the impacts of climate/environmental changes. Our approach revealed key variables for grassland birds, and highlighted the scale-dependent perceived importance of climate. At the local scale, climate effects were weak or hard to interpret. We found a general lack of consistence between supposed and observed effects at the territory level, and between landscape and territory models. Our results show the importance of predicting the potential effect of climatic factors prior to the analyses, carefully selecting ecologically meaningful variables and scales, and evaluating the nature and scale of climate-species links. We call for caution when predicting under future climates, especially when mechanistic effects and consistency across scales lack.
Data from: Detection error influences both temporal seroprevalence predictions and risk factors associations in wildlife disease models
Understanding the prevalence of pathogens in invasive species is essential to guide efforts to prevent transmission to agricultural animals, wildlife, and humans. Pathogen prevalence can be difficult to estimate for wild species due to imperfect sampling and testing (pathogens may not be detected in infected individuals and erroneously detected in individuals that are not infected). The invasive wild pig (Sus scrofa, also referred to as wild boar and feral swine) is one of the most widespread hosts of domestic animal and human pathogens in North America. We developed hierarchical Bayesian models that account for imperfect detection to estimate the seroprevalence of five pathogens (porcine reproductive and respiratory syndrome virus, pseudorabies virus, Influenza A virus in swine, Hepatitis E virus, and Brucella spp.) in wild pigs in the United States using a dataset of over 50,000 samples across nine years. To assess the effect of incorporating detection error in models, we also evaluated models that ignored detection error. Both sets of models included effects of demographic parameters on seroprevalence. We compared our predictions of seroprevalence to 40 published studies, only one of which accounted for imperfect detection. We found a range of seroprevalence among the pathogens with a high seroprevalence of pseudorabies virus, indicating significant risk to livestock and wildlife. Demographics had mostly weak effects, indicating that other variables may have greater effects in predicting seroprevalence. Models that ignored detection error led to different predictions of seroprevalence as well as different inferences on the effects of demographic parameters. Our results highlight the importance of incorporating detection error in models of seroprevalence and demonstrate that ignoring such error may lead to erroneous conclusions about the risk associated with pathogen transmission. When using opportunistic sampling data to model seroprevalence and evaluate risk factors, detection error should be included.
Data from: SIDER: an R package for predicting trophic discrimination factors of consumers based on their ecology and phylogenetic relatedness
Stable isotope mixing models (SIMMs) are an important tool used to study species' trophic ecology. These models are dependent on, and sensitive to, the choice of trophic discrimination factors (TDF) representing the offset in stable isotope delta values between a consumer and their food source when they are at equilibrium. Ideally, controlled feeding trials should be conducted to determine the appropriate TDF for each consumer, tissue type, food source, and isotope combination used in a study. In reality however, this is often not feasible nor practical. In the absence of species-specific information, many researchers either default to an average TDF value for the major taxonomic group of their consumer, or they choose the nearest phylogenetic neighbour for which a TDF is available. Here, we present the SIDER package for R, which uses a phylogenetic regression model based on a compiled dataset to impute (estimate) a TDF of a consumer. We apply information on the tissue type and feeding ecology of the consumer, all of which are known to affect TDFs, using Bayesian inference. Presently, our approach can estimate TDFs for two commonly used isotopes (nitrogen and carbon), for species of mammals and birds with or without previous TDF information. The estimated posterior probability provides both a mean and variance, reflecting the uncertainty of the estimate, and can be subsequently used in the current suite of SIMM software. SIDER allows users to place a greater degree of confidence on their choice of TDF and its associated uncertainty, thereby leading to more robust predictions about trophic relationships in cases where study-specific data from feeding trials is unavailable. The underlying database can be updated readily to incorporate more stable isotope tracers, replicates and taxonomic groups to further increase the confidence in dietary estimates from stable isotope mixing models, as this information becomes available.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.