Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

493

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

493 results for “Predictive factors”

Learn how ShareScore rates datasets ↗
edi48/100

Soil factors predict initial plant colonization on Puerto Rican landslides

Tropical storms are the principal cause of landslides in montane rainforests, such as the Luquillo Experimental Forest (LEF) of Puerto Rico . A storm in 2003 caused 30 new landslides in the LEF that we used to examine prior hypotheses that slope stability and organically enriched soils are prerequisites for plant colonization. We measured slope stability and litterfall in 1 m2 plots 8-13 months following landslide formation. At 13 months we also measured microtopography, soil characteristics (organic matter, particle size, total nitrogen, and water holding capacity), elevation, distance to forest edge, and canopy cover, as well as plant aboveground biomass, plant cover, and root biomass. Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.

openCC (other)Nov 2023View details →
zenodo44/100

Factors to predict above-ground biomass carbon carrying capacity

<p>The climate data (Mean annual temperature (&deg;C, MAT), mean annual precipitation (mm, MAP), annually accumulated temperature with days &ge; 0&deg;C (&deg;C-days, AAT0), annually accumulated temperature with days &ge; 10&deg;C (&deg;C-days, AAT10), aridity index, and humidity index ), soil properties (soil texture and soil types)&nbsp;and DEM are available from the Research Center for Eco-Environmental Sciences, Chinese Academy of Sciences (https://www.resdc.cn/); The geological elements and hydrological elements data can be found at&nbsp;http://dcc.ngac.org.cn/geologicalData/rest/geologicalData/geologicalDataDetail/402881f75d9bc077015d9bc084160000and&nbsp;https://www.webmap.cn/commres.do?method=result25W; The geomorphology data set is provided by National Tibetan Plateau Data Center (http://data.tpdc.ac.cn/zh-hans/data/63e290d7-7087-462a-acac-50195fba530b/). All data were resampled at 500m resolution.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Virtual ChIP-seq predictions of binding of 36 transcription factor in Roadmap Epigenomics Project tissues

<p>This dataset contains predictions of Virtual ChIP-seq for binding of 36&nbsp;transcription factors in Roadmap Epigenomics dataset tissues with matched DNase-seq and RNA-seq data.</p> <p>Tarball contains subfolders for each of the 36&nbsp;TFs where Virtual ChIP-seq median MCC&nbsp;in validation cell types was &gt; 0.3.</p> <p>Each subfolder contains gzipped BED files. Each file is named as &lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;_Predictions.bed.gz. Columns correspond to Chromosome, Start, End,&nbsp;&lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;, Posterior probability</p> <p>You can use the posterior probabilities provided in Virchip_PosteriorCutoffs_V3.0.0.tsv. These are posterior probability cutoffs which maximized MCC in H1-hESC cell type, or are set to 0.4 if there was no ChIP-seq data of that TF in H1-hESC (0.4 is the mode of all optimal posterior probability cutoffs in H1-hESC).</p>

opencc-zeroOct 2018View details →
zenodo44/100

QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?

<p>This dataset is described and studied in the article&nbsp;</p> <p>&quot;QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?&quot;</p> <p>published in <em>SAR and QSAR Environmental Research</em> (Taylor&amp;Francis).</p> <p>Files description:</p> <p>SI_BCFtrainset.xlsx: a collection of 1129 chemical structures and CAS identifiers with their logBCF values extracted from various literature sources.</p> <p>SI_BCFtestset.xlsx: a collection of 204 chemical structures for which the logBCF is considered of lower reliability and used as an external test set.</p> <p>SI_FullDataset_rawdata.csv: the raw data composed of 15372 entries with the following columns:&nbsp;CASRN, Tissue, Duration [d], Test organism, Exposure type, Steady state, RESPONSE, RESPONSE UNIT, Media&nbsp;type, TakenFrom, TITLE, AUTHOR, YEAR, SOURCE, SMILES</p> <p>SI_ExcludedOutliers34.csv: 34 chemical structures that have been identified as suspicious during analysis.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Risk factor prediction for Secondary Glaucoma amongst patients presenting with Pseudo exfoliation Syndrome (PEX) at Ophthalmology OPD in a Tertiary Care Centre in Ahmedabad

<p>Here we are uploading a data sheet of the<strong> &quot;Risk factor prediction for Secondary Glaucoma amongst patients presenting with Pseudo exfoliation Syndrome (PEX) at Ophthalmology OPD in a Tertiary Care Centre in Ahmedabad.&quot;&nbsp;</strong></p>

opencc-by-4.0Mar 2023View details →
dryad40/100

How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design

<p>Extra-organismal DNA (eoDNA) from material left behind by organisms (non-invasive DNA: e.g., faeces, hair) or from environmental samples (eDNA: e.g., water, soil) is a valuable source of genetic information. However, the relatively low quality and quantity of eoDNA, which can be further degraded by environmental factors, results in reduced amplification and sequencing success. This is often compensated for through cost- and time-intensive replications of genotyping/sequencing procedures. Therefore, system- and site-specific quantifications of environmental degradation are needed to maximize sampling efficiency (e.g., fewer replicates, shorter sampling durations), and to improve species detection and abundance estimates. Using ten environmentally diverse bat roosts as a case study, we developed a robust modelling pipeline to quantify the environmental factors degrading eoDNA, predict eoDNA quality, and estimate sampling-site-specific ideal exposure duration. Maximum humidity was the strongest eoDNA-degrading factor, followed by exposure duration and then maximum temperature. We also found a positive effect when hottest days occurred later. The strength of this effect fell between the strength of the effects of exposure duration and maximum temperature. With those predictors and information on sampling period (before or after offspring were born), we reliably predicted mean eoDNA quality per sampling visit at new sites with a mean squared error of 0.0349. Site-specific simulations revealed that reducing exposure duration to 2-8 days could substantially improve eoDNA quality for future sampling. Our pipeline identified high humidity and temperature as strong drivers of eoDNA degradation even in the absence of rain and direct sunlight. Furthermore, we outline the pipeline's utility for other systems and study goals, such as estimating sample age, improving eDNA-based species detection, and increasing the accuracy of abundance estimates.</p>

opencc-zeroMar 2023View details →
dryad40/100

Data and R code from: Spatiotemporal risk factors predict landscape-scale survivorship for a northern ungulate

<p>These data and computer code (written in R, https://www.r-project.org) were created to statistically evaluate a suite of spatiotemporal covariates that could potentially explain pronghorn (Antilocapra americana) mortality risk in the Northern Sagebrush Steppe (NSS) ecosystem (50.0757<sup>o</sup> N, −108.7526<sup>o</sup> W). Known-fate data were collected from 170 adult female pronghorn monitored with GPS collars from 2003-2011, which were used to construct a time-to-event (TTE) dataset with a daily timescale and an annual recurrent origin of 11 November. Seasonal risk periods (winter, spring, summer, autumn) were defined by median migration dates of collared pronghorn. We linked this TTE dataset with spatiotemporal covariates that were extracted and collated from pronghorn seasonal activity areas (estimated using 95% minimum convex polygons) to form a final dataset. Specifically, average fence and road densities (km/km2), average snow water equivalent (SWE; kg/m2), and maximum decadal normalized difference vegetation index (NDVI) were considered as predictors. We tested for these main effects of spatiotemporal risk covariates as well as the hypotheses that pronghorn mortality risk from roads or fences could be intensified during severe winter weather (i.e., interactions: SWE*road density and SWE*fence density). We also compare an analogous frequentist implementation to estimate model-averaged risk coefficients. Ultimately, the study aimed to develop the first broad-scale, spatially explicit map of predicted annual pronghorn survivorship based on anthropogenic features and environmental gradients to identify areas for conservation and habitat restoration efforts.</p> <p> </p>

opencc-zeroAug 2022View details →
zenodo40/100

Fig. 1 in Predictive factors of species composition of follower fishes in nuclear-follower feeding associations: a snapshot study

Fig. 1. Selected examples of nuclear-follower fish associations recorded in a stream in Southwestern Brazil: Leporellus vittatus begins to approach the bottom, while followed by a small group of Jupiaba acanthogaster (A); Leporinus macrocephalus disturbs the bottom, while Crenicichla vittata approaches to feed on fleeing or uncovered prey (B); L. macrocephalus stirs a little cloud, while a small group of Odontostilbe pequira and J. acanthogaster approaches (C); the same nuclear species stirs a large cloud, which congregates large numbers of O. pequira (D); Prochilodus lineatus begins to probe on the bottom, readily attracting a small group of O. pequira and J. acanthogaster (E); the same nuclear species stirs a large sediment cloud, which attracts large numbers of O. pequira (F). Photo credits: Sergio R. Floeter (A, E); José Sabino (B); Guilherme Ortigara Longo (C); Ivan Sazima (D, F).

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 2 in Predictive factors of species composition of follower fishes in nuclear-follower feeding associations: a snapshot study

Fig. 2. Two-dimensional ordination of samples, considering the composition of follower fish species in the associations from Bray-Curtis similarity coefficient, and emphasizing the size of the sediment cloud (A), the duration of the sediment cloud (B), the identity of the nuclear species (C) and the abundance of the follower species Odontostilbe pequira, represented by the circumferences' sizes (D).

opencc-by-4.0Dec 2014View details →
zenodo40/100

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 1. Factors that Affect Thyroid Function (The Institute for Functional Medicine, 2014)

<p>&nbsp;In Figure 1 are presented the main factors that affect the thyroid function. It is obvious that factors such as stress, infection, toxins, trauma and certain medication are directly responsible for the improper production of thyroid hormones. Symptoms identification and the early detection of abnormal values of thyroid hormones after clinical investigation will help in establishing the proper diagnostic and to prescribe the right medication. The patient must periodically evaluate his clinical state in order to receive the treatment as long as he needs it.&nbsp;&nbsp;</p>

opencc-by-4.0Jun 2016View details →
zenodo40/100

Fig. 2 in Predicting the risk of Alaria alata infestation in wild boar on the basis of environmental factors

Fig. 2. The prevalence of A. alata in wild boar in provinces in Poland calculated from literature values and data from the present study (A) and predicted by percentage of areas covered by WETLANDS (B) (for detailed information, see: Methods). The figure shows prevalence values for a given province and confidence intervals (lower; upper).

opencc-by-4.0Apr 2022View details →
zenodo40/100

Prediction of Humpback Whale Sighting Zones based on Environmental Factors using Tree-based Algorithms

<p>This is the datased used in the paper: Prediction of Humpback Whale Sighting Zones based on Environmental Factors using Tree-based Algorithms</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Transfer learning and DNA language models enhance transcription factor binding predictions

<p>This is the dataset for replicating the results of the paper called "Transfer learning and DNA language models enhance transcription factor binding predictions" by Ekin Deniz Aksu and Martin Vingron.</p> <p>See https://github.com/ekinda/tfbs_prediction_paper</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Metabolic pathway prediction using non-negative matrix factorization with improved precision

<p>We include samples of various data types used in the work &quot;Metabolic pathway prediction using non-negative matrix factorization with improved precision&quot;</p> <p>More information about the software package and instructions are provided in&nbsp;<a href="https://github.com/hallamlab/triUMPF">hallamlab/triUMPF</a></p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Fig. 3 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 3. Scatterplot of canonical correspondence analysis (CCA) for the fish communities of the Jogui and Iguatemi Rivers.

opencc-by-4.0Mar 2007View details →
zenodo40/100

Fig. 2 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 2. Similarity dendrogram of fish communities in Jogui River (above) and Iguatemi Rivers (below).

opencc-by-4.0Mar 2007View details →
zenodo40/100

Fig. 4 in Environmental factors predicting fish community structure in two neotropical rivers in Brazil

Fig. 4. Altitudinal distributions of the main fish species in the Jogui (A) and Iguatemi (B) rivers. Black dots represent sampling sites. Horizontal black lines represent species distribution range.

opencc-by-4.0Mar 2007View details →
zenodo40/100

Supplementary material for: RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding

<p>Supplementary Material for the Article</p> <p>Santana-Garcia, W., Rocha-Acevedo, M., Ramirez-Navarro, L., Mbouamboua, Y., Thieffry, D., Thomas-Chollier, M., Contreras-Moreira, B., van Helden, J., Medina-Rivera, A., 2019. RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding. Comput. Struct. Biotechnol. J. 17, 1415&ndash;1428.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Quasar Factor Analysis – An Unsupervised and Probabilistic Quasar Continuum Prediction Algorithm with Latent Factor Analysis

<p>Dataset used in&nbsp;<em>Quasar Factor Analysis &ndash; An Unsupervised and Probabilistic Quasar Continuum Prediction Algorithm with Latent Factor Analysis&nbsp;</em>[<a href="https://arxiv.org/abs/2211.11784">arXiv: <strong>2211.11784</strong></a>]. This dataset will be helpful to validate different continuum prediction model and study absorption systems.<br> <br> The descriptions of individual files can be found here:</p> <ul> <li><a href="/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-dr16.tar.gz?versionId=427c1aca-6f12-4e84-b91e-f9a3d678329e">sdss-dr16.tar.gz</a>&nbsp;: continuum prediction for ~100,000 quasar spectra from SDSS DR16, see Section 3.1 in <a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>;</li> <li><a href="https://zenodo.org/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-mock-with-dla-with-perturb.tar.gz?versionId=55bc466d-4cbf-43ba-9b83-209ffba4ec30">sdss-mock-with-dla-with-perturb.tar.gz&nbsp;</a>&nbsp;: ~150,000 mock quasar spectra to validate QFA performance with perturbation on &nbsp;quasar continuum from PCA template, see Section 3.2 in <a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>;</li> <li><a href="https://zenodo.org/api/files/01c3b414-7572-4a24-9c67-fbab6ae05969/sdss-mock-with-dla-without-perturb.tar.gz?versionId=efa5d05a-2a40-4a43-801e-c95b29a4efb2">sdss-mock-with-dla-without-perturb.tar.gz</a>: ~150,000 mock quasar spectra to validate QFA performance with quasar continuum directly from PCA template, see Section 3.2 in&nbsp;<a href="https://arxiv.org/abs/2211.11784">arXiv:211.11784</a>.<br> <br> <br> &nbsp;</li> </ul>

opencc-byJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record