Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

128

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

128 results for “model fit”

Learn how ShareScore rates datasets ↗
zenodo40/100

Spectral model fitting for all 4XMM-DR11 for all sources

<p>Fitting products for all sources: this deliverable includes fits with a simple model (absorbed power law) to the pipeline count rates from emldetect to all detections in the 4XMM DR11 catalogue.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Data for fitting a statistical global burned area model for seamless integration into Dynamic Global Vegetation Models

<p>The dataset is a large R data.table object saved in RDS format. It contains global, monthly data spanning the period from 2002 to 2018, with a 0.5 degrees spatial resolution. The dataset is utilized to develop and validate statistical models for predicting global burnt areas resulting from wildfires.</p>

opencc-by-4.0Nov 2024View details →
dryad40/100

Structural equation modeling reveals determinants of fitness in a cooperatively breeding bird

<p>Even in well-studied organisms, it is often challenging to uncover the social and environmental determinants of fitness. Typically, fitness is determined by a variety of factors that act in concert, thus forming complex networks of causal relationships. Moreover, even strong correlations between social and environmental conditions and fitness components may not be indicative of direct causal links, as the measured variables may be driven by unmeasured (or unmeasurable) causal factors. Standard statistical approaches, like multiple regression analyses, are not suited for disentangling such complex causal relationships. Here, we apply structural equation modeling (SEM), a technique that is specifically designed to reveal causal relationships between variables, and which also allows to include hypothetical causal factors. Therefore, SEM seems ideally suited for comparing alternative hypotheses on how fitness differences arise from differences in social and environmental factors. We apply SEM to a rich data set collected in a long-term study on the Seychelles warbler (Acrocephalus seychellensis), a bird species with facultatively cooperative breeding and a high rate of extra-group paternity. Our analysis reveals that the presence of helpers has a positive effect on the reproductive output of both female and male breeders. In contrast, per capita food availability does not affect reproductive output. Our analysis does not confirm earlier suggestions on other species that the presence of helpers has a negative effect on the reproductive output of male breeders. As such, both female and male breeders should tolerate helpers in their territories, irrespective of food availability.</p>

opencc-zeroNov 2021View details →
zenodo40/100

Are terrestrial biosphere models fit for simulating the global land carbon sink?

<p>This repository contains the data and&nbsp;code required for reproducing the results presented in the paper &quot;Are terrestrial biosphere models fit for simulating the global land carbon sink?&quot; by Seiler et al., 2021. The study evaluates an ensemble of terrestrial biosphere models&nbsp;(<a href="https://sites.exeter.ac.uk/trendy/">TRENDY</a>; v9; S3 simulations) against a wide range of reference data using the Automated Model Benchmarking R package (AMBER; version 1.1.1). The only requirement for reproducing our results is&nbsp;access to a Linux machine with <a href="https://docs.conda.io">conda</a>, an open-source package management system and environment management system,&nbsp;installed. Follow the steps described in the <em>readme</em> file to install AMBER and run the analysis. The repository also contains all output produced by our analysis.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
dryad40/100

Chronogram or phylogram for ancestral state estimation? Model-fit statistics indicate the branch lengths underlying a binary character's evolution: R scripts and simulated trees

<p>All R scripts used in this study, and the set of simulated phylogenetic trees used in the study.</p> <p>1. Modern methods of ancestral state estimation (ASE) incorporate branch length information, and it has been demonstrated that ASEs are more accurate when conducted on the branch lengths most correlated with a character's evolution; however, a reliable method for choosing between alternate branch length sets for discrete characters has not yet been proposed.<br><br>2. In this study, we simulate paired chronograms and phylograms, and generate binary characters that evolve in correlation with one of these. We then investigate (1) the effect of alternate branch lengths on ASE error, and (2) whether phylogenetic signal statistics and/or model-fit statistic can be used to select the branch lengths most correlated with a binary character.<br><br>3. In agreement with previous studies, we find that ASEs are more accurate when conducted on the branch lengths most correlated with the character. Phylogenetic signal statistics show limited utility for selecting the correct branch lengths, but model-fit statistics are found to be more accurate, with the correct branch lengths generally returning greater model-fit (lower AICc and BIC values). Using this method to choose between alternate branch length sets is more accurate when tree and character properties are more favorable for model optimization, and when shape differences between alternate phylogenies are greater.<br><br>4. Our results indicate that researchers conducting ASEs on discrete characters should carefully consider which branch lengths are appropriate, and, in the absence of other evidence, we suggest estimating model-fit values over alternate branch length sets and evolutionary models and choosing the branch length/model combination that returns better model fit.</p>

opencc-zeroMay 2022View details →
zenodo40/100

Table 5: Model Fit Summary (b) for Grade Three

<p>The model summary in Table 5 reports the strength of the relationship between the model and the<br> dependent variable, i.e., ECA exam. The multiple correlation coefficient, R, which is the linear<br> correlation between the observed and the model-predicted values of the dependent variable is .551.<br> R Square, which is also called the coefficient of determination, is .303 showing that approximately<br> 30 percent of the variation in ECA exam is explained by the model. Its interpretation is that 30<br> percent of the variation in the ECA test scores is common with the vocabulary scores. The Standard<br> Error of the Estimate of the model is approximately 3.16 out of a total of 30, meaning that the<br> prediction model produces an error range between &plusmn; [3.16]. Therefore, the prediction formula must<br> be rewritten as (AVERAGE VOCAB &times; 0.713) + 2.871&plusmn; [3.16].</p>

opencc-by-4.0Oct 2010View details →
zenodo40/100

Data for 'VespaG: Expert-guided protein language models enable accurate and blazingly fast fitness prediction'

<div>Datasets used for development of VespaG and VespaG predictions generated with <a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.&nbsp;</div> <div>&nbsp;</div> <div>Uploads contain:</div> <div> <ol> <li><strong>Performance</strong> summaries for ProteinGym [1]:<br>- Spearman and Pearson correlation for VespaG:&nbsp;<em>proteingym_performance_vespag.csv&nbsp;</em>(columns: 'DMS_id', 'Spearman', 'Pearson')<br>- Spearman correlation for evaluated methods VespaG, GEMME [2], VESPA [3], TranceptEVE [4], AlphaMissense [5], PoET [6]: <em>proteingym_spearman_allmethods.csv&nbsp;</em>(columns: 'DMS_id', 'Trancept EVE-L', 'VESPA', 'VespaG', 'GEMME', 'AlphaMissense', 'PoET', 'UniProt_ID', 'coarse_selection_type' (function), 'taxon')</li> <li><strong>Fasta</strong> files with sequences for all train sets (<em>vespag_fasta_training_datasets.zip</em> with seq_all9k.fasta, seq_human5k.fasta, seq_droso4k.fasta, seq_ecoli2k.fasta, seq_virus1k.fasta) and test set (<em>proteingym_217.fasta</em>)</li> <li><strong>VespaG</strong> <strong>Predictions</strong> for test set:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip</em> with raw_preds_ecoli.csv, raw_preds_human.csv, raw_preds_virus.csv, raw_preds_all.csv, raw_preds_droso.csv (columns: 'DMS_id', 'mutation', 'DMS_score', 'VespaG'). Predictions are based on different training data, the final model VespaG was trained on a subset of the human proteome and <strong>raw VespaG predictions</strong> <strong>for</strong> <strong>the</strong> <strong>ProteinGym benchmark are in&nbsp;raw_preds_human.csv </strong>(used to calculate the performances above).</li> <li><strong>GEMME predictions</strong> for train sets:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip&nbsp;</em>with folders 'human', 'droso', 'ecoli', 'virus', 'all' for respective fasta file (each containing GEMME mutational landscape output files named '<em>ID' + '</em>_normPred_evolCombi.txt')</li> <li><strong>ESM-2</strong> <strong>embeddings</strong> [7] for test set (<em>proteingym_217_esm2.h5</em>)</li> </ol> </div> <div>For details on VespaG see:</div> <div> <div> <div>VespaG: Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction</div> </div> <div>Celine Marquet, Julius Schlensok, Marina Abakarova, Burkhard Rost, Elodie Laine</div> <div>bioRxiv 2024.04.24.590982; doi: https://doi.org/10.1101/2024.04.24.590982</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see&nbsp;<a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <div>Exhaustive experimental annotation of the effect of all known protein variants remains daunting and expensive, stressing the need for scalable effect predictions. We introduce VespaG, a blazingly fast single amino acid variant effect predictor, leveraging embeddings of protein Language Models as input to a minimal deep learning model. To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME as a pseudo standard-of-truth. Assessed against the ProteinGym Substitution Benchmark (217 multiplex assays of variant effect with 2.5 million variants), VespaG achieved a mean Spearman correlation of 0.48 +/- 0.01, matching state-of-the-art methods such as GEMME, TranceptEVE, PoET, AlphaMissense, and VESPA. VespaG reached its top-level performance several orders of magnitude faster, predicting all mutational landscapes of the human proteome in 30 minutes on a consumer laptop (12-core CPU, 16 GB RAM).</div> <div>&nbsp;</div> <div>[1] Notin, Pascal, et al. "ProteinGym: large-scale benchmarks for protein fitness prediction and design." <em>Advances in Neural Information Processing Systems</em> 36 (2024).<br>[2] Laine, Elodie, Yasaman Karami, and Alessandra Carbone. "GEMME: a simple and fast global epistatic model predicting mutational effects." <em>Molecular biology and evolution</em> 36.11 (2019): 2604-2619.</div> <div>[3] Marquet, C&eacute;line, et al. "Embeddings from protein language models predict conservation and variant effects." <em>Human genetics</em> 141.10 (2022): 1629-1647.</div> <div>[4] Notin, Pascal, et al. "TranceptEVE: Combining family-specific and family-agnostic models of protein sequences for improved fitness prediction." <em>bioRxiv</em> (2022): 2022-12.</div> <div>[5] Cheng, Jun, et al. "Accurate proteome-wide missense variant effect prediction with AlphaMissense." <em>Science</em> 381.6664 (2023): eadg7492.</div> <div>[6] Truong Jr, Timothy, and Tristan Bepler. "PoET: A generative model of protein families as sequences-of-sequences." <em>Advances in Neural Information Processing Systems</em> 36 (2024).</div> <div>[7] Lin, Zeming, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." <em>Science</em>379.6637 (2023): 1123-1130.</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

TA B L E 1 Summary of model fit, based on the area under the curve (AUC) of the receiver operating characteristic (ROC) for training data, and the most important bioclimatic variables in past, present, and future (2070) Maxent models of 13 bat species included in this study. in Southern Africa's Great Escarpment as an amphitheater of climate-driven diversification and a buffer against future climate change in bats

TA B L E 1 Summary of model fit, based on the area under the curve (AUC) of the receiver operating characteristic (ROC) for training data, and the most important bioclimatic variables in past, present, and future (2070) Maxent models of 13 bat species included in this study.

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 2. Maximum likelihood tree from 16S rRNA data under the best-fitting model T92 in Notes on the distribution and biology of northern brown shrimp Farfantepenaeus aztecus (Ives, 1891) in the eastern Mediterranean

Figure 2. Maximum likelihood tree from 16S rRNA data under the best-fitting model T92 + G. Numbers above branches indicate bootstrap values among 1000 replicates. Branches without bootstrap numbers mean that the bootstrap values are below 50%.

opencc-by-4.0May 2015View details →
zenodo40/100

Science ready spectra, their best-fitting models and results of Jeans axisymmetric modelling described in the research paper "Transforming gas-rich low-mass discy galaxies into ultra-diffuse galaxies by ram pressure" by Grishin, Chilingarian, Afanasiev et al.

<p>This package contains data presented in the paper &quot;Transforming gas-rich low-mass discy galaxies into ultra-diffuse galaxies by ram pressure&quot; by Grishin, Chilingarian, Afanasiev et al. (2021 Nature Astronomy in press). The dataset can be used to reproduce Figures 3 and 4 from the main manuscript and Extended Data Figures 1, 2, 4, 5 from the Supplementary Information.</p> <p>(1) Python scripts and data points required to reproduce Figure 4 in the manuscript and Extended Data Figure 5 from the Supplementary Information. The data and scripts are presented in a combined .zip archive for both figures.</p> <p>(2) One-dimensional spectra extracted within 1 half-light radius from long-slit Binospec spectra and multi-wavelength far-UV-to-near-IR broadband spectral energy distributions (SEDs) assembled from the photometric measurements extracted within the same aperture for 11 galaxies from the main sample (9 in the Coma cluster and 2 in the Abell 2147 cluster) and 5 galaxies from the supplementary (auxiliary) list. The spectra and SEDs are accompanied with their best-fitting stellar population models and parameters determined by the NBursts+phot algorithm: radial velocity, velocity dispersion, truncation age, final stellar metallicity. The templates are MILES-based models with self-consistent chemical evolution presented in Grishin et al. 2019 (https://ui.adsabs.harvard.edu/abs/2019arXiv190913460G/abstract). The filenames contain the coefficient for galactic winds and the mass fraction of stars in the final starburst, e.g. _l15_60 means lambda=1.5, SSP_frac=60 per cent. The files are presented as binary FITS tables with the fields annotated using unified content descriptors (UCDs) from the list established by the International Virtual Observatory Alliance and physical units where applicable.</p> <p>(3) Two-dimensional profiles of internal kinematics (radial velocity and velocity dispersion) and stellar population properties (truncation age and final stellar metallicity) derived from the analysis of long-slit Binospec spectra for 12 galaxies after adaptive binning, 11 from the main sample and GMP3016 in the Coma cluster from the supplementary sample; best-fitting Jeans axisymmetric models without adaptive binning, i.e. full profiles along the slit. The data are presented in binary FITS tables in the two FITS extensions, one for the profiles derived from the corresponding spectra and the second one for dynamical models.</p>

opencc-by-4.0Jun 2021View details →
dryad40/100

Prior choice and data requirements of Bayesian multivariate mixed effects models fit to tag-recovery data: The need for power analyses

<p>1. Recent empirical studies have quantified correlation between survival and recovery by estimating these parameters as correlated random effects with hierarchical Bayesian multivariate models fit to tag-recovery data. In these applications, increasingly negative correlation between survival and recovery has been interpreted as evidence for increasingly additive harvest mortality. The power of these hierarchal models to detect non-zero correlations has rarely been evaluated and these few studies have not focused on tag-recovery data, which is a common data type.</p> <p>2. We assessed the power of multivariate hierarchical models to detect negative correlation between annual survival and recovery. Using three priors for multivariate normal distributions, we fit hierarchical effects models to a mallard (<em>Anas</em> <em>platyrhychos</em>) tag-recovery dataset and to simulated data with sample sizes corresponding to different levels of monitoring intensity. We also demonstrate more robust summary statistics for tag-recovery datasets than total individuals tagged.</p> <p>3. Different priors lead to substantially different estimates of correlation from the mallard data. Our power analysis of simulated data indicated most prior distribution and sample size combinations could not estimate strongly negative correlation with useful precision or accuracy. Many correlation estimates spanned the available parameter space (–1,1) and underestimated the magnitude of negative correlation. Only one prior combined with our most intensive monitoring scenario provided reliable results. Underestimating the magnitude of correlation coincided with overestimating the variability of annual survival, but not annual recovery.</p> <p>4. The inadequacy of prior distributions and sample size combinations previously assumed adequate for obtaining robust inference from tag-recovery data represents a concern in the application of Bayesian hierarchical models to tag-recovery data. Our analysis approach provides a means for examining prior influence and sample size on hierarchical models fit to capture-recapture data while emphasizing transferability of results between empirical and simulation studies.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

Data and Code for Blaszczak et al. 2023, Models of underlying autotrophic biomass dynamics fit to daily river ecosystem productivity estimates improve understanding of ecosystem disturbance and resilience

<p>Data and code for analyses in Blaszczak&nbsp;et al. 2023, Models of underlying autotrophic biomass dynamics fit to daily river ecosystem productivity estimates improve understanding of ecosystem disturbance and resilience.</p> <p>See publication&nbsp;and ReadMe file for analysis description and further details.&nbsp;</p> <p>bioRxiv pre-print:&nbsp;Blaszczak, J.R., Yackulic, C., Shriver, R., &amp; R.O. Hall, Jr. 2023. Models of underlying autotrophic biomass dynamics fit to daily river ecosystem productivity estimates improve understanding of ecosystem disturbance and resilience.&nbsp;https://doi.org/10.1101/2023.04.11.535773</p>

opencc-by-4.0May 2023View details →
dryad40/100

Data for: Considerations for fitting occupancy models to data from eBird and similar volunteer-collected data

<p>An occupancy model makes use of data that are structured as sets of repeated visits to each of many sites, in order estimate the actual probability of occupancy (i.e., proportion of occupied sites) after correcting for imperfect detection using the information contained in the sets of repeated observations. We explore the conditions under which preexisting, volunteer-collected data from the citizen science project eBird can be used for fitting occupancy models. The data archived here are used to explore two ways in which the single-visit records could be used in occupancy models. First, we use empirical data contained within this archive to assess the potential for space-for-time substitution: aggregating single-visit records from different locations within a region into pseudo-repeat visits. The archived data are used to illustrate that the locations chosen for data collection by observers were not always representative of the habitat in the surrounding area, which would lead to biased estimates of occupancy probabilities when using space-for-time substitution. Second, create a large set of simulated data (output from the simulations contained in this archive) that we used to explore the utility of including data from single-visit records to supplement sets of repeated-visit data.</p>

opencc-zeroJul 2023View details →
zenodo40/100

Data presented in "Fitting cumulus cloud size distributions from idealized cloud resolving model simulations"

<p>Updated repository containing&nbsp;the data produced and analzed in the manuscript entitled &quot;Fitting cumulus cloud size distributions from idealized cloud resolving model simulations&quot; by J. Savre and G. Craig, submitted to the Journal of Advances in Modeling Earth Systems.</p> <p>New uploads: time series file (T_S), vertical profiles (profiles_tot.nc) and horizontal slices at 2km (slice_z_2000_bis.nc). Uploaded files concern the LBA case only.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Toward greater realism in inclusive fitness models: the case of caste fate conflict in insect societies

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad40/100

Prior choice and data requirements of Bayesian multivariate mixed effects models fit to tag-recovery data: The need for power analyses

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad40/100

Structural equation modeling reveals determinants of fitness in a cooperatively breeding bird

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad40/100

Revisiting the multispecies coalescent model fit with an example from a complete molecular phylogeny of the Liolaemus wiegmannii species group (Squamata: Liolaemidae)

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad40/100

Main model fits and substitution rate predictions for: A quantitative genetic model of background selection in humans

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad40/100

Data for: Considerations for fitting occupancy models to data from eBird and similar volunteer-collected data

Open the record for dataset details and reuse information.

publicJul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record