Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “model violation”

Learn how ShareScore rates datasets ↗
zenodo48/100

Flavor-violating Higgs decays and stellar cooling anomalies in axion models

<p>We study a class of DFSZ-like models for the QCD axion that can address observed anomalies in stellar cooling. Stringent constraints from SN1987A and neutron stars are avoided by suppressed couplings to nucleons, while axion couplings to electrons and photons are sizable. All axion couplings depend on few parameters that also control the extended Higgs sector, in particular lepton flavor-violating couplings of the Standard Model-like Higgs boson&nbsp;h. This allows us to correlate axion and Higgs phenomenology, and we find that&nbsp;BR(h&nbsp;&rarr;&nbsp;&tau;e)&nbsp;can be as large as the current experimental bound of 0.22%, while&nbsp;BR(h&nbsp;&rarr;&nbsp;&mu;&mu;)&nbsp;can be larger than in the Standard Model by up to 70%. Large parts of the parameter space will be tested by the next generation of axion helioscopes such as</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Data from: Skyline fossilized birth-death model is robust to violations of sampling assumptions in total-evidence dating

<p>Several total-evidence dating studies under the fossilized birth-death (FBD) model have produced very old age estimates, which are not supported by the fossil record. This phenomenon has been termed "deep root attraction (DRA)". For two specific datasets, involving divergence time estimation for the early radiations of ants, bees and wasps (Hymenoptera) and of placental mammals (Eutheria), it has been shown that the DRA effect can be greatly reduced by accommodating the fact that extant species in these trees have been sampled to maximize diversity, so called diversified sampling. Unfortunately, current methods to accommodate diversified sampling only consider the extreme case where it is possible to identify a cut-off time such that all splits occurring before this time are represented in the sampled tree but none of the younger splits. In reality, the sampling bias is rarely this extreme, and may be difficult to model properly. Similar modeling challenges apply to the sampling of the fossil record. This raises the question of whether it is possible to find dating methods that are more robust to sampling biases. Here, we show that the skyline FBD (SFBD) process, where the diversification and fossil-sampling rates can vary over time in a piecewise fashion, provides age estimates that are more robust to inadequacies in the modeling of the sampling process and less sensitive to DRA effects. In the SFBD model we consider, rates in different time intervals are either considered to be independent and identically distributed, or assumed to be autocorrelated following an Ornstein-Uhlenbeck (OU) process. Through simulations and reanalyses of the Hymenoptera and Eutheria data, we show that both variants of the SFBD model unify age estimates under random and diversified sampling assumptions. The SFBD model can resolve DRA by absorbing the deviations from the sampling assumptions into the inferred dynamics of the diversification process over time. Although this means that the inferred diversification dynamics must be interpreted with caution, taking sampling biases into account, we conclude that the SFBD model represents the most robust approach available currently for addressing DRA in total-evidence dating.</p>

opencc-zeroApr 2022View details →
zenodo40/100

Hunting for vampires and other unlikely forms of parity violation at the Large Hadron Collider: calo-image datasets for the standard model

<p>An example calo-image dataset used in the <a href="https://arxiv.org/abs/2205.09876">paper</a>: standard model.</p> <p>Each shard is named `calo-image_sm_${DATASET}_${INDEX}.tar.gz`. Each contains one data file. DATASET is in {train,test,private_test} to label the three independent splits for training, validation, and testing respectively. INDEX labels separate batches which should be trivially combined.</p> <p>Each data file is in <a href="https://www.h5py.org/">h5</a> format. Its data are under the key &quot;entries&quot; as an array with shape (n, 32, 32) corresponding to (event_index, eta, phi) for the n calorimeter images in an unrolled eta--phi surface.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
dryad40/100

Data from: Skyline fossilized birth-death model is robust to violations of sampling assumptions in total-evidence dating

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad36/100

Supplementary Materials include results of simulation experiments to investigate the impact of phylogenetic regression with model violations.

<p>Modern comparative biology owes much to phylogenetic regression. At its conception, this technique sparked a revolution that armed biologists with phylogenetic comparative methods (PCMs) for disentangling evolutionary correlations from those arising from hierarchical phylogenetic relationships. Over the past few decades, the phylogenetic regression framework has become a paradigm of modern comparative biology that has been widely embraced as a remedy for shared ancestry. However, recent evidence has sown doubt over the efficacy of phylogenetic regression, and PCMs more generally, with the suggestion that many of these methods fail to provide an adequate defense against unreplicated evolution—the primary justification for using them in the first place. Importantly, some of the most compelling examples of biological innovation in nature result from abrupt lineage-specific evolutionary shifts, which current regression models are largely ill-equipped to deal with. Here we explore a solution to this problem by applying robust linear regression to comparative trait data. We formally introduce robust phylogenetic regression to the PCM toolkit with linear estimators that are less sensitive to model violations than the standard least-squares estimator, while still retaining high power to detect true trait associations. Our analyses also highlight an ingenuity of the original algorithm for phylogenetic regression based on independent contrasts, whereby robust estimators are particularly effective. Collectively, we find that robust estimators hold promise for improving tests of trait associations and offer a path forward in scenarios where classical approaches may fail. Our study joins recent arguments for increased vigilance against unreplicated evolution and a better understanding of evolutionary model performance in challenging–yet biologically important–settings.</p>

opencc-zeroMay 2024View details →
dryad36/100

Supplementary Materials include results of simulation experiments to investigate the impact of phylogenetic regression with model violations.

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad32/100

Data from: Impact of model violations on the inference of species boundaries under the multispecies coalescent

The use of genetic data for identifying species-level lineages across the tree of life has received increasing attention in the field of systematics over the past decade. The multispecies coalescent model provides a framework for understanding the process of lineage divergence, and has become widely adopted for delimiting species. However, because these studies lack an explicit assessment of model fit, in many cases, the accuracy of the inferred species boundaries are unknown. This is concerning given the large amount of empirical data and theory that highlight the complexity of the speciation process. Here, we seek to fill this gap by using simulation to characterize the sensitivity of inference under the multispecies coalescent to several violations of model assumptions thought to be common in empirical data. We also assess the fit of the multispecies coalescent model to empirical data in the context of species delimitation. Our results show substantial variation in model fit across datasets. Posterior predictive tests find the poorest model performance in datasets that were hypothesized to be impacted by model violations. We also show that while the inferences assuming the multispecies coalescent are robust to minor model violations, such inferences can be biased under some biologically plausible scenarios. Taken together, these results suggest that researchers can identify individual datasets in which species delimitation under the multispecies coalescent is likely to be problematic, thereby highlighting the cases where additional lines of evidence to identify species boundaries are particularly important to collect. Our study supports a growing body of work highlighting the importance of model checking in phylogenetics, and the usefulness of tailoring tests of model fit to assess the reliability of particular inferences.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Impact of model violations on the inference of species boundaries under the multispecies coalescent

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad28/100

Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods

The increasing ability to extract and sequence DNA from non-contemporaneous tissue offers biologists the opportunity to analyze ancient DNA (aDNA) together with modern DNA (mDNA) to address the taxonomy of extinct species, evolutionary origins, historical phylogeography and biogeography. Perhaps more exciting are recent developments in coalescence-based Bayesian inference that offer the potential to use temporal information from aDNA and mDNA for the estimation of substitution rates and divergence dates as an alternative to fossil and geological calibration. This comes at a time of growing interest in the possibility of time dependency for molecular rate estimates. Here we provide a critical assessment of Bayesian MCMC analysis for the estimation of substitution rate using simulated samples of aDNA and mDNA. We conclude that the current models and priors employed in Bayesian MCMC analysis of heterochronous mtDNA are susceptible to an upward bias in the estimation of substitution rates due to model misspecification when the data comes from populations with less than simple demographic histories, including sudden short-lived population bottlenecks or pronounced population structure. However when model misspecification is only mild, then the 95% HPD intervals provide adequate frequentist coverage of the true rates.

opencc-zeroDec 2009View details →
dryad28/100

Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods

Open the record for dataset details and reuse information.

publicMar 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record