Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “Extrapolation”
Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method
Open the record for dataset details and reuse information.
Data from: An updated global dataset for diet preferences in terrestrial mammals: testing the validity of extrapolation
1. Diet is a key trait of an organism's life history that influences a broad spectrum of ecological and evolutionary processes. Kissling et al. (2014) compiled a species-specific dataset of diet preferences of mammals for 38% of a total of 5364 terrestrial mammalian species assessed for the International Union for Conservation of Nature's Red List, to facilitate future studies. The authors imputed dietary data for the remaining 62% by using extrapolation from phylogenetic relatives. 2. We collected dietary information for 1261 mammalian species for which data were extrapolated by Kissling et al. (2014), in order to evaluate the success with which such extrapolation can predict true diets. 3. The extrapolation method devised by Kissling et al. (2014) performed well for broad dietary categories (consumers of plants and animals). However, the method performed inconsistently, and sometimes poorly, for finer dietary categories, varying in accuracy in both dietary categories and mammalian orders. 4. The results of the extrapolation performance serve as a cautionary tale. Given the large variation in extrapolation performance, we recommend a more conservative approach for inferring mammalian diets, whereby dietary extrapolation is implemented only when there is a high degree of phylogenetic conservatism for dietary traits. Phylogenetic comparative methods can be used to detect and measure phylogenetic signal in diet. If data for species are needed, then only the broadest feeding categories should be used. This would ensure a greater level of accuracy and provide a more robust dataset for further ecological and evolutionary analysis.
Data from: Establishing macroecological trait datasets: digitalization, extrapolation, and validation of diet preferences in terrestrial mammals worldwide
Ecological trait data are essential for understanding the broad-scale distribution of biodiversity and its response to global change. For animals, diet represents a fundamental aspect of species' evolutionary adaptations, ecological and functional roles, and trophic interactions. However, the importance of diet for macroevolutionary and macroecological dynamics remains little explored, partly because of the lack of comprehensive trait datasets. We compiled and evaluated a comprehensive global dataset of diet preferences of mammals ("MammalDIET"). Diet information was digitized from two global and cladewide data sources and errors of data entry by multiple data recorders were assessed. We then developed a hierarchical extrapolation procedure to fill-in diet information for species with missing information. Missing data were extrapolated with information from other taxonomic levels (genus, other species within the same genus, or family) and this extrapolation was subsequently validated both internally (with a jack-knife approach applied to the compiled species-level diet data) and externally (using independent species-level diet information from a comprehensive continentwide data source). Finally, we grouped mammal species into trophic levels and dietary guilds, and their species richness as well as their proportion of total richness were mapped at a global scale for those diet categories with good validation results. The success rate of correctly digitizing data was 94%, indicating that the consistency in data entry among multiple recorders was high. Data sources provided species-level diet information for a total of 2033 species (38% of all 5364 terrestrial mammal species, based on the IUCN taxonomy). For the remaining 3331 species, diet information was mostly extrapolated from genus-level diet information (48% of all terrestrial mammal species), and only rarely from other species within the same genus (6%) or from family level (8%). Internal and external validation showed that: (1) extrapolations were most reliable for primary food items; (2) several diet categories ("Animal," "Mammal," "Invertebrate," "Plant," "Seed," "Fruit," and "Leaf") had high proportions of correctly predicted diet ranks; and (3) the potential of correctly extrapolating specific diet categories varied both within and among clades. Global maps of species richness and proportion showed congruence among trophic levels, but also substantial discrepancies between dietary guilds. MammalDIET provides a comprehensive, unique and freely available dataset on diet preferences for all terrestrial mammals worldwide. It enables broad-scale analyses for specific trophic levels and dietary guilds, and a first assessment of trait conservatism in mammalian diet preferences at a global scale. The digitalization, extrapolation and validation procedures could be transferable to other trait data and taxa.
Resources for the paper: "Social Context in Political Stance Detection: Impact and Extrapolation"
<p>This repository contains the resources in our paper <strong>[Social Context in Political Stance Detection: Impact and Extrapolation]</strong><br><em>Ramon Villa-Cox, Evan Williams, Kathleen M. Carley</em></p> <p>In this work, we explore the performance and extrapolation power of political stance-detection models using an existing large-scale weakly-labeled Twitter dataset collected around the 2019 South American Protests [1]. We construct transformer-based user and tweet encoders to embed users in a low-dimensional space using their text and ego-networks. We then train heterogeneous graph attention networks to predict user stances and contrast their ability to extrapolate stance predictions to different country contexts.</p> <p>The protest dataset, which was collected between September 25 and December 24 of 2019, contains 550k labeled users split unevenly across the four countries and contains over 36 million labeled tweets. It contains an additional 1.1 million unlabeled neighbors and 40 million unlabeled tweets. This repository includes the anonymized datasets necessary to reproduce the results and tables of the paper. In addition, we include the corresponding anonymized resources for the new weakly-labeled dataset around the 2020 Chilean Referendum presented in our paper.</p> <p>Following Twitter's January 2023 User Protection Policy update, tweet or user IDs related to sensitive political events cannot be publicly shared. We respect this policy, and only share:</p> <ul> <li>The anonymized user ID, their weak-stance label, the label predicted by each model and the data split (train, validation or test) the user was assigned to.</li> <li>Anonymized user network edges used by the different network classifiers</li> <li>The type of tweet the edge represents (Original, Reply, or Quote)</li> <li>The User Embeddings produced by the User Transformer and which serve as input for the different network models.</li> </ul> <p>This repository is comprised of the following files:</p> <ol> <li>Main_Predictions.7z: Compressed folder containing anonymized user IDs their stance label and each model’s prediction for the country it was trained on. The performance metrics for each model can be obtained based on the test split for each country. This folder includes the results for the Chilean Referendum.</li> <li>Cross_Predictions.7z: Compressed folder containing the results of the cross-country experiments for each anonymized user. The performance metrics for each model, when applied on a different country can be obtained based on each complete file. Users seen during the training of each model are excluded as described in the paper.</li> <li>Tweet_Level_Edgelists.7z: Compressed folder containing anonymized tweet edge lists indicating its interaction type (Original, Reply, or Quote).</li> <li>User_Networks.7z: Compressed folder containing different anonymized user edge lists for each interaction type.</li> <li>Embeddings.zip: Compressed Pytorch tensor files containing the User Embeddings produced by the User Transformer and which serve as input for the different network models. This are provided for the main results and the cross-country and referendum experiments.</li> </ol> <p>The code developed for this study is available at: https://github.com/rvillaco/Protest_Stance_Detection</p>
Discovery of thermostable fluorescently responsive glucose biosensors by structure-assisted function extrapolation
<p>Accurate assignment of protein function from sequence remains a fascinating and difficult challenge. The periplasmic binding protein (PBP) superfamily present an interesting case of function prediction, because they are both ubiquitous in prokaryotes, and they tend to diversify through gene duplication "explosions" that can lead to large numbers of paralogs in a genome. An engineered version of the moderately thermostable glucose-binding PBP from <i>Escherichia coli</i> has been used successfully as a reagentless fluorescent biosensor both <i>in vitro</i> and <i>in vivo</i>. To develop more robust sensors that meet the challenges of real-world applications, we report the discovery of thermostable homologs that retain a glucose-mediated conformationally coupled fluorescence response. Accurately identifying a glucose-binding PBP homolog among closely related paralogs is challenging. We demonstrate that a structure-based method that filters sequences by residues that bind glucose in an archetype structure is highly effective. Using fully sequenced bacterial genomes we found that this filter reduced high paralog numbers to single hits in a genome, consistent with accurate separation of glucose binding from other functions. We expressed engineered proteins for eight homologs, chosen to represent different degrees of sequence identity and tested their glucose-mediated fluorescence responses. We accurately predicted the presence of glucose binding down to 31% sequence identity. We also have successfully identified suitable candidates for next-generation robust, fluorescent glucose sensors.</p>
MIdAS bias adjustment of extremes using Theil-Sen extrapolation: Data and plotting scripts for GMD-publication
<p>When bias adjusting climate model data using quantile mapping approaches, one needs to prescribe what to do at the tails of the distribution, where a larger range of data is likely encountered outside the calibration period. The end results is highly dependent on the method used. For the current study, submitted to the journal Geoscientific Model Development, under the name 'Robust handling of extremes in quantile mapping - "Murder your darlings"' by Berg et al. (2024), the MIdAS bias adjustment method is evaluated and extended with additional functionality to deal with issues with bias adjustment of extreme precipitation. This entry contains data for the annual precipitation sums, and annual maximum daily precipitation, for a domain over Scandinavia, including a reference data set and a large ensemble of Euro-CORDEX regional climate models before and after bias adjustment using a range of experiments, as well as scripts for analysing the data and producing the figures of the paper. Further, the entry contains the daily timeseries for the reference and climate models needed to repeat all the experiments in the paper, along with the published code for MIdAS. The readme.txt documents provides a guide to structure the data, perform the experiments and to reproduce the plots of the paper.</p>
Figure 4. Samples from secondary forest were extrapolated from 13 in Saproxylic fly diversity in a Costa Rican forest mosaic
Figure 4. Samples from secondary forest were extrapolated from 13 to 16 samples. Dotted lines represent upper and lower 95% confidence intervals.
Evaluation of Methods for Extrapolating or Estimating the Size of Children in Pediatric Intensive Care
ClinicalTrials.gov study NCT03913247. IPD Sharing: NO. Countries: 4. Publications: 1.
Data from: Establishing macroecological trait datasets: digitalization, extrapolation, and validation of diet preferences in terrestrial mammals worldwide
Open the record for dataset details and reuse information.
Data from: Extrapolating multi-decadal plant community changes based on medium-term experiments can be risky: evidence from high-latitude tundra
Open the record for dataset details and reuse information.
Data from: An updated global dataset for diet preferences in terrestrial mammals: testing the validity of extrapolation
Open the record for dataset details and reuse information.
Discovery of thermostable fluorescently responsive glucose biosensors by structure-assisted function extrapolation
Open the record for dataset details and reuse information.
Fig. 52. Species estimation–extrapolated rarefaction curve. X in The millipede genus Leucogeorgia Verhoeff, 1930 in the Caucasus, with descriptions of eleven new species, erection of a new monotypic genus and notes on the tribe Leucogeorgiini (Diplopoda: Julida: Julidae)
Fig. 52. Species estimation–extrapolated rarefaction curve. X-axis: cave location with species of Leucogeorgiini (each cave–species record counts separately); Y-axis: estimated number of species.
Data from: Quantitative cross-species extrapolation between humans and fish: the case of the anti-depressant fluoxetine
Fish are an important model for the pharmacological and toxicological characterization of human pharmaceuticals in drug discovery, drug safety assessment and environmental toxicology. However, do fish respond to pharmaceuticals as humans do? To address this question, we provide a novel quantitative cross-species extrapolation approach (qCSE) based on the hypothesis that similar plasma concentrations of pharmaceuticals cause comparable target-mediated effects in both humans and fish at similar level of biological organization (Read-Across Hypothesis). To validate this hypothesis, the behavioural effects of the anti-depressant drug fluoxetine on the fish model fathead minnow (Pimephales promelas) were used as test case. Fish were exposed for 28 days to a range of measured water concentrations of fluoxetine (0.1, 1.0, 8.0, 16, 32, 64 µg/L) to produce plasma concentrations below, equal and above the range of Human Therapeutic Plasma Concentrations (HTPCs). Fluoxetine and its metabolite, norfluoxetine, were quantified in the plasma of individual fish and linked to behavioural anxiety-related endpoints. The minimum drug plasma concentrations that elicited anxiolytic responses in fish were above the upper value of the HTPC range, whereas no effects were observed at plasma concentrations below the HTPCs. In vivo metabolism of fluoxetine in humans and fish was similar, and displayed bi-phasic concentration-dependent kinetics driven by the auto-inhibitory dynamics and saturation of the enzymes that convert fluoxetine into norfluoxetine. The sensitivity of fish to fluoxetine was not so dissimilar from that of patients affected by general anxiety disorders. These results represent the first direct evidence of measured internal dose response effect of a pharmaceutical in fish, hence validating the Read-Across hypothesis applied to fluoxetine. Overall, this study demonstrates that the qCSE approach, anchored to internal drug concentrations, is a powerful tool to guide the assessment of the sensitivity of fish to pharmaceuticals, and strengthens the translational power of the cross-species extrapolation.
Supplementary material 1 from: Oishi EM, Kattler KR, Watkins HV, Howard BR, Côté IM (2024) Substrate complexity reduces prey consumption in functional response experiments: Implications for extrapolating to the wild. NeoBiota 91: 49-66. https://doi.org/10.3897/neobiota.91.111222
Supplementary data
Models from "Applying Machine Learning to Characterize and Extrapolate the Relationship Between Seismic Structure and Surface Heat Flow"
Open the record for dataset details and reuse information.
Dataset for Increasing the Measured Effective Quantum Volume with Zero Noise Extrapolation
Open the record for dataset details and reuse information.
Figure 2. Extrapolated rarefaction curve with 95 in The native bee fauna of the Palouse Prairie (Hymenoptera: Apoidea)
Figure 2. Extrapolated rarefaction curve with 95% confidence intervals based on all collected nonHemihalictus bees. Vertical line indicates the actual number of collected non-Hemihalictus bees.
Data from: Evaluating in vitro-in vivo extrapolation of toxicokinetics
Open the record for dataset details and reuse information.
Data from: Rarefaction and extrapolation with Hill numbers: a framework for sampling and estimation in species diversity studies
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.