Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “Extrapolation”
Annual time series of global VIIRS nighttime lights for 2000-2024 at 500-m spatial resolution extrapolated using logistic regression
<p>The <a href="https://eogdata.mines.edu/products/vnl/"><strong>Annual Visible Night Light (VNL) V2</strong></a> (VIIRS) images at 500-m spatial resolution for the period 2012 to 2024 (Elvidge et al., 2021) have been used to extrapolate the values backwards for years 2000–2011. This was done by fitting a logistic regression (per pixel) and then predicting the values for the previous years (see nightlights_stack_500m.R). After consistent time-series have been produced, I also derived the difference between year 2024 and year 2000 (nightlights.difference_viirs.v21_m_500m_s_2000_2024_go_epsg4326_v20230318.tif): this shows average rate of change for the 25 years period. Use with caution: extrapolation of values can lead to artifacts. For most of the land surface, however, it appears that the growth of night lights follows exponential growth function and hence nights in the past can be represented accurately by fitting decay / logistic regression function.</p> <p>Original values from the Annual VNL V2 product have been converted from 0–200 to 0–2000 scale and are available as Cloud-Optimized GeoTIFFs.</p> <p>Principal components (PC1, PC2, PC3, PC4) were derived using SAGA GIS (sums-of-squares-and-cross-products matrix) method. The first PC1 usually matches the long-term mean value, PC2 matches the 1st derivation in values. File "nightlights_dmsp.v10_m_1km_s_19920101_20241231_go_epsg4326_v20251006.tif" contains 33 years 1992 to 2024, but at 1 km resolution.</p> <p>To cite the Annual VNL V2, please use:</p> <ul> <li>Elvidge, C. D., Zhizhin, M., Ghosh, T., Hsu, F. C., & Taneja, J. (2021). <a href="https://doi.org/10.3390/rs13050922">Annual time series of global VIIRS nighttime lights derived from monthly averages: 2012 to 2019</a>. Remote Sensing, 13(5), 922. https://doi.org/10.3390/rs13050922</li> </ul> <p>Historic night light images (1 km resolution) are also available from <a href="https://doi.org/10.6084/m9.figshare.9828827.v10">Figshare</a>:</p> <ul> <li>Li, X., Zhou, Y., Zhao, M., & Zhao, X. (2020). <a href="https://doi.org/10.1038/s41597-020-0510-y">A harmonized global nighttime light dataset 1992–2018</a>. Scientific data, 7(1), 168. https://doi.org/10.1038/s41597-020-0510-y</li> </ul>
Plasma-Prescribed Active Region Static Extrapolation Dataset
<p>A repository of extrapolations using RBF-FD Magnetohydrostatic techniques with imposed plasmas, based on the SHARP solar image photospheric magnetic field library.</p>
The Plasma-Prescribed Active Region Static Extrapolation Dataset
<p>The Plasma-Prescribed Active Region Static Extrapolation (PARSE) Dataset consists of approximately seven thousand magnetohydrostatic extrapolations of solar active regions for use in statistical or machine learning applications. The extrapolations are based on the Spaceweather HMI Active Region Patch (SHARP) library (doi <a href="https://doi.org/10.1007/s11207-014-0529-3">10.1007/s11207-014-0529-3</a>), and the magnetohydrostatic extrapolation is performed by the routine detailed in Mathews et al 2022 (doi <a href="https://doi.org/10.1016/j.jcp.2022.111214">10.1016/j.jcp.2022.111214</a>). </p>
Global Mean Sea Level, Trajectory and Extrapolation
<p>Global Mean Sea Level, Trajectory and Extrapolation</p> <p>This file contains Global Mean Sea Level (GMSL) variations along, data for the quadratic fit (trajectory) to the GMSL variations, and an extrapolation of this trajectory to 2050.</p> <p>Column 1 provides the calendar year plus the decimal fraction of the current year. The GMSL variations(column 2) are computed at the NASA Goddard Space Flight Center under the auspices of the NASA Sea Level Change program. All units for sea level are in centimeters The GMSL was generated using the NASA-SSH Simple Gridded Sea Surface Height from Standardized Reference Missions Version 1: https://podaac.jpl.nasa.gov/dataset/NASA_SSH_REF_SIMPLE_GRID_V1. It combines Sea Surface Heights from the TOPEX/Poseidon, Jason-1, OSTM/Jason-2, HDR Jason-3, and Sentinel-6 Michael Freilich missions.</p> <p>In addition, the rate and acceleration are estimated from full record of GMSL relative to the midpoint of the record and then used to generate a quadratic fit to the data. This quadratic fit is provided in column 3. The rate associated with this quadratic fit at any time in the record is also provided (column 4). </p> <p>The parameters estimated from the quadratic fit are also used to generated an extrapolated time series out to 2050 (column 5). These are provided at yearly intervals. This is not a projection and is only considered an extrapolation of the current trajectory of GMSL variations. This also differs from Nerem et al. (2022) and Sweet et al. (2022) as additional signals are not removed from GMSL prior to estimating the rate and acceleration parameters. The yearly rate associated with this extrapolation is also provided (column 6).<br><br>If you use these data please cite:<br>Willis, J.K., Hamlington, B.D., and Fournier, S., Global Mean Sea Level Time Series, Trajectory and Extrapolation. Dataset access [YYYY-MM-DD] at 10.5281/zenodo.7702314.</p> <p>References:</p> <p>Nerem, R. S., Frederikse, T., & Hamlington, B. D. (2022). Extrapolating Empirical Models of Satellite‐Observed Global Mean Sea Level to Estimate Future Sea Level Change. <em>Earth's Future</em>, <em>10</em>(4), e2021EF002290.</p> <p>Sweet, W. V., Hamlington, B. D., Kopp, R. E., Weaver, C. P., Barnard, P. L., Bekaert, D., ... & Zuzak, C. (2022). <em>Global and regional sea level rise scenarios for the United States: updated mean projections and extreme water level probabilities along US coastlines</em>. Interagency Technical Report.</p>
Data for paper "Zone extrapolations in parametric timed automata"
<p>Data for paper "<em>Zone extrapolations in parametric timed automata</em>"</p> <p>This data set comes with two zip files:</p> <ol> <li><strong>artifact.zip</strong> contains the current version of IMITATOR, all models and necessary scripts to reproduce all experiments on our benchmarks set. A file named README.md gives all instructions for reproducibility.</li> <li><strong>results.zip</strong> contains an HTML page with a table summarizing all results, and all raw results (.res) as well.</li> </ol> <p> </p>
Data from: In vitro to in vivo extrapolation from three-dimensional hiPSC-derived cardiac microtissues and physiologically based pharmacokinetic modeling to inform next-generation arrythmia risk assessment
<p>Proarrhythmic cardiotoxicity remains a substantial barrier to drug development as well as a major global health challenge. <em>In vitro</em> human pluripotent stem cell-based new approach methodologies have been increasingly proposed and employed as alternatives to existing <em>in vitro</em> and <em>in vivo</em> models that do not accurately recapitulate human cardiac electrophysiology or cardiotoxicity risk. In this study, we expanded the capacity of our previously established three-dimensional human cardiac microtissue model to perform quantitative risk assessment by combining it with a physiologically based pharmacokinetic model, allowing a direct comparison of potentially harmful concentrations predicted <em>in vitro</em> to <em>in vivo</em> therapeutic levels. This approach enabled the measurement of concentration responses and margins of exposure for two physiologically relevant metrics of proarrhythmic risk (<em>i.e.</em>, action potential duration and triangulation assessed by optical mapping) across concentrations spanning three orders of magnitude. The combination of both metrics enabled accurate proarrhythmic risk assessment of four compounds with a range of known proarrhythmic risk profiles (<em>i.e., </em>quinidine, cisapride, ranolazine, and verapamil) and demonstrated close agreement with their known clinical effects. Action potential triangulation was found to be a more sensitive metric for predicting proarrhythmic risk associated with the primary mechanism of concern for pharmaceutical-induced fatal ventricular arrhythmias, delayed cardiac repolarization due to inhibition of the rapid delayed rectifier potassium channel, or hERG channel. This study advances human induced pluripotent stem cell-based three-dimensional cardiac tissue models as new approach methodologies that enable <em>in vitro</em> proarrhythmic risk assessment with high precision of quantitative metrics for understanding clinically relevant cardiotoxicity.</p>
Fig. 4. Rarefaction curves with extrapolations and 95 in Diversity and microhabitat use of benthic invertebrates in an urban forest stream (Southeastern Brazil)
Fig. 4. Rarefaction curves with extrapolations and 95% of confidence intervals for both the total sampling and each type of microhabitat of benthic invertebrates of Tijuca River, located at the Tijuca Forest, Rio de Janeiro, Brazil.
Comparison of physiologically based pharmacokinetic modeling platforms for developmental neurotoxicity in vitro to in vivo extrapolation
Open the record for dataset details and reuse information.
Data from: In vitro to in vivo extrapolation from three-dimensional hiPSC-derived cardiac microtissues and physiologically based pharmacokinetic modeling to inform next-generation arrythmia risk assessment
Open the record for dataset details and reuse information.
Deep Well transect plant-soil feedback experiment extrapolated biomass data (2014-2016)
This project was conducted from July 2014 to July 2016 to measure plant-soil feedbacks between black grama (Bouteloua eriopoda) and blue grama (Bouteloua gracilis) in patches of different historical stability along the Deep Well transect at the Sevilleta LTER.
Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method
<p>Species distribution and ecological niche models (hereafter SDMs) are popular tools with broad applications in ecology, biodiversity conservation, and environmental science. Many SDM applications require projecting models in environmental conditions non-analog to those used for model training (extrapolation), giving predictions that may be statistically unsupported and biologically meaningless. We introduce a novel method, Shape, a model-agnostic approach that calculates the extrapolation degree for a given projection data point by its multivariate distance to the nearest training data point. Such distances are relativized by a factor that reflects the dispersion of the training data in environmental space. Distinct from other approaches, Shape incorporates an adjustable threshold to control the binary discrimination between acceptable and unacceptable extrapolation degrees. We compared Shape's performance to five extrapolation metrics based on their ability to detect analog environmental conditions in environmental space and improve SDMs suitability predictions. To do so, we used 760 virtual species to define different modeling conditions determined by species niche tolerance, distribution equilibrium condition, sample size, and algorithm. All algorithms had trouble predicting species niches. However, we found a substantial improvement in model predictions when model projections were truncated independently of extrapolation metrics. Shape's performance was dependent on extrapolation threshold used to truncate models. Because of this versatility, our approach showed similar or better performance than the previous approaches and could better deal with all modeling conditions and algorithms. Our extrapolation metric is simple to interpret, captures the complex shapes of the data in environmental space, and can use any extrapolation threshold to define whether model predictions are retained based on the extrapolation degrees. These properties make this approach more broadly applicable than existing methods for creating and applying SDMs. We hope this method and accompanying tools support modelers to explore, detect, and reduce extrapolation errors to achieve more reliable models.</p>
Data for manuscript: An extrapolation algorithm for estimating river bed grain size distributions across basins
<p>Pebble counts collected and used for the analysis presented in the manuscript: An extrapolation agorithm for estimating river bed grain size distributions across drainage basins.</p>
Data from: Extrapolating potential crop damage by insect pests based on land use data: examining inter-regional generality in agricultural landscapes_210907
<p>DamagePrediction_data_2021_210907 Data from: Extrapolating potential crop damage by insect pests based on land use data: examining inter-regional generality in agricultural landscapes</p>
Data for Floor heating pre-on/off parameters based on Model Predictive Control feature extrapolation Paper in CLIMA2022 conference proceedings
<p>This is a collection of time series results used to obtain all the results shown in the paper. The tags of the .csv or .mat files are self explanatory and easy to use.</p>
Dataset for Non-resonant Anomaly Detection with Background Extrapolation
<p>These are the datasets used in the journal version of the Non-resonant Anomaly Detection with Background Extrapolation paper. The datasets are simulated using MadGraph5 aMC@NLO, Pythia 8.310, and Delphes. There are 0.2M signal events of semi-visible jets in five sets of parameters (invisible-ratio, Z' mass) = { (1/3, 4 TeV), (1/3, 2 TeV), (1/3, 3 TeV), (0, 4 TeV), (2/3, 4 TeV) }, 18.6M background events of SM QCD jets (including background, ideal AD background, and simulated background) for training, and 21.4M background events for testing. The detailed breakdown of number of events after selections in different regions is listed in Table1 of the paper. The input parameter cards used for generating background and signal events are also included.</p>
Validating additive correction schemes against gradient-based extrapolations
<p>Supporting information for the paper.<br> <br> Python files are used to generate figures, which will be placed in the "figures" folder. The table include files in "si" are also automatically generated.</p> <p>Each of the other folders includes all data (inputs, shell scripts and outputs) of the Psi4 calculations for a given species, with the exception of the NCDT16 folder, which is a benchmark code developed previously.</p>
Dataset for "On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"
<h1>Code and Dataset for "On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"</h1> <p>This dataset accompanies the research paper titled <strong>"On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"</strong>, currently under review for the AGU Journal GRL. The study introduces a novel Regional Climate Model (RCM) emulator focusing on high-resolution climate downscaling for the New Zealand region. For additional insights and access to the codebase utilized in this research, please refer to our <a href="https://github.com/nram812/On-the-Extrapolation-of-Generative-Adversarial-Networks-for-downscaling-precipitation-extremes">Github Repository</a>.</p> <p>The code can also be found as a ".zip" file: *On-the-Extrapolation-of-Generative-Adversarial-Networks-for-downscaling-precipitation-extremes-main. </p> <h2>Aims</h2> <p>Our study focuses on two important gaps in the literature regarding the extrapolation of empirical downscaling algorithms. First, we examine how well relationships learned from a historical period extrapolate to future unobserved climates. We compare two widely used algorithms, a GAN and a deterministic CNN baseline, that use a similar architecture (i.e. convolutional layers) trained in a model-as-truth framework to downscale daily precipitation over New Zealand. We evaluate their accuracy in capturing climate change signals in mean and extreme precipitation. Second, we explore whether training on future vs. only historical periods combined with different-sized training datasets can improve extrapolation skill. </p> <h2>Geographic Focus</h2> <p>Our research focuses only on the New Zealand Region (165°E-184°W, 33°S-51°S).</p> <p> </p> <h2>Data Overview</h2> <h3>Training and Evaluation Data</h3> <p>The training data used in this study (for our RCM emulator) spans the historical period and future period (SSP370) of simulation. It comprises daily accumulated precipitation as the primary target variable, alongside large-scale predictor variables. </p> <ul> <li> <p><strong>Resolution:</strong> The target variable is presented at a 12km resolution, reflecting the highest resolution face of RCM for the New Zealand region. Predictor variables are coarsened to a 1.5-degree resolution from original CCAM outputs using conservative interpolation. </p> </li> <li> <p><strong>Period Coverage:</strong></p> <ul> <li>Training Data: 1960-2100 (Depending on Experiment, see Table 1 for list of experiment configurations)</li> <li>Validation Data: 1985-2014 + 2070-2099 (to compute the climate change signal)</li> </ul> </li> <li> <p><strong>Models:</strong></p> <ul> <li>Training on: ACCESS-CM2</li> <li>Validated on: EC-Earth3, NorESM2-MM, CNRM-CM6-1, AWI-MR-1 </li> </ul> </li> </ul> <h3>File Structure</h3> <ul> <li> <p><strong>Training Data:</strong></p> <ul> <li>Target/Ground Truth (Y): <code>target_ACCESS-CM2_hist_ssp370_pr.nc</code></li> <li>Predictor (X): <code>predictor_ACCESS-CM2_hist_ssp370.nc</code></li> </ul> </li> <li> <p><strong>Evaluation Data:<br></strong>All other GCMs can be accessed in one single file, predictor and target variables have the dimensions (time, lat, lon, GCM).</p> <ul> <li>Target/Ground Truth (Y): <code>Other_GCMs_hist_SSP370_target_fields_pr.nc</code></li> <li>Predictor (X): <code>Other_GCMs_hist_SSP370_predictor_fields.nc</code></li> </ul> </li> </ul> <h2>Methodological Insights</h2> <ul> <li> <p><strong>Regional Climate Model</strong>, Our Regional Climate Model training data is from the Conformal Cubic Atmospheric Model (CCAM) which is a global non-hydrostatic atmospheric model renowned for its variable-resolution cubic grid. . For more information about CCAM, please see the following <a href="https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2023JD038530">paper</a>.</p> </li> <li> <p><strong>Predictor and Target Variables:</strong> Daily-averaged large-scale prognostic variables, including zonal wind, meridional wind, temperature, and specific humidity, are employed as predictors at the 500mb and 850mb pressure levels. These are normalized (see the GitHub repository for the mean and standard deviation fields). Precipitation is taken as is from CCAM and accumulated for each given day. Static predictors are also used in our model, which is stored in a GitHub repository.</p> </li> <li> <p><strong>Training Framework:</strong> Our dataset benefits from the "perfect framework" training strategy, which uses CCAM-coarsened predictor variables. For more information about the perfect and imperfect training frameworks, see the following <a title="review" href="https://journals.ametsoc.org/view/journals/aies/3/2/AIES-D-23-0066.1.xml">review</a></p> </li> </ul> <table> <tbody> <tr> <td> <p><strong>Algorithm</strong></p> </td> <td> <p><strong>Training Data</strong></p> </td> <td> <p><strong>Period</strong></p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Historical</p> </td> <td> <p>1960-2014 (~21,000 days)</p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Future (SSP370)</p> </td> <td> <p>2044-2099 (~21,000 days)</p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Historical and Future (SSP370)</p> </td> <td> <p>1960-2099 (~51,000 days)</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Historical</p> </td> <td> <p>1960-2014</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Future (SSP370)</p> </td> <td> <p>2044-2099</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Historical and Future (SSP370)</p> </td> <td> <p>1960-2099</p> </td> </tr> </tbody> </table> <p><strong>Table 1:</strong> The six RCM emulator experiments performed in this study.</p>
rarefaction and extrapolation of sample coverage based on sample-based abundance data
<p>#R code.txt is the R code for plotting figures and constructing Tables</p> <p>#bciabun1010.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 10m*<em>10m</em></p> <p><em>#bciabun2020.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 20m*</em>20m</p> <p><em>#bciabun5050.txt is the species_by_plot matrix that BCI forest Plot is divided by a plot with size 50m*5</em>0m</p> <p>#fus10.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 10m*<em>10m</em></p> <p><em>#fus20.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 20m*</em>20m</p> <p>#fus50.txt is the species_by_plot matrix that Fushan forest Plot is divided by a plot with size 50m*<em>50m</em></p> <p><em>#lhc10.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 10m*</em>10m</p> <p>#lhc20.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 20m*20m</p> <p>#lhc50.txt is the species_by_plot matrix that Lianhuachi forest Plot is divided by a plot with size 50m*50m</p>
S1314, Co-expression Extrapolation (COXEN) Program to Predict Chemotherapy Response in Patients With Bladder Cancer
ClinicalTrials.gov study NCT02177695. IPD Sharing: YES. Countries: 1. Publications: 3.
Data from: Sampling strategy matters: eDNA-based assessment and extrapolation of myxozoan diversity in a model stream system
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.