Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

753

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

753 results for “metrics”

Learn how ShareScore rates datasets ↗
edi56/100

Data from: Invasion timing affects multiple scales, metrics and facets of biodiversity outcomes in ecological restoration experiments (Missouri, 2009-2016)

Vegetation responses to experimental ecological restoration treatments at Tyson Research Centre of Washington University in Missouri, USA. These data include species-level cover responses to various factorial restoration treatments. Treatments were applied starting in 2009 and were measured in 2016. Treatment responses reflect these long term responses, but the dataset is comprised to one time point.

openCC (other)May 2025View details →
edi56/100

Dataset on tree diversity metrics and aboveground carbon storage in Southeastern U.S. oak-pine forests, 2009–2019

This dataset contains measurements of tree structural and taxonomic diversity, stand attributes, and aboveground carbon storage from mixed oak-pine forests in Florida, Georgia, and Alabama, located in the southeastern United States. Data were collected from 946 mixed oak-pine, and 7224 longleaf-slash pine and oak-pine Forest Inventory and Analysis (FIA) plots respectively spanning the years 2009 to 2019. Variables include aboveground carbon, aboveground biomass, tree density, basal area, stand age, and diversity metrics such as Shannon indices for tree species and diameter-based structural classes. Functional diversity metrics—including functional dominance and functional divergence—are also included. These data were used to support a published study examining the interactive effects of diversity metrics on carbon storage using structural equation modeling. The geographic coverage represents humid subtropical forest regions of the southeastern U.S.

openCC (other)May 2025View details →
edi56/100

SBC LTER: Ocean: HFR-derived surface flow metrics, surface water retention times, and related factors in the Santa Barbara Channel (2012-2019)

This data package include three files: 1. daily maps of High-Frequency Radar (HFR) measured surface currents, indices of mesoscale eddy locations, and local retention times on a 2km grid; 2. monthly time series of wind stress, alongshore pressure gradient, surface current EOF principal components, vorticity, eddy area, eddy presence, and spatially averaged retention times from January 2012 to December 2019; 3. A MATLAB script for plotting the maps and timeseries. These data were processed in order to investigate the drivers of surface water retention in the Santa Barbara Channel, CA, details of which are available in the study: Brokaw, R.J., D.A. Siegel, and L. Washburn. Physical Drivers of Surface Water Retention in the Santa Barbara Channel. [In preparation for Journal of Geophysical Research: Oceans.]

openCC (other)Mar 2024View details →
zenodo52/100

EOSC Task Force on FAIR Metrics and Data Quality: FAIR Evaluation community survey 2023

<p>The EOSC-A FAIR Metrics and Data Quality Task Force (TF) supported the European Open Science Cloud Association (EOSC-A) by providing strategic directions on FAIRness (Findable, Accessible, Interoperable, and Reusable) and data quality. The Task Force conducted a survey&nbsp;using the <a href="https://ec.europa.eu/eusurvey/">EUsurvey tool</a> between 15.11.2022 and 18.01.2023, targeting both developers and users of FAIR assessment tools. The survey aimed at supporting the harmonisation of FAIR assessments, in terms of what it evaluated and how, across existing (and future) tools and services, as well as explore if and how a community-driven governance on these FAIR assessments would look like. The survey received 78 responses, mainly from academia, representing various domains and organisational roles. This is the anonymised survey dataset in csv format; most open-ended answers have been dropped. The codebook contains variable names, labels, and frequencies.</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Datasets for testing the robustness of LiDAR vegetation metrics to varying point densities

<p><span>The calculation of vegetation metrics from LiDAR point clouds might be affected by the available point density of a dataset. Testing how the same LiDAR vegetation metrics differ with different point densities can therefore inform about their robustness for upscaling metrics to other areas or other LiDAR point clouds. The datasets made available here were generated to test the robustness of LiDAR vegetation metrics to varying point densities and spatial resolutions (i.e., plots of 1 &times; 1 m, 2 &times; 2 m, 5 &times; 5 m and 10 &times; 10 m size). A total of 25 LiDAR vegetation metrics representing different aspects of vegetation height, vegetation cover and structural complexity were tested (see metric definition in Kissling et al. 2023, </span><span><a href="https://doi.org/10.1016/j.dib.2022.108798"><span>https://doi.org/10.1016/j.dib.2022.108798</span></a></span><span>). The metric calculation was similar to the metric calculation in the Laserchicken software (Meijer et al. 2020, </span><span><a href="https://doi.org/10.1016/j.softx.2020.100626"><span>https://doi.org/10.1016/j.softx.2020.100626</span></a></span><span>) and the Laserfarm workflow (Kissling et al. 2022, https://doi.org/10.1016/j.ecoinf.2022.101836). The Dutch AHN4 dataset from the years 2020&ndash;2022 with a point density of 20&ndash;30 points/m<sup>2</sup> was used. Initially, 100 plots (i.e., squared polygons around centre points) were randomly placed across the Netherlands in Dutch Natura 2000 sites that predominantly contain woodland habitats (using shapefiles from the European Environmental Agency). For each centre point, square polygons of the desired resolutions (i.e., 1 &times; 1 m, 2 &times; 2 m, 5 &times; 5 m or 10 &times; 10 m plot size) were generated. The square polygons were subsequently used to clip the LiDAR point clouds from the Dutch AHN4 point cloud dataset. Since not all locations of the 100 randomly placed plots contained points, the actual sample sizes were slightly smaller than 100, i.e., 94 plots for the 1 &times; 1 m, 2 &times; 2 m and 5 &times; 5 m resolution and 95 plots for the 10 &times; 10 m resolution. Metrics were calculated with the original point density of the Dutch AHN4 dataset (20&ndash;30 points/m2) and with six systematically down-sampled point clouds for the same plots (i.e., keeping 5%, 10%, 20%, 40%, 60% and 80% of the points in the original point clouds). For each clipped point cloud of a plot at a given resolution, the points were first sorted according to their GPS acquisition time (from earliest to latest). Points were then systematically discarded and only 5%, 10%, 20%, 40%, 60% and 80% of the points in the original point clouds were kept. The kept points were used for calculating the 25 LiDAR vegetation metrics. </span></p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction [dataset]

<p>This dataset contains the extension of a publicly available dataset that was published initially by Ferenc et al. in their paper:</p> <p><em>&ldquo;Ferenc, R.; Hegedus, P.; Gyimesi, P.; Antal, G.; B&aacute;n, D.; Gyim&oacute;thy, T. Challenging machine learning algorithms in predicting vulnerable javascript functions. 2019 IEEE/ACM 7th InternationalWorkshop on Realizing Artificial Intelligence Synergies in Software Engineering (RAISE). IEEE, 2019, pp. 8&ndash;14.&rdquo;</em></p> <p>The dataset contained software metrics for source code functions written in JavaScript (JS) programming language. Each function was labeled as vulnerable or clean. The authors gathered vulnerabilities from publicly available vulnerability databases.</p> <p>In our paper entitled: &ldquo;<strong>Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction</strong>&rdquo; and cited as:</p> <p><em>&ldquo;Kalouptsoglou I, Siavvas M, Kehagias D, Chatzigeorgiou A, Ampatzoglou A. Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction. Entropy. 2022; 24(5):651. <a href="https://doi.org/10.3390/e24050651">https://doi.org/10.3390/e24050651</a>&rdquo;</em></p> <p>, we presented an extended version of the dataset by extracting textual features for the labeled JS functions. In particular, we got the dataset provided by Ferenc et al. in CSV format and then we gathered all the GitHub URLs of the dataset&#39;s functions (i.e., methods). Using these URLs, we collected the source code of the corresponding JS files from GitHub. Subsequently, by utilizing the start and end line information for every function, we cut off the code of the functions. Each function was then tokenized to construct a list of tokens per function.</p> <p>To extract text features, we used a text mining technique called sequences of tokens. As a result, we created a repository with all methods&#39; source code, the token sequences of each method, and their labels. To boost the generalizability of type-specific tokens, all comments were eliminated, as well as all integers and strings, which were replaced with two unique IDs.</p> <p>The dataset contains 12,106 JavaScript functions, from which 1,493 are considered vulnerable.</p> <p>This dataset was created and utilized during the Vulnerability Prediction Task of the Horizon2020 IoTAC Project&nbsp;as training and evaluation data for the construction of vulnerability prediction models. The dataset is provided in the csv format. Each row of the csv file has the following parts:</p> <ul> <li>Label: Flag with values &lsquo;1&rsquo; for vulnerable and &lsquo;0&rsquo; for non-vulnerable methods</li> <li>Name: The name of the JavaScript method</li> <li>Longname: The longname of the JavaScript method</li> <li>Path: The path of the file of the method in the repository</li> <li>Full_repo_path: The GitHub URL of the file of the method</li> <li>TokenX: Each next row corresponds to each token included in the method</li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo52/100

Climate change velocity metrics calculated for three climate variables across Finland

<p>This dataset contains files that show the climate change velocity metrics calculated for three climate variables across Finland. The climate velocities were used to study the magnitude of projected climatic changes in a nation-wide Natura 2000 protected area (PA) network (Heikkinen et al., 2020). Using fine-resolution climate data that describes the present-day and future topoclimates and their spatio-temporal variation, the study explored the rate of climatic changes in protected areas on an ecologically relevant, but yet poorly explored scale. The velocities for the three climate variables were developed in the following work, where in-depth description of the different steps in velocity metrics calculation and a number of visualisations of their spatial variation across Finland are provided:</p><p>Risto K. Heikkinen 1, Niko Leikola 1, Juha Aalto 2,3, Kaisu Aapala 1, Saija Kuusela 1, Miska Luoto 2 &amp; Raimo Virkkala 1 2020: Fine-grained climate velocities reveal vulnerability of protected areas to climate change. Scientific Reports 10:1678. https://doi.org/10.1038/s41598-020-58638-8</p><p>1 Finnish Environment Institute, Biodiversity Centre, Latokartanonkaari 11, FI-00790 Helsinki, Finland</p><p>2 Department of Geosciences and Geography, University of Helsinki, FI-00014, Helsinki, Finland</p><p>3 Finnish Meteorological Institute, FI-00101, Helsinki, Finland&nbsp;</p><p>The dataset includes GIS compatible geotiff files describing the nine spatial climate velocity surfaces calculated across the whole of Finland at 50 m × 50 m spatial resolution. These nine different velocity surfaces consist of velocity metric values measured for each 50-m grid cell separately for the three different climate variables and in relation to the three different future climate scenarios (RCP2.6, RCP4.5 and RCP8.5). The baseline climate data for the study were the monthly temperature and precipitation data averaged for the period from 1981 to 2010 modelled at a resolution of 50-m, based on which estimates for the annual temperature sum above 5 °C (growing degree days, GDD, °C), the mean January temperature (TJan, °C) and the annual climatic water balance (WAB, the difference between annual precipitation and potential evapotranspiration; mm) were calculated. Corresponding future climate surfaces were produced using an ensemble of 23 global climate models for the years 2070–2099 (Taylor et al. 2012) and the three RCPs. The data for the three climate variables for 1981–2010 and under the three RCPs will be made available in separately via METIS - FMI's Research Data repository service (Aalto et al., in prep.).&nbsp;</p><p>The climate velocity surfaces included in the present data repository were developed using climate-analog approach (Hamann et al. 2015; Batllori et al. 2017; Brito-Morales et al. 2018), whereby velocity metrics for the 50-m grid cells were measured based on the distance between climatically similar cells under the baseline and the future climates, calculated separately for the three climate variables. In Heikkinen et al. (2020), the spatial data for the Natura 2000 protected areas were used to assess their exposure to climate change. The full data on N2K areas can be downloaded from the following link: https://ckan.ymparisto.fi/dataset/%7BED80465E-135B-4391-AA8A-FE2038FB224D%7D. However, note that the N2K areas including multiple physically separate patches were treated as separate polygons in Heikkinen et al. (2020), and a minimum size requirement of 2 hectares were requested. Moreover, the digital elevation model (DEM) data for Finland (which were dissected to Natura 2000 polygons to examine their elevational variation and its relationships to topoclimatic variation) can be downloaded from the following link:&nbsp;https://ckan.ymparisto.fi/en/dataset/dem25_astergdem25.&nbsp;</p><p>The coordinate system for the climate velocity data files is: ETRS-TM35FIN (EPSG: 3067) (or YKJ Finland/Finnish Uniform Coordinate System (EPSG: 2393)). Summary of the key settings and elements of the study are provided below. A detailed treatment is provided in Heikkinen et al. (2020).</p><p>Code to the files (four files per each velocity layer: *.tif, *.tfw. *.ovr and *.tif.aux.xml) in the dataset:&nbsp;</p><p>(a) Velocity of GDD with respect to RCP2.6 future climate (Fig 2a in Heikkinen et al. 2020). Name of the file: GDDRCP26.*</p><p>(b) Velocity of GDD with respect to RCP4.5 future climate (Fig. 2b in Heikkinen et al. 2020). Name of the file: GDDRCP45.*</p><p>(c) Velocity of GDD with respect to RCP8.5 future climate (Fig. 2c in Heikkinen et al. 2020). Name of the file: GDDRCP85.*</p><p>(d) Velocity of mean January temperature with respect to RCP2.6 future climate (Fig. 2d in Heikkinen et al. 2020). Name of the file: TJanRCP26.*</p><p>(e) Velocity of mean January temperature with respect to RCP4.5 future climate (Fig. 2e in Heikkinen et al. 2020). Name of the file: TJanRCP45.*</p><p>(f) Velocity of mean January temperature with respect to RCP8.5 future climate (Fig. 2f in Heikkinen et al. 2020). Name of the file: TJanRCP85.*</p><p>(g) Velocity of climatic water balance with respect to RCP2.6 future climate (Fig. 2g in Heikkinen et al. 2020). Name of the file: WABRCP26.*</p><p>(h) Velocity of climatic water balance with respect to RCP4.5 future climate (Fig. 2h in Heikkinen et al. 2020). Name of the file: WABRCP45.*</p><p>(i) Velocity of climatic water balance with respect to RCP8.5 future climate (Fig. 2i in Heikkinen et al. 2020). Name of the file: WABRCP85.*</p><p>Note that velocity surfaces e and f include disappearing climate conditions.</p><p><strong>Summary of the study:</strong></p><p>Climate velocity is a generic metric which provides useful information for climate-wise conservation planning to identify regions and protected areas where climate conditions are changing most rapidly, exposing them to high rates of climate displacement (Batllori et al. 2017), causing potential carry-over impacts to community structure and ecosystem functions (Ackerly et al. 2010). Climate velocity has been typically used to assess the climatic risks for species and their populations, but velocity metrics can also be used to identify protected areas which face overall difficulties in retaining ecological conditions that promote present-day biodiversity.&nbsp;</p><p>Earlier climate velocity assessments have focussed on the domains of the mesoclimate (resolutions of 1–100 km) or macroclimate (&gt;100 km scales), and fine-grained (&lt;100 m) local climatic conditions created by variation in topography ('topoclimate'; Ackerly et al. 2010; 2020) have largely been overlooked (Heikkinen et al. 2020). This omission may lead to biased exposure assessments especially in rugged terrain (Dobrowski et al. 2013; Franklin et al. 2013), as well as a limited ability to detect sites decoupled from the regional climate (Aalto et al. 2017; Lenoir et al. 2017). This study provided the first assessment of the climatic exposure risks across a national PA (Natura 2000) network based on very fine-grained velocities of three established drivers of high latitude biodiversity.&nbsp;</p><p>The produce fine-grain climate velocity measures, 50-m resolution monthly temperature and precipitation data averaged for 1981–2010 were first developed, and based on it, the three bioclimatic variables (growing degree days, mean January temperature and annual climatic water balance) were calculated for the whole study domain. In the next phase, similar future climate surfaces were produced based on data from an ensemble of 23 global climate models, extracted from the CMIP5 archives for the years 2070–2099 and the three RCP scenarios (RCP2.6, RCP4.5 and RCP8.5)26. In the final step, climate velocities for each the 50 x 50 m grid cells were measured using climate-analog velocity method (Hamann et al. 2015) and based on the distance between climatically similar cells under the baseline and future climates.</p><p>The results revealed notable spatial differences in the high velocity areas for the three bioclimatic variables, indicating contrasting exposure risks in protected areas situated in different areas. Moreover, comparisons of the 50-m baseline and future climate surfaces revealed a potential wholesale disappearance of current topoclimatic temperature conditions from almost all the studied PAs by the end of this century.</p><p><strong>Calculation of climate change velocity metrics for the three climate variables</strong></p><p>The overall process of calculation of climate velocities included three main steps.&nbsp;</p><p>(1) In the first step, we developed high-resolution monthly average temperature and precipitation data averaged over the years 1981–2010 and across the study domain at a spatial resolution of 50 × 50 m. This was done by building topoclimatic models based on climate data sourced from 313 meteorological stations (European Climate Assessment and Dataset [ECA&amp;D]) (Klok et al. 2009). Our station network and modelling domain covered the whole of Finland with an additional 100 km buffer. However, it was also extended to cover large parts of northern Sweden and Norway for areas &gt;66.5°N, as well as selected adjacent areas in Russia (for details see Heikkinen et al. 2020). This was done to capture the present-day climate spaces in Finland which are projected to move in the future beyond the country borders but have analogous climate areas in neighbouring areas; this was done to avoid developing a large number of velocity values deemed as infinite or unknown in the data for Finland.&nbsp;</p><p>The 50-m resolution average air temperature data were developed for the study domain using generalized additive modelling (GAM), as implemented in the R-package mgcv version 1.8–7 (R Development Core Team 2011; Wood 2011). In this modelling we utilised variables of geographical location (latitude and longitude, included as an anisotropic interaction), topography (elevation, potential incoming solar radiation, relative elevation) and water cover (sea and lake proximity), and subsequent leave-one-out cross-validation tests to assess model performance (for full process description, see Aalto et al. 2017; Heikkinen et al. 2020). The resulting topoclimate data effectively captured the physiographic effects of solar radiation and cold-air pooling.</p><p>To produce gridded precipitation data, we applied global kriging interpolation to the data from 343 rain gauges from the ECA&amp;D dataset. The interpolation was carried out using information on geographical location, topography (elevation and eastness index) and proximity to the sea and R package gstat. The eastness index was obtained from a sine-transforming aspect raster surface calculated from a 50 m × 50 m digital elevation model to capture the effect of prevailing westerly winds on the accumulated precipitation on windward slopes. The gridding was first run at a resolution of 500 × 500 m, whereafter gridded precipitation values were bilinearly interpolated into the same 50 × 50 m resolution as the air temperature data.&nbsp;</p><p>Next, the three bioclimatic variables ((i) growing degree days (GDD, °C days) indicating the accumulated warmth during the growing season; (ii) mean January air temperature - &nbsp;TJan, °C; (iii) climatic water balance - WAB, mm) were calculated for each 50 x 50 grid cell from the high-resolution gridded 1981–2010 ('baseline') climate data. Earlier research has demonstrated the ecological relevance of these three complementary variables which provide estimations of winter cold, seasonal warmth and moisture availability (Sykes et al. 1996; Luoto et al. 2006; Huntley et al. 2007, 2008).&nbsp;</p><p>Following Carter et al. (1991), GDD was calculated as the effective temperature sum above the base temperature of 5 °C as follows:</p><p><i>GDD</i>5 = <i>∑ni&nbsp;(Ti - Tb),&nbsp; if Ti -Tb &gt; 5</i></p><p>where Ti denotes the mean temperature at day i, Tb represents the base temperature, and n is the length of the summation period. However, because the daily air temperature data was not available, here the GDD was estimated using monthly data as in Araújo &amp; Luoto (2007). The WAB is the difference between the total annual precipitation sum and the potential evapotranspiration (PET), which was estimated from the monthly air temperatures following Skov and Svenning (2004):&nbsp;</p><p><i>PET&nbsp;</i>= 58.93 × <i>Tabove&nbsp;</i>0°<i>C </i>/ 12</p><p>(2) In the second step we developed data on future climates by using the climate projections from the ensemble of 23 global climate models (GCMs), derived from the Coupled Model Intercomparison Project phase 5 archives (Taylor et al. 2012). From these archives, we processed to predicted averaged changes in mean temperature and precipitation with respect to the baseline 1981–2010 for the years 2070–2099, and the three RCP scenarios (cf. Moss et al. 2010). As the Coupled Model Intercomparison Project phase 5 climate scenario data represent coarse-scale resolution data, we converted it to match our fine-resolution baseline climate data by interpolation. For this, the climate model data depicting the predicted change in mean temperatures and precipitation with respect to the baseline climate were bilinearly interpolated to the 50 × 50 m grid system, and the change predicted by the GCMs was added to the spatially detailed baseline climate data. After this, the bioclimatic variables were recalculated for each RCP scenario to allow the calculation of climate change velocities across the whole country and the Natura 2000 protected areas.</p><p>(3) In the third step we developed climate change velocities for the three bioclimatic&nbsp;variables using the climate-analog approach (Hamann et al. 2015) where velocity is calculated by measuring the&nbsp;distance between present-day locations with certain climatic conditions and their future climate analogues,&nbsp;divided by the number of years between the two points in time. Thus, we calculated climate-analog velocities for the 50-m resolution grid climate data by measuring the distance between climatically&nbsp;similar grid cells for the present and future climates under RCP2.6,&nbsp;RCP4.5 and RCP8.5.&nbsp;</p><p>Prior the actual climate-analog velocity measurements, the climate variable surfaces were converted from continuous values into classified variable surfaces. For this, we defined the boundary values for the variable classes so that the climatically matching grid cells had their within-class ranges as small as possible but, at the same time, avoided artefactual extreme precision. After a set of pilot reclassifications, the following within-class ranges were applied: GDD, within-class range 50 °C with 51 categories; TJan, within-class range 0.5 °C with 60 categories; WAB, within-class range 50 mm with 55 categories. Next, using the reclassified present-day and future climate surfaces the search of the minimum distances between grid cells with similar present-day and future GDD/TJan/WAB climates were executed. The search was carried out using the ArcGIS software (Desktop 10.5.1.) by employing the Euclidean distance function. The minimum distances measured for each 50-m grid cell were divided by the difference between the mean points in the two time slices,&nbsp;1981–2010&nbsp;and 2070–2099.&nbsp;</p><p>The resulting 50-m resolution climate velocity surfaces for the three climate variables are provided in the zipped files included this data&nbsp;repository. In Heikkinen et al. (2020), these climate velocity data&nbsp;were employed in a series of subsequent analyses. For example, high-velocity areas ('velocity hotspots') of the three climate variables were visually compared with each other based on maps showing their 50-m resolution velocities across mainland Finland and the degree of overlap between the present-day range and projected future range of the three climate variables were investigated in each of the 5,068 Natura 2000 polygons included in the study.</p><p><strong>References</strong></p><p>Aalto, J., Riihimäki, H., Meineri, E., Hylander, K., Luoto, M., 2017.&nbsp;Revealing topoclimatic heterogeneity using meteorological station data. International Journal of Climatology 37, 544-556.</p><p>Ackerly, D.D., Loarie, S.R., Cornwell, W.K., Weiss, S.B., Hamilton, H., Branciforte, R., Kraft, N.J.B., 2010. The geography of climate change: implications for conservation biogeography. Diversity and Distributions 16, 476-487.</p><p>Ackerly, D.D., Kling, M.M., Clark, M.L., Papper, P., Oldfather, M.F., Flint, A.L., Flint, L.E., 2020. Topoclimates, refugia, and biotic responses to climate change. Frontiers in Ecology and the Environment 18, 288-297.</p><p>Araujo, M.B., Luoto, M., 2007.&nbsp;The importance of biotic interactions for modelling species distributions under climate change. Global Ecology and Biogeography 16.</p><p>Batllori, E., Parisien, M.-A., Parks, S.A., Moritz, M.A., Miller, C., 2017. Potential relocation of climatic environments suggests high rates of climate displacement within the North American protection network. Global Change Biology 23, 3219-3230.</p><p>Brito-Morales, I., García Molinos, J., Schoeman, D.S., Burrows, M.T., Poloczanska, E.S., Brown, C.J., Ferrier, S., Harwood, T.D., Klein, C.J., McDonald-Madden, E., Moore, P.J., Pandolfi, J.M., Watson, J.E.M., Wenger, A.S., Richardson, A.J., 2018. Climate Velocity Can Inform Conservation in a Warming World. Trends in Ecology &amp; Evolution 33, 441-457.</p><p>Carter, T.R., Porter, J.H., Parry, M.L., 1991. Climatic warming and crop potential in Europe: Prospects and uncertainties. Global Environmental Change 1, 291-312.</p><p>Dobrowski, S.Z., Abatzoglou, J., Swanson, A.K., Greenberg, J.A., Mynsberge, A.R., Holden, Z.A., Schwartz, M.K., 2013. The climate velocity of the contiguous United States during the 20th century. Global Change Biology 19, 241-251.</p><p>Franklin, J., Davis, F.W., Ikegami, M., Syphard, A.D., Flint, L.E., Flint, A.L., Hannah, L., 2013. Modeling plant species distributions under future climates: how fine scale do climate projections need to be? Global Change Biology 19, 473-483.</p><p>Hamann, A., Roberts, D.R., Barber, Q.E., Carroll, C., Nielsen, S.E., 2015. Velocity of climate change algorithms for guiding conservation and management.&nbsp;Global Change Biology 21, 997-1004.&nbsp;</p><p>Heikkinen, R.K., Leikola, N., Aalto, J., Aapala, K., Kuusela, S., Luoto, M., Virkkala, R., 2020.&nbsp;Fine-grained climate velocities reveal vulnerability of protected areas to climate change. Scientific Reports 10:1678.</p><p>Huntley, B., Green, R.E., Collingham, Y.C., Willis, S.G., 2007. A climatic atlas of European breeding birds. Durham University, The RSPB and Lynx Edicions, Barcelona.</p><p>Huntley, B., Collingham, Y.C., Willis, S.G., Green, R.E., 2008. Potential Impacts of Climatic Change on European Breeding Birds.&nbsp;Plos One 3.</p><p>Klok, E.J., Klein Tank, A.M.G., 2009.&nbsp;Updated and extended European dataset of daily climate observations. International Journal of Climatology 29, 1182-1191.</p><p>Lenoir, J., Hattab, T., Pierre, G., 2017.&nbsp;Climatic microrefugia under anthropogenic climate change: implications for species redistribution.&nbsp;Ecography 40, 253-266.</p><p>Luoto, M., Heikkinen, R.K., Pöyry, J., Saarinen, K., 2006.&nbsp;Determinants of biogeographical distribution of butterflies in boreal regions. Journal of Biogeography 33, 1764-1778.</p><p>Moss, R.H., Edmonds, J.A., Hibbard, K.A., Manning, M.R., Rose, S.K., van Vuuren, D.P., Carter, T.R., Emori, S., Kainuma, M., Kram, T., Meehl, G.A., Mitchell, J.F.B., Nakicenovic, N., Riahi, K., Smith, S.J., Stouffer, R.J., Thomson, A.M., Weyant, J.P., Wilbanks, T.J., 2010. The next generation of scenarios for climate change research and assessment. Nature 463, 747-756.</p><p>R Development Core Team, 2011. R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing).</p><p>Skov, F., Svenning, J.-C., 2004.&nbsp;Potential impact of climatic change on the distribution of forest herbs in Europe. Ecography 27, 366-380.</p><p>Sykes, M.T., Prentice, I.C., Cramer, W., 1996. A bioclimatic model for the potential distributions of north European tree species under present and future climates. Journal of Biogeography 23, 203-233.</p><p>Taylor, K.E., Stouffer, R.J., Meehl, G.A., 2012. An Overview of CMIP5 and the Experiment Design. Bulletin of the American meteorological Society 93, 485-498.</p><p>Wood, S.N., 2011. Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society Series B 73, 3-36.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
edi52/100

Cardiorespiratory metrics of cobia during and following a Ucrit trial upon a 3 wk exposure to either ambient or elevated pCO2

This study examined cardiac and swimming performance during a maximum sustained swimming trial (Ucrit) of cobia (Rachycentron canadum) following a three-week exposure to either elevated (~1,600 µAtm) or ambient (~500 µAtm) pCO₂. This dataset encompasses cardiac variables (stroke volume, heart rate, and cardiac output) during swimming and recovery from a Ucrit trial, hematological variables (hemoglobin, hematocrit, mean corpuscular hemoglobin concentration) following the swimming test and hour long recovery, oxygen consumption variables (MO₂ min, MO₂ max, aerobic scope, and factorial aerobic scope) during and following the swimming test.

openCC (other)Mar 2025View details →
edi52/100

Advective nitrate fluxes, sea surface chlorophyll concentrations and other physical metrics in the Santa Barbara Channel (2012-2019)

This data package includes 6 files: (1 & 2) In-situ nitrate concentrations at the surface and mixed layer depth, and collocated remotely-sensed and reanalysis quantities of satellite sea surface temperature, 15-day cumulative wind stress, satellite sea surface chlorophyll with a 5-day lag, index of offshore position of the California Current, indices for along-channel and across-channel distance, and index for day of the year. (3) An R script for generating generalized additive models (GAMs) to predict nitrate concentrations at the surface and at the mixed layer depth using the collocated data in files 1 & 2. (4) Daily maps of satellite sea surface chlorophyll concentrations (SSChl), High-frequency radar (HFR) surface currents, weather research and forecasting (WRF) model wind-derived vertical velocities, estimated nitrate concentrations at the surface and mixed layer depth, horizontal advective nitrate fluxes at the surface and vertical advective nitrate fluxes. (5) Daily time series of spatial mean SSChl, principal component amplitude of the first mode of variability in surface currents estimated using complex empirical orthogonal function (EOF) analysis, alongshore pressure gradient, wind stress, spatial mean horizontal velocities at the western and eastern Santa Barbara Channel boundaries, spatial mean vertical velocities, spatial mean surface nitrate concentrations at the channel boundaries and across the entire channel, spatial mean mixed layer depth nitrate concentrations across the entire channel, spatial mean horizontal advective nitrate fluxes at the channel boundaries, and spatial mean vertical advective nitrate fluxes. (6) A MATLAB script for plotting examples of the daily maps and time series in files 4 & 5. These data were processed in order to investigate the impact of local nutrient delivery mechanisms on phytoplankton blooms in the Santa Barbara Channel, California, details of which are available in the study: Brokaw, R.J., D.A. Siegel, L. Washburn,

openCC (other)Jun 2025View details →
zenodo48/100

Climate Forcing due to Future Ozone Changes: An intercomparison of metrics and methods

<p>The data provided in this repository relates to a paper on ozone radiative forcing submitted for publication in Atmos. Chem. Phys., as part of the TOAR-II special issue (<a href="https://acp.copernicus.org/articles/special_issue1256.html">ACP &ndash; Special issue &ndash; Tropospheric Ozone Assessment Report Phase II (TOAR-II) Community Special Issue (ACP/AMT/BG/GMD inter-journal SI)</a>). The paper is entitled "<span>Climate Forcing due to Future Ozone Changes</span><span>: An intercomparison of metrics and methods" by authors <span><span>William J. Collins</span></span><span><span>,</span> <span>Fiona M. O&rsquo;Connor</span></span><span><span>, </span><span>Connor R. Barker</span></span><span><span>, </span><span>Rachael E. Byrom</span></span><span><span>, </span><span>Sebastian D. Eastham</span></span><span><span>,</span> <span>&Oslash;ivind Hodnebrog</span></span><span><span>, Patrick J&ouml;ckel</span></span><span><span>, </span><span>Eloise A. Marais</span></span><span><span>, </span><span>Mariano Mertens</span></span><span><span>, Gunnar Myhre</span></span><span><span>, Matthias N&uuml;tzel</span></span><span><span>, Dirk Olivi&eacute;</span></span><span><span>, Ragnhild </span><span>Bieltvedt</span><span> Skeie</span></span><span><span>5</span></span><span><span>, Laura Stecher</span></span><span><span>, Larry W. Horowitz</span></span><span><span>, Vaishali Naik</span></span><span><span>, Gregory Faluvegi</span></span><span><span>, Ulas Im</span></span><span><span>, Lee T. Murray</span></span><span><span>, Drew Shindell</span></span><span><span>, Kostas Tsigaridis</span></span><span><span>, Nathan Luke Abraham</span></span><span><span>, James Keeble.</span></span></span></p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

augMENTOR: Simulated Student Learning Profiles and their Engagement Metrics in TryHackMe Platform_V1

<p>The dataset provides simulated insights into student engagement and performance within the THM platform. It outlines mathematical representations of student learning profiles, detailing behaviors ranging from high achievers to inconsistent performers. Additionally, the dataset includes key performance indicators, offering metrics like room completion, points earned, and time spent to gauge student progress and interaction within the platform's modules.</p><p>Here are definitions of the learning profiles, along with mathematical representations of their behaviors:</p><ul><li>High Achiever: These are students who consistently perform well across all modules. Their performance can be described as a normal distribution centered at a high mean value. Their performance P in a given module can be modelled as: P = N(90, 5) where N is the normal distribution function, 90 is the mean, and 5 is the standard deviation.</li><li>Average Performer: These are students who typically perform at the average level across all modules. Their performance can be described as a normal distribution centered at a medium mean value: P = N(70, 10), where 70 is the mean, and 10 is the standard deviation.</li><li>Late Bloomer: These are students whose performance improves as they progress through the modules. Their performance can be modelled as: P = N(50 + i*10, 10), where i is the module index and shows an increasing trend.</li><li>Specialized Talent: These are students who have average performance in most modules but excel in a particular module (e.g., module5). Their performance can be described as: P = N(90, 5) if the module is module 5, else P = N(70, 10).</li><li>Inconsistent Performer: These are students whose performance varies significantly across modules. Their performance can be described as a normal distribution with a high standard deviation: P = N(70, 30), where 70 is the mean, and 30 is the high standard deviation, reflecting inconsistency.</li></ul><p>Note that the actual performances are bounded between 0 and 100 using the function max(0, min(100, performance)) to ensure valid percentages.</p><p>In these formulas, the <i>np.random.normal</i> function is used to simulate the variability in student performance around the mean values. The first argument to this function is the mean, and the second argument is the standard deviation, reflecting the level of variability around the mean. The function returns a number drawn from the normal distribution described by these parameters. Note that the proposed method is experimental and has not been validated.&nbsp;</p><p>&nbsp;</p><p>List of Key Performance Indicators (KPIs) for Student Engagement and Progress within the Platform:</p><ul><li>Room Name: This represents the unique identifier or name of a specific room (or module). Think of each room as a separate module or lesson within an educational platform. For example, Room1, Room2, etc.</li><li>Total rooms completed: Indicates the cumulative number of rooms that a student has fully completed. Completion is typically determined by meeting certain criteria, like answering all questions or achieving a certain score.</li><li>Rooms registered in: Represents the number of rooms a student has registered or enrolled in. This could be different from the total number of rooms they've completed.</li><li>Ratio of Questions completed per room: This gives an insight into a student's progress in a particular room. For instance, a ratio of 7/10 suggests the student has completed 7 out of 10 available questions in that room.</li><li>Room Completed (yes no): Indicates whether a student has fully completed a specific room or not. This could be determined by the percentage of material covered, questions answered, or a certain score achieved.</li><li>Room Last deploy (count of days): Refers to the number of days since the last update or deployment was made to that room. It can give an idea about the effort of the student.</li><li>Points in room used for the leaderboard (range 0-560): Each room assigns points based on student performance, and these points contribute to leaderboards. The range suggests that a student can earn anywhere from 0 to 560 points in a particular room.</li><li>Last answered question in a room (27th Jan 2023): This indicates the date when a student last answered a question in a specific room. It can provide insights into a student's recent activity and engagement.</li><li>Total points in all rooms (range 0-560): The cumulative score a student has achieved across all rooms.</li><li>Path Percentage completed (range 0-100): Indicates the percentage of the overall learning path that the student has completed. A path could consist of multiple modules or rooms.</li><li>Module Percentage completed (range 0-100): Represents how much of a specific module (which could have multiple lessons or topics) a student has completed.</li><li>Room Percentage completed (range 0-100): Shows the percentage of a specific room that has been completed by a student.</li><li>Time Spent on the platform (seconds): This provides an aggregate of the total time a student has spent on the entire educational platform.</li><li>Time spent on each room (seconds): Represents the amount of time a student has dedicated to a specific room. This can give insights into which rooms or modules are the most time-consuming or engaging for students.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics

<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv.&nbsp;</p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint&rsquo;s title, author, abstract, subject category and the arXiv ID (the arXiv&rsquo;s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received&nbsp;<em>c</em><sub>pre</sub>&nbsp;and&nbsp;<em>c</em><sub>pub</sub>&nbsp;citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as&nbsp;<em>c</em><sub>pre</sub>&nbsp;+&nbsp;<em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (&lsquo;astro-ph&rsquo;), Computer Science (&lsquo;comp-sci&rsquo;), Condensed Matter Physics (&lsquo;cond-mat&rsquo;), High Energy Physics (&lsquo;hep&rsquo;), Mathematics (&lsquo;math&rsquo;) and Other Physics (&lsquo;oth-phys&rsquo;), which collectively accounted for 98% of all the eprints. Those eprints&nbsp;tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints.&nbsp;</p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p>&nbsp;</p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> :&nbsp;arXiv ID</li> <li><strong>category</strong> :&nbsp;Research discipline</li> <li><strong>pre_year</strong> :&nbsp;Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> :&nbsp;Year of DOI acquisition</li> <li><strong>c_tot</strong> :&nbsp;No. of citations acquired during 1991&ndash;2019</li> <li><strong>c_pre</strong> :&nbsp;No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> :&nbsp;No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong>&nbsp;(<em>yyyy</em>&nbsp;= 1991, &hellip;, 2019) :&nbsp;No. of citations acquired in the year&nbsp;<em>yyyy</em>&nbsp;(with &lsquo;<em>yyyy</em>&rsquo; running from 1991 to 2019)</li> <li><strong>gamma</strong> :&nbsp;The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> :&nbsp;The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (&gamma;; &lsquo;gamma&rsquo;) and that of the standardised citation index (&gamma;*; &lsquo;gamma_star&rsquo;) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times.&nbsp;</p> <p>&nbsp;</p> <p><strong>Data files</strong></p> <p>A comma-separated values file (&lsquo;<strong>arXiv_impact.csv</strong>&rsquo;) and a Stata file (&lsquo;<strong>arXiv_impact.dta</strong>&rsquo;) are provided, both containing the same information.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Passive gas plume database for metrics comparison

<p>Plume database used for the evaluation of different metrics that are presented in the submitted paper &quot;New plume comparison metrics for the inversion of passive gases emissions&quot;. The synthetical CO2 plumes presented in the NetCDF file entitled &quot;Synthetical_CO2_plume_database.nc&quot; are the results of chemical transport model simulations described in the preprint available here&nbsp;<a href="https://amt.copernicus.org/preprints/amt-2022-48/">https://amt.copernicus.org/preprints/amt-2022-48/</a>.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Metagenome quality metrics and taxonomical annotation visualization through the integration of MAGFlow and BIgMAG (Sup. Material)

<p>Dataset encompassing:</p> <ul> <li>The recovered MAGs by 6 different metagenomics pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) using a mock community as input (SRR8359173 and SRR9328980), complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.&nbsp;</li> <li>The MAGs produced by nf-core/mag using rice/rhizosphere sequenced libraries (PRJNA663614, PRJNA448773 and PRJNA645385) in either single assembly/single binning or co-assembly/co-binning mode, complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.</li> <li>Scripts, commands and configuration files to run the different pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) and reproduce the experimental conditions.</li> <li>Outputs, commands and scripts to run Metabinner and Semibin in their default configuration using the rice soil samples co-assembly, along with the MAGFlow (v1.1.0) output to compare these binners against MetaBAT2.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

Multiple-benefit Conservation in Practice: Metrics Data for Quantifying Multidimensional Impacts of Landscape Change in California's Sacramento–San Joaquin Delta

<p><strong>SUMMARY</strong><br> These data represent estimated mean value, standard error, and units for a range of metrics by land cover class in the Sacramento-San Joaquin Delta. Metrics are grouped into three major categories: Agricultural Livelihoods (including metrics for gross production value, number of agricultural jobs, and annual wages per employee), Water Quality (in terms of the application rates for pesticides identified as critical pesticides, groundwater contaminants, and those posing a high or moderate risk to aquatic organisms), and Climate Change Resilience (qualitative scores representing relative tolerance for heat, drought, and flood).</p> <p><strong>DESCRIPTION</strong><br> These data were developed to facilitate projecting the net impacts of land cover change scenarios on multiple metrics of interest to the Sacramento-San Joaquin Delta, including potential benefits and trade-offs. They were used in initial analyses of scenarios representing habitat restoration and perennial crop expansion, and they are required for using the R package &quot;DeltaMultipleBenefits&quot;, which provides the code and work flow for repeating the initial analyses or analyzing new scenarios.</p> <p>For additional details about the development and applications of these data, please see: &nbsp;</p> <ul> <li>Dybala KE, et al. (In review) Multiple-benefit Conservation in Practice: A Framework for Quantifying Multi-dimensional Impacts of Landscape Change in California&rsquo;s Sacramento&ndash;San Joaquin Delta &nbsp;</li> <li>Dybala KE (2023) <em>DeltaMultipleBenefits: Projecting the Multiple Benefits of Land Cover Change in the Sacramento-San Joaquin River Delta.</em>&nbsp;R package version 1.0.0. doi: 10.5281/zenodo.7718620. Available from: https://pointblue.github.io/DeltaMultipleBenefits &nbsp;</li> </ul> <p><strong>FUNDING STATEMENT</strong><br> These data were developed as part of the project &quot;Trade-offs and Co-benefits of Landscape Change on Bird Communities and Ecosystem Services in the Sacramento&ndash;San Joaquin River Delta&quot;, funded by Proposition 1 Delta Water Quality and Ecosystem Restoration Program, Grant Agreement Number &ndash; Q1996022, administered by the California Department of Fish and Wildlife.</p> <p><strong>POINT OF CONTACT</strong><br> Kristen Dybala, Point Blue Conservation Science, kdybala@pointblue.org</p> <p><strong>SUGGESTED CITATION</strong><br> Dybala KE. 2023. Multiple-benefit Conservation in Practice: Metrics Data for Quantifying Multi-dimensional Impacts of Landscape Change in California&rsquo;s Sacramento&ndash;San Joaquin Delta. doi:10.5281/zenodo.7504874.</p> <p><strong>DATA DISTRIBUTION</strong><br> Zenodo&nbsp;(https://doi.org/10.5281/zenodo.7504874)</p> <p><strong>PROGRESS</strong><br> Complete, but note that the accompanying manuscript has not yet undergone peer-review, and thus these data may require future revision.</p> <p><strong>UPDATE FREQUENCY</strong><br> As Needed</p> <p><strong>DATE</strong><br> These data were compiled in 2022, based on data from the Quarterly Census of Employment and Wages 2014-2020 (EDD 2022), annual County Agricultural Commissioners Reports 2014-2020 (CDFA 2022),&nbsp;Pesticide Use Report Data 2014-2018 (CDPR 2022), and qualitative assessments of climate change resilience (Peterson et al. 2020, DSC 2021).</p> <p><strong>Literature Cited:</strong></p> <ul> <li>CDFA. 2022. County Ag Commissioners&rsquo; Data Listing. California Department of Food &amp; Agriculture. Available from: https://www.nass.usda.gov/Statistics_by_State/California/Publications/AgComm/index.php</li> <li>CDPR. 2022. Pesticide Use Report Data. California Department of Pesticide Regulation. Available from: https://www.cdpr.ca.gov/docs/pur/purmain.htm</li> <li>DSC. 2021. Delta Adapts: Creating a Climate Resilient Future. Public Review Draft. Delta Stewardship Council. Available from https://deltacouncil.ca.gov/delta-plan/climate-change</li> <li>EDD. 2022. Quarterly Census of Employment and Wages (QCEW). California Employment Development Department. Available from: https://data.edd.ca.gov/Industry-Information-/Quarterly-Census-of-Employment-and-Wages-QCEW-/fisq-v939</li> <li>Peterson C, Marvinney E, Dybala K. 2020. Multiple Benefits from Agricultural and Natural Land Covers in the Central Valley, CA. Migratory Bird Conservation Partnership, Sacramento, CA. Dryad Dataset doi:10.25338/B8061X</li> </ul> <p><strong>FIELD DEFINITIONS</strong></p> <ul> <li><strong>METRIC_CATEGORY:&nbsp;</strong>Broad grouping assigned to each METRIC; one of Agricultural Livelihoods, Water Quality, or Climate Change Resilience</li> <li><strong>METRIC:&nbsp;</strong>Specific metric being estimated; one of Agricultural Jobs, Annual Wages, Gross Production Value, Drought, Flood, Heat, Critical Pesticides, Groundwater Contaminant, or Risk to Aquatic Organisms</li> <li><strong>UNIT:&nbsp;</strong>The units in which the <strong>METRIC </strong>is estimated</li> <li><strong>CODE_NAME:</strong>&nbsp;The land cover class or subclass for which the <strong>METRIC </strong>is estimated</li> <li><strong>LABEL:&nbsp;</strong>A more user-friendly version of <strong>CODE_NAME</strong>, useful for creating figures and tables</li> <li><strong>SCORE_MEAN:</strong>&nbsp;The mean value of each METRIC estimated for each land cover class or subclass</li> <li><strong>SCORE_SE:&nbsp;</strong>The standard error of the mean</li> </ul> <p><strong>ABBREVIATION DEFINITIONS</strong></p> <ul> <li><strong>FTE:&nbsp;</strong>full-time equivalents; refers to converting monthly agricultural jobs data to annual estimates by dividing by 12</li> <li><strong>ha:</strong>&nbsp;hectares</li> <li><strong>kg:&nbsp;</strong>kilograms</li> <li><strong>USD:&nbsp;</strong>U.S. dollars</li> <li><strong>yr:&nbsp;</strong>year</li> </ul> <p><strong>ACCESS &amp; USE CONSTRAINTS</strong><br> CC-by-4.0 (https://creativecommons.org/licenses/by/4.0/)</p> <p><strong>KEYWORDS</strong></p> <ul> <li><strong>Themes:</strong>&nbsp;agriculture, livelihoods, economy, water quality, pesticides, climate change, resilience, multiple-benefit conservation</li> <li><strong>Place:</strong>&nbsp;Sacramento-San Joaquin River Delta, Central Valley, California<br> &nbsp;</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Experimental Factors Influence Diversity Metrics of the Gut Microbiome in Laboratory Mice

<p>Abstract<br> Introduction</p> <p>Gut microbiome studies often overlook experimental factors that could influence gut microbiome diversity and could impact findings. Large-scale studies investigating these experimental factors are lacking. Thus, we aimed to determine which experimental factors influence the gut microbiome diversity in pre-clinical animal model studies.</p> <p><br> Methods</p> <p>We extracted DNA and sequenced the V4 region of the 16S rRNA gene of a total of 538 samples from various sections of the gastrointestinal tract of 303 young and aged male and female C57BL/6J mice of three different genotypes on five diets from three animal house facilities. As a proof-of-concept in a disease model, some mice were treated with sham or angiotensin II, a commonly studied agent used as a hypertension model. Some samples were sequenced twice as a matched-comparison group.</p> <p>Results</p> <p>Using over 17 million sequencing reads, we found that experimental factors such as animal house facility, genotype, diet, age, sex, sampling site, and technical factor (i.e., sequencing batch) affected both &alpha;- and &beta;-diversity (weighted and unweighted UniFrac), and were associated with compositional changes in the microbiome at varying magnitude, with diet and sampling site having the largest effect. After adjustment by these factors, treatment with angiotensin II had no impact on &alpha;-diversity and was only significant in unweighted UniFrac (presence/absence of bacteria) analyses.</p> <p><br> Conclusion</p> <p>Our data identified several key experimental and technical factors that affect the gut microbiome in laboratory mice. Our findings support that not accounting or adjusting for these factors may lead to false-positive discoveries and non-biologically relevant findings in the gut microbiome field.</p>

opencc-by-4.0May 2023View details →
edi48/100

Integrated freshwater abundance and connectivity clusters at the Hydrologic Unit 8 scale for the Midwest and Northeast U.S.A. – freshwater metric variables and k-means cluster assignment

This dataset includes integrated freshwater abundance and connectivity cluster output, principal component scores, and lake, wetland, and stream abundance and connectivity metrics measured at the Hydrologic Unit 8 (HU8) scale for 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of the integrated freshwater landscape that includes lakes, wetlands, and streams and their surface connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). The integrated freshwater clusters were created through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics for lakes, streams, and wetlands separately, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of freshwater abundance and connectivity in the landscape.

openCC0Jul 2017View details →
edi48/100

Freshwater connectivity clusters for lakes, wetlands, and streams at the Hydrologic Unit 12 scale in the Midwest and Northeast U.S.A. – freshwater metric variables and K-means cluster assignment

This dataset includes freshwater connectivity cluster output and principal component scores for lakes, wetlands, and streams measured at the Hydrologic Unit 12 (HU12) scale in 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of freshwater connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). Freshwater connectivity clusters were created separately for lakes, wetlands, and streams through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of lake, wetland, and stream connectivity in the landscape.

openCC0Jul 2017View details →
edi48/100

Leaf leachate chemistry and uptake metrics in headwater streams July 2015

This dataset contains the results of uptake experiments conducted at Caribou-Poker Creeks Research Watershed in July 2015. The objective of the study was to determine the relative influence of molecular composition and nutrient content of organic matter on the retention of dissolved organic matter (DOM) in boreal streams. We measured in situ rates of carbon uptake in streams following introduction of leachates derived from of alder (Alnus incana ssp. tenuifolia), poplar (Populus balsamifera), and white spruce (Picea glauca) trees subject to long-term fertilization with nitrogen (N) or phosphorus (P). We measured leachate composition through chemical analysis and optical properties (absorbance and fluorescence).

openOpenApr 2017View details →
edi48/100

Fire Self-Limitation (FiSL) Experiment: Quantifying Wildfire Carbon Combustion Losses in boreal Deciduous and Mixed Forests in Interior Alaska and the Boreal Cordillera IX: metrics derived from All Raw Data Collected Plus Data from Previous Studies on the 2004 Alaska Wildfires Included in Analysis 2022

This data set includes metrics derived from field and lab data collected for deciduous and mixed deciduous-confier plots collected in the summer of 2022 (Shovel Creek (2019), Aggie Creek (2015), Hess Creek (2019), Baker (2015), Munson Creek (2021), Isom Creek (2020), 2019MA014 (2019), and 2019BC005 (2019)), as well as additional data for conifer plots from previous studies of the Taylor Highway Complex (2004), Dall Creek/Yukon Crossing (2004), and Boundary (2004) fires. Those additional data were acquired from: https://www.lter.uaf.edu/d1/d1-detail/id/773 and https://daac.ornl.gov/ABOVE/guides/ABoVE_Plot_Data_Burned_Sites.html. From this complete data set of 333 plots, 311 plots were used in analyses in Black at al. (NCC) paper: "Increased deciduous tree dominance reduces wildfire carbon losses in boreal forests". Plots excluded (from 2022 FiSL data) were poplar-dominated, mixed poplar/conifer dominated, missing soil C data, or conifer-dominated (adventituous root heights were not recorded consistently at sites in 2022 making it impossible to estimate pre-fire conifer stand organic soil C pools for 2022-collected conifer plots). Only 2005-collected conifer plots were used in NCC paper analyses. For all plots, in addition to field/lab derived site characteristics and combustion metrics, post hoc remotely sensed metrics were derived: pre-fire NDVI/EVI-2 trends, 1980-2010 climate normals, and DOB weather metrics.

openOpenOct 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record