Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.7.1
Dataset results
67 results for “sampling design”
Forest-wide bird survey at 183 sample sites the Andrews Experimental Forest from 2009 to present (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-and/4781/5. The abstract below was extracted from the Level 0 data package and is included for context: Bird occurrence data collected at 183 sample locations within the H. J. Andrews Experimental Forest (HJA) from 2009-present. We used a stratified, systematic, random design to select sample locations. We stratified across elevation, distance to road, and habitat type (plantation or mature/old-growth forest). We conduct point counts on six separate occasions from May – July, which corresponded to spring arrival and subsequent breeding period for the majority of bird species at HJA. Surveys occur between 05:15h and 10:30h and each consists of a 10-min point count where we record all birds seen or heard. The species of all birds seen and heard are recorded as well as all individual squirrels, chipmunks and pikas seen and heard. Survey-level information is also collected at each point count and includes: weather and wind conditions, stream noise, snow cover on the ground, phenology of vine maple and rhododendron. Data collection is ongoing. The H.J. Andrews Experimental Forest is a living laboratory that provides unparalleled opportunities for the study of forest and stream ecosystems in the central Cascade Range of Oregon. Since 1980, as a part of the National Science Foundation Long Term Ecological Research (NSF-LTER) program, the Andrews Experimental Forest has become a leader in the analysis of forest and stream ecosystem dynamics. Long-term field experiments and measurement programs have focused on climate dynamics, streamflow, water quality, and vegetation succession. Currently researchers are working to develop concepts and tools needed to predict effects of natural disturba
Zooplankton density for all samples collected from Toolik Lake and lakes near the Toolik Field Station, Arctic LTER 2003 - 2017 (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-arc/10272/5. The abstract below was extracted from the Level 0 data package and is included for context: Zooplankton density,were taken with a 30 cm diameter plankton net with 156 um mesh plankton netting, for all samples collected from Toolik Lake and lakes near the Toolik Field Station, Arctic LTER from 2003 - 2017. The Arctic is one of the most rapidly warming regions on Earth. Responses to this warming involve acceleration of processes common to other ecosystems around the world (e.g., shifts in plant community composition) and changes to processes unique to the Arctic (e.g., carbon loss from permafrost thaw). The objectives of the Arctic Long-Term Ecological Research (LTER) Project for 2017-2023 are to use the concepts of biogeochemical and community “openness” and “connectivity” to understand the responses of arctic terrestrial and freshwater ecosystems to climate change and disturbance. These objectives will be met through continued long-term monitoring of changes in undisturbed terrestrial, stream, and lake ecosystems in the vicinity of Toolik Lake, Alaska, observations of the recovery of these ecosystems from natural and imposed disturbances, maintenance of existing long-term experiments, and initiation of new experimental manipulations. Based on these data, carbon and nutrient budgets and indices of species composition will be compiled for each component of the arctic landscape to compare the biogeochemistry and community dynamics of each ecosystem in relation to their responses to climate change and disturbance and to the propagation of those responses across the landscape.
CBP01 Variable distance line-transect sampling of bird population numbers in different habitats on Konza Prairie (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/26/11. The abstract below was extracted from the Level 0 data package and is included for context: Records of bird species based on line transect sampling, giving perpendicular distance of sighting from the transect line on 16 separate transects. Bird surveys were conducted 2-4 times per year in January, April, June, and October for a 29-year period from 1981 to 2009. Transects were designed to determine bird communities and population numbers associated with tallgrass prairie habitats with different experimental treatments (fire frequency, grazed by bison vs. ungrazed), riparian habitats on forest edge, and gallery forests dominated by oak woodland.
Database with GRTS sampling design used for the C-Mon project.
<p> A SQLite database holding a realisation of a GRTS design using the principles of Reverse Randomized Quadrant-Recursive Raster method (Theobold et al 2007). The database covers a square 2D grid with 32768 (2^15) pixels in both dimensions. This allows aselect and spatially balanced sampling in <a href="https://www.openstreetmap.org/relation/53134">Flanders</a> (Belgium) at 10 x 10 m resolution. In the framework of the C-Mon project, selection of plots for sampling soil organic carbon (SOC) stocks over all landuses is performed based on this realisation. It is envisaged that the plots will be monitored over decades to quantify SOC stock changes over time, along with landuse changes. . </p> <p>The R script used to generate the database and to sample from the database is provided. The algorithm itself is available on <a href="https://github.com/inbo/grtsdb/releases/tag/v0.1">GitHub</a> (10.5281/zenodo.2784016).</p> <p>The C-Mon project, entitled (in Dutch): 'Actualisatie van de onderbouwing van een methodiek voor de systematische monitoring van koolstofvoorraden in de bodem' was financed by the Department Vlaams Planbureau voor Omgeving from the Flemish Government. Project-ID: OMG/VPO/BODEM/TWOL/2017/1 </p>
Multi-decade land use and land cover samples for Brazil based in a stratified sampling design and visual interpretation of Landsat data (1985 — 2018)
<p>This dataset is composed by 85,152 random points throughout the Brazilian territory selected according to a stratified sampling design, based in 127 regular regions and six slope classes (<a href="https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-shuttle-radar-topography-mission-srtm-1-arc?qt-science_center_objects=0#qt-science_center_objects">SRTM</a>). Each sample was visually inspected by three independent interpreters, which associated all the land use and land cover (LULC) changes between 1985 and 2018, on a <strong>yearly basis</strong>, using as reference two <strong>Landsat</strong> images per year, a <strong>MODIS</strong> NDVI time series and high resolution images from <strong>Google Earth</strong>. </p> <p>This process was guided by a <a href="https://www.lapig.iesa.ufg.br/chave/">reference labeling protocol</a> which established the follow LULC classes:</p> <ul> <li><strong>Annual crop:</strong> Areas occupied with short to medium-term crops, usually with a vegetative cycle of less than one year, which after harvest needs to be re-planted. </li> <li><strong>Aquaculture:</strong> Artificial lakes, where aquaculture and/or salt production activities predominate</li> <li><strong>Beach and dune (Other):</strong> Sandy areas, with bright white color, where there is no vegetation predominance of any kind.</li> <li><strong>Forest formation:</strong> Vegetation types with predominance of tree species, with continuous canopy formation</li> <li><strong>Grassland formation:</strong> Grassland formations with predominance of herbaceous stratum</li> <li><strong>Mangrove (Other):</strong> Dense and Evergreen Forest formations, often flooded by tide and associated with the mangrove coastal ecosystem.</li> <li><strong>Mining (Other):</strong> Areas where clear signs of extensive mineral extractions are present, shows clear exposure of the soil by the action of heavy machinery. Only regions surrounding the AhkBrasilien (AHK) and the CPRM digital reference data were considered.</li> <li><strong>Not observed:</strong> Areas blocked by clouds or atmospheric noise, or with absence of ground observation masked out from analysis.</li> <li><strong>Other non-forest natural formations:</strong> Marshes (with fluvio-marine influence).</li> <li><strong>Other non-vegetated area (Other):</strong> Non-permeable surface areas (infrastructure, urban expansion or mining) not mapped into their classes</li> <li><strong>Pasture:</strong> Pasture areas, natural or planted, related with farming activity. In particular in the Pampa and Pantanal biomes part of the area classified as Grassland Formation also includes pasture areas.</li> <li><strong>Perennial crop:</strong> Areas occupied with crops with a long cycle (more than one year), which allow successive harvests without the need for new crop. </li> <li><strong>Rocky outcrop (Other)</strong>: Naturally exposed rocks without soil cover, often with the partial presence of rupicolous vegetation and high slope. </li> <li><strong>Salt flat (Other):</strong> "Apicuns" or Salt flats are formations often without tree vegetation, associated to a higher, hypersaline and less flooded area in the mangrove, generally in the transition between this area and the continent.</li> <li><strong>Savanna formation:</strong> Savanna formations with defined tree and shrub-herbaceous stratum</li> <li><strong>Semi-perennial crop:</strong> Cultivated areas with sugar cane</li> <li><strong>Tree plantation:</strong> Planted tree species for commercial use (e.g. Eucalyptus, Pinus and Araucaria)</li> <li><strong>Urban infrastructure:</strong> Urban areas with predominance of non-vegetated surfaces, including roads, highways and constructions.</li> <li><strong>Water:</strong> Rivers, lakes, dams, reservoir and other water bodies</li> <li><strong>Wetland:</strong> Wetlands with fluvial influence or swampy areas</li> </ul> <p>To enable a proper area estimation and accuracy assessment (<a href="https://www.tandfonline.com/doi/abs/10.1080/01431161.2014.930207">Stehman, 2014</a>) the dataset is provided with the <strong>sampling probability</strong> for each sample (<em>brazil_lulc_samples_1985_2018</em> and <em>brazil_lulc_samples_1985_2018_row_wise</em>) and the <strong>sampling weight</strong> (<em>brazil_lulc_samples_1985_2018_row_wise</em>), which was adjusted to disregard the "<strong>Not observed" </strong>class. The number of votes for the associated LULC class (visual interpretation agreement) and an indication if the sample is between two different LULC<strong> </strong>classes (<strong>border flag</strong>) are also provided.</p> <p>The samples were used to produce several <strong><a href="https://github.com/lapig-ufg/tvi-analysis">area estimation analyses</a></strong>, including land use and land cover dynamics, historical deforestation and agricultural expansion of Brazil. A publication describing in detail the methodology and the analysis is under preparation.</p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
CGR02 Sweep sampling of Grasshoppers on Konza Prairie LTER watersheds (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-knz/29/18. The abstract below was extracted from the Level 0 data package and is included for context: Sweep samples were taken for grasshoppers (Acrididae) at two sites for each of 14 Konza Prairie LTER watersheds. Samples are taken in late July to early August. At each site on each occasion, 10 sets of 20 sweeps (200 sweeps total) are taken. Stored data include for each site on each occasion: total number of each species (all instars combined) collected and total number for each instar for each species (200 sweeps combined).
[DEPRECATED] CFC01 Kings Creek long-term fish and crayfish community sampling at Konza Prairie (Reformatted to ecocomDP Design Pattern)
This data package has been deprecated due to several issues in the L0 source dataset that prohibits the creation of an L1 ecocomDP dataset. This data package is formatted according to the "ecocomDP", a data package design pattern for ecological community surveys, and data from studies of composition and biodiversity. For more information on the ecocomDP project see https://github.com/EDIorg/ecocomDP/tree/master, or contact EDI https://environmentaldatainitiative.org. This Level 1 data package was derived from the Level 0 data package found here: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-knz&identifier=130&revision=4 The abstract below was extracted from the Level 0 data package and is included for context:
Macrobenthos Sampling data for the North Inlet Estuary, Georgetown,South Carolina, from 1981 to 1992 North Inlet LTER (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-nin/9/1. The abstract below was extracted from the Level 0 data package and is included for context: This data package consists of Macrobenthos Sampling Data for North Inlet Stations Bread and Butter Creek from 1981 to 1992, and Debidue Creek from 1981 to 1984, North Inlet LTER. The purpose of this study was to document the composition and abundance of macrobenthic subtidal populations over time at one mud and one sand site. Macrobenthos was defined here as those animals retained on a 0.5 mm mesh screen.
LTER Epibenthos Sampling Data for North Inlet Estuary, Georgetown, South Carolina from 1981 to 1992, North Inlet LTER (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-nin/7/1. The abstract below was extracted from the Level 0 data package and is included for context: This data package consists of Epibenthos Sampling for North Inlet Stations Bread and Butter Creek, from 1981 to 1992, and Debidue Creek from 1981 to 1984, The purpose of the long term monitoring of Epibenthos was to determine seasonal and inter-annual changes in the taxonomic/life stage composition and abundance of small motile epibenthic invertebrates and fishes (1-20 mm in length) in the major sub- tidal habitats of North Inlet estuary.
Fig 4 in Do different sampling designs produce differences in the metrics of curimba, Prochilodus lineatus (Characiformes: Prochilodontidae)?
Fig 4. Frequency distribution by standard length (SL) class of curimba for fixed and variable sampling sites in Volta Grande (VGR) and Jaguara (JR) reservoirs.
Fig. 3 in Do different sampling designs produce differences in the metrics of curimba, Prochilodus lineatus (Characiformes: Prochilodontidae)?
Fig. 3. Temporal variation in catch per unit effort (CPUE) of curimba for fixed and variable sampling sites in Volta Gran- de (VGR) and Jaguara (JR) reservoirs.
Global Pasture Watch - Grassland sampling design derived by Feature Space Coverage Sampling (FSCS) at 1-km spatial resolution
<p>Sampling design used in the production of the <strong>global maps of grassland dynamics 2000–2022 at 30 m spatial resolution</strong> in the scope of the Global Pasture Wath initiative. The sampling desing was based in Feature Space Coverage Sampling and resulted in 10,000 sample tiles (1x1 km) distributed across the World, which were visual interpreted in Very-High Resolution imagery thorugh the QGIS plugin <a href="https://plugins.qgis.org/plugins/qgis-fgi-plugin/">QGIS Fast Grid Inspection</a>.</p> <p>FSCS steps include:</p> <ul> <li>Short vegetation mask that includes all pixels mapped as mosaic, shrubland, grassland, and sparse vegetation in at least one year from 1993 to 2021 according to <a href="https://www.esa-landcover-cci.org/">ESA/CCI global land cover</a> (<code>gpw_short.veg.mask_esacci.lc_p_1km_s_19920101_20201231_go_epsg.3857_v1.tif</code>),</li> <li>87 input raster layers (including vegetation indices, terrain, land temperature, climate and water variable),</li> <li>Principal Components Analysis (PCA) using all input layers,</li> <li>Selection of the 10 first components (explaining 75% of variance),</li> <li> K-Means with 10,000 clusters (targeted number of samples - <br><code>gpw_grassland_fscs.kmeans.cluster_c_1km_20000101_20221231_go_epsg.3857_v1.tif</code>)</li> <li>Calculation of euclidean distance (in the principal component space) of all 1-km pixels to the centre of each cluster,</li> <li>Selection of the pixel with the shortest distance for each cluster,</li> <li>Conversion of the selected pixels into sample tiles ()</li> </ul> <p>The file <code>gpw_grassland_fscs_tile.samples_1km_20000101_20221231_go_epsg.3857_v1.gpkg</code> provides the sample tiles and include the follow collumns:</p> <ul> <li><strong>X</strong>: Latitude in Web Mercator projection (EPSG:3857),</li> <li><strong>Y</strong>: Longitude in Web Mercator projection (EPSG:3857),</li> <li><strong>cluster_id</strong>: K-Means output ranging from 0—9999,</li> <li><strong>cluster_distance</strong>: Distance from the selected sample to the centre of the cluster,</li> <li><strong>cluster_size</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Web Mercator projection (<a href="https://epsg.io/3857">EPSG:3857</a>)</li> <li><strong>cluster_size_equal_area</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Goode Homolosine Land projection (<a href="https://epsg.io/54052">ESRI:54052</a>)</li> <li><strong>cluster_size_corr</strong>: Correction factor to adjust the area distortion due to Web Mercator projection, estimated by the difference in normalized propotional values of cluster_size and cluster_size_equal_area.</li> <li><strong>rf_n_pred</strong>: Number of pixels predicted by a RF model trained to estimate probability to select the pixel closer to the centre of the KMeans cluster. The RF models were trained individually per each cluster using the 10 first components derived by PCA (<code>gpw_comps_fscs.pca_m_1km_20000101_20221231_go_epsg.3857_v1.tar.gz</code>).</li> <li><strong>rf_samp_prob</strong>: Sampling probability based on RF model (<em>rf_n_pred / cluster_size</em>)</li> <li><strong>rf_samp_wei</strong>: Sampling weight estimated in Web Mercator projection.</li> <li><strong>rf_samp_wei_coor</strong>: Corrected sampling weight estimated in Goode Homolosine Land projection.</li> </ul> <h3>Related resources</h3> <ul> <li><strong>Maps of dominant grassland:</strong><br><a href="https://zenodo.org/records/13890400">2000-2002</a> <a href="https://zenodo.org/records/13890402">2003-2005</a> <a href="https://zenodo.org/records/13890404">2006-2008</a> <a href="https://zenodo.org/records/13890408">2009-2011</a> <a href="https://zenodo.org/records/13890410">2012-2014</a> <a href="https://zenodo.org/records/13890412">2015-2017</a> <a href="https://zenodo.org/records/13890414">2018-2020</a> <a href="https://zenodo.org/records/13890416">2021-2022</a></li> <li><strong>Probability maps of cultivated grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Probability maps of natural/semi-natural grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Grassland reference samples based on VHR imagery (2000–2022):</strong><br><a href="https://doi.org/10.5281/zenodo.11281157">GeoPackage files</a></li> <li><strong>Global machine learning models (Random Forest):</strong><br><a href="https://doi.org/10.5281/zenodo.13952806">Parquet and joblib python files</a></li> <li><strong>Reference sampling design derived by FSCV:</strong><br><a href="https://doi.org/10.5281/zenodo.11391517">GeoPackage and raster files</a></li> <li><strong>Harmonized reference samples based on existing LULC dataset:</strong><br><a href="https://doi.org/10.5281/zenodo.13951976">GeoPackage and raster files</a></li> <li><strong>Source code for reproducibility:<br></strong><a href="https://doi.org/10.5281/zenodo.13952867">GitHub release</a><strong><br></strong></li> <li><strong>Mapping feedback tool:</strong><br><a href="https://geo-wiki.org">GeoWiki</a></li> <li><strong>Data catalogues:</strong><br><a href="https://stac.openlandmap.org/gpw_ggc-30m/collection.json?.language=en">OpenLandMap STAC</a> <a href="https://global-pasture-watch.projects.earthengine.app/view/ggc-30m">Google Earth Engine</a></li> </ul> <h3>Support</h3> <p>For questions of bugs/inconsistencies related to the dataset raise a GitHub issue in <a href="https://github.com/wri/global-pasture-watch">https://github.com/wri/global-pasture-watch</a></p>
Spanning Scales: The Airborne Spatial and Temporal Sampling Design of the National Ecological Observatory Network
<p>Supporting information, datasets, and R and JavaScript code for the the National Ecological Observatory Network’s Airborne Observation Platform (AOP) sampling design and publication, <em>"Spanning Scales: The Airborne Spatial and Temporal Sampling Design of the National Ecological Observatory Network"</em></p>
How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design
<p>Extra-organismal DNA (eoDNA) from material left behind by organisms (non-invasive DNA: e.g., faeces, hair) or from environmental samples (eDNA: e.g., water, soil) is a valuable source of genetic information. However, the relatively low quality and quantity of eoDNA, which can be further degraded by environmental factors, results in reduced amplification and sequencing success. This is often compensated for through cost- and time-intensive replications of genotyping/sequencing procedures. Therefore, system- and site-specific quantifications of environmental degradation are needed to maximize sampling efficiency (e.g., fewer replicates, shorter sampling durations), and to improve species detection and abundance estimates. Using ten environmentally diverse bat roosts as a case study, we developed a robust modelling pipeline to quantify the environmental factors degrading eoDNA, predict eoDNA quality, and estimate sampling-site-specific ideal exposure duration. Maximum humidity was the strongest eoDNA-degrading factor, followed by exposure duration and then maximum temperature. We also found a positive effect when hottest days occurred later. The strength of this effect fell between the strength of the effects of exposure duration and maximum temperature. With those predictors and information on sampling period (before or after offspring were born), we reliably predicted mean eoDNA quality per sampling visit at new sites with a mean squared error of 0.0349. Site-specific simulations revealed that reducing exposure duration to 2-8 days could substantially improve eoDNA quality for future sampling. Our pipeline identified high humidity and temperature as strong drivers of eoDNA degradation even in the absence of rain and direct sunlight. Furthermore, we outline the pipeline's utility for other systems and study goals, such as estimating sample age, improving eDNA-based species detection, and increasing the accuracy of abundance estimates.</p>
Optimizing sampling design for landscape genomics
Open the record for dataset details and reuse information.
How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design
Open the record for dataset details and reuse information.
Evaluating alternative study designs for optimal sampling of species' climatic niches
Ecologists have traditionally studied intraspecific variation by sampling species across their geographic ranges. However, whether this classic approach produces samples that accurately represent species' climatic niches is largely unknown. Alternative, niche-based study designs using species' climatic niches to inform sampling locations should more efficiently and completely capture the breadth of the niche, but the magnitude of this difference and how it may vary is unclear. Here we use conifers as a model system to explore these issues and reach specific recommendations for future sampling designs. Using an independent dataset of high-quality species' occurrences, we first show that recent publications examining variation across geographic space do a poor job of capturing the full breadth of species' niches, such that on average, only 22% of species' niche space was sampled. This was also true of a large compiled database, the International Tree-Ring Data Bank (ITRDB), which yielded average niche coverage of only 45%. Finally, we simulated common sampling designs (i.e., random points, grids, and transects) in both geographic and niche-based sampling frameworks. Using two sampling metrics, niche coverage and niche undersampling, we measured how completely and evenly these simulated studies characterized the niches of 64 North American conifers. Niche-based sampling better represented species' niches than geographic sampling, with the magnitude of this difference depending on study design and sample size. Niche-based gridded study designs achieved the most complete sampling at all but the smallest sample sizes, covering ~15-25% more of a species' niche than similar designs implemented geographically. With fewer than 10 samples, however, all study designs performed poorly, and niche-based transects achieved slightly higher niche coverage. Consequently, when more than a handful of samples are collected, we recommend that studies seeking to characterize variation across a species' niche consider using a gridded study design implemented in a niche-based sampling framework.
Dataset and R code from: Positive and negative effects of land abandonment on butterfly communities revealed by a hierarchical sampling design across climatic regions
<p class="MsoNormal">Land abandonment may decrease biodiversity but also provides an opportunity for rewilding. It is therefore necessary to identify areas that may benefit from traditional land management practices and those that may benefit from a lack of human intervention. In this study, we conducted comparative field surveys of butterfly occurrence in abandoned and inhabited settlements in 18 regions of diverse climatic zones in Japan to test the hypotheses that species-specific responses to land abandonment correlate with climatic niches and habitat preferences. Hierarchical models that unified species occurrence and habitat preferences revealed that negative responses to land abandonment were associated with species that have cold climatic niches and utilize open habitats, suggesting that species negatively impacted by land abandonment will decline more due to future climate warming. Maps representing species gains and losses due to land abandonment, which were created from the model estimates, showed similar geographic patterns, but some areas exhibited high species losses relative to gains. Our hierarchical modelling approach was useful for scaling up local-scale effects of land abandonment to a macro-scale assessment, which is crucial to developing spatial conservation strategies in the era of depopulation.</p>
Cone-Beam X-Ray CT Data Collection Designed for Machine Learning: Samples 38-42
<p>This upload contains samples 38 - 42 from the data collection described in</p> <p>Henri Der Sarkissian, Felix Lucka, Maureen van Eijnatten, Giulia Colacicco, Sophia Bethany Coban, Kees Joost Batenburg, "A Cone-Beam X-Ray CT Data Collection Designed for Machine Learning", <em>Sci Data</em> <strong>6, </strong>215 (2019). <a href="https://doi.org/10.1038/s41597-019-0235-y">https://doi.org/10.1038/s41597-019-0235-y</a> or <a href="https://arxiv.org/abs/1905.04787">arXiv:1905.04787</a> (2019)</p> <p>Abstract:<br> "Unlike previous works, this open data collection consists of X-ray cone-beam (CB) computed tomography (CT) datasets specifically designed for machine learning applications and high cone-angle artefact reduction: Forty-two walnuts were scanned with a laboratory X-ray setup to provide not only data from a single object but from a class of objects with natural variability. For each walnut, CB projections on three different orbits were acquired to provide CB data with different cone angles as well as being able to compute artefact-free, high-quality ground truth images from the combined data that can be used for supervised learning. We provide the complete image reconstruction pipeline: raw projection data, a description of the scanning geometry, pre-processing and reconstruction scripts using open software, and the reconstructed volumes. Due to this, the dataset can not only be used for high cone-angle artefact reduction but also for algorithm development and evaluation for other tasks, such as image reconstruction from limited or sparse-angle (low-dose) scanning, super resolution, or segmentation."</p> <p>The scans are performed using a custom-built, highly flexible X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://xre.be/">XRE nv</a>and located in the FleX-ray Lab at the <a href="https://www.cwi.nl/">Centrum Wiskunde & Informatica (CWI)</a> in Amsterdam, Netherlands. The general purpose of the FleX-ray Lab is to conduct proof of concept experiments directly accessible to researchers in the field of mathematics and computer science. The scanner consists of a cone-beam microfocus X-ray point source that projects polychromatic X-rays onto a 1536-by-1944 pixels, 14-bit flat panel detector (Dexella 1512NDT) and a rotation stage in-between, upon which a sample is mounted. All three components are mounted on translation stages which allow them to move independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete data set can be found via the following links: <a href="https://doi.org/10.5281/zenodo.2686725">1-8</a>, <a href="https://doi.org/10.5281/zenodo.2686970">9-16</a>, <a href="https://doi.org/10.5281/zenodo.2687386">17-24</a>, <a href="https://doi.org/10.5281/zenodo.2687634">25-32</a>, <a href="https://doi.org/10.5281/zenodo.2687896">33-37</a>, <a href="https://doi.org/10.5281/zenodo.2688111">38-42</a></p> <p>The corresponding Python scripts for loading, pre-processing and reconstructing the projection data in the way described in the paper can be found on <a href="https://github.com/cicwi/WalnutReconstructionCodes">github</a></p> <p>For more information or guidance in using these dataset, please get in touch with</p> <ul> <li>henri.dersarkissian [at] gmail.com</li> <li>Felix.Lucka [at] cwi.nl</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.