Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,015

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,015 results for “occurrence”

Learn how ShareScore rates datasets ↗
zenodo48/100

Monthly maps of Warm Core Ring Occupancy and occurrences of Salinity Maximum Intrusions in the Slope Sea (1990-2019)

<p>This dataset presents two important variables across the shelfbreak in the Northwest Atlantic: (i) Warm Core Ring Occupancy in the Slope sea proximate to the shelfbreak; and (ii) locations of Salinity maximum intrusions in the shelf. Monthly fields of both of these fields together are presented for the period 1990-2019 with file name format&nbsp;<em>smax_ring_mm_yyyy.jpg.&nbsp;</em>The gray and red dots on the shelf represent locations of profiles taken from the Ecosystem Monitoring Program&rsquo;s (EcoMon) hydrographic data (available from the National Centers for Environmental Information World Ocean Database accessible at&nbsp;<a href="http://www.ncei.noaa.gov/products/world-ocean-database">www.ncei.noaa.gov/products/world-ocean-database</a>). Red dots show locations of profiles which contained mid-depth salinity maximum intrusions, gray dots are profiles without any mid-depth salinity maximum intrusion. Profiles with intrusions were identified using the methodology of Gawarkiewicz et al., 2022. The ring occupancy was calculated from a Warm Core Ring Tracking dataset with ring tracks from 2000-2010 and 2011- 2020 available from Zenodo (<a href="https://doi.org/10.5281/zenodo.6436380">https://doi.org/10.5281/zenodo.6436380</a>,&nbsp;<a href="https://doi.org/10.5281/zenodo.7406675">https://doi.org/10.5281/zenodo.7406675</a>) and ring trajectories from&nbsp;1978 through 1999 available from the Bedford Institute of Oceanography, Canada. To calculate the ring occupancy the region was sub-divided into 0.1 by 0.1 degree bins. Ring trajectories and approximate geographical range (calculated from the ring area, assuming the ring is a perfect circle) were overlain on this region and the days rings are present in each bin are counted in units of ring days. A ring day is the presence of one single ring in a bin during a given day. These ring day counts were converted to percentages, dividing by days in the given month and multiplying by 100. For more details see Silver et al., 2022 and Salois et al., 2023. From these figures one can see the spatial relationship between Warm Core Rings and Salinity Maximum Intrusions, with clusters of intrusions occurring in areas adjacent to high ring occupancy.&nbsp;</p> <p>An animation of two particular years is also presented in&nbsp;<em>movie_smax_ring_1993_2012.gif</em>&nbsp;to highlight this relationship: Low ring year (1993) leading to fewer Smax intrusion and high ring year (2012) leading to more intrusions.</p> <p>&nbsp;</p> <p>Gawarkiewicz, G., Fratantoni, P., Bahr, F., &amp; Ellertson, A. (2022). Increasing Frequency of Mid‐Depth Salinity Maximum Intrusions in the Middle Atlantic Bight.&nbsp;<em>Journal of Geophysical Research: Oceans</em>,&nbsp;<em>127</em>(7), e2021JC018233.&nbsp;<a href="https://doi.org/10.1029/2021JC018233">https://doi.org/10.1029/2021JC018233</a></p> <p>Silver, A., Gangopadhyay, A., Gawarkiewicz, G., Andres, M., Flierl, G., &amp; Clark, J. (2022). Spatial Variability of Movement, Structure, and Formation of Warm Core Rings in the Northwest Atlantic Slope Sea.&nbsp;<em>Journal of Geophysical Research: Oceans</em>,&nbsp;<em>127</em>(8), e2022JC018737.&nbsp;<a href="https://doi.org/10.1029/2022JC018737">https://doi.org/10.1029/2022JC018737</a>&nbsp;</p> <p>Salois, S. L., Hyde, K. J., Silver, A., Lowman, B. A., Gangopadhyay, A., Gawarkiewicz, G., ... &amp; Lapp, M. (2023). Shelf break exchange processes influence the availability of the&nbsp;northern shortfin squid, Illex illecebrosus, in the Northwest Atlantic.&nbsp;<em>Fisheries Oceanography</em>.&nbsp;<a href="https://doi.org/10.1111/fog.12640">https://doi.org/10.1111/fog.12640</a>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Occurrence dataset for the subspecies of the American badger (Taxidea taxus berlandieri) in the north-central region of Mexico

<p>The subspecies of American badger (<em>Taxidea taxus berlandieri </em>Baird, 1858), also called tlalcoyote (Figure 1), is distributed in north-central Mexico. However, its occurrence records are scarce and the few that exist are uncertain due to incorrect georeferencing or identification of the taxonomic unit. In view of this, we disgned a spatial sampling in part of the states of Coahuila de Zaragoza, Durango, Nuevo Le&oacute;n, San Luis Potos&iacute; and Zacatecas. In this north-central protion of Mexico, we generated a grid of squares measuring 5 &times; 5 km over the entire study area using QGIS&reg; 3.10 software.&nbsp; Subsequently, we excluded squares that included urban settlements, agricultural land, or water bodies in more than 30% of their extension; we also descarted squares located at an altitude over 2,250 meters above sea level. &nbsp;To perform this filtering, we used both the land use and vegetation chart of the INEGI [Instituto Nacional de Estad&iacute;stica, Geograf&iacute;a e Inform&aacute;tica] (2018) and the Digital Elevation Model (DEM) downloaded from the USGS page [United States Geological Survey] (2019) as a basis. &nbsp;As result, we obtained 3,471 squares separated by at least 5 km. &nbsp;Then, through simple random sampling, 177 (&asymp;5%) squares were selected, where we generated centroids to be used as sampling sites.&nbsp; &nbsp;</p> <p>In field work, between 2009 and 2015, at these 177 sites we traced a 10 &times; 100 m transect, where we searched for<em> T. t. berlandieri</em> signs (i.e., burrows and scratching posts). In this case, their burrows and scratching posts are easily observed and quantified, and there is no chance of mistaking them for burrows of other species (Long 1973; Merlin 1999). Also, we recorded possible sightings, as other studies (e.g., Merlin 1999; Elbroch 2003).&nbsp; As result, we only found 33 with signs of occurrence.&nbsp;&nbsp;</p> <p><a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Figure%201.%20Taxidea%20taxus%20Berlandieri%20Baird%2C%201858.jpeg">Figure 1.</a> Individual of tlalcoyote (<em>Taxidea taxus Berlandieri</em>). Photo obtained from Naturalista (2023) and uploaded by David Molina&copy;. All rights reserved (CC BY-NC-ND).</p> <p>To increase the number of records, we included occurrence data from GBIF [Global Biodiversity Information Facility portal] (2022). We downloaded only the records that included coordinates and that their basis of registration was "preserved specimen". This, because they are correctly identified as specimens from biological collections (Maldonado&nbsp;<em>et al.</em> 2015). In addition, we only selected records for Mexico. Subsequently, we filtered the downloaded database, discarding records that were incorrectly georeferenced, with atypical and duplicate coordinates, as well as with low geospatial accuracy (e.g., less than three decimals of precision).</p> <p>We loaded the remaining data into the QGIS&reg; software and performed a spatial filtering, where we excluded data that were outside the study area, located in unlikely areas (e.g., human settlements, bodies of water, agricultural areas) and with a distance of less than 5 km from the records obtained in the field. This gave a total of 10 records from the GBIF portal. Finally, we loaded the raster layers of elevation (Elev; INEGI 2007), normalized difference vegetation index (NDVI, USGS 2019) and the slope of the terrain into the software to extract the pixel values based on the GBIF records and those obtained in the field. With this, we generated a new global dataset to which we performed environmental filtering to find environmental outliers. We plotted the normality distribution of the data for each variable and the dispersion of the data among the variables.&nbsp; In this filtering, we conserve all records. Figure 2 shows the normality distribution of the records as a function of Elev. Figure 3 shows the dispersion of the data between Elev and NDVI.</p> <p><a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Figure%202.%20Normal%20distribution.png">Figure 2.</a>&nbsp;Normality distribution of <em>T. t. berlandieri</em> occurrence records as a function of the elevation variable (Elev).</p> <p><a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Figure%203.%20Scatter%20plot.png">Figure 3.</a>&nbsp;Scatter plot of <em>T. t. berlandieri</em> occurrence records as a function of elevation (Elev) and normalized difference vegetation index (NDVI).</p> <p>For the north-central region of Mexico, we present the global database (i.e.,&nbsp;<a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Tatabe_joint.csv">Tatabe_joint.csv</a>), as well as the database that contains only the field evidence records (i.e.,&nbsp;<a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Tatabe_first_order.csv">Tatabe_first_order.csv</a>) and another one with the filtered GBIF records (i.e.,&nbsp;<a href="https://zenodo.org/api/files/9a8452c6-15e2-43cd-9c07-27b7fc1d422a/Tatabe_GBIF.csv">Tatabe_GBIF.csv</a>).</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Prevalent trends in realized probability of occurrence of main European forest tree species for 2000–2020

<p>High resolution maps resulting from a trend analysis conducted for the period 2000&ndash;2020 on the probability of occurrence maps prepared by <a href="https://doi.org/10.7717/peerj.13728">Bonannella et al. (2022)</a>. For this analysis we selected the realized distribution time series layers at 30m spatial resolution for 6 out of 16 species described in the mentioned publication:</p> <ul> <li>Silver fir (<em>Abies alba </em>Mill.)</li> <li>European beech (<em>Fagus sylvatica </em>L.)</li> <li>Norway spruce (<em>Picea abies </em>L.)</li> <li>Black pine (<em>Pinus nigra </em>J. F. Arnold)</li> <li>Scots pine (<em>Pinus sylvestris </em>L.)</li> <li>Common oak (<em>Quercus robur </em>L.)</li> </ul> <p>The trend analysis was conducted per pixel on each of these species individually. We fitted simple OLS regression models with the probability of occurrence as the dependent variable and time as the independent variable. After the model fitting, we also calculated the t-test statistics to determine the presence of an increasing (positive) or decreasing (negative) trend or no trend at all.</p> <p>By combining the regression slope coefficient (<em>&beta;</em>) and the <em>p</em>-value from the t-test statistics we assigned each pixel to one of three classes:</p> <ul> <li><em>positive</em>: <em>&beta;</em> &gt; 0.25 AND <em>p</em>-value &lt; 0.05</li> <li><em>negative</em>: <em>&beta;</em> &lt; &minus;0.25 AND <em>p</em>-value &lt; 0.05</li> <li><em>no trend / stable</em>: &minus;0.25 &le; <em>&beta;</em> &ge; 0.25 OR <em>p</em>-value &gt; 0.05</li> </ul> <p>We then aggregated the resulting classes at 1km resolution maps to capture the prevalent trend in probability of occurrence over a certain area. Files are named according to the following naming convention, e.g.:</p> <ul> <li>veg_abies.alba_slope_30m_0..0cm_epsg3035_v1.0</li> </ul> <p>with the following fields:</p> <ul> <li>theme: e.g. <strong>veg</strong>,</li> <li>species code: e.g. <strong>abies.alba</strong>,</li> <li>variable name: e.g. <strong>slope</strong>,</li> <li>resolution in meters e.g. <strong>30m</strong>,</li> <li>reference depths (vertical dimension): e.g. <strong>0..0cm</strong>,</li> <li>coordinate system: e.g. <strong>epsg3035</strong>,</li> <li>data set version: e.g. <strong>v1.0</strong>.</li> </ul> <p>For each species here we provide the following layers:</p> <ul> <li>veg_abies.alba_<strong>slope</strong>:<strong> </strong>slope coefficient (scaling factor: 10000)</li> <li>veg_abies.alba_<strong>pvalue</strong>:<strong> </strong><em>p</em>-value (scaling factor: 1000)</li> <li>veg_abies.alba_<strong>pos.trends_30m</strong>: pixels classified as <em>positive </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>pos.trends_1km</strong>: proportion of pixels of the <em>positive </em>class over a 1&times;1 km area (range 0&ndash;100)</li> <li>veg_abies.alba_<strong>neg.trends_30m</strong>: pixels classified as <em>negative </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>neg.trends_1km</strong>: proportion of pixels of the <em>negative </em>class over a 1&times;1 km area (range 0&ndash;100)</li> <li>veg_abies.alba_<strong>no.trends_30m</strong>: (pixels classified as <em>no trend / stable </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>no.trends_1km</strong>:<strong> </strong>proportion of pixels of the <em>no trend / stable </em>class over a 1&times;1 km area (range 0&ndash;100)</li> </ul> <p>Files are provided as GeoTIFFs and projected in the Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035). Styling files are provided in <em>QML</em> format</p> <p>A publication describing, in detail, all processing steps is currently in review. See at:<br> <br> Bonannella, C., Parente, L., de Bruin, S. and Herold, M. (2023). Multi-decadal trend analysis and forest disturbance assessment of European tree species: concerning signs of a subtle shift, PREPRINT (Version 1) available at Research Square [<a href="https://doi.org/10.21203/rs.3.rs-3288937/v1">https://doi.org/10.21203/rs.3.rs-3288937/v1</a>]</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
edi48/100

Data from: Heterogeneity in habitat and nutrient availability facilitate the co-occurrence of N2 fixation and denitrification across wetland - stream - lake ecotones of Lakes Superior and Huron

Great Lakes coastlines are mosaics of wetland, stream, and lake habitats, characterized by a high degree of spatial heterogeneity that may facilitate the co-occurrence of seemingly incompatible biogeochemical processes due to variation in environmental factors that favor each process. We measured nutrient limitation and rates of N2 fixation and denitrification along transects in 5 wetland - stream - lake ecotones with different nutrient loading in Lakes Superior and Huron and hypothesized that rates of both processes would be related to nutrient limitation status, habitat type, and environmental characteristics including temperature, nutrient concentrations, and organic matter quality. This data package includes information on sampling sites, dates and locations; rates of N fixation and denitrification measured at each site, date and transect location; and biomass information from nutrient diffusing substrates deployed on the study transects.

openCC (other)Jun 2023View details →
edi48/100

Freshwater insect occurrences and traits for the contiguous United States, 2001 - 2018

Freshwater insects comprise 60% of the diversity of freshwater animals and are essential indicators of ecosystem health. Yet our knowledge of the distribution of freshwater insect diversity in the United States is incomplete because we lack comprehensive data on freshwater insect distributions and traits at a continental scale. We fill this knowledge gap by presenting a database of occurrence records and functional traits for freshwater insects in the contiguous US. We compiled over 2.05 million genus occurrence records for 932 genera in the major freshwater insect orders, at 51,044 stream locations sampled between 2001 and 2018 by federal and state biological monitoring programs. We assembled life history, dispersal, morphology, and ecology traits from existing databases, scientific books, and the primary literature, and present traits and trait affinities for 1,072 insect genera. Our freshwater insects CONUS database can be used to map freshwater insect taxonomic and functional diversity, and when paired with environmental data, will provide a powerful resource for quantifying how the environment shapes diversity patterns as well as taxon-specific distributions across the contiguous US.

openCC (other)Aug 2020View details →
edi48/100

Mammal occurrence data derived from camera traps in grassland-shrubland ecotones at 24 sites in the Jornada Basin, southern New Mexico, USA, 2014-ongoing

The objective of this ongoing study is to investigate how abundance, distribution, and activity of mammals (>= 1 kg) vary across grassland to shrubland ecotones in the northern Chihuahuan Desert. This dataset includes animal occurrence data derived from camera trap images captured in 24 grassland-to-shrubland ecotone sites in the Jornada Basin, Dona Ana County, New Mexico, USA. The data set contains occurrence records from 14 mammal species with the date and time a species was detected. Also included are the number of individuals in a photo, operational dates and number of functional camera days for cameras, total number of trap nights a camera was active, and geographical coordinates of camera trap locations. Sampling is ongoing and occurs during the monsoon season from July-November. Sampling has occurred annually since 2014.

openCC (other)Aug 2024View details →
edi48/100

Records of Sargassum horneri occurrence in the eastern Pacific

Presented here are records of the occurrence of Sargassum horneri in California, USA, and Baja California, Mexico, since 2003, the year it was first discovered in the eastern Pacific. These data and their sources were published as supplementary tables in: Marks LM, Salinas-Ruiz P, Reed DC, Holbrook SJ, Culver CS, Engle JM, Kushner DJ, Caselle JE, Freiwald J, Williams JP, Smith JR, Aguilar-Rosas LE, Kaplanis NJ (2015) Range expansion of a non-native, invasive, macroalga Sargassum horneri (Turner) C. Agardh, 1820 in the eastern Pacific. BioInvasions Records 4, DOI: 10.3391/bir.2015.4.4.02

openCC (other)Oct 2022View details →
zenodo44/100

Occurrence data used to create species distribution models and apply an evaluation method

<p>These two files containing&nbsp;a table with three columns: species names, longitude, latitude. Each row of the tables represents a georeferenced presence record for the corresponding species. The original presence data were downloaded from the GBIF database and after going through a cleaning process, we ended with these records that passed all the tests.</p> <p>These datasets were used to create species distribution models (SDMs) that were then used to apply a new method to evaluate the performance of different SDMs. Jim&eacute;nez &amp; Sober&oacute;n (2020)</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Risk assessment on Glycoalkaloids in feed and food: Occurrence data in food and feed submitted to EFSA and dietary exposure assessment for humans

<p><strong>UPDATE to version 2 of this upload:</strong></p> <p>Also the raw (no data cleaning applied to it) &nbsp;occurrence dataset as extracted from EFSA DWH is provided <em>in csv format</em>. This dataset is compliant with EFSA SSD model and contains two additional columns documenting issues identified in the cleaning process (column: issue) and the action taken (column: action) to address the issue (e.g. delete record or update values in specific fields).</p> <p><strong>Description - Version 1</strong></p> <p><strong>Annex:&nbsp;&nbsp;Tables on GAs on occurrence data in food and feed, and dietary exposure assessment for humans</strong></p> <p>Table A.1. Dietary surveys used for the estimation of acute dietary exposure to GA</p> <p>Table A.2. Number of results and samples per food category submitted to EFSA through the continuous call for data</p> <p>Table A.3. Analytical results excluded from the final dataset used to estimate dietary exposure and the criteria applied for exclusion</p> <p>Table A.4. Occurrence of alpha-chaconine and alpha-solanine (UB mg/kg) in the samples included in the final dataset (left censored results highlighted in yellow)</p> <p>Table A.5. European Starch Association data on feed and potatoes for starch</p> <p>Table A.6. Details acute assessment across surveys (consumption days only)</p> <p>Table A.7. Comparison of exposure summary results obtained using the uniform vs the normal distribution for reduction factors</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

SHAPE-ID Literature Review dataset: journal occurrences with ASJC codes

<p><strong>Background and methodology:</strong></p> <p>The dataset consists of a list of 2202 journal titles represented in the <a href="https://doi.org/10.5281/zenodo.4034507">SHAPE-ID Literature Review bibliography</a>, prepared for the purposes of quantitative analysis.</p> <p>The list of journals is based on 3955 journal articles in the bibliography dataset that had an&nbsp;International Standard Serial Number&nbsp;(ISSN). To each journal title the project team attributed:</p> <p>- a weight factor based on how many articles from the given journal&nbsp; featured in bibliography dataset</p> <p>- at least one <a href="https://service.elsevier.com/app/answers/detail/a_id/15181/supporthub/scopus/">All Science Journal Classification</a> (ASJC) code, representing different scientific disciplines</p> <p>- a country of publication.&nbsp;</p> <p>In case of 1853 of those journal titles, the attribution was automatised (we matched the ISSNs of journal titles in our sample against the Scopus Sources list from February 2019). In case of the remaining 349 titles the attribution was accomplished manually, based on the information available in SCOPUS, Web of Science, JSTOR,&nbsp;Information Matrix for the Analysis of Journals&nbsp;(MIAR) and ISSN databases.</p> <p><strong>Description of the file:</strong></p> <p>This is a csv file containing a list of 2202 journal titles represented in the SHAPE-ID Literature Review bibliography, with country of publication and ASJC codes assigned.&nbsp;</p> <p>The file is formatted as follows:</p> <p>Column A: ISSN of the journal</p> <p>Column B: information on how country and ASJC codes were attributed. Value &ldquo;N&rdquo; indicates automatic attribution based on match with Scopus list of sources. Other values indicate manual attribution. Values WOS, SCOPUS, JSTOR indicate source of information. Valu &ldquo;Y&rdquo; indicates that information was compiled based on multiple sources.&nbsp;</p> <p>Column C: numeric values correspond to the weight factor, i.e. number of time articles from each journal featured in the SHAP-ID Literature Review bibliography.&nbsp;</p> <p>Column D: SHAPE-ID Zotero bibliography identifier.</p> <p>Column E: Journal title</p> <p>Column F: The country of publication</p> <p>Columns G-AD: ASJC codes (numeric and word values) associated with journal entries.&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Dataset supporting the paper: "A Case Study on the Parametric Occurrence of Multiple Steady States"

<p>Dataset supporting the paper:</p> <p>Russell Bradford, James H. Davenport, Matthew England, Hassan Errami, Vladimir Gerdt, Dima Grigoriev, Charles Hoyt, Marek Košta, Ovidiu Radulescu, Thomas Sturm, and Andreas Weber. A Case Study on the Parametric Occurrence of Multiple Steady States. In Proceedings of ISSAC ’17, Kaiserslautern, Germany, July 25-28, 2017, 8 pages. ACM, 2017.<br> https://doi.org/http://dx.doi.org/10.1145/3087604.3087622<br> Preprint: https://arxiv.org/abs/1704.08997</p> <p>We provide all the accompanying material for the main symbolic computations in Section 2.1</p> <p>###############<br> Section 2.1.1<br> ###############</p> <p>The computations were carried out within the following setting:</p> <p>1. Compiled and functional Reduce system r3606 is available and the<br> environment variable $trunk points to the trunk of Reduce source tree<br> (e.g. ~/reduce-algebra/trunk). To compile and install Reduce see:<br> http://redlog.eu/reduce-wiki/index.php/Installation</p> <p>2. Compiled and functional QEPCAD B v1.69 is installed and it is<br> possible to call QEPCAD from Reduce (via rlqepcad). To compile and<br> install QEPCAD see: https://www.usna.edu/CS/qepcadweb/B/QEPCAD.html</p> <p>To reproduce the experiments reported in the paper go:</p> <p>./generate-problems-onevar.py</p> <p>./compute-all.sh --csl ~/reduce-algebra/trunk ./problems-onevar 4</p> <p>./einsetzen.sh</p> <p>A few notes BEFORE YOU RUN:</p> <p>1. To try a different value of k19, edit the file solbiomod26.red.<br> Currently we set k19 = 500.</p> <p>2. It is advised to run these test on a multicore machine, e.g., we<br> used 32 cores (the last argument of ./compute-all.sh script) for<br> trying the 3^11 candidates (einsetzen.sh script).</p> <p>3. The script generate-problems-assignmnets.py (called from<br> einsetzen.sh) creates 3^11 test files!</p> <p>4. The potential values to be tried using interval refinement are<br> described as real algebraic numbers, i.e., pairs of polynomial and<br> isolating interval. These numbers are stored in assignment files<br> (assignment-&lt;var&gt;-&lt;possibility#&gt;) produced by ./gml-to-red.py (called<br> from einsetzen.sh)</p> <p>###############<br> Section 2.1.2<br> ###############</p> <p>These computation were conducted in commercial Computer Algebra System, Maple 2016.  We have included the Maple Worksheet file which is annotated to describe the calulations in detail.  We also include a pdf printout of the worksheet which those without access to Maple can read.</p>

opencc-by-4.0May 2017View details →
zenodo44/100

Porto Santo landscape features and endemic lichens occurrence data

<p>Landscape features of Porto Santo island of and observation data of endemic lichens belonging to Sparrius et al. 2017, Bryologist.</p>

openmit-licenseJul 2017View details →
zenodo44/100

Vertical Profiles of Convection-Permitting Simulations for Predicting Thunderstorm Occurrence

<p>This repository contains datasets for training and evaluation of the machine learning (ML) models in K. Vahid Yousefnia et al., <em>Inferring Thunderstorm Occurrence from Vertical Profiles of Convection-Permitting Simulations: Physical Insights from a Physical Deep Learning Model</em>, 2024 (submitted to <em>Artificial Intelligence for the Earth Systems,</em> preprint available at https://arxiv.org/abs/2409.20087).</p>

opengpl-3.0-or-laterOct 2024View details →
zenodo44/100

Research Data and Code for "Interdisciplinarity in the 17th Century? A Co-Occurrence Analysis of Early Modern German Dissertation Titles"

<p>This dataset documents results and code for the paper "Interdisciplinarity in the 17th Century? A Co-Occurrence Analysis of Early Modern German Dissertation Titles" by Stefan He&szlig;br&uuml;ggen-Walter, forthcoming in *Synthese*. The data to be processed are contained in four files, derived from a larger dataset related to German dissertations and sourced from the national bibliography of 17th century German prints *VD 17* that will be released at a later date. More information can be found in the file `README.md`.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Water surface occurrence and recurrence from the article "Amazon's 2023 Drought: Sentinel-1 Reveals Extreme Rio Negro River Contraction"

<p>This data package contains the 10 m spatial resolution occurrence and recurrence water surface masks from the article "Amazon's 2023 Drought: Sentinel-1 Reveals Extreme Rio Negro River Contraction" . These maps have been produced with Sentinel-1 images (10 m) and a Deep Learning method for image segmentation called U-net, methods and data are fully described in the article. Water surface occurrence is computed for the period 2022-2023 and indicates the percentage of time that a pixel is classified as water (100%: always water, 0%: never water, and values between 0 and 100 indicate seasonality). Water surface recurrence is computed for the period 2022-2023 and indicates the number of times that a pixel was classified as a water surface, i.e., 35 indicates that the pixel was classified 35 times as a water surface during the 2022-2023 period. The total size of the dataset is 158 Mo and is distributed in two Geotiffs, one for the water surface occurence and one for the water surface. When using this dataset, please cite the original article <a href="https://doi.org/10.3390/rs16061056">https://doi.org/10.3390/rs16061056</a></p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Diverse baleen whale acoustic occurrence around two sub-Antarctic Islands: A tale of residents and visitors

<p>This dataset contains the acoustic .wav file of all exemplar calls illustrated by the spectrograms in the manuscript figure, MS Excel Spreadsheet file with baleen whale call occurrence and environmental data, and the R code used for fitting the RF models. R codes must be run in the following manner:</p> <p>1. 01_tune_occ_enviro_rf_model_balance_baleen_whales</p> <p>2. 02_process_occ_enviro_rf_model_balance_baleen_whales</p> <p>The codes are self-explanatory given the comments contained therein, and the source code for fitting the codes is provided as 000_source_all.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media (01.2016-04.2021)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)

<p><strong>Sources:&nbsp;</strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Plantago patagonica occurrences from the Colorado Plateau, GBIF download 02/18/2022

<p><em>Plantago patagonica</em> occurrences from the Colorado Plateau, GBIF download 02/18/2022.&nbsp; Downloaded with <em>gbif</em> function from <em>dismo</em> package in R.</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record