Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,015
datasets available to search
ShareScore release 0.7.1
Dataset results
3,015 results for “occurrence”
Taxonomy, occurrences, phylogeny, traits and uses of the entire plant genus Scleria (Cyperaceae)
<p>This resource includes several datasets:</p> <p>(1) Taxonomy (261 species): Updated taxonomy of the genus Scleria at the species level based on Bauters et al. 2016 and 2019.</p> <p>(2) Occurrences (latitude/longitude): data was compiled using observations from the Global Biodiversity Information Facility, Red List, research grade identifications from iNaturalist (accessed 13/08/2023) and records for collections from BR, K, GENT, L, MO, NY, P, US and WAG which were georeferenced using Google Earth. The dataset includes 22,759 observations from 248 species. Methodology follows Larridon et al. (2021).</p> <p>(3) Phylogeny of the genus based on three markers (ITS, ndhF, rps16) from Larridon et al. (2021) (136 species).</p> <p>(4) Traits. (i) We measured maximum height, maximum blade length, maximum blade width, stem width, nutlet length and nutlet width from 1,254 specimens of 209 species housed at Royal Botanic Gardens, Kew and the Muséum National d'Histoire Naturelle in Paris. (ii) We also compiled another dataset of 16 continuous and categorical traits for all 261 Scleria species derived from protologues and descriptions from regional floras. (iii) We measured nutlet weight for 141 species.</p> <p>(5) Uses & ecology: ethnobotany (mostly medicinal uses) and references to its ecology in several ecosystems (e.g., pollination, dispersal, ecological role). This data was gathered from several bibliographical sources, also provided.</p> <p>(6) Pictures of nutlets from 141 species.</p>
Examination of the occurrence of drought phenomenon on the basis of the VHI coefficient from NOAA
<p>VHI (Vegetation Health Index) for the period from September 2018 were used for studies related to the occurrence of drought.</p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis
<p><strong>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</strong></p> <p><strong>Sources</strong>: Above 140k articles (01.2016-03.2019):</p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The presented tables include the most extreme co-occurring terms for the analysed social issue. The examples are chosen from the list of words with 30 most positive and 30 most negative sentiment. The presented graphs show the evolution of sentiments for social issues. The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p> <p> </p> <p><strong>Files</strong></p> <p>sentiments_mod11.csv sentiment score based on chosen unigrams</p> <p>sentiments_mod22.csv sentiment score based on chosen bigrams</p> <p>sentiments_cooc_mod11.csv, sentiments_cooc_mod12.csv, sentiments_cooc_mod21.csv, sentiments_cooc_mod22.csv combinations of co-occurrences: unigrams-unigrams, unigrams-bigrams, bigrams-unigrams, bigrams-bigrams</p> <p> </p>
Co-occurrences of trending keywords in popular tech media
<p><strong>Co-occurrences of trending keywords in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. 'metoo', 'gdpr')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p><strong>Files</strong></p> <p>unigram-unigram co-occurrences: cooc11weighted.csv</p> <p>unigram-bigram co-occurrences: cooc12weighted.csv</p> <p>bigram-unigram co-occurrences: cooc21weighted.csv</p> <p>bigram-bigram co-occurrences: cooc22weighted.csv<br> </p> <p> </p>
Global river density, seasonal and surface water occurrence and upstream area at 250 m in the Goode Homolosine projection
<p>Several layers describing density of surface water / streams projected to the <a href="https://en.wikipedia.org/wiki/Goode_homolosine_projection">Good Homolosine projection</a>. List of layers included:</p> <ul> <li>hyd_log1p.upstream.area_merit.hydro_m = Upstream Drainage Area based on the <a href="http://hydro.iis.u-tokyo.ac.jp/~yamadai/MERIT_Hydro">MERIT Hydro</a>,</li> <li>hyd_river.density_gloric_p = rasterized <a href="https://www.hydrosheds.org/page/gloric">Global River Classification (GLORIC)</a> DB,</li> <li>lcv_water.occurance_jrc.surfacewater_p = Surface Water based on the JRC's <a href="https://global-surface-water.appspot.com/">Global Surface Water</a>,</li> <li>lcv_water.seasonal_probav.glc.lc100_p = Seasonal Inland Water probability based on the <a href="https://lcviewer.vito.be/">Copernicus LC100 map</a>,</li> <li>lcv_wetlands.cw_upmc.wtd_c = composite wetland (CW) map based on <a href="https://doi.org/10.1594/PANGAEA.892657">Tootchi et al. (2019)</a>,</li> <li>Goode_Homolosine_domain_250m.tif = map domain prepared by <a href="https://doi.org/10.5281/zenodo.1475152">Luís de Sousa</a>,</li> <li>tiles_GH_100km_land.gpkg = 100 km x 100 km tiling system covering the land mass,</li> </ul> <p>Important notes: Processing steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/input_layers/WaterDensity">here</a></strong>. Antartica is not included. Reprojecting maps to Goode Homolosine projection can be cumbersome and small amount of artifacts at the edges of the map can be anticipated.</p> <p>These maps were develop in connection to the <a href="http://www.OpenLandMap.org">OpenLandMap.org</a> initiative.</p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>hyd = theme: hydrology and water dynamics,</li> <li>log1p.upstream.area = variable: log(X+1)*10 of the upstream area,</li> <li>merit.hydro = determination method: MERIT Hydro,</li> <li>m = mean value,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>b0..0cm = vertical reference: surface,</li> <li>2017 = time reference: period 2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>
Historical distribution and current drivers of guppy occurrence in Brazil
<p>This data set was used in the paper "Historical distribution and current drivers of guppy occurrence in Brazil". The file consist of occurrence records of <em>Poecilia reticulata</em> in Brazil over different time periods. In order to evaluate the historical and current distribution of <em>P. reticulata,</em> we searched for occurrence records in two major data sources. We first compiled data from all the studies cited in the recent comprehensive review of studies of Brazilian stream fish assemblages (Dias et al., 2016). We performed this search by using electronic databases and search engines (i.e., Web of Knowledge, Google Scholar, and Scielo) to look for primary studies and published papers from Brazilian journals that provide occurrence records of <em>P. reticulata</em> in Brazil. We also used combinations of the following search terms, in English and Portuguese: “guppy fish”, “non-native”, “nonindigenous aquatic species”, and “<em>Poecilia reticulata”; </em>this second search provided additional published papers mentioning this species in Brazilian territory. Both the papers and supplementary information were screened in order to find the geographical coordinates of the sampling points where this species has been detected. As a third data source, we used all the available records of <em>P. reticulata</em> from the SpeciesLink website (http://splink.cria.org.br/, accessed 2016), which aggregates species occurrence data from major biological collections worldwide, including those from Brazilian institutions. We extracted from this platform all the associated informations (e.g., the location and the associated geographical coordinates; the sampling dates containing year, month, and day; and the names of researchers who composed the sampling teams). From these three sources, we created a database of the occurrence of <em>P. reticulata</em> in Brazil.</p> <p>Some records were excluded because the geographical coordinates and/or sampling dates were not provided in detail, or could not be determined directly or from information in the publication itself (Dias et al., 2016). Overall, incomplete and discarded records comprised only 0.7% (12 out of 1649 records) of our dataset. We further used the sampling dates and associated collector information to remove duplicate records from the database. By this means, multiple records of <em>P. reticulata</em> with the identical geographical coordinates were compared in terms of the sampling day, month, year, and collector(s). If all the information was identical, only one record was included. On the other hand, multiple records of <em>P. reticulata</em> from the same location but with different dates and collectors were retained in the final database. This final database was composed of 1402 records and was used to investigate the occurrence of <em>P. reticulata</em> over time.</p> <p>References</p> <p>Dias, M. S., J. Zuanon, T. B. A. Couto, M. Carvalho, L. N. Carvalho, H. M. V. Espírito-Santo, R. Frederico, R. P. Leitão, A. F. Mortati, T. H. S. Pires, G. Torrente-Vilara, J. do Vale, M. B. dos Anjos, F. P. Mendonça, & P. A. Tedesco, 2016. Trends in studies of Brazilian stream fish assemblages. Natureza & Conservação 14: 106–111.</p>
Supplementary Data: Fungal biostarter and bacterial occurrence of dry-aged beef: the sensory quality and volatile aroma compounds after 21 days of aging
<p>This dataset contains data generated during realisation of the project Tango-IV-C/0005/2019: Biostarters development for dry aged beef production, funded by National Centre for Research and Development (Poland). These data were used to prepare paper entitled "Fungal biostarter and bacterial occurrence of dry-aged beef: the sensory quality and volatile aroma compounds after 21 days of aging" by Wiesław Przybylski , Danuta Jaworska, Paweł Kresa, Grzegorz Michał Ostrowski, Magdalena Płecha, Dorota Korsak, Dorota Derewiaka, Lech Adamczak, Urszula Siekierko, Julia Pawłowska.</p>
Mammal occurrence records (2024) in the Valparai Plateau and Anamalai Tiger Reserve, Western Ghats, India
<p>This dataset contains Mammal occurrence records (November 2023 - October 2024) in the Valparai Plateau and Anamalai Tiger Reserve, Western Ghats, India. It includes a few occurrence records from other parts of southern India. Occurrence records were gathered in the field by researchers of the Nature Conservation Foundation, India, using a mobile data collection application (EpiCollect5). Suggested citation is:<br>Nature Conservation Foundation (2024). Mammal occurrence records (2024) in the Valparai Plateau and Anamalai Tiger Reserve, Western Ghats, India. Nature Conservation Foundation, India. Dataset, Zenodo. DOI: 10.5281/zenodo.13910696<br> <br><strong>CONTACT #1</strong><br>1. Name: T. R. Shankar Raman <br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Work Phone: +91 821 2515601<br>4. Email address: trsr@ncf-india.org <br>5. ORCID: https://orcid.org/0000-0002-1347-3953</p> <p><strong>CONTACT #2</strong><br>1. Name: Divya Mudappa <br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Work Phone: +91 821 2515601<br>4. Email address: divya@ncf-india.org <br>5. ORCID: https://orcid.org/0000-0001-9708-4826</p> <p><strong>Keywords: </strong>tropical rainforest, plantations, Anamalai Hills, Western Ghats, animal distribution, mammals </p> <p><br><strong>Geographic Coverage:</strong><br>1. Location/Study Area: Valparai Plateau, Tamil Nadu, India; Anamalai Tiger Reserve, Tamil Nadu, India<br>2. GPS coordinates: Valparai Plateau (10°15'- 10°22'N, 76°52' - 76°59'E); Anamalai Tiger Reserve (10°12' - 10°35'N, 76°49' - 77°24'E)</p> <p><strong>Temporal Coverage:</strong><br>1. Begins: 2023-11-01 (Year, Month, Day)<br>2. Ends: 2024-10-01 (Year, Month, Day)</p> <p>Besides the 00_readMe.txt file containing this information, the dataset includes 23 images (photographs) and two comma-delimited text (csv) files as explained below:<br><strong>1) 01_anamalai-mammals-2024.csv </strong>-- This file has the main mammal occurrence data with relevant and renamed columns derived from the original downloaded csv file from the EpiCollect5 application website.</p> <p><strong>2) 02_nameMatch.csv</strong> -- This file matches the vernacular name as originally recorded with the correct common name and scientific name</p> <p>+23 image files (with ".jpg" file extension)</p> <p><strong>FILES INCLUDED IN DATASET</strong></p> <p><strong>01_anamalai-mammals-2024.csv</strong><br>This file has the main mammal occurrence data with relevant and renamed columns derived from the original downloaded csv file from the EpiCollect5 application website.<br>ec5_uuid: Unique ID for each observation<br>created_at: Automatic time stamp of date and time when record was created on the mobile app<br>uploaded_at: Automatic time stamp of date and time when record was uploaded using the mobile app<br>recordedBy: Name of observer<br>title: Title assigned to each record (composite of date, species, and type of observation)<br>lat_gps: Latitude in decimal degrees N<br>long_gps: Longitude in decimal degrees E<br>accuracy_gps: Horizontal accuracy of GPS location in metres<br>UTM_Northing_gps: Latitude in UTM<br>UTM_Easting_gps: Longitude in UTM<br>UTM_Zone_gps: UTM Zone<br>eventDate: Date in ISO format (yyyy-mm-dd)<br>verbatimEventDate: Date in format originally recorded (dd/mm/yyyy)<br>eventTime: Time of observation<br>vernacularName: Species common name as initially recorded<br>individualCount: Number of individuals observed<br>occurrenceRemarks: type of observation<br>habitat: Habitat type<br>photo: Filename of photo if available (NA otherwise)<br>eventRemarks: Notes or remarks about the observation</p> <p><strong>02_nameMatch.csv</strong><br>This file matches the name as originally recorded with the correct common name and scientific name.<br>vernacularName: Common or English name as initially recorded <br>scientificName: Scientific name of the species</p> <p>+23 image files (.jpg extension)</p>
Spatial clustering of Neobuccinum eatoni occurrence data for potential distribution modeling
<p>The occurrence dataset for <em>Neobuccinum eatoni</em> was compiled through filtration process, starting with records from the Global Biodiversity Information Facility (GBIF) and supplemented by museum specimens and additional sources like SOMBASE, iBOL, NIWA, ANTABIF, and SCAR-AntOBIS. Further data were sourced from the National Museum of Natural History in Paris, the University of Vigo, and recent fieldwork in Antarctica, Heard Island, and Kerguelen Island. Records were meticulously screened to remove misidentified specimens, inaccurate locations, duplicates, and outdated entries, ensuring accuracy and relevance. To address spatial autocorrelation, clustering methods divided the data into distinct geographic clusters, producing a refined dataset used to model <em>N. eatoni</em>'s potential distribution with enhanced predictive reliability by reducing spatial autocorrelation effects.</p>
CT Fishing Report Occurrences
<p>Dataset of gamefish occurrences as compiled from the Connecticut Fishing Report (2006-2018) and Trophy Fish Report (2009-2018), both published by the Connecticut Department of Energy and Environmental Protection. Compiled as thesis project by Rebecca Hedreen for a Masters of Science in Biology from Southern Connecticut State University, with advisor Dr. Sean Grace.</p>
Global Biodiversity Information Facility (GBIF): an exhaustive list of gbif record ids, dataset keys, and their associated Occurrence IDs, Institution Code, Collection Codes and Catalog Numbers. hash://sha256/ea88f03a7bfd1ba853fdbea3203d54ab81ac3cdc8e8da7c96bbbba9c4b05d933 hash://md5/c49fe34785354847b37ea4509261e130
<p>The Global Biodiversity Information Facility (GBIF) indexes thousands of biodiversity datasets from Natural History Collections, citizen science initiatives (e.g., iNaturalist, eBird), and other sources. As part of the index process, GBIF associates at least two identifiers with indexed records: a record id (aka gbifID) and a dataset id (aka dataset key). These ids are central to do lookup, reference data, and package interpreted data products.</p> <p>This publication contains an exhaustive list of GBIF IDs and ids associated by their data providers as derived from:</p> <p>GBIF.org (01 March 2023) GBIF Occurrence Download https://doi.org/10.15468/dl.pk3trq</p> <p>The resource (size: ~260GB) provided by GBIF had content id hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 and was used to generate the resource included in this publication using</p> <pre><code class="language-bash">preston cat 'zip:hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97!/0015281-230224095556074.csv'\ | cut -f 1,2,3,37,38,39\ | gzip\ > gbifid.tsv.gz </code></pre> <p>with the content id of gbifid.tsv.gz (size: ~35GB) being hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8 .</p> <p>the first 10 lines of gbifid.tsv.gz as extracted via</p> <pre><code>preston cat --remote https://zenodo.org/record/7789866/files,https://linker.bio hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8\ | gunzip\ | head</code></pre> <p>are:</p> <pre><code>gbifID datasetKey occurrenceID institutionCode collectionCode catalogNumber 2997162320 c71c8000-9fc7-422c-804a-ce6abe751771 3399442 CEPEC CEPEC CEPEC00109669 2997162309 c71c8000-9fc7-422c-804a-ce6abe751771 2733085 CEPEC CEPEC CEPEC00000818 2997162317 c71c8000-9fc7-422c-804a-ce6abe751771 2733086 CEPEC CEPEC CEPEC00000888 2997162313 c71c8000-9fc7-422c-804a-ce6abe751771 3399443 CEPEC CEPEC CEPEC00109744 2997162306 c71c8000-9fc7-422c-804a-ce6abe751771 2733087 CEPEC CEPEC CEPEC00000889 2997162316 c71c8000-9fc7-422c-804a-ce6abe751771 3399440 CEPEC CEPEC CEPEC00109605 2997162324 c71c8000-9fc7-422c-804a-ce6abe751771 2733088 CEPEC CEPEC CEPEC00000890 2997162308 c71c8000-9fc7-422c-804a-ce6abe751771 3399441 CEPEC CEPEC CEPEC00109615 2997162303 c71c8000-9fc7-422c-804a-ce6abe751771 2733089 CEPEC CEPEC CEPEC00000891</code></pre> <p>Note that at time of writing, the html resource associated with the occurrence id 2997162320, and data set key c71c8000-9fc7-422c-804a-ce6abe751771 (extracted from of the first data row example above) are available via:</p> <p>https://gbif.org/occurrence/2997162320</p> <p>and</p> <p>https://gbif.org/dataset/c71c8000-9fc7-422c-804a-ce6abe751771</p> <p>respectively.</p> <p>This resource was initially created to help integrate with Bionomia (https://bionomia.net) to help associate people identifiers provided by bionomia to their original records via their GBIF ids. Bionomia re-uses GBIF records ids as a way to define links between records and the people (e.g., curators, collectors, identifiers) that worked on them. </p> <p>In other words, this resource provides a versioned translation table from the GBIF data universe (as defined by GBIF record ids, and dataset keys) to the data collections that exist (and evolve) independent of it. </p> <p>Note that the resource identified by hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 was not included in this publication it was too big (260GB) to fit. You may be able to retrieve the resource from its original location at https://api.gbif.org/v1/occurrence/download/request/0015281-230224095556074.zip .</p>
FLEXPART 10.4 output for "Occurrence and backtracking of microplastic mass loads including tire wear particles in Northern Atlantic air"
<p>The dataset consists of three (3) files:</p> <p>-- track3h.txt shows the position of the research vessel in each of the seven (7) ship tracks/campaigns in 3-hour resolution in ascii format structured in columns as follows:</p> <p>YEAR, MONTH, DAY, HOUR, MINUTE, SECOND, LONGITUDE, LATITUDE, SHIP TRACK NUMBER</p> <p>-- FLEXPART_720x360_fine.tar.gz shows the footprint emission sensitivities for fine particles (as described in the paper) in a gridded netCDF format of 0.5 degrees resolution for 80 release points matching the coordinates and times in the track3h.txt file.</p> <p>-- FLEXPART_720x360_coarse.tar.gz shows the footprint emission sensitivities for coarse particles (as described in the paper) in a gridded netCDF format of 0.5 degrees resolution for 80 release points matching the coordinates and times in the track3h.txt file.</p>
Native North American Silene (L.) Occurrences Filtered from GBIF
<p>This is a dataset including all Native <em>Silene</em> species accepted in taxonomic nomenclature and considered to inhabit the North American range. Data was downloaded using rgbif::occ_download and accessed from R via rgbif (https://github.com/ropensci/rgbif) on 2023-03-29. The original unfiltered GBIF occurrences can be download at https://doi.org/10.15468/dl.g89y2y, and https://api.gbif.org/v1/occurrence/download/request/0128277-230224095556074.zip. The data is filtered to have coordinates in North America, no geospatial issues, no spatial duplicates filtered to the infraspecific epithet, no coordinate uncertainty greater than 100000 meters, and no occurrences lying within 1km of a college or university. This dataset is incomplete as it does not include ALL observations that occur in North American countries as observations lacking a continent field of "north_america" in GBIF are not included.</p>
Historical Occurrence of Antarctic Icebergs within Mercantile Shipping Routes and the Exceptional Events of the 1890s
<p>This is the dataset created for the Journal of Glaciology paper "<i>Historical Occurrence of Antarctic Icebergs within Mercantile Shipping Routes and the Exceptional Events of the 1890s</i>" by Robert Headland, Nick Hughes and Jeremy Wilkinson (<a href="https://doi.org/10.1017/jog.2023.80">doi:10.1017/jog.2023.80</a>). We have endeavoured to make the data as accessible as possible by providing it in a range of formats.</p><p>Please see the README.pdf for a detailed description of the files, and the paper for the dataset. Version 1.1 contains additional reports from newspaper archives.</p>
Occurrence records used to develop a climatic suitability model for emerald ash borer in DDRP
<p>Presence records used to calibrate and validate a climatic suitability model for emerald ash borer in the DDRP platform (Degree-Days, Risk, and Phenological event mapping) (Barker et al. 2023). The first sheet ("Records") of the Excel file provides the range (native or invaded), continent, country, state or province, locality, latitude, and longitude of origin for each record. The "Coords_est" column indicates whether the coordinates were estimated from city- or county-level information (1 = yes, 0 = no). The year in which the record was collected is provided if known. The second sheet of the Excel file ("References") provides a list of references for each record source.</p>
UTHSC Publication Research Category (ANZSRC 2020) Co-Occurrence
<p>This data visualization is a bibliometric analysis of University of Tennessee Health Science Center publications for the years 2018-2020. It was created for senior University leadership for the purposes of strategic planning and identifying research areas of strength.</p> <p>Bibliographic data was supplied by Dimensions by Digital Science. The chord graph was created in Tableau and demonstrates relationship pairs of Fields of Research (ANZSRC 2020) categories. Each publication record may be associated with 1+ categories. The graph highlights the frequency of category pairings within a single publication record. I.e. publications categorized as "Immunology" are most commonly also categorized with "Medical Microbiology," indicating overlap in this area of research.</p> <p>This graph was created using instructions from Marc Reid's datavis.blog entry "Creating a Chord Diagram with Tableau Prep and Desktop" (<a href="https://datavis.blog/2020/07/02/creating-chord-diagram-in-tableau/">https://datavis.blog/2020/07/02/creating-chord-diagram-in-tableau/</a>).</p> <p>The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.</p>
Exoplanet Occurrence Rates Plot
<p>Everyone knows the obligatory semi-major axis vs mass plot that every exoplanet scientist is required by law to show in every talk. We wondered how literature occurrence rates would look overlaid on this plot. So we made it. </p> <p>The script for generating this plot is located at <a href="https://github.com/logan-pearce/occurrence-rate-plot">https://github.com/logan-pearce/occurrence-rate-plot</a>, so you can adjust the colors and the plot to your heart's content.</p>
Role of fluid on earthquake occurrence: Example of the 2019 Ridgecrest and the 1997, 2009 and 2016 Central Apennines sequences
<p>This repository contains files needed to reproduce the b-value times series and stress change modeling related to the Central Apennines and Ridgecrest earthquake sequences (paper under revision, preprint available at <a href="https://doi.org/10.31223/X5MH1J">https://doi.org/10.31223/X5MH1J</a>). </p>
Blair et al. 2020: Machine learning identification of ground beetles (repackaging of occurrences published by the NEON Biorepository Data Portal)
Blair, J.; Weiser, M. D.; Kaspari, M.; Miller, M.; Siler, C.; Marshall, K. E. 2020. Robust and simplified machine learning identification of pitfall trap-collected ground beetles at the continental scale. Ecology and Evolution 10 (23): 13143-13153. https://doi.org/10.1002/ece3.6905 Additional NEON samples (not yet archived at the Biorepository) were used in this research: full list of occurrences used.
Stachewicz et al. 2021: Trait correlation, phylogenetic signal in Carabidae morphology (repackaging of occurrences published by the NEON Biorepository Data Portal)
Stachewicz JD, Fountain-Jones NM, Koontz A, Woolf H, Pearse WD, Gallinat AS. 2021. Strong trait correlation and phylogenetic signal in North American ground beetle (Carabidae) morphology bioRxiv 02.12.431029; doi: https://doi.org/10.1101/2021.02.12.431029 Many NEON samples and specimens used in this work resulted from NEON prototype data and will not be archived in the Biorepository. See the appendices in the above linked article for a full list of NEON samples and specimens and their associated collection data. Additionally, see appendices of above linked article for specimen-level morphological trait measurements and genetic sequence data.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.