Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22,922
datasets available to search
ShareScore release 0.7.1
Dataset results
22,922 results for “data collection”
Interagency Ecological Program: Over four decades of juvenile fish monitoring data from the San Francisco Estuary, collected by the Delta Juvenile Fish Monitoring Program, 1976-2025
The United States Fish and Wildlife Service Delta Juvenile Fish Monitoring Program (DJFMP) has monitored juvenile Chinook Salmon Oncorhynchus tshawytscha and other fish species within the San Francisco Estuary (Estuary) since 1976 using a combination of surface trawls and beach seines. Since 2000, three trawl sites and 58 beach seine sites have been sampled weekly or biweekly within the Estuary and lower Sacramento and San Joaquin Rivers. As part of the Interagency Ecological Program (IEP) that manages the Estuary, the DJFMP has tracked the relative abundance and distribution of naturally and hatchery produced juvenile Chinook Salmon of all races as they outmigrate through the Sacramento-San Joaquin Delta for over four decades. The data that DJFMP collected has been used not only to help inform the management of Chinook Salmon, but also to monitor the status of native species of interest such as the previously listed Sacramento Splittail Pogonichthys macrolepidotus and invasive species such as Mississippi Silverside Menidia audens and Largemouth Bass Micropterus salmoides. DATA CORRECTION/UPDATE: Previous data versions 244.6, 244.7, and 244.8 contained an error and resulted in duplicated records of hatchery Chinook Salmon in the datasets. These datasets were removed from the data repository and the error was corrected in version 244.9 and after. DNA Data: Full more details of the DNA methods and results for juvenile Chinook salmon included in this dataset, please check out Blankenship, S.M., J. Israel, E. Buttermore, and K. Reece. 2021. Knights Landing, California Department of Fish and Wildlife, Genetic Determination of Population of Origin 2017 through 2019 ver 1. Environmental Data Initiative. https://doi.org/10.6073/pasta/85fbc988c0b1362e84c318e69c7a939e. For more information on the Delta Juvenile Fish Monitoring Program: https://www.fws.gov/project/delta-juvenile-fish-monitoring-program
Interagency Ecological Program: Zooplankton catch and water quality data from the Sacramento River floodplain and tidal slough, collected by the Yolo Bypass Fish Monitoring Program, 1998-2018
Largely supported by the Interagency Ecological Program (IEP), the California Department of Water Resources (DWR) has operated a fisheries and invertebrate monitoring program in the Yolo Bypass since 1998. The main objectives of the Yolo Bypass Fish Monitoring Program (YBFMP) are to collect baseline data on lower trophic levels (phytoplankton, zooplankton and insect drift), juvenile and adult fish, hydrology, and water quality parameters. As the Yolo Bypass has been identified as a high restoration priority by numerous regulatory agencies, these baseline data are critical for evaluating success of future restoration projects. In addition, the data have already served to increase our understanding of the role of the Yolo Bypass in the life history of native fishes, and its ecological function in the San Francisco Estuary. Zooplankton are an important component in the diet of larval, juvenile, and small adult fishes within the San Francisco Estuary, including Delta Smelt, juvenile Chinook Salmon, Striped Bass, and Sacramento Splittail. The YBFMP collects zooplankton year-round from two sites. Since 2011, samples have been collected biweekly (every other week) to weekly (during floodplain inundation) using 150- and 50- micrometer mesh plankton nets. Zooplankton are identified and enumerated by contractors (currently BSA Environmental Services). The goals of the zooplankton monitoring program are to compare the seasonal variation in species densities and trends between (1) the Sacramento River channel, and (2) the Yolo Bypass, the river’s seasonal floodplain. Data on zooplankton catch and associated water quality parameters are presented in this dataset.
Toolik Lake Inlet discharge data collected during summers of 2010 to 2018, Arctic LTER, Toolik Research Station, Alaska.
Stream discharge, temperature, and conductivity of Toolik Lake Inlet stream for 2010 - 2018 study season. Water level was recorded with a Stevens PGIII Pulse Generator and Conductivity (EC) and Temperature measured with a Campbell Scientific Model 247 Conductivity and Temperature probe.
Soil temperature data collected from the Arctic LTER wet sedge experimental site Toolik Field Station North Slope, Alaska from 1994 to 2020
Soil temperature data collected every 4 hours from a wet sedge site at the Arctic Tundra LTER site at Toolik Lake. Temperatures are measured every 3 minutes and averaged every 4 hours in control, nitrogen alone, phosphorus alone, nitrogen and phosphorus, and greenhouse experimental plots soil temperatures.
Meteorological data collected on Lake E5 during the ice free season since 2000 to present, Arctic LTER, Toolik Research Station, Alaska.
Yearly file describing the metological data on Lake E5 (Lake E5 Climate station) near the Toolik Field Research Station (68 38'N, 149 36'W). Measurements include air temperature, relative humidity, wind direction, and wind speed..
Vegetation Data Collected with Point Frame for 83 Locations of 6-163 Years Old Black Spruce, Alaska Paper Birch, and Aspen Stands Across Interior Alaska. Sampled in 2008-2010 and 2013-2015.
This dataset contains point frame data for vegetation less than 1.3 m, including vascular plants, bryophytes, lichens, leaf litter and bare ground, as well as species codes used, as described in Jean et al. 2017 Canadian Journal of Forest Research. Samples of all encountered unknown species were collected for identification in the lab. Bryophyte nomenclature followed Anderson et al. (1990).
Surface Water Quality Monitoring Data collected in South Florida Coastal Waters (FCE LTER), Florida, USA, June 1989-ongoing
The Southeast Environmental Research Center at Florida International University operates a network of 331 fixed sampling sites distributed throughout the estuarine and coastal ecosystems of south Florida. The purpose of this network is to address concerns in regional water quality which cross and overlap separate political boundaries. Funding has come from different sources with individual programs being added as funding became available. Biscayne Bay, Florida Bay, Whitewater Bay, Ten Thousand Islands, Rookery Bay, Estero Bay, and Pine Island Sound are sampled monthly while the Florida Keys National Marine Sanctuary (FKNMS) and the southwest shelf are sampled quarterly. Variables currently being measured include surface and bottom temperature, salinity, dissolved oxygen, nitrate, nitrite, ammonium, total nitrogen, total organic nitrogen, total phosphorus, soluble reactive phosphorus, total organic carbon, total silicate, chlorophyll a, alkaline phosphatase activity, turbidity, and light extinction. The purpose of this network is to address concerns in regional water quality which cross and overlap separate political boundaries. One of the products is a quasi-synoptic big picture of nutrient and phytoplankton biomass distributions over the South Florida Coastal Waters. The SERC network will, in time, provide us with the data necessary to determine whether conditions within the estuaries and sanctuary are improving or declining.
Fish and consumer data collected from Northeast Shark Slough, Everglades National Park (FCE) from September 2006 to September 2008
Three 1m2 throwtrap sites were selected randomly within a few meters of the location. All fish and macroinvertebrates were sampled within each trap using a seine; the trap is determined to be clear of fish and macroinvertebrates once three empty seines are recovered. The count is reset whenever a fish or invertebrate is caught. Once the trap is determined to be clear, five sweeps are done using a couple of nets. The count resets whenever fish is caught, while an extra sweep is added whenever an invertebrate is captured. The trap is finally cleared once the fifth sweep doesn't capture any organisms. All captured organisms are anesthesized in a plastic cup filled with water and MS222; they are then preserved using a 10% formaldehyde solution for identification and archiving in the lab
Periphyton and Associated Environmental Data Relative from Samples Collected from the Greater Everglades, Florida, USA from September 2005 to November 2014
This data package contains peripihyton and environmental data collected annually during the wet season between 2005 and 2014 from sites distributed throughout the greater Everglades ecosystem. This project is part of the Comprehensive Everglades Restoration Program's Monitoring and Assessment Plan intended to document baseline variability in periphyton attributes for assessing the effectiveness of restoration projects. A total of 200 primary sampling units (PSU) of 800 m x 800 m are nested in 32 landscape units and each year, random coordinates are 'drawn' within each PSU and one sampleable draw is visited in each. Sampled periphyton is processed for diatoms, slides are prepared, and 500 frustules are enumerated and identified to the lowest possible taxonomic resolution per slide. Taxon abundances are then relativized to the total count. These data accompany environmental, periphyton biomass, and soft algal abundance datasets.
Relative Abundance Diatom Data from Periphyton Samples Collected from the Greater Everglades, Florida USA from September 2005 to November 2014
This data package contains relative diatom taxon abundances collected annually during the wet season between 2005 and 2014 from sites distributed throughout the greater Everglades ecosystem. This project is part of the Comprehensive Everglades Restoration Program's Monitoring and Assessment Plan intended to document baseline variability in periphyton attributes for assessing the effectiveness of restoration projects. A total of 200 primary sampling units (PSU) of 800 m x 800 m are nested in 32 landscape units and each year, random coordinates are 'drawn' within each PSU and one sampleable draw is visited in each. Sampled periphyton is processed for diatoms, slides are prepared, and 500 frustules are enumerated and identified to the lowest possible taxonomic resolution per slide. Taxon abundances are then relativized to the total count. These data accompany environmental, periphyton biomass, and soft algal abundance datasets. Post-2014 data are available upon request to the project PI, Evelyn Gaiser.
Vegetation data collected from Northeast Shark River Slough, Everglades National Park, Florida, USA, September 2006 - April 2025
This project was established in 2006 to document the pattern of abundance of key ecological indicators (e.g., surface water, soil, floc, periphyton and sawgrass) across the NESRS landscape. A total of 30 sites were established and monitored in 2006, 2007 and 2008. After the completion of 1-mile bridge in 2012, additional 10 new sites were established to observe the ecological impact of 1-mile bridge (known as Bridge & Census sites). In 2015, additional 40 sites were established along eight transects (T1-T8, known as near canal sites) in ENP marshes starting at, and roughly perpendicular to the L-29 canal. The purpose of these sites was to monitor the potential effects of Modified Water Deliveries (MWD) operations on changing nutrient concentrations and ratios in key ecological compartments due to increased downstream discharges from the L-29 canal beneath the 1-mile and 2.6-mile bridges and culverts along Tamiami Trail. Data collection is complete.
Long-term climate indices (SPEI and scPDSI) derived from monthly meteorology data collected at USHCN stations in the northern Chihuahuan Desert of the United States, 1911-2021
Drought indices — Standardized Precipitation Evapotranspiration Index (SPEI) and the self-calibrating Palmer Drought Severity Index (scPDSI) —where derived from 9 United States Historical Climate Network (USHCN) stations on the Chihuahuan Desert in North America for this dataset. USHCN is a subset of the NOAA Cooperative Observer Program (COOP) Network, which consists of selected sites based on spatial coverages and completeness of data. Monthly precipitation depths, minimum, maximum and mean temperature were pulled from the dataset. These drought indices were derived using the SPEI package and scPDSI packages in R. Potential evapotranspiration was also calculated in R using the Thornthwaite method. All 9 sites are within the bounds of the Chihuahuan Desert in the state of New Mexico, with a single site (EL PASO) in the state of Texas.
Nekton individual data from flume net collections along Rowley River tidal creeks associated with long term fertilization experiments, Rowley, MA.
The flume nets were deployed with the purpose of capturing salt marsh nekton. Nekton species were identified to the lowest taxonomic level using species keys. The TIDE project aims to simulate eutrophication on a large scale by the addition of NO3- aiming to reach 70μM concentrations from May to September every year during the growing season. This fertilization of the marsh has been going on at Sweeney Creek since the 2004 growing season through 2016 and at Clubhead Creek in 2005 and from 2009 till 2016. Years 2017-2020 are enrichment recovery years.
Models for "A data-driven approach to studying changing vocabularies in historical newspaper collections"
<p>NOTE: This is a badly rendered version of the README within the archive.</p> <p><strong>A data-driven approach to studying changing vocabularies in historical newspaper collections</strong></p> <p>Simon Hengchen,* Ruben Ros,** Jani Marjanen,*** Mikko Tolonen***</p> <p>*<a href="https://spraakbanken.gu.se/en/about/staff/simon">Språkbanken Text</a>, University of Gothenburg, Sweden and <a href="https://iguanodon.ai">iguanodon.ai</a>, Belgium: firstname.lastname@gu.se<br> **<a href="https://www.c2dh.uni.lu/people/ruben-ros">Centre for Contemporary and Digital History (C2DH)</a>, University of Luxembourg: firstname.lastname@uni.lu<br> ***<a href="https://www.helsinki.fi/en/researchgroups/computational-history">COMHIS</a>, University of Helsinki: <a href="mailto:firstname.lastname@helsinki.fi">firstname.lastname@helsinki.fi</a>;</p> <p>These are the supplementary materials for the DH2019 paper <em>A data-driven approach to the changing vocabulary of the ‘nation’ in English, Dutch, Swedish and Finnish newspapers, 1750-1950</em>, as well as the 2021 Digital Scholarship in the Humanities publication available in OpenAccess: <a href="https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793">https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793</a>. If you end up using whole or parts of this resource, please use the following citation(s):</p> <ul> <li>Hengchen, S., Ros, R., and Marjanen, J. (2019). A data-driven approach to the changing vocabulary of the 'nation' in English, Dutch, Swedish and Finnish newspapers, 1750-1950. In <em>Proceedings of the Digital Humanities (DH) conference 2019, Utrecht, The Netherlands</em></li> </ul> <p>and/or:</p> <ul> <li>Hengchen, S., Ros, R., Marjanen, J. and Tolonen, M., 2021. A data-driven approach to studying changing vocabularies in historical newspaper collections. Digital Scholarship in the Humanities, 36(Supplement_2), pp.ii109-ii126.</li> </ul> <p>or alternatively use one of the following <code>bib</code>s:</p> <pre><code>@inproceedings{hengchen2019nation, title="A data-driven approach to the changing vocabulary of the 'nation' in {E}nglish, {D}utch, {S}wedish and {F}innish newspapers, 1750-1950.", author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani}, year={2019}, address = "Utrecht, The Netherlands", booktitle={Proceedings of the Digital Humanities (DH) conference 2019} }</code></pre> <pre><code>@article{hengchen2021data, title={A data-driven approach to studying changing vocabularies in historical newspaper collections}, author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani and Tolonen, Mikko}, journal={Digital Scholarship in the Humanities}, volume={36}, number={Supplement\_2}, pages={ii109--ii126}, year={2021}, publisher={Oxford University Press} }</code></pre> <p> </p> <p>Files</p> <p>This archive contains two folders -- one per diachronic representation method -- as well as this README. The folders each contain four folders, which contain the models for their respective languages. As can be inferred from the small datasize, most of the earlier models are not reliable and should not be used, but are still made available. This work is licensed under a <a href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p><strong>Source material</strong></p> <p>Finnish:</p> <p>The models were created with data from the Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland (National Library of Finland, 2011). We used everything in the corpus.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h fi* 12M fi_1820_SGNS_corpus_file.gensim 89M fi_1840_SGNS_corpus_file.gensim 797M fi_1860_SGNS_corpus_file.gensim 7.0G fi_1880_SGNS_corpus_file.gensim 22G fi_1900_SGNS_corpus_file.gensim</code></pre> <p>Swedish:</p> <p>The models were created with data from the Kubhist 2 corpus (Språkbanken) -- more precisely, the data dumps available at <a href="https://spraakbanken.gu.se/lb/resurser/meningsmangder/">https://spraakbanken.gu.se</a>. After a manual evaluation of Swedish embeddings trained without pre-processing seemed to show that the embeddings were of low quality, we retrained models, only keeping sentences that were at least 10 tokens long and were constituted of at least 50% of lemmas as per the KORP processing pipeline (Borin et al, 2012).</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h sv* 1.6M sv_1740_SGNS_corpus_file.gensim 44M sv_1760_SGNS_corpus_file.gensim 124M sv_1780_SGNS_corpus_file.gensim 228M sv_1800_SGNS_corpus_file.gensim 678M sv_1820_SGNS_corpus_file.gensim 1.6G sv_1840_SGNS_corpus_file.gensim 4.5G sv_1860_SGNS_corpus_file.gensim 6.5G sv_1880_SGNS_corpus_file.gensim 113M sv_1900_SGNS_corpus_file.gensim</code></pre> <p>Dutch:</p> <p>The models were created with data from the Delpher newspaper archive (Royal Dutch Library, 2017), through data dumps for newspapers until and including 1876, and through API hits for articles from 1877 to 1899 (included).</p> <ul> <li>For anything pre-1877 we discarded full texts that had, in the metadata, anything else than exclusively <code>nl</code> or <code>NL</code> as a language tag.</li> <li>For the full texts between 1877 and 1899: we queried the API for all items in the “artikel” category that contained the determiner <code>de</code>.</li> </ul> <p>Our assumption was that most articles should contain <code>de</code> at least once, and those that didn't were too short to be deemed interesting. A subsequent study showed that was not exactly the case, but we were reassured by the fact that left-out articles were probably "shipping or financial reports" (thanks go to Melvin Wevers). We also did not include the colonial newspapers for our embeddings. This is motivated by our research questions. A list of removed newspapers is available on request.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h nl* 6.8M nl_1620_SGNS_corpus_file.gensim 7.9M nl_1640_SGNS_corpus_file.gensim 43M nl_1660_SGNS_corpus_file.gensim 78M nl_1680_SGNS_corpus_file.gensim 138M nl_1700_SGNS_corpus_file.gensim 243M nl_1720_SGNS_corpus_file.gensim 287M nl_1740_SGNS_corpus_file.gensim 431M nl_1760_SGNS_corpus_file.gensim 825M nl_1780_SGNS_corpus_file.gensim 1.2G nl_1800_SGNS_corpus_file.gensim 1.8G nl_1820_SGNS_corpus_file.gensim 3.1G nl_1840_SGNS_corpus_file.gensim 5.2G nl_1860_SGNS_corpus_file.gensim 13G nl_1880_SGNS_corpus_file.gensim</code></pre> <p>English:</p> <p>The models were created with data from the British Library Newspapers collection (<a href="https://www.gale.com/intl/primary-sources/british-library-newspapers%5D">link</a>), the Nichols collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-burney-newspapers-collection">link</a>), and the Burney collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-nichols-newspapers-collection">link</a>). We used everything in the corpora. For English, only SGNS_ALIGN models are available. We thank Gale Cengage for their help with this project.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h en* 4.3M en_1620_SGNS_corpus_file.gensim 11M en_1640_SGNS_corpus_file.gensim 11M en_1660_SGNS_corpus_file.gensim 106M en_1680_SGNS_corpus_file.gensim 409M en_1700_SGNS_corpus_file.gensim 1.7G en_1720_SGNS_corpus_file.gensim 834M en_1740_SGNS_corpus_file.gensim 2.4G en_1760_SGNS_corpus_file.gensim 5.3G en_1780_SGNS_corpus_file.gensim 5.5G en_1800_SGNS_corpus_file.gensim 15G en_1820_SGNS_corpus_file.gensim 42G en_1840_SGNS_corpus_file.gensim 65G en_1860_SGNS_corpus_file.gensim 88G en_1880_SGNS_corpus_file.gensim 26G en_1900_SGNS_corpus_file.gensim 21G en_1920_SGNS_corpus_file.gensim 6.3G en_1940_SGNS_corpus_file.gensim</code></pre> <p><strong>Word embeddings</strong></p> <p>For every language, we train diachronic embeddings as follows. We divide the data in 20-year time bins. We train SGNS_UPDATE and SGNS_ALIGN models. Current research on German (Schlechtweg et al, 2019) and English (Shoemark et al, 2019) indicates you should use the SGNS_ALIGN models. <strong>For EN, FI, NL, no tokens (including punctuation) were removed nor altered, aside from lowercasing</strong>. For SV, see above. Parameters are as follows: SGNS architecture (Mikolov et al 2013), window size of 5, frequency threshold of 100, 5 epochs, 300 dimensions (or 100 for EN).</p> <ul> <li>For SGNS_UPDATE: We first train a model for the first time bin <code>t</code>. To train the model for <code>t+1</code>, we use the <code>t</code> model to initialise the vectors for <code>t+1</code>, set the learning rate to correspond to the end learning rate of <code>t</code>, and continue training. This approach, closely following Kim et al (2014), has the advantage of avoiding the need for post-training vector space alignment.</li> </ul> <p>The Python snippet below, which makes use of gensim (Rehurek and Sojka, 2010), illustrates the approach. Special thanks go to Sara Budts.</p> <pre><code>## dict_files[key] is a dictionary with double decades as keys and a corresponding LineSentence object as value: https://radimrehurek.com/gensim/models/word2vec.html#gensim.models.word2vec.LineSentence count = 0 for key in sorted(list(dict_files.keys())): if count == 0: ## This is the first model. model = gensim.models.Word2Vec(corpus_file=dict_files[key], min_count=100, sg=1 ,size=300, workers=64, seed=1830, iter=5) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) print("Model saved, on to the next\n") count += 1 if count > 0: ## this is for the subsequent models. print("model for double decade starting in",str(key)) model = gensim.models.Word2Vec.load(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin-20)+".w2v")) print("previous model loaded") model.build_vocab(corpus_file=dict_files[key], update=True) model.train(corpus_file=dict_files[key], total_words = model.corpus_count, total_examples = model.corpus_count, start_alpha = model.alpha, end_alpha = model.min_alpha, epochs=model.epochs) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) </code></pre> <ul> <li>For SGNS_ALIGN: We independently train models for all time bins. The models in this repository are <em>NOT</em> aligned, leaving you the choice of how to align them. For example, <a href="https://gist.github.com/quadrismegistus/09a93e219a6ffc4f216fb85235535faf">here</a> is a link to code by Ryan Heuser to do just that. Models were trained with the <code>count == 0</code> scenario in the snippet above.</li> </ul> <p><strong>Acknowledgments</strong></p> <p>This work has been supported by the European Union's Horizon 2020 research and innovation programme under grant 770299 <a href="https://www.newseye.eu/">NewsEye</a>. Specials thanks go to the data providers/collection-holding institutions: the Finnish Language Bank, the Swedish Language Bank, the Royal Dutch Library, and Gale Cengage.</p> <p>The authors would like to thank the following persons and group, listed alphabetically: Antoine Doucet, Antti Kanner, Axel-Jean Caurant, Dominik Schlechtweg, Eetu Mäkelä, Elaine Zosa, Estelle Bunout, Haim Dubossarsky, Joris van Eijnatten, Krister Lindén, Lars Borin, Lidia Pivovarova, Melvin Wevers, Nina Tahmasebi, Sara Budts, Senka Drobac, Tanja Säily, the COMHIS group, and Steven Claeyssens. Computational resources were provided by CSC – IT Center for Science Ltd.</p> <p><strong>References</strong></p> <p>Borin, L., Forsberg, M., Roxendal, J. (2012). Korp-the corpus infrastructure of Spräkbanken,in: LREC. pp. 474–478.</p> <p>Kim, Y., Chiu, Y.I., Hanaki, K., Hegde, D. and Petrov, S. (2014). Temporal Analysis of Language through Neural Language Models. <em>ACL 2014</em>, p.61.</p> <p>Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space. <em>arXiv preprint arXiv:1301.3781</em>.</p> <p>National Library of Finland (2011). <em>The Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland, Kielipankki Version</em> [text corpus]. Kielipankki. Retrieved from <a href="http://urn.fi/urn:nbn:fi:lb-2016050302">http://urn.fi/urn:nbn:fi:lb-2016050302</a>.</p> <p>Rehurek, R. and Sojka, P. (2010). Software framework for topic modelling with large corpora. In <em>Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</em>.</p> <p>Royal Dutch Library (2017). <em>Delpher open krantenarchief (1.0)</em>. Den Haag, 2017.</p> <p>Schlechtweg D., Hätty A, del Tredici M., and Schulte im Walde S. (2019). A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In <em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, Florence, Italy. ACL.</p> <p>Shoemark, P., Liza, F.F., Nguyen, D., Hale, S. and McGillivray, B. (2019). Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings. In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 66-76)</em>, Hong Kong.</p> <p>Språkbanken. <em>The Kubhist Corpus</em>. Department of Swedish, University of Gothenburg. <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">https://spraakbanken.gu.se/korp/?mode=kubhist</a>.</p>
Electromagnetic data (FDEM and ERT) collected in the Venice coastland (Zennare basin)
<p>FDEM and ERT data collected southern of the Venice lagoon (Italy) in the Zennare basin in 2019-2020.</p> <p>FDEM_ZENNARE37.csv: Raw output of Quadrature and Inphase values for the 6 frequencies adopted with the GEM2 FDEM probe.<br> ERT_ROUGHoutput_Zennare.dat: Apparent resistivity data as retrieved with the 48 channels Syscal Pro georesistivimeter ERT.</p>
200 kHz pre-processed echosounder data collected on the Antarctic Circumnavigation Expedition during the austral summer of 2016/2017.
<p><strong>Dataset abstract </strong></p> <p>These data consist of pre-processed echosounder observations in the Southern Ocean collected during the Antarctic Circumnavigation Expedition (ACE; Leg2-Leg3) using an EK60 GPT operating at 200 kHz. The instrument was calibrated at South Georgia during the expedition (Leg 3) and corrections were applied prior to calculation of the volume backscattering strength (Sv). The signal-to-noise ratio (SNR) was analysed and was deemed very poor at depths greater than 100 m. Therefore, only data collected between the transducer depth (8.4 m) and 100 m were archived.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ACE-DYYYYMMDD-THHMMSS.csv, data files, comma-separated values</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This 200 kHz pre-processed echosounder data collected on ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
Summary raw meteorological data from the Southern Ocean collected on board the Antarctic Circumnavigation Expedition (ACE) during the austral summer of 2016/2017.
<p><strong>Dataset abstract</strong></p> <p>A Vaisala MAWS240 meteorological station was installed on the R/V Akademik Tryoshnikov during a circumnavigation of Antarctica in the austral summer season of 2016/2017. This dataset contains the raw meteorological data that have been extracted from the original raw text data files. Data coverage is from 17th November 2016 until 11th April 2017, with gaps where the ship was in port.</p> <p>Air temperature, relative humidity, dew point, solar radiation, ultraviolet radiation, cloud level and sky cover were recorded with a resolution of 30 seconds. Averaged wind parameter data are provided.</p> <p>Date_time should be combined with TIMEDIFF to convert it to UTC. Latitude and longitude recorded are not corrected. Underway seawater measurements were recorded as null values.</p> <p>Data from this dataset have been corrected and quality-checked in another published dataset. We recommend these data for further use (Landwehr et al., 2019; DOI 10.5281/zenodo.3379590).</p> <p><strong>Dataset contents</strong></p> <ul> <li>metdata_all_YYYYMMDD_YYYYMMDD.csv, data file, comma-separated values</li> <li>data_file_header, metadata, text format</li> <li>README.txt, metadata, text format</li> <li>ace_meteorology_raw_summary_change_log.txt</li> </ul> <p>Data files contain data for each leg of the Antarctic Circumnavigation Expedition (ACE). Dates included in the file name are the start and end dates of the legs and therefore the data within the files as well.</p> <p><strong>Change log</strong></p> <p><strong>v1.2</strong> - Added missing data from 2017-02-05 - 2017-02-08 inclusive. Updated this change log file.</p> <p><strong>v1.1</strong> - Added additional data coverage from 2016-11-17 - 2016-11-22 inclusive, into the first data file. Updated README.txt with information about data coverage. Added this change log file.</p> <p><strong>v1.0</strong> - Initial release of raw summary meteorological data.</p> <p><strong>Dataset license</strong></p> <p>This raw meteorological dataset is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
ADS-C Air Traffic Data Collected by the OpenSky Network
<p>ADS-C data collected by the OpenSky Network since 7th July 2023. </p> <p>Data underlying (Version 1.1)</p> <h1>A First Look at Exploiting the Automatic Dependent Surveillance-Contract Protocol for Open Aviation Research</h1> <p>https://journals.open.tudelft.nl/joas/article/view/7229</p>
Field data collected from pyroclastic and lahar deposits of the 472 AD (Pollena) and 1631 Vesuvius eruptions
<p><strong><span>Field data collected from pyroclastic and lahar deposits of the 472 AD (Pollena) and 1631 Vesuvius eruptions</span></strong></p> <p><span>Mauro A. Di Vito<sup>1</sup>, Ilaria Rucco<sup>2</sup>, Sandro de Vita<sup>1</sup>, Domenico M. Doronzo<sup>1</sup>, Marina Bisson<sup>3</sup>, Elena Zanella<sup>4</sup></span></p> <p><sup><span>1</span></sup><span> Istituto Nazionale di Geofisica e Vulcanologia, Osservatorio Vesuviano, Napoli, Italy</span></p> <p><sup><span>2</span></sup><span> Heriot-Watt University, School of Engineering and Physical Sciences, Edinburgh, United Kingdom</span></p> <p><sup><span>3</span></sup><span> Istituto Nazionale di Geofisica e Vulcanologia, Sezione di Pisa, Pisa, Italy</span></p> <p><sup><span>4</span></sup><span> Università di Torino, Dipartimento di Scienze della Terra, Torino, Italy</span></p> <p><span> </span></p> <p><span>This dataset is organized in an Excel file, and it includes all the data collected and reviewed during the last 20 years from drill cores, outcrops, archaeological excavations, stratigraphic trenches, and the existing literature. It focuses on the primary (pyroclastic) and secondary (lahar) deposits of the 472 AD (Pollena) and 1631 eruptions from the Somma-Vesuvius volcano. The aim is to collect stratigraphic, stratimetric, sedimentological, lithological and chronological data to generate distribution maps and to validate the numerical simulations and models for the risk assessment. In particular, this dataset is complementary to – and in support of – the full work by Di Vito et al. (2024), in which the distribution of those deposits all around the Somma-Vesuvius complex and further is presented and discussed. Such dataset was used to inform the shallow-water model of lahars by de’ Michieli Vitturi et al. (2024), which in turns was used by Sandri et al. (2024) to elaborate probabilistic maps of lahar invasion in the Somma-Vesuvius and Apennine areas.</span></p> <p><span>All the data are organized in columns: the first four aim to identify the sites, and so there is a numeric identification code (ID), the name of the site (NAME), and the metric coordinates (East-North) in the UTM WGS 84 – Zone 33 reference projection (X, Y). The last two columns are “MUNICIPALITY” and “PROVINCE” and give a spatial location to the points.</span></p> <p><span>For the two eruptions, several columns have been created:</span></p> <p><span><span>·<span> </span></span></span><span>“472_PRIM”, “1631_PRIM” and “472_ASH” and “1631_ASH” indicate, respectively, the fallout primary deposits of the eruptions and the primary ash, particularly the ash related to the last phases of the eruptions (generally phreatomagmatic).</span></p> <p><span><span>·<span> </span></span></span><span>“472_SYN” and “1631_SYN” indicate the syn-eruptive lahars related to the two eruptions, recognized from the similar composition between the primary deposit and the lahar and from the evidence of a short-term exposure between the two.</span></p> <p><span><span>·<span> </span></span></span><span>“472_POST” AND “1631_POST” indicate the post-eruptive lahars related to the two eruptions. They are considered “post” when in the deposit there are pumices belonging to older eruptions, indicating their involvement in the progressive erosion of the slopes and valleys, and when there is evidence of long periods without deposition, such as the presence of</span><span> </span><span>slightly humified surfaces or traces of human artifacts (excavations, ploughing).</span></p> <p><span><span>·<span> </span></span></span><span>“EROSION_fallout_472” refers to the sites where it was possible to find erosional unconformities between the pyroclastic deposit of the 472 AD eruption and the lahar, as well as between the lower and upper lahar flow units. The erosional features are for example the lack of one or more primary eruptive layers (eroded by the overlying deposit), a change in the granulometry, or lateral discontinuity of the deposit. </span></p> <p><span><span>·<span> </span></span></span><span>“Pdyn (kPa)”, “v (m/s)”, “C (%)” and “T” are all the parameters quantified to validate the numerical models and to assess the hazard from lahars. Pdyn is the flow dynamic pressure, which represents the capability of the flow to entrain a clast, and it depends on the velocity (v) and the flow density, which in turn results from a combination of the density of the particles and the water through the “C (%)”, that is the particle volume concentration. To calculate the flow dynamic pressure and the velocity, the parameters taken into account are the dimensions of the biggest clasts and the nature of the clasts (limestone, ceramic, brick, tephra, lava, sandstone, iron) found in the lahar deposits. The concentration is estimated considering some sites in which the flow expands in correspondence with some obstacles (for example a Roman wall). This can be assumed to be the initial height of the flow before the emplacement. Finally, “T” refers to the estimated deposition temperature of the deposit quantified by the magnetic analysis, in particular in some sites where the lahar interacted with anthropogenic structures.</span></p> <p><span><span>·<span> </span></span></span><span>“DEPOSIT (472)” indicates the type of lahar deposit (syn- or post-eruptive) of the 472 AD eruption in which the fragments were found.</span></p> <p><span><span>·<span> </span></span></span><span>“MULTIPLE LAHAR UNITS” indicates the sites in which multiple flow units are vertically identified in the lahar deposits. They are generally a result of rapid and progressive aggradation of multiple flow pulses, each one resulting from single-pulse “en masse” emplacement</span><span>.</span></p>
Raw planetary images and boulder labels data (as shapefiles) collected during the BOULDERING Marie Skłodowska-Curie Global fellowship
<p>This database contains 64 large images of craters on the lunar and martian surfaces and 3 images of boulder fields on Earth (see manuscript <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a> for more information on those terrestrial locations). The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024.</p> <p>For each image, the boulder outlines within specific tiles within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript <a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). </p> <p>For each location, you will find a raster with a .tif format, and three shapefiles:</p> <ul> <li> <p>a boulder-mapping file, which is the manually digitized outline of boulders.</p> </li> <li> <p>a tiles-completely-mapped file, which depicts the patches/tiles/windows on which the boulder mapping has been conducted.</p> </li> <li> <p>a global-tiles file, which shows all of the image patches/tiles/windows (pick the term you are the most familiar with) within a raster.</p> </li> </ul> <p>In addition you will find .pkl (which stands for pickle), which contains some information about the patches/tiles/windows if you would need to clip those windows out from the original raster. You can find more information in the way we process this raw data into a format which can be ingested in a deep learning model (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>) in the two following github repositories (<a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth</a> and <a href="https://github.com/astroNils/MLtools/tree/main" target="_blank" rel="noopener">https://github.com/astroNils/MLtools</a>). If you don't plan in adding more training data, you can directly used the pre-processed database (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>).</p> <p>There are multiple locations/images per planetary body. Cold spots are located on the Moon, but they are saved in a folder of their own. </p> <p>Note that the cold spots boulder mapping shapefiles are partially manually mapped, and partially originating from predictions made from a deep learning model (which explains the outline of boulders are predicted within one pixel).</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── raw_data/ ├── coldspots/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── earth/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── mars/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif └── moon/ └── image_name/ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp └── raster/ └── <image_name>.tif</code></pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.