Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,101

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,101 results for “historical”

Learn how ShareScore rates datasets ↗
edi52/100

Wisconsin Lake Historical Limnological Parameters 1925 - 2009

This dataset is a compilation of ten sources of data representing physical and chemical properties of 13,093 Wisconsin lakes. The goal was to compile a comprehensive resource of historical and more recent lake information which would be accessible by querying a single database. Due to the wide temporal extent (1925-2009), methods used for measuring lake parameters in this dataset have varied. A careful look at the available metadata and background information is recommended. Sampling Frequency: varies Number of sites: 13,093

openCC (other)Dec 2022View details →
edi52/100

SBC LTER: Reef: Historical Kelp Database for giant kelp (Macrocystis pyrifera) biomass in California and Mexico

ISP Alginates (formerly Kelco Co.) has collected information on the abundance of giant kelp (Macrocystis pyrifera) in California and Mexico from routine aerial surveys since 1957. The standard protocol consists of an observer visually estimating the amount of harvestable giant kelp biomass within designated kelp beds from a small fixed-wing aircraft. Observations were recorded on paper data sheets in the field and archived in notebooks housed at ISP Alginates. With cooperation from ISP Alginates, the SBCLTER converted ISP Alginates long-term records of giant kelp biomass into a digital format. The database consists of a data table containing kelp biomass, a catalog of maps. The format ISP Alginates used to report kelp abundance data changed periodically over the course of the collecting period. These details, pus descriptions of designated kelp beds are described in the protocol document.

openCC (other)Oct 2022View details →
edi52/100

SBC LTER: Reef: Benthic community structure along a gradient of historic kelp variability

These data are estimates of biomass of approximately 225 taxa of reef algae, invertebrates, and fish in transects at 11 non-core sites in the Santa Barbara Channel in summer 2018 (3 transects per site). Sites were selected along a gradient of historic kelp (Macrocystis pyrifera) variability, from sites with highly persistent kelp to sites exhibiting extensive variation in kelp biomass since 2008. See the site characteristics data table for site locations and depths. The purpose of the sampling was to explore to what extent the findings from the long-term experiment (e.g. Castorani et al. 2018) apply to natural gradients in kelp persistence. Surveys were conducted following the same methodology used in the annual surveys of kelp forest community structure, such that data from the non-core sites may be paired with annual survey data collected in summer 2018.

openCC (other)Apr 2024View details →
zenodo48/100

Models for "A data-driven approach to studying changing vocabularies in historical newspaper collections"

<p>NOTE: This is a badly rendered version of the README within the archive.</p> <p><strong>A data-driven approach to studying changing vocabularies in historical newspaper collections</strong></p> <p>Simon Hengchen,* Ruben Ros,** Jani Marjanen,*** Mikko Tolonen***</p> <p>*<a href="https://spraakbanken.gu.se/en/about/staff/simon">Spr&aring;kbanken Text</a>, University of Gothenburg, Sweden and <a href="https://iguanodon.ai">iguanodon.ai</a>, Belgium: firstname.lastname@gu.se<br> **<a href="https://www.c2dh.uni.lu/people/ruben-ros">Centre for Contemporary and Digital History (C2DH)</a>, University of Luxembourg:&nbsp;firstname.lastname@uni.lu<br> ***<a href="https://www.helsinki.fi/en/researchgroups/computational-history">COMHIS</a>, University of Helsinki:&nbsp;<a href="mailto:firstname.lastname@helsinki.fi">firstname.lastname@helsinki.fi</a>;</p> <p>These are the supplementary materials for the DH2019 paper&nbsp;<em>A data-driven approach to the changing vocabulary of the &lsquo;nation&rsquo; in English, Dutch, Swedish and Finnish newspapers, 1750-1950</em>, as well as the 2021 Digital Scholarship in the Humanities publication available in OpenAccess: <a href="https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793">https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793</a>. If you end up using whole or parts of this resource, please use the following citation(s):</p> <ul> <li>Hengchen, S., Ros, R., and Marjanen, J. (2019). A data-driven approach to the changing vocabulary of the &#39;nation&#39; in English, Dutch, Swedish and Finnish newspapers, 1750-1950. In&nbsp;<em>Proceedings of the Digital Humanities (DH) conference 2019, Utrecht, The Netherlands</em></li> </ul> <p>and/or:</p> <ul> <li>Hengchen, S., Ros, R., Marjanen, J. and Tolonen, M., 2021. A data-driven approach to studying changing vocabularies in historical newspaper collections. Digital Scholarship in the Humanities, 36(Supplement_2), pp.ii109-ii126.</li> </ul> <p>or alternatively use one of the following&nbsp;<code>bib</code>s:</p> <pre><code>@inproceedings{hengchen2019nation, title="A data-driven approach to the changing vocabulary of the 'nation' in {E}nglish, {D}utch, {S}wedish and {F}innish newspapers, 1750-1950.", author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani}, year={2019}, address = "Utrecht, The Netherlands", booktitle={Proceedings of the Digital Humanities (DH) conference 2019} }</code></pre> <pre><code>@article{hengchen2021data, title={A data-driven approach to studying changing vocabularies in historical newspaper collections}, author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani and Tolonen, Mikko}, journal={Digital Scholarship in the Humanities}, volume={36}, number={Supplement\_2}, pages={ii109--ii126}, year={2021}, publisher={Oxford University Press} }</code></pre> <p>&nbsp;</p> <p>Files</p> <p>This archive contains two folders -- one per diachronic representation method -- as well as this README. The folders each contain four folders, which contain the models for their respective languages. As can be inferred from the small datasize, most of the earlier models are not reliable and should not be used, but are still made available. This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p><strong>Source material</strong></p> <p>Finnish:</p> <p>The models were created with data from the Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland (National Library of Finland, 2011). We used everything in the corpus.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h fi* 12M fi_1820_SGNS_corpus_file.gensim 89M fi_1840_SGNS_corpus_file.gensim 797M fi_1860_SGNS_corpus_file.gensim 7.0G fi_1880_SGNS_corpus_file.gensim 22G fi_1900_SGNS_corpus_file.gensim</code></pre> <p>Swedish:</p> <p>The models were created with data from the Kubhist 2 corpus (Spr&aring;kbanken) -- more precisely, the data dumps available at&nbsp;<a href="https://spraakbanken.gu.se/lb/resurser/meningsmangder/">https://spraakbanken.gu.se</a>. After a manual evaluation of Swedish embeddings trained without pre-processing seemed to show that the embeddings were of low quality, we retrained models, only keeping sentences that were at least 10 tokens long and were constituted of at least 50% of lemmas as per the KORP processing pipeline (Borin et al, 2012).</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h sv* 1.6M sv_1740_SGNS_corpus_file.gensim 44M sv_1760_SGNS_corpus_file.gensim 124M sv_1780_SGNS_corpus_file.gensim 228M sv_1800_SGNS_corpus_file.gensim 678M sv_1820_SGNS_corpus_file.gensim 1.6G sv_1840_SGNS_corpus_file.gensim 4.5G sv_1860_SGNS_corpus_file.gensim 6.5G sv_1880_SGNS_corpus_file.gensim 113M sv_1900_SGNS_corpus_file.gensim</code></pre> <p>Dutch:</p> <p>The models were created with data from the Delpher newspaper archive (Royal Dutch Library, 2017), through data dumps for newspapers until and including 1876, and through API hits for articles from 1877 to 1899 (included).</p> <ul> <li>For anything pre-1877 we discarded full texts that had, in the metadata, anything else than exclusively&nbsp;<code>nl</code>&nbsp;or&nbsp;<code>NL</code>&nbsp;as a language tag.</li> <li>For the full texts between 1877 and 1899: we queried the API for all items in the &ldquo;artikel&rdquo; category that contained the determiner&nbsp;<code>de</code>.</li> </ul> <p>Our assumption was that most articles should contain&nbsp;<code>de</code>&nbsp;at least once, and those that didn&#39;t were too short to be deemed interesting. A subsequent study showed that was not exactly the case, but we were reassured by the fact that left-out articles were probably &quot;shipping or financial reports&quot; (thanks go to Melvin Wevers). We also did not include the colonial newspapers for our embeddings. This is motivated by our research questions. A list of removed newspapers is available on request.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h nl* 6.8M nl_1620_SGNS_corpus_file.gensim 7.9M nl_1640_SGNS_corpus_file.gensim 43M nl_1660_SGNS_corpus_file.gensim 78M nl_1680_SGNS_corpus_file.gensim 138M nl_1700_SGNS_corpus_file.gensim 243M nl_1720_SGNS_corpus_file.gensim 287M nl_1740_SGNS_corpus_file.gensim 431M nl_1760_SGNS_corpus_file.gensim 825M nl_1780_SGNS_corpus_file.gensim 1.2G nl_1800_SGNS_corpus_file.gensim 1.8G nl_1820_SGNS_corpus_file.gensim 3.1G nl_1840_SGNS_corpus_file.gensim 5.2G nl_1860_SGNS_corpus_file.gensim 13G nl_1880_SGNS_corpus_file.gensim</code></pre> <p>English:</p> <p>The models were created with data from the British Library Newspapers collection (<a href="https://www.gale.com/intl/primary-sources/british-library-newspapers%5D">link</a>), the Nichols collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-burney-newspapers-collection">link</a>), and the Burney collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-nichols-newspapers-collection">link</a>). We used everything in the corpora. For English, only SGNS_ALIGN models are available. We thank Gale Cengage for their help with this project.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h en* 4.3M en_1620_SGNS_corpus_file.gensim 11M en_1640_SGNS_corpus_file.gensim 11M en_1660_SGNS_corpus_file.gensim 106M en_1680_SGNS_corpus_file.gensim 409M en_1700_SGNS_corpus_file.gensim 1.7G en_1720_SGNS_corpus_file.gensim 834M en_1740_SGNS_corpus_file.gensim 2.4G en_1760_SGNS_corpus_file.gensim 5.3G en_1780_SGNS_corpus_file.gensim 5.5G en_1800_SGNS_corpus_file.gensim 15G en_1820_SGNS_corpus_file.gensim 42G en_1840_SGNS_corpus_file.gensim 65G en_1860_SGNS_corpus_file.gensim 88G en_1880_SGNS_corpus_file.gensim 26G en_1900_SGNS_corpus_file.gensim 21G en_1920_SGNS_corpus_file.gensim 6.3G en_1940_SGNS_corpus_file.gensim</code></pre> <p><strong>Word embeddings</strong></p> <p>For every language, we train diachronic embeddings as follows. We divide the data in 20-year time bins. We train SGNS_UPDATE and SGNS_ALIGN models. Current research on German (Schlechtweg et al, 2019) and English (Shoemark et al, 2019) indicates you should use the SGNS_ALIGN models.&nbsp;<strong>For EN, FI, NL, no tokens (including punctuation) were removed nor altered, aside from lowercasing</strong>. For SV, see above. Parameters are as follows: SGNS architecture (Mikolov et al 2013), window size of 5, frequency threshold of 100, 5 epochs, 300 dimensions (or 100 for EN).</p> <ul> <li>For SGNS_UPDATE: We first train a model for the first time bin&nbsp;<code>t</code>. To train the model for&nbsp;<code>t+1</code>, we use the&nbsp;<code>t</code>&nbsp;model to initialise the vectors for&nbsp;<code>t+1</code>, set the learning rate to correspond to the end learning rate of&nbsp;<code>t</code>, and continue training. This approach, closely following Kim et al (2014), has the advantage of avoiding the need for post-training vector space alignment.</li> </ul> <p>The Python snippet below, which makes use of gensim (Rehurek and Sojka, 2010), illustrates the approach. Special thanks go to Sara Budts.</p> <pre><code>## dict_files[key] is a dictionary with double decades as keys and a corresponding LineSentence object as value: https://radimrehurek.com/gensim/models/word2vec.html#gensim.models.word2vec.LineSentence count = 0 for key in sorted(list(dict_files.keys())): if count == 0: ## This is the first model. model = gensim.models.Word2Vec(corpus_file=dict_files[key], min_count=100, sg=1 ,size=300, workers=64, seed=1830, iter=5) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) print("Model saved, on to the next\n") count += 1 if count &gt; 0: ## this is for the subsequent models. print("model for double decade starting in",str(key)) model = gensim.models.Word2Vec.load(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin-20)+".w2v")) print("previous model loaded") model.build_vocab(corpus_file=dict_files[key], update=True) model.train(corpus_file=dict_files[key], total_words = model.corpus_count, total_examples = model.corpus_count, start_alpha = model.alpha, end_alpha = model.min_alpha, epochs=model.epochs) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) </code></pre> <ul> <li>For SGNS_ALIGN: We independently train models for all time bins. The models in this repository are&nbsp;<em>NOT</em>&nbsp;aligned, leaving you the choice of how to align them. For example,&nbsp;<a href="https://gist.github.com/quadrismegistus/09a93e219a6ffc4f216fb85235535faf">here</a>&nbsp;is a link to code by Ryan Heuser to do just that. Models were trained with the&nbsp;<code>count == 0</code>&nbsp;scenario in the snippet above.</li> </ul> <p><strong>Acknowledgments</strong></p> <p>This work has been supported by the European Union&#39;s Horizon 2020 research and innovation programme under grant 770299&nbsp;<a href="https://www.newseye.eu/">NewsEye</a>. Specials thanks go to the data providers/collection-holding institutions: the Finnish Language Bank, the Swedish Language Bank, the Royal Dutch Library, and Gale Cengage.</p> <p>The authors would like to thank the following persons and group, listed alphabetically: Antoine Doucet, Antti Kanner, Axel-Jean Caurant, Dominik Schlechtweg, Eetu M&auml;kel&auml;, Elaine Zosa, Estelle Bunout, Haim Dubossarsky, Joris van Eijnatten, Krister Lind&eacute;n, Lars Borin, Lidia Pivovarova, Melvin Wevers, Nina Tahmasebi, Sara Budts, Senka Drobac, Tanja S&auml;ily, the COMHIS group, and Steven Claeyssens. Computational resources were provided by CSC &ndash; IT Center for Science Ltd.</p> <p><strong>References</strong></p> <p>Borin, L., Forsberg, M., Roxendal, J. (2012). Korp-the corpus infrastructure of Spr&auml;kbanken,in: LREC. pp. 474&ndash;478.</p> <p>Kim, Y., Chiu, Y.I., Hanaki, K., Hegde, D. and Petrov, S. (2014). Temporal Analysis of Language through Neural Language Models.&nbsp;<em>ACL 2014</em>, p.61.</p> <p>Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space.&nbsp;<em>arXiv preprint arXiv:1301.3781</em>.</p> <p>National Library of Finland (2011).&nbsp;<em>The Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland, Kielipankki Version</em>&nbsp;[text corpus]. Kielipankki. Retrieved from&nbsp;<a href="http://urn.fi/urn:nbn:fi:lb-2016050302">http://urn.fi/urn:nbn:fi:lb-2016050302</a>.</p> <p>Rehurek, R. and Sojka, P. (2010). Software framework for topic modelling with large corpora. In&nbsp;<em>Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</em>.</p> <p>Royal Dutch Library (2017).&nbsp;<em>Delpher open krantenarchief (1.0)</em>. Den Haag, 2017.</p> <p>Schlechtweg D., H&auml;tty A, del Tredici M., and Schulte im Walde S. (2019). A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In&nbsp;<em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, Florence, Italy. ACL.</p> <p>Shoemark, P., Liza, F.F., Nguyen, D., Hale, S. and McGillivray, B. (2019). Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings. In&nbsp;<em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 66-76)</em>, Hong Kong.</p> <p>Spr&aring;kbanken.&nbsp;<em>The Kubhist Corpus</em>. Department of Swedish, University of Gothenburg.&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">https://spraakbanken.gu.se/korp/?mode=kubhist</a>.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Historical Tropical Cyclone Along-track Potential Intensity (and Derived Quantities) for Six Ocean Basins from Reanalyses

<p>Supporting derived data for Shields et al. (2020, GRL).</p> <p>Derived tropical cyclone potential intensities and associated variables across the North Atlantic (NA), Eastern North&nbsp;Pacific (EP), North Indian (NI), South Indian (SI), South Pacific (SP), and Western North Pacific (WP)&nbsp;ocean basins, from MERRA2, ERA-I, and MERRA2 with SSTs replaced by HadISSTs. NA/WP basins also have potential&nbsp;and observed intensities calculated with NCEP/NCAR and ERA-20C reanalyses over 1950-2016 and 1950-2010, respectively.</p> <p>All files are netcdf format, organized by basin, with&nbsp;suffixes on data variables to indicate reanalysis:</p> <ul> <li>&quot;_m&quot;: MERRA2 (Gelaro et al. 2017)</li> <li>&quot;_h&quot;: MERRA2-HadISSTs (Rayner et al. 2003)</li> <li>&quot;_e&quot;:&nbsp;ERA-I (Dee et al. 2011)</li> <li>&quot;_n&quot;: NCEP/NCAR (Kalnay et al. 2016)</li> <li>&quot;_c&quot;: ERA-20C (Stickler et al. 2014)</li> </ul> <p>When using this data, please include the citation:</p> <blockquote> <p><strong>Shannon Shields, Allison Wing, and Daniel M. Gilford, 2020: A Global Analysis of Interannual Variability of Potential and Actual Tropical Cyclone Intensities. Geophys. Res. Lett.</strong></p> </blockquote> <p>Potential intensities calculated with the Bister and Emanuel (2002) algorithm (<strong>pcmin.m</strong>) by Kerry Emanuel (revised by Daniel Gilford, Gilford et al. 2019), available freely at:&nbsp;ftp://texmex.mit.edu/pub/emanuel/TCMAX</p> <p>MERRA2, ERA-I, and MERRA2 with SSTs replaced by HadISSTs&nbsp;calculations were performed&nbsp;by Daniel Gilford; NCEP/NCAR and ERA-20C calculations were performed by&nbsp;Dr. Suzana Camargo&nbsp;(many thanks!).</p> <p>Please direct any questions or comments to daniel[dot]gilford[at]rutgers[dot]edu.</p>

opencc-by-4.0Jun 2020View details →
zenodo48/100

A Finite State Tranducer that models Chinese Historical Phonology

<p>This is a finite state transducer that tries to model Chinese historical phonology from Old Chinese as reconstructed by Baxter and Sagart in <em>Old Chinese: a New Reconstruction</em> (Oxford, 2014) to Middle Chinese as presented in the system of Baxter in <em>A Handbook of Old Chinese Phonology </em>(Mouton, 1992).</p>

opencc-by-4.0Jul 2020View details →
zenodo48/100

C3-EURO4M-MEDARE Mediterranean historical climate data - v.2

<p>Historical surface climate data files and meta-data for stations in Mediterranean North Africa and Middle East areas (1852-2008).</p>

opencc-zeroApr 2015View details →
zenodo48/100

Historical Weather, Load, Wind, and Solar Data for the Salt River Project

<p>We created and curated a dataset of historical (1980-2019) hourly meteorology, load, wind, and solar data for the Salt River Project (SRP) region. The data was created by PNNL's <a href="https://godeeep.pnnl.gov/">GODEEEP</a> project. Each row in the dataset is a single hour and each column is a variable. All meteorological variables are spatially-averaged over the SRP service territory. The variables and their units are as follows:</p><ol><li>"Time_UTC"; Coordinated Universal Time (UTC); Time of day.</li><li>"T2"; Fahrenheit; 2-m air temperature.</li><li>"Q2"; kg/kg; 2-m water vapor mixing ratio.</li><li>"SWDOWN"; W/m^2; Downwelling shortwave radiative flux at the surface.</li><li>"GLW"; W/m^2; Downwelling longwave radiative flux at the surface.</li><li>"WSPD"; m/s; 10-m wind speed.</li><li>"Scaled_2019_Load"; MWh; Simulated hourly demand for electricity that is scaled to 2019 levels of annual energy. This load estimate does not account for historical changes in population and economics within the SRP service territory. It is included to make it easier to isolate weather impacts on load without having to consider long-term changes.</li><li>"Load"; MWh; Simulated hourly demand for electricity.</li><li>"Agua_Fria_Solar_Capacity"; N/A; Solar capacity factor for the SRP Agua Fria project with plant configurations taken from the EIA-860 database.</li><li>"Phoenix_Solar_Capacity"; N/A; Solar capacity factor for hypothetical solar plants derived using the grid cell nearest to Phoenix, AZ.</li><li>"Flagstaff_Solar_Capacity"; N/A; Solar capacity factor for hypothetical solar plants derived using the grid cell nearest to Flagstaff, AZ.</li><li>"Phoenix_Wind_Capacity"; N/A; Wind capacity factor for hypothetical 80-m plants derived using the grid cell nearest to Phoenix, AZ.</li><li>"Flagstaff_Wind_Capacity"; N/A; Wind capacity factor for hypothetical 80-m plants derived using the grid cell nearest to Flagstaff, AZ.</li></ol>

opencc-zeroNov 2023View details →
zenodo48/100

Zero-degree isotherm latitude (ZIL) position over Antarctica: Historical and Projections

<p>This is the dataset associated to&nbsp;the research 'Southward migration of the zero-degree isotherm latitude&nbsp;over the Southern Ocean and the Antarctic Peninsula: extent and implications' published in <i>Science of the Total Environment</i>.</p><p>This repository contains:</p><ul><li><strong>ZIL_ERA5_1957-2020_position.zip:</strong>&nbsp;Historical position of the ZIL for every longitude point in ERA5 from 1957 to 2020 for different <i>seasons</i>. Files named:<ul><li>ZIL_ERA5_1957-2020<i>[season]</i>position.csv<ul><li>Dimensions:&nbsp;[lons, years]</li><li>Units: degrees latitude</li></ul></li></ul></li><li><strong>ZIL_ERA5_1957-2020_timeseries.csv:</strong>&nbsp;Historical spatially averaged&nbsp;position of the ZIL for Antartica (Ant) and the Antarctic Peninsula (AP) in ERA5 from 1957 to 2020 for different <i>seasons</i>. File named:<ul><li>ZIL_ERA5_1957-2020_timeseries.csv<ul><li>Dimensions:&nbsp;[years, season_area]</li><li>Units: degrees latitude</li></ul></li></ul></li><li><strong>ZIL_ERA5_1957-2020_meanposition.csv:</strong>&nbsp;Historical temporally averaged&nbsp;position of the ZIL&nbsp;in ERA5 from 1957 to 2020 for different <i>seasons </i>and <i>months</i>. File named:<ul><li>ZIL_ERA5_1957-2020_meanposition.csv<ul><li>Dimensions:&nbsp;[lons, season/month]</li><li>Units: degrees latitude</li></ul></li></ul></li><li><strong>ZIL_CEMIP6_Historical_position.zip:</strong>&nbsp;Mean position of the ZIL for every longitude point in Historical simulations of&nbsp;CEMIP6 from 1957 to 2014 for different <i>seasons</i>. Files named:<ul><li>ZIL_CEMIP6_Hist_[<i>season</i>]_position.csv<ul><li>Dimensions:&nbsp;[lons, models]</li><li>Units: degrees latitude</li></ul></li></ul></li><li><strong>ZIL_CEMIP6_SSP2-45.zip:</strong>&nbsp;Mean position of the ZIL for every longitude point under the SSP2-4.5 scenario in&nbsp;CEMIP6 for the period 2040-69 and 2070-90 for different <i>seasons</i>. Files named:<ul><li>ZIL_CEMIP6_SSP2-45_2040-69_[<i>season</i>]_position.csv<ul><li>Dimensions:&nbsp;[lons, models]</li><li>Units: degrees latitude</li></ul></li><li>ZIL_CEMIP6_SSP2-45_2070-99_[<i>season</i>]_position.csv<ul><li>Dimensions:&nbsp;[lons, models]</li><li>Units: degrees latitude</li></ul></li></ul></li><li><strong>ZIL_CEMIP6_SSP5-85.zip:</strong>&nbsp;Mean position of the ZIL for every longitude point under the SSP5-8.5 scenario in&nbsp;CEMIP6 for the period 2040-69 and 2070-90&nbsp;and trends for the period 2015-99 for different <i>seasons</i>. Files named:<ul><li>ZIL_CEMIP6_SSP5-85_2040-69_[<i>season</i>]_position.csv<ul><li>Dimensions:&nbsp;[lons, models]</li><li>Units: degrees latitude</li></ul></li><li>ZIL_CEMIP6_SSP5-85_2070-99_[<i>season</i>]_position.csv<ul><li>Dimensions:&nbsp;[lons, models]</li><li>Units: degrees latitude</li></ul></li></ul></li></ul><p><i><strong>seasons</strong></i> are:</p><ul><li>ANN:&nbsp;Annual mean</li><li>DJF: December-January-February (Summer)</li><li>MAM: March-April-May (Autumn)</li><li>JJA: June-July-August (Winter)</li><li>SON: September-October-November (Spring)</li></ul><p><i><strong>areas</strong></i> are:</p><ul><li>Ant:&nbsp;All Antarctica</li><li>AP: Antarctic Peninsula</li></ul><p><strong>Note:</strong> CEMIPT6 models include a column with CEMIP6 model average</p><p><strong>Version control</strong></p><p>v1.0 - Initial version<br>v1.1 - Change ERA5 dataset calculations from preliminary version of ERA5 to final version of ERA5</p><p>&nbsp;</p><p><strong>How to cite</strong></p><p>If you use this dataset, please cite the accompanying paper as:</p><p>&nbsp;</p><p><strong>Complementary code</strong></p><p>You can find the jupyter notebooks to complement the research in:&nbsp;<a href="https://doi.org/10.5281/zenodo.10063849">https://doi.org/10.5281/zenodo.10063849</a></p><p>&nbsp;</p><p><strong>Contact</strong></p><p>If you have any question, please contact with Sergi at&nbsp;<a href="mailto:sergi.gonzalez@slf.ch">sergi.gonzalez@slf.ch</a></p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Historical Travel and Communications in Finland

<p>This dataset contains a proof-of-concept GIS database of over 29,000 individual historical road polyline segments as a shapefile dataset, covering over 11,000 km<sup>2</sup>&nbsp;in the western Finland from the city of Turku to northern parts of the province of Satakunta. These polylines capture the regional layout of the overland transport infrastructure of late nineteenth and early twentieth century Finland.</p>

opencc-by-4.0Sep 2023View details →
zenodo48/100

Dataset to "Hygrothermal performance of an internally insulated masonry wall: experimentations without vapour barrier in a historic Italian Palazzo"

<p>This record contains pre-processed data of a ten-and-a-half-month monitoring period of the HeLLo project.</p> <p>The datafiles titled MeasProcessed_YYYY-MM-DD.dat correspond to the prepared data into a form for data analysis and processing as presented in &ldquo;Hygrothermal performance of an internally insulated masonry wall: experimentations without vapour barrier in a historic Italian Palazzo&rdquo;, accepted for publication in journal energy and buildings (<a href="https://doi.org/10.1016/j.enbuild.2022.111896">https://doi.org/10.1016/j.enbuild.2022.111896</a>).</p> <p>Each file, format MeasProcessed _YYYY-MM-DD.dat, corresponds to the daily registered data monitored every minute.</p> <p>Each file, format MeasProcessed_YYYY-MM-DD.dat, contains temperature (T) and relative humidity (RH) values, monitored through T-RH sensors (Telaire T9602; Amphenol). The general architecture of the acquisition system is based on a Master Slave configuration, as described in &ldquo;Development of a Compatible, Low Cost and High Accurate Conservation Remote Sensing Technology for the Hygrothermal Assessment of Historic Walls&rdquo; (doi:10.3390/electronics8060643).</p> <p>Each file, format MeasProcessed_YYYY-MM-DD.dat is a text-based DAT file and can be opened with a standard text editor.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Local Ecological Knowledge and folk medicine in historical Esthonia, Livonia, Courland and Galicia, 1805-1905

<p>Background: Historical ethnobotanical data can provide valuable information about past human-nature relationships as well as serve as a basis for diachronic analysis. This thesis aims to document medicinal plant uses in the 19th century mentioned in German-language sources in the historical regions of Esthonia, Livonia, Courland and Galicia to analyse the gathered data in regard to plant families and medicinal use categories and finally to qualitatively compare the results with various studies from the study area and surrounding regions with recently acquired data as well as historical data.</p> <p>Methods: Data was mainly obtained by systematic manual search in various relevant historical German-language works focused on the medicinal use of plants. Data about plant and non-plant constituents, their usage, the mode of administration, used plant parts and their German and local names was extracted and collected into a database in the form of Use Reports.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Historical Sea Surface Temperature (SST) data and thermal stress indices of the Tara Pacific Expedition's coral reef sampling sites, from May 1st 2002 to August 31st 2018.

<p>The Tara Pacific expedition (2016-2018) sampled coral ecosystems at 111 sampling sites around 32 islands in the Pacific Ocean, and sampled the surface of oceanic waters at 249 locations, resulting in the collection of nearly 58,000 samples (Gorsky et al. 2019, Planes et al. 2019, Flores et al. 2020). The expedition was designed to systematically study corals, fish, plankton, and seawater, and included the collection of samples for advanced biogeochemical, molecular, and imaging analysis.</p> <p>Here we provide a high-resolution historical dataset that spans from 2002 to each sites&rsquo; sampling date and gives an overview of past climate variability and heatwaves experienced by corals sampled at each site. Ocean skin temperature (11 and 12 &micro;m spectral bands longwave algorithm) was extracted from 1km resolution level-2 MODIS-Aqua and MODIS-Terra from 2002 to the sampling date and from level-2 VIIRS-SNPP from 2012 to the sampling date. Day and night overpasses were used to maximize data recovery. Following recommendations from NASA Ocean Color (OB.DAAC), only SST products of quality 0 and 1 were used. The 9 closest pixels to the sampling sites of each scene were extracted. All the extracted pixels from the 3 satellites were then averaged daily to obtain daily SST averages and standard deviations time series for each sampling site, from 2002 to the sampling date.</p> <p>Each time series was first averaged on a Julian day basis to provide a seasonal average. This yearly seasonal average was triplicated and concatenated into a 3-year seasonal cycle to apply a digital low pass filter on the middle year without generating artifacts. A digital low pass filter (filter order 3, pass band ripple 0.1; &ldquo;filfilt&rdquo; function in matlab) with 36 Julian days windows was applied to the concatenated time series to remove high frequency noise. The middle year was then extracted from the concatenated time series to recover the seasonal cycle. The sea surface temperature anomaly was calculated as the SST minus the seasonal cycle over the full time series. Considering the short periods of missing data (mean of the 95th percentile of the duration of consecutive days with missing data: 9.8 &plusmn; 4.1 days), the missing values in the SST and SST anomaly time series were linearly interpolated in order to calculate thermal stress indices. The SST anomaly frequency was calculated as the number of days over the past 52 weeks when the SST anomaly is greater than or equal to 1 &deg;C. Thermal stress indices relevant to coral reef health were then calculated using methodology developed for the Coral Reef Temperature Anomaly Database (CoRTAD) data base (Saha et al. 2019). Events of cold temperature accumulation were also reported to cause bleaching and mortality (Lirman et al. 2011; Gonz&aacute;lez-Espinosa &amp; Donner 2020), therefore, the same set of indices were calculated for cold stress adapting the CoRTAD method, but using the minimum weekly climatologies.</p> <p>A condensed table containing single values associated with each sampling site was created (&#39;TaraPacific_SST_timeseries_mean_products&#39;) extracting the minimum, maximum, sum, averages, standard deviations, and value recorded at the sampling day of each of these indices (detailed in the readme file provided with the dataset &#39;README_TaraPacific_historical_SST.md&#39;). Additional metrics of the last heating and cooling events as well as the time of recovery is also provided to represent the state of thermal stress at the day of sampling.</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

Source Code Accompanying the Paper "More on network approaches in Historical Chinese Phonology (音韻學)"

<p>First version of the source code and data accompanying the paper &quot;More on Network Approaches in Historical Chinese Phonology&quot;.</p> <p>This paper is available here:</p> <ul> <li>List, Johann-Mattis (2018): <strong>More on network approaches in Historical Chinese Phonology (音韻學)</strong>. Paper prepared for the <em>LFK Society Young Scholars Symposium</em>. Taibei: Li Fang-Kuei Society ofr Chinese Linguistics. URL: <a href="https://hal.archives-ouvertes.fr/hal-01706927">https://hal.archives-ouvertes.fr/hal-01706927</a>.</li> </ul> <pre><code>@InProceedings{List2018a, author = {List, Johann-Mattis}, title = {{More on Network Approaches in Historical Chinese Phonology (音韻學)}}, booktitle = {{LFK Society Young Scholars Symposium}}, year = {2018}, publisher = {Li Fang-Kuei Society for Chinese Linguistics}, pdf = {https://hal.archives-ouvertes.fr/hal-01706927/file/main.pdf}, url = {https://hal.archives-ouvertes.fr/hal-01706927}, address = {Taipei}, hal_id = {hal-01706927}, } </code></pre> <p>See the README.md for mor information.</p> <ul> <li>&nbsp;</li> </ul>

opencc-by-4.0Feb 2018View details →
zenodo48/100

HANZE database of historical flood impacts in Europe, 1870-2025

<p>The HANZE dataset covers riverine, pluvial, coastal and compound floods that have occurred in 42 European countries between 1870 and 31 March 2025. The data was collected by extensive data-collection from more than 1000 sources ranging from news reports through government databases to scientific papers. The dataset includes 2687 events characterized by at least one impact statistic: area inundated, fatalities, persons affected or economic loss. Economic losses are presented both in the original currencies and price levels as well as inflation and exchange-rate adjusted to 2024 value of the euro. The spatial footprint of affected areas is consistently recorded using more than 1400 subnational units corresponding, with minor exceptions, to the European Union&rsquo;s Nomenclature of Territorial Units for Statistics (NUTS), level 3. Daily start and end dates, information on causes of the event, notes on data quality issues or associated non-flood impacts, and full bibliography of each record supplement the dataset. Apart from the possibility to download the data, the database can be viewed, filtered and visualized online: <a href="https://naturalhazards.eu">https://naturalhazards.eu</a>. The dataset is designed to be complimentary to HANZE-Exposure, a high-resolution model of historical exposure changes (such as population and asset value), and be easily usable in statistical and spatial analyses.</p> <p><strong>This is a preliminary update of HANZE v2.1, adding 169 floods for years 2021-2025 (until 31 March 2025). It makes only minor revisions to previous data (adds 11 pre-2021 events, revises 17 records and removes 3 events that were newly reassessed as non-flood events). A more extensive revision of the data is planned for 2026.</strong></p> <p>The dataset contains the following files (CSV comma-delimited, UTF8, and ESRI shapefiles in zipped folders)</p> <p><strong>HANZE flood events database&nbsp;&nbsp;&nbsp;</strong></p> <p>HANZE_events.csv - Flood event data</p> <p>HANZE_references.csv - List of all references</p> <p>HANZE3_events_regions_2010.zip - Flood event data as GIS file (regions v2010)</p> <p>HANZE3_events_regions_2021.zip - Flood event data as GIS file (regions v2021)</p> <p>HANZE3_events_regions_2021.zip - Flood event data as GIS file (regions v2021)</p> <p><strong>Supplementary data&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</strong></p> <p>S1_countries_codes_and_names.csv - Country codes/names</p> <p>S2_regions_codes_and_names_v2010.csv - Region codes/names, v2010</p> <p>S3_regions_codes_and_names_v2021.csv - Region codes/names, v2021</p> <p>S3a_regions_codes_and_names_v2024.csv - Region codes/names, v2024</p> <p>S4_list_of_all_currencies_by_country.csv - Data on all currencies used in the study area since 1870</p> <p>S5_currency_conversion_rates.csv - Conversion rates applied to compute losses in 2024 euros</p> <p>S6_GDP_deflators_by_country.csv - Gross domestic product deflator by country, 1870-2025</p> <p>S7_floods_removed_from_HANZE.csv - Flood events in HANZE v1 and v2, which were excluded from v3</p> <p>Regions_v2010_simplified.zip - Map of subnational regions used in the database, v2010</p> <p>Regions_v2021_simplified.zip - Map of subnational regions used in the database, v2021</p> <p>Regions_v2024_simplified.zip - Map of subnational regions used in the database, v2024</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Historical and modelled renewable energy production for India

<p>This archive contains all the datasets produced for the paper:<br><br><span>Hunt,&nbsp;K. M. R.</span>, &amp;&nbsp;<span>Bloomfield,&nbsp;H. C.</span>&nbsp;(<span>2024</span>).&nbsp;<span>Quantifying renewable energy potential and realized capacity in India: Opportunities and challenges</span>.&nbsp;<em>Meteorological Applications</em>,&nbsp;<span>31</span>(<span>3</span>), e2196.&nbsp;<a href="https://doi.org/10.1002/met.2196">https://doi.org/10.1002/met.2196</a></p> <p>&nbsp;</p> <table style="border-collapse: collapse; width: 99.9642%;"><colgroup><col style="width: 31.0476%;"><col style="width: 17.1785%;"><col style="width: 37.7477%;"><col style="width: 14.0133%;"></colgroup> <tbody> <tr> <td><strong>Data Description&nbsp;</strong></td> <td><strong>Figure/Table</strong></td> <td><strong>&nbsp;File Name</strong></td> <td><strong>Dates Valid</strong></td> </tr> <tr> <td>Installed capacity by type in each state</td> <td>Table 1</td> <td>installed-by-state-oct2022.csv</td> <td>Oct 2022</td> </tr> <tr> <td>All-India installed capacity by type</td> <td>Figure 2</td> <td>tabulated-installed-by-date.csv</td> <td>2017&ndash;2023</td> </tr> <tr> <td>Hourly wind capacity factor</td> <td>Figure 4</td> <td>wind capacity factor.zip</td> <td>1979&ndash;2022</td> </tr> <tr> <td>Hourly solar capacity factor</td> <td>Figure 6</td> <td>solar capacity factor.zip&nbsp;</td> <td>1979&ndash;2022</td> </tr> <tr> <td>Present-day installation locations</td> <td>Figure 11</td> <td>OSM[hydropower,wind_turbine,solar]_ installations.geojson</td> <td>Mar 2022</td> </tr> <tr> <td>Gridded 1&deg;&times;1&deg; estimate of installed wind/solar capacity</td> <td>Figure 12a/13a</td> <td>CEA_1x1_gridded_installed_[wind,solar]_cap.nc</td> <td>May 2021</td> </tr> <tr> <td>Gridded 1&deg;&times;1&deg; estimate of installed wind capacity</td> <td>Figure 12b</td> <td>TWP_1x1_gridded_installed_wind_cap.nc</td> <td>May 2021</td> </tr> <tr> <td>Gridded 1&deg;&times;1&deg; estimate of installed solar capacity</td> <td>Figure 13b</td> <td>K21_1x1_gridded installed solar cap.nc</td> <td>Sep 2018</td> </tr> <tr> <td>Reported daily wind/solar/hydro production</td> <td>Figure 14/S3</td> <td>POSOCO_reported_[wind,solar,hydro]_MU_ daily.csv</td> <td>2012&ndash;2023</td> </tr> <tr> <td>Modelled &lsquo;historical&rsquo; production</td> <td>Figure 14/16/S4a/b</td> <td>modelled-historical-[daily,hourly]-renewable output.nc</td> <td>1979&ndash;2022</td> </tr> <tr> <td>Recommended locations for new wind/solar installations</td> <td>Figure 17</td> <td>areas-for-exploration.nc</td> <td>--</td> </tr> </tbody> </table> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Digitalised dry heathlands on historical Ferraris Map of Flanders

<p>Dataset description</p> <p>Dataset contains GIS vector layers with the polygons of regions, in the Esri shapefile format. Polygons have been drawn manually using QGIS software according to the borders of paths and roads, capturing schematic patterns of dry heathland from De Coene et al. (2012).<br> CRS: EPSG:3857 - WGS 84 / Pseudo-Mercator<br> Charset Encoding: UTF-8</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Bibliographic Data from the Computational Methods Applied to Earthen Historical Structures Review

<p>This database contains all the&nbsp;bibliographic&nbsp;information about the 293 records found after applying the Search Strategy used for the&nbsp;Computational Methods Applied to Earthen Historical Structures Review.&nbsp;Such strategy consisted on using relevant keywords grouped into three different search queries within &rdquo;TITLE-ABS-KEY&rdquo;, for the years 2019-2023:</p> <ol> <li>(&rdquo;earthen heritage&rdquo; OR &rdquo;earthen historical building*&rdquo; OR &rdquo;earthen historical structure*&rdquo; OR &rdquo;earthen&nbsp;architect*&rdquo; OR &rdquo;earthen monument*&rdquo;).</li> <li>(adobe OR &rdquo;rammed earth&rdquo; OR cob ) AND (&rdquo;computational method*&rdquo; OR &rdquo;numerical analy*&rdquo;).</li> <li>(adobe OR &rdquo;rammed earth&rdquo; OR cob ) AND (fem OR dem OR la OR &rdquo;finite element&rdquo; OR &rdquo;discrete&nbsp;element&rdquo; OR &rdquo;limit analysis&rdquo;).</li> </ol> <p>The search was conducted on April 7, 2023.</p>

opencc-by-4.0May 2023View details →
edi48/100

Data in support of Primack et al. 2022 Frontiers in Ecology & Environment: Historically excluded groups in ecology are undervalued and poorly treated

Hostile workplaces undermine efforts to make the ecological sciences more inclusive and welcoming. A survey sent to the Ecological Society of America membership and ECOLOG-L listserv subscribers provides a snapshot of a range of workplace experiences in ecology. The results of this survey are published as Primack et al. 2022. Historically excluded groups in ecology are undervalued and poorly treated. Frontiers in Ecology and the Environment. This dataset includes the survey results and code for data analysis.

openCC (other)Oct 2022View details →
edi48/100

Historical and future Lake Surface Water Temperature for 80 major lakes in Southeast Asia [LSWT-SEA]

The present dataset is part of a study delving into the intricate relationship between lake surface temperature (LSWT) and the broader context of climate change in the ecologically diverse region of Southeast Asia (SEA). Recognizing LSWT as a highly responsive indicator of climatic shifts, the research aims to shed light on the region's vulnerability to these changes. Using a suite of predictive models (namely Multilinear Regression (MLR), Multilayer perceptron (MLP), Random Forest (RF), eXtreme Gradient Boosting (XGB), Multilayer perceptron (MLP)) the study reconstructs historical LSWT trends from 1986 to 2020 and projects future scenarios until 2100, contingent upon various Representative Concentration Pathway (RCP) trajectories. Using MODIS-derived LSWT as predicted variable. The dataset package includes the data used to carry out the research: ECMWF ERA5 and CHIRPS climatic predicting variables, MODIS-derived daytime and nighttime LSWT, historically predicted daily daytime and nighttime LSWT, future predictions of LSWT for multiple Representative Concentration Pathways (RCPs), long term historical and future trends.

openCC (other)Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record