Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
647
datasets available to search
ShareScore release 0.9.0
Dataset results
647 results for “historical_data”
Historical GIS Data for Harvard Forest Properties from 1908 to Present
Since 1908, the Harvard Forest has conducted forest surveys approximately every 10-20 years on its three largest tracts (total 1033 ha). These maps have been digitized along with maps of environmental factors (topography, soils), disturbance (1938 hurricane, historical land-use), and silvicultural treatments. These datalayers will allow researchers to understand the influence of environment factors, disturbances, and silviculture on the structure and composition of modern forest stands as well as assisting in locating and describing research sites. The dataset also includes an elevation grid (NED 30 meter cells), and a shapefile of linear features (trails, stonewalls, etc). Original maps were transcribed to standardized basemaps by various researchers. These basemaps were then scanned and digitized as shapefiles in ArcView GIS 3.2. The shapefiles were then transformed to Massachusetts State Plane Meters NAD83 projection in ArcGIS and rubbersheeted to align better with aerial photographs downloaded from MassGIS. Locations of control points will be permanently archived at the Harvard Forest to facilitate transformation of future datalayers.
Massachusetts Historical Landcover and Census Data 1640-1999
An appreciation of historical landuse and its effects is crucial when interpreting the structure, composition, and spatial characteristics of modern forests. The Harvard Forest has compiled many different historical data sources in an ongoing effort to understand how anthropogenic disturbances have shaped our modern landscapes. Estimates of town land use and land cover were gathered from a variety of sources, including tax valuations (1801-1860) and state agricultural census records (1865-1905). Data prior to 1801 rarely cover the entire state and are excluded from these datasets. Data on forest structure are available for several time periods, including 1885 and 1895 (Agricultural Censuses) and 1916-1920s (State Forester’s reports).
Historical and Ecological GIS Data from Manuel F. Correllus State Forest on Martha’s Vineyard 1830-1994
Sand-plain ecosystems are a priority for conservation because they are uncommon, support numerous rare or uncommon plant and animal species, serve as groundwater recharge areas, and are threatened by land development. The 5,200-acre Manuel F. Correllus State Forest, in the central part of Martha’s Vineyard, is part of one of the larger sand-plain ecosystems in New England. This GIS data package was created as part of a study on the history and ecology of Martha’s Vineyard and the state forest as part of an effort to understand sandplain landscapes and make management recommendations for their maintenance.
Historical GIS Data for Prospect Hill Tract at Harvard Forest 1733-1986
This dataset contains elevation, 1986 forest type, land-use history, and soils maps for the Prospect Hill Tract, digitized from paper maps in the Harvard Forest Archives. File format = Idrisi 4.1 binary. Resolution = 10m x 10m. Coordinates = UTM zone 18. Datum = 1927 North American. This dataset has been replaced with a new vector series for the entire Harvard Forest (see HF110).
Historical and future land use and land cover data for the STARS4Water river basins
<p>Dataset contains data on historical and future land use and land cover for seven European river basins (Danube, Drammen, Duero, East Anglia, Messara, Rhine and Seine) being case study basin in the STARS4Water, and a shapefile with river basin boundaries. The average area fraction of five general land use classes (crop, forest, grass, urban and other) within the project river basins was calculated at five-year intervals starting in 2016 and ending in 2051. This dataset was prepared based on the data available in the "LUCAS LUC future land use and land cover change dataset for Europe (Version 1.1)" repository (Hoffmann et al., 2022, DOI: 10.26050/WDCC/LUC_future_EU_v1.1).</p>
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - SARI SET Surface Water level data from Salt River Bay National Historical Park and Ecological Preserve, St. Croix, US Virgin Islands.
Surface water level data (m) was collected in Salt River Bay National Historic Park and Ecological Preserve (SARI) by the South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2017 to 2024 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This data-package is complete.
Models for "A data-driven approach to studying changing vocabularies in historical newspaper collections"
<p>NOTE: This is a badly rendered version of the README within the archive.</p> <p><strong>A data-driven approach to studying changing vocabularies in historical newspaper collections</strong></p> <p>Simon Hengchen,* Ruben Ros,** Jani Marjanen,*** Mikko Tolonen***</p> <p>*<a href="https://spraakbanken.gu.se/en/about/staff/simon">Språkbanken Text</a>, University of Gothenburg, Sweden and <a href="https://iguanodon.ai">iguanodon.ai</a>, Belgium: firstname.lastname@gu.se<br> **<a href="https://www.c2dh.uni.lu/people/ruben-ros">Centre for Contemporary and Digital History (C2DH)</a>, University of Luxembourg: firstname.lastname@uni.lu<br> ***<a href="https://www.helsinki.fi/en/researchgroups/computational-history">COMHIS</a>, University of Helsinki: <a href="mailto:firstname.lastname@helsinki.fi">firstname.lastname@helsinki.fi</a>;</p> <p>These are the supplementary materials for the DH2019 paper <em>A data-driven approach to the changing vocabulary of the ‘nation’ in English, Dutch, Swedish and Finnish newspapers, 1750-1950</em>, as well as the 2021 Digital Scholarship in the Humanities publication available in OpenAccess: <a href="https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793">https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793</a>. If you end up using whole or parts of this resource, please use the following citation(s):</p> <ul> <li>Hengchen, S., Ros, R., and Marjanen, J. (2019). A data-driven approach to the changing vocabulary of the 'nation' in English, Dutch, Swedish and Finnish newspapers, 1750-1950. In <em>Proceedings of the Digital Humanities (DH) conference 2019, Utrecht, The Netherlands</em></li> </ul> <p>and/or:</p> <ul> <li>Hengchen, S., Ros, R., Marjanen, J. and Tolonen, M., 2021. A data-driven approach to studying changing vocabularies in historical newspaper collections. Digital Scholarship in the Humanities, 36(Supplement_2), pp.ii109-ii126.</li> </ul> <p>or alternatively use one of the following <code>bib</code>s:</p> <pre><code>@inproceedings{hengchen2019nation, title="A data-driven approach to the changing vocabulary of the 'nation' in {E}nglish, {D}utch, {S}wedish and {F}innish newspapers, 1750-1950.", author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani}, year={2019}, address = "Utrecht, The Netherlands", booktitle={Proceedings of the Digital Humanities (DH) conference 2019} }</code></pre> <pre><code>@article{hengchen2021data, title={A data-driven approach to studying changing vocabularies in historical newspaper collections}, author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani and Tolonen, Mikko}, journal={Digital Scholarship in the Humanities}, volume={36}, number={Supplement\_2}, pages={ii109--ii126}, year={2021}, publisher={Oxford University Press} }</code></pre> <p> </p> <p>Files</p> <p>This archive contains two folders -- one per diachronic representation method -- as well as this README. The folders each contain four folders, which contain the models for their respective languages. As can be inferred from the small datasize, most of the earlier models are not reliable and should not be used, but are still made available. This work is licensed under a <a href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p><strong>Source material</strong></p> <p>Finnish:</p> <p>The models were created with data from the Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland (National Library of Finland, 2011). We used everything in the corpus.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h fi* 12M fi_1820_SGNS_corpus_file.gensim 89M fi_1840_SGNS_corpus_file.gensim 797M fi_1860_SGNS_corpus_file.gensim 7.0G fi_1880_SGNS_corpus_file.gensim 22G fi_1900_SGNS_corpus_file.gensim</code></pre> <p>Swedish:</p> <p>The models were created with data from the Kubhist 2 corpus (Språkbanken) -- more precisely, the data dumps available at <a href="https://spraakbanken.gu.se/lb/resurser/meningsmangder/">https://spraakbanken.gu.se</a>. After a manual evaluation of Swedish embeddings trained without pre-processing seemed to show that the embeddings were of low quality, we retrained models, only keeping sentences that were at least 10 tokens long and were constituted of at least 50% of lemmas as per the KORP processing pipeline (Borin et al, 2012).</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h sv* 1.6M sv_1740_SGNS_corpus_file.gensim 44M sv_1760_SGNS_corpus_file.gensim 124M sv_1780_SGNS_corpus_file.gensim 228M sv_1800_SGNS_corpus_file.gensim 678M sv_1820_SGNS_corpus_file.gensim 1.6G sv_1840_SGNS_corpus_file.gensim 4.5G sv_1860_SGNS_corpus_file.gensim 6.5G sv_1880_SGNS_corpus_file.gensim 113M sv_1900_SGNS_corpus_file.gensim</code></pre> <p>Dutch:</p> <p>The models were created with data from the Delpher newspaper archive (Royal Dutch Library, 2017), through data dumps for newspapers until and including 1876, and through API hits for articles from 1877 to 1899 (included).</p> <ul> <li>For anything pre-1877 we discarded full texts that had, in the metadata, anything else than exclusively <code>nl</code> or <code>NL</code> as a language tag.</li> <li>For the full texts between 1877 and 1899: we queried the API for all items in the “artikel” category that contained the determiner <code>de</code>.</li> </ul> <p>Our assumption was that most articles should contain <code>de</code> at least once, and those that didn't were too short to be deemed interesting. A subsequent study showed that was not exactly the case, but we were reassured by the fact that left-out articles were probably "shipping or financial reports" (thanks go to Melvin Wevers). We also did not include the colonial newspapers for our embeddings. This is motivated by our research questions. A list of removed newspapers is available on request.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h nl* 6.8M nl_1620_SGNS_corpus_file.gensim 7.9M nl_1640_SGNS_corpus_file.gensim 43M nl_1660_SGNS_corpus_file.gensim 78M nl_1680_SGNS_corpus_file.gensim 138M nl_1700_SGNS_corpus_file.gensim 243M nl_1720_SGNS_corpus_file.gensim 287M nl_1740_SGNS_corpus_file.gensim 431M nl_1760_SGNS_corpus_file.gensim 825M nl_1780_SGNS_corpus_file.gensim 1.2G nl_1800_SGNS_corpus_file.gensim 1.8G nl_1820_SGNS_corpus_file.gensim 3.1G nl_1840_SGNS_corpus_file.gensim 5.2G nl_1860_SGNS_corpus_file.gensim 13G nl_1880_SGNS_corpus_file.gensim</code></pre> <p>English:</p> <p>The models were created with data from the British Library Newspapers collection (<a href="https://www.gale.com/intl/primary-sources/british-library-newspapers%5D">link</a>), the Nichols collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-burney-newspapers-collection">link</a>), and the Burney collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-nichols-newspapers-collection">link</a>). We used everything in the corpora. For English, only SGNS_ALIGN models are available. We thank Gale Cengage for their help with this project.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h en* 4.3M en_1620_SGNS_corpus_file.gensim 11M en_1640_SGNS_corpus_file.gensim 11M en_1660_SGNS_corpus_file.gensim 106M en_1680_SGNS_corpus_file.gensim 409M en_1700_SGNS_corpus_file.gensim 1.7G en_1720_SGNS_corpus_file.gensim 834M en_1740_SGNS_corpus_file.gensim 2.4G en_1760_SGNS_corpus_file.gensim 5.3G en_1780_SGNS_corpus_file.gensim 5.5G en_1800_SGNS_corpus_file.gensim 15G en_1820_SGNS_corpus_file.gensim 42G en_1840_SGNS_corpus_file.gensim 65G en_1860_SGNS_corpus_file.gensim 88G en_1880_SGNS_corpus_file.gensim 26G en_1900_SGNS_corpus_file.gensim 21G en_1920_SGNS_corpus_file.gensim 6.3G en_1940_SGNS_corpus_file.gensim</code></pre> <p><strong>Word embeddings</strong></p> <p>For every language, we train diachronic embeddings as follows. We divide the data in 20-year time bins. We train SGNS_UPDATE and SGNS_ALIGN models. Current research on German (Schlechtweg et al, 2019) and English (Shoemark et al, 2019) indicates you should use the SGNS_ALIGN models. <strong>For EN, FI, NL, no tokens (including punctuation) were removed nor altered, aside from lowercasing</strong>. For SV, see above. Parameters are as follows: SGNS architecture (Mikolov et al 2013), window size of 5, frequency threshold of 100, 5 epochs, 300 dimensions (or 100 for EN).</p> <ul> <li>For SGNS_UPDATE: We first train a model for the first time bin <code>t</code>. To train the model for <code>t+1</code>, we use the <code>t</code> model to initialise the vectors for <code>t+1</code>, set the learning rate to correspond to the end learning rate of <code>t</code>, and continue training. This approach, closely following Kim et al (2014), has the advantage of avoiding the need for post-training vector space alignment.</li> </ul> <p>The Python snippet below, which makes use of gensim (Rehurek and Sojka, 2010), illustrates the approach. Special thanks go to Sara Budts.</p> <pre><code>## dict_files[key] is a dictionary with double decades as keys and a corresponding LineSentence object as value: https://radimrehurek.com/gensim/models/word2vec.html#gensim.models.word2vec.LineSentence count = 0 for key in sorted(list(dict_files.keys())): if count == 0: ## This is the first model. model = gensim.models.Word2Vec(corpus_file=dict_files[key], min_count=100, sg=1 ,size=300, workers=64, seed=1830, iter=5) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) print("Model saved, on to the next\n") count += 1 if count > 0: ## this is for the subsequent models. print("model for double decade starting in",str(key)) model = gensim.models.Word2Vec.load(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin-20)+".w2v")) print("previous model loaded") model.build_vocab(corpus_file=dict_files[key], update=True) model.train(corpus_file=dict_files[key], total_words = model.corpus_count, total_examples = model.corpus_count, start_alpha = model.alpha, end_alpha = model.min_alpha, epochs=model.epochs) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) </code></pre> <ul> <li>For SGNS_ALIGN: We independently train models for all time bins. The models in this repository are <em>NOT</em> aligned, leaving you the choice of how to align them. For example, <a href="https://gist.github.com/quadrismegistus/09a93e219a6ffc4f216fb85235535faf">here</a> is a link to code by Ryan Heuser to do just that. Models were trained with the <code>count == 0</code> scenario in the snippet above.</li> </ul> <p><strong>Acknowledgments</strong></p> <p>This work has been supported by the European Union's Horizon 2020 research and innovation programme under grant 770299 <a href="https://www.newseye.eu/">NewsEye</a>. Specials thanks go to the data providers/collection-holding institutions: the Finnish Language Bank, the Swedish Language Bank, the Royal Dutch Library, and Gale Cengage.</p> <p>The authors would like to thank the following persons and group, listed alphabetically: Antoine Doucet, Antti Kanner, Axel-Jean Caurant, Dominik Schlechtweg, Eetu Mäkelä, Elaine Zosa, Estelle Bunout, Haim Dubossarsky, Joris van Eijnatten, Krister Lindén, Lars Borin, Lidia Pivovarova, Melvin Wevers, Nina Tahmasebi, Sara Budts, Senka Drobac, Tanja Säily, the COMHIS group, and Steven Claeyssens. Computational resources were provided by CSC – IT Center for Science Ltd.</p> <p><strong>References</strong></p> <p>Borin, L., Forsberg, M., Roxendal, J. (2012). Korp-the corpus infrastructure of Spräkbanken,in: LREC. pp. 474–478.</p> <p>Kim, Y., Chiu, Y.I., Hanaki, K., Hegde, D. and Petrov, S. (2014). Temporal Analysis of Language through Neural Language Models. <em>ACL 2014</em>, p.61.</p> <p>Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space. <em>arXiv preprint arXiv:1301.3781</em>.</p> <p>National Library of Finland (2011). <em>The Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland, Kielipankki Version</em> [text corpus]. Kielipankki. Retrieved from <a href="http://urn.fi/urn:nbn:fi:lb-2016050302">http://urn.fi/urn:nbn:fi:lb-2016050302</a>.</p> <p>Rehurek, R. and Sojka, P. (2010). Software framework for topic modelling with large corpora. In <em>Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</em>.</p> <p>Royal Dutch Library (2017). <em>Delpher open krantenarchief (1.0)</em>. Den Haag, 2017.</p> <p>Schlechtweg D., Hätty A, del Tredici M., and Schulte im Walde S. (2019). A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In <em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, Florence, Italy. ACL.</p> <p>Shoemark, P., Liza, F.F., Nguyen, D., Hale, S. and McGillivray, B. (2019). Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings. In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 66-76)</em>, Hong Kong.</p> <p>Språkbanken. <em>The Kubhist Corpus</em>. Department of Swedish, University of Gothenburg. <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">https://spraakbanken.gu.se/korp/?mode=kubhist</a>.</p>
C3-EURO4M-MEDARE Mediterranean historical climate data - v.2
<p>Historical surface climate data files and meta-data for stations in Mediterranean North Africa and Middle East areas (1852-2008).</p>
Historical Weather, Load, Wind, and Solar Data for the Salt River Project
<p>We created and curated a dataset of historical (1980-2019) hourly meteorology, load, wind, and solar data for the Salt River Project (SRP) region. The data was created by PNNL's <a href="https://godeeep.pnnl.gov/">GODEEEP</a> project. Each row in the dataset is a single hour and each column is a variable. All meteorological variables are spatially-averaged over the SRP service territory. The variables and their units are as follows:</p><ol><li>"Time_UTC"; Coordinated Universal Time (UTC); Time of day.</li><li>"T2"; Fahrenheit; 2-m air temperature.</li><li>"Q2"; kg/kg; 2-m water vapor mixing ratio.</li><li>"SWDOWN"; W/m^2; Downwelling shortwave radiative flux at the surface.</li><li>"GLW"; W/m^2; Downwelling longwave radiative flux at the surface.</li><li>"WSPD"; m/s; 10-m wind speed.</li><li>"Scaled_2019_Load"; MWh; Simulated hourly demand for electricity that is scaled to 2019 levels of annual energy. This load estimate does not account for historical changes in population and economics within the SRP service territory. It is included to make it easier to isolate weather impacts on load without having to consider long-term changes.</li><li>"Load"; MWh; Simulated hourly demand for electricity.</li><li>"Agua_Fria_Solar_Capacity"; N/A; Solar capacity factor for the SRP Agua Fria project with plant configurations taken from the EIA-860 database.</li><li>"Phoenix_Solar_Capacity"; N/A; Solar capacity factor for hypothetical solar plants derived using the grid cell nearest to Phoenix, AZ.</li><li>"Flagstaff_Solar_Capacity"; N/A; Solar capacity factor for hypothetical solar plants derived using the grid cell nearest to Flagstaff, AZ.</li><li>"Phoenix_Wind_Capacity"; N/A; Wind capacity factor for hypothetical 80-m plants derived using the grid cell nearest to Phoenix, AZ.</li><li>"Flagstaff_Wind_Capacity"; N/A; Wind capacity factor for hypothetical 80-m plants derived using the grid cell nearest to Flagstaff, AZ.</li></ol>
Historical Sea Surface Temperature (SST) data and thermal stress indices of the Tara Pacific Expedition's coral reef sampling sites, from May 1st 2002 to August 31st 2018.
<p>The Tara Pacific expedition (2016-2018) sampled coral ecosystems at 111 sampling sites around 32 islands in the Pacific Ocean, and sampled the surface of oceanic waters at 249 locations, resulting in the collection of nearly 58,000 samples (Gorsky et al. 2019, Planes et al. 2019, Flores et al. 2020). The expedition was designed to systematically study corals, fish, plankton, and seawater, and included the collection of samples for advanced biogeochemical, molecular, and imaging analysis.</p> <p>Here we provide a high-resolution historical dataset that spans from 2002 to each sites’ sampling date and gives an overview of past climate variability and heatwaves experienced by corals sampled at each site. Ocean skin temperature (11 and 12 µm spectral bands longwave algorithm) was extracted from 1km resolution level-2 MODIS-Aqua and MODIS-Terra from 2002 to the sampling date and from level-2 VIIRS-SNPP from 2012 to the sampling date. Day and night overpasses were used to maximize data recovery. Following recommendations from NASA Ocean Color (OB.DAAC), only SST products of quality 0 and 1 were used. The 9 closest pixels to the sampling sites of each scene were extracted. All the extracted pixels from the 3 satellites were then averaged daily to obtain daily SST averages and standard deviations time series for each sampling site, from 2002 to the sampling date.</p> <p>Each time series was first averaged on a Julian day basis to provide a seasonal average. This yearly seasonal average was triplicated and concatenated into a 3-year seasonal cycle to apply a digital low pass filter on the middle year without generating artifacts. A digital low pass filter (filter order 3, pass band ripple 0.1; “filfilt” function in matlab) with 36 Julian days windows was applied to the concatenated time series to remove high frequency noise. The middle year was then extracted from the concatenated time series to recover the seasonal cycle. The sea surface temperature anomaly was calculated as the SST minus the seasonal cycle over the full time series. Considering the short periods of missing data (mean of the 95th percentile of the duration of consecutive days with missing data: 9.8 ± 4.1 days), the missing values in the SST and SST anomaly time series were linearly interpolated in order to calculate thermal stress indices. The SST anomaly frequency was calculated as the number of days over the past 52 weeks when the SST anomaly is greater than or equal to 1 °C. Thermal stress indices relevant to coral reef health were then calculated using methodology developed for the Coral Reef Temperature Anomaly Database (CoRTAD) data base (Saha et al. 2019). Events of cold temperature accumulation were also reported to cause bleaching and mortality (Lirman et al. 2011; González-Espinosa & Donner 2020), therefore, the same set of indices were calculated for cold stress adapting the CoRTAD method, but using the minimum weekly climatologies.</p> <p>A condensed table containing single values associated with each sampling site was created ('TaraPacific_SST_timeseries_mean_products') extracting the minimum, maximum, sum, averages, standard deviations, and value recorded at the sampling day of each of these indices (detailed in the readme file provided with the dataset 'README_TaraPacific_historical_SST.md'). Additional metrics of the last heating and cooling events as well as the time of recovery is also provided to represent the state of thermal stress at the day of sampling.</p>
Bibliographic Data from the Computational Methods Applied to Earthen Historical Structures Review
<p>This database contains all the bibliographic information about the 293 records found after applying the Search Strategy used for the Computational Methods Applied to Earthen Historical Structures Review. Such strategy consisted on using relevant keywords grouped into three different search queries within ”TITLE-ABS-KEY”, for the years 2019-2023:</p> <ol> <li>(”earthen heritage” OR ”earthen historical building*” OR ”earthen historical structure*” OR ”earthen architect*” OR ”earthen monument*”).</li> <li>(adobe OR ”rammed earth” OR cob ) AND (”computational method*” OR ”numerical analy*”).</li> <li>(adobe OR ”rammed earth” OR cob ) AND (fem OR dem OR la OR ”finite element” OR ”discrete element” OR ”limit analysis”).</li> </ol> <p>The search was conducted on April 7, 2023.</p>
Data in support of Primack et al. 2022 Frontiers in Ecology & Environment: Historically excluded groups in ecology are undervalued and poorly treated
Hostile workplaces undermine efforts to make the ecological sciences more inclusive and welcoming. A survey sent to the Ecological Society of America membership and ECOLOG-L listserv subscribers provides a snapshot of a range of workplace experiences in ecology. The results of this survey are published as Primack et al. 2022. Historically excluded groups in ecology are undervalued and poorly treated. Frontiers in Ecology and the Environment. This dataset includes the survey results and code for data analysis.
Datasets For "Estimating Maximum Extent of Auroral Equatorward Boundary using Historical and Simulated Surface Magnetic Field Data", Blake et al. (2020), JGR
<p>Datasets and sample Python codes for the 2020 paper <em>"Estimating Maximum Extent of Auroral Equatorward Boundary using Historical and Simulated Surface Magnetic Field Data"</em>, by Blake et al., submitted to the Journal of Gephysical Research, Space Physics. </p> <p>Up-to-date Python codes can be found at <a href="https://github.com/TerminusEst/Auroral_Boundary_Geomag">https://github.com/TerminusEst/Auroral_Boundary_Geomag</a></p> <p>The complete SWMF simulation folders (including parameter and log files etc.) can be requested from <a href="https://ccmc.gsfc.nasa.gov/index.php">NASA's Community Coordinated Modeling Center</a>.</p> <p>#########</p> <p><strong>Data/ </strong>contains the following:</p> <p><strong>Data/HIST_DATA.txt </strong>contains the minimum Dst values and calculated maximum extents of the auroral equatorward boundaries for 25 years of INTERMAGNET data (1991-2016). The fourth column is the standard deviation of the calculated auroral boundary in degrees. </p> <p><strong>Data/Boundary_Fits.csv </strong>contains the calculated minimum Dst values, and calculated auroral boundaries using Method 1 and Method 2 (see main paper's ttext), for each of the 15 SWMF simulations. Also included are the uncertainties for each calculation.</p> <p><strong>Data/SWMF_outputs/ </strong>contains 15<strong> </strong>.txt files,<strong> </strong>each of which correspond to an SWMF simulation of the same name given in Table 1 in the main text. These data are for the magnetic longitude, magnetic latitude and maximum calculated <em>E<sub>H</sub> </em>(V/km) for each simulation.</p> <p>#########</p> <p><strong>Codes/ </strong>contains two python scripts, and some sample data. These scripts correspond to Section 2 in the main text:</p> <p>1) <strong>Boundary_Calc.py</strong> calculates the extent of the auroral boundary using magnetic latitudes and maximum calculated <em>E<sub>H</sub></em> values from multiple INTERMAGNET sites. </p> <p>2) <strong>Efield_Calc.py </strong>calculates the E-field for a single INTERMAGNET site using the Quebec 1-D resistivity model.</p> <p>A more detailed description of these codes can be found here: <a href="https://github.com/TerminusEst/Auroral_Boundary_Geomag">https://github.com/TerminusEst/Auroral_Boundary_Geomag</a></p> <p> </p>
C3-EURO4M-MEDARE Mediterranean historical climate data
<p>Historical surface climate data files and meta-data for stations in Mediterranean North Africa and Middle East areas (1852-2008)</p>
Map data of historical global estimates of soil respiration
<p>The map data of global soil respiration converted to NetCDF format.</p><p>All open access available estimates were collated.</p><p>Shoji Hashimoto, Akihiko Ito, Kazuya Nishina (2023) "Divergent data-driven estimates of global soil respiration". Communications Earth & Environment, 4 Article number: 460</p><p><a href="https://doi.org/10.1038/s43247-023-01136-2 ">https://doi.org/10.1038/s43247-023-01136-2</a> </p><p>Refer to Table 1 for the study ID and data source or the attributions of the NetCDF file. </p>
Data from: Solar energy resource availability under extreme and historical wildfire smoke conditions
<p>The data in this repository are used to generate the figures in the article "Solar energy resource availability under extreme and historical wildfire smoke conditions" by Corwin et al. (accepted 2024) in <em>Nature Communications</em>. Data are the final processessed and merged datasets sourced from the following publicly available data products:</p> <ul> <li>National Renewable Energy Laboratory’s (NREL) National Solar Radiation Database (NSRDB) (<a href="https://nsrdb.nrel.gov/)">https://nsrdb.nrel.gov/)</a>. <ul> <li>Bulk download in July 2023 via AWS: <a href="https://registry.opendata.aws/nrel-pds-nsrdb/">https://registry.opendata.aws/nrel-pds-nsrdb/</a></li> <li>Variables: modeled irradiance (clear-sky and all-sky direct normal (DNI) and global horizontal (GHI) irradiance, aerosol optical depth, and cloud optical depth</li> </ul> </li> <li>National Oceanic and Atmospheric Administration’s (NOAA) National Environmental Satellite, Data, and Information Service (NESDIS) Hazard Mapping System (HMS) smoke product. <ul> <li>Access: <a href="https://www.ospo.noaa.gov/Products/land/hms.html#maps">https://www.ospo.noaa.gov/Products/land/hms.html#maps</a></li> <li>Variables: smoke plume locations</li> </ul> </li> <li>National Aeronautics and Space Administration's (NASA) Multi-Angle Implementation of Atmospheric Correction (MAIAC) aerosol product (MCD19A2 MODIS/Terra + Aqua land aerosol optical depth daily L2G Global 1km SIN Grid V006). <ul> <li>Access: <a href="https://lpdaac.usgs.gov/products/mcd19a2v006/">https://lpdaac.usgs.gov/products/mcd19a2v006/</a></li> <li>Variables: aerosol optical depth and cloud mask</li> </ul> </li> <li>NASA's Clouds and the Earth’s Radiant Energy System (CERES) cloud data product (SYN1deg-1Hour Edition 4.1) <ul> <li>Access: <a href="https://ceres-tool.larc.nasa.gov/ord-tool/jsp/SYN1degEd41Selection.jsp">https://ceres-tool.larc.nasa.gov/ord-tool/jsp/SYN1degEd41Selection.jsp</a></li> <li>Variables: cloud optical depth</li> </ul> </li> </ul> <p>A detailed description of the data processing methods used to produce the final merged data are available in the article by Corwin et al. </p> <p>Associated code scripts are located in the linked code repository.</p>
Estimating historical air-sea CO2 fluxes: Incorporating physical knowledge within a data-only approach
<p>Reconstructed surface ocean pCO2 and air-sea CO2 fluxes for 1990-2019 using the pCO2-Residual Approach (JAMES 2021MS002960, in review)</p> <p>Surface ocean pCO2 (spo2) and resulting estimates of the air-sea CO2 flux (fCO2) are included in the netcdf file at monthly temporal resolution and for 1x1 grid cell spatial resolution. SeaFlux (https://zenodo.org/record/5482547#.YlT72y-B0_U) variables are used to calculate the fluxes from surface ocean pCO2.</p>
Historical phenology data from Hough (1864) and datasheets for phenometric analyses
<p>This dataset contains a zip file of phenology data from Hough (1864) converted from tables in that source to a format usable for analyses. This dataset also contains the summary datasheet for a phenometric analysis for an in process manuscript, 'Phenological response to climatic change depends on seasonal warming velocity and species traits' by Robert Guralnick, Erin Grady, Theresa Crimmins, and Lindsay Campbell.</p>
Historical tone data for Tai languages
<p>Comma-separated values (CSV) of historical tone data for 300+ Tai doculects (languages and dialects). Gives historical categories (Gedney 1972) and tone numerals (Chao 1930). Contact author for source citations. I recommend getting in touch if you'd like to use this dataset There is a good chance I have a newer (bigger, cleaner, better) version of it you could use!</p>
Downscaled 20CRv2c (#37) gridded historical climate data over China (1851-2010)
<p><strong>Gridded historical climate </strong><strong>data over China, spanning 1851 to 2010. Dynamically downscaled to 25km resolution using the PRECIS2.0 (HadRM3P) Met Office regional climate model, driven by 20th century reanalysis (20CRv2c, NOAA/ESRL PSD 20th Century Reanalysis version 2c, ensemble member 37).</strong></p> <p>This data has been un-rotated to true latitude longitude coordinates from its original rotate pole frame of reference. For more information on the PRECIS regional climate model, visit <a href="http://www.metoffice.gov.uk/precis">www.metoffice.gov.uk/precis</a>. Data near the boundaries should be used with caution due to model configuration aspects of regional climate modelling, and the interpolation method applied.</p> <p><strong>Domain</strong>: 17N to 58.84N, 73E to 135.7E</p> <p><strong>Countries covered</strong>: China, Nepal, Bhutan, Bangladesh, Taiwan, Mongolia, North Korea, South Korea, Kyrgzstan, and northern parts of India, Myanmar, Lao PDR & Vietnam.</p> <p><strong>Variables</strong>: pr (mean precipitation flux), tm (mean surface temperature), tn (minimum surface temperature) & tx (maximum surface temperature)</p> <p><strong>Time averaging</strong>: monthly</p> <p> </p> <p><em>This data set supplements the equivalent downscaled ERA-Interim data set: <a href="https://zenodo.org/record/2600192#.XJj3uKD7RWE">Downscaled ERA-Interim gridded historical climate data over China (1980-2010)</a> doi: 10.5281/zenodo.2600192</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.