Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,101

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,101 results for “historic”

Learn how ShareScore rates datasets ↗
edi48/100

Historical Records of the H.J. Andrews Experimental Forest Program, 1947 to present

This database contains historical records of the H.J. Andrews Experimental Forest (a U.S. Forest Service property near Blue River, OR, within the Willamette National Forest) consisting of three parts: 1.) inventory of the physical records and information about locating those physical records and three collections of digital records created from the physical records: 2.) records concerning the forest itself (HJATheForest), and 3.) records of the personal history of H.J. Andrews (HJATheMan).

openCC (other)Feb 2021View details →
edi48/100

Hohokam canals as multi-use facilities: Pre-historic canal system in the central Arizona-Phoenix metropolitan area

This is the digitized version of a map of the Hohokam canal system in what is now the Phoenix metropolitan area. It is based on the thesis research by J. B. Howard (Howard, J. (1990). Paleohydraulics : techniques for modeling the operation and growth of prehistoric canal systems. Thesis (M.A.)--Arizona State University, 1990). The original paper map is based on previous archaeological data, overlayed onto USGS 7.5 minute quadrangle maps to recreate the canal pattern.

openCC0Feb 2021View details →
edi48/100

Historical shorelines for the Atlantic barrier islands of Virginia south of Chincoteague Inlet, 1949-2006

Historical shorelines for the Virginia barrier islands from Fishermans Island to Wallops Island were compiled from various remote sensing and ground-based sources. The COAST dataset (Dolan et al. 1978, Dolan et al. 1990) tabulated shoreline position based on historic aerial orthophotographs at transects spaced 50 meters apart located along the mid-Atlantic coast of the United States (South Carolina to New Jersey); only data for the Virginia barrier islands (1949-1988) are presented here. COASTS data are supplemented with more recent shoreline data digitized from aerial imagery (USGS 1994 and VGIN 2002) and collected with GPS (Fenster 2006). Baselines and transects used both to reconstruct the COASTS data and to produce shoreline positions and calculate shoreline rates of change using DSAS are also included. Note: there are some unresolved georeferencing issues that cause inconsistencies in the shorelines from different base data frames. Data from different data frames should be integrated with caution. Future versions of this dataset will resolve these issues.

openCustomFeb 2022View details →
edi48/100

Marsh migration land use inferred from historical T-Sheets of the Chesapeake Bay

Detailed methods are listed in the associated publication (Schieder et al., 2018 https://doi.org/10.1007/s12237-017-0336-9). Briefly, we compared the spatial distribution of marshes in nineteenth-century maps to modern aerial photographs for the areas included in 40 NOS topographic sheets ("T-sheets") that included information on simple land types (e.g., marsh, farmland, forests) from the tidal portions of the Chesapeake Bay. Tidal marsh extent was digitized by hand by tracing the boundary between marsh and open water and the boundary between marsh and upland. The marsh-forest boundary was identified as the line between the dense tree canopy and marsh, the marsh-agriculture boundary was identified as the line between agriculture and marsh, and the marsh-water boundary was identified as the line between open water and adjacent land excluding beaches. The areas of agricultural land converted to marsh, forestland converted to marsh, and total upland conversion to marsh were summarized for each T-Sheet.

openCustomMay 2024View details →
zenodo44/100

Datasets For "Estimating Maximum Extent of Auroral Equatorward Boundary using Historical and Simulated Surface Magnetic Field Data", Blake et al. (2020), JGR

<p>Datasets and sample Python codes for the 2020 paper <em>&quot;Estimating Maximum Extent of Auroral Equatorward Boundary using Historical and Simulated Surface Magnetic Field Data&quot;</em>, by Blake et al.,&nbsp;submitted to the Journal of Gephysical Research, Space Physics.&nbsp;</p> <p>Up-to-date Python codes can be found at&nbsp;<a href="https://github.com/TerminusEst/Auroral_Boundary_Geomag">https://github.com/TerminusEst/Auroral_Boundary_Geomag</a></p> <p>The complete SWMF simulation folders (including parameter and log files etc.) can be requested from <a href="https://ccmc.gsfc.nasa.gov/index.php">NASA&#39;s Community Coordinated Modeling Center</a>.</p> <p>#########</p> <p><strong>Data/&nbsp;</strong>contains the following:</p> <p><strong>Data/HIST_DATA.txt&nbsp;</strong>contains the minimum Dst values and calculated maximum extents of the auroral equatorward boundaries for 25 years of INTERMAGNET data (1991-2016). The fourth column is the standard deviation of the calculated auroral boundary in&nbsp;degrees.&nbsp;</p> <p><strong>Data/Boundary_Fits.csv&nbsp;</strong>contains the calculated minimum Dst values, and calculated auroral boundaries using Method 1 and Method 2 (see main paper&#39;s ttext), for each of the 15 SWMF simulations. Also included are&nbsp;the uncertainties for each calculation.</p> <p><strong>Data/SWMF_outputs/&nbsp;</strong>contains 15<strong>&nbsp;</strong>.txt&nbsp;files,<strong>&nbsp;</strong>each of which correspond to an SWMF simulation of the same name given in Table 1 in the main text. These data are for the magnetic longitude, magnetic latitude and maximum calculated <em>E<sub>H</sub>&nbsp;</em>(V/km) for each simulation.</p> <p>#########</p> <p><strong>Codes/&nbsp;</strong>contains two python scripts, and some sample data. These scripts correspond to Section 2 in the main text:</p> <p>1)&nbsp;<strong>Boundary_Calc.py</strong>&nbsp;calculates the extent of the auroral boundary using magnetic latitudes and maximum calculated <em>E<sub>H</sub></em> values from multiple INTERMAGNET sites.&nbsp;&nbsp;</p> <p>2) <strong>Efield_Calc.py&nbsp;</strong>calculates the E-field for a single INTERMAGNET site using the Quebec 1-D resistivity model.</p> <p>A more detailed description of these codes can be found here:&nbsp;<a href="https://github.com/TerminusEst/Auroral_Boundary_Geomag">https://github.com/TerminusEst/Auroral_Boundary_Geomag</a></p> <p>&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

MALDI-TOF-MS spectra of historical whale skeletons from the Museum of Zoology, Strasbourg

<p>Spectra data from historical whale skeletons from the Museum.&nbsp; Samples were&nbsp;acid demineralized&nbsp;followed by gelatinization, digestion with&nbsp;trypsin, and peptide purification on C18 filters.&nbsp; They were run on on Bruker autoflex MALDI-TOF-MS.&nbsp; Mzml file formats for the raw data are provided here along with a file information csv file which provides the identification of the samples.<br> <br> For more information see the associated publciation.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Datasets and Models for Historical Newspaper Article Segmentation

<p>This record contains the datasets and models used and produced for the work reported in the paper &quot;<em>Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers</em>&quot; (<a href="https://infoscience.epfl.ch/record/282863?ln=en">link</a>).</p> <p>Please cite this paper if you are using the models/datasets or find it relevant to your research:</p> <pre><code>@article{barman_combining_2020, title = {{Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers}}, author = {Raphaël Barman and Maud Ehrmann and Simon Clematide and Sofia Ares Oliveira and Frédéric Kaplan}, journal= {Journal of Data Mining \&amp; Digital Humanities}, volume= {HistoInformatics} DOI = {10.5281/zenodo.4065271}, year = {2021}, url = {https://jdmdh.episciences.org/7097}, }</code></pre> <p><br> <strong>Please note that this record contains data under different licenses.</strong><br> <br> <strong>1. DATA</strong></p> <ul> <li><strong>Annotations (json files)</strong>: JSON files contains image annotations, with one file per newspaper containing region annotations (label and coordinates) in VIA format. The following licenses apply: <ul> <li>&nbsp;luxwort.json: those annotations are under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 license</a>. Please refer to the right statement specified for each image in the file.</li> <li>GDL.json, IMP.json and JDG.json: those annotations are under a <a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">CC BY-SA 4.0 license</a>.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>Image files: </strong>The archive images.zip contains the Swiss titles image files (GDL, IMP, JDG) used for the experiments described in the paper. Those images are under copyright (property of the journal <em>Le Temps </em>and of <em>ArcInfo</em>) and can be used <em>for academic research or educational purposes only</em>. Redistribution, publication or commercial use are not permitted. These terms of use are similar to the following right statement: <a href="http://rightsstatements.org/vocab/InC-EDU/1.0/">http://rightsstatements.org/vocab/InC-EDU/1.0/</a></li> </ul> <p>&nbsp;</p> <p><strong>2. MODELS</strong></p> <p>Some of the best models are released under a <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC BY-SA 4.0</a> license (they are also available as assets of the current Github <a href="https://github.com/dhlab-epfl/dhSegment-text/releases/tag/0.1">release</a>).</p> <ul> <li><strong>JDG_flair-FT</strong>: this model was trained on JDG using french Flair and FastText embeddings. It is able to predict the four classes presented in the paper (<code>Serial</code>, <code>Weather</code>, <code>Death notice</code> and <code>Stocks</code>).</li> <li><strong>Luxwort_obituary_flair-bpemb</strong>: this model was trained on Luxwort using multilingual Flair and Byte-pair embeddings. It is able to predict the <code>Death notice</code> class.</li> <li><strong>Luxwort_obituary_flair-FT_indomain</strong>: this model was trained on Luxwort using in-domain Flair and FastText embeddings (trained on Luxwort data). It is also able to predict the <code>Death notice</code> class.</li> </ul> <p>Those models can be used to predict probabilities on new images using the same code as in the original <a href="https://github.com/dhlab-epfl/dhSegment">dhSegment</a> repository. One needs to adjust three parameters to the <code>predict</code> function: 1) <code>embeddings_path</code> (the path to the embeddings list), 2) <code>embeddings_map_path</code>(the path to the compressed embedding map), and 3) <code>embeddings_dim</code> (the size of the embeddings).</p> <p>Please refer to the paper for further information or contact us.</p> <p>&nbsp;</p> <p><strong>3. CODE:&nbsp;</strong></p> <p><a href="https://github.com/dhlab-epfl/dhSegment-text">https://github.com/dhlab-epfl/dhSegment-text</a></p> <p><br> <strong>4. ACKNOWLEDGEMENTS</strong><br> We warmly thank the journal <a href="https://letemps.ch">Le Temps</a> (owner of <em>La Gazette de Lausanne</em> and the <em>Journal de Gen&egrave;ve</em>) and the group <a href="https://www.arcinfo.ch/">ArcInfo</a> (owner of <em>L&#39;Impartial</em>) for accepting to share the related datasets for academic purposes. We also thank the <a href="https://bnl.public.lu/fr.html">National Library of Luxembourg</a> for its support with all steps related to the <em>Luxemburger Wort</em> annotation release.<br> This work was realized in the context of the <a href="https://impresso-project.ch"><em>impresso</em> - Media Monitoring of the Past</a> project and supported by the Swiss National Science Foundation under grant CR- SII5_173719.<br> <br> <strong>5. CONTACT</strong><br> Maud Ehrmann (EPFL-DHLAB)<br> Simon Clematide (UZH)</p>

openother-ncJan 2021View details →
zenodo44/100

Historical uncertainty in Gregory of Tours's History of the Franks (book 7)

<p>Our goal was to create a research dataset based on geographical and chronological uncertainties in the work of Gregory of Tours&#39;s *History of the Franks* (book 7). We used and modified a topology of geographical and chronological uncertainty based on a rudimentary schema that would be universal when analysing an historical source :</p> <p>Chronological :<br> * uncertain dating<br> * uncertain method of dating<br> * lack of dating<br> * precise dating</p> <p>Geographical:&nbsp;<br> * uncertain location<br> * general location (region, country)<br> * lack of location&nbsp;<br> * precise location</p> <p>After working on book 7 for a while, that schema was reworked as those 9 types of uncertainty :&nbsp;</p> <p>Chronological :<br> * uncertain_dating<br> * uncertain_method_dating<br> * event_dating_null<br> * precise_dating</p> <p>Geographical:&nbsp;<br> * uncertain_location<br> * general_location&nbsp;<br> * event_location_null<br> * uncertain_method_location<br> * precise_location</p> <p><br> The geographical and chronological focus makes it possible to identify where and when, in a source, the historical uncertainty is higher.&nbsp;</p> <p>Using python, that dataset was then automatically cleaned and enhanced with bounding box based on geo-mapping information for the entries of geographical uncertainty. Those were classified as either precise_location or general_location.&nbsp;</p> <p>For example, anything relating to a city general area (like &nbsp;&quot;in the Rouen area&quot;) creates a general_location bounding box encompassing the *current* geographical space occupied by the municipality of Rouen (in the format&nbsp;&#39;LongMin&#39;, &#39;LongMax&#39;, &#39;LatMin&#39;, &#39;LatMax&#39;&nbsp;in a single column &quot;bbox&quot;). Anything described as a unique point in space (like &quot;in Paris&quot;) creates a precise_location and its corresponding lat/long system of coordinates.&nbsp;</p> <p>This is an arbitrary way to translate slightly undefined geographical concepts&nbsp;of uncertainty into formal data, but at least it can be fully explained explicitly.<br> &nbsp;&nbsp;</p> <p>Translation used:&nbsp;Tours G. <em>et alii</em>, <em>The history of the franks</em>, Penguin Books Limited, 1974, <a href="https://books.google.ch/books?id=4Lx-M2RHGgoC">https://books.google.ch/books?id=4Lx-M2RHGgoC</a>.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

C3-EURO4M-MEDARE Mediterranean historical climate data

<p>Historical surface climate data files and meta-data for stations in Mediterranean North Africa and Middle East areas (1852-2008)</p>

opencc-by-sa-4.0Nov 2013View details →
zenodo44/100

Historically Irish Surnames Dataset

<p>This dataset provides a list of surnames that are reliably Irish and that can be used for identifying textual references to Irish individuals in the London area and surrounding countryside within striking distance of the capital. This classification of the Irish necessarily includes the Irish-born and their descendants. The dataset has been validated for use on records up to the middle of the nineteenth century, and should only be used in cases in which a few mis-classifications of individuals would not undermine the results of the work, such as large-scale analyses. These data were created through an analysis of the 1841 Census of England and Wales, and validated against the Middlesex Criminal Registers (National Archives HO 26) and the Vagrant Lives Dataset (Crymble, Adam et al. (2014). Vagrant Lives: 14,789 Vagrants Processed by Middlesex County, 1777-1786. Zenodo. 10.5281/zenodo.13103). The sample was derived from the records of the Hundred of Ossulstone, which included much of rural and urban Middlesex, excluding the City of London and Westminster. The analysis was based upon a study of 278,949 adult males. Full details of the methodology for how this dataset was created can be found in the following article, and anyone intending to use this dataset for scholarly research is strongly encouraged to read it so that they understand the strengths and limits of this resource:</p> <p>&nbsp;&nbsp;&nbsp; Adam Crymble, &#39;A Comparative Approach to Identifying the Irish in Long Eighteenth Century London&#39;, _Historical Methods: A Journal of Quantitative and Interdisciplinary History_, vol. 48, no. 3 (2015): 141-152.</p> <p>The data here provided includes all 283 names listed in Appendix I of the above paper, but also an additional 209 spelling variations of those root surnames, for a total of 492 names.</p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Supplementary Material for "Sequence Comparison in Historical Linguistics"

<p>This is the official version of the supplementary material which was used as the basis for the study on &quot;Sequence Comparison in Historical Linguistics&quot; (List, D&uuml;sseldorf, D&uuml;sseldorf University Press).</p>

opencc-by-4.0Sep 2014View details →
zenodo44/100

Dataset for the paper "Historical model biases in monthly high temperature anomalies indicate under-projection of future temperature extremes"

<div> <div>This repository holds data and scripts related to the revision of the paper entitled: <span>"Historical model biases in monthly high temperature anomalies indicate under-projection of future temperature extremes" </span>by Lei Duan, Lyssa M. Freese, Govindasamy Bala, and Ken Caldeira. <span>The paper is currently submitted for peer review. </span>Any questions regarding the data and paper could be sent to the corresponding author: Lei Duan (leiduan@carnegiescience.edu).&nbsp;</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo44/100

PARESv3 : PArish REgistry Survey − Historical Census Table Dataset (19th, 20th centuries) − France

<h2>PARES Dataset v3</h2> <p>PARES (PArish REcord Survey) contains<strong> 535 images of handwritten census tables</strong> for years ranging from around <strong>1650 A.D. until 1850 A.D.</strong>.They come from two <strong>French cities</strong>, Vic-sur-Seille (French department of Moselle) and Echevronne (French department of C&ocirc;te d'Or). While they mention very ancient times, the documents are handwritten transcriptions of even older documents and are quite recent, copied from original documents during the 1950's and 1960's for demographic studies led by the INED in France (<em>Institut National des &eacute;tudes d&eacute;mographiques</em> &minus; National Institute for Demographic Studies). These copies were made by only a few different writers.</p> <p>In this updated version of the dataset, each table row has been carefully annotated and transcribed. Please note that for each row transcription, we have specified the attribute to which each value corresponds.</p> <p>We published a paper, <a href="https://link.springer.com/article/10.1007/s10032-025-00531-z">The PARES Database: Information Extraction over Historical Parish Records,</a> in which we better describe the dataset and the tasks it's possible to run on it.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Map data of historical global estimates of soil respiration

<p>The map data of global soil respiration converted to NetCDF format.</p><p>All open access available estimates were collated.</p><p>Shoji Hashimoto, Akihiko Ito, Kazuya Nishina (2023)&nbsp;"Divergent data-driven estimates of global soil respiration". Communications Earth &amp;&nbsp;Environment, 4 Article number: 460</p><p><a href="https://doi.org/10.1038/s43247-023-01136-2 ">https://doi.org/10.1038/s43247-023-01136-2</a>&nbsp;</p><p>Refer to Table 1 for the study ID and data source&nbsp;or the attributions of the NetCDF file.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

HyG: A hydraulic geometry dataset derived from historical stream gage measurements across the conterminous United States

<p>Regional- and continental-scale models predicting variations in the magnitude and timing of streamflow are important tools for forecasting water availability as well as flood inundation extent and associated damages. Such models must define the geometry of stream channels through which flow is routed. These channel parameters, such as width, depth, and hydraulic resistance, exhibit substantial variability in natural systems. While hydraulic geometry relationships have been extensively studied in the United States, they remain unquantified for thousands of stream reaches across the country. Consequently, large-scale hydraulic models frequently take simplistic approaches to channel geometry parameterization. Over-simplification of channel geometries directly impacts the accuracy of streamflow estimates, with knock-on effects for water resource and hazard prediction.</p> <p>Here, we present a hydraulic geometry dataset derived from long-term measurements at U.S. Geological Survey (USGS) stream gages across the conterminous United States (CONUS). This dataset includes (a) at-a-station hydraulic geometry parameters following the methods of Leopold and Maddock (1953), (b) at-a-station Manning's n calculated from the Manning equation, (c) daily discharge percentiles, and (d) downstream hydraulic geometry regionalization parameters based on HUC4 (Hydrologic Unit Code 4). This dataset is referenced in Heldmyer et al. (2022); further details and implications for CONUS-scale hydrologic modeling are available in that article (https://doi.org/10.5194/hess-26-6121-2022).&nbsp;</p> <p><strong>At-a-station Hydraulic Geometry</strong></p> <p>We calculated hydraulic geometry parameters using historical USGS field measurements at individual station locations. Leopold and Maddock (1953) derived the following power law relationships:</p> <p>\(w={aQ^b}\)</p> <p>\(d=cQ^f\)</p> <p>\(v=kQ^m\)</p> <p>where Q is discharge, w is width, d is depth, v is velocity, and a, b, c, f, k, and m are at-a-station hydraulic geometry (AHG) parameters. We downloaded the complete record of USGS field measurements from the USGS NWIS portal (https://waterdata.usgs.gov/nwis/measurements). This raw dataset includes 4,051,682 individual measurements from a total of 66,841 stream gages within CONUS. Quantities of interest in AHG derivations are Q, w, d, and v. USGS field measurements do not include d--we therefore calculated d using d=A/w, where A is measured channel area. We applied the following quality control (QC) procedures in order to ensure the robustness of AHG parameters derived from the field data:</p> <ol> <li>We considered only measurements which reported Q, v, w and A.</li> <li>For each gage, we excluded measurements older than the most recent five years, so as to minimize the effects of long-term channel evolution on observed hydraulic geometry relationships.</li> <li>We excluded gages for which measured Q disagreed with the product of measured velocity and measured area by more than 5%. Gages for which&nbsp; \( Q\neq vA\) are often tidally influenced and therefore may not conform to expected channel geometry relationships.</li> <li>Q, v, w, and d from field measurements at each gage were log-transformed. We performed robust linear regressions on the relationships between log(Q) and log(w), log(v), and log(d). AHG parameters were derived from the regressed explanatory variables. <ol> <li>We applied an iterative outlier detection procedure to the linear regression residuals. Values of log-transformed w, v, and d residuals falling outside a three median absolute deviation (MAD) envelope were excluded. Regression coefficients were recalculated and the outlier detection procedure was reapplied until no new outliers were detected.</li> <li>Gages for which one or more regression had p-values &gt;0.05 were excluded, as the relationships between log-transformed Q and w, v, or d lacked statistical significance.</li> <li>Gages were omitted if regressed AHG parameters did not fulfill two additional relationships derived by Leopold and Maddock: \(b+f+m=1{\displaystyle \pm }0.1\) and \(a{\displaystyle \times }c{\displaystyle \times }k=1{\displaystyle \pm }0.1\).</li> </ol> </li> <li>If the number of field measurements for a given gage was less than 10, either initially or after individual measurements were removed via steps 1-4, the gage was excluded from further analysis.</li> </ol> <p>Application of the QC procedures described above removed 55,328 stream gages, many of which were short-term campaign gages at which very few field measurements had been recorded. We derived AHG parameters for the remaining 11,513 gages which passed our QC.</p> <p><strong>At-a-station Manning's n</strong></p> <p>We calculated hydraulic resistance at each gage location by solving Manning's equation for Manning's n, given by</p> <p>\(n = {{R^{2/3}S^{1/2}} \over v}\)</p> <p>where v is velocity, R is hydraulic radius and S is longitudinal slope. We used smoothed reach-scale longitudinal slopes from the NHDPlusv2 (National Hydrography Dataset Plus, version 2) ElevSlope data product. We note that NHDPlusv2 contains a minimum slope constraint of 10<sup>-5</sup> m/m--no reach may have a slope less than this value. Furthermore, NHDPlusv2 lacks slope values for certain reaches. As such, we could not calculate Manning's n for every gage, and some Manning's n values we report may be inaccurate due to the NHDPlusv2 minimum slope constraint. We report two Manning's n values, both of which take stream depth as an approximation for R. The first takes the median stream depth and velocity measurements from the USGS's database of manual flow measurements for each gage. The second uses stream depth and velocity calculated for a 50th percentile discharge (Q<sub>50</sub>; see below). Approximating R as stream depth is an assumption which is generally considered valid if the width-to-depth ratio of the stream is greater than 10<span>&mdash;</span>which was the case for the vast majority of field measurements. Thus, we report two Manning's n values for each gage, which are each intended to approximately represent median flow conditions.</p> <p><strong>Daily discharge percentiles</strong></p> <p>We downloaded full daily discharge records from 16,947 USGS stream gages through the NWIS online portal. The data includes records from both operational and retired gages. Records for operational gages were truncated at the end of the 2018 water year (September 30, 2018) in order to avoid use of preliminary data. To ensure the robustness of daily discharge percentiles, we applied the following QC:</p> <ol> <li>For a given gage, we removed blocks of missing discharge values longer than 6 months. These long blocks of missing data generally correspond to intervals in which a gage was temporarily decommissioned for maintenance.</li> <li>A gage was omitted from further analysis if its discharge record was less than 10 years (3,652 days) long, and/or less than 90% complete (&gt;10% missing values after removal of long blocks in step 1.</li> </ol> <p>We calculated discharge percentiles for each of the 10,871 gages which passed QC. Discharge percentiles were calculated at increments of 1% between Q<sub>1</sub> and Q<sub>5</sub>, increments of 5% (e.g. Q<sub>10</sub>, Q<sub>15</sub>, Q<sub>20</sub>, etc.) between Q<sub>5</sub> and Q<sub>95</sub>, increments of 1% between Q<sub>95</sub> and Q<sub>99</sub>, and increments of 0.1% between Q<sub>99</sub> and Q<sub>100</sub> in order to provide higher resolution at the lowest and highest flows, which occur much less frequently.</p> <p><strong>HG Regionalization</strong></p> <p>We regionalized AHG parameters from gage locations to all stream reaches in the conterminous United States. This downstream hydraulic geometry regionalization was performed using all gages with AHG parameters in each HUC4, as opposed to traditional downstream hydraulic geometry--which involves interpolation of parameters of interest to ungaged reaches on individual streams. We performed linear regressions on log-transformed drainage area&nbsp;and Q at a number of flow percentiles as follows:</p> <p>\(log(Q_i) = \beta_1log(DA) + \beta_0\)</p> <p>where Q<sub>i</sub> is streamflow at percentile i, DA is drainage area and \(\beta_1\) and \(\beta_0\) are regression parameters. We report \(\beta_1\),&nbsp; \(\beta_0\) , and the r<sup>2</sup> value of the regression relationship for Q percentiles Q<sub>10</sub>, Q<sub>25</sub>, Q<sub>50</sub>, Q<sub>75</sub>, Q<sub>90</sub>, Q<sub>95</sub>, Q<sub>99</sub>, and Q<sub>99.9</sub>. Further discussion and additional analysis of HG regionalization are presented in Heldmyer et al. (2022).</p> <p><strong>Dataset description</strong></p> <p>We present the HyG dataset in a comma-separated value (csv) format. Each row corresponds to a different USGS stream gage. Information in the dataset includes gage ID (column 1), gage location in latitude and longitude (columns 2-3), gage drainage area (from USGS; column 4), longitudinal slope of the gage's stream reach (from NHDPlusv2; column 5), AHG parameters derived from field measurements (columns 6-11), Manning's n calculated from median measured flow conditions (column 12), Manning's n calculated from Q50 (column 13), Q percentiles (columns 14-51), HG regionalization parameters and r<sup>2</sup> values (columns 52-75), and geospatial information for the HUC4 in which the gage is located (from USGS; columns 76-87). Users are advised to exercise caution when opening the dataset. Certain software, including Microsoft Excel and Python, may drop the leading zeros in USGS gage IDs and HUC4 IDs if these columns are not explicitly imported as strings.</p> <p>&nbsp;</p> <p><strong>Errata</strong></p> <p>In version 1, drainage area was mistakenly reported in cubic meters but labeled in cubic kilometers. This error has been corrected in version 2.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Scaling landscape fire history in sagebrush: Wildfires not historically frequent in the main population of threatened Gunnison Sage-grouse

<p>The main population of &sim;5,000 Threatened Gunnison sage-grouse (GUSG; Centrocercus minimus) in Colorado depends on sagebrush that are killed by wildfires, with recovery taking decades, so frequent fire is a threat, but did it occur historically? Early land surveys showed that the historical (preindustrial) fire rotation (FR), the expected period to burn area equal to a focal land area, was 90-143 years in GUSG ranges, which is not frequent fire (&le;25 years). However, recent research, based on fire scars on trees at ten sites near sagebrush, suggested some frequent fire historically in the main population. That study was not spatial, essential to estimate FR, so spatial data were created in GIS with land-survey reconstructions, survey dates, fire-scar sites, Thiessen polygons around sites, and sagebrush. The previous study assumed fires that burned 2+ sites likely burned across sagebrush. Historical FRs were calculated several ways over a common period. A recovery estimate of FR was 90-135 years, a land-survey estimate 82-131 years, and three spatial scar-based estimates 93-107 years, showing agreement. However, comparing land-survey and fire-scar results showed that using fire scars spatially only 43% matched land surveys. Detailed analysis showed that 10 fire-scar sites were insufficient to detect historical fire sizes and distributions across the large 168,753 ha sagebrush area. An adequate historical fire reconstruction could require &sim;45-60 fire-scar sites, making only &sim;30,000 ha of sagebrush feasible. Using the two remaining methods, which cross-validate, showed frequent fire did not occur historically in the study area, as historical FRs were 82-135 years.&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Data from: Solar energy resource availability under extreme and historical wildfire smoke conditions

<p>The data in this repository are used to generate the figures in the article "Solar energy resource availability under extreme and historical wildfire smoke conditions" by Corwin et al. (accepted 2024) in <em>Nature Communications</em>. Data are the final processessed and merged datasets sourced from the following publicly available data products:</p> <ul> <li>National Renewable Energy Laboratory&rsquo;s (NREL) National Solar Radiation Database (NSRDB) (<a href="https://nsrdb.nrel.gov/)">https://nsrdb.nrel.gov/)</a>. <ul> <li>Bulk download in July 2023 via AWS:&nbsp;<a href="https://registry.opendata.aws/nrel-pds-nsrdb/">https://registry.opendata.aws/nrel-pds-nsrdb/</a></li> <li>Variables: modeled irradiance (clear-sky and all-sky direct normal (DNI) and global horizontal (GHI) irradiance, aerosol optical depth, and cloud optical depth</li> </ul> </li> <li>National Oceanic and Atmospheric Administration&rsquo;s (NOAA) National Environmental Satellite, Data, and Information Service (NESDIS) Hazard Mapping System (HMS) smoke product. <ul> <li>Access: <a href="https://www.ospo.noaa.gov/Products/land/hms.html#maps">https://www.ospo.noaa.gov/Products/land/hms.html#maps</a></li> <li>Variables: smoke plume locations</li> </ul> </li> <li>National Aeronautics and Space Administration's (NASA) Multi-Angle Implementation of Atmospheric Correction (MAIAC) aerosol product (MCD19A2 MODIS/Terra + Aqua land aerosol optical depth daily L2G Global 1km SIN Grid V006).&nbsp; <ul> <li>Access: <a href="https://lpdaac.usgs.gov/products/mcd19a2v006/">https://lpdaac.usgs.gov/products/mcd19a2v006/</a></li> <li>Variables: aerosol optical depth and cloud mask</li> </ul> </li> <li>NASA's Clouds and the Earth&rsquo;s Radiant Energy System (CERES) cloud data product (SYN1deg-1Hour Edition 4.1) <ul> <li>Access: <a href="https://ceres-tool.larc.nasa.gov/ord-tool/jsp/SYN1degEd41Selection.jsp">https://ceres-tool.larc.nasa.gov/ord-tool/jsp/SYN1degEd41Selection.jsp</a></li> <li>Variables: cloud optical depth</li> </ul> </li> </ul> <p>A detailed description of the data processing methods used to produce the final merged data are available in the article by Corwin et al.&nbsp;</p> <p>Associated code scripts are located in the linked code repository.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

A list of Swedish words that have experienced historical semantic changes

<p>This list contains a set of Swedish words that have experienced semantic change during the past centuries. The list has been collected during the VR funded project <a href="https://languagechange.org/">Towards Computational Lexical Semantic Change Detection</a>, ( 2018-01184) and is a work in progress. Because the work is currently on pause, we have chosen to release the list as is in the hope of facilitating collaboration and use in other research.</p> <p>The list has four columns in the following format:</p> <pre><code>Ord*, Vilken betydelseförändring har skett*, När skedde förändringen, Källor (exv SAOL) word, what change occurred, when the change occur, reference</code></pre> <p><br> Not all fields are filled for every word. Where there are multiple change periods, there are multiple lines with empty values for word and what change occurred, see the example with <em>egendomlig </em>below.</p> <pre><code>egendomlig,som utgör (ngns) egendom &gt; karaktäristisk (positiv) &gt; speciell (negativt!!),"A. Sen 1600tal, ",SAOB ,,B. Sen 1850, ,,"C. Sen ??, efter 190",</code></pre> <p>The .xlsx file contains links to the references.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Historical Film Shot Dataset V1 (HistShotDS V1)

<p><em>Paper title: </em></p> <p><strong>HistShot: A Shot Type Dataset based on Historical Documentation during WWII</strong></p> <p><em>Conference title: </em></p> <p>International Conference on Pattern Recognition Applications and Methods (ICPRAM 2022)</p> <p>&nbsp;</p> <p><em>Description:</em></p> <p>Automated shot type classification plays a significant role in film preservation and indexing of film datasets. In this paper a historical shot type dataset (HistShot) is presented, where the frames have been extracted from original historical documentary films. A center frame of each shot has been chosen for the dataset and is annotated according to the following shot types: Close-Up (CU), Medium-Shot (MS), Long-Shot (LS), Extreme-Long-Shot (ELS), Intertitle (I), and Not Available/None (NA). The validity to choose the center frame is shown by a user study. Additionally, standard CNN-based methods (ResNet50, VGG16) have been applied to provide a baseline for the HistShot dataset.</p> <p>&nbsp;</p> <p><em>References: </em></p> <p>Github Repository: <a href="https://github.com/dahe-cvl/ICPRAM2022_histshotV1">https://github.com/dahe-cvl/ICPRAM2022_histshotV1</a></p> <p>VHH-MMSI: <a href="https://vhh-mmsi.eu/">https://vhh-mmsi.eu/</a></p> <p>VHH-project Page: <a href="https://www.vhh-project.eu/">https://www.vhh-project.eu/</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Datasets for Ecohydrological Model for Grassland Lacking Historical Measurements

<p>We proposed the distributed dynamic process model (DDPM) and here are the example datasets of our model.</p> <p>Datasets I : Eva_module_datasets.zip: Here are the datasets for running the DDPM Eva module.</p> <p>Datasets II: RP&amp;FLC_module_datasets.zip: Here are the datasets for running the DDPM RP &amp; FLC module.</p> <p>&nbsp;</p> <p>Also, you can find our model at https://github/myli1993/DDPM_ver1.0.</p> <p>You can watch my report on MODSIM 2021 at https://www.bilibili.com/video/BV1vR4y1s7vc.</p> <p>-----------------------------------------------------------------------------------------</p> <p>When using this dataset, please cite:</p> <p>[1] <strong>Li, MY.</strong>; Liu, TX.; Duan, LM.; et al. Confluence simulations based on dynamic channel parameters in the grasslands lacking historical measurements. <em><strong>Journal of Hydrology</strong></em>, 2023, Volume 627, 130425. <a href="https://doi.org/10.1016/j.jhydrol.2023.130425">https://doi.org/10.1016/j.jhydrol.2023.130425</a></p> <p>[2] <strong>Li, MY.</strong>; Liu, TX.; Duan, LM.; et al. A novel evapotranspiration downscaling approach based on dynamic sensitive parameters and deep learning in the grassland lacking historical measurements. <em><strong>Ecological Indicators</strong></em>, 2025, Volume 178, 113839. https://doi.org/10.1016/j.ecolind.2025.113839</p> <p>-----------------------------------------------------------------------------------------</p> <p>What's new:</p> <p>Version 1.01:</p> <p>Here, we added the datasets for Eva module, and renamed the datasets for RP &amp; FLC module. If you have downloaded the datasets in Version 1.0, you don't have to download the RP&amp;FLC_module_datasets.zip for another time.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record