Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27,923

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

27,923 results for “model”

Learn how ShareScore rates datasets ↗
edi52/100

Modeling the effect of explicit vs implicit representation of grazing on ecosystem carbon and nitrogen cycling in response to elevated carbon dioxide and warming in arctic tussock tundra, Alaska - Dataset B

We use a simple model of coupled carbon and nitrogen cycles in terrestrial ecosystems to examine how explicitly representing grazers versus having grazer effects implicitly aggregated in with other biogeochemical processes in the model alters predicted responses to elevated carbon dioxide and warming. The aggregated approach can affect model predictions because grazer-mediated processes can respond differently to changes in climate from the processes with which they are typically aggregated. We use small-mammal grazers in arctic tundra as an example and find that the typical three-to-four-year cycling frequency is too fast for the effects of cycle peaks and troughs to be fully manifested in the ecosystem biogeochemistry. We conclude that implicitly aggregating the effects of small-mammal grazers with other processes results in an underestimation of ecosystem response to climate change relative to estimations in which the grazer effects are explicitly represented. The magnitude of this underestimation increases with grazer density. We therefore recommend that grazing effects be incorporated explicitly when applying models of ecosystem response to global change.

openCC (other)Mar 2022View details →
edi52/100

Model Simulations of The Effects of Shifts in High-frequency Weather Variability (No Long-term Weather Trend) Control Carbon Loss from Land to the Atmosphere, Toolik Lake, Alaska, 2022-2122

Climate change is increasing extreme weather events, but effects on high-frequency weather variability and the resultant impacts on ecosystem function are poorly understood. We assessed ecosystem responses of arctic tundra to changes in day-to-day weather variability using a biogeochemical model and stochastic simulations of daily temperature, precipitation, and light. Changes in weather variability altered ecosystem carbon, nitrogen, and phosphorus stocks and cycling rates. Some responses of processes (e.g., respiration) were inconsistent with expectations, indicating that whole-ecosystem interactions and feedbacks moderate or even reverse responses to weather variability. More weather variability led to greater carbon losses from land to atmosphere, and less variability led to higher carbon sequestration on land. The magnitude of response to weather variability was similar to that predicted from climate mean trend effects. This dataset consists of the MEL parameter file, driver files and output files for simulations without a long term weather trend.

openCC (other)Aug 2022View details →
edi52/100

Model Simulations of The Effects of Shifts in High-frequency Weather Variability (With a Long-term Trend) on Carbon Loss from Land to the Atmosphere, Toolik Lake, Alaska, 2022-2122

Climate change is increasing extreme weather events, but effects on high-frequency weather variability and the resultant impacts on ecosystem function are poorly understood. We assessed ecosystem responses of arctic tundra to changes in day-to-day weather variability using a biogeochemical model and stochastic simulations of daily temperature, precipitation, and light. Changes in weather variability altered ecosystem carbon, nitrogen, and phosphorus stocks and cycling rates. Some responses of processes (e.g., respiration) were inconsistent with expectations, indicating that whole-ecosystem interactions and feedbacks moderate or even reverse responses to weather variability. More weather variability led to greater carbon losses from land to atmosphere, and less variability led to higher carbon sequestration on land. The magnitude of response to weather variability was similar to that predicted from climate mean trend effects. This dataset consists of the MEL parameter file, driver files and output files for simulations with a long-term weather trend.

openCC (other)Aug 2022View details →
edi52/100

Model estimates of runoff, dissolved organic carbon, soil temperature and moisture for Elson Lagoon watershed, Alaska, 1981-2020

This dataset contains model estimates of dissolved organic carbon (DOC) yield (mg C/m^2) and runoff (mm), for surface and subsurface flows, soil temperature (degree C), and soil moisture (% of soil volume) for grid cells spanning the Elson Lagoon watershed in northwest Alaska. Daily air temperature, precipitation, and wind speed data from Utqiagvik airport were used for meteorological forcings for the daily simulation by the Permafrost Water Balance Model (PWBM) from 1981 to 2020. The DOC and runoff data files are organized by grid cell and month. The soil temperature and soil moisture files are organized by grid cell and day of year (DOY), and contain values for the first eight model soil layers, with centers of the layers at 1, 3, 8, 13, 23, 33, 45, 55 cm depth. The estimates are most useful for analyses of the dynamics of the watershed’s surface and subsurface runoff and DOC yield. Leachate DOC concentrations can be obtained using the gridded runoff and yield values. A manuscript describing the data and associated analysis has been accepted for publication in Environmental Research Letters (Rawlins et al., 2021).

openCC0Sep 2021View details →
edi52/100

The dataset and model code pertinent to the Everglades Peat Elevation Model (EvPEM): The salinity and inundation mesocosm experiment in freshwater and brackish water sawgrass wetlands in Florida Coastal Everglades (2015-2017).

This is an assembled data and Everglades Peat Elevation Model (EvPEMv1.0) Stella code used to estimate and simulate net ecosystem carbon balance (NECB) and peat elevation change in response to saltwater intrusion and level of inundations. Data from several studies were combined for the estimation of NECB, model parameterization, and calibration (Wilson, 2018; Wilson et al., 2018, 2019; Charles et al., 2019; Servais et al., 2020). The reported data includes aboveground net primary productivity (ANPP), belowground net primary productivity (BNPP), peat elevation change, and decomposition rates that were collected from outdoor laboratory mesocosm experiments conducted at the Florida Bay Interagency Science Center in Key Largo, Florida during 2015-17. The plant-soil monoliths were obtained from a freshwater peat and a brackish water peat marsh located within the Florida Coastal Everglades and transported to the Key Largo facility for the experimental manipulations. In experiments focused on the brackish water marsh, three experiments were carried out reflecting the combined effect of salinity, inundation, and peat exposure to air. The brackish water experiments characterized submerged (SUB), exposed (EXP), and extended depth of exposure of peat surface (EXTEXP) conditions, as we varied water depth relative to the peat surface. Each experiment was subjected to two salinity manipulations: (1) ambient (~10 ppt) porewater salinity (AMB) and (2) elevated (~20 ppt) salinity (SALT). The experimental design included six (2 X 3) treatments: (1) submerged ambient salinity (AMB.SUB), (2) submerged elevated salinity (SALT.SUB.), (3) exposed ambient salinity (AMB.EXP), (4) exposed elevated salinity (SALT.EXP), (5) exposed with extended exposure/dry-down ambient salinity (AMB.EXTEXP), and (6) exposed with extended exposure/dry-down elevated salinity (SALT.EXTEXP). The water level was kept 4 cm above the peat surface for the brackish water SUB treatments. Exposure for the EXP treatment

openCC (other)Feb 2022View details →
edi52/100

Monthly Spartina alterniflora marsh vegetation data for additional sites along the Georgia coast used in the Belowground Ecosystem Resiliency Model

Study plots (1-m2) were established in three Spartina alterniflora-dominated marshes - 2 on Sapelo Island, Georgia, and 1 on Skidaway Island, Georgia, and sampled once each during May, July, August, September, and October of 2016. Nine replicate plots were placed in vegetated marsh along transects that spanned 3 Landsat-8 pixel footprints, with 3 plots per pixel foot print. In each plot, measurements included plant biomass, plant species, stem density, and height. Aboveground biomass was calculated using allometric relationships between plant height, flowering status and mass from plant clipping studies. During these surveys, destructive core sampling was also performed in the proximity of the plots (n = 1 per plot) to measure above and below ground biomass. Chlorophyll, foliar N, and Leaf Area Index measurements were taken in the proximity of the plots.

openCC (other)Jul 2021View details →
edi52/100

Species Distribution Modeling of Carnivorous Plants Worldwide

Forecasting how carnivorous plant species will respond to climatic change is a key issue in their conservation and management but presents a number of challenges. These challenges derive from interactions between the relatively simplistic statistical methods typically used to forecast species responses to climatic change, which to date have been limited mainly to species distribution models (“SDMs) and particular aspects of the ecology of carnivorous plants, including their rarity, habitat specialization, and limited dispersal ability. The small ranges and oftentimes low local abundance of carnivorous plants provide few occurrence records, which increase the potential for poorly or over-fitted SDMs and misspecification of relationships with their “optimal” environments. The unique habitats in which carnivorous plants often grow also are difficult to characterize using the basic temperature and precipitation data that often undergird SDMs. Rather, habitats in which carnivorous plants are common often are decoupled from broader climatic patterns (e.g., many retain high soil moisture even during seasonal drought) and may be associated with frequent disturbance. Last, dispersal limitation also may constrain range shifts of carnivorous plants as the climate changes. These three issues raise two related questions that are critical for understanding and forecasting the future of carnivorous plants. First, to what extent are current carnivorous plants distributions constrained by climate; and second, how readily, if at all, might carnivorous plants disperse to colonize new habitat as it becomes climatically suitable? We estimated the vulnerability of carnivorous plants to climatic change in light of challenges identified with SDMs in general and their particular application to these unique species. We combined two approaches: “ensembles of small models”, which attempt to deal with the challenges of fitting SDMs for data-limited species; and “bioclimatic velocity”, which is

openCC0Dec 2023View details →
edi52/100

MCSE Model input data at the Kellogg Biological Station, Hickory Corners, MI (1988 to 2020)

Dataset AbstractConsolidated dataset for the ARDEN crop modeling effort. This pulls together several useful data tables into one dataset. Further information can be found at https://agmip.github.io/ARDN/original data source http://lter.kbs.msu.edu/datasets/195

openCC (other)Jan 2021View details →
edi52/100

MCR LTER: Coral Reef: Material legacy disturbance type model; data for Kopecky et al., 2023 Ecology

This data package contains the code necessary to create a mathematical model of coral reef recovery dynamics following different types and intensities of disturbances that either remove dead coral skeletons (e.g., tropical storms) or leave standing dead skeletons (e.g., coral bleaching) and run associated analyses. We explored the sensitivity of the model to variation in key parameters, such as the strength of herbivory, and the degree to which dead skeletons protect algae from herbivory. Further, we assessed disturbance intensities and values of these parameters that lead to shifts between coral and macroalgae-dominated reefs. This code was published in Ecology and were a part of the thesis of K. Kopecky (2023). Analyses and full methods descriptions of this model can be found in the manuscript “Material legacies can degrade resilience: Structure-retaining disturbances promote regime shifts on coral reefs” (DOI: https://doi.org/10.1002/ecy.4006). No novel data were used or generated in this study. This manuscript uses data collected by the U.S. National Science Foundation's (NSF) Moorea Coral Reef Long Term Ecological Research (MCR LTER) site under Grant No. OCE 2224354 (and earlier awards). Additional financial support to the MCR LTER site was provided through a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2023).

openCC (other)Sep 2023View details →
edi52/100

Modeling the effects of lake morphology on chloride retention and salt-driven stratification in two urban lakes in St. Paul, MN

Road salt inputs have caused widespread salinization of urban lakes in northern temperate regions. Watershed characteristics are known to be important drivers of lake chloride concentrations, but there has been less focus on how lake morphometry influences seasonal and interannual dynamics in lake chloride, and how these chloride levels may alter mixing in the water column. We analyzed chloride retention for two urban lakes (Como Lake and Lake McCarrons) in Saint Paul, Minnesota, that are in adjacent watersheds and have similar surface areas, but differ in depth and water residence time. Summer chloride concentrations were negatively related to total summer precipitation for Como Lake (maximum depth 2.2 m), but the relationship was less strong for Lake McCarrons (maximum depth 7.6 m). We used a zero-dimensional model to simulate chloride dynamics in both lakes and tracked the fate of chloride over time. In Como Lake, the mass of chloride in the lake turns over within three years, whereas chloride inputs are retained for >10 years in Lake McCarrons. We then used a one-dimensional hydrodynamic lake model (GLM-AED) to examine how lake depth affects how current chloride loading rates alter lake mixing. Salt inputs significantly extended the duration of summer stratification for simulated lakes with depths of 8 m or more, and salt inputs increased the number of days of hypoxia and anoxia across all depths. These results underscore the importance of considering lake morphometry in understanding the effects of salt inputs on lake ecosystems.

openCC (other)Aug 2025View details →
edi52/100

Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine: otolith data, model data, and post-processed model data products

This dataset includes hatch and larval period for sand lance collected in 2019 and results from particle tracking runs of simulated sand lance larvae throughout the Northeast U.S. Shelf as part of Long-Term Ecological Research (NES-LTER). Release dates vary by region, corresponding to hatch and settlement dates of settling sand lance collected in 2019. Particles were depth-keeping throughout the upper 40 m to best replicate our understanding of the vertical distribution of sand lance larvae. Data were used to determine the average particle transport pathways from these sand lance habitats, including connectivity among the three hotspots, and spatial variability of connectivity within each hotspot. Further information can be found within the manuscript: Suca, J. J., Ji, R., Baumann, H., Pham, K., Silva, T. L., Wiley, D. N., Feng, Z., & Llopiz, J. K. (2022). Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine. Fisheries Oceanography, 31( 3), 333-352. https://doi.org/10.1111/fog.12580

openCC (other)Jun 2022View details →
edi52/100

Modeled Organic Carbon, Dissolved Oxygen, and Secchi for six Wisconsin Lakes, 1995-2014

This data package contains model output data, driving data, and supplemental information for a two-layer modeling study that investigated organic carbon and oxygen dynamics within six Wisconsin lakes over a twenty-year period (1995-2014). The six lakes are Lake Mendota, Lake Monona, Trout Lake, Allequash Lake, Big Muskellunge Lake, and Sparkling Lake. The model output includes daily predictions of six state variables: labile particulate organic carbon, recalcitrant particulate organic carbon, labile dissolved organic carbon, recalcitrant dissolved organic carbon, dissolved oxygen, and Secchi depth. The output also includes daily predictions of physical and metabolism fluxes that were used in the prediction of the state variables. This data package also contains model driving data for each lake and other supplemental information that was calculated during the modeling runs.

openCC0Nov 2022View details →
edi52/100

SBC LTER: Daily averages of modeled significant wave height (Hs) and peak wave period (Tp) in the Santa Barbara Coastal area from the Coastal Data Information Program - Monitoring and Prediction System (CDIP MOP)

From http://cdip.ucsb.edu: The Coastal Data Information Program (CDIP) is a research group at Scripps Institution of Oceanography that monitors coastal waves and nearshore sand levels on regional scales. CDIP maintains a network of optimally-placed, directional wave buoys from San Diego to Eureka. The buoy measurements are used to initialize a high spatial resolution (100m x 100m) linear spectral wave propagation model. The resulting hourly hindcasts and nowcasts of CA coastal wave conditions have a level of accuracy that is not possible with more traditional wind-wave generation models that are initialized with modeled wind fields.

openCC (other)Jun 2025View details →
OpenNeuro48/100

Model-based fMRI reveals co-existing specific and generalized concept representations

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro48/100

In vivo T1w MRI of a TDP-43 knock-in mouse model of ALS-FTD

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

SERVS lightcone of model galaxies.

<p>This SERVS lightcone of model galaxies has been constructed using the Lagos12 Galform model using the techniques described in Merson et al. 2013. The lightcone covers the redshift range z = 0.0 to z = 6.0 and has a sky coverage of 18.09 deg<sup>2</sup>, centred on a sky position of (RA,DEC) = (93.5◦, 7.5◦ ).</p> <p>The SERVS lightcone contains 1518854 model galaxies with apparent, dust attenuated magnitudes in the Spizter 3.6 microns bands down to 2micro Jy (AB=23.1).</p> <p>The lightcone was constructed on the Millennium dark matter only N-body simulation.</p> <p>For each model galaxy, model multi-wavelength coverage is provided together with a range of global properties. Further details are provided in the SERVS.Lagos12.DB.Mill1.lightcone.readme.pdf document, within this dataset.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Models for "A data-driven approach to studying changing vocabularies in historical newspaper collections"

<p>NOTE: This is a badly rendered version of the README within the archive.</p> <p><strong>A data-driven approach to studying changing vocabularies in historical newspaper collections</strong></p> <p>Simon Hengchen,* Ruben Ros,** Jani Marjanen,*** Mikko Tolonen***</p> <p>*<a href="https://spraakbanken.gu.se/en/about/staff/simon">Spr&aring;kbanken Text</a>, University of Gothenburg, Sweden and <a href="https://iguanodon.ai">iguanodon.ai</a>, Belgium: firstname.lastname@gu.se<br> **<a href="https://www.c2dh.uni.lu/people/ruben-ros">Centre for Contemporary and Digital History (C2DH)</a>, University of Luxembourg:&nbsp;firstname.lastname@uni.lu<br> ***<a href="https://www.helsinki.fi/en/researchgroups/computational-history">COMHIS</a>, University of Helsinki:&nbsp;<a href="mailto:firstname.lastname@helsinki.fi">firstname.lastname@helsinki.fi</a>;</p> <p>These are the supplementary materials for the DH2019 paper&nbsp;<em>A data-driven approach to the changing vocabulary of the &lsquo;nation&rsquo; in English, Dutch, Swedish and Finnish newspapers, 1750-1950</em>, as well as the 2021 Digital Scholarship in the Humanities publication available in OpenAccess: <a href="https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793">https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793</a>. If you end up using whole or parts of this resource, please use the following citation(s):</p> <ul> <li>Hengchen, S., Ros, R., and Marjanen, J. (2019). A data-driven approach to the changing vocabulary of the &#39;nation&#39; in English, Dutch, Swedish and Finnish newspapers, 1750-1950. In&nbsp;<em>Proceedings of the Digital Humanities (DH) conference 2019, Utrecht, The Netherlands</em></li> </ul> <p>and/or:</p> <ul> <li>Hengchen, S., Ros, R., Marjanen, J. and Tolonen, M., 2021. A data-driven approach to studying changing vocabularies in historical newspaper collections. Digital Scholarship in the Humanities, 36(Supplement_2), pp.ii109-ii126.</li> </ul> <p>or alternatively use one of the following&nbsp;<code>bib</code>s:</p> <pre><code>@inproceedings{hengchen2019nation, title="A data-driven approach to the changing vocabulary of the 'nation' in {E}nglish, {D}utch, {S}wedish and {F}innish newspapers, 1750-1950.", author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani}, year={2019}, address = "Utrecht, The Netherlands", booktitle={Proceedings of the Digital Humanities (DH) conference 2019} }</code></pre> <pre><code>@article{hengchen2021data, title={A data-driven approach to studying changing vocabularies in historical newspaper collections}, author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani and Tolonen, Mikko}, journal={Digital Scholarship in the Humanities}, volume={36}, number={Supplement\_2}, pages={ii109--ii126}, year={2021}, publisher={Oxford University Press} }</code></pre> <p>&nbsp;</p> <p>Files</p> <p>This archive contains two folders -- one per diachronic representation method -- as well as this README. The folders each contain four folders, which contain the models for their respective languages. As can be inferred from the small datasize, most of the earlier models are not reliable and should not be used, but are still made available. This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p><strong>Source material</strong></p> <p>Finnish:</p> <p>The models were created with data from the Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland (National Library of Finland, 2011). We used everything in the corpus.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h fi* 12M fi_1820_SGNS_corpus_file.gensim 89M fi_1840_SGNS_corpus_file.gensim 797M fi_1860_SGNS_corpus_file.gensim 7.0G fi_1880_SGNS_corpus_file.gensim 22G fi_1900_SGNS_corpus_file.gensim</code></pre> <p>Swedish:</p> <p>The models were created with data from the Kubhist 2 corpus (Spr&aring;kbanken) -- more precisely, the data dumps available at&nbsp;<a href="https://spraakbanken.gu.se/lb/resurser/meningsmangder/">https://spraakbanken.gu.se</a>. After a manual evaluation of Swedish embeddings trained without pre-processing seemed to show that the embeddings were of low quality, we retrained models, only keeping sentences that were at least 10 tokens long and were constituted of at least 50% of lemmas as per the KORP processing pipeline (Borin et al, 2012).</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h sv* 1.6M sv_1740_SGNS_corpus_file.gensim 44M sv_1760_SGNS_corpus_file.gensim 124M sv_1780_SGNS_corpus_file.gensim 228M sv_1800_SGNS_corpus_file.gensim 678M sv_1820_SGNS_corpus_file.gensim 1.6G sv_1840_SGNS_corpus_file.gensim 4.5G sv_1860_SGNS_corpus_file.gensim 6.5G sv_1880_SGNS_corpus_file.gensim 113M sv_1900_SGNS_corpus_file.gensim</code></pre> <p>Dutch:</p> <p>The models were created with data from the Delpher newspaper archive (Royal Dutch Library, 2017), through data dumps for newspapers until and including 1876, and through API hits for articles from 1877 to 1899 (included).</p> <ul> <li>For anything pre-1877 we discarded full texts that had, in the metadata, anything else than exclusively&nbsp;<code>nl</code>&nbsp;or&nbsp;<code>NL</code>&nbsp;as a language tag.</li> <li>For the full texts between 1877 and 1899: we queried the API for all items in the &ldquo;artikel&rdquo; category that contained the determiner&nbsp;<code>de</code>.</li> </ul> <p>Our assumption was that most articles should contain&nbsp;<code>de</code>&nbsp;at least once, and those that didn&#39;t were too short to be deemed interesting. A subsequent study showed that was not exactly the case, but we were reassured by the fact that left-out articles were probably &quot;shipping or financial reports&quot; (thanks go to Melvin Wevers). We also did not include the colonial newspapers for our embeddings. This is motivated by our research questions. A list of removed newspapers is available on request.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h nl* 6.8M nl_1620_SGNS_corpus_file.gensim 7.9M nl_1640_SGNS_corpus_file.gensim 43M nl_1660_SGNS_corpus_file.gensim 78M nl_1680_SGNS_corpus_file.gensim 138M nl_1700_SGNS_corpus_file.gensim 243M nl_1720_SGNS_corpus_file.gensim 287M nl_1740_SGNS_corpus_file.gensim 431M nl_1760_SGNS_corpus_file.gensim 825M nl_1780_SGNS_corpus_file.gensim 1.2G nl_1800_SGNS_corpus_file.gensim 1.8G nl_1820_SGNS_corpus_file.gensim 3.1G nl_1840_SGNS_corpus_file.gensim 5.2G nl_1860_SGNS_corpus_file.gensim 13G nl_1880_SGNS_corpus_file.gensim</code></pre> <p>English:</p> <p>The models were created with data from the British Library Newspapers collection (<a href="https://www.gale.com/intl/primary-sources/british-library-newspapers%5D">link</a>), the Nichols collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-burney-newspapers-collection">link</a>), and the Burney collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-nichols-newspapers-collection">link</a>). We used everything in the corpora. For English, only SGNS_ALIGN models are available. We thank Gale Cengage for their help with this project.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h en* 4.3M en_1620_SGNS_corpus_file.gensim 11M en_1640_SGNS_corpus_file.gensim 11M en_1660_SGNS_corpus_file.gensim 106M en_1680_SGNS_corpus_file.gensim 409M en_1700_SGNS_corpus_file.gensim 1.7G en_1720_SGNS_corpus_file.gensim 834M en_1740_SGNS_corpus_file.gensim 2.4G en_1760_SGNS_corpus_file.gensim 5.3G en_1780_SGNS_corpus_file.gensim 5.5G en_1800_SGNS_corpus_file.gensim 15G en_1820_SGNS_corpus_file.gensim 42G en_1840_SGNS_corpus_file.gensim 65G en_1860_SGNS_corpus_file.gensim 88G en_1880_SGNS_corpus_file.gensim 26G en_1900_SGNS_corpus_file.gensim 21G en_1920_SGNS_corpus_file.gensim 6.3G en_1940_SGNS_corpus_file.gensim</code></pre> <p><strong>Word embeddings</strong></p> <p>For every language, we train diachronic embeddings as follows. We divide the data in 20-year time bins. We train SGNS_UPDATE and SGNS_ALIGN models. Current research on German (Schlechtweg et al, 2019) and English (Shoemark et al, 2019) indicates you should use the SGNS_ALIGN models.&nbsp;<strong>For EN, FI, NL, no tokens (including punctuation) were removed nor altered, aside from lowercasing</strong>. For SV, see above. Parameters are as follows: SGNS architecture (Mikolov et al 2013), window size of 5, frequency threshold of 100, 5 epochs, 300 dimensions (or 100 for EN).</p> <ul> <li>For SGNS_UPDATE: We first train a model for the first time bin&nbsp;<code>t</code>. To train the model for&nbsp;<code>t+1</code>, we use the&nbsp;<code>t</code>&nbsp;model to initialise the vectors for&nbsp;<code>t+1</code>, set the learning rate to correspond to the end learning rate of&nbsp;<code>t</code>, and continue training. This approach, closely following Kim et al (2014), has the advantage of avoiding the need for post-training vector space alignment.</li> </ul> <p>The Python snippet below, which makes use of gensim (Rehurek and Sojka, 2010), illustrates the approach. Special thanks go to Sara Budts.</p> <pre><code>## dict_files[key] is a dictionary with double decades as keys and a corresponding LineSentence object as value: https://radimrehurek.com/gensim/models/word2vec.html#gensim.models.word2vec.LineSentence count = 0 for key in sorted(list(dict_files.keys())): if count == 0: ## This is the first model. model = gensim.models.Word2Vec(corpus_file=dict_files[key], min_count=100, sg=1 ,size=300, workers=64, seed=1830, iter=5) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) print("Model saved, on to the next\n") count += 1 if count &gt; 0: ## this is for the subsequent models. print("model for double decade starting in",str(key)) model = gensim.models.Word2Vec.load(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin-20)+".w2v")) print("previous model loaded") model.build_vocab(corpus_file=dict_files[key], update=True) model.train(corpus_file=dict_files[key], total_words = model.corpus_count, total_examples = model.corpus_count, start_alpha = model.alpha, end_alpha = model.min_alpha, epochs=model.epochs) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) </code></pre> <ul> <li>For SGNS_ALIGN: We independently train models for all time bins. The models in this repository are&nbsp;<em>NOT</em>&nbsp;aligned, leaving you the choice of how to align them. For example,&nbsp;<a href="https://gist.github.com/quadrismegistus/09a93e219a6ffc4f216fb85235535faf">here</a>&nbsp;is a link to code by Ryan Heuser to do just that. Models were trained with the&nbsp;<code>count == 0</code>&nbsp;scenario in the snippet above.</li> </ul> <p><strong>Acknowledgments</strong></p> <p>This work has been supported by the European Union&#39;s Horizon 2020 research and innovation programme under grant 770299&nbsp;<a href="https://www.newseye.eu/">NewsEye</a>. Specials thanks go to the data providers/collection-holding institutions: the Finnish Language Bank, the Swedish Language Bank, the Royal Dutch Library, and Gale Cengage.</p> <p>The authors would like to thank the following persons and group, listed alphabetically: Antoine Doucet, Antti Kanner, Axel-Jean Caurant, Dominik Schlechtweg, Eetu M&auml;kel&auml;, Elaine Zosa, Estelle Bunout, Haim Dubossarsky, Joris van Eijnatten, Krister Lind&eacute;n, Lars Borin, Lidia Pivovarova, Melvin Wevers, Nina Tahmasebi, Sara Budts, Senka Drobac, Tanja S&auml;ily, the COMHIS group, and Steven Claeyssens. Computational resources were provided by CSC &ndash; IT Center for Science Ltd.</p> <p><strong>References</strong></p> <p>Borin, L., Forsberg, M., Roxendal, J. (2012). Korp-the corpus infrastructure of Spr&auml;kbanken,in: LREC. pp. 474&ndash;478.</p> <p>Kim, Y., Chiu, Y.I., Hanaki, K., Hegde, D. and Petrov, S. (2014). Temporal Analysis of Language through Neural Language Models.&nbsp;<em>ACL 2014</em>, p.61.</p> <p>Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space.&nbsp;<em>arXiv preprint arXiv:1301.3781</em>.</p> <p>National Library of Finland (2011).&nbsp;<em>The Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland, Kielipankki Version</em>&nbsp;[text corpus]. Kielipankki. Retrieved from&nbsp;<a href="http://urn.fi/urn:nbn:fi:lb-2016050302">http://urn.fi/urn:nbn:fi:lb-2016050302</a>.</p> <p>Rehurek, R. and Sojka, P. (2010). Software framework for topic modelling with large corpora. In&nbsp;<em>Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</em>.</p> <p>Royal Dutch Library (2017).&nbsp;<em>Delpher open krantenarchief (1.0)</em>. Den Haag, 2017.</p> <p>Schlechtweg D., H&auml;tty A, del Tredici M., and Schulte im Walde S. (2019). A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In&nbsp;<em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, Florence, Italy. ACL.</p> <p>Shoemark, P., Liza, F.F., Nguyen, D., Hale, S. and McGillivray, B. (2019). Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings. In&nbsp;<em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 66-76)</em>, Hong Kong.</p> <p>Spr&aring;kbanken.&nbsp;<em>The Kubhist Corpus</em>. Department of Swedish, University of Gothenburg.&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">https://spraakbanken.gu.se/korp/?mode=kubhist</a>.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Input Runoff Data for RAPID Model Pre-Processor (RRR) from ECMWF ERA-Interim/Land

<p>This database can be used as the input runoff files in the RAPID model [<em>David et al.,</em> 2011] pre-processor (RRR). The runoff files were acquired/derived from the ECMWF ERA-Interim/Land [<em>Balsamo et al.,</em> 2015] outputs, available from ECMWF Data Server. The ERA-Interim/Land outputs are available in daily temporal resolution. The database contains the following files;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ECMWF_Interim_Land_<strong><em>yyyy</em></strong>.tar.gz&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (Note: <strong><em>yyyy</em></strong> = 2000 to 2009)</p> <p>&nbsp;</p> <p>Note: These runoff data were used by <em>Sikder et al.</em> [2019] to assess the performance of available global LSM runoffs in South and Southeast Asian river basins.</p> <p>&nbsp;</p> <p>Other necessary links associated with this database:</p> <p>RAPID model: <a href="https://github.com/c-h-david/rapid">https://github.com/c-h-david/rapid</a></p> <p>RAPID model pre-processor (rrr): <a href="https://github.com/c-h-david/rrr">https://github.com/c-h-david/rrr</a></p> <p>ECMWF outputs: <a href="https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land">https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land</a></p> <p>&nbsp;</p> <p>References:</p> <p>Balsamo, G., Albergel, C., Beljaars, A., Boussetta, S., Brun, E., Cloke, H., et al. [2015], ERA-Interim/Land: a global land surface reanalysis data set, Hydrol. Earth Syst. Sci., 19, 389&ndash;407, <a href="https://doi.org/10.5194/hess-19-389-2015">https://doi.org/10.5194/hess-19-389-2015</a></p> <p>David, C. H., D. R. Maidment, G. Y. Niu, Z. L. Yang, F. Habets, and V. Eijkhout [2011], River network routing on the NHDPlus dataset, J. Hydrometeorol., 12, 913&ndash;934, <a href="https://doi.org/10.1175/2011JHM1345.1">https://doi.org/10.1175/2011JHM1345.1</a></p> <p>Sikder, M. S., C. H. David, G. H. Allen, X. Qiao, E. J. Nelson, and M. A. Matin [2019], Evaluation of Available Global Runoff Datasets Through a River Model in Support of Transboundary Water Management in South and Southeast Asia, Front. Environ. Sci., 7:171, <a href="https://doi.org/10.3389/fenvs.2019.00171">https://doi.org/10.3389/fenvs.2019.00171</a></p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

RAPID Model Input Files for Mekong-Indus-Ganges-Brahmaputra-Megna (MIGBM) River Basins

<p>This database contains Inputs and intermediate files of the RAPID model pre-processor (RRR), and also outputs from the RRR (<em>i.e.</em>, Inputs for RAPID); which were used by <em>Sikder et al.</em> [2019] to assess the performance of available global LSM runoffs in South and Southeast Asian river basins. If you use this RAPID Model Input Files for Mekong-Indus-Ganges-Brahmaputra-Megna (MIGBM) River Basins in your work, please cite: <em>Sikder et al.</em>, [2019], Evaluation of Available Global Runoff Datasets Through a River Model in Support of Transboundary Water Management in South and Southeast Asia, Front. Environ. Sci., 7:171, <a href="https://doi.org/10.3389/fenvs.2019.00171">https://doi.org/10.3389/fenvs.2019.00171</a>.</p> <p>The database contains;</p> <ul> <li>Global River basin and Network Shapefiles: HydroSHEDS.tar.gz</li> <li>Extracted Basin Shapefile: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; MIGBM_basin.tar.gz</li> <li>Extracted River Network Shapefiles:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; MIGBM_<strong><em>res</em></strong>_ntwk.tar.gz&nbsp;&nbsp;&nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Catchment Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; rapid_catchment_as_<strong><em>riv</em></strong>_res.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; &nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Connectivity Files: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; rapid_connect_<strong><em>res</em></strong>_MIGBM.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp; &nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Coordinate Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; coords_<strong><em>res</em></strong>_MIGBM.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Base Parameter Files: &nbsp; &nbsp; <strong><em>p</em></strong>fac_<strong><em>res</em></strong>_MIGBM_1km_hour.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; (Note: <strong><em>p</em></strong> = k or x; <strong><em>res</em></strong> = fine or coarse)</li> <li>Sort Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sort_<strong><em>res</em></strong>_MIGBM_topo.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Sorted Basin Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; riv_bas_id_<strong><em>res</em></strong>_MIGBM_topo.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Coupling Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; rapid_coupling.tar.gz</li> <li>Parameter Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; rapid_param.tar.gz</li> <li>Volume Files:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; m3_riv_<strong><em>res</em></strong>_MIGBM_20000101_20091231_<strong><em>prj</em></strong>_<strong><em>LSMsr</em></strong>_<strong><em>tr</em></strong>_utc.nc&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (Note: <strong><em>res</em></strong> = fine or coarse; <strong><em>prj</em></strong> = GLDAS or GLDAS.2.0 or GLDAS.2.1 or ECMWF; <strong><em>LSM</em></strong> = CLM, MOS, NOAH, VIC, ERAint; <strong><em>sr</em></strong> = 10 or 025; <strong><em>tr</em></strong> = 3H or D)</li> </ul> <p>&nbsp;</p> <p>Other necessary links associated with this database:</p> <p>RAPID model: <a href="https://github.com/c-h-david/rapid">https://github.com/c-h-david/rapid</a></p> <p>RAPID model pre-processor (rrr): <a href="https://github.com/c-h-david/rrr">https://github.com/c-h-david/rrr</a></p> <p>GLDAS outputs: <a href="https://disc.gsfc.nasa.gov/datasets?keywords=GLDAS">https://disc.gsfc.nasa.gov/datasets?keywords=GLDAS</a></p> <p>ECMWF outputs: <a href="https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land">https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land</a></p> <p>&nbsp;</p> <p>References:</p> <p>Balsamo, G., Albergel, C., Beljaars, A., Boussetta, S., Brun, E., Cloke, H., et al. [2015], ERA-Interim/Land: a global land surface reanalysis data set, Hydrol. Earth Syst. Sci., 19, 389&ndash;407, <a href="https://doi.org/10.5194/hess-19-389-2015">https://doi.org/10.5194/hess-19-389-2015</a></p> <p>David, C. H., D. R. Maidment, G. Y. Niu, Z. L. Yang, F. Habets, and V. Eijkhout [2011], River network routing on the NHDPlus dataset, J. Hydrometeorol., 12, 913&ndash;934, <a href="https://doi.org/10.1175/2011JHM1345.1">https://doi.org/10.1175/2011JHM1345.1</a></p> <p>Rodell, M., P. R. Houser, U. Jambor, J. Gottschalck, K. Mitchell, C.-J. Meng, et al. [2004], The global land data assimilation system, Bull. Am. Meteorol. Soc. 85, 381&ndash;394, <a href="https://doi.org/10.1175/BAMS-85-3-381">https://doi.org/10.1175/BAMS-85-3-381</a></p> <p>Sikder, M. S., C. H. David, G. H. Allen, X. Qiao, E. J. Nelson, and M. A. Matin [2019], Evaluation of Available Global Runoff Datasets Through a River Model in Support of Transboundary Water Management in South and Southeast Asia, Front. Environ. Sci., 7:171, <a href="https://doi.org/10.3389/fenvs.2019.00171">https://doi.org/10.3389/fenvs.2019.00171</a></p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Replication package of "Search-based Crash Reproduction using Behavioral Model Seeding"

<p>Search-based crash reproduction approaches assist developers during debugging by generating a test case which reproduces a crash given its stack trace. One of the fundamental steps of this approach is creating objects needed to trigger the crash. One way to overcome this limitation is seeding: using information about the application during the search process. With seeding, the existing usages of classes can be used in the<br> search process to produce realistic sequences of method calls which create the required objects. In this study, we introduce behavioral model seeding: a new seeding method which learns class usages from both<br> the system under test and existing test cases. Learned usages are then synthesized in a behavioral model (state machine). Then, this model serves to guide the evolutionary process. To assess behavioral model-seeding, we evaluate it against test-seeding (the state-of-the-art technique for seeding realistic objects) and no-seeding (without seeding any class usage). For this evaluation, we use a benchmark of 122 hard-to-reproduce crashes stemming from six open-source projects. Our results indicate that behavioral model-seeding outperforms both test seeding and no-seeding by a minimum of 6% without any notable negative impact on efficiency.</p>

opencc-by-4.0Oct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record