Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Additional Figures for winning models for sample in A Comparative L-dwarf Sample Exploring the Interplay Between Atmospheric Assumptions and Data Properties
<p>Additional Figures for winning models for sample in <em>A Comparative L-dwarf Sample Exploring the Interplay Between Atmospheric Assumptions and Data Properties (<a href="https://arxiv.org/pdf/2209.02754.pdf">https://arxiv.org/pdf/2209.02754.pdf</a>).</em></p> <p>Model naming key: NC = cloud-free, d2_89 = power-law deck cloud</p> <p>SDSS J1416+1348A: Winning model: power-law deck cloud</p> <p>Spectral Type Comparison J1526+2043 Winning model: Cloud-free</p> <p>Temperature Comparisons</p> <p>J1539-0520 Winning model: Power-law deck cloud and cloud-free tied.</p> <p>J0539-0059 Winning model: Power-law deck cloud and cloud-free tied. </p> <p><br> </p>
Data for: Sticker-and-Spacer Model for Amyloid Beta Condensation and Fibrillation
<p>Data for: Sticker-and-Spacer Model for Amyloid Beta Condensation and Fibrillation</p> <p>Preprint of the paper on bioRxiv:</p> <p>https://www.biorxiv.org/content/10.1101/2022.06.04.494837v1</p>
Data from: Hindcast-validated species distribution models reveal future vulnerabilities of mangroves and salt marsh species
<p>Rapid climate change threatens biodiversity via habitat loss, range shifts, increases in invasive species, novel species interactions, and other unforeseen changes. Coastal and estuarine species are especially vulnerable to the impacts of climate change due to sea level rise and may be severely impacted in the next several decades. Species distribution modeling can project the potential future distributions of species under scenarios of climate change using bioclimatic data and georeferenced occurrence data. However, models projecting suitable habitat into the future are impossible to ground truth. One solution is to develop species distribution models for the present and project them to periods in the recent past where distributions are known to test model performance before making projections into the future. Here, we develop models using abiotic environmental variables to quantify the current suitable habitat available to eight Neotropical coastal species: four mangrove species and four salt marsh species. Using a novel model validation approach that leverages newly available monthly climatic data from 1960-2018, we project these niche models into two time periods in the recent past (i.e., within the past half-century) when either mangrove or salt marsh dominance was documented via other data sources. Models were hindcast-validated and then used to project the suitable habitat of all species at four time periods in the future under a model of climate change. For all future time periods, the projected suitable habitat of mangrove species decreased, and suitable habitat declined more severely in salt marsh species.</p>
Climate change effects on deep-water corals – habitat suitability model input data
<p>Deep-water corals are protected in the seas around New Zealand by legislation that prohibits intentional damage and removal, and by marine protected areas where bottom trawling is prohibited. However, these measures do not protect them from the impacts of a changing climate and ocean acidification. To enable adequate future protection from these threats we require knowledge of the present distribution of corals and the environmental conditions that determine their preferred habitat, as well as the likely future changes in these conditions, so that we can identify areas for potential refugia.</p> <p>In this study, we built habitat suitability models for 12 taxa of deep-water corals using a comprehensive set of sample data and predicted present and future seafloor environmental conditions from an earth system model specifically tailored for the South Pacific. These models predicted that for most taxa there will be substantial shifts in the location of the most suitable habitat and decreases in the area of such habitat by the end of the 21st century, driven primarily by decreases in seafloor oxygen concentrations, shoaling of aragonite and calcite saturation horizons, and increases in nitrogen concentrations. The current network of protected areas in the region appear to provide little protection for most coral taxa, as there is little overlap with areas of highest habitat suitability, either in the present or the future. We recommend an urgent re-examination of the spatial distribution of protected areas for deep-water corals in the region, utilising spatial planning software that can balance protection requirements against value from fishing and mineral resources, take into account the current status of the coral habitats after decades of bottom trawling, and consider connectivity pathways for colonisation of corals into potential refugia.</p>
Data and code availibility for the paper "Ice-nucleating agents in sea spray aerosol identified and quantified with a holistic multi-modal freezing model" by Alpert et al.
<p>The data and codes used in the paper "Ice-nucleating agents in sea spray aerosol identified and quantified with a holistic multi-modal freezing model" by Alpert et al., published in <em>Science Advances</em> are included in this collection. Detailed descriptions of the files are given in the readme file.</p>
LaMEM source code and input files corresponding to Present‐day upper‐mantle architecture of the Alps: Insights from data‐driven dynamic modelling
<p>This repository contains LaMEM source code and input files for the models presented in Kumar, A., Cacace, M., Scheck-Wenderoth, M., Götze, H.-J., & Kaus, B. J. P. (2022). Present-day upper-mantle architecture of the Alps: Insights from data-driven dynamic modeling. Geophysical Research Letters, 49, e2022GL099476. https://doi. org/10.1029/2022GL099476</p>
Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"
<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States. </p> <p>The set of 47 news media outlets analysed (listed in Figure 1 of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets’ ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript. </p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets’ online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles’ HTML raw data using outlet-specific XPath expressions. </p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000’s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000. Figure S 1 in the SI reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small. Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth. </p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely: </p> <p>Siebert/sentiment-roberta-large-english from https://huggingface.co/siebert/sentiment-roberta-large-english. This model is a fine-tuned checkpoint of <a href="https://huggingface.co/roberta-large">RoBERTa-large</a> (<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with sentiment-roberta-large-english </p> <p>DistilRoberta j-hartmann/emotion-english-distilroberta-base from https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of <a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with Ekman's 6 basic emotions, plus a neutral class. The model was trained on 6 diverse datasets. Please refer to the original author at https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the siebert/sentiment-roberta-large-english Transformer model. https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset. https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the j-hartmann/emotion-english-distilroberta-base Transformer model. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>
Research data for "Indirect learning and physically guided validation of interatomic potential models"
<p>This dataset contains structural data, potential parameter files, and data shown in the plots for the publication "Indirect learning and physically guided validation of interatomic potential models". Details of the contents can be found in README.txt.</p>
Supplementary data: "Open modeling of electricity and heat demand curves for all residential buildings in Germany"
<p>This repository contains supplementary data for the paper <a href="https://doi.org/10.1186/s42162-022-00201-y"><em> "Open modeling of electricity and heat demand curves for all residential buildings in Germany"</em></a>.</p> <p>See <em>README.md</em> / <em>README.pdf</em> for further details.</p> <p><strong>Citing</strong></p> <p>Please cite as:</p> <p><em>Büttner, C., Amme, J., Endres, J. et al. Open modeling of electricity and heat demand curves for all residential buildings in Germany. Energy Inform 5 (Suppl 1), 21 (2022).</em></p> <p><strong>Funding</strong></p> <p>The authors thank the Federal Ministry for Economic Affairs and Climate Action for funding the research project eGon (funding code: 03EI1002).</p> <p> </p>
ODP Site 1249, ODP Site 1252, and IODP Site U1325: X-ray fluoresence core scanning, laser diffraction grain size, CNS elemental/isotopic, environmental magnetism, and age model data
<p>We present data used as part of an integrative early diagenesis study (submitted September 2022 to Marine Geology) focused on identifying zones of magnetite dissolution and pyrite precipitation in which magnetic susceptibility records are altered in gas-hydrate bearing sediments on the Cascadia Margin using archived cores from the Ocean Drilling Program (ODP) and Integrated Ocean Drilling Program (IODP). We analyzed the upper 85 to 100 m below seafloor (mbsf) from ODP Sites 1249 and 1252, and IODP Site U1325. ODP 1249 is at the summit of Hydrate Ridge in an area of active methane seepage and massive gas hydrate accumulations and ODP 1252 is in a nearby slope basin with little occurrence of hydrate. IODP Site U1325 is on the northern Cascadia Margin in a slope basin, with turbidite-hosted gas hydrate. We also include XRF data from the upper sections of ODP Site 1251, IODP U1327, and U1328.</p> <p>We measured X-ray fluorescence using an Avaatech core scanner at the IODP Gulf Coast Repository at Texas A&M University. We measured total carbon, total organic carbon (TOC), total nitrogen, and total sulfur using a Perkin Elmer 2400 Series CHNS/O Analyzer at the university of New Hampshire (ODP Site 1249 and 1252 only). A subset was analyzed for δ<sup>13</sup>C-TOC using a Costech ECS 4010 elemental analyzer interfaced with a Thermo Finnegan Delta Plus XP continuous flow isotope ratio mass spectrometer at Washington State University. Grain size was measured with a Malvern Mastersizer 2000 laser diffraction particle size analyzer and Hydro 2000G dispersal unit at the University of New Hampshire. Mass frequency-dependent magnetic susceptibility was measured using a Bartington MS2 Magnetic Susceptibility Meter and Bartington MS2B dual frequency sensor (Site U1325 only). Isothermal remanent magnetization and thermal demagnetization curves were measured using a HSM2 SQUID-based Spinner Magnetometer with an ASC Scientific IM-10-30 Impulse Magnetizer and ASC Scientific TD-48SC magnetically-shielded oven (Site U1325 only). Radiocarbon was measured on mixed planktic foraminifers at the Radiocarbon was measured at National Ocean Sciences Accelerator Mass Spectrometry (NOSAMS) facility at Woods Hole Oceanographic Institution (ODP Site 1252 and IODP Site U1325).. Radiocarbon ages were calibrated to calendar ages using CALIB 8.2 and the Marine20 calibration curve. For ODP Site 1252 we used a reservoir correction of 230 ± 50 years (Yaquina Bay, Oregon, USA) and for IODP Site U1325 we used a reservoir correction of 202 ± 50 years (Amphitrite Point, British Columbia, Canada). δ<sup>18</sup>O was measured on benthic foraminifer <em>Uvigerina peregrina</em> tests using a Finnegan MAT 252 isotope ratio mass spectrometer with Kiel III device at the Oregon State University Stable Isotope Laboratory (ODP 1252) and a Finnegan MAT 253 isotope ratio mass spectrometer with Kiel IV device at the University of Michigan Stable Isotope Laboratory (ODP Site 1249 and IODP Site U1325).</p>
FESOM model data and particle tracking data used in publication 'Cross-shelf transport of Barents Sea dense water as a sink for CO2 in the Arctic Ocean'
<p>FESOM model data and particle tracking data used in the paper 'Cross-shelf transport of Barents Sea dense water as a sink for CO2 in the Arctic Ocean" by Andreas Rogge at al.</p> <p>1) FESOM velocity fields averaged over the top 200 m water depth, averaged over the time period 2015-2018.</p> <p>2) FESOM transect at 95°E in the Arctic Ocean (temperature, salinity and velocity), averaged over the time period 2015-2018.</p> <p>3a) Particle back-tracking data based on daily FESOM velocity fields in netcdf format. Particles were released at 95°E every 14 days during the year 2018 and tracked until they reached the surface. Three different constant sinking velocities were used, representative for small and large non-ballasted particles and small ballasted particles.</p> <p>3b) Distribution of particles at the surface for the experiments with three different sinking velocities as mat files. </p>
Data of "Ductile fracture of high entropy alloys: from the design of an experimental campaign to the development of a micromechanics-based modeling framework"
<p>Data related to the publication (we would be grateful if you could cite the paper in the case in which you are using the data):</p> <p>title = "Ductile fracture of high entropy alloys: from the design of an experimental campaign to the development of a micromechanics-based modeling framework",<br> journal = "Engineering Fracture Mechanics",<br> year = "2022",<br> volume = "275",<br> pages = "108844 ",<br> doi = "https://doi.org/10.1016/j.engfracmech.2022.108844",<br> author = "Antoine Hilhorst, Julien Leclerc, Thomas Pardoen, Pascal J. Jacques, Ludovic Noels, Van-Dung Nguyen"</p> <p>New version following review.</p> <p> </p> <p> </p>
Data from: Petrogenetic studies of Permian pegmatites in the Chinese Altay: implications for a two-stage post-collisional magmatism model
<p>Understanding the petrogenesis of rare-metal pegmatites is important for understanding ore-forming processes and their tectonic settings. In this study, we performed zircon U-Pb geochronological and Hf-O isotopic analyses of the Xiaokalasu (XKLS), Dakalasu (DKLS) and Yelaman (YLM) pegmatites in the Chinese Altay orogen. These pegmatites have low εHf(t) values (-0.6 ~ +4.3), two-stage model ages of 989 ~ 1293 Ma, and high δ<sup>18</sup>O values (+6.52 ~ +11.31), indicating that they may have been derived from the anatexis of mature sedimentary rocks in the deep crust, with a small amount of mantle-derived or juvenile material. Geochronological and Hf-O isotopic data for granitic intrusions in the Chinese Altay Mountains indicate that the εHf(t) values decreased from the Permian to the Triassic, which implies that two-stage post-collisional magmatism occurred in this region. During the Permian, the thin lower crust was cold; thus, magmatism likely originated in the deep crust close to the Moho surface and involved intense mantle-crust interactions. During the Triassic, asthenospheric upwelling provided heat to the lower crust, which increased the geothermal gradient and led to the anatexis of shallow crustal material.</p>
Training data and models for microphysics emulation
<p>Training data and models for microphysics emulation</p> <p>The training data is a subset of the full dataset described in the NeurIPS submission. Roughly speaking, 30 day runs with FV3GFS, with zhao carr microphysics. To keep the data reasonable in size, 1000 random netCDFs are sampled from the over 7000 files in the full training dataset. 200 test files are sampled.</p> <p>Also contains the trained ML models at models/</p> <p>Data behind the plots and tables is at plot-data/.</p> <p> </p>
How well does a convection-permitting climate model represent the reverse orographic effect of extreme hourly precipitation? - Observed precipitation data
<p>The dataset contains the rain gauge hourly rainfall series used in the paper "How well does a convection-permitting climate model represent the reverse orographic effect of extreme hourly precipitation?". Each rain gauge series is saved in one Matlab variable, organized as a structure S with five fields:</p> <p>S.name: the identification name of the rain gauge station</p> <p>S.vals_mm: series of hourly rainfall in millimeter</p> <p>S.time_utc: time steps series, in UTC time</p> <p>S.elev_m: elevation of the station, in m a.s.l.</p> <p>S.xy_utm: station coordinates X and Y in meter in the Reference system WGS84/UTM zone 32N</p>
Wage data for benchmarking CROMP model
<p>Data used for benchmarking the CROMP model (doi: https://doi.org/10.5281/zenodo.7152807).</p> <p>This is a synthetic wage data created from Data scientist salary dataset in Kaggle (https://www.kaggle.com/datasets/nikhilbhathi/data-scientist-salary-us-glassdoor). The detailed methodology behind creation of this dataset is explained in the unpublished paper associated with the CROMP model cited above, titled "Constrained Regression with Ordered and Margin-sensitive Parameters: Application in improving interpretability for wage models with prior knowledge" written by this author.</p> <p>The "hc_data" sheet is used for benchmarking the CROMP model.</p>
Data from: Nature versus nurture: Structural equation modeling indicates that parental care does not mitigate consequences of poor environmental conditions in Eastern bluebirds (Sialia sialis)
<p>1. How organisms respond to variation in environmental conditions and whether behavioral responses can mitigate negative consequences on growth, condition and other fitness measures are critical to our ability to conserve populations in changing environments. Offspring development is affected by environmental conditions and parental care behavior. When adverse environmental conditions are present, parents may alter behaviors to mitigate the impacts of poor environmental conditions on offspring.</p> <p>2. We determined if parental behavior (provisioning rates, attentiveness, nest temperature) varied in relation to environmental conditions (e.g., food availability, ectoparasites) and if parental behavior mitigated negative consequences of the environment on their offspring in Eastern bluebirds (<i>Sialia sialis</i>).</p> <p>3. We found that offspring on territories with lower food availability had higher hematocrit, and when bird blow flies (<i>Protocalliphora</i> spp.) were present growth rates were reduced. Parents increased provisioning and nest attendance in response to increased food availability but did not alter behavior in response to parasitism by blow flies. While parents altered behavior in response to resource availability, parents were unable to override the direct effects of negative environmental conditions on offspring growth and hematocrit.</p> <p>4. Our work highlights the importance of the environment on offspring development and suggests that parents may not be able to sufficiently alter behavior to ameliorate challenging environmental conditions.</p>
New Findings on Existing Resilient Modulus Constitutive Models through Performance Comparison on LTPP Data
<p>This dataset is the result of the study entitled, "New Findings on Existing Resilient Modulus Constitutive Models through Performance Comparison on LTPP Data".</p>
Supporting data for article comparison and Uncertainty Analysis of Species Distribution Models
<p>Downloaded from Web of Science for the supporting data of article comparison and Uncertainty Analysis of Species Distribution Models.</p>
Preprocessed Data for "Comparing Storm Resolving Models and Climates via Unsupervised Machine Learning"
<p>Preprocessed Data (training and test) for 3 SRMs used in "Comparing Storm Resolving Models and Climates via Unsupervised Machine Learning". Here we included ICON, SPCAM, and SPCAM with sea surface tem[eratures warmed by +4K. Additionally we include lat/lon information for the test data.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.