Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
90
datasets available to search
ShareScore release 0.9.0
Dataset results
90 results for “model transformation”
MAgPIE model - Land use change and carbon emissions of a transformation to timber cities
<p>MAgPIE source code for reproducibility</p> <p>Land use change and carbon emissions of a transformation to timber cities<br> (Nature Communications, 2022)</p> <p>DOI: 10.1038/s41467-022-32244-w</p> <p>Abhijeet Mishra1,2,*, Florian Humpenöder1, Galina Churkina1, Christopher P.O. Reyer1, Felicitas Beier1,2, Benjamin Leon Bodirsky1, Hans Joachim Schellnhuber1, Hermann Lotze-Campen1,2, and Alexander Popp1</p> <p>1 Potsdam Institute for Climate Impact Research (PIK), Member of Leibniz Association, P.O.Box 60 12 03, 14412,6<br> Potsdam, Germany<br> 2 Humboldt University of Berlin, Department of Agricultural Economics, Unter den Linden 6, 10099 Berlin,8<br> Germany</p> <p>Abhijeet Mishra<br> *mishra@pik-potsdam.de<br> May 2022</p> <p>See README.md for further details.</p>
Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"
<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States. </p> <p>The set of 47 news media outlets analysed (listed in Figure 1 of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets’ ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript. </p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets’ online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles’ HTML raw data using outlet-specific XPath expressions. </p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000’s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000. Figure S 1 in the SI reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small. Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth. </p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely: </p> <p>Siebert/sentiment-roberta-large-english from https://huggingface.co/siebert/sentiment-roberta-large-english. This model is a fine-tuned checkpoint of <a href="https://huggingface.co/roberta-large">RoBERTa-large</a> (<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with sentiment-roberta-large-english </p> <p>DistilRoberta j-hartmann/emotion-english-distilroberta-base from https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of <a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with Ekman's 6 basic emotions, plus a neutral class. The model was trained on 6 diverse datasets. Please refer to the original author at https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the siebert/sentiment-roberta-large-english Transformer model. https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset. https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the j-hartmann/emotion-english-distilroberta-base Transformer model. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>
PubChem and ChEMBL-series processed dataset used in Exhaustive local chemical space exploration using a transformer model
<p>PubChem and ChEMBL-series processed dataset used in <span>Exhaustive local chemical space exploration using </span><span>a transformer model</span></p>
Data package for paper "Transformer models for astrophysical time series and the GRB prompt-afterglow relation"
<p>This is a data package accompanying the paper "Transformer models for astrophysical time series and the GRB<br>prompt-afterglow relation". The code used to acquire the data is in the "data" folder. The code used to analyse the data is in the "analysis" folder.</p> <p>DOI paper: <a href="https://doi.org/10.1093/rasti/rzae026">10.1093/rasti/rzae026</a></p>
Data Set for Predicting the Performance of ATL Model Transformations
<p>Model transformation languages are special-purpose languages, which are designed to define transformations as comfortably as possible, i.e., often in a declarative way. With the increasing use of transformations in various domains, the complexity and size of input models are also increasing. However, developers often lack suitable models for performance testing. We have therefore conducted experiments in which we predict the performance of model transformations based on characteristics of input models using machine learning approaches. This dataset contains our raw and processed input data, the scripts necessary to repeat our experiments, and the results we obtained.</p> <p>Our input data consists of the time measurements for six different transformations defined in the Atlas Transformation Language (ATL), as well as the collected characteristics of the real-world input models that were transformed. We provide the script that implements our experiments. We predict the execution time of ATL transformations using the machine learning approaches linear regression, random forests and support vector regression using a radial basis function kernel. We also investigate different sets of characteristics of input models as input for the machine learning approaches. These are described in detail in the provided documentation.pdf. The results of the experiments are provided as raw data in individual cvs files. Additionally, we calculated the mean absolute percentage error in % and the 95th percentile of the absolute percentage error in % for each experiment and provide these results. Furthermore, we provide our Eclipse plugin, which collects the characteristics for a set of given models, the Java projects used to measure the execution time of the transformations, and other supporting scripts, e.g. for the analysis of the results.</p> <p>A short introduction with a quick start guide can be found in README.md and a detailed documentation in documentaion.pdf.</p>
A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary Reversing
<p>This is a dataset and pre-trained model for the official implementation of <a href="https://github.com/agwaBom/AsmDepictor"><strong>AsmDepictor</strong></a>, "A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary Reversing", In the 18th ACM Asia Conference on Computer and Communications Security <a href="https://asiaccs2023.org/">AsiaCCS '2023</a></p> <p> </p>
Data Set for Enhanced Performance Prediction of ATL Model Transformations
<p>Model transformation languages are domain-specific languages, which are designed to comfortably define transformations. With the increasing use of transformations in various domains, the complexity and size of input models are also increasing. However, developers often lack suitable models for performance testing. We have therefore conducted experiments in which we predict the performance of model transformations based on characteristics of input models using machine learning approaches. In particular, we focused on how to predict the performance of transformations that also transform attributes whose values can have arbitrary size. This dataset contains our raw and processed input data, the scripts necessary to repeat our experiments, and the results we obtained.</p> <p>Our input data consists of the time measurements for six different transformations defined in the Atlas Transformation Language (ATL), as well as the collected characteristics of the real-world input models we used. In this data set, we provide the script that implements our experiments. We predict the execution time of ATL transformations using the machine learning approaches linear regression, random forests and support vector regression using a radial basis function kernel. We also investigate different sets of characteristics of input models as input for the machine learning approaches. These are described in detail in the provided documentation.pdf. The results of the experiments are provided as raw data in individual cvs files. Furthermore, we provide our Eclipse plugin, which collects the characteristics for a set of given models.</p> <p>A detailed documentation is available in documentaion.pdf.</p>
Modeling Multi-level Dyadic Behavior to Transform the Science and Practice of Psychotherapy Process and Outcome.
ClinicalTrials.gov study NCT03594773. IPD Sharing: NO. Countries: 1. Publications: 2.
Numerical model code, input files and output data for publication ``Mixing and Transformation in a Deep Western Boundary Current: a case study''
<p>Contains numerical model data (code, input files, selected output) to supplement publication ``Mixing and Transformation in a Deep Western Boundary current'', by Spingys and co-authors. All umerical model data, including any errors, is the responsibility of Sonya Legg. This data set will allow reproduction of simulations used in the above-referenced paper.</p>
SAMPLER representations of FFPE TCGA-lung WSIs using tile-level features of the MMIL-Transformer model
<p>Here we provide single-scale SAMPLER representations of the FFPE TCGA-lung (LUAD and LUSC) WSIs using tile-level features provided in https://github.com/hustvl/MMIL-Transformer. To learn more about SAMPLER please visit https://github.com/TheJacksonLaboratory/SAMPLER.</p><p>The SAMPLER representations are provided as a single python pickle file. This pickle file contains a dictionary where each key is a WSI ID and each entry is the SAMPLER representation of the WSI.</p>
Model code and output for "Transformation Processes in the Oder Lagoon as seen from a Model Perspective"
<p>This dataset contains the numerical model output files that were used in the analysis, the numerical model input files that were used to produce the model simulation, and the specific source code for the model components presented in "Transformation Processes in the Oder Lagoon as seen from a Model Perspective", submitted to Biogeoscience.</p> <p>data1.tar is the direct diagnostic model output. data2.tar are condensed data which are used for the analysis.</p> <p>forcing.tar contains data to run the model. Most files should be placed inside a directory named INPUT/ that resides in the working directory where the model is being run. The following files should be at the top level in the working directory: data_table, diag_table, field_table, input.nml.</p> <p>For large files, a subset in time (1995) is provided for the first year of the simulation. This dataset does not include the coastdat atmospheric data, which can be downloaded directly from https://www.wdc-climate.de/ui/entry?acronym=coastDat-2_COSMO-CLM.</p> <p>src.tar contains the model source code.</p>
Explaining Non-Entailment by Model Transformation for the Description Logic EL - IJCKG22 - Resources
<p><strong>resources.zip</strong> contains the experiment data and results, and a README file explaining how to rerun the experiment.</p>
Incremental Model Transformations with Triple Graph Grammars for Multi-version Models and Multi-version Pattern Matching Evaluation Data
<p>Java abstract syntax graphs for two software development projects in non-recreating multi-version model encoding.</p>
A Unified Theoretical Framework for the Synergistic Integration of Transformers and Diffusion Models
<p>This paper introduces a novel, comprehensive theoretical framework for the synergistic integration of Transformer and Diffusion models, two paradigms that have independently revolutionized machine learning. We establish a fundamental correspondence between these models through a unified representation and a generalized dynamics equation, bridging the gap between their seemingly disparate architectures. Our key contributions include:</p> <p>(1) A unified mathematical formulation that encapsulates both Transformer and Diffusion processes;</p> <p>(2) A novel Diffusion-Enhanced Attention mechanism that incorporates Diffusion dynamics into Transformer attention;</p> <p>(3) Rigorous theoretical analyses including convergence guarantees, generalization bounds, and sample efficiency proofs for the integrated model.</p> <p>We provide detailed mathematical derivations and empirical validations across various tasks, demonstrating significant improvements over standalone models and existing hybrid approaches. This work lays the foundation for a new class of AI models that leverage the strengths of both paradigms, potentially leading to more powerful, efficient, and versatile AI systems. Our framework opens up new avenues for research in areas such as enhanced language modeling, advanced image generation, and multi-modal learning, paving the way for the next generation of AI technologies.</p>
Supplementary Material: Effectiveness of Performance Visualizations for Declarative Model Transformations
<p>We performed a case study to evaluate whether our performance visualizations developed for the declarative transformation language Henshin are suitable for performing a root cause analysis. Our study consisted of four parts. 1) Participants completed a questionnaire that collected data on their demographics and knowledge of model transformations. 2) The participants watched a video explaining the basics of models, Henshin, and our performance visualizations. 3) The study participants solved four different tasks one after the other. Guided by a questionnaire, they carried out a root cause analysis. 4) Finally, in a short interview session, we asked the participants about their assessment of the comprehensibility and usefulness of the visualizations.</p> <p>In total, 18 participants took part in our study. Our results show that most participants could correctly read and interpret the information provided by the visualizations. The majority of our participants could propose a performance optimization based on the visualizations that optimized the execution of a transformation.</p> <p>This data set contains our study material, which is necessary to repeat the study, our raw and processed data.</p>
zMAP toolset: model-based analysis of large-scale proteomic data via a variance stabilizing z-transformation
<p>Data and code used to generate the analyses and figures in paper "zMAP toolset: model-based analysis of large-scale proteomic data via a variance stabilizing z-transformation" are provided here.</p>
Dataset accompanying "Integrated nowcasting of convective precipitation with Transformer-based models using multi-source data"
<p>Dataset accompanying the article <a href="https://arxiv.org/abs/2409.10367" target="_blank" rel="noopener">Integrated nowcasting of convective precipitation with Transformer-based models using multi-source data</a>. </p> <p>Contains almost 8000 events that are sampled from the summer months (where convective precipitation events are most likely to occur) of 2019-2023, centred over Austria.</p> <p>Each sample has a temporal span of 4 hours with a spatial extent of 400 x 700 km, with following data streams:</p> <ul> <li>4 MSG infrared channels with central wavelengths of 6.2, 7.3, 8.7, and 10.8 μm</li> <li>Rain rates mosaicked from ground-based radar observations</li> <li>Lightning data from ground-based observations</li> <li>INCA precipitation analysis</li> <li>INCA convective available potential energy (CAPE) estimates</li> </ul> <p>The dataset is accompanied by elevation and coordinate information. </p> <p>Please refer to the manuscript and the <a href="https://github.com/caglarkucuk/earthformer-multisource-to-inca">GitHub repository</a> for further information and helper code for reading the data files.</p>
Dataset for downscaling, used in downscaled Spatiotemporal Precipitation Model Based on a Transformer Attention Mechanism
<p>In this research, we introduce a novel method leveraging the Transformer architecture to generate high-fidelity precipitation model outputs. This technique emulates the statistical characteristics of high-resolution datasets while substantially lowering computational expenses. The core concept involves utilizing a blend of coarse and fine-grained simulated precipitation data, encompassing diverse spatial resolutions and geospatial distributions, to instruct the neural network in the transformation process. We have crafted an innovative ST-Transformer encoder component that dynamically concentrates on various regions, allocating heightened focus to critical spatial zones or sectors. This tailored module is instrumental in enhancing the model's ability to generate outcomes that are not only more true-to-life but also more consistent with physical laws. It adeptly mirrors the temporal and spatial fluctuations in precipitation data and adeptly represents extreme weather events, such as heavy and enduring storms. The efficacy and superiority of our proposed approach are substantiated through a comparative analysis with several cutting-edge forecasting techniques. This evaluation is conducted on two distinct datasets, each derived from simulations run by regional climate models over a period of four months. The datasets vary in their spatial resolutions, with one featuring a 50-kilometer resolution and the other a 12-kilometer resolution, both sourced from the Weather Research and Forecasting (WRF) Model.</p>
Data from: Swapping birth and death: symmetries and transformations in phylodynamic models
Stochastic birth--death models provide the foundation for studying and simulating evolutionary trees in phylodynamics. A curious feature of such models is that they exhibit fundamental symmetries when the birth and death rates are interchanged. In this paper, we first provide intuitive reasons for these known transformational symmetries. We then show that these transformational symmetries (encoded in algebraic identities) are preserved even when individuals at the present are sampled with some probability. However, these extended symmetries require the death rate parameter to sometimes take a negative value. In the last part of this paper, we describe the relevance of these transformations and their application to computational phylodynamics, particularly to maximum likelihood and Bayesian inference methods, as well as to model selection.
Multiscale modelling of flow, heat transfer and transformation during thermal treatment of starch suspensions
<p>Multiscale modelling of flow, heat transfer and transformation during thermal treatment of starch suspensions</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.