Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

141

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

141 results for “Sentiment”

Learn how ShareScore rates datasets ↗
zenodo36/100

Lexicon and example extensions from paper "Lexicon-based comments-oriented news sentiment analyzer system"

<p>Lexicon and example extensions for the article &quot;Moreo, Alejandro, et al. &quot;Lexicon-based comments-oriented news sentiment analyzer system.&quot;&nbsp;<em>Expert Systems with Applications</em>&nbsp;39.10 (2012): 9166-9180.&quot;&nbsp; Founded by Ministerio de Educaci&oacute;n y Ciencia and Junta de Andaluc&iacute;a with&nbsp;Projects: TIN2007-60199, TIC2009-5011 and TIN2007-67984</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Bag Brands Sentiment Dataset

<p>The Bag Brand Sentiment Dataset is a collection of tweet data from Twitter that focuses on several popular designer bag brands. The dataset includes tweets related to seven specific keywords: &quot;gucci bag&quot;, &quot;chanel bag&quot;, &quot;dior bag&quot;, &quot;louis vuitton bag&quot;, &quot;prada bag&quot;, &quot;hermes bag&quot;, and &quot;supreme bag&quot;.</p> <p>The data was obtained using the Twitter API, which is a tool used to extract data from Twitter. The dataset consists of a total of 2881 tweets that were obtained through Twitter crawling. Before the dataset was compiled, a pre-processing process was conducted to remove duplicate data, ensuring that the dataset contains only unique tweets.</p> <p>The Bag Brand Sentiment Dataset is useful for analyzing consumer sentiment towards popular designer bag brands. It can be used by marketers to gain insights into consumer preferences and attitudes towards specific brands. Additionally, researchers can use the dataset to study trends in consumer sentiment towards luxury goods or to explore how social media platforms are used to discuss designer bag brands.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Sentiment analysis data and word embeddings for Erzya, Komi-Zyrian, Moksha and Udmurt

<p>The aligned sentiment annotated data is in setiment_eval_data.json, vectors.zip has the word embeddings in a textual Gensim format, code.zip has the code and models.zip the sentiment analysis model.</p> <p>Please cite the following paper:</p> <p><strong>Alnajjar, K., H&auml;m&auml;l&auml;inen, M., &amp; Rueter, J, (2023)&nbsp;Sentiment Analysis Using Aligned Word Embeddings for Uralic Languages. In <em>Proceedings of the Second Workshop on Resources and Representations for Under-resourced Languages and Domains (RESOURCEFUL-2023)</em></strong></p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Forex News Annotated Dataset for Sentiment Analysis

<p>This dataset contains&nbsp;news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY. The data was extracted from reputable platforms <a href="https://www.forexlive.com">Forex Live</a>&nbsp;and <a href="https://www.fxstreet.com/">FXstreet</a>&nbsp;over a period of 86 days, from January to May 2023. The&nbsp;dataset comprises 2,291 unique news headlines. Each headline includes an associated forex pair, timestamp, source, author, URL, and the corresponding article text. Data was collected using web scraping techniques executed via a custom service on a virtual machine. This service periodically retrieves the latest news for a specified forex pair (ticker) from each platform, parsing all available information. The collected data is then processed to extract details such as the article&#39;s timestamp, author, and URL. The URL is further used to retrieve the full text of each article. This data acquisition process repeats approximately every 15 minutes.</p> <p>To ensure the reliability of the dataset, we manually annotated each headline for sentiment. Instead of solely focusing on the textual content, we <strong>ascertained sentiment based on the potential short-term impact of the headline on its corresponding forex pair</strong>. This method recognizes the currency market&#39;s acute sensitivity to economic news, which significantly influences many trading strategies. As such, this dataset could serve as an invaluable resource for fine-tuning sentiment analysis models in the financial realm.</p> <p>We used three categories for annotation: &#39;positive&#39;, &#39;negative&#39;, and &#39;neutral&#39;, which correspond to bullish, bearish, and hold sentiments, respectively, for the forex pair linked to each headline. The following&nbsp;Table&nbsp;provides examples of annotated headlines along with brief explanations of the assigned sentiment.&nbsp;</p> Examples of Annotated Headlines Forex Pair Headline Sentiment Explanation GBPUSD&nbsp; Diminishing bets for a move to 12400&nbsp; Neutral Lack of strong sentiment in either direction GBPUSD&nbsp; No reasons to dislike Cable in the very near term as long as the Dollar momentum remains soft&nbsp;&nbsp; Positive Positive sentiment towards GBPUSD (Cable) in the near term GBPUSD&nbsp; When are the UK jobs and how could they affect GBPUSD &nbsp; Neutral Poses a question and does not express a clear sentiment JPYUSD Appropriate to continue monetary easing to achieve 2% inflation target with wage growth&nbsp;&nbsp; Positive Monetary easing from Bank of Japan (BoJ) could lead to a weaker JPY in the short term due to increased money supply USDJPY Dollar rebounds despite US data. Yen gains amid lower yields &nbsp; Neutral Since both the USD and JPY are gaining, the effects on the USDJPY forex pair might offset each other USDJPY USDJPY to reach 124 by Q4 as the likelihood of a BoJ policy shift should accelerate Yen gains &nbsp; Negative USDJPY is expected to reach a lower value, with the USD losing value against the JPY AUDUSD <p>RBA Governor Lowe&rsquo;s Testimony High inflation is damaging and corrosive &nbsp;</p> Positive Reserve Bank of Australia (RBA) expresses concerns about inflation. Typically, central banks combat high inflation with higher interest rates, which could strengthen AUD. <p>Moreover, the dataset includes two columns with the predicted sentiment class and score as predicted by the <a href="https://huggingface.co/ProsusAI/finbert">FinBERT</a> model. Specifically, the FinBERT model outputs a set of probabilities for each sentiment class (positive, negative, and neutral), representing the model&#39;s confidence in associating the input headline with each sentiment category. These probabilities are used to determine the predicted class and a sentiment score&nbsp;for each headline. The sentiment score is computed by subtracting the negative class probability from the positive one.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Is the sentiment priced? Evidence from the Korean stock market

<p>SAS files.</p> <p>1) dtset_ks.KOSPI&nbsp;</p> <p>2) dtset_kq: KOSDAQ</p> <p>Data description: This datasets are&nbsp;the common stock of non-financial companies traded on KOSPI(dtset_ks)&nbsp;and KOSDAQ(dtset_kq), including delisted stocks. KOSPI comprises primarily of large, established firms, whereas KOSDAQ comprises young, entrepreneurial firms.&nbsp;The spans from February 2000 to June 2022 and is based on the stock excess return.&nbsp;This dataset utilizes data from FnGuide and the yields of CD (91-day) provided by the Bank of Korea as a risk-free rate.&nbsp;</p> <p>Please look at the paper&nbsp;to&nbsp;see specific variables.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

Dataset and scripts for "Sentiment Analysis over Collaborative Relationships in Open Source Software Projects"

<p>Dataset and scripts for &quot;Sentiment Analysis over Collaborative Relationships in Open Source Software Projects&quot;.</p> <p>README is included in the files</p>

opencc-by-4.0Jul 2019View details →
dryad32/100

Data from: Understanding sentiment of national park visitors from social media data

<p>National parks are key for conserving biodiversity and supporting people´s well-being. However, anthropogenic pressures challenge the existence of national parks and their conservation effectiveness. Therefore, it is crucial to assess how people perceive national parks in order to enhance socio-political support for conservation. User-generated data shared by visitors on social media provide opportunities to understand how people perceive (e.g. preferences, feelings, opinions) national parks during nature-based recreational experiences. In this study, we applied methods from automated natural language processing to assess visitors' sentiment when describing experiences in Instagram posts geolocated inside four national parks in South Africa. We found that visitors' sentiment was positive, and mostly included emotions such as joy, anticipation, trust and surprise, with only a small occurrence of posts with negative feelings. Appreciation of nature, in association with a diverse set of other aspects, such as activities, geographical features and tourist attractions, was used to describe experiences related to nature, wilderness, traveling, holidays and adventures. The type of nature-based experience described by visitors was park specific, revealing different profiles of parks providing wildlife or scenery experiences. Findings support and highlight the societal role of national parks in providing visitors with opportunities to develop positive connections with nature. Social media data may be used to understand visitors' perceptions, and how the image of national parks is constructed by users in the virtual social environment. This may help inform management for promoting a high quality tourism experience, as well as conservation marketing aimed at fostering socio-political support for national parks and their long-term conservation effectiveness.</p>

opencc-zeroJul 2020View details →
zenodo32/100

SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets

<p>There are 2 sub-tasks: sentiment analysis for Spanglish (Spanish-English) and for Hinglish (Hindi-English).</p> <p>The sentiment classes are Positive, negative, neutral.&nbsp;</p> <p>Hinglish dataset has 20k instances.</p> <p>Spanglish dataset has ~19k instances.&nbsp;</p> <p>Website:&nbsp;<a href="https://ritual-uh.github.io/sentimix2020/">https://ritual-uh.github.io/sentimix2020/</a></p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Aspect-based Sentiment Analysis of Scientific Reviews - Openreview dataset

<p>The dataset contains all the data used in the JCDL 2020 research paper: <a href="https://dl.acm.org/doi/10.1145/3383583.3398541">Aspect-based Sentiment Analysis of Scientific Reviews</a></p> <p>The dataset is split into multiple files containing&nbsp;all the sentence annotations and the ICLR open review dataset (with reviews and scores and the confidence scores, final recommendation, etc.) for the last three years.</p> <p>The file &quot;iclr_conf.p&quot; is a pickle file which contains a NumPy array object.<br> The array contains 2681 rows corresponding to each accepted or rejected paper of 2017,2018,2019<br> Each row contains 4 columns.<br> The first column is the link of the paper in openreview.net, from where the data related to the paper is collected.<br> The second column is either 0 or 1, corresponding to the final decision: rejection or acceptance respectively.<br> The third column is the year of the conference for the particular submission.<br> The fourth column is another NumPy array containing 3 reviews in 3 rows. Each row of this array contains 3 columns containing the list of sentences in the same sequence as it appears in the text of the review, the confidence(ranging from 1-5), and the rating(ranging(1-10)) respectively.</p> <p>Each line of the file &quot;sentences.csv&quot; contains one sentence whose corresponding annotation is provided in the corresponding line in the file &quot;annotations.csv&quot;<br> The file &quot;annotations.csv&quot; is a file containing 8 comma-separated integers in each line.<br> Each column corresponds to the following aspects: Appropriateness, Clarity, Originality, Empirical/Theoretical Soundness, Meaningful Comparison, Substance,<br> Impact of Dataset/Software/Ideas and Recommendation.<br> An integer 0,1,2,3 corresponds to the following sentiment labels of the sentence on that aspect: Absent, Positive, Negative, Neutral</p> <p>Please cite our paper published in JCDL-2020 if you use our data: <a href="https://dl.acm.org/doi/10.1145/3383583.3398541">https://dl.acm.org/doi/10.1145/3383583.3398541</a></p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

PolSentiLex: Sentiment Detection in Socio-Political Discussions on Russian Social Media

<p>A Russian-language sentiment lexicon for social media discussions on political and social issues.</p> <p>The file contains raw markings collected with LINIS coding service <a href="https://linis-crowd.org">https://linis-crowd.org</a> [in Russian].</p> <p>Learn more about PolSentiLex in our papers:</p> <ul> <li> <p>Koltsova, O., &amp; Alexeeva, S. (2015). Linis-crowd.org: A lexical resource for Russian sentiment analysis of social media [Linis-crowd.org: Lexichesk resurs dl&rsquo;a analiza tonal&rsquo;nosti sotsial&rsquo;no-politicheskix tekstov]. Computational Linguis- Tics and Computantional Ontologies: Proceedings of the XVIII Joint Conference &ldquo;Internet and Modern Society (IMS-2015)&rdquo; [Kompyuternaya Lingvistika i Vyichis- Litelnyie Ontologii: Sbornik Nauchnyih Statey. Trudyi XVIII Ob&rsquo;edinennoy Konferen- Tsii &laquo;Internet i Sovremennoe Obschestvo&raquo; (IMS-2015)], 25&ndash;34. [in Russian] URL: <a href="https://scila.hse.ru/data/2020/06/02/1603986481/koltsovaoyuetal.pdf">https://scila.hse.ru/data/2020/06/02/1603986481/koltsovaoyuetal.pdf</a></p> </li> <li> <p>Koltsova, O., Alexeeva, S., &amp; Koltsov, S. (2016). An Opinion Word Lexicon and a Training Dataset for Russian Sentiment Analysis of Social Media. Computational Linguistics and Intellectual Technologies: Proceedings of the International Conference &ldquo;Dialogue 2016&rdquo;, 277&ndash;287. URL: <a href="http://www.dialog-21.ru/media/3400/koltsovaoyuetal.pdf">http://www.dialog-21.ru/media/3400/koltsovaoyuetal.pdf</a></p> </li> <li> <p>Koltsova O., Alexeeva S., Pashakhin S., Koltsov S. (2020) PolSentiLex: Sentiment Detection in Socio-Political Discussions on Russian Social Media. In: Filchenkov A., Kauttonen J., Pivovarova L. (eds) Artificial Intelligence and Natural Language. AINL 2020. Communications in Computer and Information Science, vol 1292. Springer, Cham. <a href="https://doi.org/10.1007/978-3-030-59082-6_1">https://doi.org/10.1007/978-3-030-59082-6_1</a></p> </li> </ul>

opencc-by-nc-4.0Oct 2020View details →
zenodo32/100

Emoji Sentiment Lexicons

<p>These are the emoji sentiment lexica derived from valence scores from cooccurrence with sentiment-carrying messages. One lexicon is based on a Twitter corpus and contains Unicode emojis, the other is based on a collection of Twitch chat logs and mainly contains valence values for Twitch emotes.</p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

Datasets of the article "From Classification to Quantification in Tweet Sentiment Analysis"

<p>Datasets used for the following SNAM paper:<br> ---------------------------------------------------------------------------------------------------<br> Title: From Classification to Quantification in Tweet Sentiment Analysis<br> Authors: Wei Gao and Fabrizio Sebastiani<br> Organization: Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar<br> ---------------------------------------------------------------------------------------------------</p> <p>[Content]</p> <p>* SemEval2013, SemEval2014, SemEval2015 datasets:<br> &nbsp; - semeval.train.feature.txt: Training set for learning sentiment models at development stage<br> &nbsp; - semeval.dev.feature.txt: Held-out set for tuning parameters<br> &nbsp; - semeval.train+dev.feature.txt: Training set for learning the final sentiment model<br> &nbsp; - semeval13.test.feature.txt: SemEval2013 test set<br> &nbsp; - semeval14.test.feature.txt: SemEval2014 test set<br> &nbsp; - semeval15.test.feature.txt: SemEval2015 test set<br> &nbsp;&nbsp;<br> * Other datasets: semeval2016, sanders, sst, omd, hcr, gasp, wa, wb<br> &nbsp; - X.train.feature.txt: Training set for learning sentiment models at development stage<br> &nbsp; - X.dev.feature.txt: Held-out set for tuning parameters<br> &nbsp; - X.train+dev.feature.txt: Training set for learning the final sentiment model<br> &nbsp; - X.test.feature.txt (or X.dev-test.feature.txt for semeval2016 only): Test set<br> where X is one of semeval2016, sanders, sst, omd, hcr and gasp.</p> <p>* Training files are saved in ./data/train directory, and held-out and test files are in ./data/test directory</p> <p><br> For more details, please refer to the paper.</p> <p><br> [Citation]<br> You can cite the following paper when referring to the dataset:</p> <pre>@article{gao2016classification, title={From classification to quantification in tweet sentiment analysis}, author={Gao, Wei and Sebastiani, Fabrizio}, journal={Social Network Analysis and Mining}, volume={6}, number={1}, pages={19}, year={2016}, publisher={Springer} }</pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2015View details →
zenodo32/100

URDU Dataset for Multi-modal Sentiment Analysis

<p>The "Multi-modal Sentiment Analysis Dataset for Urdu Language Opinion Videos" is a valuable resource aimed at advancing research in sentiment analysis, natural language processing, and multimedia content understanding. This dataset is specifically curated to cater to the unique context of Urdu language opinion videos, a dynamic and influential content category in the digital landscape.</p> <p><strong>Dataset Description:</strong></p> <ul> <li><strong>Size and Diversity:</strong>&nbsp;This dataset comprises an extensive collection of Urdu language opinion videos, encompassing a wide spectrum of topics and sentiments. It consists of a total of 214 videos, each of varying lengths, offering a diverse and comprehensive representation of the Urdu language content landscape.</li> <li><strong>Sentiment Annotations:</strong>&nbsp;The dataset is meticulously annotated with sentiment labels, providing information on the emotional tone expressed in each video. The sentiment labels include "positive," "negative," and "neutral," offering a nuanced understanding of the sentiment conveyed in these multimedia opinion pieces.</li> <li><strong>Multi-modal Approach:</strong>&nbsp;A unique feature of this dataset is its multi-modal approach. It combines text, audio, and visual data to enable researchers to delve into the various dimensions of sentiment analysis within the context of opinion videos. The multi-modal annotations encompass the textual content of spoken words, the auditory characteristics of the videos, and the visual cues from the video frames.</li> </ul> <p><strong>Significance and Applications:</strong></p> <p>This dataset holds significant value for both the research community and practical applications:</p> <ul> <li><strong>Research Advancement:</strong>&nbsp;Researchers can employ this dataset to investigate the complex landscape of sentiment analysis within the context of opinion videos. It facilitates inquiries into sentiment trends, the development of sentiment analysis models, and the creation of sentiment-aware multimedia content analysis tools.</li> <li><strong>Content Recommendation:</strong> The dataset can play a pivotal role in the development of content recommendation systems that cater to viewers' emotional preferences. Understanding sentiment in opinion videos is crucial for improving content engagement and user experience.</li> <li><strong>User Engagement Analysis:</strong> The dataset can empower studies on user engagement and interaction with multimedia content. It is an essential resource for researchers aiming to decode the factors influencing viewer reactions and engagement in multimedia.</li> </ul> <p>&nbsp;</p> <p>Researchers are encouraged to explore and utilize this dataset for various academic and commercial purposes, fostering innovation in sentiment analysis and multimedia understanding. The dataset is made available with open access to facilitate collaborative research and to contribute to the broader knowledge in the field.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Does Investor Sentiment Predict Bitcoin Return and Volatility? - A Quantile Regression Approach

<p>This dataset was used in generating findings for the paper titled &quot;<strong>Does Investor Sentiment Predict Bitcoin Return and Volatility? - A Quantile Regression Approach&quot;.</strong></p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

One million articles from five post socialist countries with extracted features: sentiment, basic emotions, LDA topics and presence of influential domestic politicians

<p>This is a replication data for my paper under blind review.<br> <br> This paper develops a new prediction model for media content presence on a website. It analyses a new corpus of one million articles from five countries: Poland, Russia, Belarus, Kazakhstan and Ukraine, in two languages, Polish and Russian. These articles were scraped daily from seventeen websites in 2017-2020 period. The research applies a wide range of natural language processing methods to automatically derive several properties of each article: its topic, sentiment, basic emotions, mentions of influential domestic politicians. The articles&rsquo; embeddings and their cosine similarity are used to calculate the news context, such as how an article differs from the daily issue main themes. These features are used to estimate a logistic regression assessing the likelihood that the same or slightly modified, as measured by cosine similarity, article will remain on the main web page the next day. The key, and somewhat unexpected result is that articles with negative sentiment polarity are less likely to be published for more than one day. This result holds for all countries analyzed. It means that the negative news bias documented in the literature is partly offset by their shorter life cycle.<br> <br> Data is in the Python pickle format. Should be read into Python using the pickle.load() function. Each element (row) is the data frames or list represents one news article. Each file has the same format. Loading a pickle file returns a list of four elements:<br> 1. A dummy variable equal to 1 when the article was published the next day, with the text being identical<br> 2. A dummy variable equal to 1 when the article was published the next day, but we allow for small text modifications (cosine similarity &gt; 0.99)<br> 3. Dataframe with extracted features, described below.<br> 4. List with texts of articles in Polish or Russian<br> <br> Ad 3. The columns of the dataframe are as follows (we refer to row number i in description):<br> - pandas index (may appear once or twice in the datafame)<br> - maxcosine: maximum cosine similarity between art i and all articles published next day<br> - cosine_diff: cosine similarity between article i and the elementwise average of embeddings of all articles in the current issue. Measure how similar is the article i to the core narrative of the current issue<br> - cosine_std: std. dev. of cosine similarity measures between all pairs of articles in the current issue. Measures how focused or dispersed is the current issue news coverage<br> - thirteen LDA topic groups: politics, legislation and legal affairs (POL); economy, finance, various sectors of the economy (ECO); military, war, protests, crime, security threats (MIL); international affairs, specific issues concerning foreign countries (INT); technology (TECH); family issues, culture, sport, education (FAM); regional issues and housing (REG); health issues and the Covid-19 pandemic (HEA); media (MED); accidents (ACC); religion (REL); the Soviet Union (USSR); and articles for which no topic could be determined (MISC).<br> - rsent.c: relative sentiment that is dictionary based sentiment of articles i minus the average sentiment of the newspaper. This approach eliminates newspaper or country idiosyncratic sentiment factors. c stands for Covid, the sentiment lexicon was augmented with Covid related terms<br> - dip_*: Variable measuring if influential domestic politicians are mentioned in article i, * represent a country acronym. If N is equal to the number of occurrences of the names of influential domestic politicians in the article i, dip_* = 0 if N=0, dip_* = 1+ log(N) if N&gt;0.<br> - three or four names of news portals from which the data was scraped.<br> - names of six basic emotions and the article i emotion scores calculated using zero-shot learning and the large version of the XLM (Conneau et al., 2019) model from the huggingface transformers library available at https://huggingface.co/vicgalle/xlm-roberta-large-xnli-anli<br> Names of the politicians used to calculate dip variables<br> Russia<br> &quot;putin&quot; &quot;medvedev&quot; &quot;vaino&quot; &quot;shoigu&quot; &quot;bortnikov&quot; &quot;lavrov&quot; &quot;mishustin&quot; &quot;kirienko&quot; &quot;sechin&quot;<br> Ukraine<br> &quot;zelensky&quot; &quot;shmygal&quot; &quot;akhmetov&quot; &quot;avakov&quot; &quot;ermak&quot; &quot;poroshenko&quot; &quot;medvedchuk&quot; &quot;groisman&quot;<br> Kazakhstan<br> &quot;sagyntaev&quot; &quot;mamin&quot; &quot;tokayev&quot; &quot;nnazarbayev&quot; &quot;dnazarbayeva&quot; &quot;kulibayev&quot; &quot;masimov&quot;<br> Belarus<br> &quot;alukashenko&quot; &quot;vakulchik&quot; &quot;vlukashenko&quot; &quot;kobyakov&quot; &quot;makei&quot; &quot;myasnikovich&quot;&nbsp; &quot;rumas&quot; &quot;golovchenko&quot;<br> Poland<br> &quot;kaczynski&quot; &quot;duda&quot; &quot;morawiecki&quot; &quot;ziobro&quot;<br> Data coverage<br> Country, news portal, numbr of articles<br> Russia iz.ru 43,782<br> Russia kommersant.ru 46,070<br> Russia novayagazeta.ru 29,357<br> Russia vedomosti.ru 27,797<br> Kazakhstan informburo.kz 29,375<br> Kazakhstan nur.kz 67,350<br> Kazakhstan tengrinews.kz 44,285<br> Kazakhstan zakon.kz 109,442<br> Belarus bdg.by 33,447<br> Belarus belgazeta.by 21,995<br> Belarus sb.by 83,685<br> Ukraine kp.ua 194,792<br> Ukraine segodnya.ua 45,835<br> Ukraine vesti.ua 90,559<br> Poland gazeta.pl 53,321<br> Poland rp.pl 49,587<br> Poland wpolityce.pl 76,625<br> <br> In the provided dataframes the number of observations is smaller, because the issues for which there was no next day issue, were removed.<br> <br> Data was scraped daily between 2017 or 2018 (depending on the country) and January 2021.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Replication package for "Investor Sentiment, Sovereign Debt Mispricing, and Economic Outcomes"

<p>Replication package for &quot;Investor Sentiment, Sovereign Debt Mispricing, and Economic Outcomes&quot; by Ramzy Al-Amine and Tim Willems</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Senti-Pol-sr Sentiment Lexicon

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo32/100

Relative Investor Sentiment (01-01-1996 - 31.12.2022, monthly and daily)

<p>The data set contains the daily and the monthly Relative Investor Sentiment for the time 1996 to 2022 which is described and used in:</p> <p>Gao, Xiang and Koedijk, Kees and Walther, Thomas and Wang, Zhan, Relative Investor Sentiment (May 24, 2022). Available at SSRN:&nbsp;<a href="https://ssrn.com/abstract=4122594" target="_blank" rel="noopener">https://ssrn.com/abstract=4122594</a>&nbsp;or&nbsp;<a href="https://dx.doi.org/10.2139/ssrn.4122594" target="_blank" rel="noopener">http://dx.doi.org/10.2139/ssrn.4122594</a></p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset for Sentiment Analysis of X platform about "MK Hasil Pemilu"

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis&rdquo;, Proceedings of the 7th International Conference on Machine Learning and Natural Language Processing (MLNLP 2024), Chengdu, China, October 18-20, 2024 (Paper accepted for publication, Preprint available at: https://arxiv.org/abs/2410.03293)</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>The outbreak of COVID-19 served as a catalyst for content creation and dissemination on social media platforms, as such platforms serve as virtual communities where people can connect and communicate with one another seamlessly. While there have been several works related to the mining and analysis of COVID-19-related posts on social media platforms such as Twitter (or X), YouTube, Facebook, and TikTok, there is still limited research that focuses on the public discourse on Instagram in this context. Furthermore, the prior works in this field have only focused on the development and analysis of datasets of Instagram posts published during the first few months of the outbreak. The work presented in this paper aims to address this research gap and presents a novel multilingual dataset of <strong>500,153 Instagram posts about COVID-19 published between January 2020 and September 2024</strong>. This dataset contains Instagram posts in <strong>161 different languages</strong>. After the development of this dataset, multilingual sentiment analysis was performed using VADER and twitter-xlm-roberta-base-sentiment. This process involved classifying each post as positive, negative, or neutral. The results of sentiment analysis are presented as a separate attribute in this dataset.</p> <p><em><strong>For each of these posts, the Post ID, Post Description, Date of publication, language code, full version of the language, and sentiment label are presented as separate attributes in the dataset.</strong></em></p> <p>The Instagram posts in this dataset are present in <strong>161 different languages</strong> out of which the top 10 languages in terms of frequency are English (343041 posts), Spanish (30220 posts), Hindi (15832 posts), Portuguese (15779 posts), Indonesian (11491 posts), Tamil (9592 posts), Arabic (9416 posts), German (7822 posts), Italian (5162 posts), Turkish (4632 posts)</p> <p>There are <strong>535,021 distinct hashtags in this dataset</strong> with the top 10 hashtags in terms of frequency being #covid19 (169865 posts), #covid (132485 posts), #coronavirus (117518 posts), #covid_19 (104069 posts), #covidtesting (95095 posts), #coronavirusupdates (75439 posts), #corona (39416 posts), #healthcare (38975 posts), #staysafe (36740 posts), #coronavirusoutbreak (34567 posts)</p> <p>The following is a description of the attributes present in this dataset</p> <ul> <li><em><strong>Post ID</strong></em>:&nbsp;Unique ID of each Instagram post</li> <li><em><strong>Post Description</strong></em>:&nbsp;Complete description of each post in the language in which it was originally published</li> <li><em><strong>Date</strong></em>: Date of publication in MM/DD/YYYY format</li> <li><em><strong>Language code</strong></em>: Language code (for example: &ldquo;en&rdquo;) that represents the language of the post as detected using the Google Translate API&nbsp;</li> <li><em><strong>Full Language</strong></em>: Full form of the language (for example: &ldquo;English&rdquo;) that represents the language of the post as detected using the Google Translate API&nbsp;</li> <li><em><strong>Sentiment</strong></em>: Results of sentiment analysis (using the preprocessed version of each post) where each post was classified as positive, negative, or neutral</li> </ul> <p><strong>Open Research Questions</strong></p> <p>This dataset is expected to be helpful for the investigation of the following research questions and even beyond:</p> <ol> <li>How does sentiment toward COVID-19 vary across different languages?</li> <li>How has public sentiment toward COVID-19 evolved from 2020 to the present?</li> <li>How do cultural differences affect social media discourse about COVID-19 across various languages?</li> <li>How has COVID-19 impacted mental health, as reflected in social media posts across different languages?</li> <li>How effective were public health campaigns in shifting public sentiment in different languages?</li> <li>What patterns of vaccine hesitancy or support are present in different languages?</li> <li>How did geopolitical events influence public sentiment about COVID-19 in multilingual social media discourse?</li> <li>What role does social media discourse play in shaping public behavior toward COVID-19 in different linguistic communities?</li> <li>How does the sentiment of minority or underrepresented languages compare to that of major world languages regarding COVID-19?</li> <li>What insights can be gained by comparing the sentiment of COVID-19 posts in widely spoken languages (e.g., English, Spanish) to those in less common languages?</li> </ol> <p>All the Instagram posts that were collected during this data mining process to develop this dataset were publicly available on Instagram and did not require a user to log in to Instagram to view the same (at the time of writing this paper).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record