Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1 result for “basic emotions”

Learn how ShareScore rates datasets ↗
zenodo32/100

One million articles from five post socialist countries with extracted features: sentiment, basic emotions, LDA topics and presence of influential domestic politicians

<p>This is a replication data for my paper under blind review.<br> <br> This paper develops a new prediction model for media content presence on a website. It analyses a new corpus of one million articles from five countries: Poland, Russia, Belarus, Kazakhstan and Ukraine, in two languages, Polish and Russian. These articles were scraped daily from seventeen websites in 2017-2020 period. The research applies a wide range of natural language processing methods to automatically derive several properties of each article: its topic, sentiment, basic emotions, mentions of influential domestic politicians. The articles&rsquo; embeddings and their cosine similarity are used to calculate the news context, such as how an article differs from the daily issue main themes. These features are used to estimate a logistic regression assessing the likelihood that the same or slightly modified, as measured by cosine similarity, article will remain on the main web page the next day. The key, and somewhat unexpected result is that articles with negative sentiment polarity are less likely to be published for more than one day. This result holds for all countries analyzed. It means that the negative news bias documented in the literature is partly offset by their shorter life cycle.<br> <br> Data is in the Python pickle format. Should be read into Python using the pickle.load() function. Each element (row) is the data frames or list represents one news article. Each file has the same format. Loading a pickle file returns a list of four elements:<br> 1. A dummy variable equal to 1 when the article was published the next day, with the text being identical<br> 2. A dummy variable equal to 1 when the article was published the next day, but we allow for small text modifications (cosine similarity &gt; 0.99)<br> 3. Dataframe with extracted features, described below.<br> 4. List with texts of articles in Polish or Russian<br> <br> Ad 3. The columns of the dataframe are as follows (we refer to row number i in description):<br> - pandas index (may appear once or twice in the datafame)<br> - maxcosine: maximum cosine similarity between art i and all articles published next day<br> - cosine_diff: cosine similarity between article i and the elementwise average of embeddings of all articles in the current issue. Measure how similar is the article i to the core narrative of the current issue<br> - cosine_std: std. dev. of cosine similarity measures between all pairs of articles in the current issue. Measures how focused or dispersed is the current issue news coverage<br> - thirteen LDA topic groups: politics, legislation and legal affairs (POL); economy, finance, various sectors of the economy (ECO); military, war, protests, crime, security threats (MIL); international affairs, specific issues concerning foreign countries (INT); technology (TECH); family issues, culture, sport, education (FAM); regional issues and housing (REG); health issues and the Covid-19 pandemic (HEA); media (MED); accidents (ACC); religion (REL); the Soviet Union (USSR); and articles for which no topic could be determined (MISC).<br> - rsent.c: relative sentiment that is dictionary based sentiment of articles i minus the average sentiment of the newspaper. This approach eliminates newspaper or country idiosyncratic sentiment factors. c stands for Covid, the sentiment lexicon was augmented with Covid related terms<br> - dip_*: Variable measuring if influential domestic politicians are mentioned in article i, * represent a country acronym. If N is equal to the number of occurrences of the names of influential domestic politicians in the article i, dip_* = 0 if N=0, dip_* = 1+ log(N) if N&gt;0.<br> - three or four names of news portals from which the data was scraped.<br> - names of six basic emotions and the article i emotion scores calculated using zero-shot learning and the large version of the XLM (Conneau et al., 2019) model from the huggingface transformers library available at https://huggingface.co/vicgalle/xlm-roberta-large-xnli-anli<br> Names of the politicians used to calculate dip variables<br> Russia<br> &quot;putin&quot; &quot;medvedev&quot; &quot;vaino&quot; &quot;shoigu&quot; &quot;bortnikov&quot; &quot;lavrov&quot; &quot;mishustin&quot; &quot;kirienko&quot; &quot;sechin&quot;<br> Ukraine<br> &quot;zelensky&quot; &quot;shmygal&quot; &quot;akhmetov&quot; &quot;avakov&quot; &quot;ermak&quot; &quot;poroshenko&quot; &quot;medvedchuk&quot; &quot;groisman&quot;<br> Kazakhstan<br> &quot;sagyntaev&quot; &quot;mamin&quot; &quot;tokayev&quot; &quot;nnazarbayev&quot; &quot;dnazarbayeva&quot; &quot;kulibayev&quot; &quot;masimov&quot;<br> Belarus<br> &quot;alukashenko&quot; &quot;vakulchik&quot; &quot;vlukashenko&quot; &quot;kobyakov&quot; &quot;makei&quot; &quot;myasnikovich&quot;&nbsp; &quot;rumas&quot; &quot;golovchenko&quot;<br> Poland<br> &quot;kaczynski&quot; &quot;duda&quot; &quot;morawiecki&quot; &quot;ziobro&quot;<br> Data coverage<br> Country, news portal, numbr of articles<br> Russia iz.ru 43,782<br> Russia kommersant.ru 46,070<br> Russia novayagazeta.ru 29,357<br> Russia vedomosti.ru 27,797<br> Kazakhstan informburo.kz 29,375<br> Kazakhstan nur.kz 67,350<br> Kazakhstan tengrinews.kz 44,285<br> Kazakhstan zakon.kz 109,442<br> Belarus bdg.by 33,447<br> Belarus belgazeta.by 21,995<br> Belarus sb.by 83,685<br> Ukraine kp.ua 194,792<br> Ukraine segodnya.ua 45,835<br> Ukraine vesti.ua 90,559<br> Poland gazeta.pl 53,321<br> Poland rp.pl 49,587<br> Poland wpolityce.pl 76,625<br> <br> In the provided dataframes the number of observations is smaller, because the issues for which there was no next day issue, were removed.<br> <br> Data was scraped daily between 2017 or 2018 (depending on the country) and January 2021.</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record