Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

29 results for “News Data”

Learn how ShareScore rates datasets ↗
zenodo48/100

South African News Data

<p>This repository hosts a valuable collection of South African local language news compiled and enriched by the University of Pretoria's Data Science for Social Impact research group - https://dsfsi.github.io/. Spanning across South Africa's official languages, these localised news datasets aim to support natural language processing and social computing research focused on domestic challenges.</p><h2>Related Publication</h2><p><strong>Please cite</strong></p><p>@inproceedings{marivate2020investigating, title = {Investigating an Approach for Low Resource Language Dataset Creation, Curation and Classification: Setswana and Sepedi}, author = {Marivate, Vukosi &nbsp;and Sefara, Tshephisho &nbsp;and Chabalala, Vongani &nbsp;and Makhaya, Keamogetswe &nbsp;and Mokgonyane, Tumisho &nbsp;and Mokoena, Rethabile &nbsp;and Modupe, Abiodun}, booktitle = {Proceedings of the first workshop on Resources for African Indigenous Languages}, year = {2020}, address = {Marseille, France}, publisher = {European Language Resources Association (ELRA)}, url = {https://aclanthology.org/2020.rail-1.3}, pages = {15--20}, language = {English}, ISBN = {979-10-95546-60-3}, preprint_url ={https://arxiv.org/abs/2003.04986}, dataset_url = {https://zenodo.org/record/3668495}, keywords = {NLP} }</p><h2>Datasets</h2><h3><a href="http://www.sabc.co.za/">SABC&nbsp;</a></h3><p>We claim no copyright of the SABC original content.&nbsp;</p><ul><li>Motsweding FM (An SABC Setswana radio station) Facebook Page'</li><li>Dikgang Tsa Setswana (SABC Setswana News)</li><li>Thobela FM (An SABC Sepedi radio station) Facebook Page</li><li>Ditaba Tsa Sepedi (SABC Sepedi News)</li></ul><h3>Disclaimer</h3><p>This dataset contains extracted news content from a different sources including, but not limited to: SABC. While efforts were made to ensure the accuracy and completeness of this data, there may be errors or discrepancies between the original publications and this dataset. No warranties, guarantees or representations are given in relation to the information contained in the dataset. The members of the Data Science for Societal Impact Research Group bear no responsibility and/or liability for any such errors or discrepancies in this dataset. The original owners bear no responsibility and/or liability for any such errors or discrepancies in this dataset. It is recommended that users verify all information contained herein before making decisions based upon this information.</p><h2>&nbsp;</h2>

opencc-by-sa-4.0Feb 2020View details →
zenodo44/100

How do Google News' top 100 sources visually represent the data centres' energy footprint?

<p><strong>By querying &quot;data centres&#39; energy footprint&quot; on Google News in incognito mode, the candidate has selected and mapped the top 100 results according to the ranking on May 15, 2022.&nbsp;</strong></p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

South African Disinformation [Fake News] Website Data - 2020

<p>See publication:&nbsp;<strong>Is it Fake? News Disinformation Detection on South African News Websites</strong></p> <p>We used, as sources, investigations by the news websites MyBroadband (<a href="https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/">https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/</a>) and News24 (<a href="https://exposed.news24.com/the-website-blacklist/">https://exposed.news24.com/the-website-blacklist/</a>). These articles covered investigations into disinformation websites in South Africa in 2018. They compiled lists of websites that were suspected to be disinformation. During the period from those articles to present, a number of the websites have become inaccessible or offline. We attempted to use the internet archives <a href="https://archive.org/web/">WayBack Machine</a>&nbsp;we could only get partial snapshots and error messages.</p> <p>A web-scraper only worked for one of the sources although manual editing was still required to clean the text from Javascript code and some paragraph duplicates. On most of the other websites, a web-scraper did not work well as there were too many advertisements and broken parts of pages. Because of all these problems, most of the articles were manually copied and pasted and cleaned in flat files. In some cases, the text of articles could not be copied and was not made part of the South African disinformation corpus.</p> <p><strong>Citing the dataset</strong></p> <blockquote> <p>@inproceedings{de2021fake, title={Is it Fake? News Disinformation Detection on South African News Websites}, author={de Wet, Harm and Marivate, Vukosi}, booktitle={2021 IEEE AFRICON}, pages={1--6}, year={2021}, organization={IEEE} }</p> </blockquote>

opencc-by-sa-4.0Jul 2021View details →
zenodo40/100

Data for PAN at SemEval 2019 Task 4: Hyperpartisan News Detection

<p>Training, validation, and test data for the <a href="https://webis.de/events/semeval-19/">PAN @ SemEval 2019 Task 4: Hyperpartisan News Detection</a>.</p> <p>See the README for details.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Propaganda and fake news on the war in Ukraine: data from Russian-speaking social media communities

<p>The data set contains posts from social media networks popular among Russian-speaking communities. Information was searched based on pre-defined keywords (&quot;war&quot;, &quot;special military operation&quot;,&nbsp;etc.) and is mainly related to the ongoing war in Ukraine with Russia. After a thorough review and analysis of the data, both propaganda and fake news were identified.&nbsp;The collected data is anonymized. Feature engineering and text preprocessing can be applied to obtain new insights and knowledge from this data set. The data set is useful for the study of information wars and propaganda identification.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

News data for studying media exposure to the Boston Marathon bombings

<p>The news data sets released here have been used to study the relationship between media exposure and individuals&#39; threat perception. Media exposure to mass violence has been shown to have a detrimental impact on people&#39;s threat perception and mental wellness, but little has been done to explore how exposure to different news content may impact mental health in people&#39;s everyday lives. In our study, we empirically test how emotionally potent media coverage of a real-world threat, namely, the Boston Marathon bombings occurred in 2013, alters threat perception of the community members over the first and the third anniversaries (in 2014 and 2016).</p> <p>The data were collected using a wave-based longitudinal design. There are two data sets, and each covers three waves:</p> <ul> <li>Dataset (I) -- news coverage before (Wave 1), during (Wave 2), and after (Wave 3) 2014 anniversary</li> <li>Dataset (II) -- news coverage before (Wave 1), during (Wave 2), and after (Wave 3) 2016 anniversary</li> </ul> <p>The collection procedure was informed by our survey study. Based on the survey completed by our subjects, we identified the four most frequent news outlets in the response: Metro (MT), New York Times (NY), Boston Globe (BG), and Boston Herald (BH). Other outlets, such as USA Today and Wall Street Journal, were reported by less than ten respondents. Therefore, our data collection focused on the news published by the four most frequent outlets.</p> <p>The data sets include the metadata of the news coverage over the aforementioned six waves. The raw content of the news stories was removed to respect the copyright owners.</p> <p><strong>Summary of the data collection procedure </strong></p> <p>We used news aggregators including Google and Yahoo news, to retrieve news articles published by the four outlets on a daily basis. We first collected the URLs of the news articles from the news aggregators and retrieved and parsed the news content using an HTML parser. In total, we collected over 38.5K and 54.1K news articles in dataset I and II, respectively.</p> <p>There are six files; each correspond to news coverage from the outlets in each wave. In these files, each line contains four columns: outlet, time, title, url which indicate the outlet of each news article, the time of publishing, the title of the article, and the URL to the article.</p> <p>We are making the data sets available for academic researchers and public use, to enable the discovery of new insights and development of better techniques to improve crisis communication and mental wellness.</p>

openother-openDec 2017View details →
zenodo40/100

Robberies for cigarettes news reports 2009-2018 data

<p>Dataset used to analyze New Zealand news reports of robberies of stores for tobacco during 2009-2018.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Supporting Information (software and data) for: Client-side energy and GHGs assessment of advertising and tracking in the news websites

<p>This is the open data and free/libre and open source software repository for the article &quot;Client-side energy and GHGs assessment of advertising and tracking in the news websites&quot; by Fabio Pesari, Giovanni Lagioia, Annarita Paiano.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Documents used in the PLANET4B D1.1 analysis of biodiversity discourse by news outlets - 2010 and 2022 Data

<p>Data used to analyse biodiversity discourse in news outlet as part of Deliverable D1.1. of the Planet4B Project.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Data for manuscript "Reciprocal Radicalization: The Rise of Culture War Terminology in British and American News Coverage"

<p>This data set contains frequency counts of target words in 16&nbsp;million news and opinion articles from 10 popular news media outlets in the United Kingdom: The Guardian, The Times, The Independent, The Daily Mirror, BBC, Financial Times, Metro, Telegraph, The and The Daily Mail plus a few additional American-based outlets used for comparison reference. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice, social justice related terms or counterreaction to it. A&nbsp;few additional&nbsp;words are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 3&nbsp;of the main manuscript is available in the outlet&#39;s online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We derived relative frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions.&nbsp;Tokens were lowercased prior to estimating frequency counts.&nbsp;To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1&nbsp;million words of article content from an outlet.&nbsp;</p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript&nbsp;</p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>-targetWordsInArticlesCountsGuardianExampleWords&nbsp;contains counts of target words in outlets articles as well as total counts of words in articles for illustrative Figure 1 in main manuscript</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail&nbsp;to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word &quot;Facebook&quot; in an article footnote encouraging article readers to follow the journalist&rsquo;s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. To conclude, in a data analysis of 16 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 of main manuscript for supporting evidence).</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Data for manuscript "The Prevalence of Terms Denoting Far-right and Far-left Political Extremism in U.S. and U.K. News Media"

<p>This data set belongs to an academic manuscript examining longitudinally (2000-2019) the prevalence of terms denoting far-right and far-left political extremism in a large corpus of more than 32 million written news and opinion articles from 54 news media outlets popular in the United States and the United Kingdom.</p> <p>The textual content of news and opinion articles from the 54 outlets listed in the main manuscript is available in the outlet&#39;s online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions.&nbsp;Tokens were lowercased prior to estimating frequency counts.&nbsp;To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year.&nbsp;</p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript&nbsp;</p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>&nbsp;</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word &quot;Facebook&quot; in an article footnote encouraging article readers to follow the journalist&rsquo;s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in &quot;The Hill&quot; and &quot;Newsmax&quot; news outlets where XPath expressions had some shortfalls at precisely capturing articles&rsquo; content. For &quot;The Hill&quot;, in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles&rsquo; words that is about 40% smaller than the total population of articles words for those three years. In the case of &quot;NewsMax&quot;, the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words.</p> <p>To conclude, in a data analysis of 32 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 in the main manuscript for illustration of the accuracy of the frequency counts).</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"

<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States.&nbsp;</p> <p>The set of 47 news media outlets analysed (listed in Figure 1&nbsp;of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets&rsquo; ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript.&nbsp;</p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets&rsquo; online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles&rsquo; HTML raw data using outlet-specific XPath expressions.&nbsp;</p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000&rsquo;s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000.&nbsp;Figure S 1 in the SI&nbsp;reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small.&nbsp;Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million&nbsp;headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth.&nbsp;</p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely:&nbsp;</p> <p>Siebert/sentiment-roberta-large-english from&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english.&nbsp;This model is a fine-tuned checkpoint of&nbsp;<a href="https://huggingface.co/roberta-large">RoBERTa-large</a>&nbsp;(<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library&nbsp;https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with&nbsp;sentiment-roberta-large-english&nbsp;</p> <p>DistilRoberta&nbsp;j-hartmann/emotion-english-distilroberta-base from&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of&nbsp;<a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with&nbsp;&nbsp;Ekman&#39;s 6 basic emotions, plus a neutral class.&nbsp;The model was trained on 6 diverse datasets. Please refer to the original author at&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning.&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the&nbsp;siebert/sentiment-roberta-large-english Transformer model.&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar&nbsp;URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset.&nbsp;https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the&nbsp;j-hartmann/emotion-english-distilroberta-base Transformer model.&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>

opencc-by-4.0Jul 2021View details →
dryad36/100

Data from: Fake news? The impact of information mismatch in mating behaviour

<p>Multiple cues are often used for mate choice in complex environments, potentially entailing mismatches between different sources of information. We address the consequences thereof for receivers using the spider mite <em>Tetranychus urticae</em>, in which virgin females are highly valuable mates compared to mated females, given first male sperm precedence. Accordingly, males prefer virgins and distinguish them using cues from the females and/or that are present on the substrate. Whereas the former are more reliable, the latter may allow for a faster or more long-distance response. However, there can be mismatched information between cues as females move and/or mate. Here, we tested the consequences thereof by exposing males to mated or virgin females on patches previously impregnated with cues deposited by females of either mating status. Male mating attempts were solely affected by substrate cues while female acceptance and the number of mating events were independently affected by both cues. Copulation duration, in contrast, depended mainly on the mating status of the female, with the number of copulations and the total time spent mating being intermediate in environments with mismatched information. Ultimately, male survival costs mirrored male investment in mating. These results suggest that, in environments with mismatched information, the substrate cues left by females are instrumental for males to find their mates, but they can also lead to males paying survival costs without the associated benefit of mating effectively, or suffering reduced costs at the expense of losing effective mating opportunities. The benefit of using multiple cues will then hinge upon the frequency of information mismatch, which itself should vary with the dynamics of populations.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Crunchbase in RDF: A Large Data Set About Jobs, Websites, Organizations, News, People, Products, and Acquisitions

<p><strong>CrunchBase</strong> in an online platform providing information about startups and technology companies, including related entities such as the products they sell, key people they employ, and investments they made and received.</p> <p>We provide here an <strong>RDF data set of Crunchbase</strong> as of October 2015. The data set contains information about</p> <ul> <li>1,946,435 jobs</li> <li>1,348,449 websites</li> <li>567,937 organizations</li> <li>519,763 news</li> <li>430,093 people</li> <li>60,076 products, and</li> <li>33,127 acquisitions.</li> </ul> <p>The data set has been used, among other things, for data integration with financial data sources to evaluate the performance of particular companies and for monitoring news to find statements that are not in Crunchbase as an RDF knowledge graph yet.</p> <p>Note that the provided data set was created in October 2015 when all Crunchbase data was <strong>licensed under Creative Commons Attribution-NonCommercial License 4.0 (CC-BY-NC) and partly under Creative Commons Attribution License 4.0 (CC-BY)</strong>. Also the provied<strong> data set is licensed under these licenses.</strong> Concerning licensing of current Crunchbase data, we can refer to <a href="https://about.crunchbase.com/terms-of-service/">https://about.crunchbase.com/terms-of-service/</a>.</p> <p>For <strong>more information</strong> about the data set, see our paper <a href="http://dbis.informatik.uni-freiburg.de/content/team/faerber/papers/CrunchBaseWrapper_SWJ2017.pdf">A Linked Data Wrapper for CrunchBase.</a></p> <p>When you use the data set, please <strong>cite</strong> us as follows:</p> <blockquote> <p>Michael F&auml;rber, Carsten Menne, Andreas Harth. &ldquo;A Linked Data Wrapper for CrunchBase&rdquo;. In: Semantic Web Journal 9(4). IOS Press, 2018, pp. 505&ndash;5015. (<a href="https://dblp.org/rec/bibtex/journals/semweb/FarberMH18">BibTeX entry at DBLP</a>)</p> </blockquote>

opencc-by-nc-4.0Aug 2016View details →
zenodo36/100

PIE News Early User Survey data

<p>The PIE News &ldquo;Early User Survey&rdquo; is an evaluation-related activity which is grounded on the adoption of a subjective understanding of the term &ldquo;precarious&rdquo;, as a condition and not as a contractual status. This means that one can be precarious also with a permanent work contract, because the many possible dimensions of uncertainty (e.g.: employer under risk, crisis of the job market) drive an objective or just perceived individual precarious feeling. The more objective understanding of the precarious experience will probably be a project outcome.</p> <p>The survey ended the 25 October 2016 involving people contacted by the PIE News pilot leaders in Italy, Croatia and the Netherlands.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Data for manuscript "Prevalence in News Media of two Competing Hypotheses about COVID-19 Origins"

<p>The Covid-19 pandemic has been one of the most disruptive and painful phenomena of the last few decades. As of July 2021, the origins of the SARS-CoV-2 virus that caused the outbreak remain a mystery. This work analyzes the prevalence in news media articles of two popular hypotheses about SARS-CoV-2 virus origins: the natural emergence and the lab-leak hypotheses.&nbsp;</p> <p>This data set contains frequency counts of target words in news and opinion articles from 12&nbsp;popular news media outlets. The target words are listed in the associated manuscript and are mostly words associated with the Covid-19 pandemic.&nbsp;</p> <p>The list of compressed files in this data set is listed next:</p> <p>targetWordsInArticlesCounts.rar&nbsp;contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>targetWordsFrequencies.rar daily, weekly, monthly&nbsp;word frequencies</p> <p>wordEmbeddingModels.rar monthly embedding models of news outlets content</p> <p>analysisScripts.rar analysis notebooks</p> <p>The textual content of news and opinion articles from the outlets is available in the outlet&#39;s online domains and/or public cache repositories such as Google cache, The Internet Wayback Machine, and Common Crawl. We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific XPath expressions.&nbsp;Tokens were lowercased prior to estimating frequency counts.&nbsp;</p> <p>Yearly frequency usage of a target word in an outlet in any given temporal interval ( daily, weekly, monthly) was estimated by dividing the total number of occurrences of the target word in all articles of a given temporal interval by the number of all words in all articles of that temporal interval. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.&nbsp;Most of the incorrect frequency counts are minor deviations from the actual counts such as for instance counting a word in an article footnote encouraging article readers to find related articles and that the XPath expression might mistakenly include&nbsp;as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in the WSJ where in less than 5% of the articles XPath expressions failed to capture the article&#39;s main text content. Other outlets articles samples sizes might not be comprehensive but, to the best of our knowledge, they are representative and include tens of thousands of articles per outlet/year. To conclude, in a data analysis of over 1.5&nbsp;million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 of main manuscript for supporting evidence).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Raw data for manuscript: Use of immunology in news and YouTube videos in the context of COVID-19: politicization and information bubbles

<p>Coding of newsarticles and videos related to immunology and COVID-19 in Italian and English</p>

opencc-by-4.0Sep 2023View details →
dryad36/100

Data from: Fake news? The impact of information mismatch in mating behaviour

Open the record for dataset details and reuse information.

publicMay 2024View details →
zenodo32/100

Data set for "Frequency of Prejudice Coverage in News Media Worldwide"

<p>Previous research has identified a post-2010 sharp increase of terms used to denounce prejudice (i.e. racism, sexism, homophobia, Islamophobia, anti-Semitism, etc.) in U.S. and U.K. news media content. Here, we extend previous analysis to an international sample of news media organizations. Thus, we quantify the prevalence of prejudice-denouncing terms and social justice associated terminology (diversity, inclusion, equality, etc.) in over 98 million news and opinion articles across 124 popular news media outlets from 36 countries representing 6 different world regions: English-speaking West, continental Europe, Latin America, sub-Saharan Africa, Persian Gulf region and Asia. We find that the post-2010 increasing prominence in news media of the studied terminology is not circumscribed to the U.S. and the U.K. but rather appears to be a mostly global phenomenon starting in the first half of the 2010s decade in pioneering countries yet largely prevalent around the globe post-2015. However, different world regions&rsquo; news media emphasize distinct types of prejudice with varying degrees of intensity. We find no evidence of U.S. news media having been first in the world in increasing the frequency of prejudice coverage in their content. The large degree of temporal synchronicity with which the studied set of terms increased in news media across a vast majority of countries raises important questions about the root causes driving this phenomenon.</p> <p>We provide here a reproducibility data set of counts of target terms and total number of unigrams in each article analyzed plus a Google searchable prefix of the headline of the article for verification of the integrity of the data set and the frequency counts.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Data for manuscript: "Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets"

<p>This data set contains material for the purpose of scientific reproducibility of the accompanying manuscript &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot;.</p> <p>Note that this data set is distributed with an Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) License. NonCommercial means you&nbsp;may not use the material for commercial purposes. NoDerivatives means if you remix, transform, or build upon the material, you may not distribute the modified material. Attribution means you must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. See attached license terms for details.</p> <p>The work &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot; describes an analysis of political associations in 27 million diachronic (1975-2019) news and opinion articles from 47 news media outlets popular in the United States. We use embedding models trained on individual outlets content to quantify outlet-specific latent associations between positive/negative sentiment words and terms loaded with political connotations such as those describing political orientation, party affiliation, names of influential politicians and ideologically aligned public figures.&nbsp;</p> <p>News and opinion articles from the outlets listed in Figure 3 are available in the outlet&#39;s online domains and/or public cache repositories such as Google cache, The Internet Wayback Machine [31] and Common Crawl [32]. This work has not analyzed video or audio content of news media organizations, except when the outlet explicitly provides a transcript of such content in article form.<br> The temporal coverage of articles from different news outlets is not uniform. For most media organizations, news articles availability in their online domains or Internet cache backups becomes sparse as a function of articles&rsquo; age. This is not the case for some news outlets, where availability of news articles goes back to the 1970s. The Supplementary Material (SM) illustrates the time ranges of article data analyzed based on news outlets articles online availability.</p> <p>Textual content included in our analysis is circumscribed to the articles&rsquo; headlines and main text and does not include other article elements such as figure captions. Targeted textual content was located in HTML raw data using outlet specific XPath expressions. Tokens were lowercased prior to estimating embedding models. Markup language tags, URLs, nonalphanumeric characters, punctuation, digits, 330 common stop words and multiple spaces were removed prior to estimating word embeddings models.<br> All the analysis scripts and the diachronic word embedding models built from each of the 47 news media outlets analyzed in this work are available in this repository.</p> <p>For the purpose of reproducibility, we also provide in the above repository the articles&rsquo; text used to train the news outlets embedding models with the caveat that outlets articles not accessible without a subscription have been excluded. Also, for the included articles, stop words have been removed and the remaining words have been randomly scrambled within a sliding window of size 10 to render the articles incomprehensible to a human reader. These steps have been taken to not infringe articles copyright. These preprocessing steps have only minor impact on Continuous Bag of Words (CBOW) word2vec and the results reported in this work are similar when using the scrambled articles text to train outlet-specific embedding models.</p> <p>We derived outlet-specific word embedding models at every five-year time intervals within the 1975-2019 time range. The gensim [33] implementation of word2vec was used to train the embedding models. The continuous bag of words (CBOW) architecture performed slightly better than the Skip-Gram architecture in commonly used validation metrics so it was used for all subsequent analysis.&nbsp;</p> <p>For training the word embedding models, the following parameters were used: vector dimensions=300, window size=10, negative sampling=10, down sampling frequent words = 0.0001, minimum frequency count of 5 (only terms that appear more than 5 times in the corpus were included into the word embedding model vocabulary), number of training iterations (epochs) through the corpus=5. The exponent used to shape the negative sampling distribution was the default 0.75.&nbsp;</p> <p>Outlet-specific embedding models performance across a range of commonly used semantic, syntactic and analogy tasks was similar to popular pre-trained embedding models trained on corpora such as Twitter or Google books on similarity, association and word analogy tasks, see Supplemeentary Material of the manuscript for detailed validation tests results.</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record