Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
AG's news corpus (AGNEWS)
<p>AG’s news corpus (AGNEWS): This AG’s corpus of news articles was collected from the web. The whole corpus contains 496,835 categorized news articles from more than 2000 news sources. Four largest classes (World, Sports, Business and Sci/Tech) from this corpus were chosen to construct the dataset used in the experiments, using only the title and description fields</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>
Argentina news from Pagina12 Newspaper
<p>This dataset contains information from the web page of the Argentinian newspaper "Pagina12" (https://www.pagina12.com.ar/). The dataset was developed with a Python code in the context of academical studies (Data Science master). The final purpose was only for academic research.</p> <p>The dataset includes different fields of an article link of the cited newspaper, specifically the code was implemented in Python using Srapy library. Hence, the code scrap the final link of each notice or new from the different sections. For each notice or new, the following fields have been extracted: title, date, author, section and url of the new.</p>
Occitan Corpus from Lo Congrès news
<p>This corpus contains was automatically compiled from the news on locongres.org bilingual website. This work was made by Lo Congrès permanent de la lenga occitana (https://locongres.org) as a part of its project "Còrpus" (http://abrac.at/corpusproject).</p> <p>It contains csv files with occitan sentences aligned with their french translations.</p>
A dataset with news messages from a Russian and a Ukrainian TV news channels
<p>This data was used in the following publications:</p> <p>1. Koltsova, O., & Pashakhin, S. (2019). Agenda divergence in a developing conflict: Quantitative evidence from Ukrainian and Russian TV newsfeeds. Media, War & Conflict, 1750635 21982987. <a href="https://doi.org/10.1177/1750635219829876">https://doi.org/10.1177/1750635219829876</a></p> <p>2.Pashakhin S. Topic Modeling for Frame Analysis of News Media // Proceedings of the AINL FRUCT 2016. С. 103-105 – URL: <a href="http://fruct.org/publications/abstract-AINL-FRUCT-2016/files/Pas.pdf">http://fruct.org/publications/abstract-AINL-FRUCT-2016/files/Pas.pdf</a></p> <p>The dataset contains 45,009 news messages collected from official websites of a Russian (Channel One) and a Ukrainian (Channel 5) TV channels. Ukrainian news items were translated into Russian.</p> <p>The dataset has six variables:</p> <ul> <li>text -- a news item;</li> <li>channel -- a source of an item ('first' -- Russian TV channel, 'five' -- Ukrainian TV channel);</li> <li>date -- the date of publishing;</li> <li>url -- links to original news messages.</li> </ul>
ISPON: A New Dataset for Identifying Sources in Political Online News
<p>This dataset contains a set of annotations for informational news sources (such as eyewitnesses, public officials, academic experts, reports, or other documentation) that provide support for claims made within online political news articles. Our dataset contains fine-grained annotations on the sources cited within each article, including in-text notations highlighting the words or phrases signaling a source. The dataset comprises annotations for nearly 2,500 articles covering 47 outlets. In addition, the dataset includes a larger set of >150,000 URLs from 92 outlets.</p>
BreXLiMe: A Semantically Enriched Dataset With News Articles, Micro-Posts, and TV Shows Related to the Brexit
<p>We provide a <strong>large data set of media content metadata</strong> from various media sources (including online news sites, social media, and live-TV) in three languages (<strong>English, German, and Spanish</strong>). Overall, the data set contains rich metadata for about <strong>240 thousand news articles, 12 million micro-posts, and 900 TV shows</strong>. All media content information has been semantically enriched with annotations of both entities and categories from DBpedia.</p> <p>The data can be used as a valuable data basis for applications and studies of various disciplines (e.g., social studies, political science, and humanities) on the case of Brexit, particularly on the <strong>media landscape before the Brexit referendum held on June 23, 2016</strong>.</p> <p>We provide the data set in the RDF serialization format Turtle (.ttl) as well as in XML.</p> <p>If you use our data set, please <strong>cite</strong> it as follows:</p> <pre><code>Lei Zhang, Maribel Acosta, Michael Färber, Steffen Thoma and Achim Rettinger. "BreXearch: Exploring Brexit Data Using Cross-Lingual and Cross-Media Semantic Search". In: Proceedings of the ISWC 2017 Posters & Demonstrations Track within the 16th International Semantic Web Conference (ISWC 2017). Vienna, Austria, 2017.</code></pre> <p> </p>
MIDAS hand-annotated news articles
<p>This dataset was produced in 2020 from the data collected throughout 2019 for the development of the MIDAS project (http://www.midasproject.eu/)</p> <p>The data is distributed throughout 5 topics:</p> <p>- EUS: Childhood Obesity (UC Basque Country)<br> - FIN: Mental Health (UC Finland)<br> - IRE: Diabetes (UC Ireland)<br> - NIR: Children in Care (UC Northern Ireland)<br> - INF: Infectious Diseases including Coronavirus (UC Influenzanet)</p> <p>The available data comes in 3 kinds and file formats:</p> <p>TXT - the source of news including ID, title and body of text<br> CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br> JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p> <p>The CSV files with name starting in "f1_", "pr_", "re_" are the results of the F1/Precision/Recall evaluation for each of the cases.</p> <p>## AUTHORS</p> <p>Joao Pita Costa, Anthony Staines, Jarmo Pääkkönen, Jenni Konttila, Joseba Bidaurrazaga, Oihana Belar, Christine Henderson</p> <p>## ACKNOWLEDGMENTS</p> <p>This work was supported by the European Commission H2020 project MIDAS (G.A. nr. 727721). </p> <p><br> ## LICENSE</p> <p>This dataset is licensed over Creative Commons.</p>
Meneame news dataset
<p>This dataset contains 150 news scrapped from Meneame from time range 17-10-2020 to 20-10-2020</p>
COVID Fake News Dataset
<p><strong>Context</strong></p> <p>The dataset contains the list of COVID Fake News/Claims which is shared all over the internet.</p> <p><strong>Content</strong></p> <ol> <li>Headlines: String attribute consisting of the headlines/fact shared.</li> <li>Outcome: It is binary data where 0 means the headline is fake and 1 means that it is true.</li> </ol> <p><strong>Inspiration</strong></p> <p>In many research portals, there was this common question in which the combined fake news dataset is available or not. This led to the publication of this dataset.</p>
Contextualizing Trending Entities in News Stories
<p>This repository contains the enrichments for the dataset <a href="https://catalog.ldc.upenn.edu/LDC2008T19">The New York Times Annotated Corpus</a> developed for the paper:</p> <p>“Marco Ponza, Diego Ceccarelli, Paolo Ferragina, Edgar Meij, Sambhav Kothari. Contextualizing Trending Entities in News Stories. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining (WSDM 2021).”</p> <p>It includes a total of 149 trends constituted by 120K entities. The goal is to retrieve a set of entities ranked with respect to their usefulness in explaining why a given trending entity is actually trending.</p> <p><strong>Format</strong></p> <p>The repository contains the enrichments in JSON format.</p> <p>The news stories of the New York Times from which these enrichments have been developed are available from <a href="https://catalog.ldc.upenn.edu/LDC2008T19">LDC</a>.</p> <p><strong>Data Splits</strong></p> <p>We perform two kinds of evaluation.</p> <ol> <li>Unsupervised evaluation, where we use the complete dataset of 149 trends as a benchmark.</li> <li>Supervised evaluation, where we train/tune our models on a training/development set and we test them on a test set.</li> </ol> <ul> <li>The training set contains 50 trends constituted by 36.3K entities from 1996 to 2000.</li> <li>The development set contains 34 trends constituted by 26.7K entities from 2000 to 2002.</li> <li>The test set contains 65 trends constituted by 57K entities from 2002 to 2007.</li> </ul> <p>Use</p> <p>Please cite the data set and the accompanying paper if you found the resources in this repository useful:</p> <p>@inproceedings{ponza2021,<br> Title = {Contextualizing Trending Entities in News Stories},<br> author = {Ponza, Marco and Ceccarelli, Diego and Ferragina, Paolo and Meij, Edgar and Kothari, Sambhav},<br> Booktitle = {Proceedings of the 14th ACM International Conference on Web Search and Data Mining},<br> Year = {2021},<br> }</p>
News of CanalUGR tracked on Google News, Yahoo! News and Bing News
Dataset contains 613 news of CanalUGR (University of Granada Communication Office) tracked on the main online news aggregators (Google News, Yahoo! News and Bing News). We include: number in CanalUGR, media, country, type.
Hacker News Curated Comments Dataset
<p>A curated dataset from fh-bigquery:hackernews.stories</p> <p>Only HN stories with more than 10 comments are included, and only comments from users with more than 10 comments are included.</p>
Cropped News Article from University of Potsdam News Site
<p>This data was cropped from the <a href="https://www.uni-potsdam.de/de/nachrichten">University of Potsdam news website</a>. The annotated labels are made by the writers of the articles.</p>
Rare Diseases hand-annotated news articles and research articles
<p>This dataset was produced in 2023 from the data collected throughout 2022 from MEDLINE (scientific articles) and from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p><p>The data is distributed across 16 diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p><p>The available data comes in 2 kinds and file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br>JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p><p>The CSV files with name starting in "f1_", "pr_", "re_" are the results of the F1/Precision/Recall evaluation for each of the cases.</p><p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdolšek Draksler (domain expert) </p>
News Ninja Dataset
<p><strong>About</strong><br>Recent research shows that visualizing linguistic media bias mitigates its negative effects. However, reliable automatic detection methods to generate such visualizations require costly, knowledge-intensive training data. To facilitate data collection for media bias datasets, we present News Ninja, a game employing data-collecting game mechanics to generate a crowdsourced dataset. Before annotating sentences, players are educated on media bias via a tutorial. Our findings show that datasets gathered with crowdsourced workers trained on News Ninja can reach significantly higher inter-annotator agreements than expert and crowdsourced datasets. As News Ninja encourages continuous play, it allows datasets to adapt to the reception and contextualization of news over time, presenting a promising strategy to reduce data collection expenses, educate players, and promote long-term bias mitigation.<br><br></p> <p><strong>General<br></strong>This dataset was created through player annotations in the News Ninja Game made by ANON. Its goal is to improve the detection of linguistic media bias. Support came from ANON. None of the funders played any role in the dataset creation process or publication-related decisions.</p> <p>The dataset includes sentences with binary bias labels (processed, biased or not biased) as well as the annotations of single players used for the majority vote. It includes all game-collected data. All data is completely anonymous. The dataset does not identify sub-populations or can be considered sensitive to them, nor is it possible to identify individuals.</p> <p>Some sentences might be offensive or triggering as they were taken from biased or more extreme news sources. The dataset contains topics such as violence, abortion, and hate against specific races, genders, religions, or sexual orientations.</p> <p> </p> <p><strong>Description of the Data Files<br></strong>This repository contains the datasets for the anonymous News Ninja submission. The tables contain the following data:</p> <p><strong>ExportNewsNinja.csv</strong>: Contains 370 <a href="https://www.kaggle.com/datasets/timospinde/babe-media-bias-annotations-by-experts?resource=download">BABE sentences</a> and 150 new sentences with their text (sentence), words labeled as biased (words), BABE ground truth (ground_Truth), and the sentence bias label from the player annotations (majority_vote). The first 370 sentences are re-annotated BABE sentences, and the following 150 sentences are new sentences.</p> <p><strong>AnalysisNewsNinja.xlsx</strong>: Contains 370 <a href="https://www.kaggle.com/datasets/timospinde/babe-media-bias-annotations-by-experts?resource=download">BABE sentences</a> and 150 new sentences. The first 370 sentences are re-annotated BABE sentences, and the following 150 sentences are new sentences. The table includes the full sentence (Sentence), the sentence bias label from player annotations (isBiased Game), the new expert label (isBiased Expert), if the game label and expert label match (Game VS Expert), if differing labels are a false positives or false negatives (false negative, false positive), the ground truth label from BABE (isBiasedBABE), if Expert and BABE labels match (Expert VS BABE), and if the game label and BABE label match (Game VS BABE). It also includes the analysis of the agreement between the three rater categories (Game, Expert, BABE).</p> <p><strong>demographics.csv</strong>: Contains demographic information of News Ninja players, including gender, age, education, English proficiency, political orientation, news consumption, and consumed outlets.</p> <p> </p> <p><strong>Collection Process<br></strong>Data was collected through interactions with the NewsNinja game. All participants went through a tutorial before annotating 2x10 BABE sentences and 2x10 new sentences. For this first test, players were recruited using Prolific. The game was hosted on a costume-built responsive website. The collection period was from 20.02.2023 to 28.02.2023. Before starting the game, players were informed about the goal and the data processing. After consenting, they could proceed to the tutorial.</p> <p>The dataset will be open source. A link with all details and contact information will be provided upon acceptance. No third parties are involved.</p> <p>The dataset will not be maintained as it captures the first test of NewsNinja at a specific point in time. However, new datasets will arise from further iterations. Those will be linked in the repository. Please cite the NewsNinja paper if you use the dataset and contact us if you're interested in more information or joining the project.</p>
Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News
<p>Data supporting "Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News". <br><br></p> <p> If you use this dataset in your own research, please cite this paper:</p> <p>```<br>@misc{abels2024mitigating,<br> title={Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News}, <br> author={Axel Abels and Elias Fernandez Domingos and Ann Nowé and Tom Lenaerts},<br> year={2024},<br> eprint={2403.08829},<br> archivePrefix={arXiv},<br> primaryClass={cs.HC}<br>}<br>```</p> <p> </p> <table> <tbody> <tr> <td><strong>column name</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>treatment</td> <td>identifier for the set of headlines presented to the participant</td> </tr> <tr> <td>trial</td> <td>trial/round in which the headline was presented </td> </tr> <tr> <td>arm</td> <td>which "arm" the headline was presented as (0=left, 1=middle, 2=right)</td> </tr> <tr> <td>advice</td> <td>the participant's response (0=very unlikely, 0.25=unlikely, 0.5=undecided, 0.75=likely, 1=very likely)</td> </tr> <tr> <td>genuine</td> <td>whether the headline was genuine (1) or altered (0)</td> </tr> <tr> <td>headline</td> <td>the headline as shown to the participant</td> </tr> <tr> <td>original</td> <td>the headline before a possible alteration</td> </tr> <tr> <td>expert_id</td> <td>participant's identifier</td> </tr> <tr> <td>sentiment</td> <td>whether the headline reported a negative (-1) or positive (1) outcome</td> </tr> <tr> <td>expert:ethnicity</td> <td>the participant's ethnicity</td> </tr> <tr> <td>expert:sex</td> <td>the participant's sex</td> </tr> <tr> <td>expert:age</td> <td>the participant's age</td> </tr> <tr> <td>outcome:white, outcome:black, outcome:young, outcome:old, outcome:male, outcome:female</td> <td>whether the headline reported a negative (-1) or positive (1) or neutral (0) outcome for the specified group</td> </tr> <tr> <td>trial_time</td> <td>how long the participant took to respond to the trial/round</td> </tr> </tbody> </table> <p><strong>abstract</strong><br>Individual and social biases undermine the effectiveness of human advisers by inducing judgment errors which can disadvantage protected groups. In this paper, we study the influence these biases can have in the pervasive problem of fake news by evaluating human participants' capacity to identify false headlines. By focusing on headlines involving sensitive characteristics, we gather a comprehensive dataset to explore how human responses are shaped by their biases. Our analysis reveals recurring individual biases and their permeation into collective decisions. We show that demographic factors, headline categories, and the manner in which information is presented significantly influence errors in human judgment. We then use our collected data as a benchmark problem on which we evaluate the efficacy of adaptive aggregation algorithms. In addition to their improved accuracy, our results highlight the interactions between the emergence of collective intelligence and the mitigation of participant biases. </p>
CORAPE News
<p>The dataset was obtained from <a href="https://radio.corape.org.ec/">https://radio.corape.org.ec/</a> (CORAPE: Coordinator of Popular and Educational Community Media of Ecuador). Register news information like, tag, category, headline, number of visualizations and the news links about a set of news and interviews in Ecuador.</p>
TNCD (Twitter News Cascade Dataset)
<p>The TNCD (Twitter News Cascade Dataset) 1,200 news items posted on Twitter. It contains post URLs retrieved from sportsmen, politicians and news channels accounts, most from September 2020. For more information<br> please see <a href="https://www.sciencedirect.com/science/article/pii/S1568494621003367">https://www.sciencedirect.com/science/article/pii/S1568494621003367</a></p>
Mapping 'the constructive turn' in comment sections of news websites
<p>The research project critically examines the guidelines of the comments sections of the twenty largest online news outlets over the last ten years. Rather than focusing on the familiar negative comments of news consumers and their narratives, we analyze and compare the news outlets’ guidelines and how they have led in what we call ‘a constructive turn’. We propose our own theoretical framework to analyze what is encouraged and what is discouraged in news outlets’ guidelines. Results show an increasing focus on constructiveness in the guidelines of the comment sections and a shift to more positivity, rather than on deleting and filtering negative or toxic comments. Although platforms differ in their views on the role of commenting and the definition of constructiveness, the turn towards the constructive design of the commenting platform is shared among them.</p> <p>This dataset contains the commentary guidelines in the top 20 English-language online news websites of December 2020 based on research conducted by Similar Web (Source: Similar Web for Gazette). For each news publication, the current commentary guidelines were scrapped from the internet, alongside earlier versions of their guidelines. In total, three moments were used to map the guidelines: 2021, 2015 and 2010. The content was analysed through coding using Nvivo software. We applied a bottom-up approach - by creating simple codes and eventually grouping them together. Each set of guidelines was coded on what behaviour was <em>encouraged</em> and what was <em>discouraged</em> by the news outlet, and what kind of <em>discussion environment </em>the news outlet expects from their commenters in general (e.g. entertaining, healthy, inclusive etc.). </p> <p>This dataset contains coded content for the project. Following logic was used in uploading the documents:</p> <p>1 - Nvivo project file - can be opened using Nvivo for Mac - contains all information (files, codes, etc.)</p> <p>We also upload more user-friendly data (the following documents are uploaded in MS Word format):</p> <p>2 - Codebook (provides the logical structure of coding applied + number of codes for each category)<br> 3 - Code excerpts for discouraged elements found in the content<br> 4 - Code excerpts for encouraged elements found in the content<br> 5 - Code excerpts for discussion environment elements found in the content</p> <p>Disclaimer: The user-generated content guidelines of news media companies are their own intellectual property and we do not own any rights to it. </p> <p><br> </p>
Data for PAN at SemEval 2019 Task 4: Hyperpartisan News Detection
<p>Training, validation, and test data for the <a href="https://webis.de/events/semeval-19/">PAN @ SemEval 2019 Task 4: Hyperpartisan News Detection</a>.</p> <p>See the README for details.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.