Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
The impact of news exposure on collective attention in the United States during the 2016 Zika epidemic
<p>This repository contains the data of the study "The impact of news exposure on collective attention in the United States during the 2016 Zika epidemic".</p> <p><strong>Epidemiological data</strong></p> <p>The folder <em>zika_USA_weekly_cases_2016.zip </em>contains weekly ZIKV incidence counts reported by the US Centers for Disease Control and Prevention in 2016, by state. Data were extracted from reports made publicly available by the CDC at: <a href="https://zenodo.org/record/584136#.Xk07-RNKjOQ">https://zenodo.org/record/584136#.Xk07-RNKjOQ</a> </p> <p><strong>Web news data</strong></p> <p>The file <em>news_GDELT_data.csv.gz </em>contains all news items extracted from the GDELT platform (<a href="https://www.gdeltproject.org/">https://www.gdeltproject.org/</a>) matching <em>TAX_DISEASE_ZIKA </em>as a Theme, and <em>United_States</em> as a Location in the GDELT platform. </p> <p><strong>TV closed captions</strong></p> <p>The file <em>zika_TV_mentions_dataframe.csv </em>contains all the TV news items of 2016 matching the word ``Zika" in the TV News Archive https://archive.org/details/tv</p> <p><strong>Wikipedia pageview counts</strong></p> <p>Dataset 1: <em>wikipedia_dataset1_zika_daily_pageview_usa.csv</em></p> <p>Content of each line of the dataset: day, pageview_count</p> <p>The dataset contains the daily number of pageview counts of 128 different Wikipedia pages related to the Zika virus (aggregated and summed to total) originated in the United States, from January 1st to December 31st, 2016.</p> <p>Dataset 2: <em>wikipedia_dataset2_zika_daily_pageview_bystate.zip</em></p> <p>Content of each line of the dataset: day, pageview_count, state</p> <p>The dataset contains the daily number of pageview counts of 128 different Wikipedia pages related to the Zika virus (aggregated and summed to total) originated in the United States, disaggregated by state, from January 1st to December 31st, 2016.</p> <p>Dataset 3: <em>wikipedia_dataset3_zika_pagecount_by_city.csv</em></p> <p>Content of each line of the dataset: US_city, pageview_count_Zika,pageview_count_total</p> <p>The dataset contains the total number of pageview counts of 128 different Wikipedia pages related to the Zika virus (pageview_count_Zika) originated in 788 cities (US_city) of the United States with a population larger than 40,000 in 2016.The dataset also contains the total number of pageview counts to all Wikipedia pages (all Wikipedia projects, pageview_count_total) originated in 788 cities (US_city) of the United States with a population larger than 40,000 in 2016."</p>
SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles
<p>This dataset contains the files and annotations for <a href="https://propaganda.qcri.org/semeval2020-task11/index.html">SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles</a>. The task was composed by two subtasks: span identification (SI) and technique classification (TC). This dataset includes the following:</p> <ul> <li>The text files for training, development, and testing sets for both the SI and the TC tasks.</li> <li>The gold-standard files for the training sets, for both the SI and the TC task</li> </ul> <p>Our propaganda identification initiative remains active. We keep a <a href="https://propaganda.qcri.org/ptc/leaderboard.php">live leader-board</a> reporting the performance of models submited up to date.</p> <p><strong>Reference</strong></p> <p>Giovanni Da San Martino, Alberto Barrón-Cedeño, Henning Wachsmuth, Rostislav Petrov, and Preslav Nakov. 2020. <a href="https://propaganda.qcri.org/">Task 11: Detection of Propaganda Techniques in News Articles</a>. In Proceedings of the 14th International Workshop on Semantic Evaluation (SemEval 2020). Barcelona, Spain (2020)</p> <p> </p> <pre><code>@InProceedings{SemEval20-11-DaSanMartino, author = "Da San Martino, Giovanni and Barr\'{o}n-Cede\~no, Alberto and Wachsmuth, Henning and Petrov, Rostislav and Nakov, Preslav", title = "{SemEval}-2020 Task 11: {D}etection of Propaganda Techniques in News Articles", pages = "", abstract = "We describe the outcome of the SemEval 2020 Task 11 on the detection of propaganda in news articles. We present two tasks. In the first task, systems are asked to identify specific text spans in a free text where propaganda is being applied. In the second task, systems are asked to identify the propaganda technique being applied in a text span. We describe the construction of the evaluation framework (dataset and evaluation metrics) as well as the approaches explored by the different participants. ", crossref = "SemEval20" }</code></pre> <p> </p>
A dataset of media releases (Twitter, News and Comments, Youtube, Facebook) form Poland related to COVID-19 for open research
<p>Social behavior has a fundamental impact on the dynamics of infectious diseases (such as COVID-19), challenging public health mitigation strategies and possibly the political consensus. The widespread use of the traditional and social media on the Internet provides us with an invaluable source of information on societal dynamics during pandemics. With this dataset, we aim to understand mechanisms of COVID-19 epidemic-related social behavior in Poland deploying methods of computational social science and digital epidemiology. We have collected and analyzed COVID-19 perception on the Polish language Internet during 15.01-31.07(06.08) and labeled data quantitatively (Twitter, Youtube, Articles) and qualitatively (Facebook, Articles and Comments of Article) in the Internet by infomediological approach.</p> <p>- manually labelled1,449 articles / Facebook posts from Lower Silesia (facebook_articles_lower_silesia.zip) and 111 texts from outside this region;</p> <p>-manually labelled 1000 most popular tweets (twits_annotated.xlsx) with cathegories is_fake (categorical and numeric) topic and sentiment; </p> <p>-extracted 57,306 representative articles (articles_till_06_08.zip) in Polish using Eventregitry.org tool in language Polish and topic "Coronavirus" in article body;</p> <p>- extracted 1,015,199 (tweets_till_31_07_users.zip and tweets_till_31_07_text.zip) and Tweets from #Koronawirus in language Polish using Twitter API.</p> <p>- collected 1,574 videos (youtube_comments_till_31_07.zip and youtube_movie.csv) with keyword: Koronawirus on YouTube and 247,575 comments on them using Google API;</p> <p>- We supplemented the media observations with an analysis of 244 social empirical studies till 25.05 on COVID-19 in Poland (empirical_social_studies.csv).</p> <p>Reports and analyzes and coding books can be found in Polish at: <a href="http://www.infodemia-koronawirusa.pl/percepcja-koronowirusa-na-dolnym-slasku/">http://www.infodemia-koronawirusa.pl</a></p> <p>Main report (in Polish) https://depot.ceon.pl/handle/123456789/19215 </p>
Day 026: News Paper Box
3D Scan of a news paper box in Soho, Manhattan. It was tricky to scan the back as it was but against a traffic signal. And scanning all four sides would require standing on a busy traffic lane. Scanned with Polycam Source: Objaverse 1.0 / Sketchfab
News about Andalusian universities in Google News
In this dataset we show the total account of news for each university of Andalusia in Google News from 2011.
Online news and scientific production of Andalusian universities for 2011
Excel file with online News about the Andalusian universities retrieved from Google News during 2011, the scientific production of those universities and the comparison of both datasets.
Hacker News lda2vec pretrained word vectors
<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900 and https://zenodo.org/record/49902</p>
Hacker News lda2vec model pretrained word vectors
<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900</p>
Hindi News Article Text Dataset
<p>The Hindi News Article Dataset (HNAD) comprises over a million meticulously curated news articles in the Hindi language, sourced from diverse online platforms. Covering a wide range of topics, including politics, economics, culture, sports, and more, it offers a comprehensive representation of the Hindi news landscape. </p>
GPT annotated news dataset in IPTC news taxonomy
<p>The dataset consists of 4,672 news articles covering 17 first-level categories and 51 second-level categories. Each News item is annotated in IPTC news taxonomy using Gpt3.5 Turbo model in zero-shot setting. The prompting method can be found in our paper "<a href="https://ieeexplore.ieee.org/abstract/document/10367969">Evaluating the Effectiveness of GPT Large Language Model for News Classification in the IPTC News Ontology</a>"</p>
NLP and machine learning to measure peace from news media
<p>"Hate speech" can mobilize violence and destruction. What are the characteristics of "peace speech" that reflect and support the social processes that maintain peace? In this study we used a data driven, machine learning approach to identify the words most associated with lower-peace versus higher-peace countries. Logistic regression and random forest classifiers were trained using five respected, traditional peace indices: Global Peace Index, Positive Peace Index, World Happiness Index, Fragile States Index, and Human Development Index. The feature inputs into the machine learning model were the word frequencies from the news media in each country and the output classifications were the level of peace in that country. The machine learning model was successful in properly classifying the level of peace from the news media in a country (both accuracy and F1: 96% - 100%). We also used that trained machine model to create a machine learning peace index that measured the level of peace in countries, including countries not in the training set, which correlated with the average of those five traditional peace indices (r-squared = 0.8349). Using the random forest feature importance method we found that the words in news media in lower-peace countries were characterized by words related to government, order, control and fear (such as government, state, law, security and court), while higher-peace countries were characterized by an increased prevalence of words related to optimism for the future and fun (such as time, like, home, believe and game).</p>
Conservative News Media and Criminal Justice: Evidence from Exposure to Fox News Channel
<p>Conservative News Media and Criminal Justice: Evidence from Exposure to Fox News Channel" (Elliott Ash and Michael Poyker), <i>Economic Journal</i>, 2023</p>
Dataset Indonesian News
<p>Dataset for paper "Identifikasi Topik Hangat Di Media Berita Menggunakan Pendekatan Natural Language Processing". This dataset include news data from Antara, CNBC, CNN Indonesia JPNN, Kumparan, Merdeka, Okezone, Republika, Sindonews, Suara, Tempo, and Tribun. This dataset retrived at 16 september 2023</p>
Rare Diseases hand-annotated news articles: Angelman, De Lange, Fragile X, Kleefstra
<p>This dataset was produced in 2023 from the data collected throughout 2023 from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p> <p>The data is distributed across 4 specific diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p> <p>The available data comes in the file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings</p> <p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdolšek Draksler (domain expert) </p>
COVID-19 Fake News Detection Dataset
<p>Please note that this data set was originally shared by Patwa et al. (2021) on GitHub. </p> <p><strong>Reference </strong></p> <p>Patwa, P., Sharma, S. Pykl, S., Guptha, V., Kumari, G., Akhtar, M. S., Ekbal, A., Das A. & Chakraborty, T. (2021). Fighting an Infodemic: COVID-19 Fake News Dataset. Combating Online Hostile Posts in Regional Languages during Emergency Situation, Cham, Springer International Publishing. https://doi.org/10.1007/978-3-030-73696-5_3</p>
Abstractive News Captions with High- level cOntext Representation (ANCHOR) dataset
<p><span>The</span> <span>Abstractive News Captions with High-</span><span>level cOntext Representation</span> <span>(</span><span>ANCHOR</span><span>) dataset contains 70K+ samples </span><span>sourced from 5 different news media organizations. This dataset can be utilized for Vision & Language tasks such as Text-to-Image Generation, Image Caption Generation, etc.</span></p>
Documents used in the PLANET4B D1.1 analysis of biodiversity discourse by news outlets - 2010 and 2022 Data
<p>Data used to analyse biodiversity discourse in news outlet as part of Deliverable D1.1. of the Planet4B Project.</p>
Data for manuscript "Reciprocal Radicalization: The Rise of Culture War Terminology in British and American News Coverage"
<p>This data set contains frequency counts of target words in 16 million news and opinion articles from 10 popular news media outlets in the United Kingdom: The Guardian, The Times, The Independent, The Daily Mirror, BBC, Financial Times, Metro, Telegraph, The and The Daily Mail plus a few additional American-based outlets used for comparison reference. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice, social justice related terms or counterreaction to it. A few additional words are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 3 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We derived relative frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>-targetWordsInArticlesCountsGuardianExampleWords contains counts of target words in outlets articles as well as total counts of words in articles for illustrative Figure 1 in main manuscript</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. To conclude, in a data analysis of 16 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 of main manuscript for supporting evidence).</p>
Japanese Fake News Dataset
<p>Japanese Fake News Dataset</p> <p>Read our project page for more details: https://hkefka385.github.io/dataset/fakenews-japanese/</p>
Dataset: "I Can't Keep It Up." A Dataset from the Defunct Voat.co News Aggregator
<p>This is the dataset released with the <a href="https://arxiv.org/abs/2201.05933">paper </a>titled: "I Can’t Keep It Up." A Dataset from the Defunct Voat.co News Aggregator. <br> The dataset consists of 15,133 <a href="http://ndjson.org/">Newline delimited JSON</a> files (ndjson). More specifically, 7,616 files for submission data, 7,515 for comment data, 1 for user data, and 1 for subverse data. Each line in the ndjson files consists of a JSON object. The JSON objects contain all the key/values we collect through the Voat API and the custom parser of the Internet Archive Wayback Machine Voat snapshot release.<br> For the detailed description of every <em>key </em>in the JSON structure, along with the type of the <em>value</em>, please read the readme.pdf file provided with this dataset.</p> <p> </p> <p>If you find our dataset useful, please cite our paper:</p> <blockquote> <pre>@inproceedings{mekacher2022can, title={"I Can't Keep It Up." A Dataset from the Defunct Voat.co News Aggregator}, author={Mekacher, Amin and Papasavva, Antonis}, booktitle={16th International Conference on Web and Social Media}, year={2022} }</pre> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.