Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
Multilingual news article similarity dataset
<p>This dataset contains the extended version of the authors' earlier work: <a href="../records/6507872">https://zenodo.org/records/6507872,</a> where pairs of news articles drawn from the first half of 2020 are annotated for seven aspects of similarity in the original version as well as an additional FRAME aspect:</p> <ul> <li><strong>GEO</strong>: How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong> How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong> Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong> How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong> Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong> Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> <li><strong>FRAME</strong> Do the articles have similar framing and express similar opinions?</li> </ul>
Dataset: The Role of News Consumption on Influencers' Facebook Pages in Threat Perception and Political Conservatism During Times of COVID-19: A Comparative Study between the USA, Spain, and Egypt
<p>Este archivo ofrece los datos en bruto de una encuesta examina el impacto del consumo de noticias en las páginas de Facebook de los influencers en la motivación del conservadurismo político durante amenazas como el terrorismo o las pandemias. Muestra: N=1309, jóvenes de entre 18 y 35 años en Estados Unidos, España y Egipto. Trabajo de campo realizado entre el 10 de agosto de 2021 y el 5 de septiembre de 2021.</p> <p><span>Dataset correspondiente al proyecto El rol de la ciudadanía en la comunicación política digital CI-COMPOL (PID2020-119492GB-I00) financiado por MCIN/AEI/10.13039/501100011033/. IP: Andreu Casero-Ripollés, Departamento de Ciencias de la Comunicación, Universitat Jaume I de Castellón</span></p>
SemEval-2022 Task 8: Multilingual news article similarity
<p>This dataset contains pairs of news articles drawn from the first half of 2020 and annotated for seven aspects of similarity:</p> <ul> <li><strong>GEO</strong>: How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong> How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong> Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong> How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong> Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong> Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> </ul> <p>Further details are provided in</p> <blockquote> <p>Chen et al. (2022). SemEval-2022 Task 8: Multilingual news article similarity. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022). <a href="https://aclanthology.org/2022.semeval-1.155/">https://aclanthology.org/2022.semeval-1.155/</a></p> </blockquote> <p>The data in this repository includes pairs of URLs and annotations. The text of webpages is generally via the Internet Archive in this special collection: https://archive.org/details/2020-multilingual-news-article-similarity . A script to download and process the webpages is available at https://github.com/euagendas/semeval_8_2022_ia_downloader . </p>
Twitter Crawling for "Viral or Heboh" News
<p>The data we take is about tweet form 10 official account of news portal on twitter that mentioned word "Viral" or "Heboh". The dataset contain 5 columns and more than 500 tweet sorted by time.</p>
South African News Data
<p>This repository hosts a valuable collection of South African local language news compiled and enriched by the University of Pretoria's Data Science for Social Impact research group - https://dsfsi.github.io/. Spanning across South Africa's official languages, these localised news datasets aim to support natural language processing and social computing research focused on domestic challenges.</p><h2>Related Publication</h2><p><strong>Please cite</strong></p><p>@inproceedings{marivate2020investigating, title = {Investigating an Approach for Low Resource Language Dataset Creation, Curation and Classification: Setswana and Sepedi}, author = {Marivate, Vukosi and Sefara, Tshephisho and Chabalala, Vongani and Makhaya, Keamogetswe and Mokgonyane, Tumisho and Mokoena, Rethabile and Modupe, Abiodun}, booktitle = {Proceedings of the first workshop on Resources for African Indigenous Languages}, year = {2020}, address = {Marseille, France}, publisher = {European Language Resources Association (ELRA)}, url = {https://aclanthology.org/2020.rail-1.3}, pages = {15--20}, language = {English}, ISBN = {979-10-95546-60-3}, preprint_url ={https://arxiv.org/abs/2003.04986}, dataset_url = {https://zenodo.org/record/3668495}, keywords = {NLP} }</p><h2>Datasets</h2><h3><a href="http://www.sabc.co.za/">SABC </a></h3><p>We claim no copyright of the SABC original content. </p><ul><li>Motsweding FM (An SABC Setswana radio station) Facebook Page'</li><li>Dikgang Tsa Setswana (SABC Setswana News)</li><li>Thobela FM (An SABC Sepedi radio station) Facebook Page</li><li>Ditaba Tsa Sepedi (SABC Sepedi News)</li></ul><h3>Disclaimer</h3><p>This dataset contains extracted news content from a different sources including, but not limited to: SABC. While efforts were made to ensure the accuracy and completeness of this data, there may be errors or discrepancies between the original publications and this dataset. No warranties, guarantees or representations are given in relation to the information contained in the dataset. The members of the Data Science for Societal Impact Research Group bear no responsibility and/or liability for any such errors or discrepancies in this dataset. The original owners bear no responsibility and/or liability for any such errors or discrepancies in this dataset. It is recommended that users verify all information contained herein before making decisions based upon this information.</p><h2> </h2>
llustrated London News Illustration Dataset (1842-1890)
<p><strong>Description</strong></p> <p>This dataset contains comprehensive metadata for 72,081 illustrations extracted from the Illustrated London News (ILN) between 1842-1890. The ILN was the first and most influential illustrated newspaper of the Victorian era, making this dataset a valuable resource for researchers in digital humanities, media history, and visual culture studies. The dataset provides detailed information about each illustration, enabling large-scale analysis of Victorian visual culture and the evolution of newspaper illustration practices. </p> <p><strong>Content</strong></p> <p>The dataset consists of a CSV file containing the following information for each illustration:</p> <ul> <li>Publication date (YYYY-MM-DD format)</li> <li>Volume and issue number</li> <li>Page number within issue</li> <li>Bounding box coordinates (in YOLO format)</li> <li>Model confidence score from the detection model</li> <li>llustration sequence number on page (indicating reading order)</li> <li>OCR-extracted caption text</li> <li>Original Internet Archive item identifier</li> <li>Page URL for accessing the original scan</li> </ul> <p>This dataset also contains a pt file with multimodal embeddings (Open-CLIP) for all the illustrations</p> <p><strong>Methods</strong></p> <p>The illustrations were systematically extracted using several computational steps:</p> <ol> <li>Collection of 56,699 digitized pages from the Internet Archive's Serials in Microfilm Collection</li> <li>Fine-tuning of YOLOv8 object detection model on 908 manually annotated pages (mAP50: 0.964, mAP95: 0.92)</li> <li>Automated extraction of illustrations using the fine-tuned model</li> <li>Caption text extraction using Tesseract OCR</li> <li>Generation of multimodal embeddings using LAION OpenCLIP model (ViT-L-14-DataComp.XL-s13B-b90K</li> </ol> <p><strong>Code Availability</strong></p> <p>All code used to create this dataset is available in two GitHub repositories:</p> <p>Repository: https://github.com/tpsmi/multimodaliln </p> <ul> <li>Jupyter notebooks for downloading ILN pages</li> <li>YOLOv8 fine-tuning code</li> <li>Illustration extraction pipeline</li> <li>OCR processing scripts</li> <li>Embedding generation code</li> </ul> <p>Repository: https://github.com/tpsmi/ilnmultimodalsearch </p> <ul> <li>Multimodal search implementation</li> <li>Text-to-image and image-to-image retrieval</li> <li>User interface code</li> <li>API endpoints for search functionality</li> </ul> <p><strong>Original Data Source</strong></p> <p>The original page scans are freely available through the Internet Archive's Serials in Microfilm Collection. This dataset builds upon these public domain materials by providing structured metadata and computational annotations.</p> <p><strong>Citation</strong></p> <p>Please cite our dataset paper (will be added) or this dataset (Zenodo DOI)</p> <p><strong>Related Publications</strong></p> <p>[Publication details when available]</p>
BuzzFeed-Webis Fake News Corpus 2016
<p>The corpus comprises the output of 9 publishers in a week close to the US elections. Among the selected publishers are 6 prolific hyperpartisan ones (three left-wing and three right-wing), and three mainstream publishers (see Table 1). All publishers earned Facebook’s blue checkmark, indicating authenticity and an elevated status within the network. For seven weekdays (September 19 to 23 and September 26 and 27), every post and linked news article of the 9 publishers was fact-checked by professional journalists at BuzzFeed. In total, 1,627 articles were checked, 826 mainstream, 256 left-wing and 545 right-wing. The imbalance between categories results from differing publication frequencies.</p>
Croatian Coronavirus News Comments Corpus News-CommHR
<p>A corpus of readers' news comments posted below news articles on the topic of the covid-19 pandemic, published in major Croatian daily newspapers and news portals in the six-month early pandemic period (March 2020 to September 2020).</p> <div>The corpus is designed to facilitate research on crisis discourses, crisis communication, as well as pandemic-time linguistic innovation. It is available in plain text version and XML with full metadata. The corpus complements a separate corpus of news articles Croatian Coronavirus Corpus NewsHR. Parallel versions from Slovenia and Serbia are also available.</div> <div> </div> <div>The project leading to this publication has received funding from the European Union’s Horizon 2020 research and innovation programme under the <a href="https://cordis.europa.eu/programme/id/H2020-EU.4./en">H2020-EU.4. - SPREADING EXCELLENCE AND WIDENING PARTICIPATION </a>programme Widening fellowships grant agreement No 101038047.</div>
Supplementary File: Entertainment interspersed with propaganda: How non-legacy-news accounts deliver explicitly political content to mass audiences on Russia's most popular social network VK
<p>Supplementary file and dataset for the paper "Entertainment interspersed with propaganda: How non-legacy-news accounts deliver explicitly political content to mass audiences on Russia’s most popular social network VK"</p>
Swahili : News Classification Dataset
<p>Swahili is spoken by 100-150 million people across East Africa. In Tanzania, it is one of two national languages (the other is English) and it is the official language of instruction in all schools. News in Swahili is an important part of the media sphere in Tanzania.</p> <p>News contributes to education, technology, and the economic growth of a country, and news in local languages plays an important cultural role in many Africa countries. In the modern age, African languages in news and other spheres are at risk of being lost as English becomes the dominant language in online spaces.<br> <br> The Swahili news dataset was created to reduce the gap of using the Swahili language to create NLP technologies and help AI practitioners in Tanzania and across Africa continent to practice their NLP skills to solve different problems in organizations or societies related to Swahili language. Swahili News were collected from different websites that provide news in the Swahili language. I was able to find some websites that provide news in Swahili only and others in different languages including Swahili.<br> <br> The dataset was created for a specific task of text classification, this means each news content can be categorized into six different topics (Local news, International news , Finance news, Health news, Sports news, and Entertainment news). The dataset comes with a specified train/test split. The train set contains 75% of the dataset and test set contains 25% of the dataset.</p> <p><strong>Acknowledgment</strong>: This project was supported by the <a href="https://www.k4all.org/project/language-dataset-fellowship/">AI4D language dataset fellowship</a> through K4All and <a href="https://zindi.africa/">Zindi Africa</a>.</p>
INTRODUCTION OF COVID-NEWS-US-NNK AND COVID-NEWS-BD-NNK DATASET
<p>Introduction</p> <p>There are several works based on Natural Language Processing on newspaper reports. Mining opinions from headlines [ 1 ] using Standford NLP and SVM by Rameshbhaiet. Al.compared several algorithms on a small and large dataset. Rubinet. al., in their paper [ 2 ], created a mechanism to differentiate fake news from real ones by building a set of characteristics of news according to their types. The purpose was to contribute to the low resource data available for training machine learning algorithms. Doumitet. al.in [ 3 ] have implemented LDA, a topic modeling approach to study bias present in online news media.</p> <p>However, there are not many NLP research invested in studying COVID-19. Most applications include classification of chest X-rays and CT-scans to detect presence of pneumonia in lungs [ 4 ], a consequence of the virus. Other research areas include studying the genome sequence of the virus[ 5 ][ 6 ][ 7 ] and replicating its structure to fight and find a vaccine. This research is crucial in battling the pandemic. The few NLP based research publications are sentiment classification of online tweets by Samuel et el [ 8 ] to understand fear persisting in people due to the virus. Similar work has been done using the LSTM network to classify sentiments from online discussion forums by Jelodaret. al.[ 9 ]. NKK dataset is the first study on a comparatively larger dataset of a newspaper report on COVID-19, which contributed to the virus’s awareness to the best of our knowledge.</p> <p> </p> <p>2 Data-set Introduction</p> <p>2.1 Data Collection</p> <p>We accumulated 1000 online newspaper report from United States of America (USA) on COVID-19. The newspaper includes The Washington Post (USA) and StarTribune (USA). We have named it as “Covid-News-USA-NNK”. We also accumulated 50 online newspaper report from Bangladesh on the issue and named it “Covid-News-BD-NNK”. The newspaper includes The Daily Star (BD) and Prothom Alo (BD). All these newspapers are from the top provider and top read in the respective countries. The collection was done manually by 10 human data-collectors of age group 23- with university degrees. This approach was suitable compared to automation to ensure the news were highly relevant to the subject. The newspaper online sites had dynamic content with advertisements in no particular order. Therefore there were high chances of online scrappers to collect inaccurate news reports. One of the challenges while collecting the data is the requirement of subscription. Each newspaper required $1 per subscriptions. Some criteria in collecting the news reports provided as guideline to the human data-collectors were as follows:</p> <ul> <li>The headline must have one or more words directly or indirectly related to COVID-19.</li> <li>The content of each news must have 5 or more keywords directly or indirectly related to COVID-19.</li> <li>The genre of the news can be anything as long as it is relevant to the topic. Political, social, economical genres are to be more prioritized.</li> <li>Avoid taking duplicate reports.</li> <li>Maintain a time frame for the above mentioned newspapers.</li> </ul> <p>To collect these data we used a google form for USA and BD. We have two human editor to go through each entry to check any spam or troll entry.</p> <p>2.2 Data Pre-processing and Statistics</p> <p>Some pre-processing steps performed on the newspaper report dataset are as follows:</p> <ul> <li>Remove hyperlinks.</li> <li>Remove non-English alphanumeric characters.</li> <li>Remove stop words.</li> <li>Lemmatize text.</li> </ul> <p>While more pre-processing could have been applied, we tried to keep the data as much unchanged as possible since changing sentence structures could result us in valuable information loss. While this was done with help of a script, we also assigned same human collectors to cross check for any presence of the above mentioned criteria.</p> <p>The primary data statistics of the two dataset are shown in Table 1 and 2.</p> <pre><code>Table 1: Covid-News-USA-NNK data statistics </code></pre> <pre><code>No of words per headline </code></pre> <pre><code>7 to 20 </code></pre> <pre><code>No of words per body content </code></pre> <pre><code>150 to 2100 </code></pre> <pre><code>Table 2: Covid-News-BD-NNK data statistics No of words per headline </code></pre> <pre><code>10 to 20 </code></pre> <pre><code>No of words per body content </code></pre> <pre><code>100 to 1500 </code></pre> <p>2.3 Dataset Repository</p> <p>We used GitHub as our primary data repository in account name NKK^1. Here, we created two repositories USA-NKK^2 and BD-NNK^3. The dataset is available in both CSV and JSON format. We are regularly updating the CSV files and regenerating JSON using a py script. We provided a python script file for essential operation. We welcome all outside collaboration to enrich the dataset.</p> <p> </p> <p> </p> <p>3 Literature Review</p> <p>Natural Language Processing (NLP) deals with text (also known as categorical) data in computer science, utilizing numerous diverse methods like one-hot encoding, word embedding, etc., that transform text to machine language, which can be fed to multiple machine learning and deep learning algorithms.</p> <p>Some well-known applications of NLP includes fraud detection on online media sites[ 10 ], using authorship attribution in fallback authentication systems[ 11 ], intelligent conversational agents or chatbots[ 12 ] and machine translations used by Google Translate[ 13 ]. While these are all downstream tasks, several exciting developments have been made in the algorithm solely for Natural Language Processing tasks. The two most trending ones are BERT[ 14 ], which uses bidirectional encoder-decoder architecture to create the transformer model, that can do near-perfect classification tasks and next-word predictions for next generations, and GPT-3 models released by OpenAI[ 15 ] that can generate texts almost human-like. However, these are all pre-trained models since they carry huge computation cost. Information Extraction is a generalized concept of retrieving information from a dataset. Information extraction from an image could be retrieving vital feature spaces or targeted portions of an image; information extraction from speech could be retrieving information about names, places, etc[ 16 ]. Information extraction in texts could be identifying named entities and locations or essential data. Topic modeling is a sub-task of NLP and also a process of information extraction. It clusters words and phrases of the same context together into groups. Topic modeling is an unsupervised learning method that gives us a brief idea about a set of text. One commonly used topic modeling is Latent Dirichlet Allocation or LDA[17].</p> <p>Keyword extraction is a process of information extraction and sub-task of NLP to extract essential words and phrases from a text. TextRank [ 18 ] is an efficient keyword extraction technique that uses graphs to calculate the weight of each word and pick the words with more weight to it.</p> <p>Word clouds are a great visualization technique to understand the overall ’talk of the topic’. The clustered words give us a quick understanding of the content.</p> <p> </p> <p> </p> <p>4 Our experiments and Result analysis</p> <p>We used the wordcloud library^4 to create the word clouds. Figure 1 and 3 presents the word cloud of Covid-News-USA- NNK dataset by month from February to May. From the figures 1,2,3, we can point few information:</p> <ul> <li>In February, both the news paper have talked about China and source of the outbreak.</li> <li>StarTribune emphasized on Minnesota as the most concerned state. In April, it seemed to have been concerned more.</li> <li>Both the newspaper talked about the virus impacting the economy, i.e, bank, elections, administrations, markets.</li> <li>Washington Post discussed global issues more than StarTribune.</li> <li>StarTribune in February mentioned the first precautionary measurement: wearing masks, and the uncontrollable spread of the virus throughout the nation.</li> <li>While both the newspaper mentioned the outbreak in China in February, the weight of the spread in the United States are more highlighted through out March till May, displaying the critical impact caused by the virus.</li> </ul> <p>We used a script to extract all numbers related to certain keywords like ’Deaths’, ’Infected’, ’Died’ , ’Infections’, ’Quarantined’, Lock-down’, ’Diagnosed’ etc from the news reports and created a number of cases for both the newspaper. Figure 4 shows the statistics of this series. From this extraction technique, we can observe that April was the peak month for the covid cases as it gradually rose from February. Both the newspaper clearly shows us that the rise in covid cases from February to March was slower than the rise from March to April. This is an important indicator of possible recklessness in preparations to battle the virus. However, the steep fall from April to May also shows the positive response against the attack. We used Vader Sentiment Analysis to extract sentiment of the headlines and the body. On average, the sentiments were from -0.5 to -0.9. Vader Sentiment scale ranges from -1(highly negative to 1(highly positive). There were some cases</p> <p>where the sentiment scores of the headline and body contradicted each other,i.e., the sentiment of the headline was negative but the sentiment of the body was slightly positive. Overall, sentiment analysis can assist us sort the most concerning (most negative) news from the positive ones, from which we can learn more about the indicators related to COVID-19 and the serious impact caused by it. Moreover, sentiment analysis can also provide us information about how a state or country is reacting to the pandemic. We used PageRank algorithm to extract keywords from headlines as well as the body content. PageRank efficiently highlights important relevant keywords in the text. Some frequently occurring important keywords extracted from both the datasets are: ’China’, Government’, ’Masks’, ’Economy’, ’Crisis’, ’Theft’ , ’Stock market’ , ’Jobs’ , ’Election’, ’Missteps’, ’Health’, ’Response’. Keywords extraction acts as a filter allowing quick searches for indicators in case of locating situations of the economy, how states are defending against the pandemic, the condition of the health care system etc.</p> <p> </p> <p>5 Conclusion</p> <p>This dataset can demonstrate how news reports could speculate the situation differently based on the news source. The different types of experiments are possible to assert the importance of Natural Language Processing in newspaper report analysis. We are looking for more collaborators in GitHub to enrich the dataset, which will make it possible to run extensive deep learning experiments.</p> <p> </p> <p>References</p> <p>[1] Chaudhary Jashubhai Rameshbhai and Joy Paulose. Opinion mining on newspaper headlines using svm and nlp. International Journal of Electrical & Computer Engineering (2088-8708), 9(3), 2019. [2]Victoria L Rubin, Yimin Chen, and Nadia K Conroy. Deception detection for news: three types of fakes. Proceedings of the Association for Information Science and Technology, 52(1):1–4, 2015. [3]Sarjoun Doumit and Ali Minai. Online news media bias analysis using an lda-nlp approach. InInternational Conference on Complex Systems, 2011. [4]Md Manjurul Ahsan, Kishor Datta Gupta, Mohammad Maminur Islam, Sajib Sen, Md Rahman, Moham- mad Shakhawat Hossain, et al. Study of different deep learning approach with explainable ai for screening patients with covid-19 symptoms: Using ct scan and chest x-ray image dataset.arXiv preprint arXiv:2007.12525, 2020. [5]Gurjit S Randhawa, Maximillian PM Soltysiak, Hadi El Roz, Camila PE de Souza, Kathleen A Hill, and Lila Kari. Machine learning using intrinsic genomic signatures for rapid classification of novel pathogens: Covid-19 case study.Plos one, 15(4):e0232391, 2020.</p> <p>[6]Ahmad Alimadadi, Sachin Aryal, Ishan Manandhar, Patricia B Munroe, Bina Joe, and Xi Cheng. Artificial intelligence and machine learning to fight covid-19, 2020. [7]Shreshth Tuli, Shikhar Tuli, Rakesh Tuli, and Sukhpal Singh Gill. Predicting the growth and trend of covid- pandemic using machine learning and cloud computing.Internet of Things, page 100222, 2020. [8]Jim Samuel, GG Ali, Md Rahman, Ek Esawi, Yana Samuel, et al. Covid-19 public sentiment insights and machine learning for tweets classification.Information, 11(6):314, 2020. [9] Hamed Jelodar, Yongli Wang, Rita Orji, and Hucheng Huang. Deep sentiment classification and topic discovery on novel coronavirus or covid-19 online discussions: Nlp using lstm recurrent neural network approach.arXiv preprint arXiv:2004.11695, 2020.</p> <p>[10]Nafiz Sadman, Kishor Datta Gupta, Ariful Haque, Subash Poudyal, and Sajib Sen. Detect review manipulation by leveraging reviewer historical stylometrics in amazon, yelp, facebook and google reviews. InProceedings of the 2020 The 6th International Conference on E-Business and Applications, pages 42–47, 2020.</p> <p>[11]Nafiz Sadman, Kishor Datta Gupta, Ariful Haque, Subash Poudyal, and Sajib Sen. Stylometry as a reliable method for fallback authentication. InProceedings of the 2020 17th International Conference on Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology, 2020.</p> <p>[12]Ethan Fast, Binbin Chen, Julia Mendelsohn, Jonathan Bassen, and Michael S Bernstein. Iris: A conversational agent for complex tasks. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pages 1–12, 2018.</p> <p>[13] Philipp Koehn.Statistical machine translation. Cambridge University Press, 2009.</p> <p>[14]Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805, 2018.</p> <p>[15]Will Douglas Heavenarchive page. Openai’s new language generator gpt-3 is shockingly good—and completely mindless <a href="https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-">https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-</a> gpt-3-nlp/. Technical report.</p> <p>[16]Chin-Hui Lee and Sabato Marco Siniscalchi. An information-extraction approach to speech processing: Analysis, detection, verification, and recognition.Proceedings of the IEEE, 101(5):1089–1115, 2013.</p> <p>[17]Hamed Jelodar, Yongli Wang, Chi Yuan, Xia Feng, Xiahui Jiang, Yanchao Li, and Liang Zhao. Latent dirich- let allocation (lda) and topic modeling: models, applications, a survey. Multimedia Tools and Applications, 78(11):15169–15211, 2019.</p> <p>[18]Monica Bianchini, Marco Gori, and Franco Scarselli. Inside pagerank.ACM Transactions on Internet Technology (TOIT), 5(1):92–128, 2005.</p>
Meneame news dataset
<p>This dataset contains a list of news published in Meneame.net between 2020-10-17 and 2020-10-20.</p>
Mapping Ottoman Damascus Through News Reports: A Practical Approach,
<p>This release corresponds to the publication of the book chapter "Mapping Ottoman Damascus Through News Reports: A Practical Approach," in <em>Digital Humanities and Islamic & Middle East Studies</em>, ed. Elias Muhanna (Boston, Berlin: De Gruyter, 2015), pp. 175-198.</p>
Navigating News Narratives: A Media Bias Analysis Dataset
<p>The prevalence of bias in the news media has become a critical issue, affecting public perception on a range of important topics such as political views, health, insurance, resource distributions, religion, race, age, gender, occupation, and climate change. The media has a moral responsibility to ensure accurate information dissemination and to increase awareness about important issues and the potential risks associated with them. This highlights the need for a solution that can help mitigate against the spread of false or misleading information and restore public trust in the media.</p><p><strong>Data description: </strong>This is a dataset for news media bias covering different dimensions of the biases: political, hate speech, political, toxicity, sexism, ageism, gender identity, gender discrimination, race/ethnicity, climate change, occupation, spirituality, which makes it a unique contribution. The dataset used for this project does not contain any personally identifiable information (PII).</p><p><strong>Data Format: </strong>The format of data is:</p><ul><li>ID: Numeric unique identifier.</li><li>Text: Main content.</li><li>Dimension: Categorical descriptor of the text.</li><li>Biased_Words: List of words considered biased.</li><li>Aspect: Specific topic within the text.</li><li>Label: Neutral, Slightly Biased , Highly Biased</li></ul><p><br><strong>Annotation Scheme: </strong>The annotation scheme is based on Active learning, which is Manual Labeling --> Semi-Supervised Learning --> Human Verifications (iterative process)</p><ul><li>Bias Label: Indicate the presence/absence of bias (e.g., no bias, mild, strong).</li><li>Words/Phrases Level Biases: Identify specific biased words/phrases.</li><li>Subjective Bias (Aspect): Capture biases related to content aspects.</li></ul><p><br><strong>List of datasets used : </strong>We curated different news categories like Climate crisis news summaries , occupational, spiritual/faith/ general using RSS to capture different dimensions of the news media biases. The annotation is performed using active learning to label the sentence (either neural/ slightly biased/ highly biased) and to pick biased words from the news.</p><p>We also utilize publicly available data from the following links. Our Attribution to others.</p><p> <strong>MBIC (media bias): </strong>Spinde, Timo, Lada Rudnitckaia, Kanishka Sinha, Felix Hamborg, Bela Gipp, and Karsten Donnay. "MBIC--A Media Bias Annotation Dataset Including Annotator Characteristics." arXiv preprint arXiv:2105.11910 (2021). <a href="https://zenodo.org/records/4474336">https://zenodo.org/records/4474336</a> </p><p><strong>Hyperpartisan news: </strong>Kiesel, Johannes, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast. "Semeval-2019 task 4: Hyperpartisan news detection." In Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 829-839. 2019. <a href="https://huggingface.co/datasets/hyperpartisan_news_detection">https://huggingface.co/datasets/hyperpartisan_news_detection</a> </p><p><strong>Toxic comment classification: </strong>Adams, C.J., Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski. 2017. "Toxic Comment Classification Challenge." Kaggle. <a href="https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge">https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge</a>.</p><p><strong>Jigsaw Unintended Bias: </strong>Adams, C.J., Daniel Borkan, Inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and Nithum. 2019. "Jigsaw Unintended Bias in Toxicity Classification." Kaggle. <a href="https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification">https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification</a>.</p><p><strong>Age Bias : </strong>Díaz, Mark, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. "Addressing age-related bias in sentiment analysis." In Proceedings of the 2018 chi conference on human factors in computing systems, pp. 1-14. 2018. <a href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/F6EMTS">Age Bias Training and Testing Data - Age Bias and Sentiment Analysis Dataverse (harvard.edu)</a></p><p><strong>Multi-dimensional news Ukraine: </strong>Färber, Michael, Victoria Burkard, Adam Jatowt, and Sora Lim. "A multidimensional dataset based on crowdsourcing for analyzing and detecting news bias." In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pp. 3007-3014. 2020. <a href="https://zenodo.org/records/3885351#.ZF0KoxHMLtV">https://zenodo.org/records/3885351#.ZF0KoxHMLtV</a> </p><p><strong>Social biases: </strong>Sap, Maarten, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. "Social bias frames: Reasoning about social and power implications of language." arXiv preprint arXiv:1911.03891 (2019). <a href="https://maartensap.com/social-bias-frames/">https://maartensap.com/social-bias-frames/</a> </p><p> </p><p><strong>Goal of this dataset :</strong>We want to offer open and free access to dataset, ensuring a wide reach to researchers and AI practitioners across the world. The dataset should be user-friendly to use and uploading and accessing data should be straightforward, to facilitate usage.</p><p><strong>If you use this dataset, please cite us.</strong></p><p>Navigating News Narratives: A Media Bias Analysis Dataset © 2023 by <a href="https://www.linkedin.com/in/shainaraza/">Shaina Raza, Vector Institute </a>is licensed under <a href="http://creativecommons.org/licenses/by-nc/4.0/?ref=chooser-v1">CC BY-NC 4.0 </a></p><p> </p>
Cyberhate that targets people who are plus-size in the news: The role of bystanders in mitigating social pathologies (CYBERPLUS)
<p>The dataset was created for the project "Cyberhate that targets people who are plus-size in the news: The role of bystanders in mitigating social pathologies (CYBERPLUS)". The data was collected between July 12 and July 26, 2024, from 1,030 young Czech people aged 16-25. The survey asked young people about their sociodemographic information, attitudes toward and perceptions of entitativity of three groups (overweight people, underweight people, people with physical disabilities), group identification, bystander appraisals and behavioural intentions, hate speech perception, and internet use. It included an experimental part in which the participants were exposed as bystanders to social media news posts about overweight people and comments under the posts. The dataset is accompanied by a data dictionary and a technical report.</p>
How do Google News' top 100 sources visually represent the data centres' energy footprint?
<p><strong>By querying "data centres' energy footprint" on Google News in incognito mode, the candidate has selected and mapped the top 100 results according to the ranking on May 15, 2022. </strong></p>
FA-KES: A Fake News Dataset around the Syrian War
<p>We have produced a labeled dataset that presents fake news surrounding the conflict in Syria. The dataset consists of a set of articles/news labeled by 0 (fake) or 1 (credible). Credibility of articles are computed with respect to a ground truth information obtained from the Syrian Violations Documentation Center (VDC). In particular, for each article, we crowdsource the information extraction (e.g., date, location, Number of casualties) job using the crowdsourcing platform Figure Eight (formally CrowdFlower). Then, we match those articles against the VDC database to be able to deduce whether an article is fake or not. The dataset can be used to train machine learning models to detect fake news. </p> <p> </p> <p> </p> <p> </p>
South African Disinformation [Fake News] Website Data - 2020
<p>See publication: <strong>Is it Fake? News Disinformation Detection on South African News Websites</strong></p> <p>We used, as sources, investigations by the news websites MyBroadband (<a href="https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/">https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/</a>) and News24 (<a href="https://exposed.news24.com/the-website-blacklist/">https://exposed.news24.com/the-website-blacklist/</a>). These articles covered investigations into disinformation websites in South Africa in 2018. They compiled lists of websites that were suspected to be disinformation. During the period from those articles to present, a number of the websites have become inaccessible or offline. We attempted to use the internet archives <a href="https://archive.org/web/">WayBack Machine</a> we could only get partial snapshots and error messages.</p> <p>A web-scraper only worked for one of the sources although manual editing was still required to clean the text from Javascript code and some paragraph duplicates. On most of the other websites, a web-scraper did not work well as there were too many advertisements and broken parts of pages. Because of all these problems, most of the articles were manually copied and pasted and cleaned in flat files. In some cases, the text of articles could not be copied and was not made part of the South African disinformation corpus.</p> <p><strong>Citing the dataset</strong></p> <blockquote> <p>@inproceedings{de2021fake, title={Is it Fake? News Disinformation Detection on South African News Websites}, author={de Wet, Harm and Marivate, Vukosi}, booktitle={2021 IEEE AFRICON}, pages={1--6}, year={2021}, organization={IEEE} }</p> </blockquote>
Sogou News Corpus (SOGOU)
<p>Sogou news corpus (SOGOU): It is a Chinese dataset from the combination of the SogouCA and SogouCS news corpora, containing 500K news articles in various topic channels. They were labeled by manually classifying their domain names. Five categories were defined: ‘‘sports’’, ‘‘finance’’, ‘‘entertainment’’, ‘‘automobile’’ and ‘‘technology’’. The models for English can be applied to this dataset without change. The fields used are title and content (Zhang et al., 2016).</p>
20 news group (20ng)
<p>20 Newsgroups (20NG) is a classical and popular dataset for experiments in text applications of machine learning techniques. It contains 18,846 newsgroups documents, partitioned almost evenly across 20 different newsgroups categories. </p> <p>http://qwone.com/~jason/20Newsgroups/</p> <p><strong>The files:</strong><br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p> <p><strong>Label Definition: (Score File)</strong></p> <p>0 atheist resources<br> 1 computer graphics<br> 2 computer os ms windows misc<br> 3 computer system ibm pc hardware<br> 4 computer system mac hardware<br> 5 computer windows x<br> 6 misc miscellaneous for sale<br> 7 rec autos<br> 8 rec motorcycles<br> 9 rec sport baseball<br> 10 rec sport hockey<br> 11 science crypt<br> 12 science electronics<br> 13 science med<br> 14 science space<br> 15 society religion christian<br> 16 talk politics guns<br> 17 talk politics mideast<br> 18 talk politics misc miscellaneous<br> 19 talk religion misc miscellaneous</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.