Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “media bias”

Learn how ShareScore rates datasets ↗
zenodo44/100

Navigating News Narratives: A Media Bias Analysis Dataset

<p>The prevalence of bias in the news media has become a critical issue, affecting public perception on a range of important topics such as political views, health, insurance, resource distributions, religion, race, age, gender, occupation, and climate change. The media has a moral responsibility to ensure accurate information dissemination and to increase awareness about important issues and the potential risks associated with them. This highlights the need for a solution that can help mitigate against the spread of false or misleading information and restore public trust in the media.</p><p><strong>Data description: </strong>This is a dataset for news media bias covering different dimensions of the biases: political, hate speech, political, toxicity, sexism, ageism, gender identity, gender discrimination, race/ethnicity, climate change, occupation, spirituality, which makes it a unique contribution. The dataset used for this project does not contain any personally identifiable information (PII).</p><p><strong>Data Format: </strong>The format of data is:</p><ul><li>ID: Numeric unique identifier.</li><li>Text: Main content.</li><li>Dimension: Categorical descriptor of the text.</li><li>Biased_Words: List of words considered biased.</li><li>Aspect: Specific topic within the text.</li><li>Label: Neutral, Slightly Biased , Highly Biased</li></ul><p><br><strong>Annotation Scheme: </strong>The annotation scheme is based on Active learning, which is Manual Labeling --&gt; Semi-Supervised Learning --&gt; Human Verifications (iterative process)</p><ul><li>Bias Label: Indicate the presence/absence of bias (e.g., no bias, mild, strong).</li><li>Words/Phrases Level Biases: Identify specific biased words/phrases.</li><li>Subjective Bias (Aspect): Capture biases related to content aspects.</li></ul><p><br><strong>List of datasets used : </strong>We curated different news categories like Climate crisis news summaries , occupational, spiritual/faith/ general using RSS to capture different dimensions of the news media biases. The annotation is performed using active learning to label the sentence (either neural/ slightly biased/ highly biased) and to pick biased words from the news.</p><p>We also utilize publicly available data from the following links. Our Attribution to others.</p><p>&nbsp;<strong>MBIC (media bias): &nbsp;</strong>Spinde, Timo, Lada Rudnitckaia, Kanishka Sinha, Felix Hamborg, Bela Gipp, and Karsten Donnay. "MBIC--A Media Bias Annotation Dataset Including Annotator Characteristics." arXiv preprint arXiv:2105.11910 (2021).&nbsp;<a href="https://zenodo.org/records/4474336">https://zenodo.org/records/4474336</a>&nbsp;&nbsp;</p><p><strong>Hyperpartisan&nbsp; news: </strong>Kiesel, Johannes, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast. "Semeval-2019 task 4: Hyperpartisan news detection." In Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 829-839. 2019.&nbsp;<a href="https://huggingface.co/datasets/hyperpartisan_news_detection">https://huggingface.co/datasets/hyperpartisan_news_detection</a>&nbsp;</p><p><strong>Toxic comment classification: </strong>Adams, C.J., Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski. 2017. "Toxic Comment Classification Challenge." Kaggle.&nbsp;<a href="https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge">https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge</a>.</p><p><strong>Jigsaw Unintended Bias: </strong>Adams, C.J., Daniel Borkan, Inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and Nithum. 2019. "Jigsaw Unintended Bias in Toxicity Classification." Kaggle.&nbsp;<a href="https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification">https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification</a>.</p><p><strong>Age Bias : </strong>Díaz, Mark, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. "Addressing age-related bias in sentiment analysis." In Proceedings of the 2018 chi conference on human factors in computing systems, pp. 1-14. 2018.&nbsp;<a href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/F6EMTS">Age Bias Training and Testing Data - Age Bias and Sentiment Analysis Dataverse (harvard.edu)</a></p><p><strong>Multi-dimensional news Ukraine: </strong>Färber, Michael, Victoria Burkard, Adam Jatowt, and Sora Lim. "A multidimensional dataset based on crowdsourcing for analyzing and detecting news bias." In Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management, pp. 3007-3014. 2020.&nbsp;<a href="https://zenodo.org/records/3885351#.ZF0KoxHMLtV">https://zenodo.org/records/3885351#.ZF0KoxHMLtV</a>&nbsp;</p><p><strong>Social biases: </strong>Sap, Maarten, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. "Social bias frames: Reasoning about social and power implications of language." arXiv preprint arXiv:1911.03891 (2019).&nbsp;<a href="https://maartensap.com/social-bias-frames/">https://maartensap.com/social-bias-frames/</a>&nbsp;</p><p>&nbsp;</p><p><strong>Goal of this dataset :</strong>We want to offer open and free access to dataset, ensuring a wide reach to researchers and AI practitioners across the world. The dataset should be user-friendly to use and uploading and accessing data should be straightforward, to facilitate usage.</p><p><strong>If you use this dataset, please cite us.</strong></p><p>Navigating News Narratives: A Media Bias Analysis Dataset&nbsp;© 2023&nbsp;by&nbsp;<a href="https://www.linkedin.com/in/shainaraza/">Shaina Raza, Vector Institute&nbsp;</a>is licensed under&nbsp;<a href="http://creativecommons.org/licenses/by-nc/4.0/?ref=chooser-v1">CC BY-NC 4.0&nbsp;</a></p><p>&nbsp;</p>

opencc-by-nc-4.0Dec 2022View details →
zenodo44/100

Qbias – A Dataset on Media Bias in Search Queries and Query Suggestions

<p>We present Qbias, two novel datasets&nbsp;that promote the investigation of bias in online news search as described in</p> <blockquote> <p>Fabian Haak and Philipp Schaer. 2023. 𝑄𝑏𝑖𝑎𝑠 - A Dataset on Media Bias in Search Queries and Query Suggestions. In Proceedings of ACM Web Science Conference (WebSci&rsquo;23). ACM, New York, NY, USA, 6 pages.&nbsp;<a href="https://doi.org/10.1145/3578503.3583628">https://doi.org/10.1145/3578503.3583628</a>.</p> </blockquote> <p><strong>Dataset 1: AllSides Balanced News Dataset (allsides_balanced_news_headlines-texts.csv)</strong></p> <p>The dataset contains 21,747 news articles collected from <a href="https://www.allsides.com/headline-roundups">AllSides balanced news headline</a> roundups in November 2022 as presented in our publication. The AllSides balanced news feature three expert-selected U.S. news articles from sources of different political views (left, right, center), often featuring spin bias, and slant other forms of non-neutral reporting on political news. All articles are tagged with a bias label by four expert annotators based on the expressed political partisanship, left, right, or neutral. The AllSides balanced news aims to offer multiple political perspectives on important news stories, educate users on biases, and provide multiple viewpoints. Collected data further includes headlines, dates, news texts, topic tags (e.g., &quot;Republican party&quot;, &quot;coronavirus&quot;, &quot;federal jobs&quot;), and the publishing news outlet. We also include AllSides&#39; neutral description of the topic of the articles.<br> Overall, the dataset contains 10,273 articles tagged as left, 7,222 as right, and 4,252 as center.</p> <p>To provide easier access to the most recent and complete version of the dataset for future research, we provide a scraping tool and a regularly&nbsp;updated version of the dataset at <a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>. The repository also contains regularly updated more recent versions of the dataset with additional tags (such as the URL to the article). We chose to publish the version used for fine-tuning the models on Zenodo to enable the reproduction of the results of our study.&nbsp;</p> <p>&nbsp;</p> <p><strong>Dataset 2: Search Query Suggestions&nbsp;(suggestions.csv)</strong></p> <p>The second dataset we provide consists of 671,669 search query suggestions for root queries based on tags of the AllSides biased news dataset. We collected search query suggestions from Google and Bing for the 1,431 topic tags, that have been used for tagging AllSides news at least five times, approximately half of the total number of topics.&nbsp;The topic tags include names, a wide range of political terms, agendas, and topics (e.g., &quot;communism&quot;, &quot;libertarian party&quot;, &quot;same-sex marriage&quot;), cultural and religious terms (e.g., &quot;Ramadan&quot;, &quot;pope Francis&quot;), locations and other news-relevant terms.&nbsp;On average, the dataset contains 469 search queries for each topic.&nbsp;In total, 318,185 suggestions have been retrieved from Google and 353,484 from Bing.</p> <p>The file contains a &quot;root_term&quot; column based on the AllSides topic tags. The &quot;query_input&quot; column contains the search term submitted to the search engine (&quot;search_engine&quot;). &quot;query_suggestion&quot; and &quot;rank&quot; represents the search query suggestions at the respective positions returned by the search engines at the given time of search &quot;datetime&quot;. We scraped our data from a US server saved in &quot;location&quot;.</p> <p>We retrieved ten search query suggestions provided by the Google and Bing search autocomplete systems for the input of each of these root queries, without&nbsp;performing a search. Furthermore, we extended the root queries by the letters a to z (e.g., &quot;democrats&quot; (root term) &gt;&gt; &quot;democrats a&quot; (query input) &gt;&gt;&nbsp;&quot;democrats and recession&quot; (query suggestion)) to simulate a user&#39;s input during information search and generate a total of up to 270 query suggestions per topic and search engine. The dataset we provide contains columns for root term, query input, and query suggestion for each suggested query. The location from which the search is performed is the location of the Google servers running Colab, in our case Iowa in the United States of America, which is added to the dataset.&nbsp;</p> <p><strong>AllSides Scraper</strong></p> <p>At&nbsp;<a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>, we provide a scraping tool, that allows for the automatic retrieval of all available articles at the AllSides balanced news headlines.&nbsp;</p> <p>We want to provide an easy means of retrieving the news and all corresponding information. For many tasks it is relevant to have the most recent documents available. Thus, we provide this Python-based scraper, that scrapes all available AllSides news articles and gathers available information. By providing the scraper we facilitate access to a recent version of the dataset for other researchers.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Media Bias Aware Simulation Dataset

<p>We utilized the Hyperpartisan News Detection Dataset, released with the SemEval-2019 Task 4 Hyperpartisan detection task, due to its extensive bias labels. To ensure accurate bias labels, we used the Overlap-checking (1:1) model, retaining only articles where the distant supervision bias labels matched the model's predictions. This validation process resulted in 409,757 articles.</p> <p>These articles span from 1960 to 2018, with a sparse distribution in earlier years. We focused on articles from May 1, 2017, to December 31, 2017, resulting in a subset of 72,940 news articles, ensuring a consistent daily news flow. We processed this subset by removing HTML tags and special characters and generating news summaries using PEGASUS.&nbsp;We used Latent Dirichlet Allocation (LDA) to categorize the articles into 20 news themes, based on perplexity scores.</p> <p>This dataset was then fed into the simulation framework, with a cut-off date of June 24, 2017.</p> <ul> <li><strong>News Recommendation Dataset</strong>: Includes user-item interaction records from May 1 to June 24, providing users' reading histories and interacted news articles. <ul> <li><strong>Training Split</strong>: Data from May 1 to June 17, used to train news recommendation algorithms.</li> <li><strong>Evaluation Split</strong>: Data from June 17 to June 24, used to evaluate the trained recommendation algorithms.</li> </ul> </li> <li> <p><strong>Candidate News Dataset</strong>: News articles published from June 25 to December 31, presented to users during simulations.</p> </li> </ul> <p>For more information, please visit <a href="https://github.com/ruanqin0706/UserRecSimulation" target="_new" rel="noreferrer">https://github.com/ruanqin0706/UserRecSimulation</a>.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Raw data and media: Tetraspanins are unevenly distributed across single extracellular vesicles and bias sensitivity to multiplexed cancer biomarkers

<p>Raw datasets and media accompanying the manuscript:&nbsp;T<strong>etraspanins are unevenly distributed across single extracellular vesicles and bias sensitivity to multiplexed cancer biomarkers</strong>, published in the Journal of Nanobiotechnology&nbsp;</p>

opencc-zeroAug 2021View details →
dryad32/100

Data from: Exploring media representation of the exotic pet trade, with a focus on welfare: Taxonomic, framing, and language biases in peer-reviewed publications and newspaper articles

Open the record for dataset details and reuse information.

publicJun 2025View details →
zenodo24/100

Sentiment analysis of media's political bias in micro-blogging - A computational framework for optimized recommendation systems

<p>This dataset contains tweets from four Pakistani news channels, namely DAWN, GEO, ARY, and 24News, collected using the twarc command line tool with a Twitter academic researcher account. The tweets were collected between Nov, 2015, and April 4, 2022, and relate to three major political parties in Pakistan, PTI, PMLN, and PPP through their official names in the relevant news channel page tweets.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo20/100

Rewritten Media Bias News Headlines

<p>This dataset was used in the paper "Rewriting Bias: Mitigating Media Bias in News Recommender Systems through Automated Rewriting" by Qin Ruan, Jin Xu, Susan Leavy, Brian Mac Namee, and Ruihai Dong, presented at the 32nd ACM Conference on User Modeling, Adaptation and Personalization (UMAP'24), July 1-4, 2024, in Cagliari, Italy. (ACM, New York, NY, USA, 10.1145/3627043.3659541).</p> <p>The dataset includes seven distinct rewritten versions of news headlines generated using the following methods proposed in the paper: RADJ, RADV, RNOUN, RVERN, RALL, RG3.5, and RG4.0.</p> <p>The rewriting approaches are categorised into two main categories: <strong>Word Replacement </strong>approaches and <strong>Large Language Models </strong>approaches.</p> <p>1.Sentence Rewriting Using Word Replacement :</p> <ul> <li><strong>Replace Adjectives (RADJ)</strong>: This approach replaces adjectives identified as contributing to bias with neutral or opposite ones.</li> <li><strong>Replace Adverbs (RADV)</strong>: This approach replaces adverbs identified as contributing to bias with neutral or opposite ones.</li> <li><strong>Replace Nouns (RNOUN)</strong>: This approach replaces nouns identified as contributing to bias with neutral or more general ones.</li> <li><strong>Replace Verbs (RVERB)</strong>: This approach replaces verbs identified as contributing to bias with neutral or more factual ones.</li> <li><strong>Replace All (RALL)</strong>: This is a combination of all the above approaches.</li> </ul> <p>2.Sentence Rewriting Using Large language Models:</p> <ul> <li><strong>Sentence Rewriting using GPT-3.5 (RG3.5)</strong>: This method rephrases sentences using GPT-3.5 based on a specific prompt.</li> <li><strong>Sentence Rewriting using GPT-4.0 (RG4.0)</strong>: This method rephrases sentences using GPT-4.0 based on a specific prompt.</li> </ul> <p>The dataset consists of original news headlines and their corrsponding rewritten versions prodcued by each of the seven approaches mentioned above. Each rewritten version aims to reduce bias while maintaining the original meaning.</p> <p>&nbsp;</p> <p><strong>Dataset Structure:</strong></p> <p>The dataset consists of a single file with the following columns:</p> <ul> <li>NewsId: A unique identifier for each news article.</li> <li>Original: The original headline of the news article.</li> <li>RG3.5: The headline rewritten using the <a href="https://platform.openai.com/docs/models/gpt-3-5">GPT-3.5</a> model.</li> <li>RG4.0: The headline rewritten using the <a href="https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo">GPT-4.0</a> model.</li> <li>RADJ: The headline with adjectives replaced to reduce bias.</li> <li>RADV: The headline with adverbs replaced to reduce bias.</li> <li>RVERB: The headline with verbs replaced to reduce bias.</li> <li>RNOUN: The headline with nouns replaced to reduce bias.</li> <li>RALL: The headline with adjectives, adverbs, verbs. noun replaced to reduce bias.</li> </ul> <p><strong>Usage:</strong></p> <p>This dataset can be used to study the impact of different sentence rewriting approaches on reudicng bias in enws healines. It is also suitable for researchers interested in natural language processing, media bias reduction, and the application of large language models in generative AI.</p> <p>For more information, please visit <a href="https://github.com/ruanqin0706/MediaBiasinNewsRec">https://github.com/ruanqin0706/MediaBiasinNewsRec</a></p> <p><strong>Contact Information:&nbsp;</strong>qin.ruan@ucdconnect.ie</p>

restrictedcc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record