Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
Forecasting with news sentiment: Evidence with UK newspapers
<p>These are datasets of economic sentiments derived from Uk newspapers using a dictionary and support vector machines. For more information on the application refer to : </p> <p>Rambaccussing, D. and Kwiatkowski, A., 2020. Forecasting with news sentiment: Evidence with UK newspapers. <em>International Journal of Forecasting</em>, <em>36</em>(4), pp.1501-1516.</p> <p>https://www.sciencedirect.com/science/article/pii/S0169207020300595</p> <p> </p> <p> </p>
Lexicon and example extensions from paper "Lexicon-based comments-oriented news sentiment analyzer system"
<p>Lexicon and example extensions for the article "Moreo, Alejandro, et al. "Lexicon-based comments-oriented news sentiment analyzer system." <em>Expert Systems with Applications</em> 39.10 (2012): 9166-9180." Founded by Ministerio de Educación y Ciencia and Junta de Andalucía with Projects: TIN2007-60199, TIC2009-5011 and TIN2007-67984</p>
Changes in the use of 'refugees' and 'migrants' in international news media and academic publications 2010–2022
<p>The dataset contains source data for Figure 2 in the article, based on search results from Factiva and Web of Science.</p>
An Evaluation Framework for Mapping News Headlines to Event Classes in a Knowledge Graph
<p>News headline to event classes corpus derived from Wikidata.</p> <p>Github: https://github.com/mbouadeus/news-headline-event-linking</p>
Forex News Annotated Dataset for Sentiment Analysis
<p>This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY. The data was extracted from reputable platforms <a href="https://www.forexlive.com">Forex Live</a> and <a href="https://www.fxstreet.com/">FXstreet</a> over a period of 86 days, from January to May 2023. The dataset comprises 2,291 unique news headlines. Each headline includes an associated forex pair, timestamp, source, author, URL, and the corresponding article text. Data was collected using web scraping techniques executed via a custom service on a virtual machine. This service periodically retrieves the latest news for a specified forex pair (ticker) from each platform, parsing all available information. The collected data is then processed to extract details such as the article's timestamp, author, and URL. The URL is further used to retrieve the full text of each article. This data acquisition process repeats approximately every 15 minutes.</p> <p>To ensure the reliability of the dataset, we manually annotated each headline for sentiment. Instead of solely focusing on the textual content, we <strong>ascertained sentiment based on the potential short-term impact of the headline on its corresponding forex pair</strong>. This method recognizes the currency market's acute sensitivity to economic news, which significantly influences many trading strategies. As such, this dataset could serve as an invaluable resource for fine-tuning sentiment analysis models in the financial realm.</p> <p>We used three categories for annotation: 'positive', 'negative', and 'neutral', which correspond to bullish, bearish, and hold sentiments, respectively, for the forex pair linked to each headline. The following Table provides examples of annotated headlines along with brief explanations of the assigned sentiment. </p> Examples of Annotated Headlines Forex Pair Headline Sentiment Explanation GBPUSD Diminishing bets for a move to 12400 Neutral Lack of strong sentiment in either direction GBPUSD No reasons to dislike Cable in the very near term as long as the Dollar momentum remains soft Positive Positive sentiment towards GBPUSD (Cable) in the near term GBPUSD When are the UK jobs and how could they affect GBPUSD Neutral Poses a question and does not express a clear sentiment JPYUSD Appropriate to continue monetary easing to achieve 2% inflation target with wage growth Positive Monetary easing from Bank of Japan (BoJ) could lead to a weaker JPY in the short term due to increased money supply USDJPY Dollar rebounds despite US data. Yen gains amid lower yields Neutral Since both the USD and JPY are gaining, the effects on the USDJPY forex pair might offset each other USDJPY USDJPY to reach 124 by Q4 as the likelihood of a BoJ policy shift should accelerate Yen gains Negative USDJPY is expected to reach a lower value, with the USD losing value against the JPY AUDUSD <p>RBA Governor Lowe’s Testimony High inflation is damaging and corrosive </p> Positive Reserve Bank of Australia (RBA) expresses concerns about inflation. Typically, central banks combat high inflation with higher interest rates, which could strengthen AUD. <p>Moreover, the dataset includes two columns with the predicted sentiment class and score as predicted by the <a href="https://huggingface.co/ProsusAI/finbert">FinBERT</a> model. Specifically, the FinBERT model outputs a set of probabilities for each sentiment class (positive, negative, and neutral), representing the model's confidence in associating the input headline with each sentiment category. These probabilities are used to determine the predicted class and a sentiment score for each headline. The sentiment score is computed by subtracting the negative class probability from the positive one.</p>
Quotatives Indicate Decline in Objectivity in U.S. Political News
<p>Data to reproduce results for the paper </p> <p>Quotatives Indicate Decline in Objectivity in U.S. Political News (ICWSM 2023)</p> <p>Code: https://github.com/epfl-dlab/quotative_bias</p>
A small dataset of news from NYT
<p>This dataset has been collected during the ISWS2023 PhD School of Semantic Web at Bertinoro for the Project Work of the Dragon Team Research Group.</p> <p>It contains a collection of 10 news (not the entire articles, just the first few paragraphs) get from the New York Times.</p> <p>The news have been published in a date that is after the data of the pre-train of the gpt-3.5-turbo LLM so contains informations that are not known by this model and can be useful to study the inference of new triples in KG generation from texts process.</p>
Indian Broadcast News Debate (IBND) Corpus
<p>The Indian Broadcast News Debate (IBND) corpus is created by collecting news debates from two popular English Indian news channels The corpus contains audio data of 15 TV news debates obtained from two Indian English news channels. The total duration of the corpus is 12 hours and 47 minutes. The corpus contains 94 unique speakers (83 male and 11 female), excluding field reporters. The typical duration of each debate varies from 20 to 60 minutes. The IBND corpus consists of annotations for shouted vs. normal speech, overlapped vs. single speaker's speech and competitive vs. non-competitive speech. </p>
Network embedding for understanding the National Park System through the lenses of news media, scientific communication and biogeography
<p>The United States national parks encompass a variety of biophysical and historical resources important for national cultural heritage. Yet how these resources are socially constructed often depends upon the beholder. Parks tend to be conceptualized according to their (fixed) geographic context, so our understanding of this system of systems is dominated by this geographic lens. To expose the systemic structure that exists beyond their geographic embedding, we analyze three representations of the national park system using park-park similarity networks according to their co-occurrence in: (a) ~423,000 news media articles; (b) ~11,000 research publications; and (c) ~60,000 species inhabiting parks. We quantify structural variation between network representations by leveraging similarity measures at different scales: park-level (park-park correlations) and system-level (network communities' consistency). Because parks are governed and experienced at multiple scales, cross-network comparison informs how management should account for the varying objectives and constraints that dominate at each scale. Our results identify an interesting paradox: whereas park-level correlations depend strongly on the representative lens, the network communities are remarkably robust and consistent with the underlying geographic embedding. Our data-driven methodology is generalizable to other geographically embedded socio-environmental systems and supports the holistic analysis of systems-level structure that may elude other approaches.</p>
Raw data for manuscript: Use of immunology in news and YouTube videos in the context of COVID-19: politicization and information bubbles
<p>Coding of newsarticles and videos related to immunology and COVID-19 in Italian and English</p>
Sample - Donald Trump's tweets mentioning the term 'fake news' (2017-2021)
<p>This is a dataset of the sample used for the paper "<strong>Disintermediation and disinformation as a political strategy: using AI to analyse Trump’s “fake news” discourse on Twitter" </strong>by Diez-Gracia, A., Sánchez-García, P. & Martín-Román, J. (2023) published in the Journal 'El Profesional de la Información' (pending DOI). Obtained by filtering the open-source database http://www.thetrumparchive.com</p> <p>Article full reference: Diez-Gracia, A., Sánchez-García, P. & Martín-Román, J. (2023). Disintermediation and disinformation as a political strategy: using AI to analyse Trump's "fake news" discourse on Twitter. El profesional de la información, vol., n.</p>
Fake News Dataset with In-content Annotations and Detailed Lying Excerpts
<p>This dataset contains 95 fake news collected from two Brazilian fact-checking services (E-farsas and Boatos). We carefully annotated some excerpts of the fake news to enable deeper analyses of their falsehood. From these annotations, we divided the news into three groups: real news, totally fake news and fake news but which contain only a few lying snippets.</p><p>The annotated fragments were grouped based on four categories of falsehood: untrue, unverifiable fact, incorrectly named entity, and exaggeration. Other parts of the fake news, such as the passages around the fragments, were separated to add more information to the results of this research.</p>
NLP and machine learning to measure peace from news media
Open the record for dataset details and reuse information.
Incentivizing news consumption on social media platforms using large language models and realistic bot accounts
Open the record for dataset details and reuse information.
Data from: Fake news? The impact of information mismatch in mating behaviour
Open the record for dataset details and reuse information.
A content analysis of Vietnam online news about a pentavaent vaccine in the EPI
Open the record for dataset details and reuse information.
Network embedding for understanding the National Park System through the lenses of news media, scientific communication and biogeography
Open the record for dataset details and reuse information.
News article mentions of 4chan from three weeks in 2017 and posts with links to these articles on 4chan/pol/
<p>This upload includes two dataset. The first comprises articles that mentioning "4chan" in its post body as retreived by Nexis Uni. We identified the three weeks in which the highest spikes in terms of the amount of articles occured in 2017 (see the image mentions_nexis_uni_2017_weeks.png ) and extracted the articles that were published within these. The dataset includes the article source, the article title, and the sentences where they mention 4chan to analyse the framing. We deleted the article bodies for copyright reasons. Two columns indicate the URL to the article if it appeared online. Additionally, we added a column with URLs to archive.is links of the same articles, in case they were archived.</p> <p>The second dataset consists of posts on the imageboard 4chan/pol/ that refer to one of the aforementioned URLs.</p>
alfalifr/twitterCrawlViral: Twitter Crawling for "Viral" or "Heboh" News
<p>The data we take is about tweet form 10 official account of news portal on twitter that mentioned word "Viral" or "Heboh". The dataset contain 5 columns and sorted by time.</p>
SemEval-2020 Task 7: Assessing Humor in Edited News Headlines
<p>This is the task dataset for SemEval-2020 Task 7: Assessing Humor in Edited News Headlines.</p> <p>The task’s dataset contains news headlines in which short edits were applied to make them funny, and the funniness of these edited headlines was rated using crowdsourcing. This task includes two subtasks, the first of which is to estimate the funniness of headlines on a humor scale in the interval 0-3. The second subtask is to predict, for a pair of edited versions of the same original headline, which is the funnier version.</p> <p>CodaLab page hosting the competition:<br> <a href="https://competitions.codalab.org/competitions/20970">https://competitions.codalab.org/competitions/20970</a></p> <p>Starter Github code (scripts for running baseline and evaluation):<br> <a href="https://github.com/n-hossain/semeval-2020-task-7-humicroedit">https://github.com/n-hossain/semeval-2020-task-7-humicroedit</a></p> <p>Task mailing list:<br> <a href="https://groups.google.com/forum/#!forum/semeval-2020-task-7-all">https://groups.google.com/forum/#!forum/semeval-2020-task-7-all</a><br> ----------------------------------------------------------------------</p> <p>ZIP contents:<br> -------------</p> <p>Folders:<br> - subtask-1: Dataset for the funniness regression subtask.<br> - subtask-2: Dataset for the "Funnier of the Two" classification subtask.</p> <p>Files:<br> - {train, dev, test}.csv: the task's dataset including labels<br> - train_funlines.csv: additional training data gathered from the FunLines competition (https://funlines.co)<br> - baseline.zip: contains csv file which is the output of the BASELINE system. This is a template of the output format that can be submitted to CodaLab for scoring.</p> <p><strong>Reference</strong></p> <p>Please cite the task paper when using this dataset:</p> <p>Nabil Hossain, John Krumm, Michael Gamon and Henry Kautz. 2020. Semeval-2020 Task 7: Assessing Humor in Edited News Headlines. In Proceedings of International Workshop on Semantic Evaluation (SemEval-2020).</p> <pre><code>BIBTEX: @InProceedings{hossainSemEval2020Task7, author = {Hossain, Nabil and Krumm, John and Gamon, Michael and Kautz,Henry}, title = {SemEval-2020 {T}ask 7: {A}ssessing Humor in Edited News Headlines}, booktitle = {Proceedings of the 14th International Workshop on Semantic Evaluation ({S}em{E}val-2020)}, address = {Barcelona, Spain}, year = {2020}}</code></pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.