Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
Greek News Sign Language Dataset - Part A
<p>Part A entails 1.000 signed phrases of crime-related news stories broadcasted in Greece.</p>
FaCov Dataset: COVID-19 Viral News and Rumors Fact-Check Articles Dataset
<p>The data were collected by web-scraping pages from the websites collected earlier, using the <a href="https://webscraper.io/">Web Scraper browser extension</a>.</p> <p>More specifically, the sections of these websites that dealt exclusively with COVID-19 related content were scraped. In cases where the website did not have such a specified section, the search functionality within the website was used to query terms related to COVID-19 and the articles in the search results were scraped. Also in some cases, all articles were scraped and those unrelated to COVID-19 were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV</p> <p>The following information was extracted along with the articles:</p> <p>Title of the fact check article</p> <p>URL of the fact check article</p> <p>Claim being discussed in the article (if available)</p> <p>Summary of the fact check article (if available)</p> <p>Content of the fact check article• Label assigned by the article to the claim</p> <p>Author of the fact check article (if available)</p> <p>Date of publication of the article (if available)</p>
covid-19 news stories China, South Korea and the U.S.
<p>This dataset is news stories on the COVID-19 pandemic published by national news agencies in China, Korea, and the U.S. The news stories were collected on<em> Factiva</em> by keyword searching, published within a one-month time frame after a national break in each country. </p>
Greek News Sign Language Dataset - Part B
<p>Part B entails 989 signed phrases of crime-related news stories broadcasted in Greece.</p>
Dataset: Sentiment Analysis annotation of News headlines covering the Olympic legacy of Rio 2016 and London 2012 published by the Brazilian and British online media
<p>Dataset of 464 news headlines with sentiment manually annotated by a domain expert using the labels positive, negative and neutral. Data contains URLs for news articles published between 2004-2020 by the British and Brazilian media in English and Brazilian Portuguese covering the Olympic legacies of London 2012 and Rio 2016. Articles were collected from the news outlets’ websites using Google search engine.</p> <p>News outlets:</p> <ul> <li>The Guardian</li> <li>Daily Mail</li> <li>Globo</li> <li>Estadao</li> </ul>
Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text
<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label ‘other’ was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article's publication (DD/MM/YYYY)</li> <li>Link: article's link</li> <li>Title: article's title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article's body text</li> </ul> <p><em>PS: Documents do not include articles' body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label 'other' was used. Therefore, 'other' identifies contradictory results.</li> </ul> <p> </p>
News headlines of BBC articles published by @BBCBreaking twitter account
<p>The dataset consists of a list of news articles headlines retrieved from tweets published by @BBCBreaking profile in specific years (2012, 2015, 2017, 2019 and 2022).</p> <p>The dataset is in <code>.csv</code> format and is organised as follows:</p> <ul> <li>Columns: <ul> <li>ID (tweet ID)</li> <li>created_at (tweet publication's date)</li> <li>url (url of the news article attached to the tweet)</li> <li>Titles (news headline)</li> </ul> </li> <li>Rows: Each row contains a single news article headline sorted by date of publication (created_at). Total number of entries: 7213.</li> </ul> <p>For more details about data collection refer to <a href="https://github.com/caiocmello/news-mood">Github</a>.</p>
Propaganda and fake news on the war in Ukraine: data from Russian-speaking social media communities
<p>The data set contains posts from social media networks popular among Russian-speaking communities. Information was searched based on pre-defined keywords ("war", "special military operation", etc.) and is mainly related to the ongoing war in Ukraine with Russia. After a thorough review and analysis of the data, both propaganda and fake news were identified. The collected data is anonymized. Feature engineering and text preprocessing can be applied to obtain new insights and knowledge from this data set. The data set is useful for the study of information wars and propaganda identification.</p>
IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022
<p>IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022</p> <p>Reference paper</p> <p>Madodonga, A., Marivate, V., & Adendorff, M. (2023). Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati. <em>Journal of the Digital Humanities Association of Southern Africa</em>, <em>4</em>(01). https://doi.org/10.55492/dhasa.v4i01.4449</p> <p> </p> <p>> @article{Madodonga_Marivate_Adendorff_2023, title={Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati}, volume={4}, url={https://upjournals.up.ac.za/index.php/dhasa/article/view/4449}, DOI={10.55492/dhasa.v4i01.4449}, author={Madodonga, Andani and Marivate, Vukosi and Adendorff, Matthew}, year={2023}, month={Jan.} }</p>
News data for studying media exposure to the Boston Marathon bombings
<p>The news data sets released here have been used to study the relationship between media exposure and individuals' threat perception. Media exposure to mass violence has been shown to have a detrimental impact on people's threat perception and mental wellness, but little has been done to explore how exposure to different news content may impact mental health in people's everyday lives. In our study, we empirically test how emotionally potent media coverage of a real-world threat, namely, the Boston Marathon bombings occurred in 2013, alters threat perception of the community members over the first and the third anniversaries (in 2014 and 2016).</p> <p>The data were collected using a wave-based longitudinal design. There are two data sets, and each covers three waves:</p> <ul> <li>Dataset (I) -- news coverage before (Wave 1), during (Wave 2), and after (Wave 3) 2014 anniversary</li> <li>Dataset (II) -- news coverage before (Wave 1), during (Wave 2), and after (Wave 3) 2016 anniversary</li> </ul> <p>The collection procedure was informed by our survey study. Based on the survey completed by our subjects, we identified the four most frequent news outlets in the response: Metro (MT), New York Times (NY), Boston Globe (BG), and Boston Herald (BH). Other outlets, such as USA Today and Wall Street Journal, were reported by less than ten respondents. Therefore, our data collection focused on the news published by the four most frequent outlets.</p> <p>The data sets include the metadata of the news coverage over the aforementioned six waves. The raw content of the news stories was removed to respect the copyright owners.</p> <p><strong>Summary of the data collection procedure </strong></p> <p>We used news aggregators including Google and Yahoo news, to retrieve news articles published by the four outlets on a daily basis. We first collected the URLs of the news articles from the news aggregators and retrieved and parsed the news content using an HTML parser. In total, we collected over 38.5K and 54.1K news articles in dataset I and II, respectively.</p> <p>There are six files; each correspond to news coverage from the outlets in each wave. In these files, each line contains four columns: outlet, time, title, url which indicate the outlet of each news article, the time of publishing, the title of the article, and the URL to the article.</p> <p>We are making the data sets available for academic researchers and public use, to enable the discovery of new insights and development of better techniques to improve crisis communication and mental wellness.</p>
Updates applied to Flemish online news and their associated change types
<p>This dataset contains 291,666 news articles produced by six different Flemish online news outlets (VRT NWS, Knack, Het Laatste Nieuws, Het Nieuwsblad, De Morgen, De Standaard), together with (in total 197,979) updates applied to these news articles in the first 24 hours after publication. The respective article versions (one row per article version that has been put online over time) can be found in the 'article_versions_vrt.csv', 'article_versions_knack.csv', 'article_versions_hln.csv', 'article_versions_nieuwsblad.csv', 'article_versions_demorgen.csv' and 'article_versions_standaard.csv'. Documentation regarding the meaning of the attributes in these files is provided in 'article_versions_README.txt'.</p> <p>Next to this dataset, we also provide a coded set of changes made during a subset of the news updates in 'coded_article_change_types.csv'. The file contains all text extracts that are part of a specific change, together with the type of the change to which the text extract belongs. Corresponding documentation is provided in 'coded_article_change_types_README.txt'.</p>
Fake News
<p>This data collection focuses on capturing user-generated content from the popular social network Reddit in 2024. The dataset “Fake News” comprises collected data from 3636 users of Reddit. This dataset consists of .csv .xls, and .xlsx files, containing textual data associated with fake news.</p> <p>Funded by the EU NextGeneration EU through the Recovery and Resilience Plan for Slovakia under the project No. 09I03-03-V01-000153</p>
University of Notre Dame News: A Reading
I have done a bit of analysis -- reading -- against the set of news distributed by the University of Notre Dame, and below is some of what I learned.
Figure 5 in Rotifers of Bahia State, Brazil: News records and limitations to studies
Figure 5. Numbers of Rotifera species per locality in Bahia State, Brazil. The codes follow Table 1.
Figure 1 in Rotifers of Bahia State, Brazil: News records and limitations to studies
Figure 1. Map of Bahia State, Brazil, highlighting in the 13 sampling sites. Sampling sites described in Table 1.
Figure 3 in Rotifers of Bahia State, Brazil: News records and limitations to studies
Figure 3. Rotifers from Bahia State, Brazil, sampled from 2010 to 2016. M. Hexarthra intermedia brasiliensis Hauer, 1953. N. Filinia opoliensis (Zacharias, 1898). O.Filinia terminalis (Plate, 1886). P.Lecane aquila Harring & Myers, 1926. Q. Lecane bulla bulla (Gosse, 1851). R. Lecane cornuta (Müller, 1786). S. Lecane quadridentata (Ehrenberg, 1830). T. Lecane monostyla (Daday, 1897). U. Lecane hornemanni (Ehrenberg, 1834). V.Lecane leontina (Turner, 1892). W.Lecane ludwigii (Eckstein, 1883). X.Lecane lunaris crenata (Harring, 1913). Species stained with bengal rose. Scale bars= 100 μm.
Figure 2 in Rotifers of Bahia State, Brazil: News records and limitations to studies
Figure 2. Rotifers from Bahia State, Brazil, sampled from 2010 to 2016.A.Anuraeopsis fissa Gosse, 1851. B.Brachionus calyciflorus Pallas, 1766. C. Brachionus caudatus f. austrogenitus Ahlstrom, 1940. D.Brachionus falcatus Zacharias, 1898. E.Brachionus quadridentatus quadridentatus Hermann, 1783. F.Brachionus urceolaris urceolaris Müller, 1773. G.Keratella cochlearis (Gosse, 1851). H.Platyias quadricornis (Ehrenberg, 1832). I. Dipleuchlanis propatula (Gosse, 1886). J. Testudinella dendradena de Beauchamp, 1955. K. Trichocerca pusilla (Jennings, 1903). L. Squatinella mutica (Ehrenberg, 1832). Species stained with bengal rose. Scale bars= 100 μm.
Dataset: News Corporation (NWS) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: News Corporation (NWSA) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: How do news about a heatwave affect public prioritization of climate change adaptation and mitigation behaviors?
<p><span>These datasets contain survey data that was used to evaluate the effect of the exposure to heatwave news texts on people’s preference for climate mitigation and adaptation actions, as presented in the manuscript titled “<em>How do news about a heatwave affect public prioritization of climate change adaptation and mitigation behaviors?</em>”. Three versions of the dataset are available:</span></p> <ol> <li><strong>Original dataset</strong>: This version contains choice text as data points and includes all finished survey responses that passed the attention check questions (n=1209).</li> <li><strong>Original recoded dataset</strong>: This version was generated by recoding choice text into numerical values. The 'Income' variable, representing household income levels for both Canadian and US residents, was added by converting reported income ranges to a unified scale based on exchange rate equivalencies. The "Income_Canadians" and "Income_US" columns were subsequently removed to avoid repetitions. </li> <li><strong>Final dataset</strong>: This version excludes observations from participants who completed the survey in under four minutes and those who selected the same response for every item within each matrix-style question (also known as straight-lining). Additionally, responses with missing values in questions regarding political views, gender, and household income, as well as responses where participants identified as non-binary or indicated that their gender was not listed, were omitted (see “Methods” for more details). Dependent variables have been added based on the original responses, including personal-level mitigation and adaptation likelihoods, personal-level mitigation preference, and both non-weighted and weighted collective-level mitigation preference. Furthermore, the dataset includes a 'Climate Change Concern' variable, derived through principal component analysis of thirteen variables expressing participants’ climate change attitudes and efficacy beliefs concerning climate actions. Variables not used in the subsequent data analysis were removed. Age, political views, education, and income columns were standardized. The final dataset was used for the data analysis presented in the manuscript.</li> </ol> <p>The following variables/columns can be found across the three versions of the dataset:</p> <ul> <li>Dependent variables: <ul> <li>Starting with “<em>Personal_Mitigation</em>”: participant’s self-reported likelihood of taking selected personal-level climate change mitigation actions</li> <li>Starting with “<em>Personal_Adaptation</em>”: participant’s self-reported likelihood of taking selected personal-level climate change adaptation actions</li> <li>Starting with “<em>Collective_Mitigation</em>”: participant’s ranking of the collective-level climate change mitigation initiatives</li> <li>Starting with “<em>Collective_Adaptation</em>”: participant’s ranking of the collective-level climate change adaptation initiatives</li> <li><em>Personal_Mitigation_Likelihood</em>: personal-level mitigation likelihood (present only in the final dataset)</li> <li><em>Personal_Adaptation_Likelihood</em>: personal-level adaptation likelihood (present only in the final dataset)</li> <li><em>Personal_Preference</em>: personal-level mitigation preference (present only in the final dataset)</li> <li><em>Collective_Preference_Unweighted</em>: non-weighted collective-level mitigation preference (present only in the final dataset)</li> <li><em>Collective_Preference_Weighted</em>: weighted collective-level mitigation preference (present only in the final dataset)</li> </ul> </li> <li>Independent variables: <ul> <li><em>Group</em>: group that the participant was assigned to as part of the experimental intervention</li> <li><em>Distance</em>: indicates whether the participant was assigned to read about a heatwave occurring in their community or a city 6,000 km away (for experimental groups only)</li> <li><em>Severity</em>: indicates whether the participant was prompted to read about a heatwave without or with the mention of associated causalities (for experimental groups only)</li> </ul> </li> <li>Covariates and supporting variables: <ul> <li><em>Gender</em>: gender identity</li> <li><em>Identity</em>: ethnic and/or racial identity</li> <li><em>Age</em>: age</li> <li><em>Political_Views</em>: position on the liberal-conservative continuum</li> <li><em>Education</em>: highest level of education</li> <li><em>Country</em>: country of residence</li> <li><em>Canada_Province</em>: province or territory of residence (for Canadian participants only)</li> <li><em>US_State</em>: state of residence (for US participants only)</li> <li><em>Duration_Residence</em>: duration of residence in the current community</li> <li><em>Income_Canadians</em>: annual household income in Canadian dollars (for Canadian participants only)</li> <li><em>Income_US</em>: annual household income in US dollars (for US participants only)</li> <li><em>Income</em>: annual household income for both Canadian and US residents derived by converting reported income ranges to a unified scale based on exchange rate equivalencies</li> <li><em>Efficacy_Mitigation_Personal</em>: belief regarding the response efficacy of personal-level climate change mitigation actions</li> <li><em>Efficacy_Mitigation_Collective</em>: belief regarding the response efficacy of collective-level climate change mitigation actions</li> <li><em>Efficacy_Adaptation_Personal</em>: belief regarding the response efficacy of personal-level climate change adaptation actions</li> <li><em>Efficacy_Adaptation_Collective</em>: belief regarding the response efficacy of collective-level climate change adaptation</li> <li><em>Climate_Change_Importance:</em> perception of climate change as a personally important issue</li> <li>Climate_Change_Worry: level of worry about climate change</li> <li>Starting with “<em>Climate_Risk</em>”: beliefs regarding the degree of harm that climate change will cause to plants and animal species (Climate_Risk_Animals_Plants), future generations of people (Climate_Risk_Future_Generations), people in developing countries (Climate_Risk_Developing_Countries), people in participant’s country (Climate_Risk_Country), people in participant’s community (Climate_Risk_Community), and the participant personally (Climate_Risk_Personal)</li> <li>Climate_Change_Onset_Time: belief regarding when climate change will start harming people in their community</li> <li><em>Six_Americas_Segment</em>: the Global Warming's Six Americas segment participant aligns with derived based on the Six Americas Short SurveY (SASSY) Group Scoring Tool</li> <li><em>Climate_Change_Concern</em>: variable derived through PCA of thirteen variables expressing participants' climate change attitudes and efficacy beliefs pertaining to climate actions (present only in the final dataset)</li> <li><em>Survey_Duration_Seconds</em>: The amount of time it took the respondent to complete the survey</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.