Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “discourse analysis”
Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 5
<p>Dataset containing four .xlsx and .csv files for the exercises in Episode 5 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>
Data set 1 discourse analysis BRAD research project
<p>Discourse analysis data set with excerpts of press articles generated in the coding (coded with keywords ‘Brexit’ and ‘deportations’). This data set connects to the WP3 of the BRAD research project.</p>
Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 2
<p>Dataset containing three subgenre-specific .xlsx files for the exercises in Episode 2 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>
Dataset from "Merging Digital Humanities and Discourse Analysis in the Study of COVID-19 Vaccine Distribution in Norwegian Newspapers" (Sverdljuk et al. 2022)
<p>Contains URNs (identifiers) for the newspapers used in the corpus study "Merging Digital Humanities and Discourse Analysis in the Study of COVID-19 Vaccine Distribution in Norwegian Newspapers".</p> <p>For each subcorpus there is an Excel file containing references to the objects used, together with basic metadata.</p> <p>The corpus definitions can be used in various webapps of the DH-LAB at the National Library of Norway, e.g.:</p> <p><a href="https://beta.nb.no/dhlab/concordances/">https://beta.nb.no/dhlab/concordances/</a></p> <p><a href="https://beta.nb.no/dhlab/collocations/">https://beta.nb.no/dhlab/collocations/</a></p> <p>See more at <a href="https://www.nb.no/dh-lab/">https://www.nb.no/dh-lab/</a></p>
Dataset of the linguistic analysis of the Eastern German Crisis Discourse from 1976 to 1986
<p>The dataset consists of a ZIP file with the speeches contained in the five volumes of the protocol of the party congress of the Socialist Unity Party of Germany (SED) between 1976 and 1986. The texts have been digitized as PDF files and then converted into machine-readable TEXT files using an OCR software. These TEXT files have been parsed with TagAnt (v. 2.0.4 Windows 10 64-bit), an annotation software. According to the data returned by AntConc (v. 4.0.5 Windows 10 64-bit), four corpora have been created: a main corpus with a total of 184,750 tokens and 16,143 types, a 'corpus A' with 70,533 tokens and 11,964 types, a 'corpus B' with 65,757 tokens and 11,967 types, and a 'corpus C' with 48,460 tokens and 8,145 types. The main corpus includes all speeches from the five volumes, while the three additional corpora have been created based on specific criteria or topics. The 'corpus A' and 'corpus B' have similar token counts and types, and likely differ based on a specific subset of speeches or themes. The 'corpus C' is the smallest corpus, with a focus on a specific aspect of the discourse. This dataset is suitable for exploring and analyzing the Eastern German Crisis Discourse from 1976 to 1986, particularly for scholars who may be interested in political and historical analysis.</p>
Documents used in the PLANET4B analysis of biodiversity discourse by environmental NGOs
<p>These files include press releases that have been published on the internet by European environmental NGOs, and which were used in the PLANET4B project analysis of the discourse on biodiversity.</p>
Documents used in the PLANET4B analysis of biodiversity discourse by political parties
<p>These files include press releases that have been published on the internet by European political parties, and which were used in the PLANET4B project analysis of the discourse on biodiversity.</p>
Documents used in the PLANET4B D1.1 analysis of biodiversity discourse by news outlets - 2010 and 2022 Data
<p>Data used to analyse biodiversity discourse in news outlet as part of Deliverable D1.1. of the Planet4B Project.</p>
Discourse Analysis of Spanish Geographical Indications
<p>The data were collected through a questionnaire administered between November 2019 and August 2020. The dataset presents 44 statements organized by columns, in accordance with the distribution of the grid used in the Q methodology.</p>
Analysis of Network of Discourse, Nature
<p><strong>Note:</strong> This dataset is also available as a Git repository, currently located at <a href="https://codeberg.org/cpence/nature-debate.charlespence.net">https://codeberg.org/cpence/nature-debate.charlespence.net</a>.</p> <p>This is the raw data accompanying a network analysis of biologists participating in the "biometry-Mendelism" debate, as it occurred in the correspondence pages of the journal <em>Nature,</em> between 1890 and 1915. It accompanies the following book chapter:</p> <p>Pence, Charles H. forthcoming (likely 2021). "How Not to Fight About Theory: The Debate Between Biometry and Mendelism in <em>Nature</em>, 1890–1915." In De Block, A. and Ramsey, G., eds., <em>The Dynamics of Science: Computational Frontiers in History and Philosophy of Science.</em> Pittsburgh, PA: Univ. of Pittsburgh Press.</p>
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis”, Proceedings of the 7th International Conference on Machine Learning and Natural Language Processing (MLNLP 2024), Chengdu, China, October 18-20, 2024 (Paper accepted for publication, Preprint available at: https://arxiv.org/abs/2410.03293)</p> <p> </p> <p><strong>Abstract</strong></p> <p>The outbreak of COVID-19 served as a catalyst for content creation and dissemination on social media platforms, as such platforms serve as virtual communities where people can connect and communicate with one another seamlessly. While there have been several works related to the mining and analysis of COVID-19-related posts on social media platforms such as Twitter (or X), YouTube, Facebook, and TikTok, there is still limited research that focuses on the public discourse on Instagram in this context. Furthermore, the prior works in this field have only focused on the development and analysis of datasets of Instagram posts published during the first few months of the outbreak. The work presented in this paper aims to address this research gap and presents a novel multilingual dataset of <strong>500,153 Instagram posts about COVID-19 published between January 2020 and September 2024</strong>. This dataset contains Instagram posts in <strong>161 different languages</strong>. After the development of this dataset, multilingual sentiment analysis was performed using VADER and twitter-xlm-roberta-base-sentiment. This process involved classifying each post as positive, negative, or neutral. The results of sentiment analysis are presented as a separate attribute in this dataset.</p> <p><em><strong>For each of these posts, the Post ID, Post Description, Date of publication, language code, full version of the language, and sentiment label are presented as separate attributes in the dataset.</strong></em></p> <p>The Instagram posts in this dataset are present in <strong>161 different languages</strong> out of which the top 10 languages in terms of frequency are English (343041 posts), Spanish (30220 posts), Hindi (15832 posts), Portuguese (15779 posts), Indonesian (11491 posts), Tamil (9592 posts), Arabic (9416 posts), German (7822 posts), Italian (5162 posts), Turkish (4632 posts)</p> <p>There are <strong>535,021 distinct hashtags in this dataset</strong> with the top 10 hashtags in terms of frequency being #covid19 (169865 posts), #covid (132485 posts), #coronavirus (117518 posts), #covid_19 (104069 posts), #covidtesting (95095 posts), #coronavirusupdates (75439 posts), #corona (39416 posts), #healthcare (38975 posts), #staysafe (36740 posts), #coronavirusoutbreak (34567 posts)</p> <p>The following is a description of the attributes present in this dataset</p> <ul> <li><em><strong>Post ID</strong></em>: Unique ID of each Instagram post</li> <li><em><strong>Post Description</strong></em>: Complete description of each post in the language in which it was originally published</li> <li><em><strong>Date</strong></em>: Date of publication in MM/DD/YYYY format</li> <li><em><strong>Language code</strong></em>: Language code (for example: “en”) that represents the language of the post as detected using the Google Translate API </li> <li><em><strong>Full Language</strong></em>: Full form of the language (for example: “English”) that represents the language of the post as detected using the Google Translate API </li> <li><em><strong>Sentiment</strong></em>: Results of sentiment analysis (using the preprocessed version of each post) where each post was classified as positive, negative, or neutral</li> </ul> <p><strong>Open Research Questions</strong></p> <p>This dataset is expected to be helpful for the investigation of the following research questions and even beyond:</p> <ol> <li>How does sentiment toward COVID-19 vary across different languages?</li> <li>How has public sentiment toward COVID-19 evolved from 2020 to the present?</li> <li>How do cultural differences affect social media discourse about COVID-19 across various languages?</li> <li>How has COVID-19 impacted mental health, as reflected in social media posts across different languages?</li> <li>How effective were public health campaigns in shifting public sentiment in different languages?</li> <li>What patterns of vaccine hesitancy or support are present in different languages?</li> <li>How did geopolitical events influence public sentiment about COVID-19 in multilingual social media discourse?</li> <li>What role does social media discourse play in shaping public behavior toward COVID-19 in different linguistic communities?</li> <li>How does the sentiment of minority or underrepresented languages compare to that of major world languages regarding COVID-19?</li> <li>What insights can be gained by comparing the sentiment of COVID-19 posts in widely spoken languages (e.g., English, Spanish) to those in less common languages?</li> </ol> <p>All the Instagram posts that were collected during this data mining process to develop this dataset were publicly available on Instagram and did not require a user to log in to Instagram to view the same (at the time of writing this paper).</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Data for manuscript: "Prevalence of prejudice denoting words in news media discourse: a chronological analysis"
<p>This data set contains frequency counts of target words in 27 million news and opinion articles from 47 popular news media outlets in the United States. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 4 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1.25 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year such as Reason, Alternet or The American Spectator. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>-cableNews.rar contains prevalence of target words in TV cable news. Data is from Stanford Cable TV News Analyzer (https://tvnews.stanford.edu/)</p> <p>-surveyData.rar contains longitudinal survey data used in the manuscript and links to original sources</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in "The Hill" and "Newsmax" news outlets where XPath expressions had some shortfalls at precisely capturing articles’ content. For "The Hill", in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles’ words that is about 40% smaller than the total population of articles words for those three years. In the case of "NewsMax", the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words. To conclude, in a data analysis of 27 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 and Figure 2 of main manuscript for supporting evidence).</p>
COGNITIVE METHODS: THE LEVEL ANALYSIS OF DISCOURSE MARKERS
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.