Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

27 results for “News articles”

Learn how ShareScore rates datasets ↗
zenodo52/100

Multilingual news article similarity dataset

<p>This dataset contains the extended version of the authors' earlier work:&nbsp;<a href="../records/6507872">https://zenodo.org/records/6507872,</a> where pairs of news articles drawn from the first half of 2020 are annotated for seven aspects of similarity in the original version as well as an additional FRAME aspect:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> <li><strong>FRAME</strong> Do the articles have similar framing and express similar opinions?</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo52/100

SemEval-2022 Task 8: Multilingual news article similarity

<p>This dataset contains pairs of news articles drawn from the first half of 2020 and annotated for seven aspects of similarity:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong>&nbsp;Do the articles have similar tones?</li> </ul> <p>Further details are provided in</p> <blockquote> <p>Chen et al. (2022). SemEval-2022 Task 8: Multilingual news article similarity. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022).&nbsp;<a href="https://aclanthology.org/2022.semeval-1.155/">https://aclanthology.org/2022.semeval-1.155/</a></p> </blockquote> <p>The data in this repository includes pairs of URLs and annotations. The text of webpages is generally&nbsp;via the Internet Archive in this special collection: https://archive.org/details/2020-multilingual-news-article-similarity . A script to download and process the webpages is available at&nbsp;https://github.com/euagendas/semeval_8_2022_ia_downloader .&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

BreXLiMe: A Semantically Enriched Dataset With News Articles, Micro-Posts, and TV Shows Related to the Brexit

<p>We provide a <strong>large data set of media content metadata</strong> from various media sources (including online news sites, social media, and live-TV) in three languages (<strong>English, German, and Spanish</strong>). Overall, the data set contains rich metadata for about <strong>240 thousand news articles, 12 million micro-posts, and 900 TV shows</strong>. All media content information has been semantically enriched with annotations of both entities and categories from DBpedia.</p> <p>The data can be used as a valuable data basis for applications and studies of various disciplines (e.g., social studies, political science, and humanities) on the case of Brexit, particularly on the <strong>media landscape before the Brexit referendum held on June 23, 2016</strong>.</p> <p>We provide the data set in the RDF serialization format Turtle (.ttl) as well as in XML.</p> <p>If you use our data set, please <strong>cite</strong> it as follows:</p> <pre><code>Lei Zhang, Maribel Acosta, Michael Färber, Steffen Thoma and Achim Rettinger. "BreXearch: Exploring Brexit Data Using Cross-Lingual and Cross-Media Semantic Search". In: Proceedings of the ISWC 2017 Posters &amp; Demonstrations Track within the 16th International Semantic Web Conference (ISWC 2017). Vienna, Austria, 2017.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

MIDAS hand-annotated news articles

<p>This dataset was produced in 2020 from the data collected throughout 2019 for the development of the MIDAS project (http://www.midasproject.eu/)</p> <p>The data is distributed throughout 5 topics:</p> <p>- EUS: Childhood Obesity (UC Basque Country)<br> - FIN: Mental Health (UC Finland)<br> - IRE: Diabetes (UC Ireland)<br> - NIR: Children in Care (UC Northern Ireland)<br> - INF: Infectious Diseases including Coronavirus (UC Influenzanet)</p> <p>The available data comes in 3 kinds and file formats:</p> <p>TXT - the source of news including ID, title and body of text<br> CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br> JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p> <p>The CSV files with name starting in &quot;f1_&quot;, &quot;pr_&quot;, &quot;re_&quot; are the results of the F1/Precision/Recall evaluation for each of the cases.</p> <p>## AUTHORS</p> <p>Joao Pita Costa, Anthony Staines, Jarmo P&auml;&auml;kk&ouml;nen, Jenni Konttila, Joseba Bidaurrazaga, Oihana Belar, Christine Henderson</p> <p>## ACKNOWLEDGMENTS</p> <p>This work was supported by the European Commission H2020 project MIDAS (G.A. nr. 727721).&nbsp;</p> <p><br> ## LICENSE</p> <p>This dataset is licensed over Creative Commons.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Cropped News Article from University of Potsdam News Site

<p>This data was cropped from the <a href="https://www.uni-potsdam.de/de/nachrichten">University of Potsdam news website</a>. The annotated labels are made by the writers of the articles.</p>

opencc-by-sa-4.0Nov 2022View details →
zenodo40/100

Rare Diseases hand-annotated news articles and research articles

<p>This dataset was produced in 2023 from the data collected throughout 2022 from MEDLINE (scientific articles) and from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p><p>The data is distributed across 16 diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p><p>The available data comes in 2 kinds and file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br>JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p><p>The CSV files with name starting in "f1_", "pr_", "re_" are the results of the F1/Precision/Recall evaluation for each of the cases.</p><p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdolšek Draksler (domain expert) &nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

FaCov Dataset: COVID-19 Viral News and Rumors Fact-Check Articles Dataset

<p>The data were collected by web-scraping pages from the websites collected earlier, using the <a href="https://webscraper.io/">Web Scraper browser extension</a>.</p> <p>More specifically, the sections of these websites that dealt exclusively with COVID-19 related content were scraped. In cases where the website did not have such a specified section, the search functionality within the website was used to query terms related to COVID-19 and the articles in the search results were scraped. Also in some cases, all articles were scraped and those unrelated to COVID-19 were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV</p> <p>The following information was extracted along with the articles:</p> <p>Title of the fact check article</p> <p>URL of the fact check article</p> <p>Claim being discussed in the article (if available)</p> <p>Summary of the fact check article (if available)</p> <p>Content of the fact check article&bull; Label assigned by the article to the claim</p> <p>Author of the fact check article (if available)</p> <p>Date of publication of the article (if available)</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

News headlines of BBC articles published by @BBCBreaking twitter account

<p>The dataset consists of a list of news articles headlines retrieved from tweets published by @BBCBreaking profile in specific years (2012, 2015, 2017, 2019 and 2022).</p> <p>The dataset is in&nbsp;<code>.csv</code>&nbsp;format and is organised as follows:</p> <ul> <li>Columns: <ul> <li>ID (tweet ID)</li> <li>created_at (tweet publication&#39;s date)</li> <li>url (url of the news article attached to the tweet)</li> <li>Titles (news headline)</li> </ul> </li> <li>Rows: Each row contains a single news article headline sorted by date of publication (created_at). Total number of entries: 7213.</li> </ul> <p>For more details about data collection refer to <a href="https://github.com/caiocmello/news-mood">Github</a>.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022

<p>IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022</p> <p>Reference paper</p> <p>Madodonga, A., Marivate, V., &amp; Adendorff, M. (2023). Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati.&nbsp;<em>Journal of the Digital Humanities Association of Southern Africa</em>,&nbsp;<em>4</em>(01). https://doi.org/10.55492/dhasa.v4i01.4449</p> <p>&nbsp;</p> <p>&gt;&nbsp;@article{Madodonga_Marivate_Adendorff_2023, title={Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati}, volume={4}, url={https://upjournals.up.ac.za/index.php/dhasa/article/view/4449}, DOI={10.55492/dhasa.v4i01.4449}, author={Madodonga, Andani and Marivate, Vukosi and Adendorff, Matthew}, year={2023}, month={Jan.} }</p>

opencc-by-sa-4.0Oct 2022View details →
zenodo40/100

A Scientific Journal List at Japanese News Articles

<p><strong>Abstract</strong> (our paper)</p> <p>In Japanese scientific news articles, although the research results are described clearly, the article&#39;s sources tend to be uncited. This makes it difficult for readers to know the details of the research. In this paper, we address the task of extracting journal names from Japanese scientific news articles. We hypothesize that a journal name is likely to occur in a specific context. To support the hypothesis, we construct a character-based method and extract journal names using this method. This method only uses the left and right context features of journal names. The results of the journal name extractions suggest that the distribution hypothesis plays an important role in identifying the journal names.</p> <p><strong>Data</strong></p> <p>list.txt.gz:<br> The first column is the extraction text by our method (journal name), the second column is the cleaned text, the third column is the news date, and the fourth column is the news URL.</p> <p><strong>Publication</strong></p> <p>This data set is part of our experimental results. If you make use of this data set, please cite:</p> <ul> <li>Masato Kikuchi, Kento Kawakami, Mitsuo Yoshida, Kyoji Umemura. <a href="https://doi.org/10.14923/transinfj.2018DEP0007">Conservative Direct Estimation for Likelihood Ratios Based on Observed Frequencies</a>. <em>The IEICE Transactions on Information and Systems (Japanese edition)</em>. vol.J102-D, no.4, pp.289-301, 2019.</li> <li>Masato Kikuchi, Mitsuo Yoshida, Kyoji Umemura. <a href="http://www.apsipa.org/proceedings/2018/pdfs/0000143.pdf">Journal Name Extraction from Japanese Scientific News Articles</a>. <em>Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference 2018</em>. pp.143-148, 2018. [<a href="https://doi.org/10.23919/APSIPA.2018.8659765">DOI</a>]</li> </ul>

opencc-zeroMar 2019View details →
zenodo40/100

CommonCrawl News Articles by Political Orientation

<p><strong>Dataset description &amp; reproduction steps</strong></p> <p>The dataset includes news articles gathered from CommonCrawl for media outlets that were selected based on their political orientation. The news articles span publication dates from 2010 to 2021. For more details, please check out our <a href="https://aclanthology.org/2022.findings-emnlp.152/">Paper</a> and <a href="https://github.com/webis-de/emnlp22-social-bias-representation-accuracy">GitHub repository</a>.</p> <p>The database file containing the news articles has two main tables, <em>article_urls</em> and <em>article_contents</em>. The tables have the following columns:</p> <p><em>article_urls</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this URL entry. This column is used as primary key for the table.</li> <li><code>url</code>: The plain text URL for the news article, as found in CommonCrawl.</li> <li><code>outlet_name</code>: The name of the news outlet that published the article.</li> </ul> <p><em>article_contents</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this content entry. This column is used as primary key for the table. The key is the same key used in the <em>article_urls</em> table to allow for cross-referencing.</li> <li><code>date</code>: The automatically extracted publishing date of the article. If it was not possible to automatically extract the date, this field remains empty.</li> <li><code>content</code>: The plain text content of the article automatically extracted from the crawled HTML document.</li> <li><code>content_preprocessed</code>: The articles content split by sentences.</li> <li><code>langauge</code>: The langauge of the article, as identified by the langdetect module (as ISO 639-1 code).</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo40/100

SDG Knowledge Hub Dataset of SDG-labeled News Articles

<p>Dataset of articles published on the IISD SDG Knowledge Hub (<a href="http://sdg.iisd.org/">sdg.iisd.org</a>). The SDG Knowledge Hub is an online resource publishing news and commentary regarding the implementation of the United Nations&rsquo; 2030 Agenda for Sustainable Development and the Sustainable Development Goals (SDGs). Labels assigned by the authors and validated by SDG Knowledge Hub editors indicate which of the 17 SDGs an article addresses.</p> <p>The data set was generated for the following publications. Please consider citing the publication if you use the data.&nbsp;</p> <p>Wulff, D. U., Meier, D. S., &amp; Mata, R. (2023). Using novel data and ensemble models to improve automated labeling of Sustainable Development Goals.&nbsp;<em>arXiv preprint arXiv:2301.11353</em>.</p> <p>The data contain 9,172 articles downloaded on September 15th, 2021, and are shared with permission from the SDG Knowledge Hub.</p> <p>The comma-separated data file&nbsp;includes the following columns:</p> <p>url - URL of the article.</p> <p>title - Title of the article.</p> <p>type - Type of article. Either "News", "Policy briefs", "Guest articles", or "Generation 30".&nbsp;</p> <p>text - Text of the article.&nbsp;</p> <p>date - Publishing date of article.</p> <p>sdgs - SDG labels assigned by authors and editors.&nbsp;</p> <p>SDG-01 to&nbsp;SDG-17 - SDG indicators&nbsp;extracted from the author and editor label.&nbsp;&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles

<p>This dataset contains the files and annotations for <a href="https://propaganda.qcri.org/semeval2020-task11/index.html">SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles</a>. The task was composed by two subtasks: span identification (SI) and technique classification (TC). This dataset includes the following:</p> <ul> <li>The text files for training, development, and testing sets for both the SI and the TC tasks.</li> <li>The gold-standard files for the training sets, for both the SI and the TC task</li> </ul> <p>Our propaganda identification initiative remains active. We keep a <a href="https://propaganda.qcri.org/ptc/leaderboard.php">live leader-board</a> reporting the performance of models submited up to date.</p> <p><strong>Reference</strong></p> <p>Giovanni Da San Martino, Alberto Barr&oacute;n-Cede&ntilde;o, Henning Wachsmuth, Rostislav Petrov, and Preslav Nakov. 2020. <a href="https://propaganda.qcri.org/">Task 11: Detection of Propaganda Techniques in News Articles</a>. In Proceedings of the 14th International Workshop on Semantic Evaluation (SemEval 2020). Barcelona, Spain (2020)</p> <p>&nbsp;</p> <pre><code>@InProceedings{SemEval20-11-DaSanMartino, author = "Da San Martino, Giovanni and Barr\'{o}n-Cede\~no, Alberto and Wachsmuth, Henning and Petrov, Rostislav and Nakov, Preslav", title = "{SemEval}-2020 Task 11: {D}etection of Propaganda Techniques in News Articles", pages = "", abstract = "We describe the outcome of the SemEval 2020 Task 11 on the detection of propaganda in news articles. We present two tasks. In the first task, systems are asked to identify specific text spans in a free text where propaganda is being applied. In the second task, systems are asked to identify the propaganda technique being applied in a text span. We describe the construction of the evaluation framework (dataset and evaluation metrics) as well as the approaches explored by the different participants. ", crossref = "SemEval20" }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Hindi News Article Text Dataset

<p>The Hindi News Article Dataset (HNAD) comprises over a million meticulously curated news articles in the Hindi language, sourced from diverse online platforms. Covering a wide range of topics, including politics, economics, culture, sports, and more, it offers a comprehensive representation of the Hindi news landscape.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Rare Diseases hand-annotated news articles: Angelman, De Lange, Fragile X, Kleefstra

<p>This dataset was produced in 2023 from the data collected throughout 2023 from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p> <p>The data is distributed across 4 specific diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p> <p>The available data comes in the file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings</p> <p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdol&scaron;ek Draksler (domain expert) &nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

News article mentions of 4chan from three weeks in 2017 and posts with links to these articles on 4chan/pol/

<p>This upload includes two dataset. The first comprises articles that mentioning &quot;4chan&quot; in its post body as retreived by Nexis Uni. We identified the three weeks in which the highest spikes in terms of the amount of articles occured in 2017 (see the image mentions_nexis_uni_2017_weeks.png ) and extracted the articles that were published within these. The dataset includes the article source, the article title, and the sentences where they mention 4chan to analyse the framing. We deleted the article bodies for copyright reasons. Two columns indicate the URL to the article if it appeared online. Additionally, we added a column with URLs to archive.is links of the same articles, in case they were archived.</p> <p>The second dataset consists of posts on the imageboard 4chan/pol/ that refer to one of the aforementioned URLs.</p>

opencc-byFeb 2020View details →
zenodo32/100

Marathi News Article Text Dataset

<p>The Marathi News Article Text Dataset (MNATD) consists of meticulously curated news articles in the Marathi language, sourced from various online platforms. With a substantial collection exceeding over 6.5 Lakh articles, this dataset provides a thorough exploration of the Marathi news landscape. Encompassing diverse subjects such as Politics, Astrology , sports, and more, MNATD offers a comprehensive representation of the Marathi news domain. Researchers and enthusiasts will find this dataset valuable for delving into the linguistic and thematic richness of Marathi journalism. MNATD aims to contribute to a nuanced understanding of the Marathi news landscape and serve as a valuable resource for linguistic analysis and research.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Data for manuscript: "Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets"

<p>This data set contains material for the purpose of scientific reproducibility of the accompanying manuscript &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot;.</p> <p>Note that this data set is distributed with an Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) License. NonCommercial means you&nbsp;may not use the material for commercial purposes. NoDerivatives means if you remix, transform, or build upon the material, you may not distribute the modified material. Attribution means you must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. See attached license terms for details.</p> <p>The work &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot; describes an analysis of political associations in 27 million diachronic (1975-2019) news and opinion articles from 47 news media outlets popular in the United States. We use embedding models trained on individual outlets content to quantify outlet-specific latent associations between positive/negative sentiment words and terms loaded with political connotations such as those describing political orientation, party affiliation, names of influential politicians and ideologically aligned public figures.&nbsp;</p> <p>News and opinion articles from the outlets listed in Figure 3 are available in the outlet&#39;s online domains and/or public cache repositories such as Google cache, The Internet Wayback Machine [31] and Common Crawl [32]. This work has not analyzed video or audio content of news media organizations, except when the outlet explicitly provides a transcript of such content in article form.<br> The temporal coverage of articles from different news outlets is not uniform. For most media organizations, news articles availability in their online domains or Internet cache backups becomes sparse as a function of articles&rsquo; age. This is not the case for some news outlets, where availability of news articles goes back to the 1970s. The Supplementary Material (SM) illustrates the time ranges of article data analyzed based on news outlets articles online availability.</p> <p>Textual content included in our analysis is circumscribed to the articles&rsquo; headlines and main text and does not include other article elements such as figure captions. Targeted textual content was located in HTML raw data using outlet specific XPath expressions. Tokens were lowercased prior to estimating embedding models. Markup language tags, URLs, nonalphanumeric characters, punctuation, digits, 330 common stop words and multiple spaces were removed prior to estimating word embeddings models.<br> All the analysis scripts and the diachronic word embedding models built from each of the 47 news media outlets analyzed in this work are available in this repository.</p> <p>For the purpose of reproducibility, we also provide in the above repository the articles&rsquo; text used to train the news outlets embedding models with the caveat that outlets articles not accessible without a subscription have been excluded. Also, for the included articles, stop words have been removed and the remaining words have been randomly scrambled within a sliding window of size 10 to render the articles incomprehensible to a human reader. These steps have been taken to not infringe articles copyright. These preprocessing steps have only minor impact on Continuous Bag of Words (CBOW) word2vec and the results reported in this work are similar when using the scrambled articles text to train outlet-specific embedding models.</p> <p>We derived outlet-specific word embedding models at every five-year time intervals within the 1975-2019 time range. The gensim [33] implementation of word2vec was used to train the embedding models. The continuous bag of words (CBOW) architecture performed slightly better than the Skip-Gram architecture in commonly used validation metrics so it was used for all subsequent analysis.&nbsp;</p> <p>For training the word embedding models, the following parameters were used: vector dimensions=300, window size=10, negative sampling=10, down sampling frequent words = 0.0001, minimum frequency count of 5 (only terms that appear more than 5 times in the corpus were included into the word embedding model vocabulary), number of training iterations (epochs) through the corpus=5. The exponent used to shape the negative sampling distribution was the default 0.75.&nbsp;</p> <p>Outlet-specific embedding models performance across a range of commonly used semantic, syntactic and analogy tasks was similar to popular pre-trained embedding models trained on corpora such as Twitter or Google books on similarity, association and word analogy tasks, see Supplemeentary Material of the manuscript for detailed validation tests results.</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Jul 2021View details →
zenodo32/100

Translating Non-Binary Coming-Out Reports: Gender-Fair Language Strategies and Use in News Articles

<p>Materials used for the &quot;Translating Non-Binary Coming-Out Reports: Gender-Fair Language Strategies and Use in News Articles&quot; paper appeared in the Journal of Specialised Translation, issue 40.&nbsp;</p> <p>The .xlsx files contain&nbsp;links to the analysed texts as well as a first annotation that was refined in SPSS (see .sav files).&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Italian and US news articles on COVID-19

<p>Sample of Italian and US news &nbsp;1/1/2020-20/10/2021obtqained&nbsp;searching: convalescent plasma, hydroxychloroquine, ivermectin, lockdown, mask, vaccine and vitamin D combined in a Boolean AND search with &ldquo;covid AND&nbsp; (published OR publication OR journal)&rdquo;. (In Italian,: plasma convalescenti, idrossiclorochina, ivermectina, lockdown, mascherine, vaccino, vitamina D AND &ldquo;covid AND (ricerca OR pubblicato OR pubblicazione)</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record