Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
297
datasets available to search
ShareScore release 0.9.0
Dataset results
297 results for “News”
Challenge or Empower: Revisiting Argumentation Quality in a News Editorial Corpus
<p>The Webis-Editorial-Quality-18 corpus is a novel corpus with 1000 news editorials. The aim of this Corpus is to study a new notion for news editorials quality. It contains the quality assessments of 1000 news editorials, each annotated by three liberals and three conservatives. The annotators also reported free-text reasons for the effects they observed.</p>
A Scientific Journal List at Japanese News Articles
<p><strong>Abstract</strong> (our paper)</p> <p>In Japanese scientific news articles, although the research results are described clearly, the article's sources tend to be uncited. This makes it difficult for readers to know the details of the research. In this paper, we address the task of extracting journal names from Japanese scientific news articles. We hypothesize that a journal name is likely to occur in a specific context. To support the hypothesis, we construct a character-based method and extract journal names using this method. This method only uses the left and right context features of journal names. The results of the journal name extractions suggest that the distribution hypothesis plays an important role in identifying the journal names.</p> <p><strong>Data</strong></p> <p>list.txt.gz:<br> The first column is the extraction text by our method (journal name), the second column is the cleaned text, the third column is the news date, and the fourth column is the news URL.</p> <p><strong>Publication</strong></p> <p>This data set is part of our experimental results. If you make use of this data set, please cite:</p> <ul> <li>Masato Kikuchi, Kento Kawakami, Mitsuo Yoshida, Kyoji Umemura. <a href="https://doi.org/10.14923/transinfj.2018DEP0007">Conservative Direct Estimation for Likelihood Ratios Based on Observed Frequencies</a>. <em>The IEICE Transactions on Information and Systems (Japanese edition)</em>. vol.J102-D, no.4, pp.289-301, 2019.</li> <li>Masato Kikuchi, Mitsuo Yoshida, Kyoji Umemura. <a href="http://www.apsipa.org/proceedings/2018/pdfs/0000143.pdf">Journal Name Extraction from Japanese Scientific News Articles</a>. <em>Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference 2018</em>. pp.143-148, 2018. [<a href="https://doi.org/10.23919/APSIPA.2018.8659765">DOI</a>]</li> </ul>
RIA News corpus 2001-2018
<p>Files lema_year.pkl contain text corpus build from the articles published by RIA News. Each files contains sentences from the articles of the selected year. Words are lematized. This data is used for training word2vec models.</p> <p>File url_index.pkl has index of lematized words to hte URL of article containing selected word.</p>
Datasets from Costa Rican news sources for fake news detection
<p>Today, technology has changed the way information is propagated and how the message is received. The interpretation of the news may have different angles depending on the source of origin. Because of this, there has been an increase in misinformation, in the way of influencing public opinion and in how we perceive or estimate reality.<br> The objective of this beta dataset is to be used for the evaluation of data mining models that allow the classification of true or potentially fake news that are generated by Costa Rican news sites only. This is intended to assess the level of reliability of the models and extend the scope of this research in future work.</p> <p>The dataset has been pre-processed (standarized using lower cases, lemmatized and removed any possible noise from it) and analyzed using LIWC dictionaries. One version has the news text in Spanish and was processed using LIWC2007 dictionary in Spanish. The second version was processed using LIWC2015 English dictionary and has the news text in English. The reason to having two versions is to be able to test using the newer LIWC dictionary which includes more Summary Language Variables that the Spanish version doesn't have and analyze how this and other variables can contribute to different results when creating models. </p> <p>The file "DescripcionVariables" provides a description of all variables used.</p> <p> </p>
Slovenian Coronavirus News Comments Corpus News-CommSLO
<p>A corpus of readers' news comments posted below news articles on the topic of the covid-19 pandemic, published in major Slovenian daily newspapers and news portals in the six-month early pandemic period (March 2020 to September 2020).</p> <div>The corpus is designed to facilitate research on crisis discourses, crisis communication, as well as pandemic-time linguistic innovation. It is available in plain text version and XML with full metadata. The corpus complements a separate corpus of news articles Slovenian Coronavirus Corpus NewsSLO. Parallel versions from Croatia and Serbia are also available.</div> <div> </div> <div>The project leading to this publication has received funding from the European Union’s Horizon 2020 research and innovation programme under the <a href="https://cordis.europa.eu/programme/id/H2020-EU.4./en">H2020-EU.4. - SPREADING EXCELLENCE AND WIDENING PARTICIPATION </a>programme Widening fellowships grant agreement No 101038047.</div>
Serbian Coronavirus News Comments Corpus News-CommSR
<p>A corpus of readers' news comments posted below news articles on the topic of the covid-19 pandemic, published in major Serbian daily newspapers and news portals in the six-month early pandemic period (March 2020 to September 2020).</p> <div>The corpus is designed to facilitate research on crisis discourses, crisis communication, as well as pandemic-time linguistic innovation. It is available in plain text version and XML with full metadata. The corpus complements a separate corpus of news articles Serbian Coronavirus Corpus NewsSR. Parallel versions from Croatia and Slovenia are also available.</div> <div> </div> <div>The project leading to this publication has received funding from the European Union’s Horizon 2020 research and innovation programme under the <a href="https://cordis.europa.eu/programme/id/H2020-EU.4./en">H2020-EU.4. - SPREADING EXCELLENCE AND WIDENING PARTICIPATION </a>programme Widening fellowships grant agreement No 101038047.</div>
Robberies for cigarettes news reports 2009-2018 data
<p>Dataset used to analyze New Zealand news reports of robberies of stores for tobacco during 2009-2018. </p>
Reporting uncertified science in the news media during the Covid-19 pandemic
<p>Preprints have established a stable position in the dissemination of scientific findings. This position has been reinforced by the Covid-19 pandemic, which required the rapid dissemination of new scientific information. However, in most cases, preprints have not undergone peer review and, as a consequence, lack the scientific rigor of other scientific publications such as journal articles. This presents a challenge for journalists who are tasked with keeping the public informed about the latest scientific developments in the context of great uncertainty during a global pandemic. Having to report on rapidly changing circumstances under increasing pressure from social media while also having to compete for attention in a saturated media landscape might place strain on the adherence to journalistic norms. This does not only have implications in terms of the case-by-case accuracy of reporting, but on the public perception of science at large.</p> <p>This paper investigates the reporting of scientific information from pre-preprints based on a sample of 2,877 online news articles related to Covid-19 in the South African news media. Our results show that despite the publication of guidelines for reporting on preprints in the media, there is still a way to go regarding the judicious use of scientific information from preprints by journalists.</p> <p>The dataset includes the table of the 2,877 online news articles as well as the analysis done (as separate tables).</p> <p>Fields in the dataset: Full Text, Headline, ID, Publication Date, Site Name,URL,Syndicated, Covid-related scientific research findings,<br> Preprint, Title, Authors,DOI, Archive / journal, URL</p>
A study on real graphs of fake news spreading on Twitter
<p><strong>*** Fake News on Twitter</strong> <strong>***</strong></p> <p>These 5 datasets are the results of an empirical study on the spreading process of newly fake news on Twitter. Particularly, we have focused on those fake news which have given rise to a truth spreading simultaneously against them. The story of each fake news is as follow:</p> <p>1- FN1: A Muslim waitress refused to seat a church group at a restaurant, claiming "religious freedom" allowed her to do so.</p> <p>2- FN2: Actor Denzel Washington said electing President Trump saved the U.S. from becoming an "Orwellian police state."</p> <p>3- FN3: Joy Behar of "The View" sent a crass tweet about a fatal fire in Trump Tower.</p> <p>4- FN4: The animated children's program 'VeggieTales' introduced a cannabis character in August 2018.</p> <p>5- FN5: In September 2018, the University of Alabama football program ended its uniform contract with Nike, in response to Nike's endorsement deal with Colin Kaepernick.</p> <p>The data collection has been done in two stages that each provided a new dataset: 1- attaining Dataset of Diffusion (DD) that includes information of fake news/truth tweets and retweets 2- Query of neighbors for spreaders of tweets that provides us with Dataset of Graph (DG). </p> <p><strong>DD </strong></p> <p>DD for each fake news story is an excel file, named FNx_DD where x is the number of fake news, and has the following structure:</p> <p>The structure of excel files for each dataset is as follow:</p> <ul> <li>Each row belongs to one captured tweet/retweet related to the rumor, and each column of the dataset presents a specific information about the tweet/retweet. These columns from left to right present the following information about the tweet/retweet: </li> <li>User ID (user who has posted the current tweet/retweet)</li> <li>The number of published tweet/retweet by the user at the time of posting the current tweet/retweet</li> <li>Language of the tweet/retweet</li> <li>Number of followers </li> <li>Number of followings (friends)</li> <li>Date and time of posting the current tweet/retweet</li> <li>Number of like (favorite) the current tweet had been acquired before crawling it</li> <li>Number of times the current tweet had been retweeted before crawling it</li> <li>Is there any other tweet inside of the current tweet/retweet (for example this happens when the current tweet is a quote or reply or retweet)</li> <li>The source (OS) of device by which the current tweet/retweet was posted</li> <li>Tweet/Retweet ID</li> <li>Retweet ID (if the post is a retweet then this feature gives the ID of the tweet that is retweeted by the current post)</li> <li>Quote ID (if the post is a quote then this feature gives the ID of the tweet that is quoted by the current post)</li> <li>Reply ID (if the post is a reply then this feature gives the ID of the tweet that is replied by the current post)</li> <li>Frequency of tweet occurrences which means the number of times the current tweet is repeated in the dataset (for example the number of times that a tweet exists in the dataset in the form of retweet posted by others)</li> <li>State of the tweet which can be one of the following forms (achieved by an agreement between the annotators):</li> </ul> <ul> <li>r : The tweet/retweet is a fake news post</li> <li>a : The tweet/retweet is a truth post</li> <li>q : The tweet/retweet is a question about the fake news, however neither confirm nor deny it</li> <li>n : The tweet/retweet is not related to the fake news (even though it contains the queries related to the rumor, but does not refer to the given fake news)</li> </ul> <p> </p> <p><strong>DG</strong></p> <p>DG for each fake news contains two files:</p> <ul> <li>A file in graph format (.graph) which includes the information of graph such as who is linked to whom. (This file named FNx_DG.graph, where x is the number of fake news)</li> <li>A file in Jsonl format (.jsonl) which includes the real user IDs of nodes in the graph file. (This file named FNx_Labels.jsonl, where x is the number of fake news)</li> </ul> <p>Because in the graph file, the label of each node is the number of its entrance in the graph. For example if node with user ID 12345637 be the first node which has been entered into the graph file then its label in the graph is 0 and its real ID (12345637) would be at the row number 1 (because the row number 0 belongs to column labels) in the jsonl file and so on other node IDs would be at the next rows of the file (each row corresponds to 1 user id). Therefore, if we want to know for example what the user id of node 200 (labeled 200 in the graph) is, then in jsonl file we should look at row number 202.</p> <p> </p> <p>The user IDs of spreaders in DG (those who have had a post in DD) would be available in DD to get extra information about them and their tweet/retweet. The other user IDs in DG are the neighbors of these spreaders and might not exist in DD.</p>
Fake News Text Collections
<p>The description are in: <a href="https://github.com/GoloMarcos/FKTC">https://github.com/GoloMarcos/FKTC</a></p>
Financial News dataset for text mining
<p>please cite this dataset by :</p> <p>Nicolas Turenne, Ziwei Chen, Guitao Fan, Jianlong Li, Yiwen Li, Siyuan Wang, Jiaqi Zhou (2021) Mining an English-Chinese parallel Corpus of Financial News, BNU HKBU UIC, technical report</p> <p> </p> <p>The dataset comes from Financial Times news website (https://www.ft.com/)</p> <p>news are written in both languages Chinese and English.</p> <p><a href="https://zenodo.org/api/files/74136928-3c77-4388-aa47-b1079efa1650/FTIE.zip?versionId=440c065e-1528-4f01-800d-541a0611ea7f">FTIE.zip</a> contains all documents in a file individually</p> <p><a href="https://zenodo.org/api/files/74136928-3c77-4388-aa47-b1079efa1650/FT-en-zh.rar">FT-en-zh.rar</a> contains all documents in one file</p> <p>Below is a sample document in the dataset defined by these fields and syntax : </p> <p>id;time;english_title;chinese_title;integer;english_body;chinese_body</p> <p> </p> <p>1021892;2008-09-10T00:00:00Z;FLAW IN TWIN TOWERS REVEALED;科学家发现纽约双子塔倒塌的根本原因;1;Scientists have discovered the fundamental reason the Twin Towers collapsed on September 11 2001. The steel used in the buildings softened fatally at 500?C – far below its melting point – as a result of a magnetic change in the metal. @ The finding, announced at the BA Festival of Science in Liverpool yesterday, should lead to a new generation of steels capable of retaining strength at much higher temperatures.;科学家发现了纽约世贸双子大厦(Twin Towers)在2001年9月11日倒塌的根本原因。由于磁性变化,大厦使用的钢在500摄氏度——远远低于其熔点——时变软,从而产生致命后果。 @ 这一发现在昨日利物浦举行的BA科学节(BA Festival of Science)上公布。这应会推动能够在更高温度下保持强度的新一代钢铁的问世。<br> </p> <p>The dataset contains 60,473 bilingual documents.</p> <p>Time range is from 2007 and 2020. </p> <p>This dataset has been used for parallel bilingual news mining in Finance domain.</p>
Supporting Information (software and data) for: Client-side energy and GHGs assessment of advertising and tracking in the news websites
<p>This is the open data and free/libre and open source software repository for the article "Client-side energy and GHGs assessment of advertising and tracking in the news websites" by Fabio Pesari, Giovanni Lagioia, Annarita Paiano.</p>
CommonCrawl News Articles by Political Orientation
<p><strong>Dataset description & reproduction steps</strong></p> <p>The dataset includes news articles gathered from CommonCrawl for media outlets that were selected based on their political orientation. The news articles span publication dates from 2010 to 2021. For more details, please check out our <a href="https://aclanthology.org/2022.findings-emnlp.152/">Paper</a> and <a href="https://github.com/webis-de/emnlp22-social-bias-representation-accuracy">GitHub repository</a>.</p> <p>The database file containing the news articles has two main tables, <em>article_urls</em> and <em>article_contents</em>. The tables have the following columns:</p> <p><em>article_urls</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this URL entry. This column is used as primary key for the table.</li> <li><code>url</code>: The plain text URL for the news article, as found in CommonCrawl.</li> <li><code>outlet_name</code>: The name of the news outlet that published the article.</li> </ul> <p><em>article_contents</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this content entry. This column is used as primary key for the table. The key is the same key used in the <em>article_urls</em> table to allow for cross-referencing.</li> <li><code>date</code>: The automatically extracted publishing date of the article. If it was not possible to automatically extract the date, this field remains empty.</li> <li><code>content</code>: The plain text content of the article automatically extracted from the crawled HTML document.</li> <li><code>content_preprocessed</code>: The articles content split by sentences.</li> <li><code>langauge</code>: The langauge of the article, as identified by the langdetect module (as ISO 639-1 code).</li> </ul>
SDG Knowledge Hub Dataset of SDG-labeled News Articles
<p>Dataset of articles published on the IISD SDG Knowledge Hub (<a href="http://sdg.iisd.org/">sdg.iisd.org</a>). The SDG Knowledge Hub is an online resource publishing news and commentary regarding the implementation of the United Nations’ 2030 Agenda for Sustainable Development and the Sustainable Development Goals (SDGs). Labels assigned by the authors and validated by SDG Knowledge Hub editors indicate which of the 17 SDGs an article addresses.</p> <p>The data set was generated for the following publications. Please consider citing the publication if you use the data. </p> <p>Wulff, D. U., Meier, D. S., & Mata, R. (2023). Using novel data and ensemble models to improve automated labeling of Sustainable Development Goals. <em>arXiv preprint arXiv:2301.11353</em>.</p> <p>The data contain 9,172 articles downloaded on September 15th, 2021, and are shared with permission from the SDG Knowledge Hub.</p> <p>The comma-separated data file includes the following columns:</p> <p>url - URL of the article.</p> <p>title - Title of the article.</p> <p>type - Type of article. Either "News", "Policy briefs", "Guest articles", or "Generation 30". </p> <p>text - Text of the article. </p> <p>date - Publishing date of article.</p> <p>sdgs - SDG labels assigned by authors and editors. </p> <p>SDG-01 to SDG-17 - SDG indicators extracted from the author and editor label. </p>
Migration Reframed - Multilingual stance annotated Twitter news replies on migration in Europe in the context of the Ukrainian crisis
<p><em>The corresponding paper for this dataset "Migration Reframed? A multilingual analysis on the stance shift in Europe during the Ukrainian crisis" has been published in the ACM Web Conference 2023 (WWW'23), and can be accessed here: </em><a href="https://doi.org/10.1145/3543507.3583442">https://doi.org/10.1145/3543507.3583442</a> <em>. Please cite this when using the dataset.</em></p> <p>Twitter dataset of European news and replies to investigate public stance on refugees/migrants around the Ukrainian Crisis.</p> <p>September 2021 to August 2022.</p> <p>Countries:</p> <ul> <li>France</li> <li>Germany</li> <li>Italy</li> <li>Poland</li> <li>Spain</li> </ul> <p>Dataset contains:</p> <ul> <li>Usernames of news outlet accounts on Twitter</li> <li>Tweet IDs of these news accounts during the mentioned period filtered for the migration topic + respective replies from the public</li> <li>Tweet IDs of stance annotated replies</li> <li>8,242 tweet/reply pairs labeled with the stance (positive / negative / neutral) on migrants/refugees (on request)</li> </ul> <table> <caption>Dataset overview by the numbers</caption> <thead> <tr> <th scope="col">Country</th> <th scope="col">News Outlets</th> <th scope="col">News Tweets</th> <th scope="col">Replies</th> <th scope="col">Stance Annotated</th> </tr> </thead> <tbody> <tr> <td>France</td> <td>37</td> <td>2,020</td> <td>32,839</td> <td>500</td> </tr> <tr> <td>Germany</td> <td>72</td> <td>3,752</td> <td>55,317</td> <td>500</td> </tr> <tr> <td>Italy</td> <td>21</td> <td>1,305</td> <td>9,892</td> <td>500</td> </tr> <tr> <td>Poland</td> <td>35</td> <td>3,138</td> <td>27,892</td> <td>6,242</td> </tr> <tr> <td>Spain</td> <td>35</td> <td>1,263</td> <td>20,771</td> <td>500</td> </tr> <tr> <td> </td> <td>200</td> <td>11,478</td> <td>146,711</td> <td>8,242</td> </tr> </tbody> </table> <p>Please note: Stance labels are not included and are only available on request.</p>
Teenager Safe News Dataset
<p>The provided News dataset consists of 12,800 news titles that have been categorized into two groups: safe and unsafe, specifically with regard to their appropriateness for teenagers and kids. This dataset likely aims to facilitate the development of models or algorithms that can classify news titles based on their suitability for young audiences.</p> <p>Categorizing news titles as safe or unsafe for teenagers and kids suggests a concern for the content's age-appropriateness and potential impact on young readers. The term "safe" in this context implies that the news title contains content that is considered suitable, non-offensive, and aligned with ethical guidelines for teenagers and kids. Conversely, the term "unsafe" suggests that the news title may include content that could be harmful, inappropriate, or unsuitable for young audiences.</p>
Arabic news credibility on Twitter using sentiment analysis and ensemble learning
<p>Arabic news credibility on Twitter using sentiment analysis and ensemble learning.</p> <p> </p> <p>WHAT IS IT?</p> <p>-----------</p> <p>an Arabic news credibility model on Twitter using sentiment analysis and ensemble learning.</p> <p>Here we include the Collected dataset and the source code of the proposed model written in Python language and using Keras library with Tensorflow backend.</p> <p> </p> <p>Required Packages</p> <p>------------------</p> <ol> <li>Keras (<a href="https://keras.io/">https://keras.io/</a>).</li> <li>Scikit-learn (<a href="http://scikit-learn.org/)">http://scikit-learn.org/)</a></li> <li>Imnlearn (<a href="https://imbalanced-learn.org/stable/">imbalanced-learn documentation — Version 0.10.1</a>)</li> </ol> <p> </p> <p> </p> <p>To Run the model</p> <p>---------------</p> <p>One data file is required to run the model which are:</p> <p> </p> <ol> <li>The data that were used are the collected dataset in the file, set the path of the required data file in the code.</li> </ol> <p> </p> <p>The dataset</p> <p>---------------</p> <ol> <li>There are the dataset file with all features, you can choose the features that you need and apply it on the model.</li> <li>There are a description file that describe each feature in the news credibility dataset</li> <li>The file Tweet_ID contains the list of tweets id in the dataset.</li> <li>The annotated replies based on credibility is provided.</li> </ol> <p> </p> <p> </p> <p> </p> <p> </p> <p>CONTACTS</p> <p>--------</p> <ul> <li>If you want to report bugs or have general queries email to <duha_atif@yahoo.com></li> </ul> <p> </p> <p> </p> <p> </p>
Enabling Roll-up and Drill-down Operations in News Exploration
<p>This dataset contains 200k news articles. News entities are linked to DBPedia via named entity linking.</p>
Dataset and Models for Detection of News Agency Releases in Historical Newspapers
<p>This record contains the annotated datasets and models used and produced for the work reported in the Master Thesis "<em>Where Did the News come from? Detection of News Agency Releases in Historical Newspapers</em> " (<a href="https://infoscience.epfl.ch/record/305129?&ln=en">link</a>).</p> <p>Please cite this report if you are using the models/datasets or find it relevant to your research:</p> <pre><code>@article{Marxen:305129, title = {Where Did the News Come From? Detection of News Agency Releases in Historical Newspapers}, author = {Marxen, Lea}, pages = {114p}, year = {2023}, url = {http://infoscience.epfl.ch/record/305129}, }</code></pre> <p><br> <strong>1. DATA</strong></p> <p>The <strong>newsagency-dataset</strong> contains historical newspaper articles with annotations of news agency mentions. The articles are divided into French (fr) and German (de) subsets and a train, dev and test set respectively. The data is annotated at token-level in the CoNLL format with IOB tagging format.</p> <p>The distribution of articles in the different sets is as follows:</p> <table> <caption>Dataset Statistics</caption> <thead> <tr> <th scope="row"> </th> <th scope="col">Lg.</th> <th scope="col">Docs</th> <th scope="col">Agency Mentions</th> </tr> </thead> <tbody> <tr> <th scope="row">Train</th> <td>de</td> <td>333</td> <td>493</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>903</td> <td>1,122</td> </tr> <tr> <th scope="row">Dev</th> <td>de</td> <td>32</td> <td>26</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>110</td> <td>114</td> </tr> <tr> <th scope="row">Test</th> <td>de</td> <td>32</td> <td>58</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>120</td> <td>163</td> </tr> </tbody> </table> <p>Due to an error, there are seven duplicated articles in the French test set (<em>article IDs: courriergdl-1847-10-02-a-i0002, courriergdl-1852-02-14-a-i0002, courriergdl-1860-10-31-a-i0016, courriergdl-1864-12-15-a-i0005, lunion-1860-11-27-a-i0004, lunion-1865-02-05-a-i0012, lunion-1866-02-16-a-i0009</em>).</p> <p> </p> <p><strong>2. MODELS</strong></p> <p>The two agency detection and classification models used for the inference on the <em><a href="https://impresso-project.ch/">impresso</a> </em>Corpus are released as well:</p> <ul> <li><strong>newsagency-model-de</strong>: based on <a href="https://www.deepset.ai/german-bert">German BERT</a> (with maximum sequence length 128), fine-tuned with the German training set of the newsagency-dataset</li> <li><strong>newsagency-model-fr</strong>: based on <a href="https://huggingface.co/dbmdz/bert-base-french-europeana-cased">French Europeana BERT</a> (with maximum sequence length 128), fine-tuned with the French training set of the newsagency-dataset</li> </ul> <p>The models perform multitask classification with two prediction heads, one for token-level agency entity classification and one for sentence-level (<em>has_agency: yes/no</em>). They can be run with TorchServe, for details see the <a href="https://github.com/impresso/newsagency-classification/tree/main/lib/bert_classification">newsagency-classification</a> repository.</p> <p> </p> <p>Please refer to the report for further information or contact us.</p> <p> </p> <p><strong>3. CODE</strong></p> <p><a href="https://github.com/impresso/newsagency-classification">https://github.com/impresso/newsagency-classification</a></p> <p> </p> <p><strong>4. CONTACT</strong></p> <p>Maud Ehrmann (EPFL-DHLAB)<br> Emanuela Boros (EPFL-DHLAB)</p>
Web Archive of Independent News Sites on Turkish Affairs derivatives
<p>Derivatives of the <a href="https://archive-it.org/collections/12911">Web Archive of Independent News Sites on Turkish Affairs</a> collection from the <a href="https://archive-it.org/home/IvyPlus">Ivy Plus Libraries Confederation</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a> and <a href="https://cloud.archivesunleashed.org/">Archives Unleashed Cloud</a>.</p> <p>The <strong>ivy-12911-parquet.tar.gz</strong> derivatives are in the <a href="https://parquet.apache.org/">Apache Parquet format</a>, which is a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb">this</a> notebook for examples.</p> <p><strong>Domains</strong></p> <pre><code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>domain</li> <li>count</li> </ul> <p><strong>Web Pages</strong></p> <pre><code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>url</li> <li>mime_type_web_server</li> <li>mime_type_tika</li> <li>content</li> </ul> <p><strong>Web Graph</strong></p> <pre><code class="language-java">.webgraph()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>crawl_date</li> <li>src</li> <li>dest</li> <li>anchor</li> </ul> <p><strong>Image Links</strong></p> <pre><code class="language-java">.imageLinks()</code></pre> <p>Produces a DataFrame with the following columns:</p> <ul> <li>src</li> <li>image_url</li> </ul> <p><a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"><strong>Binary Analysis</strong></a></p> <ul> <li>Audio</li> <li>Images</li> <li>PDFs</li> <li>Presentation program files</li> <li>Spreadsheets</li> <li>Text files</li> <li>Word processor files<br> </li> </ul> <p>The <strong>ivy-12911-auk.tar.gz </strong>derivatives<strong> </strong>are the <a href="https://cloud.archivesunleashed.org/derivatives">standard set of web archive derivatives</a> produced by the Archives Unleashed Cloud.</p> <ul> <li><strong>Gephi </strong>file, which can be loaded into <a href="https://gephi.org/">Gephi</a>. It will have basic characteristics already computed and a basic layout.</li> <li><strong>Raw Network</strong> file, which can also be loaded into <a href="https://gephi.org/">Gephi</a>. You will have to use that network program to lay it out yourself.</li> <li><strong>Full text</strong> file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.</li> <li><strong>Domains count</strong> file. A text file containing the frequency count of domains captured within your web archive.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.