Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10,623

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

10,623 results for “COVID”

Learn how ShareScore rates datasets ↗
zenodo44/100

COVID-19 Tweets : A dataset contaning more than 600k tweets on the novel CoronaVirus

<p>This dataset contains&nbsp;653 996&nbsp;tweets related to the Coronavirus topic and highlighted by hashtags such&nbsp;as: #COVID-19, #COVID19, #COVID, #Coronavirus, #NCoV and #Corona. The tweets&#39; crawling period started on the 27<sup>th</sup> of February and ended on the 25<sup>th</sup> of March 2020, which is spread over four weeks.&nbsp;</p> <p>The tweets were generated by 390 458 users from 133 different countries and were written in 61 languages. English being the most used language with almost 400k tweets, followed by Spanish with around 80k tweets.&nbsp;</p> <p>The data is stored in as a CSV file, where each line represents a tweet. The CSV file provides information on the following fields:</p> <ul> <li>Author: the user who posted the tweet</li> <li>Recipient: contains the name of the user in case of a reply, otherwise it would have the same value as the previous field</li> <li>Tweet: the full content of the tweet</li> <li>Hashtags: the list of hashtags present in the tweet</li> <li>Language: the language of the tweet</li> <li>Relationship: gives information on the type of the tweet, whether it is a retweet, a reply, a tweet with a mention, etc.&nbsp;</li> <li>Location: the country of the author of the tweet, which is unfortunately not always available</li> <li>Date: the publication date of the tweet</li> <li>Source: the device or platform used to send the tweet</li> </ul> <p>The dataset can as well be used to construct a social graph since it includes the relations &quot;Replies to&quot;, &quot;Retweet&quot;, &quot;MentionsInRetweet&quot; and&nbsp;&quot;Mentions&quot;.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

STROBE checklist for a set of scientific works about COVID-19

<p>STROBE checklist for a set of scientific works about COVID-19. This dataset&nbsp;is the result of the expert-based assessment carried out in&nbsp;<a href="https://arxiv.org/abs/2004.06179">arXiv:2004.06179</a>.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

CMU-MisCov19: A Novel Twitter Dataset for Characterizing COVID-19 Misinformation

<p>From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. Detection and characterization of misinformation requires an availability of annotated datasets. Most of the published COVID-19 Twitter datasets are generic, lack annotations or labels, employ automated annotations using transfer learning or semi-supervised methods, or are not specifically designed for misinformation. Annotated datasets are either only focused on &quot;fake news&quot;, are small in size, or have less diversity in terms of classes.</p> <p>Here, we present a novel Twitter misinformation dataset called <strong>&quot;CMU-MisCov19&quot;</strong> with 4573 annotated tweets over 17 themes around the COVID-19 discourse.&nbsp;We also present our annotation codebook for the different COVID-19 themes&nbsp;on Twitter, along with their descriptions and examples,&nbsp;for the community to use for collecting further annotations. Further details related to the dataset, and our analysis based on this dataset can be found at&nbsp;<a href="https://arxiv.org/abs/2008.00791">https://arxiv.org/abs/2008.00791</a>. In adherence to the Twitter&rsquo;s terms and conditions, we&nbsp;do not provide&nbsp;the full tweet JSONs but provide a &quot;.csv&quot; file with the tweet IDs so that the tweets&nbsp;can be rehydrated. We also provide the annotations, and the date of creation for each tweet for the reproduction of the results of our analyses.</p> <p><strong>Note: If for any reason, you are not able to rehydrate all the tweets, reach out to&nbsp;Shahan Ali Memon at (shahan@nyu.edu).</strong></p> <p>If you use this data, please cite our paper as follows:&nbsp;</p> <p><em>&quot;Shahan Ali Memon and Kathleen M. Carley. Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset, In Proceedings of The 5th International Workshop on Mining Actionable Insights from Social Networks (MAISoN 2020), co-located with CIKM, virtual event due to COVID-19, 2020.&quot;</em></p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

INTRODUCTION OF COVID-NEWS-US-NNK AND COVID-NEWS-BD-NNK DATASET

<p>Introduction</p> <p>There are several works based on Natural Language Processing on newspaper reports. Mining opinions from headlines [ 1 ] using Standford NLP and SVM by Rameshbhaiet. Al.compared several algorithms on a small and large dataset. Rubinet. al., in their paper [ 2 ], created a mechanism to differentiate fake news from real ones by building a set of characteristics of news according to their types. The purpose was to contribute to the low resource data available for training machine learning algorithms. Doumitet. al.in [ 3 ] have implemented LDA, a topic modeling approach to study bias present in online news media.</p> <p>However, there are not many NLP research invested in studying COVID-19. Most applications include classification of chest X-rays and CT-scans to detect presence of pneumonia in lungs [ 4 ], a consequence of the virus. Other research areas include studying the genome sequence of the virus[ 5 ][ 6 ][ 7 ] and replicating its structure to fight and find a vaccine. This research is crucial in battling the pandemic. The few NLP based research publications are sentiment classification of online tweets by Samuel et el [ 8 ] to understand fear persisting in people due to the virus. Similar work has been done using the LSTM network to classify sentiments from online discussion forums by Jelodaret. al.[ 9 ]. NKK dataset is the first study on a comparatively larger dataset of a newspaper report on COVID-19, which contributed to the virus&rsquo;s awareness to the best of our knowledge.</p> <p>&nbsp;</p> <p>2 Data-set Introduction</p> <p>2.1 Data Collection</p> <p>We accumulated 1000 online newspaper report from United States of America (USA) on COVID-19. The newspaper includes The Washington Post (USA) and StarTribune (USA). We have named it as &ldquo;Covid-News-USA-NNK&rdquo;. We also accumulated 50 online newspaper report from Bangladesh on the issue and named it &ldquo;Covid-News-BD-NNK&rdquo;. The newspaper includes The Daily Star (BD) and Prothom Alo (BD). All these newspapers are from the top provider and top read in the respective countries. The collection was done manually by 10 human data-collectors of age group 23- with university degrees. This approach was suitable compared to automation to ensure the news were highly relevant to the subject. The newspaper online sites had dynamic content with advertisements in no particular order. Therefore there were high chances of online scrappers to collect inaccurate news reports. One of the challenges while collecting the data is the requirement of subscription. Each newspaper required $1 per subscriptions. Some criteria in collecting the news reports provided as guideline to the human data-collectors were as follows:</p> <ul> <li>The headline must have one or more words directly or indirectly related to COVID-19.</li> <li>The content of each news must have 5 or more keywords directly or indirectly related to COVID-19.</li> <li>The genre of the news can be anything as long as it is relevant to the topic. Political, social, economical genres are to be more prioritized.</li> <li>Avoid taking duplicate reports.</li> <li>Maintain a time frame for the above mentioned newspapers.</li> </ul> <p>To collect these data we used a google form for USA and BD. We have two human editor to go through each entry to check any spam or troll entry.</p> <p>2.2 Data Pre-processing and Statistics</p> <p>Some pre-processing steps performed on the newspaper report dataset are as follows:</p> <ul> <li>Remove hyperlinks.</li> <li>Remove non-English alphanumeric characters.</li> <li>Remove stop words.</li> <li>Lemmatize text.</li> </ul> <p>While more pre-processing could have been applied, we tried to keep the data as much unchanged as possible since changing sentence structures could result us in valuable information loss. While this was done with help of a script, we also assigned same human collectors to cross check for any presence of the above mentioned criteria.</p> <p>The primary data statistics of the two dataset are shown in Table 1 and 2.</p> <pre><code>Table 1: Covid-News-USA-NNK data statistics </code></pre> <pre><code>No of words per headline </code></pre> <pre><code>7 to 20 </code></pre> <pre><code>No of words per body content </code></pre> <pre><code>150 to 2100 </code></pre> <pre><code>Table 2: Covid-News-BD-NNK data statistics No of words per headline </code></pre> <pre><code>10 to 20 </code></pre> <pre><code>No of words per body content </code></pre> <pre><code>100 to 1500 </code></pre> <p>2.3 Dataset Repository</p> <p>We used GitHub as our primary data repository in account name NKK^1. Here, we created two repositories USA-NKK^2 and BD-NNK^3. The dataset is available in both CSV and JSON format. We are regularly updating the CSV files and regenerating JSON using a py script. We provided a python script file for essential operation. We welcome all outside collaboration to enrich the dataset.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>3 Literature Review</p> <p>Natural Language Processing (NLP) deals with text (also known as categorical) data in computer science, utilizing numerous diverse methods like one-hot encoding, word embedding, etc., that transform text to machine language, which can be fed to multiple machine learning and deep learning algorithms.</p> <p>Some well-known applications of NLP includes fraud detection on online media sites[ 10 ], using authorship attribution in fallback authentication systems[ 11 ], intelligent conversational agents or chatbots[ 12 ] and machine translations used by Google Translate[ 13 ]. While these are all downstream tasks, several exciting developments have been made in the algorithm solely for Natural Language Processing tasks. The two most trending ones are BERT[ 14 ], which uses bidirectional encoder-decoder architecture to create the transformer model, that can do near-perfect classification tasks and next-word predictions for next generations, and GPT-3 models released by OpenAI[ 15 ] that can generate texts almost human-like. However, these are all pre-trained models since they carry huge computation cost. Information Extraction is a generalized concept of retrieving information from a dataset. Information extraction from an image could be retrieving vital feature spaces or targeted portions of an image; information extraction from speech could be retrieving information about names, places, etc[ 16 ]. Information extraction in texts could be identifying named entities and locations or essential data. Topic modeling is a sub-task of NLP and also a process of information extraction. It clusters words and phrases of the same context together into groups. Topic modeling is an unsupervised learning method that gives us a brief idea about a set of text. One commonly used topic modeling is Latent Dirichlet Allocation or LDA[17].</p> <p>Keyword extraction is a process of information extraction and sub-task of NLP to extract essential words and phrases from a text. TextRank [ 18 ] is an efficient keyword extraction technique that uses graphs to calculate the weight of each word and pick the words with more weight to it.</p> <p>Word clouds are a great visualization technique to understand the overall &rsquo;talk of the topic&rsquo;. The clustered words give us a quick understanding of the content.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>4 Our experiments and Result analysis</p> <p>We used the wordcloud library^4 to create the word clouds. Figure 1 and 3 presents the word cloud of Covid-News-USA- NNK dataset by month from February to May. From the figures 1,2,3, we can point few information:</p> <ul> <li>In February, both the news paper have talked about China and source of the outbreak.</li> <li>StarTribune emphasized on Minnesota as the most concerned state. In April, it seemed to have been concerned more.</li> <li>Both the newspaper talked about the virus impacting the economy, i.e, bank, elections, administrations, markets.</li> <li>Washington Post discussed global issues more than StarTribune.</li> <li>StarTribune in February mentioned the first precautionary measurement: wearing masks, and the uncontrollable spread of the virus throughout the nation.</li> <li>While both the newspaper mentioned the outbreak in China in February, the weight of the spread in the United States are more highlighted through out March till May, displaying the critical impact caused by the virus.</li> </ul> <p>We used a script to extract all numbers related to certain keywords like &rsquo;Deaths&rsquo;, &rsquo;Infected&rsquo;, &rsquo;Died&rsquo; , &rsquo;Infections&rsquo;, &rsquo;Quarantined&rsquo;, Lock-down&rsquo;, &rsquo;Diagnosed&rsquo; etc from the news reports and created a number of cases for both the newspaper. Figure 4 shows the statistics of this series. From this extraction technique, we can observe that April was the peak month for the covid cases as it gradually rose from February. Both the newspaper clearly shows us that the rise in covid cases from February to March was slower than the rise from March to April. This is an important indicator of possible recklessness in preparations to battle the virus. However, the steep fall from April to May also shows the positive response against the attack. We used Vader Sentiment Analysis to extract sentiment of the headlines and the body. On average, the sentiments were from -0.5 to -0.9. Vader Sentiment scale ranges from -1(highly negative to 1(highly positive). There were some cases</p> <p>where the sentiment scores of the headline and body contradicted each other,i.e., the sentiment of the headline was negative but the sentiment of the body was slightly positive. Overall, sentiment analysis can assist us sort the most concerning (most negative) news from the positive ones, from which we can learn more about the indicators related to COVID-19 and the serious impact caused by it. Moreover, sentiment analysis can also provide us information about how a state or country is reacting to the pandemic. We used PageRank algorithm to extract keywords from headlines as well as the body content. PageRank efficiently highlights important relevant keywords in the text. Some frequently occurring important keywords extracted from both the datasets are: &rsquo;China&rsquo;, Government&rsquo;, &rsquo;Masks&rsquo;, &rsquo;Economy&rsquo;, &rsquo;Crisis&rsquo;, &rsquo;Theft&rsquo; , &rsquo;Stock market&rsquo; , &rsquo;Jobs&rsquo; , &rsquo;Election&rsquo;, &rsquo;Missteps&rsquo;, &rsquo;Health&rsquo;, &rsquo;Response&rsquo;. Keywords extraction acts as a filter allowing quick searches for indicators in case of locating situations of the economy, how states are defending against the pandemic, the condition of the health care system etc.</p> <p>&nbsp;</p> <p>5 Conclusion</p> <p>This dataset can demonstrate how news reports could speculate the situation differently based on the news source. The different types of experiments are possible to assert the importance of Natural Language Processing in newspaper report analysis. We are looking for more collaborators in GitHub to enrich the dataset, which will make it possible to run extensive deep learning experiments.</p> <p>&nbsp;</p> <p>References</p> <p>[1] Chaudhary Jashubhai Rameshbhai and Joy Paulose. Opinion mining on newspaper headlines using svm and nlp. International Journal of Electrical &amp; Computer Engineering (2088-8708), 9(3), 2019. [2]Victoria L Rubin, Yimin Chen, and Nadia K Conroy. Deception detection for news: three types of fakes. Proceedings of the Association for Information Science and Technology, 52(1):1&ndash;4, 2015. [3]Sarjoun Doumit and Ali Minai. Online news media bias analysis using an lda-nlp approach. InInternational Conference on Complex Systems, 2011. [4]Md Manjurul Ahsan, Kishor Datta Gupta, Mohammad Maminur Islam, Sajib Sen, Md Rahman, Moham- mad Shakhawat Hossain, et al. Study of different deep learning approach with explainable ai for screening patients with covid-19 symptoms: Using ct scan and chest x-ray image dataset.arXiv preprint arXiv:2007.12525, 2020. [5]Gurjit S Randhawa, Maximillian PM Soltysiak, Hadi El Roz, Camila PE de Souza, Kathleen A Hill, and Lila Kari. Machine learning using intrinsic genomic signatures for rapid classification of novel pathogens: Covid-19 case study.Plos one, 15(4):e0232391, 2020.</p> <p>[6]Ahmad Alimadadi, Sachin Aryal, Ishan Manandhar, Patricia B Munroe, Bina Joe, and Xi Cheng. Artificial intelligence and machine learning to fight covid-19, 2020. [7]Shreshth Tuli, Shikhar Tuli, Rakesh Tuli, and Sukhpal Singh Gill. Predicting the growth and trend of covid- pandemic using machine learning and cloud computing.Internet of Things, page 100222, 2020. [8]Jim Samuel, GG Ali, Md Rahman, Ek Esawi, Yana Samuel, et al. Covid-19 public sentiment insights and machine learning for tweets classification.Information, 11(6):314, 2020. [9] Hamed Jelodar, Yongli Wang, Rita Orji, and Hucheng Huang. Deep sentiment classification and topic discovery on novel coronavirus or covid-19 online discussions: Nlp using lstm recurrent neural network approach.arXiv preprint arXiv:2004.11695, 2020.</p> <p>[10]Nafiz Sadman, Kishor Datta Gupta, Ariful Haque, Subash Poudyal, and Sajib Sen. Detect review manipulation by leveraging reviewer historical stylometrics in amazon, yelp, facebook and google reviews. InProceedings of the 2020 The 6th International Conference on E-Business and Applications, pages 42&ndash;47, 2020.</p> <p>[11]Nafiz Sadman, Kishor Datta Gupta, Ariful Haque, Subash Poudyal, and Sajib Sen. Stylometry as a reliable method for fallback authentication. InProceedings of the 2020 17th International Conference on Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology, 2020.</p> <p>[12]Ethan Fast, Binbin Chen, Julia Mendelsohn, Jonathan Bassen, and Michael S Bernstein. Iris: A conversational agent for complex tasks. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pages 1&ndash;12, 2018.</p> <p>[13] Philipp Koehn.Statistical machine translation. Cambridge University Press, 2009.</p> <p>[14]Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805, 2018.</p> <p>[15]Will Douglas Heavenarchive page. Openai&rsquo;s new language generator gpt-3 is shockingly good&mdash;and completely mindless <a href="https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-">https://www.technologyreview.com/2020/07/20/1005454/openai-machine-learning-language-generator-</a> gpt-3-nlp/. Technical report.</p> <p>[16]Chin-Hui Lee and Sabato Marco Siniscalchi. An information-extraction approach to speech processing: Analysis, detection, verification, and recognition.Proceedings of the IEEE, 101(5):1089&ndash;1115, 2013.</p> <p>[17]Hamed Jelodar, Yongli Wang, Chi Yuan, Xia Feng, Xiahui Jiang, Yanchao Li, and Liang Zhao. Latent dirich- let allocation (lda) and topic modeling: models, applications, a survey. Multimedia Tools and Applications, 78(11):15169&ndash;15211, 2019.</p> <p>[18]Monica Bianchini, Marco Gori, and Franco Scarselli. Inside pagerank.ACM Transactions on Internet Technology (TOIT), 5(1):92&ndash;128, 2005.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Longitudinal high-throughput TCR repertoire profiling reveals the dynamics of T cell memory formation after mild COVID-19 infection

<p>Processed TCRbeta and TCRalpha repertoires after mild COVID-19 (Version 2.0: day 85 timepoints added) infection,&nbsp;see&nbsp;preprint:&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2020.05.18.100545v3">https://www.biorxiv.org/content/10.1101/2020.05.18.100545v3</a></p> <p>and GitHub repository:&nbsp;<a href="https://github.com/pogorely/Minervina_COVID">https://github.com/pogorely/Minervina_COVID</a></p> <p>Two donors (M and W), two biological replicates of PBMC&nbsp;(F1 and F2), CD4+, CD8+, and Memory subpopulations&nbsp;for each post-infection time points (day 15, 30, 37, 45, 85 post-infection), and pre-infection PBMC repertoires sampled in 2019 and 2018.&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Data from: The impact of human mobility networks on the global spread of COVID-19

<p>This is&nbsp;empirical dataset from the paper &quot;The impact of human mobility networks on the global spread of COVID-19&quot;. Specifically, the dataset includes several files: (a) the COVID-19 network - an origin/destination matrix (i.e., &quot;covid_network.csv&quot;); (b) the common language network - edgelist format (i.e. &quot;edge_list_comlang.csv&quot;); (c) the same continent network - edgelist format (i.e., &quot;edge_list_continent.csv&quot;; (d) the contiguity network (i.e., &quot;edge_list_contig.csv&quot;);&nbsp; (e) the migration network - edgelist format (i.e., &quot;edge_list_migration_in.csv&quot;; (f) the tourism network - edgelist format (i.e., edge_list_tourism_in.csv&quot;); (g) the list of nodes (countries) corresponding to files (b)-(e) (i.e., &quot;nodes.csv&quot;).&nbsp;Additionally, we uploaded the Rcode used in the paper (i.e. &quot;code&quot;), as a .pdf file format,&nbsp;the&nbsp;data source for the figures included in the paper (i.e., &quot;covid_network_matrix.csv&quot;, &quot;matrix_migration_out.csv&quot;, &quot;matrix_tourism.csv&quot; - Figure 1; &quot;Fig_2_a_matrix_comlang.csv&quot;, Fig_2_b_matrix_contig.csv&quot;, &quot;Fig_2_c_matrix_continent.csv&quot; - Figure 2; &quot;Fig_3.graphmlz - Figure 3; Fig_4.graphmlz - Figure 4)&nbsp;and the &quot;global network of COVID-19 onset&quot; (an individual-level data) (i.e., &quot;global_covid_network.csv&quot;).&nbsp;</p> <p>For details, please, see the Methods section of the paper:&nbsp;The impact of human mobility networks on the global spread of COVID-19&nbsp;(Hancean, M.-G., Slavinec, M., Perc, M).&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

The Daily Life of Software Engineers during the COVID-19 Pandemic -- Replication Package

<p>Following the onset of the COVID-19 pandemic and subsequent lockdowns, software engineers&#39; daily life was disrupted and abruptly forced into remote working from home. &nbsp;This change deeply impacted typical working routines, affecting both well-being and productivity.&nbsp;Moreover, this pandemic will have long-lasting effects in the software industry, with several tech companies allowing their employees to work from home indefinitely if they wish to do so. &nbsp;Therefore, it is crucial to analyze and understand how a typical working day looks like when working from home and how individual activities affect software developers&#39; well-being and productivity.&nbsp;We performed a two-wave longitudinal study involving almost 200 globally carefully selected software professionals, inferring daily activities with perceived well-being, productivity, and other relevant psychological and social variables.&nbsp;Results suggest that the time software engineers spent doing specific activities from home was similar when working in the office. (e.g., coding &gt; emails &gt; code review &gt; networking). &nbsp;However, we also found some meaningful mean differences.&nbsp;The amount of time developers spent on each activity was unrelated to their well-being, perceived productivity, and other variables.&nbsp;We conclude that working remotely is not per se&nbsp;a challenge for organizations or developers.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Covid-on-the-Web dataset

<p>This RDF dataset provides two main knowledge graphs produced by processing the scholarly articles of the <a href="https://www.semanticscholar.org/cord19">COVID-19 Open Research Dataset</a> (CORD-19), a resource of articles about COVID-19 and the coronavirus family of viruses.</p> <p>The <em>CORD-19 Named Entities Knowledge Graph</em> describes named entities identified and disambiguated by NCBO BioPortal annotator, Entity-fishing and DBpedia Spotlight. The <em>CORD-19 Argumentative Knowledge Graph</em> describes argumentative components&nbsp;and PICO elements extracted from the articles by the Argumentative Clinical Trial Analysis platform (ACTA).</p> <p>Homepage: <a href="https://github.com/Wimmics/CovidOnTheWeb">https://github.com/Wimmics/CovidOnTheWeb</a></p> <p>License: see the LICENCE file in the archive.</p>

openother-openMay 2020View details →
Figshare44/100

Dataset for the paper "Prolonged prothrombin time as an early prognostic indicator of severe acute respiratory distress syndrome in patients with COVID-19 related pneumonia"

<p>The dataset contains the data on ICU-transferred (N=100) and Stable (N=131) patients with COVID-19 (N=156) and Non-COVID-19 viral pneumonia (N=75). Among COVID-19 patients of this study, 82 patients developed Refractory Respiratory Failure (RRF) or Severe Acute Respiratory Distress Syndrome (SARDS) and were transferred to Intensive Care Unit (ICU), 74 patients had a Stable course of disease and were not transferred to ICU. Collected data are presented as a table with columns:<br> - Gender;<br> - Age (years);<br> - SARS-CoV-2 RT-PCR testing results;<br> - Time between the disease onset and admission to the hospital (days);<br> - Time between admission to the hospital and transfer to ICU (days);<br> - Artificial lung ventilation in ICU needed;<br> - C-reactive protein (CRP) upon admission (mg/L);<br> - International Normalized Ratio (INR) upon admission;<br> - Prothrombin Time (PT) upon admission (sec.);<br> - Fibrinogen upon admission (mg/L);<br> - Chest Computed Tomography (CT) upon admission: lung tissue affected (%);<br> - Platelet count upon admission (10^9/L);<br> - Chest CT, 1 week after admission: lung tissue affected (%);<br> - CRP, 1 week after admission (mg/L);<br> - Platelet count, 1 week after admission (10^9/L).</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Artificial COVID-19 Cases in Paris and Geographic Data Useful for Geomasking

<p>Artificial dataset of addresses of COVID-19 cases in Paris. The dataset was created to test geomasking techniques to be used on the real data collected by the French health administration. The dataset was used in the paper &quot;Geographically Masking Addresses to Study COVID-19 Clusters&quot; by Walid Houfaf-Khoufaf and Guillaume Touya. The dataset contains the following files:</p> <ul> <li>roads_paris_IGN.shp contains the road lines from IGN France in Paris;</li> <li>buildings_paris_IGN.shp contains the building polygons from IGN France in Paris (useful to aggregate points to building groups);</li> <li>faces.shp contains the blocks built from the roads (useful to aggregate points to blocks);</li> <li>ban_75.shp contains all the address points in Paris from the open BAN database.</li> <li>artificial_COVID_cases.csv&nbsp; contains the artificial COVID cases generated from 3 months in 2020 in the Paris area.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

International datasets behavior effects COVID-19

<p>This dataset stems from the project &lsquo;Beprepared&rsquo;: (<a href="https://be-prepared-consortium.nl/">https://be-prepared-consortium.nl/</a>) which aims to provide in-depth analyses of mixed-method behavioural science data collected throughout the unprecedented COVID-19 pandemic and inform preparedness strategies for future outbreaks. In approaching the research from a behavioural and social science perspective, researchers focus on four main themes:</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Prevention behaviour, psychosocial and contextual determinants, and (communication) interventions</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Resilience and engagement of citizens, communities and organisations</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Research methodology and preparedness</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Effective and integrated policy advice</p> <p>&nbsp;</p> <p>This resource links to the theme &lsquo;research methodology&rsquo; and provides an overview of datasets that have been used internationally to study the behavioral effects of the Covid-19 pandemic. These datasources can be used to study how people behave in a variety of settings during the Covid pandemic and so to inform policy-makers, but also to study the effects of behavioral interventions. It includes datasources that for example study mobility behavior at a regional or national level, physical distancing in public, health adherence behaviors (like handwashing, mask wearing), social contacts on- and offline, purchasing behaviors (shopping) etc.</p> <p>&nbsp;</p> <p>The resource consists of two datasets:</p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp; A dataset (in .xlsx and .csv format) of the search strategy used to come to the list of datasources called &ldquo;search strategy&rdquo;</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp; A dataset (in .xslx and .csv format) of the results of the search, called &ldquo;search results&rdquo;</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp; A dataset (in .xslx and .csv format) of a step where duplicate studies are identified</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp; A dataset (in .xslx and .csv format) where for 131 studies the data quality was assessed</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Impact of vaccinations, boosters and lockdowns on COVID-19 waves in French Polynesia

<p>COVID-19 case, hospitalisation, death, seroprevalence, vaccination and population data, and age-dependent contact rate, severe burden risk and vaccine effectiveness parameter estimates, required to fit model and run simulations in article &quot;Impact of vaccinations, boosters and lockdowns on COVID-19 waves in French Polynesia&quot;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Covid-19 CT dataset for Body Part Regression Tutorial

<p>The dataset is a subset of CT scans from the <a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70226443">Covid-19-AR</a> dataset from the Cancer Image Archive. The data were converted from the DICOM file format to the NifTI&nbsp;file format for better and easier handling. The converted dataset was created for a tutorial of the <a href="https://github.com/MIC-DKFZ/BodyPartRegression">bpreg</a> python package.</p> <p>Acknowledgment:<br> The dataset was funded with federal funds from the National Center for Advancing Translational Sciences&nbsp;&nbsp;UL1 TR003107 and the National Cancer Institute, Contract No. 75N91019D00024, Subcontract 20X023F.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

COVID-19 vaccination data in Israel by age over time until August 2021

<p>COVID-19 vaccination data in Israel processed to show vaccination by age over time. These&nbsp;datasets are&nbsp;derived from publicly available Ministry of Health data, but processed for analytics about uptake in different age groups over time. They cover the mass vaccination campaign for COVID-19 until August 2021. The campaign consisted of the administration of multiple doses of the Pfizer vaccine.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Bibliography on COVID-19 and ischemic stroke

<p>Search on PubMed literature on COVID-19 and ischemic stroke.</p> <p>The uploaded database was generated on November 2, 2021.&nbsp;</p> <p>The database contains the following attributes:</p> <p>- PMID: PubMed identifier of the article.&nbsp;<br> - Autors: list of authors.&nbsp;<br> - Refer&egrave;ncia: bibliographic reference of the article.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Audio recordings of COVID-19 positive individuals from the prospective Predi-COVID cohort study with their ageusia and anosmia status

<p>We uploaded 1636 audio recordings originating from 259 distinct participants in the prospective Predi-COVID cohort study recruited between May 2020 and May 2021. The audios have been converted from their original format into WAV files. The audio name structure integrates the participant ID, the recording date and time of the audio recording, the type of audio (Type 1: reading of a text, Type2: hold the [a] vowel), the original audio format, and the symptomatic status for ageusia and anosmia (1: symptomatic, 0: asymptomatic) as such:</p> <p>predi-covid_{participant}{recording date and time}{type of audio}{original format}{sympyomatic status}.wav</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data for: COVID-19 patents/patent applications (Jan. 2020 – Oct. 2021)

<p>This dataset contains information regarding both applications and granted patents on COVID-19 disease</p> <p>Derwent Innovation database was used for data mining (accessed on Nov. 20, 2021).</p> <p>The patent search was carried out on by means of a precise set of keywords and performed in the title/abstract/claims search field.</p> <p>6,148 Inpadoc patent families were retrieved.&nbsp;</p> <p>The&nbsp;<em>XLS file</em>&nbsp;contains information related to Title, Abstract - DWPI&nbsp;, First Claim, Priority Number, Priority Date, Application Number, Application Date, Publication Number, Publication Date, IPC &ndash; Current, CPC &ndash; Current, Assignee/Applicant, Optimized Assignee, INPADOC Family Members.&nbsp;</p> <p>The top countries/regions are China (3,271), WO (1,057), India (487), United States (455).</p> <p>The top IPC codes are listed in the following table:</p> <p>&nbsp;</p> <table align="center"> <tbody> <tr> <td> <p><strong>IPC</strong></p> </td> <td> <p><strong>Definition</strong></p> </td> <td> <p><strong>No. of patents/applications</strong></p> </td> </tr> <tr> <td> <p>A61P 31/14</p> </td> <td> <p><em>Antivirals for RNA viruses</em></p> </td> <td> <p>1565</p> </td> </tr> <tr> <td> <p>G01N 33/569</p> </td> <td> <p><em>Biological material &bull;&bull;&nbsp;Chemical analysis of biological material &bull;&bull;&bull; Immunoassay; Biospecific binding assay; Materials therefor &bull;&bull;&bull;&bull; for microorganisms</em></p> </td> <td> <p>783</p> </td> </tr> <tr> <td> <p>C12Q 1/70</p> </td> <td> <p><em>Measuring or testing processes &bull; involving virus or bacteriophage</em></p> </td> <td> <p>642</p> </td> </tr> <tr> <td> <p>A61P 11/00</p> </td> <td> <p><em>Drugs for disorders of the respiratory system</em></p> </td> <td> <p>582</p> </td> </tr> <tr> <td> <p>A61K 39/215</p> </td> <td> <p><em>Medicinal preparations containing antigens or antibodies &bull; Viral antigens &bull;&bull; Coronaviridae, e.g., avian infectious bronchitis virus</em></p> </td> <td> <p>377</p> </td> </tr> </tbody> </table> <p><strong>Value of the dataset</strong>: prior art searches; patent landscape analysis&nbsp;</p> <p><strong>Steps to reproduce data</strong>:&nbsp;</p> <p>CTB=(&quot;covid-19&quot; OR &quot;covid 19&quot; OR &quot;covid19&quot; ADJ &quot;SARS-CoV-2&quot; OR &quot;SARS-CoV2&quot; OR &quot;sarscov2&quot; ADJ &quot;2019 ncov&quot; OR &quot;2019-nCoV&quot; OR &quot;2019nCoV&quot; ADJ &quot;covid-2019&quot; OR &quot;covid 2019&quot; OR &quot;COVID2019&quot; OR &quot;severe acute respiratory syndrome coronavirus 2&quot; OR &quot;2019 novel coronavirus&quot; OR &quot;coronavirus disease 2019&quot; OR &quot;novel corona virus&quot; OR &quot;novel coronavirus&quot; OR &quot;new corona virus&quot; OR &quot;new coronavirus&quot; OR &quot;Wuhan coronavirus&quot;)</p> <p>CTB=title/abstract/claims</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Unexposed populations and potential COVID-19 burden in European countries as of 21st November 2021

<p>Estimates of numbers of SARS-CoV-2 infections by country and age group over time, current proportions in different immune states, and potential remaining burden of hospitalisations and deaths for 19 European countries from article &quot;Unexposed populations and potential COVID-19 burden in European countries as of 21st November 2021&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)

<p><strong>Sources:&nbsp;</strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record