Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,674
datasets available to search
ShareScore release 0.7.1
Dataset results
9,674 results for “COVID-19”
The Daily Life of Software Engineers during the COVID-19 Pandemic -- Replication Package
<p>Following the onset of the COVID-19 pandemic and subsequent lockdowns, software engineers' daily life was disrupted and abruptly forced into remote working from home. This change deeply impacted typical working routines, affecting both well-being and productivity. Moreover, this pandemic will have long-lasting effects in the software industry, with several tech companies allowing their employees to work from home indefinitely if they wish to do so. Therefore, it is crucial to analyze and understand how a typical working day looks like when working from home and how individual activities affect software developers' well-being and productivity. We performed a two-wave longitudinal study involving almost 200 globally carefully selected software professionals, inferring daily activities with perceived well-being, productivity, and other relevant psychological and social variables. Results suggest that the time software engineers spent doing specific activities from home was similar when working in the office. (e.g., coding > emails > code review > networking). However, we also found some meaningful mean differences. The amount of time developers spent on each activity was unrelated to their well-being, perceived productivity, and other variables. We conclude that working remotely is not per se a challenge for organizations or developers.</p>
Dataset for the paper "Prolonged prothrombin time as an early prognostic indicator of severe acute respiratory distress syndrome in patients with COVID-19 related pneumonia"
<p>The dataset contains the data on ICU-transferred (N=100) and Stable (N=131) patients with COVID-19 (N=156) and Non-COVID-19 viral pneumonia (N=75). Among COVID-19 patients of this study, 82 patients developed Refractory Respiratory Failure (RRF) or Severe Acute Respiratory Distress Syndrome (SARDS) and were transferred to Intensive Care Unit (ICU), 74 patients had a Stable course of disease and were not transferred to ICU. Collected data are presented as a table with columns:<br> - Gender;<br> - Age (years);<br> - SARS-CoV-2 RT-PCR testing results;<br> - Time between the disease onset and admission to the hospital (days);<br> - Time between admission to the hospital and transfer to ICU (days);<br> - Artificial lung ventilation in ICU needed;<br> - C-reactive protein (CRP) upon admission (mg/L);<br> - International Normalized Ratio (INR) upon admission;<br> - Prothrombin Time (PT) upon admission (sec.);<br> - Fibrinogen upon admission (mg/L);<br> - Chest Computed Tomography (CT) upon admission: lung tissue affected (%);<br> - Platelet count upon admission (10^9/L);<br> - Chest CT, 1 week after admission: lung tissue affected (%);<br> - CRP, 1 week after admission (mg/L);<br> - Platelet count, 1 week after admission (10^9/L).</p>
Artificial COVID-19 Cases in Paris and Geographic Data Useful for Geomasking
<p>Artificial dataset of addresses of COVID-19 cases in Paris. The dataset was created to test geomasking techniques to be used on the real data collected by the French health administration. The dataset was used in the paper "Geographically Masking Addresses to Study COVID-19 Clusters" by Walid Houfaf-Khoufaf and Guillaume Touya. The dataset contains the following files:</p> <ul> <li>roads_paris_IGN.shp contains the road lines from IGN France in Paris;</li> <li>buildings_paris_IGN.shp contains the building polygons from IGN France in Paris (useful to aggregate points to building groups);</li> <li>faces.shp contains the blocks built from the roads (useful to aggregate points to blocks);</li> <li>ban_75.shp contains all the address points in Paris from the open BAN database.</li> <li>artificial_COVID_cases.csv contains the artificial COVID cases generated from 3 months in 2020 in the Paris area.</li> </ul> <p> </p>
International datasets behavior effects COVID-19
<p>This dataset stems from the project ‘Beprepared’: (<a href="https://be-prepared-consortium.nl/">https://be-prepared-consortium.nl/</a>) which aims to provide in-depth analyses of mixed-method behavioural science data collected throughout the unprecedented COVID-19 pandemic and inform preparedness strategies for future outbreaks. In approaching the research from a behavioural and social science perspective, researchers focus on four main themes:</p> <p>· Prevention behaviour, psychosocial and contextual determinants, and (communication) interventions</p> <p>· Resilience and engagement of citizens, communities and organisations</p> <p>· Research methodology and preparedness</p> <p>· Effective and integrated policy advice</p> <p> </p> <p>This resource links to the theme ‘research methodology’ and provides an overview of datasets that have been used internationally to study the behavioral effects of the Covid-19 pandemic. These datasources can be used to study how people behave in a variety of settings during the Covid pandemic and so to inform policy-makers, but also to study the effects of behavioral interventions. It includes datasources that for example study mobility behavior at a regional or national level, physical distancing in public, health adherence behaviors (like handwashing, mask wearing), social contacts on- and offline, purchasing behaviors (shopping) etc.</p> <p> </p> <p>The resource consists of two datasets:</p> <p>1. A dataset (in .xlsx and .csv format) of the search strategy used to come to the list of datasources called “search strategy”</p> <p>2. A dataset (in .xslx and .csv format) of the results of the search, called “search results”</p> <p>3. A dataset (in .xslx and .csv format) of a step where duplicate studies are identified</p> <p>4. A dataset (in .xslx and .csv format) where for 131 studies the data quality was assessed</p>
Impact of vaccinations, boosters and lockdowns on COVID-19 waves in French Polynesia
<p>COVID-19 case, hospitalisation, death, seroprevalence, vaccination and population data, and age-dependent contact rate, severe burden risk and vaccine effectiveness parameter estimates, required to fit model and run simulations in article "Impact of vaccinations, boosters and lockdowns on COVID-19 waves in French Polynesia"</p>
Covid-19 CT dataset for Body Part Regression Tutorial
<p>The dataset is a subset of CT scans from the <a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=70226443">Covid-19-AR</a> dataset from the Cancer Image Archive. The data were converted from the DICOM file format to the NifTI file format for better and easier handling. The converted dataset was created for a tutorial of the <a href="https://github.com/MIC-DKFZ/BodyPartRegression">bpreg</a> python package.</p> <p>Acknowledgment:<br> The dataset was funded with federal funds from the National Center for Advancing Translational Sciences UL1 TR003107 and the National Cancer Institute, Contract No. 75N91019D00024, Subcontract 20X023F. </p>
COVID-19 vaccination data in Israel by age over time until August 2021
<p>COVID-19 vaccination data in Israel processed to show vaccination by age over time. These datasets are derived from publicly available Ministry of Health data, but processed for analytics about uptake in different age groups over time. They cover the mass vaccination campaign for COVID-19 until August 2021. The campaign consisted of the administration of multiple doses of the Pfizer vaccine.</p>
Bibliography on COVID-19 and ischemic stroke
<p>Search on PubMed literature on COVID-19 and ischemic stroke.</p> <p>The uploaded database was generated on November 2, 2021. </p> <p>The database contains the following attributes:</p> <p>- PMID: PubMed identifier of the article. <br> - Autors: list of authors. <br> - Referència: bibliographic reference of the article. </p>
Audio recordings of COVID-19 positive individuals from the prospective Predi-COVID cohort study with their ageusia and anosmia status
<p>We uploaded 1636 audio recordings originating from 259 distinct participants in the prospective Predi-COVID cohort study recruited between May 2020 and May 2021. The audios have been converted from their original format into WAV files. The audio name structure integrates the participant ID, the recording date and time of the audio recording, the type of audio (Type 1: reading of a text, Type2: hold the [a] vowel), the original audio format, and the symptomatic status for ageusia and anosmia (1: symptomatic, 0: asymptomatic) as such:</p> <p>predi-covid_{participant}{recording date and time}{type of audio}{original format}{sympyomatic status}.wav</p>
Data for: COVID-19 patents/patent applications (Jan. 2020 – Oct. 2021)
<p>This dataset contains information regarding both applications and granted patents on COVID-19 disease</p> <p>Derwent Innovation database was used for data mining (accessed on Nov. 20, 2021).</p> <p>The patent search was carried out on by means of a precise set of keywords and performed in the title/abstract/claims search field.</p> <p>6,148 Inpadoc patent families were retrieved. </p> <p>The <em>XLS file</em> contains information related to Title, Abstract - DWPI , First Claim, Priority Number, Priority Date, Application Number, Application Date, Publication Number, Publication Date, IPC – Current, CPC – Current, Assignee/Applicant, Optimized Assignee, INPADOC Family Members. </p> <p>The top countries/regions are China (3,271), WO (1,057), India (487), United States (455).</p> <p>The top IPC codes are listed in the following table:</p> <p> </p> <table align="center"> <tbody> <tr> <td> <p><strong>IPC</strong></p> </td> <td> <p><strong>Definition</strong></p> </td> <td> <p><strong>No. of patents/applications</strong></p> </td> </tr> <tr> <td> <p>A61P 31/14</p> </td> <td> <p><em>Antivirals for RNA viruses</em></p> </td> <td> <p>1565</p> </td> </tr> <tr> <td> <p>G01N 33/569</p> </td> <td> <p><em>Biological material •• Chemical analysis of biological material ••• Immunoassay; Biospecific binding assay; Materials therefor •••• for microorganisms</em></p> </td> <td> <p>783</p> </td> </tr> <tr> <td> <p>C12Q 1/70</p> </td> <td> <p><em>Measuring or testing processes • involving virus or bacteriophage</em></p> </td> <td> <p>642</p> </td> </tr> <tr> <td> <p>A61P 11/00</p> </td> <td> <p><em>Drugs for disorders of the respiratory system</em></p> </td> <td> <p>582</p> </td> </tr> <tr> <td> <p>A61K 39/215</p> </td> <td> <p><em>Medicinal preparations containing antigens or antibodies • Viral antigens •• Coronaviridae, e.g., avian infectious bronchitis virus</em></p> </td> <td> <p>377</p> </td> </tr> </tbody> </table> <p><strong>Value of the dataset</strong>: prior art searches; patent landscape analysis </p> <p><strong>Steps to reproduce data</strong>: </p> <p>CTB=("covid-19" OR "covid 19" OR "covid19" ADJ "SARS-CoV-2" OR "SARS-CoV2" OR "sarscov2" ADJ "2019 ncov" OR "2019-nCoV" OR "2019nCoV" ADJ "covid-2019" OR "covid 2019" OR "COVID2019" OR "severe acute respiratory syndrome coronavirus 2" OR "2019 novel coronavirus" OR "coronavirus disease 2019" OR "novel corona virus" OR "novel coronavirus" OR "new corona virus" OR "new coronavirus" OR "Wuhan coronavirus")</p> <p>CTB=title/abstract/claims</p>
Unexposed populations and potential COVID-19 burden in European countries as of 21st November 2021
<p>Estimates of numbers of SARS-CoV-2 infections by country and age group over time, current proportions in different immune states, and potential remaining burden of hospitalisations and deaths for 19 European countries from article "Unexposed populations and potential COVID-19 burden in European countries as of 21st November 2021"</p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)
<p><strong>Sources: </strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p> </p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology is modified relative to the regular trend analysis due to the short period of analysis (weekly freqiencies)</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every week) </li> <li>Several media sources: all articles are treated equally</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of weeks since the beginning of the analysed period (January 2020) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed week (marginal change of the frequency), revealing which keywords had the biggest weekly growth</li> </ul> <p>Columns</p> <p>freq_2020_weeks (e.g. freq_2020_ww0): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p>
Students' perceived obstacles with Forced Online Distance Learning during the CoVID-19 outbreak and their preferences to continue with the introduced teaching methods after the reopening of the University of Maribor [Project documentation]
<p>The outbreak of COVID -19 forced most universities into distance education. Three didacticians and researchers from the University of Maribor, Slovenia: Kosta Dolenc, Mateja Ploj Virtič and Andrej Šorgo formed a self-initiated initiative project group during the COVID -19 epidemic and started the first project with the working title: The Side Effects of Forced Online Distance Education (FODE).</p> <p>The aim of the second study, conducted during the first wave of the epidemic in March 2020, was to investigate the response of university students to the new situation. The project documentation provided for the Forced Online Distance Learning (FODL) consists of:</p> <ul> <li>abstract,</li> <li>instrument,</li> <li>copy of the descriptive statistics,</li> <li>and SPSS dataset.</li> </ul>
Forced Continuance Intention Model of Distance Online Teaching during CoVID-19 outbreak at University of Maribor, Slovenia [Project documentation]
<p>The outbreak of COVID -19 forced most universities into distance education. Three didacticians and researchers from the University of Maribor, Slovenia: Kosta Dolenc, Mateja Ploj Virtič and Andrej Šorgo formed a self-initiated initiative project group during the COVID -19 epidemic and started the project with the working title: The Side Effects of Forced Online Distance Education (FODE).</p> <p>The aim of the first study, conducted during the first wave of the epidemic in March 2020, was to investigate the response of university teachers to the new situation. The project documentation provided for the Forced Online Distance Teaching (FODT) consist of:</p> <ul> <li>abstract,</li> <li>instrument,</li> <li>copy of the descriptive statistics, and</li> <li>SPSS dataset.</li> </ul>
Covid-19 Worldwide Data 2021 Cases and Vaccination
<p>This dataset includes information about active cases, accumulative cases, accumulative deaths, daily information, vaccination classified by country of year 2021. </p>
A Multilingual Dataset of COVID-19 Vaccination Attitudes on Twitter
<p>This dataset consists of the IDs of 2,198,090 tweets collected from Western Europe, of which 17,934 are annotated with labels indicating the originators' affective vaccination stances, including Positive (PO), Negative (NG), Positive but dissatisfaction (PD), Neutral (NE) and Off-topic (OT).</p> <p>all_tweets.txt contains all the ids of the collected tweets, annotated_tweets.txt contains the ids of the annotated tweets and the categories they are annotated to.</p> <p> </p>
Covid-19 et Intention d'utiliser les chèques psycho par les étudiants : la prise en compte du contexte dans la théorie du comportement planifié
<p>Cette base de données est issue d’une enquête quantitative par questionnaire (nombre d’observations = 460, période : mars 2021). Elle est construite sur la base de la théorie du comportement planifié. La variable finale que le modèle cherche à expliquer (la variable dépendante) est l’intention d’utiliser les chèques psycho, dispositif proposé par le gouvernement français en février 2021, par les étudiants. Ces chèques permettent aux étudiants de bénéficier de 3 consultations auprès d’un psychologue conventionné, pour un montant total de 96 €.</p> <p>Les objectifs (et les utilisations possibles) de cette BDD sont autant orientés vers les besoins des acteurs (notamment ceux en charge de la mise en œuvre du dispositif « chèque psycho » pour les étudiants ou de dispositifs similaires) que des chercheurs.</p> <p>L’objectif opérationnel est de mesurer (statistiques descriptives) et de comprendre (statistiques explicatives) les facteurs qui conduisent des étudiants à envisager d’utiliser ce dispositif ; et donc à pouvoir comprendre comment agir pour optimiser cette utilisation. L’extension récente (juin 2021) de ce dispositif aux enfants et adolescents de 3 à 17 ans (dispositif « PsyEnfantAdo ») induit un deuxième objectif opérationnel : la possibilité pour les acteurs en charge de ce nouveau dispositif de s’inspirer de la méthodologie présentée dans cette base pour piloter au mieux ce projet.</p> <p>L’objectif théorique réside a) dans la prise en compte de l’impact d’un contexte (le vécu des étudiants durant la pandémie de la Covid-19 (t leur antécédents au niveau psychologique dans la théorie du comportement planifié (TCP) et b) dans la confirmation de la place de l’identité personnelle en tant que mesure alternative de l’intention comportementale (et non en tant que variable explicative de cette intention). Les chercheurs pourront également utiliser cette base dans des méta-analyses sur la TCP, la prise en compte du contexte dans la compréhension de l’intention comportementale et les impacts de la Covid-19.</p>
Dataset for evaluation of unrealistic optimism in time of pandemic COVID-19 on a Polish sample
<p>This dataset contains the data used in a project called "Unrealistic optimism in the eye of the storm. Positive bias towards the consequences of COVID-19 during the second and third waves of the pandemic. ". The project concerns the occurrence of a cognitive bias - unrealistic optimism - with regard to contracting the coronavirus. The following information are attached to the dataset: codebooks with variables' names; analysis codes to replicate our results; supplementary materials with the description of procedures. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.