Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,173
datasets available to search
ShareScore release 0.9.0
Dataset results
1,173 results for “Pandemic”
Data for "The effects of weather and mobility on respiratory viruses dynamics before and during the COVID-19 pandemic in the USA and Canada".
<p>Epidemiological and mobility data analysed in the paper "The effects of weather and mobility on respiratory viruses dynamics before and during the COVID-19 pandemic in the USA and Canada".</p>
A comparative dataset on public perceptions of multiple risks during the COVID-19 pandemic in Italy and Sweden
<p>These datasets are the result of two nation-wide surveys conducted in Italy and Sweden in August 2020 and in november 2020. The surveys (which are identical in the two rounds) explore the respondents' risk perception, preparedness, knowledge, and experience regarding a set of hazards, namely: epidemics, floods, droughts, earthquakes, wildfires, terror attacks, domestic violence, economic crises, and climate change. </p> <p>The data files include the questionnaire survey (the Italian and Swedish versions as well as the English translation) and the two datasets of all the answers to the two surveys. Each column in the dataset refers to an item in the survey (e.g. a question or a sub-question), and each row represents a single respondent. </p> <p>For additional information on the August 2020 dataset, see <a href="https://www.nature.com/articles/s41597-020-00778-7">Mondino et al. (2020)</a>.</p>
Comparison of pandemic excess mortality in 2020-2021 across different empirical calculations
<p>Different modeling approaches can be used to calculate excess deaths for the COVID-19 pandemic period. We compared 6 calculations of excess deaths (4 previously published and two new ones that we performed with and without age-adjustment) for 2020-2021. With each approach, we calculated excess deaths metrics and the ratio R of excess deaths over recorded COVID-19 deaths. The main analysis focused on 33 high-income countries with weekly deaths in the Human Mortality Database (HMD at mortality.org) and reliable death registration. Secondary analyses compared calculations for other countries, whenever available. Across the 33 high-income countries, excess deaths were 2.0-2.8 million without age-adjustment, and 1.6-2.1 million with age-adjustment with large differences across countries. In our analyses after age-adjustment, 8 of 33 countries had no overall excess deaths; there was a death deficit in children; and 0.478 million (29.7%) of the excess deaths were in people <65 years old. In countries like France, Germany, Italy, and Spain excess death estimates differed 2 to 4-fold between highest and lowest figures. The R values’ range exceeded 0.3 in all 33 countries. In 16 of 33 countries, the range of R exceeded 1. In 25 of 33 countries some calculations suggest R>1 (excess deaths exceeding COVID-19 deaths) while others suggest R<1 (excess deaths smaller than COVID-19 deaths). Inferred data from 4 evaluations for 42 countries and from 3 evaluations for another 98 countries are very tenuous Estimates of excess deaths are analysis-dependent and age-adjustment is important to consider. Excess deaths may be lower than previously calculated. </p>
Data for the article "Professionalism, emotional wellbeing, and dropout intention in health professions students during the pandemic"
<p>Dataset from a study of attitudes and perceptions of medicine and nursing students in Peru during the COVID-19 pandemic. Survey was applied from 2020-07-24 to 2021-04-16.</p> <p>This dataset is described in the article: </p> <p>Castagnetto, J.M., Hancco-Monrroy, D.E., Caballero-Apaza, L.M. <em>et al.</em> Professionalism, emotional wellbeing, and dropout intention in health professions students during the pandemic. <em>Sci Data</em> <strong>12</strong>, 1259 (2025). <a href="https://doi.org/10.1038/s41597-025-05508-5">https://doi.org/10.1038/s41597-025-05508-5</a> (<a href="https://www.nature.com/articles/s41597-025-05508-5">https://www.nature.com/articles/s41597-025-05508-5</a>)</p>
TomoBreast randomized clinical trial's lung-heart outcomes and mortality through the 2020 COVID-19 pandemic: data and software
<p>Dataset and R script to reproduce the analyses of the manuscript:</p> <p>Vinh-Hung V, Gorobets O, Adriaenssens N, Van Parijs H, Storme G, Verellen D, Nguyen NP, Magne N, De Ridder M.</p> <p><strong>Lung-heart outcomes and mortality through the 2020 COVID-19 pandemic in a prospective cohort of breast cancer radiotherapy patients.</strong></p> <p>Cancers 2022; 14(24):6241. https:// doi.org/10.3390/cancers14246241</p> <p>https://www.mdpi.com/2072-6694/14/24/6241</p> <p>PubMed: PMID: 36551726</p> <p>PMCID: PMC9777311</p> <p>Info on the variables in file "aelq6_public.R"</p> <p>reproduced in "aelq_2_3_readme.txt":</p> <p>"aelq2_base2.txt" = baseline characteristics.</p> <p>"aelq3.txt" = longitudinal maesurements.</p> <p>Variables in "aelq2_base2.txt":</p> <p>"<strong>aelq2_base2.txt</strong>" = baseline characteristics. <br># Age at randomization, years. <br># RTdose: cf TomoBreast papers. <br># 51 Gy = hypofractionated, simultaneous integrated boost<br># 42 Gy = hypofractionated, no boost, mastectomy cases only<br># 50 Gy = conventional, no boost, mastectomy cases only<br># 66 Gy = conventional, sequential boost<br># Weight kg, Height cm, <br># Detection 1=found by screening (senology follow-up/controle)<br># 2=found by symptoms (pain, palpable)<br># 9=unknown<br># Smoker 0= Not smoker<br># 1= Smoker<br># 2=ex-smoker<br># Mastectomy (and other binary coded) 1= yes<br># chemosched 0=none<br># 1= planned after RT (sequential)<br># 2= prior to RT and is finished (sequential)<br># 3= chemo is on-going or is planned to start with RT (concomitant)<br># hormonetherapy 0=no<br># 1=tamoxifen (nolvadex)<br># 2=Femara (Letrozole)<br># 3=zoladex<br># 4=tamoxifen + zoladex<br># Laterality 1,=Right, 2=Left, 3=Bilateral<br># LengthFU: length of follow-up, days from randomization</p> <p>"<strong>aelq3.txt</strong>" = longitudinal maesurements.<br># "Nr" = Case ID<br># "Time" in days from origin (origin =date of randomization), <br># if negative =before randomization<br># "KPS" "Weight" <br># "Died" "LocalRec" "Metast" "NewPrim" = binary code, 0=no, 1=yes<br># "fAEBreast" "fAEHeart" "fAELung" "fAEOther" <br># fAE = freedom from breast, heart, lung, other adverse event score<br># "LVEF2" = ejection fraction, %<br># "MacIver" = estimated cardiac strain</p> <p># the following are pulmonary function tests, untransformed units<br># "FVC", "FEV1", "PEF", "VC", "TLC", "RV", "FRC", "Raw", "sRaw", "DLCO",<br># "VA", "PF"</p> <p># "fDY", "fFA", "fPA" = freedom from dyspnea, from fatigue, from pain<br># range 0 to 100 (best)<br># see papers:</p> <p># Van Parijs, H.; Vinh-Hung, V.; Fontaine, C.; Storme, G.; Verschraegen, C.;<br># Nguyen, D.M.; Adriaenssens, N.; Nguyen, N.P.; Gorobets, O.; De Ridder, M.<br># Cardiopulmonary-related patient-reported outcomes in a randomized clinical<br># trial of radiation therapy for breast cancer. BMC Cancer 2021, 21, 1177,<br># doi:10.1186/s12885-021-08916-z.</p> <p># preprint:<br># Van Parijs, H.; Cecilia-Joseph, E.; Gorobets, O.; Storme, G.; <br># Adriaenssens, N.; Heyndrickx, B.; Verschraegen, C.; Nguyen, N.P.;<br># De Ridder, M.; Vinh-Hung, V. Lung-heart toxicity in a randomized <br># clinical trial of hypofractionated image guided radiation therapy for<br># breast cancer. Preprints 2022, 202212, 0214.<br># https://doi.org/10.20944/preprints202212.0214.v1</p> <p># <br># "Year" = year of the observation<br># example: randomized 1/1/2011, measurement done 1/31/2011, time = 30 days,<br># Year =2011<br>#<br> </p>
The Impact of the COVID-19 Pandemic On Cities. A Scoping Review Protocol
<p>The aim of the scoping review is to map out evidence based research on the Covid-19 pandemic impact on the European cities. The review questions touch three broad areas of interest:</p> <ol> <li>the aspects of urban life described and analysed in publications on the impact of the pandemic on cities</li> <li>the aspects of urban life that are described in terms of crisis, breakdown, turnaround, etc. (crisis, disruption, slump, shift…) in such studies</li> <li>theoretical and methodological approaches applied in such studies</li> </ol> <p>The search was conducted in June 2022, with the final body of literature consisting of 3,994 publication references from EBSCOhost, APA Psyc, Scopus, Web of Science, Proquest, Wiley, Sage, JSTOR, Tailor&Francis, Oxford Journals databases (Fig. 1). The following English words were searched for in titles, abstracts and keywords in the databases: (pandemic OR ‘Covid-19’) AND (city OR cities OR urban*). We used the following criteria for articles to be included in the study: 1) peer and non-peer-reviewed empirical papers in journals published in English from January 2019 to June 2022; 2) included studies where the impact of COVID-19 pandemic on European city/cities was an explicit variable of interest; 3) contained analysis of empirical data on cities or urban life retrieved or collected within and explicitly addressing the COVID-19 pandemic; 4) addressed the social, cultural, economic, political and socio-geographical aspects of a city. We excluded from our sample papers that were: 1) theoretical and opinion literature, media press releases, reports, MA dissertations and PhD theses; 2) secondary research papers (reviews, meta-analyses); 3) papers not in English; 4) studies about non-European cities; 5) studies which do not explicitly address the impact of the COVID-19 pandemic on cities; 6) studies addressing a city as a variable of secondary importance; 7) studies outside the scope of the COVID-19 pandemic, published before December 2019; 8) studies not addressing the social or human aspects of urban life.</p> <p>The final database of coded documents consisted of 138 empirical articles presenting findings on the impact of the COVID-19 pandemic on European cities. </p>
Pandemic severity indicator for COVID-19 in Germany dataset
<p>The datasets included in this repository represent a pandemic severity indicator for the COVID-19 pandemic in Germany based on a composite indicator for the years 2020 and 2021. The pandemic severity index consists of three indicators: the incidence of patients tested positive for COVID-19, the incidence of patients with COVID-19 in intensive care, and the incidence of registered deaths due to COVID-19. The datasets have been developed within the CODIFF project (Socio-Spatial Diffusion of COVID-19 in Germany) at Leibniz Insitute for Research on Society and Space. The project received funding by Deutsche Forschungsgemeinschaft (DFG, project number 492338717). The datasets have been used in the following publications, in which further methodological details on the indicator can be found:</p> <ul> <li><a href="https://doi.org/10.1101/2023.02.17.23286084">Stabler, M., & Kuebart, A. (2023). Tempo-spatial dynamics of COVID-19 in Germany: A phase model based on a pandemic severity indicator. <em>medRxiv</em>, 2023-02</a>.</li> <li><a href="https://doi.org/10.1016/j.sste.2023.100605">Kuebart, A., & Stabler, M. (2023). Waves in time, but not in space – An analysis of pandemic severity of COVID-19 in Germany. <em>Spatial and Spatio-temporal Epidemiology</em>, 2023.</a></li> </ul> <p>This repository consists of two files:</p> <p><strong>pandemic_severity_germany </strong></p> <p>This table contains the composite indicator for daily pandemic severity for Germany on the national scale as well as the three sub-indicators for each day between 2020-03-01 and 2021-12-31. The sub-indicators were sourced from the <a href="https://github.com/robert-koch-institut">Robert Koch Institute</a>, the German government agency responsible for disease control and prevention.</p> <p><strong>pandemic_severity_counties</strong></p> <p>This table contains the composite indicator for daily pandemic severity for Germany on the level of the 400 individual counties, as well as the three sub-indicators for each day between 2020-03-01 and 2021-12-31. The sub-indicators were sourced from the <a href="https://github.com/robert-koch-institut">Robert Koch Institute</a>, the German government agency responsible for disease control and prevention. The counties can be identified by name (kreis) or by county identification number (ags5)</p>
PANDEM-2 European COVID-19 training data set
<p>The PANDEM-2 COVID-19 European training dataset is a large collection of time series either of real or realistic synthetic (generated) data and indicators associated with the European pandemic response to the COVID-19 pandemic. It is intended to be used for training in pandemic management.</p> <p> </p> <p>This dataset is the result of a data gathering requirement process for pandemic management involving feedback and inputs from several public health and first responder professionals as well as researchers and military personnel directly involved in the European COVID-19 pandemic response. This work is part of the PANDEM-2 project funded by the <em>Horizon 2020 Secure Societies</em> program. To collect this data, an open source software was developed named PANDEM-Source allowing reproducibility and customisation of this dataset. </p> <p> </p> <p>The dataset includes indicators for cases, deaths, hospitalisation, testing and laboratory data including pathogen genomic information, vaccination, non-pharmaceutical interventions, participatory surveillance, social media, flights resources (human and material such as beds or vaccines), and contact tracing activities. When no open available data was found, realistic synthetic data and indicators were generated with the goal of producing a data set to be used for pandemic management training. </p> <p>The project received funding from the European Union’s Horizon 2020 Research and Innovation programme under the Grant Agreement No. 883285. The material presented and views expressed here are the responsibility of the author(s) only. The EU Commission takes no responsibility for any use made of the information set out.<br> References</p> <p> </p> <p> </p>
Adjoint-based Data Assimilation of an Epidemiology Model for the Covid-19 Pandemic in 2020 --- Data Files
<p>New data on github:</p> <p>https://github.com/sesterhenn/Corona-DataAssimilation</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p>doi://10.5281/zenodo.3732292</p> <p>https://zenodo.org/record/3733244</p>
The Daily Life of Software Engineers during the COVID-19 Pandemic -- Replication Package
<p>Following the onset of the COVID-19 pandemic and subsequent lockdowns, software engineers' daily life was disrupted and abruptly forced into remote working from home. This change deeply impacted typical working routines, affecting both well-being and productivity. Moreover, this pandemic will have long-lasting effects in the software industry, with several tech companies allowing their employees to work from home indefinitely if they wish to do so. Therefore, it is crucial to analyze and understand how a typical working day looks like when working from home and how individual activities affect software developers' well-being and productivity. We performed a two-wave longitudinal study involving almost 200 globally carefully selected software professionals, inferring daily activities with perceived well-being, productivity, and other relevant psychological and social variables. Results suggest that the time software engineers spent doing specific activities from home was similar when working in the office. (e.g., coding > emails > code review > networking). However, we also found some meaningful mean differences. The amount of time developers spent on each activity was unrelated to their well-being, perceived productivity, and other variables. We conclude that working remotely is not per se a challenge for organizations or developers.</p>
Dataset for: Pre-pandemic artificial MERS analog of polyfunctional SARS-CoV-2 S1/S2 furin cleavage site domain is unique among spike proteins of genus Betacoronavirus
<table> <tbody> <tr> <th> </th> <td> <div> <h3><strong>Data File Descriptions and Methods</strong></h3> <ol> <li><strong>Data file 1 [betacov_matching_IPR042578.fasta]</strong>: Representative set of 2,465 betacoronavirus S protein overlapping homologous superfamily sequences retrieved in fasta format on 4 December 2022 from the InterPro repository at https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/.<br><br></li> <li><strong>Data File 2 [betacov_matching_IPR042578_motif.fasta]</strong>: With Data File 1 as input, extracted 98,122 furin cleavage site (FCS) output motifs of 20 amino acids length, including overlapping and redundant sequences, produced with the FindFur algorithm with preset parameters as described by (Gu, 2020). FindFur as used was deposited on 15 December 2020 at the GitHub software repository at https://github.com/chwisteeng/FindFur.<br><br></li> <li><strong>Data File 3 [table_s1s2_hits_betacov_polyf.pdf]</strong>: Compiled summary table of sequence hits (PDF) of spike S1/S2 domains across genus <em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></li> <li> <p><strong>Data File 4 [table_s1s2_hits_betacov_polyf.xlsx]</strong>: Compiled summary table of sequence hits (MS Excel) of spike S1/S2 domains across genus <em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></p> </li> <li> <p><strong>Data File 5 [betacov_s1s2_nls_pat7_furin_psort.txt]: </strong>Nuclear localization signal (NLS) detection output for 5 representative betacoronavirus spike sequence domains, including the positive hits for pat7 in SARS-CoV-2 and for MERS-MA30 CoV. NLS predictions used the PSORT algorithm available as a webservice at https://wolfpsort.hgc.jp/ which is based on the work of Nakai and Horton (Nakai and Horton, 1999). Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 6 [betacov_s1s2_oglyc_netogly.txt]: </strong>Detection output for 5 representative betacoronavirus spike sequence domains tested for Thr/Ser O-glycosite residue pairs with the standard prediction software NetOGlyc4.0 (Steentoft et al., 2013) as available at https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/. Positive hits have scores above 0.5. Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 7 [betacov_s1s2_nls_pat7_furin_blastp.txt]</strong>: Comprehensive sequence database searches using were performed using the NCBI protein BLAST (blastp) algorithm with webservice available at https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins. The following blastp search parameters and settings were used: Word size=2; Expect value=200000; Hitlist size=500; Gapcosts=9,1; Matrix=PAM30; Filter string=F; Genetic Code=1;Window Size=40; Threshold=11; Composition-based stats=0; Database Posted date=Jan 19, 2023 2:59 AM; Number of letters=17,117,563; Number of sequences=10,766; Entrez query: Includes: Betacoronavirus (taxid:694002); Excludes: SARS-CoV-2 (taxid:2697049). The six polyfunctional input query consensus motif sequences were TXXPR(K/H/R)XRSX and TXXPRX(K/H/R)RSX.</p> </li> </ol> <h3><strong>References</strong></h3> <p>Gu, C., 2020. FindFur: A Tool for Predicting Furin Cleavage Sites of Viral Envelope Substrates. Master’s Thesis, San Jose State University, CA, USA. doi: <a href="https://doi.org/10.31979/etd.4ahv-9jya">10.31979/etd.4ahv-9jya</a> </p> <p>Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID Core and Curation Team; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat Methods. 2023. 20(4):512-522. doi: <a href="https://doi.org/10.1038/s41592-023-01769-3">10.1038/s41592-023-01769-3</a>.</p> <p>Nakai, K., Horton, P., 1999. PSORT: a program for detecting sorting signals in proteins and predicting their subcellular localization. Trends Biochem Sci 24, 34–36. doi: <a href="https://doi.org/10.1016/s0968-0004(98)01336-x">10.1016/s0968-0004(98)01336-x</a></p> <p>Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. Precision mapping of the human O-GalNAc glycoproteome through SimpleCell technology. EMBO J 32, 1478–1488. doi: <a href="https://doi.org/10.1038/emboj.2013.79">10.1038/emboj.2013.79</a></p> </div> </td> </tr> </tbody> </table>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)
<p><strong>Sources: </strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p> </p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology is modified relative to the regular trend analysis due to the short period of analysis (weekly freqiencies)</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every week) </li> <li>Several media sources: all articles are treated equally</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of weeks since the beginning of the analysed period (January 2020) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed week (marginal change of the frequency), revealing which keywords had the biggest weekly growth</li> </ul> <p>Columns</p> <p>freq_2020_weeks (e.g. freq_2020_ww0): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p>
Dataset for evaluation of unrealistic optimism in time of pandemic COVID-19 on a Polish sample
<p>This dataset contains the data used in a project called "Unrealistic optimism in the eye of the storm. Positive bias towards the consequences of COVID-19 during the second and third waves of the pandemic. ". The project concerns the occurrence of a cognitive bias - unrealistic optimism - with regard to contracting the coronavirus. The following information are attached to the dataset: codebooks with variables' names; analysis codes to replicate our results; supplementary materials with the description of procedures. </p>
Dataset: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter
<p>This is the dataset, trained model, and software companion for the paper titled: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter accepted for the Workshop on Data for the Wellbeing of Most Vulnerable of the ICWSM 2022 conference.</p> <p>The COVID-19 pandemic has shown a measurable increase in the usage of sinophobic comments or terms on online social media platforms. In the United States, Asian Americans have been primarily targeted by violence and hate speech stemming from negative sentiments about the origins of the novel SARS-CoV-2 virus. While most published research focuses on extracting these sentiments from social media data, it does not connect the specific news events during the pandemic with changes in negative sentiment on social media platforms. In this work we combine and enhance publicly available resources with our own manually annotated set of tweets to create machine learning classification models to characterize the sinophobic behavior. We then applied our classifier to a pre-filtered longitudinal dataset spanning two years of pandemic related tweets and overlay our findings with relevant news events.</p>
Pandemic Border Discourses Dataset and Codebook - Swiss Case
<p>The Pandemic Border Discourses project identifies and compares the evolution of discourses restricting internal and external mobility in Europe as the Covid-19 pandemic is unfolding. It is designed to show how political actors use discourses to justify their decisions in emergency situations, and analyse whether and how unforeseen systemic pressure disrupts bordering discourses and practices. It contributes to a better understanding of the political, social and economic issues driving policy decision in times of crisis, above all the tension between national interest and transnational solidarity.</p> <p>This coding manual explains our data collection strategy and introduces the variables of the dataset. Building on a core-sentence analysis method, we collect and analyse institutional discourses about mobility during the Covid-19 crisis on Twitter.</p> <p> </p>
Supplementary Material for "'A Certain Enemy Robbed Me of My Life': Medieval Riddles, Digital Transformations, and Pandemic Pedagogy"
<p>Assignment, games, and illustrations associated with "'A Certain Enemy Robbed Me of My Life': Medieval Riddles, Digital Transformations, and Pandemic Pedagogy"</p>
Pollinator-flower interactions in gardens during the COVID-19 pandemic lockdown of 2020
<p>During the main COVID-19 pandemic lockdown period of 2020 an impromptu set of pollination ecologists came together via social media and personal contacts to carry out standardised surveys of the flower visits and plants in their gardens. The surveys involved 67 rural, suburban and urban gardens, of various sizes, ranging from 61.18<sup>o</sup> North in Norway to 37.96<sup>o</sup> South in Australia and resulted in a data set of 25,174 rows long and comprising almost 47,000 visits to flowers, as well as records of plants that were not visited by pollinators. In this first publication from the project we present a brief description of the data and make it freely available for any researchers to use in the future, the only restriction being that they cite this paper in the first instance. As well as producing a data set that we hope will be widely used in the future, the project helped enormously with the health and mental wellbeing of the participants, a by-product of ecological field work that cannot be over-estimated.</p>
Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022.
<p>Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022. These data were used to estimate the time-varying reproduction number (R) in South Africa, as described in https://www.medrxiv.org/content/10.1101/2022.07.22.22277932v1.full.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.