Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,674
datasets available to search
ShareScore release 0.7.1
Dataset results
9,674 results for “COVID-19”
Data for "Paris Agreement requires substantial, broad, and sustained policy efforts beyond COVID-19 recovery packages"
<p>This dataset contains the underlying data for the following publication: Tanaka, K., C. Azar, O. Boucher, P. Ciais, Y. Gaucher, D. J. A. Johansson (2022) Paris Agreement requires substantial, broad, and sustained policy efforts beyond COVID-19 public stimulus packages. <em>Climatic Change</em> <strong>172, </strong>1 (2022). https://doi.org/10.1007/s10584-022-03355-6</p> <p>Earlier manuscripts were published as a preprint. https://arxiv.org/abs/2104.08342</p>
Modeling robust COVID-19 intensive care unit occupancy thresholds for imposing mitigation to prevent exceeding capacities
<p>Simulation output files for 'Modeling robust COVID-19 intensive care unit occupancy thresholds for imposing mitigation to prevent exceeding capacities'.</p> <p>Simulating COVID-19 transmission and hospital burden to assess at which intensive care unit (ICU) occupancies mitigation, that reduces transmission, needs to be triggered to avoid exceeding ICU capacity limits, using the city of Chicago, Illinois as an example.</p> <p>Manuscript is under review for scientific publication, (see <a href="https://www.medrxiv.org/content/10.1101/2021.06.27.21259530v1">preprint on medRxiv</a>) and scripts are available from the GitHub repository at https://github.com/numalariamodeling/ICUtrigger_covid_chicago_paper_2021. </p> <p>Simulation output files uploaded per scenario including projected COVIID-19 transmission and burden trajectories for Chicago city for March 2020 to May 2021 per day.</p> <p>Simulation scenarios:</p> <p><reopening % above ICU capacity>_<delay after reaching ICU threshold>_<%mitigation>_<common simulation name> i.e. `50perc_1daysdelay_pr6_triggeredrollback_reopen`</p> <ul> <li>`emodl` file <ul> <li>required file for COVID-19 transmission model in the <a href="https://docs.idmod.org/projects/cms/en/latest/index.html">Compartmental Modeling Software</a> (see <a href="https://github.com/numalariamodeling/ICUtrigger_covid_chicago_paper_2021">GitHub repository</a> for details)</li> </ul> </li> <li>sampled_parameters.csv <ul> <li>simulation input and scenario parameters, (nrow=4400, 400 unique parameter combinations * 11 scenario values)</li> </ul> </li> <li>rt_trajectoriescovidregion_11.csv <ul> <li>estimated reproductive numbers per trajectory for complete timeline per day</li> </ul> </li> <li>trajectoriesDat_region_11_traces.csv <ul> <li>filtered to include top 100 trajectories fitted to ICU data</li> </ul> </li> <li>trajectoriesDat_region_trimfut.csv <ul> <li>truncated to only include projections after September 1st 2020</li> </ul> </li> </ul> <p>The folder `mainfigures_csvs.zip` includes processed simulation output data for the publication figures.</p>
Deciphering the Neurosensory Olfactory Pathway and Associated Neo-Immunometabolic Vulnerabilities Implicated in COVID-Associated Mucormycosis (CAM) and COVID-19 in a Diabetes Backdrop—A Novel Perspective
<p>Raw data files of transcriptomic profiling experiments, which form the basis for our publication (https://www.mdpi.com/2673-4540/3/1/13).</p>
Where2Test Saxony-Czechia COVID-19 new cases dataset
<p>Data in the repository were used in the study "Fine-scale variation in the effect of national border on COVID-19 spread: A case study of the Saxon-Czech border region", published in <a href="https://www.sciencedirect.com/journal/spatial-and-spatio-temporal-epidemiology">Spatial and Spatio-temporal Epidemiology</a>.</p> <p>This repository consists of two files:</p> <p><strong>saxony-westczechia_cases7</strong></p> <p>Weekly numbers of new COVID-19 cases in all municipalities in Saxony and Northwestern Czechia (Liberec, Ústí nad Labem, and Karlovy Vary regions) in the first half of 2021. Data are extracted from the websites <a href="https://www.coronavirus.sachsen.de">coronavirus.sachsen</a> and <a href="https://onemocneni-aktualne.mzcr.cz/covid-19">onemocneni-aktualne.mzcr.cz/covid-19</a>. The missing values were interpolated, and daily values were recalculated to weekly values.</p> <p><strong>municipalities</strong></p> <p>The second file consists of a list of all municipalities with their names, geometries, and population values. For Germany, we used the dataset <a href="https://hub.arcgis.com/datasets/esri-de-content::gemeindegrenzen-2018-mit-einwohnerzahl/about">"Gemeindegrenzen 2018 mit Einwohnerzahl"</a> (© GeoBasis-DE / BKG, Statistisches Bundesamt (Destatis) (2020), <a href="http://www.govdata.de/dl-de/by-2-0">dl-de/by-2-0</a>) as a source of geometries and population sizes of the municipalities (“<em>Gemeinde”</em>) in Saxony. Czech population numbers on the municipality level ("obec") were taken from the <a href="https://www.czso.cz/csu/czso/population-of-municipalities-1-january-2021">Czech Statistical Office</a>, while the geometries were obtained from <a href="https://www.cuzk.cz/ruian/RUIAN.aspx">RÚIAN</a> (@<a href="http://geoportal.cuzk.cz">Czech Office for Surveying, Mapping and Cadastre</a>, 2021). To keep the same geometry detail on both sides of the borders, we applied the Douglas-Peucker simplification algorithm implemented in the Python library <a href="https://github.com/mattijn/topojson">TopoJSON</a>.</p>
Data from: Elevated fires during COVID-19 lockdown and the vulnerability of protected areas
<p><strong>Related article:</strong> Johanna Eklund, Julia P G Jones, Matti Räsänen, Jonas Geldmann, Ari-Pekka Jokinen, Adam Pellegrini, Domoina Rakotobe, O. Sarobidy Rakotonarivo, Tuuli Toivonen, and Andrew Balmford. Elevated fires during COVID-19 lockdown and the vulnerability of protected areas. Nature Sustainability (2022) https://doi.org/10.1038/s41893-022-00884-x.</p> <p><strong>In this dataset:</strong></p> <p>This dataset contains information about monthly fire incidence and precipitation for the protected areas of Madagascar from January 2012 to December 2020. The fire data is sourced from NASA’s Visible Infrared Imaging Radiometer Suite (VIIRS) 375 m active fire product and the precipitation data from the Global Precipitation Measurement (GPM) mission (for years 2016-2020) and its predecessor The Tropical Rainfall Measuring Mission (TRMM) (for years 2011-2015) at spatial resolution 10 km. The fire and precipitation data was overlayed with the protected area polygons of the June 2020 release of the World Database of Protected Areas. For sources and more details on how the data was compiled see the related article. The data can be used to inspect temporal dynamics of wildfires inside protected areas and for informing adaptive protected area management and planning.</p> <p><strong>Please cite this dataset as:</strong></p> <p>Johanna Eklund, Julia P G Jones, Matti Räsänen, Jonas Geldmann, Ari-Pekka Jokinen, Adam Pellegrini, Domoina Rakotobe, O. Sarobidy Rakotonarivo, Tuuli Toivonen, and Andrew Balmford. Elevated fires during COVID-19 lockdown and the vulnerability of protected areas. Nature Sustainability (2022) https://doi.org/10.1038/s41893-022-00884-x.</p> <p><strong>Column names</strong></p> <p>NAME: Name of protected area</p> <p>Fires_sum: Number of observed fires (VIIRS)</p> <p>Month: Month</p> <p>Year: Year</p> <p>Precipitation: Precipitation (mm)</p> <p>Plag_1:Plag_12: Precipitation during previous month; 2 months ago; 3 months ago…12 months ago</p> <p>YEAR_CREAT: Year of establishment of protected area</p> <p>Biome: Biome</p> <p>REP_AREA: Area of protected area (km<sup>2</sup>)</p> <p>Fires_per_km2: Fires per km<sup>2</sup></p> <p>Prec_acc_12m: Accumulated precipitation during the last 12 months</p> <p>fBiome: Biome as factor</p> <p>fNAME: Name as factor</p> <p>sPrecipitation: Precipitation (scaled; see Methods section of article)</p> <p>sPlag_1: Precipitation in previous month (scaled; see Methods section of article)</p> <p>sPrec_acc_12m: Accumulated precipitation during the last 12 months (scaled; see Methods section of article)</p> <p>Pred_Zinb_1a: Predicted fires (see Methods section of article)</p> <p>Diff_Zinb_1a: Difference: Observed fires - predicted fires</p> <p>Year_pred: Year for prediction</p> <p><strong>License</strong><br> Creative Commons Attribution 4.0 International.</p>
PANACEA dataset - Heterogeneous COVID-19 Claims
<p>The peer-reviewed publication for this dataset has been presented in the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), and can be accessed here: https://arxiv.org/abs/2205.02596. Please cite this when using the dataset.</p> <p> </p> <p>This dataset contains a heterogeneous set of True and False COVID claims and online sources of information for each claim.</p> <p> </p> <p>The claims have been obtained from online fact-checking sources, existing datasets and research challenges. It combines different data sources with different foci, thus enabling a comprehensive approach that combines different media (Twitter, Facebook, general websites, academia), information domains (health, scholar, media), information types (news, claims) and applications (information retrieval, veracity evaluation).</p> <p> </p> <p>The processing of the claims included an extensive de-duplication process eliminating repeated or very similar claims. The dataset is presented in a LARGE and a SMALL version, accounting for different degrees of similarity between the remaining claims (excluding respectively claims with a 90% and 99% probability of being similar, as obtained through the MonoT5 model). The similarity of claims was analysed using BM25 (Robertson et al., 1995; Crestani et al., 1998; Robertson and Zaragoza, 2009) with MonoT5 re-ranking (Nogueira et al., 2020), and BERTScore (Zhang et al., 2019).</p> <p> </p> <p>The processing of the content also involved removing claims making only a direct reference to existing content in other media (audio, video, photos); automatically obtained content not representing claims; and entries with claims or fact-checking sources in languages other than English.</p> <p> </p> <p>The claims were analysed to identify types of claims that may be of particular interest, either for inclusion or exclusion depending on the type of analysis. The following types were identified: (1) Multimodal; (2) Social media references; (3) Claims including questions; (4) Claims including numerical content; (5) Named entities, including: PERSON − People, including fictional; ORGANIZATION − Companies, agencies, institutions, etc.; GPE − Countries, cities, states; FACILITY − Buildings, highways, etc. These entities have been detected using a RoBERTa base English model (Liu et al., 2019) trained on the OntoNotes Release 5.0 dataset (Weischedel et al., 2013) using Spacy.</p> <p> </p> <p>The original labels for the claims have been reviewed and homogenised from the different criteria used by each original fact-checker into the final True and False labels.</p> <p> </p> <p>The data sources used are:</p> <p>- The CoronaVirusFacts/DatosCoronaVirus Alliance Database. https://www.poynter.org/ifcn-covid-19-misinformation/</p> <p>- CoAID dataset (Cui and Lee, 2020) https://github.com/cuilimeng/CoAID</p> <p>- MM-COVID (Li et al., 2020) https://github.com/bigheiniu/MM-COVID</p> <p>- CovidLies (Hossain et al., 2020) https://github.com/ucinlp/covid19-data</p> <p>- TREC Health Misinformation track https://trec-health-misinfo.github.io/</p> <p>- TREC COVID challenge (Voorhees et al., 2021; Roberts et al., 2020) https://ir.nist.gov/covidSubmit/data.html</p> <p> </p> <p>The LARGE dataset contains 5,143 claims (1,810 False and 3,333 True), and the SMALL version 1,709 claims (477 False and 1,232 True).</p> <p> </p> <p>The entries in the dataset contain the following information:</p> <p>- Claim. Text of the claim.</p> <p>- Claim label. The labels are: False, and True.</p> <p>- Claim source. The sources include mostly fact-checking websites, health information websites, health clinics, public institutions sites, and peer-reviewed scientific journals.</p> <p>- Original information source. Information about which general information source was used to obtain the claim.</p> <p>- Claim type. The different types, previously explained, are: Multimodal, Social Media, Questions, Numerical, and Named Entities.</p> <p> </p> <p>Funding. This work was supported by the UK Engineering and Physical Sciences Research Council (grant no. EP/V048597/1, EP/T017112/1). ML and YH are supported by Turing AI Fellowships funded by the UK Research and Innovation (grant no. EP/V030302/1, EP/V020579/1).</p> <p> </p> <p>References</p> <p>- Arana-Catania M., Kochkina E., Zubiaga A., Liakata M., Procter R., He Y.. Natural Language Inference with Self-Attention for Veracity Assessment of Pandemic Claims. NAACL 2022 https://arxiv.org/abs/2205.02596</p> <p>- Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at trec-3. Nist Special Publication Sp,109:109.</p> <p>- Fabio Crestani, Mounia Lalmas, Cornelis J Van Rijsbergen, and Iain Campbell. 1998. “is this document relevant?. . . probably” a survey of probabilistic models in information retrieval. ACM Computing Surveys (CSUR), 30(4):528–552.</p> <p>- Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Now Publishers Inc.</p> <p>- Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pre-trained sequence-to-sequence model. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, pages 708–718.</p> <p>- Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.</p> <p>- Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.</p> <p>- Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23.</p> <p>- Limeng Cui and Dongwon Lee. 2020. Coaid: Covid-19 healthcare misinformation dataset. arXiv preprint arXiv:2006.00885.</p> <p>- Yichuan Li, Bohan Jiang, Kai Shu, and Huan Liu. 2020. Mm-covid: A multilingual and multimodal data repository for combating covid-19 disinformation.</p> <p>- Tamanna Hossain, Robert L. Logan IV, Arjuna Ugarte, Yoshitomo Matsubara, Sean Young, and Sameer Singh. 2020. COVIDLies: Detecting COVID-19 misinformation on social media. In Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020, Online. Association for Computational Linguistics.</p> <p>- Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. Trec-covid: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum, volume 54, pages 1–12. ACM New York, NY, USA.</p>
Dataset: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter
<p>This is the dataset, trained model, and software companion for the paper titled: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter accepted for the Workshop on Data for the Wellbeing of Most Vulnerable of the ICWSM 2022 conference.</p> <p>The COVID-19 pandemic has shown a measurable increase in the usage of sinophobic comments or terms on online social media platforms. In the United States, Asian Americans have been primarily targeted by violence and hate speech stemming from negative sentiments about the origins of the novel SARS-CoV-2 virus. While most published research focuses on extracting these sentiments from social media data, it does not connect the specific news events during the pandemic with changes in negative sentiment on social media platforms. In this work we combine and enhance publicly available resources with our own manually annotated set of tweets to create machine learning classification models to characterize the sinophobic behavior. We then applied our classifier to a pre-filtered longitudinal dataset spanning two years of pandemic related tweets and overlay our findings with relevant news events.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Literature sources
<p>The spreadsheet in the present dataset (CSV format) includes the sources considered during the literature review stage for the report: From intent to impact: Investigating the effects of open sharing commitments. Please note that not all sources in this deposit have been referenced in the above-mentioned report and that the report may include additional sources</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Survey responses
<p>The spreadsheets in the present dataset (CSV format) include the anonymised responses to our online survey of signatories of the Joint Statement on open research and data sharing. Responses have been split into quantitative responses (i.e., closed survey questions) and qualitative responses (i.e., free text survey questions).</p> <p>This data has been used to inform our final report, which is available in our <a href="https://zenodo.org/communities/data-sharing-in-public-health-emergencies">Zenodo Project Community</a>.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Thematic coding of qualitative research findings
<p>The spreadsheet in the present dataset (CSV format) includes the anonymised thematic coding that has been applied to our interview and literature review findings to inform the preparation of the report: From intent to impact: Investigating the effects of open sharing commitments.</p> <p>The thematic coding has been applied by using <a href="https://www.qsrinternational.com/nvivo-qualitative-data-analysis-software/home">NVivo</a>, a professional qualitative analysis software, and then exported in spreadsheet form for public sharing.</p> <p>Find out more about this project in our dedicated <a href="https://zenodo.org/communities/data-sharing-in-public-health-emergencies">Zenodo project community</a>.</p>
Exames de Pacientes no Diagnóstico do Covid-19
<p>O conjunto de dados usado neste estudo, foi fornecido pelo <strong><em>Hospital Adventista de Manaus</em></strong>, situado na 32<sup>a</sup> posição no Ranking Nacional dos Hospitais em Todo o Brasil para o ano de 2020, o Quadro 1 apresenta uma lista completa dos atributos do conjunto de dados.</p> <p>Quadro 1 - Lista dos atributos do dataset</p> <table> <tbody> <tr> <td><strong>Atributo</strong></td> <td><strong>Descrição</strong></td> <td><strong>Natureza</strong></td> </tr> <tr> <td>idade</td> <td>Idade do Paciente</td> <td>Numérico</td> </tr> <tr> <td>rt_pcr</td> <td>Exame Covid-19</td> <td>Categórico</td> </tr> <tr> <td>leucócitos</td> <td>Exame Laboratoriais</td> <td>Numérico</td> </tr> <tr> <td>basofilos</td> <td>Exame Laboratoriais</td> <td>Numérico</td> </tr> <tr> <td>creatinina</td> <td>Exame Laboratoriais</td> <td>Numérico</td> </tr> <tr> <td>proteina_c</td> <td>Exame Laboratoriais</td> <td>Numérico</td> </tr> <tr> <td>hemoglobina</td> <td>Exame Laboratoriais</td> <td>Numérico</td> </tr> </tbody> </table> <p>As amostras são de pacientes que foram admitidos para internação com suspeita de Convid-19 e foram submetidos a exames que seguem o protocolo estabelecido de diagnóstico adotado naquela instituição, além do exame <strong>rt_pcr <em>(covid-19)</em></strong>, também sistematicamente eram realizados outros exames que dão suporte ao diagnóstico do Covid-19.</p>
Translated Emission Pathways (TEPs): Long-Term Simulations of COVID-19 CO2 Emissions and Thermosteric Sea Level Rise Projections - Supplementary Materials
<p>Supplementary materials for Gonzalez, A. R., & Lin, T. (2022). Translated Emission Pathways (TEPs): Long-Term Simulations of COVID-19 CO<sub>2</sub> Emissions and Thermosteric Sea Level Rise Projections. <em>Earth's Future</em>. In Press.</p> <p><strong>Summary: This study introduces climate science to a broader audience by presenting an accessible research framework and environmental data related to the ongoing COVID-19 pandemic. A series of translated emission pathways (TEPs) were constructed based on the CO<sub>2</sub> emission patterns from the various phases of COVID-19 response. In addition to resembling the forcing scenarios used within climate research, a thermosteric sea level rise analysis was incorporated to further emphasize the environmental benefits that can be obtained from long-term sustainability. As a promising start for including the general public in climate change discussion, this research promotes collective environmental action that mirrors the recommendations of the scientific community.</strong></p>
Life table data for "Bounce backs amid continued losses: Life expectancy changes since COVID-19"
<p><strong>Life table data for "Bounce backs amid continued losses: Life expectancy changes since COVID-19"</strong></p> <p><em>cc-by Jonas Schöley, José Manuel Aburto, Ilya Kashnitsky, Maxi S. Kniffka, Luyin Zhang, Hannaliis Jaadla, Jennifer B. Dowd, and Ridhi Kashyap. "Bounce backs amid continued losses: Life expectancy changes since COVID-19".</em></p> <p>These are CSV files of life tables over the years 2015 through 2021 across 29 countries analyzed in the paper "Bounce backs amid continued losses: Life expectancy changes since COVID-19".</p> <p><strong>40-lifetables.csv</strong></p> <p>Life table statistics 2015 through 2021 by sex, region and quarter with uncertainty quantiles based on Poisson replication of death counts. Actual life tables and expected life tables (under the assumption of pre-COVID mortality trend continuation) are provided.</p> <p><strong>30-lt_input.csv</strong></p> <p>Life table input data.</p> <ul> <li>`id`: unique row identifier</li> <li>`region_iso`: iso3166-2 region codes</li> <li>`sex`: Male, Female, Total</li> <li>`year`: iso year</li> <li>`age_start`: start of age group</li> <li>`age_width`: width of age group, Inf for age_start 100, otherwise 1</li> <li>`nweeks_year`: number of weeks in that year, 52 or 53</li> <li>`death_total`: number of deaths by any cause</li> <li>`population_py`: person-years of exposure (adjusted for leap-weeks and missing weeks in input data on all cause deaths)</li> <li>`death_total_nweeksmiss`: number of weeks in the raw input data with at least one missing death count for this region-sex-year stratum. missings are counted when the week is implicitly missing from the input data or if any NAs are encounted in this week or if age groups are implicitly missing for this week in the input data (e.g. 40-45, 50-55)</li> <li>`death_total_minnageraw`: the minimum number of age-groups in the raw input data within this region-sex-year stratum</li> <li>`death_total_maxnageraw`: the maximum number of age-groups in the raw input data within this region-sex-year stratum</li> <li>`death_total_minopenageraw`: the minimum age at the start of the open age group in the raw input data within this region-sex-year stratum</li> <li>`death_total_maxopenageraw`: the maximum age at the start of the open age group in the raw input data within this region-sex-year stratum</li> <li>`death_total_source`: source of the all-cause death data</li> <li> <p>`death_total_prop_q1`: observed proportion of deaths in first quarter of year</p> </li> <li> <p>`death_total_prop_q2`: observed proportion of deaths in second quarter of year</p> </li> <li> <p>`death_total_prop_q3`: observed proportion of deaths in third quarter of year</p> </li> <li> <p>`death_total_prop_q4`: observed proportion of deaths in fourth quarter of year</p> </li> <li> <p>`death_expected_prop_q1`: expected proportion of deaths in first quarter of year</p> </li> <li> <p>`death_expected_prop_q2`: expected proportion of deaths in second quarter of year</p> </li> <li> <p>`death_expected_prop_q3`: expected proportion of deaths in third quarter of year</p> </li> <li> <p>`death_expected_prop_q4`: expected proportion of deaths in fourth quarter of year</p> </li> <li>`population_midyear`: midyear population (July 1st)</li> <li>`population_source`: source of the population count/exposure data</li> <li>`death_covid`: number of deaths due to covid</li> <li>`death_covid_date`: number of deaths due to covid as of <date></li> <li>`death_covid_nageraw`: the number of age groups in the covid input data</li> <li>`ex_wpp_estimate`: life expectancy estimates from the World Population prospects for a five year period, merged at the midpoint year</li> <li>`ex_hmd_estimate`: life expectancy estimates from the Human Mortality Database</li> <li>`nmx_hmd_estimate`: death rate estimates from the Human Mortality Database</li> <li>`nmx_cntfc`: Lee-Carter death rate projections based on trend in the years 2015 through 2019</li> </ul> <p><em>Deaths</em></p> <ul> <li>source: <ul> <li>STMF input data series (https://www.mortality.org/Public/STMF/Outputs/stmf.csv)</li> <li>ONS for GB-EAW pre 2020</li> <li>CDC for US pre 2020</li> </ul> </li> <li>STMF: <ul> <li>harmonized to single ages via pclm</li> <li>pclm iterates over country, sex, year, and within-year age grouping pattern and converts irregular age groupings, which may vary by country, year and week into a regular age grouping of 0:110</li> <li>smoothing parameters estimated via BIC grid search seperately for every pclm iteration</li> <li>last age group set to [110,111)</li> <li>ages 100:110+ are then summed into 100+ to be consistent with mid-year population information</li> <li>deaths in unknown weeks are considered; deaths in unknown ages are not considered</li> </ul> </li> <li>ONS: <ul> <li>data already in single ages</li> <li>ages 100:105+ are summed into 100+ to be consistent with mid-year population information</li> <li>PCLM smoothing applied to for consistency reasons</li> </ul> </li> <li>CDC: <ul> <li>The CDC data comes in single ages 0:100 for the US. For 2020 we only have the STMF data in a much coarser age grouping, i.e. (0, 1, 5, 15, 25, 35, 45, 55, 65, 75, 85+). In order to calculate life-tables in a manner consistent with 2020, we summarise the pre 2020 US death counts into the 2020 age grouping and then apply the pclm ungrouping into single year ages, mirroring the approach to the 2020 data</li> </ul> </li> </ul> <p><em>Population</em></p> <ul> <li>source: <ul> <li>for years 2000 to 2019: World Population Prospects 2019 single year-age population estimates 1950-2019</li> <li>for year 2020: World Population Prospects 2019 single year-age population projections 2020-2100</li> </ul> </li> <li>mid-year population <ul> <li>mid-year population translated into exposures: <ul> <li>if a region reports annual deaths using the Gregorian calendar definition of a year (365 or 366 days long) set exposures equal to mid year population estimates</li> <li>if a region reports annual deaths using the iso-week-year definition of a year (364 or 371 days long), and if there is a leap-week in that year, set exposures equal to 371/364\*mid_year_population to account for the longer reporting period. in years without leap-weeks set exposures equal to mid year population estimates. further multiply by fraction of observed weeks on all weeks in a year.</li> </ul> </li> </ul> </li> </ul> <p><em>COVID deaths</em></p> <ul> <li>source: COVerAGE-DB (https://osf.io/mpwjq/)</li> <li>the data base reports cumulative numbers of COVID deaths over days of a year, we extract the most up to date yearly total</li> </ul> <p><em>External life expectancy estimates</em></p> <ul> <li>source: <ul> <li>World Population Prospects (https://population.un.org/wpp/Download/Files/1_Indicators%20(Standard)/CSV_FILES/WPP2019_Life_Table_Medium.csv), estimates for the five year period 2015-2019</li> <li>Human Mortality Database (https://mortality.org/), single year and age tables</li> </ul> </li> </ul>
Data for Figures and Tables in "Bounce backs amid continued losses: Life expectancy changes since COVID-19"
<p><strong>Data for Figures and Tables in "Bounce backs amid continued losses: Life expectancy changes since COVID-19"</strong></p> <p><em>cc-by Jonas Schöley, José Manuel Aburto, Ilya Kashnitsky, Maxi S. Kniffka, Luyin Zhang, Hannaliis Jaadla, Jennifer B. Dowd, and Ridhi Kashyap. "Bounce backs amid continued losses: Life expectancy changes since COVID-19".</em></p> <p>These are CSV files of data in the figures and tables published in the paper "Bounce backs amid continued losses: Life expectancy changes since COVID-19".</p> <p><strong>50-e0diffT.csv</strong></p> <p>Figure 1: Life expectancy changes 2019/20 and 2020/21 across countries. The countries are ordered by increasing cumulative life expectancy losses since 2019. Grey dots indicate the average annual LE changes over the years 2015 through 2019.</p> <p><strong>51-arriagaT.csv</strong></p> <p>Figure 2: Age contributions to life expectancy changes since 2019 separated for 2020 and 2021. The position of the arrowhead indicates the total contribution of mortality changes in a given age group to the change in life expectancy at birth since 2019. The discontinuity in the arrow indicates those contributions separately for the years 2020 and 2021. Annual contributions can compound or reverse. The total life expectancy change from 2019 to 2021 in a given country is the sum of the arrowhead positions across age.</p> <p><strong>52-sexdiff.csv</strong></p> <p>Figure 3: Change in the female life expectancy advantage from 2019 through 2021. Blue colors indicate an increase and red colors a decrease in the female life expectancy advantage. Muted colors indicate non-significant changes.</p> <p><strong>53-e0diffcodT.csv</strong></p> <p>Figure 4: Life expectancy deficit in 2021 decomposed into contributions by age and cause of death. LE deficit is defined as observed minus expected life expectancy had pre-pandemic mortality trends continued.</p> <p><strong>55-vaxe0.csv</strong></p> <p>Figure 5: Years of life expectancy deficit during October through December 2021 contributed by ages <60 and 60+ against % of population twice vaccinated by October 1st in the respective age groups. LE deficit is defined as the counterfactual LE from a Lee-Carter mortality forecast based on death rates for the fourth quarter of the years 2015 to 2019 minus observed LE.</p> <p><strong>54-tab_arriaga.csv</strong></p> <p>Table 1: Months of life expectancy (LE) changes and deficits (labelled ES) since the start of the pandemic attributed to age-specific mortality changes (labelled AT). LE deficit is defined as observed minus expected life expectancy had pre-pandemic mortality trends continued.</p>
Pollinator-flower interactions in gardens during the COVID-19 pandemic lockdown of 2020
<p>During the main COVID-19 pandemic lockdown period of 2020 an impromptu set of pollination ecologists came together via social media and personal contacts to carry out standardised surveys of the flower visits and plants in their gardens. The surveys involved 67 rural, suburban and urban gardens, of various sizes, ranging from 61.18<sup>o</sup> North in Norway to 37.96<sup>o</sup> South in Australia and resulted in a data set of 25,174 rows long and comprising almost 47,000 visits to flowers, as well as records of plants that were not visited by pollinators. In this first publication from the project we present a brief description of the data and make it freely available for any researchers to use in the future, the only restriction being that they cite this paper in the first instance. As well as producing a data set that we hope will be widely used in the future, the project helped enormously with the health and mental wellbeing of the participants, a by-product of ecological field work that cannot be over-estimated.</p>
Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022.
<p>Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022. These data were used to estimate the time-varying reproduction number (R) in South Africa, as described in https://www.medrxiv.org/content/10.1101/2022.07.22.22277932v1.full.</p>
Dysglycemias in patients admitted to ICUs with severe acute respiratory syndrome due to COVID-19 versus other causes – A cohort study - Dataset
<p>Dataset of a cohort whose summary is described below.</p> <p>Abstract</p> <p>Importance: Dysglycemias have been associated with worse prognosis in critically ill patients with or without diabetes, but data on their association with severe COVID-19 and outcomes are lacking. Objectives: To analyze the relationship of dysglycemias with COVID-19 in hospitalized patients with severe acute respiratory syndrome (SARS) and assess the influence of dysglycemias on mortality. Design, Setting and Participants: Cohort of consecutive patients with SARS and suspected COVID-19 hospitalized in intensive care units (ICUs) across eight hospitals in Curitiba-Brazil. Main Outcomes and Measures: The primary outcome was the influence of COVID-19 on the variation of the following parameters of dysglycemia: highest glucose level at admission, mean and highest glucose levels during ICU stay, mean glucose variation, and percentage of days with hyperglycemia. The secondary outcome was the influence of COVID-19 and each of the five parameters of dysglycemia on hospital mortality within 30 days from ICU admission. Results: We compared 703 patients with COVID-19 and 138 without COVID-19 admitted to the ICUs due to SARS. Compared with patients without COVID-19, those with COVID-19 had significantly higher glucose peaks at admission (198.1mg/dL vs. 167.8mg/dL, respectively) and during ICU stay (285.9mg/Dl vs. 230.9md/dL), higher mean daily glucose values (167.9mg/dL vs. 149.8mg/dL), higher percentage of days with hyperglycemia during ICU stay (vs. 45.0 vs. 31.5), and greater mean daily glucose variations (85.3mg/dL vs. 63.5mg/dL). However, these associations were lost after adjustment for APACHE II scores, SOFA scores, CRP level, corticosteroid use and nosocomial infection. Dysglycemia and COVID-19 were each independent risk factors for mortality. Conclusions and Relevance: Patients with SARS due to COVID-19 had higher mortality and more frequent dysglycemia than patients with SARS due to other causes. This association seemed to be related to disease severity and inflammation and was independent of corticosteroid use, suggesting no specific relationship with the SARS-CoV-2 infection.</p>
Survey on the Effects of COVID-19 on the Wellbeing of Mexico City Households (ENCOVID- 19 CDMX – JULY 2020)
<p>Amid the COVID-19 outbreak, the ENCOVID-19 CDMX provides information on the well-being of Mexico City households in four main domains: labor, income, mental health, and food insecurity. It offers timely information to understand the social consequences of the pandemic and the lockdown measures. It is a cross-sectional telephone survey that, in addition to the four main domains and a set of COVID-19 related questions, includes key indicators to capture the impact of the pandemic on issues like education, social programs, and crime. This is the first dataset of the project, corresponding to July 2020, collected four months after the lockdown began in Mexico. Data collection was performed between the 8th and the 17th of July.</p>
Introducing the COVID-19 YouTube (COVYT) speech dataset featuring the same speakers with and without infection
<p>The COVYT dataset contains speech samples from individuals who self-reported their COVID-19 infection on public social media platforms (YouTube, Xiaohongshu). These videos, as well as accompanying videos of the same people prior to infection, were mined in an attempt to gather publicly-available data for COVID-19 research. This release includes the links to the original videos along with the accompanying manual segmentation and diarisation that identifies the utterances of the target individuals. We are additionally releasing features derived from the segmented utterances. Finally, the dataset includes partitioning information according to 4 different cross-validation schemes. See the arxiv pre-print for more details: https://arxiv.org/abs/2206.11045</p>
Longitudinal characterization of circulating neutrophils uncovers distinct phenotypes associated with severity in hospitalized COVID-19 patients
<p>Code and data for the manuscript "Longitudinal characterization of circulating neutrophils uncovers distinct phenotypes associated with severity in hospitalized COVID-19 patients".</p> <p>Contains all code located at <a href="https://github.com/lasalletj/COVID_Neutrophils">https://github.com/lasalletj/COVID_Neutrophils</a> as well as additional data files needed to run the code.</p> <p>Three additional publicly available data objects are required to run the code from start to finish. The first, covid.combined_final.Robj, from the Sinha et al. Nature Medicine 2022 paper (<a href="https://doi.org/10.1038/s41591-021-01576-3">https://doi.org/10.1038/s41591-021-01576-3</a>), is downloadable from the following link: <a href="https://figshare.com/ndownloader/files/31562957">https://figshare.com/ndownloader/files/31562957</a>. The other two required objects, seurat_COVID19_Neutrophils_cohort2_rhapsody_jonas_FG_2020-08-18.rds and seurat_COVID19_freshWB-PBMC_cohort2_rhapsody_jonas_FG_2020-08-18.rds, are from the Schulte-Schrepping et al. Cell 2020 paper (<a href="https://doi.org/10.1016/j.cell.2020.08.001">https://doi.org/10.1016/j.cell.2020.08.001</a>), and can be downloaded from <a href="https://beta.fastgenomics.org/datasets/detail-dataset-ee4b1a0f339140ad82f861aea35076f1#Files">https://beta.fastgenomics.org/datasets/detail-dataset-ee4b1a0f339140ad82f861aea35076f1#Files</a> and <a href="https://beta.fastgenomics.org/datasets/detail-dataset-1ad2967be372494a9fdba621610ad3f3#Files">https://beta.fastgenomics.org/datasets/detail-dataset-1ad2967be372494a9fdba621610ad3f3#Files</a>, respectively.</p> <p>Any additional information required to reanalyze the data reported in this work paper is available from the Lead Contact, Moshe Sade-Feldman (msade-feldman@mgh.harvard.edu) upon request.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.