Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,674
datasets available to search
ShareScore release 0.7.1
Dataset results
9,674 results for “COVID-19”
Monitoring feedback to authors on the quality of trials evaluating interventions aimed at preventing and treating COVID-19
<p>We aimed to assess transparency of reporting and risk of bias of randomized trials evaluating interventions aimed at preventing and treating COVID-19.</p> <p>This review is part of a larger project: the COVID-NMA project (Boutron 2020a). The COVID-NMA project aims to provide decision-makers with a complete, high-quality and up-to-date synthesis of evidence on interventions for the prevention and treatment of COVID 19. For this purpose, we perform a living mapping of all registered randomized controlled trials and a living evidence synthesis of data from RCTs. We developed a master protocol on the effect of all interventions for the prevention and treatment of COVID-19 (first published on April 8, 2020; an update on May 11, 2020, June 17, 2020, and September 8, 2020) (Boutron 2020b). We set-up a platform (<a href="https://covid-nma.com/">https://covid-nma.com</a>) where all our results are made available and updated weekly.</p>
Diagnostic accuracy of a set of clinical and radiological criteria for screening of COVID-19 using RT-PCR as the reference standard - Dataset
<p>Dataset of a cohort whose summary is described below.</p> <p>Abstract</p> <p><strong>Objective:</strong> To evaluate the accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of a set of clinical-radiological criteria for COVID-19 screening in patients with severe acute respiratory failure (SARF) admitted to intensive care units (ICUs), using reverse-transcriptase polymerase chain reaction (RT-PCR) as the reference standard. <strong>Method: </strong>Diagnostic accuracy study including a historical cohort of 1009 patients consecutively admitted to ICUs across six hospitals in Curitiba (Brazil) from March to September, 2020. The sample was stratified into groups by the strength of suspicion for COVID-19 (strong <em>versus</em> weak) using parameters based on three clinical and radiological (chest computed tomography) criteria. The diagnosis of COVID-19 was confirmed by RT-PCR (referent). <strong>Results:</strong> With respect to RT-PCR, the proposed criteria had 98.5% (95% confidence interval [95% CI] 97.5–99.5%) sensitivity, 70% (95% CI 65.8–74.2%) specificity, 85.5% (95% CI 83.4–87.7%) accuracy, PPV of 79.7% (95% CI 76.6–82.7%) and NPV of 97.6% (95% CI 95.9–99.2%). <strong>Conclusion: </strong>The proposed set of clinical-radiological criteria were accurate in identifying patients with strong <em>versus</em> weak suspicion for COVID-19 and had high sensitivity and considerable specificity with respect to RT-PCR. These criteria may be useful for screening COVID-19 in patients presenting with SARF.</p>
Risk and symptoms of COVID-19 in health professionals according to baseline immune status and booster vaccination during the Delta and Omicron waves in Switzerland – a multicentre cohort study
<p>For details, see publication</p>
Altered infective proficiency of the gut microbiome following COVID-19
<p><strong>The effects of SARS-CoV-2 infections comprise of many heterogeneous symptoms including several involving the human gastrointestinal tract. We assess the effects of COVID-19 on the host microbiome</strong></p>
Knowledge of Social Networks for Health is Associated with COVID-19 Health Protective Behaviors
<p>This is the dataset and stata code for the paper "Knowledge of Social Networks for Health is Associated with COVID-19 Health Protective Behaviors” submitted to Plos One May 1st, 2024.</p>
The UK COVID-19 Vocal Audio Dataset
<p>The UK COVID-19 Vocal Audio Dataset is designed for the training and evaluation of machine learning models that classify SARS-CoV-2 infection status or associated respiratory symptoms using vocal audio. The UK Health Security Agency recruited voluntary participants through the national Test and Trace programme and the REACT-1 survey in England from March 2021 to March 2022, during dominant transmission of the Alpha and Delta SARS-CoV-2 variants and some Omicron variant sublineages. Audio recordings of volitional coughs, exhalations, and speech (speech not available in open access version) were collected in the 'Speak up to help beat coronavirus' digital survey alongside demographic, self-reported symptom and respiratory condition data, and linked to SARS-CoV-2 test results. The UK COVID-19 Vocal Audio Dataset represents the largest collection of SARS-CoV-2 PCR-referenced audio recordings to date. PCR results were linked to 70,794 of 72,999 participants and 24,155 of 25,776 positive cases. Respiratory symptoms were reported by 45.62% of participants. This dataset has additional potential uses for bioacoustics research, with 11.30% participants reporting asthma, and 27.20% with linked influenza PCR test results.</p> <h3>Contents</h3> <ul> <li><strong>participant_metadata.csv</strong> row-wise, participant identifier indexed information on participant demographics and health status. Please see <a href="https://arxiv.org/pdf/2212.07738.pdf">A large-scale and PCR-referenced vocal audio dataset for COVID-19</a> for a full description of the dataset.</li> <li><strong>audio_metadata.csv</strong> row-wise, participant identifier indexed information on three recorded audio modalities, including audio filepaths. Please see <a href="https://arxiv.org/pdf/2212.07738.pdf">A large-scale and PCR-referenced vocal audio dataset for COVID-19</a> for a full description of the dataset.</li> <li><strong>train_test_splits.csv</strong> row-wise, participant identifier indexed information on train test splits for the following sets: 'Randomised' train and test set, Standard' train and test set, Matched' train and test sets, 'Longitudinal' test set and 'Matched Longitudinal' test set. Please see <a href="https://arxiv.org/abs/2212.08570">Audio-based AI classifiers show no evidence of improved COVID-19 screening over simple symptoms checkers</a> for a full description of the train test splits.</li> <li><strong>audio/ </strong>directory containing all the recordings in .wav format <ul> <li>Due to the large size of the dataset, to assist with ease of download, the audio files have been zipped into <strong>covid_data.z{ip, 01-24}.</strong> This enables the dataset to be downloaded in short periods, reducing the chances of a dropped internet connection scuppering progress. To unzip, first, ensure that all zip files are in the same directory. Then run the command 'unzip covid_data.zip' or right-click on 'covid_data.zip' and use a programme such as 'The Unarchiver' to open the file.</li> <li>Once extracted, to check the validity of the download, please run the 'python Turing-RSS-Health-Data-Lab-Biomedical-Acoustic-Markers/data-paper/unit-tests.py. All tests should pass with no exceptions. Please clone the GitHub repo detailed below.</li> </ul> </li> <li><strong>README.md</strong> full dataset descriptor.</li> <li><strong>DataDictionary_UKCOVID19VocalAudioDataset_OpenAccess.xlsx </strong>descriptor of each dataset attribute with the percentage coverage.</li> </ul> <h3>Code Base</h3> <p>The accompanying code can be found here: https://github.com/alan-turing-institute/Turing-RSS-Health-Data-Lab-Biomedical-Acoustic-Markers</p> <h3>Citations:</h3> <p>Please cite.</p> <p>@article{coppock2024audio,</p> <p> author = {Coppock, Harry and Nicholson, George and Kiskin, Ivan and Koutra, Vasiliki and Baker, Kieran and Budd, Jobie and Payne, Richard and Karoune, Emma and Hurley, David and Titcomb, Alexander and Egglestone, Sabrina and Cañadas, Ana Tendero and Butler, Lorraine and Jersakova, Radka and Mellor, Jonathon and Patel, Selina and Thornley, Tracey and Diggle, Peter and Richardson, Sylvia and Packham, Josef and Schuller, Björn W. and Pigoli, Davide and Gilmour, Steven and Roberts, Stephen and Holmes, Chris},</p> <p> title = {Audio-based AI classifiers show no evidence of improved COVID-19 screening over simple symptoms checkers},</p> <p> journal = {Nature Machine Intelligence},</p> <p> year = {2024},</p> <p> doi = {https://doi.org/10.1038/s42256-023-00773-8}</p> <p>}</p> <p>@article{budd2024,</p> <p> author={Jobie Budd and Kieran Baker and Emma Karoune and Harry Coppock and Selina Patel and Ana Tendero Cañadas and Alexander Titcomb and Richard Payne and David Hurley and Sabrina Egglestone and Lorraine Butler and George Nicholson and Ivan Kiskin and Vasiliki Koutra and Radka Jersakova and Peter Diggle and Sylvia Richardson and Bjoern Schuller and Steven Gilmour and Davide Pigoli and Stephen Roberts and Josef Packham Tracey Thornley Chris Holmes},</p> <p> title={A large-scale and PCR-referenced vocal audio dataset for COVID-19},</p> <p> journal={Scientific Data},</p> <p> year={2024},</p> <p> doi = {https://doi.org/10.1038/s41597-024-03492-w}</p> <p>}</p> <p>@article{Pigoli2022,</p> <p> author={Davide Pigoli and Kieran Baker and Jobie Budd and Lorraine Butler and Harry Coppock and Sabrina Egglestone and Steven G.\ Gilmour and Chris Holmes and David Hurley and Radka Jersakova and Ivan Kiskin and Vasiliki Koutra and George Nicholson and Joe Packham and Selina Patel and Richard Payne and Stephen J.\ Roberts and Bj\"{o}rn W.\ Schuller and Ana Tendero-Ca$\tilde{n}$adas and Tracey Thornley and Alexander Titcomb},</p> <p>title={Statistical Design and Analysis for Robust Machine Learning: A Case Study from Covid-19},</p> <p> year={2022},</p> <p> journal={arXiv},</p> <p> doi = {10.48550/ARXIV.2212.08571}</p> <p>}</p> <p> </p> <h3>The Dublin Core™ Metadata Initiative</h3> <p> </p> <p>- Title: The UK COVID-19 Vocal Audio Dataset, Open Access Edition.</p> <p>- Creator: The UK Health Security Agency (UKHSA) in collaboration with The Turing-RSS Health Data Lab.</p> <p>- Subject: COVID-19, Respiratory symptom, Other audio, Cough, Asthma, Influenza.</p> <p>- Description: The UK COVID-19 Vocal Audio Dataset Open Access Edition is designed for the training and evaluation of machine learning models that classify SARS-CoV-2 infection status or associated respiratory symptoms using vocal audio. The UK Health Security Agency recruited voluntary participants through the national Test and Trace programme and the REACT-1 survey in England from March 2021 to March 2022, during dominant transmission of the Alpha and Delta SARS-CoV-2 variants and some Omicron variant sublineages. Audio recordings of volitional coughs and exhalations were collected in the 'Speak up to help beat coronavirus' digital survey alongside demographic, self-reported symptom and respiratory condition data, and linked to SARS-CoV-2 test results. The UK COVID-19 Vocal Audio Dataset Open Access Edition represents the largest collection of SARS-CoV-2 PCR-referenced audio recordings to date. PCR results were linked to 70,794 of 72,999 participants and 24,155 of 25,776 positive cases. Respiratory symptoms were reported by 45.62% of participants. This dataset has additional potential uses for bioacoustics research, with 11.30% participants reporting asthma, and 27.20% with linked influenza PCR test results.</p> <p>- Publisher: The UK Health Security Agency (UKHSA).</p> <p>- Contributor: The UK Health Security Agency (UKHSA) and The Alan Turing Institute.</p> <p>- Date: 2021-03/2022-03</p> <p>- Type: Dataset</p> <p>- Format: Waveform Audio File Format audio/wave, Comma-separated values text/csv</p> <p>- Identifier: <strong>10.5281/zenodo.10043977</strong></p> <p>- Source: The UK COVID-19 Vocal Audio Dataset Protected Edition, accessed via application to <a href="https://www.gov.uk/government/publications/accessing-ukhsa-protected-data/accessing-ukhsa-protected-data">Accessing UKHSA protected data</a>.</p> <p>- Language: eng</p> <p>- Relation: The UK COVID-19 Vocal Audio Dataset Protected Edition, accessed via application to <a href="https://www.gov.uk/government/publications/accessing-ukhsa-protected-data/accessing-ukhsa-protected-data">Accessing UKHSA protected data</a>.</p> <p>- Coverage: United Kingdom, 2021-03/2022-03.</p> <p>- Rights: Open Government Licence version 3 (OGL v.3), © Crown Copyright UKHSA 2023.</p> <p>- accessRights: When you use this information under the Open Government Licence, you should include the following attribution: The UK COVID-19 Vocal Audio Dataset Open Access Edition, UK Health Security Agency, 2023, licensed under the <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/">Open Government Licence v3.0</a> and cite the papers detailed above.</p> <p> </p>
Datos diarios sobre la COVID-19 en Andalucía (España) a escala de provincia y municipio
<p>Data on daily COVID-19 confirmed cases, hospitalised people (both in conventional and intensive care units), and deaths, for Andalucía region in southern Spain, at the scale of both provinces and municipalities. Data cover the period from the very start of the pandemic (26 February 2020) up to 4 April 2022. The data were captured daily by <a href="https://frodriguezsanchez.net/" target="_blank" rel="noopener">Francisco Rodríguez-Sánchez</a> from the <a href="https://www.juntadeandalucia.es/institutodeestadisticaycartografia/salud/COVID19.html">official website </a>of Junta de Andalucía (wayback machine capture from 1 February 2022 <a href="https://web.archive.org/web/20220119092327/https://www.juntadeandalucia.es/institutodeestadisticaycartografia/badea/informe/anual?CodOper=b3_2314&idNode=42348" target="_blank" rel="noopener">here</a>). Note the official data were often changed retrospectively by the government as numbers were revised continuously. The full history of changes of both datasets can be checked at this GitHub repository: <a href="https://github.com/Pakillo/COVID19-Andalucia" target="_blank" rel="noopener">https://github.com/Pakillo/COVID19-Andalucia</a>. There you can also get the R code used to analyse those data, which were visualised in an online daily report here: <a href="https://pakillo.github.io/COVID19-Andalucia/evolucion-coronavirus-andalucia.html" target="_blank" rel="noopener">https://pakillo.github.io/COVID19-Andalucia/evolucion-coronavirus-andalucia.html</a>.</p> <p> </p>
COVID-19 Vaccines Database: 2020-2022
<p>The attached databases were generated and used in for the analysis of the research article <em>'Which roads lead to access? A global landscape of six COVID-19 vaccine innovation models'. </em>They contain data related to COVID-19 vaccines' registration status, prices, production, purchases, deliveries, and investments between 2020 and 2022. <em><br></em></p> <p>These databases were compiled by collecting and revising data from two sources: UNICEF's COVID Market Dasboard, and the COVID-19 vaccine R&D investments tracker from the Geneva Graduate Institute's Global Health Centre: https://www.knowledgeportalia.org/covid-19-vaccine-r-d-funding.<em><br></em></p>
Effect of the Increased Nursing Attrition Rate on Nursing Administration Process during the Covid-19 Pandemic in a Selected Tertiary Care Hospital
<p><span>During<span> </span>the<span> </span>COVID-19<span> </span>outbreak,<span> </span>healthcare<span> </span>professionals,<span> </span>particularly<span> </span>nurses,<span> </span>were<span> </span>more<span> </span>prone to<span> </span>diseases.<span> </span>Globally<span> </span>attrition<span> </span>rate<span> </span>was<span> </span>high<span> </span>among<span> </span>nurses<span> </span>and<span> </span>during<span> </span>the<span> </span>pandemic,<span> </span>it<span> </span>increased because<span> </span>of<span> </span>various<span> </span>reasons<span> </span>such<span> </span>as<span> </span>the<span> </span>risk<span> </span>of<span> </span>infection,<span> </span>occupational<span> </span>and<span> </span>psychological<span> </span>stress, causing risk to their loved ones. This led to a chaotic situation where nurse managers were forced to implement specific strategic plans to deal with increased nurse attrition. This study aims<span> </span>to<span> </span>describe<span> </span>the<span> </span>impact<span> </span>of<span> </span>nurse<span> </span>attrition<span> </span>rate<span> </span>on<span> </span>nursing<span> </span>administration<span> </span>during<span> </span>COVID-19 at a selected tertiary care hospital. The research approach adopted in this study is descriptive cross-sectional. A total sample of 66 nurses involved in nursing administration. The data is collected through a structured questionnaire and the nurse attrition data during the COVID-19 pandemic period was collected from the interview method during the survey. Statistical tests used were frequency, percentage, mean, Standard Deviation (S.D). The study showed that there is a moderate impact of increased nurse attrition on nursing administration during the COVID-19 pandemic. The study led to the identification of gaps that need to be addressed in a similar crisis.</span></p>
Short-lived air pollutants and climate forcers through the lens of the COVID-19 pandemic
<p>The data in this repository is part of the paper titled "Short-lived air pollutants and climate forcers through the lens of the COVID-19 pandemic". The data is required to obtain a detrended lockdown effects on air quality. The raw data was downloaded from the European Centre for Medium-Range Weather Forecasts Atmospheric Composition Reanalysis 4 (EAC4) product portal. More details of the data are listed below:</p> <p>"ozone_data.nc": Global mixing ratio of ozone (monthly)</p> <p>"pm_data.nc": Global mass concentration of fine particulate matters, and aerosol optical depth (AOD) at 550 nm (monthly)</p> <p>"BAU_clean_latest.csv": The pollution level under a business-as-usual (BAU) scenario, inferred from the historical pollution data by Theil-Sen linear regression (monthly)</p>
The complete corpus of #COVID-19 Twitter dataset
<p><br> COVID-19 pandemic initiated over a year ago continues to spread around the globe and the ongoing research regarding COVID-19 is on a continues growth as well. The online discourse on social media regarding COVID-19 has been growing along with the timeline of the pandemic.</p> <p>Open data on Twitter have been released and offer the research community the opportunity for new findings and resolving this new threat. In this dataset, we open a corpus of Twitter's data from March 2020 till today, that is being updated every day based on the two most important hashtags regarding COVID-19. This dataset will offer the research community the opportunity to explore the social extensions of this pandemic including topic analysis, hate speech sentiment analysis, regarding either the opinion of the users on the pandemic, the comments on the public discourse, or the vaccination releases. The dataset has been collected by retrieving all the tweets that contain the hashtags: #coronavirus and #COVID19 including approximately 208M tweets for hashtags #coronavirus and 392M tweets for hashtag #COVID-19, resulting in a total of 600M tweets. </p>
Coffee Consumption per Capita and Covid-19 Mortality Rate
<p>There is a correlation between average of "Coffee Consumption per capita" and average of "Covid-19 Mortality Rate" for countries with high coffee consumption per capita (the countries that has more than 5.4 kg per capita per year consumption).</p> <p>The Details of computations and data are provided in an attached supplementary file (Excel File Format).</p> <p>Data gathered on 10 Aug 2021</p> <p> </p>
Datasets Cured and Enriched with Provenance from the National Vaccination Campaign Against COVID-19
<p>The COVID-19 pandemic is a global threat. If, on the one hand, weaccount for many losses, on the other hand, the generation of datasets and ur-gent analytical demands has accelerated. Among the combat strategies, vacci-nation and data-centered epidemiological investigations stand out. This datasetpaper presents the process of building cured and annotated datasets with prove-nance metadata. The main dataset is based on the registration data of the Vacci-nation Campaign against COVID-19 in Brazil. The dataset contains thousandsof records processed up to March 2021. The data were analyzed, investigated,treated and cross-checked with other sources, in order to correct and comple-ment them, resulting in cured datasets and aligned to the FAIR principles.</p>
Proactive COVID-19 testing in a partially vaccinated population.
<p>Complete simulation-generated datasets analyzed in McGee et al. (2021) Proactive COVID-19 testing in a partially vaccinated population. medRxiv 2021.08.15.21262095.</p> <p>Data is uploaded in comma-separated .csv files which have been compressed using gzip. Descriptions of data columns can be found in the column_descriptions.csv file.</p>
Model-driven mitigation measures for reopening schools during the COVID-19 pandemic.
<p>Complete simulation-generated datasets analyzed in McGee et al. (2021) Model-driven mitigation measures for reopening schools during the COVID-19 pandemic. PNAS. In press at time of upload. (medRxiv 2021.01.22.21250282).</p> <p>Data is uploaded in tab-separated .csv files which have been compressed using gzip. Descriptions of data columns can be found in the column_descriptions.csv file.</p>
Fighting COVID-19 with computational tools: an AI guided review of 17,000 studies - The CSCoV database.
<p>CSCoV (Computational Studies about COVID-19) is a dataset containing COVID-19 related studies extracted from PubMed, bioRxiv, medRxiv, and arXiv, together with article and author related metrics obtained from Semantic Scholar (plus page views from bioRxiv and medRxiv). Using machine learning, the articles are categorized in six topics (Pharmacology, Genomics, Epidemiology, Healthcare, Clinical Medicine, Clinical Imaging) and prioritized. The database is periodically updated.</p> <ul> <li>Publication: TBA</li> <li>Files included in this release: <ul> <li>cscov_09_2021.png: dataset statistics for the current CSCoV release.</li> <li>cscov_09_2021.tsv: CSCoV database.</li> <li>schema.json: metadata.</li> <li>cscov_09_2021.tar.gz: Doc2Vec and DeepWalk features used for the DL model</li> </ul> </li> <li> <p>Source code: <a href="https://github.com/SFB-KAUST/covid-review">https://github.com/SFB-KAUST/covid-review</a></p> </li> </ul>
Protein tagged from Covid-19 trial records in ClinicalTrials.gov
<p>Top 200 proteins appeared in the Covid-19 trial records in https://clinicaltrials.gov/, grouped by CATH classification, with URL linking back to the trial record at https://clinicaltrials.gov/.</p>
Extended data for "TeenCovidLife: A resource to understand the impact of the Covid-19 pandemic on adolescents in Scotland"
<p>Extended data for "TeenCovidLife: A resource to understand the impact of the Covid-19 pandemic on adolescents in Scotland" Wellcome Open Research submission</p>
Predicting COVID-19 Incidence Through Spatiotemporal Human Interactions
<p>This repository contains data (features) necessary to run STXGB model. STXGB is a spatiotemporal autoregressive model that predicts county-level new cases of COVID-19 in the coterminous US in 1- to 4-week prediction horizons using spatiotemporal lags of infection rates, human interactions, human mobility, and socioeconomic composition of counties as predictive features.</p>
COVRIN D0.3.1: Database of COVID-19 research activities
<p>OHEJP project: COVRIN "One Health research integration on SARS-CoV-2 emergence, risk assessment and preparedness".</p> <p>Since the start of the pandemic in early 2020, a huge number of research projects have been initiated on SARS-CoV-2/COVID-19; additionally, many pre-existing networks and infrastructures have turned their attention to the virus, setting up SARS-CoV-2-specific services. To avoid overlaps and ensure optimal use of resources, a scoping review was performed of European Union-supported SARS-CoV-2 research activities that overlap with COVRIN in terms of focus.</p> <p>This database is associated with OHEJP Deliverable report "D0.3.1: Scoping review of European Union-supported COVID-19 research activities" available at https://doi.org/10.5281/zenodo.5537781</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.