Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,862
datasets available to search
ShareScore release 0.7.1
Dataset results
1,862 results for “Tuberculosis”
Species-specific proteotypic peptides for characterization of non-tuberculosis mycobacteria
<p>Non-tuberculous mycobacteria are opportunistic bacteria that closely resemble <i>Mycobacterium tuberculosis,</i> causing respiratory infections in humans. While genotyping through genome sequencing is found accurate in detecting mycobacterial species but struggles with distinguishing bacterial co-infection. These challenges lead to delayed therapeutic intervention, drug resistance, and disease complications. Lately, mass spectrometry-based (MALDI-TOF-MS) proteomics has routinely been used in diagnosing mycobacterial species in clinical samples. However, it suffers accurate species detection owing to extensive bacterial cultures and poor specificity in polymicrobial infections. In contrast, due to its sensitivity, LC-MS/MS based proteomics is widely employed for accurate bacterial proteome mapping. In this study, in-depth proteomics of 9 NTM species with proteome database searches of 26 datasets was used. In total 20 million peptide spectrum matches were identified aiding to ≥40% proteome coverage in 7 NTMs with highest in <i>M. abscessus</i>. Further, metaproteomic analysis and rescoring of peptides resulted in high-confidence species-specificity proteotypic peptides in <i>M. smegmatis</i> (2342), <i>M. vaccae</i> (960), <i>M. abscessus</i> (75), <i>M. avium</i> subsp. <i>paratuberculosis</i> (3) and <i>M. fortuitum</i> (1). Finally, database search results were converted to spectral library format for easier future usage in targeted proteomic workflows. This workflow in deriving species-specific peptides with high confidence can be extended in distinguishing closely related bacterial species for enhancing microbial diagnostics.</p>
Post-trial access practice in Malaria, Tuberculosis, and NTDs Clinical Trial studies in Sub-Saharan African countries, quantitative study
<p>This is the data set used <span>to evaluate post trial access plan and implementation practice on TB, Malaria and NTD clinical trial studies conducted in the sub-Saharan African countries. </span></p>
The within-host population dynamics of Mycobacterium tuberculosis vary with treatment efficacy.
<p>Data used for the publication of a paper entitled: <strong>The within-host population dynamics of <em>Mycobacterium tuberculosis</em> vary with treatment efficacy.</strong></p> <p>The data were derived from:</p> <p>1. the deep sequencing of serial sputum samples from 12 TB patients,</p> <p>2. the deep sequencing of liquid cultures derived from the expansion of individual colonies <em>in vitro</em>,</p> <p>3. <em>In </em><em>silico</em> simulations of DNA sequencing, populations and mutagenesis.</p> <p>The analytical scripts associated with the generation of the data can be found at:</p> <p>https://github.com/swisstph/TBRU_serialTB/</p> <p><strong>Paper Abstract:</strong></p> <p><strong>Background:</strong></p> <p>Combination therapy is one of the most effective tools for limiting the emergence of drug resistance. Despite the widespread adoption of combination therapy across diseases, drug resistance rates continue to rise, leading to failing treatment regimens. The mechanisms underlying treatment failure are well studied, but the processes governing successful combination therapy are poorly understood. We addressed this question by studying the population dynamics of <em>Mycobacterium tuberculosis</em> within tuberculosis patients undergoing treatment with different combinations of antibiotics.</p> <p><strong>Results:</strong></p> <p>By combining very deep whole genome sequencing (~1,000-fold genome-wide coverage) with sequential sputum sampling, we were able to detect transient genetic diversity driven by the apparently continuous turnover of minor alleles, which could serve as the source of drug-resistant bacteria. However, we report that treatment efficacy had a clear impact on the population dynamics: sufficient drug pressure bore a clear signature of purifying selection leading to apparent genetic stability. In contrast, <em>M. tuberculosis</em> populations subject to less drug pressure showed markedly different dynamics, including cases of acquisition of additional drug resistance.</p> <p><strong>Conclusions:</strong></p> <p>Our findings show that for a pathogen like <em>M. tuberculosis</em>, which is well adapted to the human host, purifying selection constrains the evolutionary trajectory to resistance in effectively treated individuals. Nonetheless, we also report a continuous turnover of minor variants, which could give rise to the emergence of drug resistance in cases of drug pressure weakening. Monitoring bacterial population dynamics could therefore provide an informative metric for assessing the efficacy of novel drug combinations.</p>
A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases
<p>The following data and scripts are part of the manuscript titled 'A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases' which is currently going through the peer-review process and has already been published as a preprint. Please read the README.txt file for information on the files uploaded.</p>
Distributable, Metabolic PET Reporting of Tuberculosis
<p>Tuberculosis remains a large global disease burden for which treatment regimens are protracted and monitoring of disease activity difficult. Existing detection methods rely almost exclusively on bacterial culture from sputum which limits sampling to organisms on the pulmonary surface. Advances in monitoring tuberculous lesions have utilized the common glucoside [18 F]FDG, yet lack specificity to the causative pathogen Mycobacterium tuberculosis (Mtb) and so do not directly correlate with pathogen viability. Here we show that a close mimic that is also positron-emitting of the non-mammalian Mtb disaccharide trehalose – 2-[ 18 F]fluoro-2-deoxytrehalose ([18 F]FDT) – is a mechanism-based reporter of Mycobacteria-selective enzyme activity in vivo. Use of [18 F]FDT in the imaging of Mtb in diverse models of disease, including non-human primates, successfully co-opts Mtb-specific processing of trehalose to allow the specific imaging of TB-associated lesions and to monitor the effects of treatment. A pyrogen-free, direct enzyme-catalyzed process for its radiochemical synthesis allows the ready production of [18 F]FDT from the most globally-abundant organic 18F-containing molecule, [18 F]FDG. The full, pre-clinical validation of both production method and [18 F]FDT now creates a new, bacterium selective, clinical diagnostic candidate for clinical evaluation. We anticipate that this distributable technology to generate clinical-grade [18 F]FDT directly from the widelyavailable clinical reagent [18 F]FDG, without need for either custom-made radioisotope generation or specialist chemical methods and/or facilities, could now usher in global, democratized access to a TB-specific PET tracer.</p>
The Role of the World Bank in Financing Tuberculosis Control- Dataset
<p>The completed dataset for the World Bank's financing of TB control initiatives from 1986-2017 (coded_questionnaire.csv). This is accompanied by the variable description guide (Glossary and Variable Description Guide.pdf). </p> <p>Co-financiers_database.csv is the data used in the analysis of co-financing of World Bank tuberculosis projects. </p> <p>IHME_DAH_DATABASE 24-10-17.csv is the data used in the analysis of Development Assitance for Health and Tuberculosis.</p> <p>World Bank and TB.R is the source code used to encode and analyse all of the above data. Data analysis on R version 3.4.3.</p> <p>Extended Data.pdf is the supplementary data for publication. </p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
A modified decontamination and storage method for sputum from patients with tuberculosis
<p>Sputum sampling is a cheap and non-invasive method for diagnosis Tuberculosis. A modified method is devised to enhance handling capacity and minimize risk of contamination in culture.</p> <p>"22_samples_TTP_GU_method_comparison_dataset.csv" is a dataset for comparing standard method and modified method of sputum handling and storing procedures before being cultured in MGIT. MGIT is used in BD BACTEC 960 MGIT system which generates "Time to positive" hours for a culture to growth and "Growth Unit" for estimating the amount of growth.</p> <p>"348_samples_TTP_GU_modified method_dataset.csv " is a dataset for applying modified method on selected 348 sputum samples for culture in MGIT. The dataset contains "Time to positive", "Growth Unit", "ZN smear grade", and "Duration of frozen:.</p>
Multiple Sequence Alignment of a diverse dataset with 1788 Mycobacterium tuberculosis isolates
<p><strong>Multiple Sequence Alignment of a diverse dataset with 1788 <em>Mycobacterium tuberculosis</em> isolates used for <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> benchmarking</strong></p> <p>The dataset comprises whole-genome sequence data published by <a href="https://doi.org/10.1016/S1473-3099(15)00062-6">Walker et al. 2015</a>. For the multiple sequence analysis, we proceeded as follows:</p> <ol> <li>Reads were downloaded from ENA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA282721">PRJNA282721</a> (accessed on March 16<sup>th</sup>, 2023) and trimmed using Trimmomatic (<a href="https://pubmed.ncbi.nlm.nih.gov/24695404/">Bolger et al., 2014</a>) with <a href="https://github.com/B-UMMI/INNUca">INNUca</a> default settings;</li> <li>Quality-processed reads were individually mapped against the H37Rv reference genome (Genbank accession: <a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_000962.3/">NC_000962.3</a>) using <a href="https://github.com/tseemann/snippy">Snippy</a> v4.5.1 and SNP-calling was performed on variant sites with the following criteria: a minimum proportion of reads differing from the reference of 70%, a minimum mapping quality of 30 and a minimum coverage for SNP calling of 10;</li> <li>A full alignment was extracted using Snippy’s core module (snippy-core), with masking of SNPs falling within known <em>M. tuberculosis</em> genomic regions with high GC content, repetitive elements and resistance-associated positions (corresponding to ~8% of the genome), as previously described for surveillance purposes (<a href="https://pubmed.ncbi.nlm.nih.gov/30948181/">Macedo et al., 2019</a>);</li> <li><em>M. tuberculosis </em>lineages were determined using tb-profiler v4.4.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31234910/">Phelan et al., 2019</a>), with samples from the <em>M. tuberculosis</em> complex other than <em>M. tuberculosis</em>, representing a mix of multiple lineages, or with less than 95% of mapped positions in the reference, being excluded;</li> <li>A filtered alignment comprising the maximum number of informative sites (88,562 nucleotide sites with at least one mutation in a given sequence) was extracted from the full alignment using the alignment_processing.py v1.1.0 (default settings) of <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a>, and then used as input for the benchmarking.</li> </ol> <p>In this repository, we provide two alignment files:</p> <ul> <li>Core_MTB_1787_strs.full.aln: this corresponds to the full multiple sequence alignment comprising 1787 samples and the reference (corresponding to the point 4 of the methodology).</li> <li>MTb_original_align_profile.fasta: this corresponds to the multiple sequence alignment comprising 1787 samples and the reference and only presenting the alignment informative sites (corresponding to the point 5 of the methodology)</li> </ul>
TBValid collection: Pulmonary tuberculosis validation collection
<p>The TBValid dataset comprises 870 digital patients with different profiles, each with fixed Age, BMI and MtbSputum. Each is identified by a vector of features involving biological and pathophysiological parameters to roughly represent different profiles in the population and initial bacterial load. Individual patient data collected during a clinical trial have been transformed into aggregated data, which are already irreversibly anonymised. Subsequently, these aggregated data have been sampled through the procedure described in “Generation of digital patients for the simulation of tuberculosis with UISS-TB”, doi: 10.1186/s12859-020-03776-z. The obtained derivative dataset, owned by its creators, does not constitute sensitive data according to European laws, and it is impossible with this dataset to re-establish the identity of the patients enrolled in the original clinical trial.</p>
Dataset - A lung-on-chip model reveals an essential role for alveolar epithelial cells in controlling bacterial growth during early M. tuberculosis infection
<p>Description of the sub-folders<br> Name, type of data, corresponding Figure in the manuscript<br> 3D view of the LoC model - .tiff image stack, Figure 1.</p> <p>Bacterial Growth Rate Data - .tiff image stacks, .csv files and MATLAB code to extract the fluorescence intensity over time, Figure 2, Figure 2 - figure supplement 2, Figure 2 - figure supplement 4, Figure 3, Figure 3 - figure supplement 2, Figure 4.</p> <p>AT Characterization - .tiff image stacks and MATLAB code to extract the number and volume of lamellar bodies from the stack of confocal images, Figure 1, Figure 1 - figure supplement 1, Figue 1 - figure supplement 2.</p> <p>AT Infection in LoC model - .tiff image stacks, Figure 2 - figure supplement 1.</p> <p>AT Infection in vivo - .tiff image stacks, Figure 1 - figure supplement 3.</p> <p>Simulations of in vivo infections - .dat files of growth rates in macrophages for the WT and ESX-1 deficient populations and MATLAB code to simulate an infection from this data, Figure 4.</p> <p> </p>
Association between meteorological factors and the number of tuberculosis notifications: a time-series study in Hong Kong
<p> Using a 22-year consecutive surveillance data in Hong Kong, including monthly averages of meteorological factors, air pollution concentrations , total number of TB cases notified, to analyze the association of monthly average temperature and relative humidity with temporal dynamics of monthly total number of TB cases notified. </p>
Data from: Effect of culling on individual badger (Meles meles) behaviour: potential implications for bovine tuberculosis transmission
1. Culling wildlife as a form of disease management can have unexpected and sometimes counterproductive outcomes. In the UK, badgers (Meles meles) are culled in efforts to reduce badger-to-cattle transmission of Mycobacterium bovis, the causative agent of bovine tuberculosis (TB). However, culling has previously been associated with both increased and decreased incidence of M. bovis infection in cattle. 2. The adverse effects of culling have been linked to cull-induced changes in badger ranging, but such changes are not well documented at the individual level. Using GPS-collars, we characterised individual badger behaviour within an area subjected to widespread industry-led culling, comparing it with the same area before culling and with three unculled areas. 3. Culling was associated with a 61% increase (95% CI 27-103%) in monthly home range size, a 39% increase (95% CI 28-51%) in nightly maximum distance from the sett, and a 17% increase (95% CI 11-24%) in displacement between successive GPS-collar locations recorded at 20-minute intervals. Despite travelling further, we found a 91.2 minute (95% CI 67.1-115.3 minute) reduction in the nightly activity time of individual badgers associated with culling. These changes became apparent while culls were ongoing and persisted after culling ended. 4. Expanded ranging in culled areas was associated with individual badgers visiting 45% (95% CI 15-80%) more fields each month, suggesting that surviving individuals had the opportunity to contact more cattle. Moreover, surviving badgers showed a 19.9-fold increase (95% CI 10.8-36.4 increase) in the odds of trespassing into neighbouring group territories, increasing opportunities for intergroup contact. 5. Synthesis and Applications: Badger culling was associated with behavioural changes among surviving badgers which potentially increased opportunities for both badger-to-badger and badger-to-cattle transmission of M. bovis. Furthermore, by reducing the time badgers spent active, culling may have reduced badgers' accessibility to shooters, potentially undermining subsequent population control efforts. Our results specifically illustrate the challenges posed by badger behaviour to cull-based TB control strategies and furthermore, they highlight the negative impacts culling can have on integrated disease control strategies.
Migration and tuberculosis in Berlin (dataset)
<p>This is the supporting data file for the manuscript entitled "Higher rate of tuberculosis in second generation migrants compared to native residents in a metropolitan setting in Western Europe" (Marx et al., PLoS ONE). The dataset includes anonymized, routinely collected notification data (variables labeled as "nd") for 314 individuals and anonymized survey data (i.e. data obtained through interviews; variables labeled as "sd") for a subset of 154 individuals. The data are published open-access, in accordance with the PLoS ONE data policy (2014).</p>
Supporting dataset for manuscript: "Higher rate of tuberculosis in second generation migrants compared to native residents in a metropolitan setting in Western Europe" (PLoS ONE)
<p>This is the supporting datafile for the manuscript entitled "Higher rate of tuberculosis in second generation migrants compared to native residents in a metropolitan setting in Western Europe" (Marx et al., PLoS ONE). The dataset includes anonymized, routinely collected notification data (variables labeled as "nd") for 314 individuals and anonymized survey data (i.e. data obtained through interviews; variables labeled as "sd") for a subset of 154 individuals. The data are published open-access, in accordance with the PLoS ONE data policy (2014).</p>
COMBAT TB Tuberculosis genome annotation database
<p>A Neo4j (version 2.3.3) format graph database containing annotation related to the M. tuberculosis H37Rv genome, created as part of the COMBAT TB project at the South African National Bioinformatics Institute.</p>
X-ray diffraction images for the H145E mutant of the iron-dependent superoxide dismutase from Mycobacterium tuberculosis.
<p>X-ray diffraction images of the H145E mutant (prefixed h145e) which were collected in October 1995 using a graphite-monochromated copper K-alpha rotating anode source (wavelength 1.5418 Å) with a Marresearch 90 cm image plate detector at a distance of 120 mm from the crystal. The data were collected at room temperature in two passes, each consisting of 100 one degree rotations of the crystal. Each image had an exposure time of 20 minutes. The crystal was rotated in the capillary tube prior to collection of the second pass in order to record the 'blind' region of the diffraction pattern and this set of images is prefixed h145eb. </p>
X-ray diffraction images for the H145Q mutant of the iron-dependent superoxide dismutase from Mycobacterium tuberculosis.
<p>X-ray diffraction images collected from one crystal at room temperature using a rotating anode copper source (wavelength 1.5418 Å) and a 30 cm Marresearch image plate detector. The crystal-to-detector distance was 150 mm and a 90 mm image plate scan radius was used. Each of the 60 images had an exposure time of 20 minutes and corresponds to a 3 degree phi-rotation of the crystal. Diffraction extends to about 3.3 Å resolution. </p>
User Guide – Dashboard on Zoonotic tuberculosis: Mycobacterium
<p>User Guide – Dashboard on Zoonotic tuberculosis focusing on Mycobacterium bovis and M. caprae</p>
Dataset from Remote analysis of Sputum Smears for Mycobacterium Tuberculosis Quantification using Digital Crowdsourcing
<p>Worldwide, TB is one of the top 10 causes of death and the leading cause from a single infectious agent. Although the development and roll out of Xpert MTB/RIF has recently become a major breakthrough in the field of TB diagnosis, smear microscopy remains the most widely used method for TB diagnosis, especially in low- and middle-income countries.</p> <p>This is a minimal dataset to reproduce our research that tests the feasibility of a crowdsourced approach to tuberculosis image analysis. In particular, we investigated whether anonymous volunteers with no prior experience would be able to count acid-fast bacilli in digitized images of sputum smears by playing an online game. Following this approach 1790 people identified the acid-fast bacilli present in 60 digitized images, the best overall performance was obtained with a specific number of combined analysis from different players and the performance was evaluated with the F1 score, sensitivity and positive predictive value, reaching values of 0.933, 0.968 and 0.91, respectively.</p> <p>The dataset includes 24 digitized images of sputum smears and the corresponding gameplays clicks. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.