Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,108

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,108 results for “pathogen”

Learn how ShareScore rates datasets ↗
zenodo44/100

Simulated NGS read datasets for prediction of novel fungal pathogens and multiple pathogen classes

<p>This repository contains simulated Illumina read datasets for novel fungal pathogen prediction and real-time detection of multiple pathogen classes. They were used to train the models hosted at <a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>.<br> The reads were simulated with Mason (<a href="https://www.seqan.de/apps/mason/">https://www.seqan.de/apps/mason/</a>) from genomes downloaded from NCBI, based on metadata stored in a manually curated database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>).</p> <p>We provide the following:</p> <p>1) An rds file describing assignment of fungal species from the database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>) to training, validation and test sets (TrainValTest_fungi.rds). A second rds file (TrainValTest_temporal.rds) includes species added within 12 weeks after the original datasets were compiled. Those species were used for a temporal benchmark.</p> <p>2) Fungal validation and test sets. Each contains 1.25 million, 250bp-long reads simulated from non-overlapping sets of human (&quot;pathogenic&quot;) or non-human (&quot;nonpathogenic&quot;) pathogens. The test set contains paired reads (&quot;_1&quot; and &quot;_2&quot; for the first and second mate). The number of reads per species is proportional to the respective genome length. An additional, temporal test set (*temporal*fasta.gz) includes 15 species added after 12 weeks from the consturction of the original datasets.</p> <p>3) Fungal training sets. They contain 250bp-long reads simulated from species not present in the validation or test sets. There are four variants:<br> 3a) &quot;low-coverage, linear&quot; - 20 million reads, number of reads per species proportional to genome length<br> 3b) &quot;low-coverage, logarithmic&quot; - 20 million reads, number of reads per species proportional to the logarithm of genome length (&quot;log&quot;)<br> 3c) &quot;high-coverage, linear&quot; - 240 million reads, number of reads per species proportional to genome length&nbsp; (&quot;24&quot;)<br> 3d) &quot;high-coverage, logarithmic&quot; - 240 million reads, number of reads per species proportional to the logarithm of genome length&nbsp; (&quot;24log&quot;)</p> <p>4) Training, validation and test sets for the multiclass models. They should be used together with the &quot;pathogenic&quot; read sets hosted at <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>. Here, we share sets for two of the four total classes:<br> 4a) The &#39;non-pathogen&#39; class is a mixture of &quot;nonpathogenic&quot; biacterial and viral read sets, concatenated and downsampled to the original read number (20M for training, 1.25M for validation and test). The training and validation sets contain mixed-length (25-20bp) simulated subreads (original sets hosted here: <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>). The test set contains 250bp long reads based on the test sets from here: <a href="https://zenodo.org/record/3678563">https://zenodo.org/record/3678563</a> and here: <a href="https://zenodo.org/record/4312525">https://zenodo.org/record/4312525</a>; it was also sorted by species.<br> 4b) Mixed-length versions of the &quot;pathogenic&quot; fungal training and validation sets, prepared by random shortening of the &quot;low-coverage&quot; read sets in the &quot;linear&quot; (_rn_) and &quot;logarithmic&quot; (_rn_*log_) flavours.</p> <p>See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a></p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Pathogen-sugar interactions revealed by universal saturation transfer analysis

<p>Supporting data for the &quot;Pathogen-sugar interactions revealed by universal saturation transfer analysis&quot; manuscript.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria

<p>Supplementary dataset from &quot;<strong><em>Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria</em></strong>&quot;</p> <p>&nbsp;</p> <p><strong>Supplementary Figure legends</strong></p> <p><strong>Figure S1. Illustration of the transposon insertions around the <em>M. bovis </em>genome. </strong>Sequencing of the input library showed that transposon insertions were evenly distributed around the genome and 27,419 of the permissible 66,931 thymine&ndash;adenine dinucleotide (TA) sites contained an insertion representing an insertion density of ~41%. The outer ring are the genomic coordinates, the blue lines represent transposon insertions and the gray boxes indicate regions of that did not have any insertions. Plot made with Circlize (Gu et al, 2014).</p> <p>&nbsp;</p> <p><strong>Figure S2. Diversity of the output library isolated from lung and thoracic lymph node lesions compared to the input library. </strong>On average, libraries recovered from lung lesions contained 14,456 unique mutants and those recovered from the lymph nodes contained an average of 16,210 unique mutants. Insertion density is represented as a proportion of the TA sites that contained insertions. The numbers on the x-axis refer to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p>&nbsp;</p> <p><strong>Figure S3. Volcano plots showing the distribution of log<sub>2</sub> fold-changes and -log<sub>10</sub> of adjusted p-values for representative lung (A) and lymph node (B) samples. </strong>Adjusted p-values (BH-fdr correction) &lt; 0.000001 cluster at the limits of the plot and precision reflects the number of resampling iterations (10,000).</p> <p>&nbsp;</p> <p><strong>Figure S4. Scatterplot of mean log<sub>2</sub> fold change per gene for all lung samples against all thoracic lymph node samples</strong>. Correlation between mean log<sub>2</sub> fold change among genes between the tissues was calculated with Spearman&#39;s ranked correlation, = 0.878, p-value &lt; 2.2e-16.</p> <p>&nbsp;</p> <p><strong>Figure S5. Fold-changes caused by transposon insertions in <em>RD1<sup>BCG</sup></em> and <em>RD1<sup>MIC</sup> </em>in the lungs and lymph nodes of infected cattle. </strong>Boxplot for log<sub>2 </sub>fold-changes in genes of the RD1<sup>BCG</sup> region. Samples with adjusted p-values (BH-fdr corrected) &lt;0.05 are indicated with purple points. Gene names highlighted in magenta have fewer than 5 TA sites located in the gene; too few to determine the statistical significance of changes in insertion levels with this method.</p> <p>&nbsp;</p> <p><strong>Supplementary Tables </strong></p> <p><strong>Table S1. Sequencing statistics of the input and output transposon libraries. </strong>The numbers in the column labelled &ldquo;filename&rdquo; refers to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p>&nbsp;</p> <p><strong>Table S2. Tissues collected and scored for gross pathology. </strong>Tissues from head and neck lymph nodes (from the right and left sub-mandibular lymph nodes, the right and left medial retropharyngeal lymph nodes), thoracic lymph nodes (the right and left bronchial lymph nodes, the cranial tracheobronchial lymph nodes, the cranial and caudal mediastinal lymph nodes) and from lung lesions, were collected and scored.</p> <p>&nbsp;</p> <p><strong>Table S3. Log<sub>2</sub> fold-changes for insertions across the entire genome of <em>M. bovis</em> AF2122/97. </strong>Cells are coloured according to log<sub>2</sub> fold-change. Refer to the text for the gene groups in individual tabs.</p> <p>&nbsp;</p> <p><strong>Table S4. </strong>Custom transposon sequencing primers and adaptors used in sequencing of the transposon libraries.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix

<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation:&nbsp;https://doi.org/10.1101/2022.07.29.502101</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Antibiotic resistant pathogen outbreak investigation: an interdisciplinary module to teach fundamentals of evolutionary biology

<p>The evolution of resistance to antibiotics provides a timely and relevant topic for teaching undergraduate students evolutionary biology. Here, we present a module incorporating modified sequencing data from eight antibiotic resistant pathogen outbreaks in hospital settings with bioinformatics and phylogenetic analyses. This module uses whole genome sequencing data from hospital outbreaks investigated by the Centers for Disease Control and Prevention to provide examples of antibiotic resistance spread. Students work in groups to analyze outbreak data to identify the bacterial species and antibiotic resistance genes, to infer a phylogenetic tree examining relatedness among isolates, and to determine a possible source of the outbreak. Students then compile their results in individual reports and provide recommendations for preventing the further spread of antibiotic resistant organisms. In addition to providing genomic outbreak data, we include a teaching concepts guide discussing three integral components of the module: how evolutionary biology concepts of natural selection and competition impact antibiotic resistance; outbreak investigation information to aid in phylogenetic analysis and creation of recommendations; and instructions for the bioinformatics protocol. Completion of this module provides students an opportunity to think critically about the evolution of resistance, practice bioinformatics techniques, and relate evolutionary biology to current events.</p>

opencc-by-4.0Jan 2018View details →
zenodo44/100

Data for 'Genetic variation in trophic avoidance shows fruit flies are generally attracted to bacterial pathogens'

<p>Raw data dn R code for the analysis of data dn generation of all figures in the above referenced paper. Descriptions of each data file are included wihtin the R script.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Row sequcenes data for assessing the risks of potential pathogens and antibiotic resistance genes among heterogeneous habitats in a temperate estuary wetland

<p>The study included 118 usable samples within three different habitats (water, soil, and sediment) across the Liaohe River basin to the Red Beach wetland collected from seven papers, and all of the sequence files were uploaded for availability.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Supplementary dataset for the publication "Prevalence of tick-borne bacterial pathogens in Germany – has the situation changed after a decade?"

<p>The dataset supplements the journal article &nbsp;"Prevalence of tick-borne bacterial pathogens in Germany &ndash; has the situation changed after a decade?" published by Katja Mertens-Scholz, Bernd Hoffmann, J&ouml;rn M. Gethmann, Hanka Brangsch, Mathias W. Pletz and Christine Klaus in the journal mdpi microorganisms (DOI: &nbsp;<a href="https://doi.org/10.3389/fcimb.2024.1429667">https://doi.org/10.3389/fcimb.2024.1429667&nbsp;</a><a>)</a></p> <p>The file "Tick samples" contains all data regarding time, location and detected pathogens". The files "Rickettsia sequences" and Borrelia sequences" contains all sequence information of positive samples. The file "Rickettsia reference genes" contains information of used reference genes for phylogenetic analysis.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

BacSPaD: A robust bacterial strains' pathogenicity resource based on integrated and curated genomic metadata

<p>The vast array of omics data in microbiology presents significant opportunities for studying bacterial pathogenesis and creating computational tools for predicting pathogenic potential. However, the field lacks a comprehensive, curated resource that catalogs bacterial strains and their ability to cause human infections. Current methods for identifying pathogenicity determinants often introduce biases and miss critical aspects of bacterial pathogenesis.<br>In response to this gap, we introduce BacSPaD (Bacterial Strains&rsquo; Pathogenicity Database), a thoroughly curated database focusing on pathogenicity annotations for a wide range of high-quality, complete bacterial genomes. Our rule-based annotation workflow combines metadata from trusted sources with automated keyword matching, extensive manual curation, and detailed literature review. Our analysis classified 5,502 genomes as pathogenic to humans (HP) and 490 as non-pathogenic to humans (NHP), encompassing 532 species, 193 genera, and 96 families. Statistical analysis demonstrated a significant but moderate correlation between virulence factors and HP classification, highlighting the complexity of bacterial pathogenicity and the need for ongoing research. This resource is poised to enhance our understanding of bacterial pathogenicity mechanisms and aid in the development of predictive models. To improve accessibility and provide key visualization statistics, we developed a user-friendly web interface, accessible at<a href="https://bacspad.altrabio.com/"> </a><a href="https://bacspad.altrabio.com/"><u>https://bacspad.altrabio.com</u></a>.</p>

opencc-by-nc-sa-4.0Aug 2024View details →
zenodo44/100

Revealing real-time 3D in vivo pathogen dynamics in plants by label-free optical coherence tomography

<p>This repository contains all data and code underlying the publication: J. de Wit et al. "<em>Revealing real-time 3D in vivo pathogen dynamics in plants by label-free optical coherence tomography</em>" in Nature Communications (2024) (https://doi.org/10.1038/s41467-024-52594-x)</p> <p><strong>--------------Code description------------------</strong></p> <p>The set of scripts is largely organized around the figures. For each (sub)figure, also from supplementary materials, that involves data and plotting, there is a script that generates the plot from data that can be found in the different zip files that are present in the Zenodo repository under https://doi.org/10.5281/zenodo.11428245.</p> <p>The scripts use the data that is contained in the ZIP folders. The ZIP folders are organized by experiment (Experiment 1, including contrast optimization; Experiment 2), one for the other data (OtherData, the validation for with Trypan blue, and the Arabidopsis, Radish and nematode) and one as a smaller dataset to explain the method on a single B-scan (Example_Bscan_dynamicOCT).</p> <p>IMPORTANT: The folder where the ZIP files are unzipped should be put in the file '<em>basepath.txt</em>', such that the data can be automatically loaded.</p> <p>Besides the figures that mention 'MakeFig...' there are a few more scripts:</p> <ul> <li><em>pointcloud_generation_experiment1.py</em>: this file makes the point clouds from the dynamic OCT images as described in Fig 2b. The resulting data is saved as maximum intensity projections and axial sums(forming the basis for Fig S4, S6 and S7) and as voxel counts (forming the basis of Fig.2c and Fig S5)</li> <li><em>pointcloud_generation_timelapses.py</em>: this file does the segmentation for experiment 2 and saves the maximum intensity projections and axial sums of the different stages in the segmentation (forming the basis of Fig3a,d,e and FigS9a,b), and saves the point clouds of the data. These point clouds were refined manually in CloudCompare as described in methods. These segmented point clouds are contained in the data zip folder of experiment 2.</li> <li><em>StatisticalTests.R</em>: This R file calculates the statistical tests for Fig.2cd and Fig.S5d. Here the path is not automatically updated, and should be manually set. The input file is contained in "Experiment1/SegmentationData/segmentationdata_samples.csv" and the output of the file is "D:/DataZenodo/Experiment1/SegmentationData/data_combined_Rstats_output.csv"</li> <li><em>example_dynamic_Bscan.py</em>: This script gives an example of the dynamic OCT processing as proposed in this paper. First it shows the process from an OCT interference spectrum to a B-scan. Then it loads 100 B-scans and applies dynamic OCT, including normalization with histograms. Finally it gives a dynamic B-scan and plots this against the average normal OCT image. This script can be used with only the zip folder "Example_Bscan_dynamicOCT", which reduces the amount of data needed to download/unzip.</li> </ul> <p>The list of other script files to load the data and generate the figures (guiding to the path of uncropped figures) is:</p> <ul> <li><em>MakeFig1bce_Fig2e.py</em></li> <li><em>MakeFig1d.py</em></li> <li><em>MakeFig1agraphs_FigureS1.py</em></li> <li><em>MakeFig3acde_S9ab.py</em></li> <li><em>MakeFigS2_determine_dynamic_range_experiment1.py</em></li> <li><em>MakeFigS8.py</em></li> <li><em>MakeFigureS4-S6-S7.py</em></li> <li><em>MakeFigureS5.py</em></li> <li><em>MakeHistFig2b_makeFigS3b-e.py</em></li> <li><em>MakePlotsFig2ab.py</em></li> <li><em>MakePlotsFig2cd.py</em></li> </ul> <p>Code was all run in Python 3 using Anaconda Spyder.</p> <p>Moreover, the zip file with the code contains the folder '<em>figures</em>' with all subfigures. Some of them are automatically saved from the scripts, others (like photos, icons, but also the Trypan blue microscopy figure) are added in the respective folder. The are logically organized by figure number.</p> <p><strong>--------------Dataset Description-----------------</strong></p> <p>As mentioned above, the data is organized in four zip folders for both experiments, the other data (validation with Trypan blue, other plant-pathogens) and one for the dynamic OCT B-scan example. The data contain the following:</p> <p><strong>Experiment 1:&nbsp;</strong></p> <ul> <li>DynamicOCTimages whose subfolders (organized by date) contain a folder per volume dataset in experiment 1 with a z-stack of .tif files that form the imaged volume. The lateral sampling is 3 um and the axial sampling is 1.37 um.&nbsp;</li> <li>ContrastOptimization: This folder contains&nbsp; <ul> <li><em>Bscans_with_segmentation</em>: segmented B-scans for contrast optimization (Fig S3)</li> <li><em>Bscan_figS1_fig1</em>: The B-scans and segementation for Figure S1.</li> <li><em>histogramdata_dynamicrange</em>: The histograms, bins and deducted reference data for determining the dynamic range per color channel for experiment 1 (Fig S2)</li> <li><em>logcompressed_3value_dOCT_example</em>: An example data stack for obtaining histograms (see script MakeFigS2_determine_dynamic_range_experiment1.py)</li> <li><em>overlaps_threshold-100-98-95-92-90-85-80-75-70-65-60-55-50-45perc_red-1_2_blue_-3_0_green1_filt.npy</em>: A file with intermediate data for the contrast optimization, which can also be generated with the script "<em>MakeHistFig2b_makeFigS3b-e.py</em>"</li> </ul> </li> <li>SegmentationData: This folder contains&nbsp; <ul> <li><em>MIP_segmentation_stages</em>: maximum intensityp projections and axial sums for all images at different stages in the segmentation (basis for Fig S4,6,7)</li> <li><em>processed_masks and StackMasks</em>: the manually obtained masks (segmented in StackMasks, made into masks in the folder 'processed_masks') for filtering out stomata, veins and artefacts.</li> <li><em>Unmasked_axialsum_th34_formanualsegmentation</em>: This folder contains the images of Fig.S4 and were used for the segmentation (we addes a small offset, such that the in segmentation we could set it to 0 and have a unique mask).&nbsp;</li> <li><em>overview_samples_bremiayn.csv</em>: A dataframe with the data for all the samples in experiment 1 that is used as input for the segmentation. It also contains the result of the manual check whether it has infection (Fig2c, left).</li> <li><em>segmentationdata_samples.csv</em>: This supplements the file of overview_samples_bremiayn.csv with the results from the segmentation and is output to script "<em>pointcloud_generation_experiment1.py</em>". It forms the basis of Fig.2a-c, and FigS5, as well as for the R-script to do the statistical testing.</li> <li><em>qPCR_dOCT.csv</em>: This script contains the qPCR data and is input to Fig2d.&nbsp;</li> </ul> </li> </ul> <p><strong>Experiment 2:</strong></p> <ul> <li><em>DynamicOCTimages</em>: This contains the z-stacks of .tif files of the volumes for experiment 2 (and one extra, where a z-slice is used in Fig.1b, bottom). Sampling step size is here again 3 um in lateral direction and 1.37 um in axial direction.</li> <li><em>.npy files </em>with the histograms (with same bins as Experiment 1), maxvalues and reference values for the dynamic range calculation.</li> <li><em>segmentation_data</em>: this folder contains: <ul> <li><em>quantification_volume_disc160_33_10.csv</em> and <em>quantification_volume_disc160_33_10.xlsx</em>: data from the manually segmented point clouds that form the basis of Fig.3c.</li> <li><em>timelapse_sampleoverview.csv</em>: overview of the samples that is used as input in the file "<em>pointcloud_generation_timelapses.py</em>"</li> <li><em>pointclouds</em>: Folder with segmented point clouds for the three leaf discs. These files could &nbsp;be loaded in CloudCompare.</li> <li><em>overviewMIPs</em>: folder with overview maximum intensity projections for the different steps in segmentation, which also forms the input of Fig.3a, FigS9ab.</li> <li>rawpointclouds: folder with the automatically generated point clouds from file&nbsp;<em>pointcloud_generation_timelapses.py&nbsp;</em>which were imported into CloudCompare as the basis for the segmented point clouds.</li> </ul> </li> </ul> <p><strong>OtherData:</strong></p> <p>This folder contains the z-stacks of dynamic OCT tif images for Arabidopsis (here both a normal contrast and one that has been increased to only contain the original 0-180 range); nematodes, radish (called radijs_test_PP_py_0002), spores for Fig1c (SporesImaging) and the dynamic OCT image of Fig1d.&nbsp;</p> <p><strong>example_Bscan_dynamicOCT:</strong></p> <p>This folder contains data to run the script example_dynamic_Bscan.py to show the dynamic OCT imaging process from raw OCT spectra.</p> <ul> <li><em>raw_spectra_exampleframe:</em> contains interference spectra, a reference spectrum and interpolation grid to show how to get from a raw OCT spectrum to a normal single B-scan.</li> <li><em>abs_images:</em> contains 100 subsequent B-scans that can be used to generate a dynamic OCT image as done in example_dynamic_Bscan.py</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Global Cluster Test Results: Pathogenic Fungi in Decayed Norway Spruce Stands

<p><strong>Accessing the Results:</strong> Users can retrieve the test results by opening the dataset <code>Global_cluster_test_results.RData</code> in an R session and using the functions inside the script of the same name.</p> <p><strong>Description: </strong>This dataset provides global cluster test results analyzing the spatial distribution of pathogenic fungi in 273 Norway spruce stands in Norway (Lara et al., 2024). The stands, composed mainly of Norway spruce (27% to 100%), also include Scots pine and birch. It focuses on spatial patterns of decayed spruce trees, offering p-values, clustering metrics, and other parameters from statistical analyses.</p> <p><strong>Analysis:</strong> The dataset includes results from three global cluster tests (Tango, 2010):</p> <ul> <li>Tango's Nearest Neighbors (TNN)</li> <li>Tango's Double Exponential Clinal (TCN)</li> <li>Diggle and Chetwynd&rsquo;s (DC)</li> </ul> <p><strong>Methodology:</strong> Cluster testing employed 1,000 Monte Carlo simulations for each test across all stands to establish null distributions and adjusted p-values, ensuring robust statistical assessments under the random labeling hypothesis: H0: the observed n0 decayed trees are a random sample from the&nbsp;entire sample of size n = n0 + n1 (decayed trees + healthy trees) (Tango, 2010).</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics

<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table.&nbsp;</p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands &ldquo;--cut&rdquo; and &ldquo;--duplicate-proteins&rdquo; are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

A curated database of fungal pathogens and their host range

<p>This database contains a manually curated set of human, animal and plant pathogens, annotated with their confirmed host range and relevant sources. In addition to that, we include additional sets of plant-associated fungi (which may include non-pathogens), as well as fungi with an automatically assigned, putative human, animal or plant host. The labelled fungal species are linked to their representative GenBank genomes wherever possible. Genomes that were screened, but no label was found, are also included.</p> <p><strong>[Last update on: 11 Dec 2022]</strong><br> [Home page: <a href="https://dacs-hpi.gitlab.io/pathogenic-fungi/">https://dacs-hpi.gitlab.io/pathogenic-fungi/</a>]<br> <br> The database is stored in a flat-file format. All metadata are stored in all_data_[date].csv, and all_data_[date].rds contains the same data in a compressed format that can be easily loaded in R. The database was first compiled on 9 Oct 2021 (v1.0), and then updated on 2 Jan 2022 (v1.1) and 11 Dec 2022 (v1.2).</p> <p>The core database is limited to manually confirmed human, animal and plant pathogens with available genomes as of 9 Oct 2021. Those data are a subset of all_data, and are stored in core_fungal_pathogens.csv and core_fungal_pathogens.rds.</p> <p>The temporal-test subset contains confirmed pathogens with genomes added to GenBank between 9 Oct 2021 and 2 Jan 2022.</p> <p>You may also be interested in trained neural network models predicting pathogenic potentials of novel fungi from DNA sequences (<a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>) and simulated Illumina read sets used to train them (<a href="https://zenodo.org/record/5846397">https://zenodo.org/record/5846397</a>).<br> <br> See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a> and <strong>the paper</strong> presented at ECCB &#39;22 and published in <em>Bioinformatics:</em> <a href="https://doi.org/10.1093/bioinformatics/btac495">https://doi.org/10.1093/bioinformatics/btac495.</a></p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Genome-wide screen reveals Rab12 GTPase as a critical activator of pathogenic LRRK2 kinase

<p>Primary data associated with the figure 1 of the manuscript &quot;<strong>Genome wide screen reveals Rab12 GTPase as a critical activator of pathogenic LRRK2 kinase&quot;&nbsp;</strong>(Herschel S. Dhekne, Francesca Tonelli, Wondwossen M. Yeshaw,&nbsp;Claire Y. Chiang, Charles Limouse, Ebsy Jaimon, Elena Purlyte,&nbsp;Dario Alessi, and Suzanne Pfeffer).&nbsp;</p> <p>These include&nbsp;</p> <p>- data (annotated .tiff exports) from Metamorph acquired spinning disk confocal microscope</p> <p>- sequencing data as fastq files .gz files from Hiseq or Miseq next generation sequencing,&nbsp;&nbsp;</p> <p>- graphs made using GraphPad Prism&nbsp;(.pzf files).</p> <p>- flow cytometry .fcs files and workspace files</p> <p>-&nbsp;<a href="https://zenodo.org/api/files/de4d1da1-785a-4f55-a542-876647a2a5ad/Figure%201-%20file%20names%20legend.xlsx">Figure 1- file names legend.xlsx</a>&nbsp;excel sheet explaining the details</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Data from: Pathogen community composition and co-infection patterns in a wild community of rodents

<p><strong>ABSTRACT</strong></p> <p>Rodents are major reservoirs of pathogens that can cause disease in humans and livestock. It is therefore important to know what pathogens naturally circulate in rodent populations, and to understand the factors that may influence their distribution in the wild. Here, we describe the incidence and distribution patterns of a range of endemic and zoonotic pathogens circulating among rodent communities in northern France. The community sample consisted of 713 rodents, including 11&nbsp; host species &nbsp;from diverse habitats. Rodents were screened for virus exposure (hantaviruses, cowpox virus, Lymphocytic choriomeningitis virus, Tick-borne encephalitis virus) using antibody assays. Bacterial communities were characterized using 16S rRNA amplicon sequencing of splenic samples. Multiple correspondence (MCA), regression and association screening (SCN) analyses were used to determine the degree to which extrinsic factors contributed to pathogen community structure, and to identify patterns of associations between pathogens within hosts. We found a rich diversity of bacterial genera, with 36 known or suspected to be pathogenic. We revealed that host species is the most important determinant of pathogen community composition, and that hosts that share habitats can have very different pathogen communities. Pathogen diversity and co-infection rates also vary among host species. Aggregation of pathogens responsible for zoonotic diseases suggests that some rodent species may be more important for transmission risk than others. Moreover we detected positive associations between several pathogens, including <em>Bartonella</em>, <em>Mycoplasma</em> species, Cowpox virus (CPXV) and hantaviruses, and these patterns were generally specific to particular host species. Altogether, our results suggest that host and pathogen specificity is the most important driver of pathogen community structure, and that interspecific pathogen-pathogen associations also depend on host species.</p> <p><strong>FILE DESCRIPTION:</strong></p> <p><strong>MiSeq raw sequences of the 16Sv4 rRNA gene from spleen rodent samples</strong></p> <p>This ZIP file contains the FASTQ files of the paired-end reads (R1: reads 1; R2: reads 2) produced for each spleen rodent sample using the MiSeq platform. The 749 multiplexed PCR products were indexed using both forward and reverse indices. Information of the multiplexed samples (<em>n</em>=363 in replicate) and positive (<em>n</em>= 6) &amp; negative controls (<em>n</em>= 17) is provided in the following XLSX file titled: 16S_raw_abundance_data.xlsx</p> <p>File name: <strong>MiSeq raw sequences of the V4 region 16S rRNA gene.zip</strong></p> <p><strong>Raw input and output files generated by the mothur program</strong></p> <p>This ZIP file contains all the input and output files generated during the MiSeq sequence analysis with the mothur program.</p> <p>File name: <strong>Raw input and output files generated by the mothur program.zip</strong></p> <p><strong>Log file generated by the mothur program</strong></p> <p>This TXT file contains is the history of all the command lines and parameters used during the MiSeq sequence analysis with the mothur program.</p> <p>File name: <strong>mothur.1428506786.logfile</strong></p> <p><strong>Raw abundance table of the 16v4 rRNA gene from spleen rodent samples before data filtering</strong></p> <p>This XLSX file contains the number of reads for each distinct Operational Taxonomic Unit (OTU) and each of the PCR products, including the 332 spleen rodent samples analyzed in the study and the negative &amp; positive controls, sequenced in the MiSeq run before the data filtering. This file contains also the following information: Study_site, Study_year, Sample_habitat, Host_species, Host_age, Host_sex, PCR_ID and the taxonomic classification (Kingdom to Genus) of each OTU.</p> <p>File name: <strong>16S_raw_abundance_data.xlsx</strong></p> <p><strong>Occurrence table of the 16v4 rRNA gene from spleen rodent samples after data filtering</strong></p> <p>This XLSX file contains the occurrences (presence: 1 ; absence: 0) after data filtering of each putative pathogenic Operational Taxonomic Unit (OTU) for each of the 332 spleen rodent samples analyzed in the study.</p> <p>File name: <strong>16S_presence_absence_data.xlsx</strong></p> <p><strong>Statistical Analysis Scripts and Data File</strong></p> <p>This ZIP file contains the R scripts for performing statistical analyses reported in the main text and supplemental materials. There is one main file (Analyses.R), as well as two source scripts required for association screening analyses (SCN.txt and FctTestScreenENV.txt). It also includes an R-legible data file containing occurrences (presence: 1 ; absence: 0) for all pathogen exposure variables on which statistical analyses were conducted (PA_DATA.csv) for each of the 332 spleen rodent samples analyzed in the study. The column names for bacterial exposures correspond to the &ldquo;Pathogen Code&rdquo; given in the 16S_presence_absence_data.xlsx file.</p> <p>File name:&nbsp;<strong>Statistical Analysis Scripts and Data File.zip</strong></p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

H4K20me3 is important for Ash1-mediated H3K36me3 and transcriptional silencing in facultative heterochromatin in a fungal pathogen

<p>Normalized ChIP-seq datasets&nbsp;for visualization in IGV. The tracks contain means of pooled replicate datasets.</p> <p>ChIP-seq data were quality-filtered and adapters removed with trimmomatic v.0.39&nbsp;(Bolger et al., 2014). Mapping was performed with bowtie2 v.2.4.4&nbsp;(Langmead and Salzberg, 2012), and sorting and indexing with samtools v.1.9&nbsp;(Li, 2011). Normalized coverage bigwig files and heatmaps were created with deeptools v.3.5.1&nbsp;(Ram&iacute;rez et al., 2016). Wiggletools v.1.2 and the UCSC Genome Browser tools were used to calculate means for replicates and converting wig to bigwig files.</p> <p>Reference genome file is modified from Goodwin et al., 2011. Chromosome 18 was removed from the genome as our reference isolate&nbsp;is missing chromosome 18.&nbsp;</p> <p>Gene annotation file was obtained from FungiDB (release&nbsp;53) and is based on the annotation published by Grandaubert et al., 2015.</p> <p>In this version, we have added new ChIP-seq bw tracks for ∆ash1::ash1-gfp-V5&nbsp;and ∆kmt5::kmt5 complementation experiments. All tracks coming from this experiment are labeled *_compl_exp_mean.bw.</p> <p>We also added ChIP peak files (peaks called with HOMER: Heinz et&nbsp;al., 2010) for H4K20me3, H3K36me3 and H3K27me3 in WT, ∆kmt5 and ∆ash1, as well as H3K36me3 peak files for Set2- and Ash1-mediated H3K36me3.</p> <p>We have also added bed files (500 bp windows)&nbsp;containing facultative heterochromatin clusters 1 (Zt09_500bp_K27filtered_K36_K20_cluster1.bed) and 2 (Zt09_500bp_K27filtered_K36_K20_cluster2.bed).&nbsp;</p>

opencc-by-4.0Jan 2023View details →
edi44/100

Decomposition of Microstegium vimineum litter, plants grew through the Big Oaks National Wildlife Refuge in 2019. Litter used in this experiment naturally senesced in the fall 2019, decomposition data collected through 2020. Plants were infected or not-infected with the foliar fungal pathogen Bipolaris gigantea during the 2019 growing season.

Decomposition of plant litter, facilitated primarily by microbial decomposers, plays a critical role in biogeochemical cycling and ecosystem function. Emerging pathogens have the potential to impact litter decomposition by altering the chemical composition and associated microbial community of host tissue. Here, we compared litter decomposition of the invasive grass Microstegium vimineum collected from sites with Bipolaris leaf spot symptoms and sites with no apparent disease symptoms in a common garden experiment. Our results revealed that leaf tissue from litter from non-infected sites decomposed more rapidly through the spring than litter from infected sites. Differences in fungal composition between infected and non-infected litter at the start of the experiment largely persisted through the summer. Our work demonstrates that pathogen colonization may facilitate the persistence of infected host litter, potentially slowing the return of nutrients to the environmental pool while also promoting the survival and dispersal of primary inoculum the following season.

openCC (other)Jun 2023View details →
edi44/100

Emerging fungal pathogen of an invasive grass: Implications for competition with native plant species

This data package includes data and code from an experiment testing the effects of a leaf spot fungal infection and competition from the invasive (to the U.S.) grass Microstegium vimineum on the performance of three native grass species: Dichanthelium clandestinum, Elymus virginicus, and Eragrostis spectabilis. The experiment was performed between June and September of 2019 in a greenhouse on the University of Florida campus in Gainesville, FL, USA. The leaf spot infection is caused by the fungal pathogen Bipolaris gigantea, which has recently emerged on populations of M. vimineum in the U.S. We tested the hypothesis that infection of B. gigantea would both directly and indirectly affect the native grass species by measuring the change in biomass of each species with and without pathogen inoculation (direct effects) and by measuring the effect of pathogen inoculation on M. vimineum competition through changes in native grass biomass across a density gradient of M. vimneum (indirect effects). The code includes statistical analyses and figures. The code was run using R (version 4.0.1).

openCC (other)Feb 2021View details →
zenodo40/100

Genomic determinants of pathogenicity in SARS-CoV-2 and other human coronaviruses

<p><strong>Dataset S1.</strong>Complete nucleotide sequence alignment of all human CoV used for region identification.&nbsp;</p> <p><strong>Dataset S2.</strong>Complete nucleotide sequence alignment of all CoV (of human and non-human hosts).</p> <p><strong>Dataset S3.</strong>Distances between leaves (each CoV strain in Dataset S2 was considered), from every reference genome of each of the seven human CoV.</p> <p><strong>Dataset S4.</strong>Alignment of strains used for zoonotic jump analysis.</p>

opencc-by-4.0May 2020View details →
dryad40/100

Data from: Using genetic relatedness to understand heterogeneous distributions of urban rat-associated pathogens

<p>Urban Norway rats (<i>Rattus norvegicus</i>) carry several pathogens transmissible to people. However, pathogen prevalence can vary across fine spatial scales (i.e., by city block). Using a population genomics approach, we sought to describe rat movement patterns across an urban landscape, and to evaluate whether these patterns align with pathogen distributions. We genotyped 605 rats from a single neighborhood in Vancouver, Canada and used 1,495 genome-wide single nucleotide polymorphisms to identify parent-offspring and sibling relationships using pedigree analysis. We resolved 1,246 pairs of relatives, of which only 1% of pairs were captured in different city blocks. Relatives were primarily caught within 33 meters of each other leading to a highly leptokurtic distribution of dispersal distances. Using binomial generalized linear mixed models we evaluated whether family relationships influenced rat pathogen status with the bacterial pathogens <i>Leptospira interrogans</i>, <i>Bartonella tribocorum</i>, and <i>Clostridium difficile</i>, and found that an individual's pathogen status was not predicted any better by including disease status of related rats. The spatial clustering of related rats and their pathogens lends support to the hypothesis that spatially restricted movement promotes the heterogeneous patterns of pathogen prevalence evidenced in this population. <span>Our findings also highlight the utility of evolutionary tools to understand movement and rat-associated health risks in urban landscapes.</span></p>

opencc-zeroDec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record