Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10,950

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10,950 results for “Host”

Learn how ShareScore rates datasets ↗
zenodo44/100

Host Network Traffic 2019

<p><strong><em>Dataset Summary</em></strong></p> <ul> <li><strong>Timespan</strong>: 2019-01-01 : 2019-12-31</li> <li><strong>Granularity:&nbsp;</strong>1-hour disjoint time windows</li> <li><strong># of&nbsp;characteristics observed:&nbsp;</strong>9</li> <li><strong>Hosts observed: </strong>65536</li> <li><strong>Labels:&nbsp;</strong>included</li> <li><strong>Unzipped volume:&nbsp;</strong>approx. 10 GB</li> </ul> <p><strong><em>Dataset Origins</em></strong></p> <p>Dataset&nbsp;was collected over the <strong>whole year</strong>&nbsp;<strong>&nbsp;2019</strong>. The observation points for the collection of IP flows were located at the borders of the university campus network. The campus university network has /16 CIDR IPv4 network range at disposal and contains various network segments from segments connecting dormitories, over server segments, to a segment containing working stations of university administrative workers.&nbsp;<strong>A host in our dataset is identified by its source IPv4 address. &nbsp;</strong></p> <p><em><strong>Variables</strong></em></p> <p>The dataset contains the following variables:</p> <ul> <li><strong>Aggregations</strong>&nbsp;- created sums of the individual variables over a one-hour interval: <ul> </ul> <ul> <li><strong># of flows &nbsp;</strong>- number of flows for a given source IP&nbsp;</li> <li><strong># of packets </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong># of bytes </strong>&nbsp;-&nbsp;number of packets for a given source IP</li> <li><strong>flow duration </strong>&nbsp;- average flow duration in seconds</li> </ul> </li> <li><strong>Distinct Counts&nbsp;</strong>- count of distinct values for each variable over a one-hour window <ul> <li><strong># of peers </strong>&nbsp;- number of distinct communication peers for a given source IP</li> <li><strong># of ports </strong>&nbsp;- number of distinct destination ports&nbsp;for a given source IP</li> <li><strong># of protocols</strong>&nbsp;- number of distinct communication protocols&nbsp;for a given source IP</li> <li><strong># of AS numbers</strong>&nbsp;- number of distinct destination AS numbers for a given source IP</li> <li><strong># of countries </strong>&nbsp;- number of distinct destination countries&nbsp;for a given source&nbsp;</li> </ul> </li> </ul> <p><em><strong>Dataset Structure</strong></em></p> <ul> <li><strong>Dataset Files</strong> - each variable is contained in one <strong>Comma-Separated File (.csv)&nbsp;</strong>file <ul> <li><strong>Row index&nbsp;-&nbsp;</strong>&nbsp;timestamp of the observation window (8760 rows)</li> <li><strong>Columns index -&nbsp;</strong>&nbsp;anonymized IP addresses (65536&nbsp;columns)</li> </ul> </li> <li><strong>Label File -&nbsp;</strong>contains labels of the individual IP addresses from the Dataset Files <ul> <li><strong>Row index </strong>- anonymized IP addresses (65536 rows)</li> <li><strong>Columns index </strong>- labels for the IP addresses <ul> <li><strong>Subnet </strong>- ID&nbsp;of a subnet - hosts belonging to the same subnet have the same Id.</li> <li><strong>Subnet_range&nbsp;</strong>- CIDR range of a&nbsp;subnet</li> <li><strong>Unit -&nbsp;</strong>an ID of&nbsp;&nbsp;administrative unit owning the network range</li> <li><strong>Sub-unit </strong>&nbsp;- an ID of&nbsp;&nbsp;administrative sub-unit owning the network range</li> <li><strong>Subnet_label -&nbsp;&nbsp;</strong>subnet label <ul> <li><strong>Servers - </strong>selected subnets containing mostly servers (133.250.178.0/24, 133.250.163.0/24)</li> <li><strong>Workstations - </strong>selected subnets containing mostly workstations&nbsp;(133.250.146.0/24,&nbsp;133.250.157.128/25)</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p><strong><em>Further notes</em></strong></p> <ul> <li><strong>N/A values </strong> <ul> <li><strong>Variables&nbsp;</strong>- means that in a given observation window, the host did not communicate</li> <li><strong>Labels -&nbsp;</strong>no additional information on this IP is available</li> </ul> </li> <li><strong>Dataset load&nbsp;</strong> <ul> <li> <pre><code class="language-python">df = pd.read_csv(&lt;filename&gt;,header=[0], index_col=[0])</code></pre> </li> </ul> </li> </ul>

opencc-by-4.0Apr 2020View details →
zenodo44/100

liampshaw/Pathogen-host-range: Pathogen-host-range initial code release

<p>Release of code and dataset for publication of associated paper: &quot;The phylogenetic range of bacterial and viral pathogens of vertebrates&quot; (doi: 10.1111/mec.15463).</p>

openmit-licenseMay 2020View details →
zenodo44/100

Coevolving plasmids drive gene flow and genome plasticity in host-associated intracellular bacteria

<p>Comparative genomics and modeling of plasmids of the obligate host-associated intracellular phylum chlamydiae.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

A dissymmetric [Gd2] coordination molecular dimer hosting six addressable spin qubits. Open data sets

<p>Includes data relevant for publication with DOI&nbsp;<a href="https://doi.org/10.1038/s42004-020-00422-w">10.1038/s42004-020-00422-w</a>&nbsp;plus a table with information about how the data were obtained and processed.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Data from: Choosy beetles: how host trees and southern boreal forest naturalness may determine dead wood beetle communities

<p>See methods section of paper for detailed information on dataset&nbsp;and sources; briefly, these .csv&nbsp;files includes numbers of each beetle species captured at all sites used in the project, as well as information about each site and about each species.</p> <p>&nbsp;</p> <p>Data from:</p> <p><strong>Choosy beetles: how host trees and southern boreal forest naturalness may determine dead wood beetle communitie</strong><strong>s</strong></p> <p>Ryan C. Burner, Tone Birkemoe, J&ouml;rg G. Stephan, Lukas Drag, J&ouml;rg Muller, Otso Ovakainen, M&aacute;ria Potterf, Olav Skarpaas, Tord Snall, Anne Sverdrup-Thygeson</p> <p>Forest Ecology and Management, 2021</p> <p>&nbsp;</p> <p>From abstract of paper:</p> <p>Wood-living beetles make up a large proportion of forest biodiversity, and contribute to important ecosystem services, including decomposition. Beetle communities in managed southern boreal forests are less species rich than in natural and near-natural forest stands. In addition, many beetle species rely primarily on specific tree species. Yet, the associations between individual beetle species, forest management category, and tree species are seldom quantified, even for red-listed beetles. We compiled a beetle capture dataset from flight intercept traps placed in Norway spruce (<em>Picea abies</em>), oak (<em>Quercus sp.</em>), and Eurasian aspen (<em>Populus tremulae</em>) trees in 413 sites in mature managed forest, near-natural forest, and clear-cuts in southeastern Norway. We used joint species distribution models to estimate the strength of associations for 368 saproxylic beetle species (including 20 vulnerable, endangered, or critical red-listed species) for each forest management category and tree species. Tree species on which traps were mounted had the largest effect on beetle communities; oaks had the most highly associated beetle species, including most of the red-listed species, followed by Norway spruce and Eurasian aspen. Most beetle species were more likely to be captured in near-natural than in mature managed forest. Our estimated associations were compatible &ndash; for many species &ndash; with categorical classifications found in several existing databases of saproxylic beetle preferences. These quantitative beetle-habitat associations will improve future analyses that have typically relied on categorical classifications. Our results highlight the need to prioritize conservation of near-natural forests and oak trees in Scandinavia to protect the habitat of many red-listed species in particular. Furthermore, we underline the importance of carefully considering the species of trees on which traps are mounted in order to representatively sample beetle communities in forest stands.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

The within-host population dynamics of Mycobacterium tuberculosis vary with treatment efficacy.

<p>Data used for the publication of a paper entitled: <strong>The within-host population dynamics of <em>Mycobacterium tuberculosis</em> vary with treatment efficacy.</strong></p> <p>The data were derived from:</p> <p>1. the deep sequencing of serial sputum samples from 12 TB patients,</p> <p>2. the deep sequencing of liquid cultures derived from the expansion of individual colonies <em>in vitro</em>,</p> <p>3. <em>In </em><em>silico</em> simulations of DNA sequencing, populations and mutagenesis.</p> <p>The analytical scripts associated with the generation of the data can be found at:</p> <p>https://github.com/swisstph/TBRU_serialTB/</p> <p><strong>Paper Abstract:</strong></p> <p><strong>Background:</strong></p> <p>Combination therapy is one of the most effective tools for limiting the emergence of drug resistance. Despite the widespread adoption of combination therapy across diseases, drug resistance rates continue to rise, leading to failing treatment regimens. The mechanisms underlying treatment failure are well studied, but the processes governing successful combination therapy are poorly understood. We addressed this question by studying the population dynamics of <em>Mycobacterium tuberculosis</em> within tuberculosis patients undergoing treatment with different combinations of antibiotics.</p> <p><strong>Results:</strong></p> <p>By combining very deep whole genome sequencing (~1,000-fold genome-wide coverage) with sequential sputum sampling, we were able to detect transient genetic diversity driven by the apparently continuous turnover of minor alleles, which could serve as the source of drug-resistant bacteria. However, we report that treatment efficacy had a clear impact on the population dynamics: sufficient drug pressure bore a clear signature of purifying selection leading to apparent genetic stability. In contrast, <em>M. tuberculosis</em> populations subject to less drug pressure showed markedly different dynamics, including cases of acquisition of additional drug resistance.</p> <p><strong>Conclusions:</strong></p> <p>Our findings show that for a pathogen like <em>M. tuberculosis</em>, which is well adapted to the human host, purifying selection constrains the evolutionary trajectory to resistance in effectively treated individuals. Nonetheless, we also report a continuous turnover of minor variants, which could give rise to the emergence of drug resistance in cases of drug pressure weakening. Monitoring bacterial population dynamics could therefore provide an informative metric for assessing the efficacy of novel drug combinations.</p>

opencc-by-sa-4.0Dec 2016View details →
zenodo44/100

Update of the Xylella spp. host plant database

<p>Following a request from the European Commission, in 2018 EFSA released a renovated database of host plant species of <em>Xylella</em> spp. (<em>including both species</em> <em>X. fastidiosa </em>and <em>X.&nbsp;taiwanensis</em><em>) together with a scientific report</em> (EFSA, 2018). EFSA was tasked to maintain and update this database periodically. The mandate now covers the period 2021-2026 and EFSA is requested to release an update of the database twice per year.</p> <p>In July 2025 EFSA released the twelfth update of the&nbsp;<em>Xylella</em> spp. host plant database (VERSION 12) with information retrieved from literature search up to December 2024 and recent Europhyt outbreak notifications (EFSA, 2025). The protocol applied for the extensive literature review, data collection and reporting, as well as results and lists of host plants are described in detail in the related scientific report (EFSA, 2025).</p> <p>The overall number of <em>Xylella</em> spp. host plants determined with at least two different detection methods or positive with one method (between: sequencing, pure culture isolation) reaches now 463 plant species, 210 genera and 71 families (category A &ndash; see section 2.4.2 of EFSA (2025)). Such numbers rise to 727 plant species, 319 genera and 91 families if considered regardless of the detection method applied (category E, see section 2.4.2 of EFSA (2025)).</p> <p>The Excel files here attached represent the VERSION 12 of the <em>Xylella</em> spp. host plants database. For a detailed description of the information included in the database, please consult the related scientific report (EFSA, 2025).</p> <p>The Excel file &ldquo;<em>Xylella</em> spp. host plants database &ndash; VERSION 12&rdquo; contains several sheets: the LEGENDA (with extensive description of each table), the full detailed raw data of the <em>Xylella</em> spp. host plant database (sheet &ldquo;observation&rdquo;) and several examples of data extraction.</p> <p>Additional Excel files contain the lists of host plant species of <em>X. fastidiosa</em> (subsp. unknown (i.e. not reported), <em>fastidiosa</em>, <em>multiplex</em>, <em>pauca</em>, <em>morus</em>, <em>sandyi</em>, <em>tashke</em>, <em>fastidiosa/sandyi</em>) and <em>X. taiwanensis</em> infected naturally, artificially and in not specified conditions, and according to different categories (A, B, C, D, E &ndash; see section 2.4.2 of EFSA (2025)). The Excel file &ldquo;new_host_plant_species_v12&rdquo; contain the list of new host plant species added to the database in this new update.</p> <p><strong>Question number: EFSA-Q-2025-00045</strong></p> <p><strong>Output number: EN-9564</strong></p> <p><strong>Contacts: plants@efsa.europa.eu</strong></p> <p><em>Bibliography:</em></p> <p>EFSA (European Food Safety Authority). (2018). Scientific report on the update of the&nbsp;<em>Xylella</em> spp. host plant database. <em>EFSA Journal 2018</em>, <em>16</em>(9), 5408, 87 pp. <a href="https://doi.org/10.2903/j.efsa.2018.5408">https://doi.org/10.2903/j.efsa.2018.5408</a>&nbsp;</p> <p>EFSA (European Food Safety Authority), Cavalieri, V., Fasanelli, E., Furnari, G., Gibin, D., Gutierrez Linares, A., La Notte, P., Pasinato, L., &amp; Stancanelli, G. (2025). Update of the&nbsp;<em>Xylella</em> spp. host plant database &ndash; Systematic literature search up to 31 December 2024. <em>EFSA Journal</em>, <em>23</em>(7), e9563. <a href="https://doi.org/10.2903/j.efsa.2025.9563">https://doi.org/10.2903/j.efsa.2025.9563</a></p>

opencc-by-4.0Sep 2018View details →
zenodo44/100

PHI-base: the Pathogen-Host Interactions Database, version 5.1

<p><strong>Download the dataset here: <a title="Download PHI-base 5.1" href="https://zenodo.org/records/16738930/files/phi-base_v5.1.zip?download=1">phi-base_v5.1.zip</a></strong></p> <p>The Pathogen&ndash;Host Interactions Database (PHI-base) is an online database that catalogues experimentally-verified pathogenicity, virulence and effector genes from fungal, oomycete, and bacterial pathogens, which infect animal, plant, fungal, and insect hosts. PHI-base is a valuable resource in the discovery of genes in medically and agronomically important pathogens, which may be potential targets for chemical intervention.</p> <p>Information in PHI-base is manually curated by domain experts and is supported by strong experimental evidence (for example, gene disruption and gene complementation experiments), as well as references to the literature in which the original experiments are described. Annotations are made using terms from ontologies and controlled vocabularies, including the <a href="https://www.geneontology.org/">Gene Ontology</a> (GO), <a href="https://pubmed.ncbi.nlm.nih.gov/21030441/">Brenda Tissue Ontology</a> (BTO), and the <a href="https://obofoundry.org/ontology/phipo.html">Pathogen&ndash;Host Interaction Phenotype Ontology</a> (PHIPO).</p> <p>PHI-base 5 includes data that was curated using a new curation process described in&nbsp;<a href="https://doi.org/10.7554/eLife.84658">Cuzick et. al</a>&nbsp;(2023). Data releases for PHI-base 5 do not use the same schema as data releases from PHI-base 4, but all data records from PHI-base 4 that can be made compatible with the new schema are included with this release. Data releases from PHI-base 4 and PHI-base 5 will occur in parallel until such time that all data from PHI-base 4 can be migrated to PHI-base 5. The PHI-base 4 data releases are available on Zenodo at&nbsp;<a href="https://zenodo.org/doi/10.5281/zenodo.5356870">https://zenodo.org/doi/10.5281/zenodo.5356870</a>.</p> <p>For more information about the planned transition from PHI-base 4 to PHI-base 5, see the&nbsp;<a href="https://phi5.phi-base.org/#/help">Help</a>&nbsp;and&nbsp;<a href="https://phi5.phi-base.org/#/announcements">Announcements</a>&nbsp;page on the PHI-base 5 website.</p> <h2>Release statistics</h2> <p>This version of the PHI-base 5 dataset contains the following types of information:</p> <table style="border-collapse: collapse; border-width: 1px; width: 40.2168%; height: 362.8px;"> <thead> <tr style="height: 19.6px;"> <th style="border-width: 1px; width: 84.846%; height: 19.6px;">Data type</th> <th style="border-width: 1px; width: 15.4054%; height: 19.6px;">Count</th> </tr> </thead> <tbody> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Genes</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">9457</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Interactions</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">31094</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Pathogen species</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">303</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Host species</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">237</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Diseases</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">343</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">References</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">5202</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;"><strong>Annotations</strong></td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">&nbsp;</td> </tr> <tr style="height: 10px;"> <td style="border-width: 1px; width: 84.846%; height: 10px;">Pathogen-host interaction phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 10px;">18260</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Gene-for-gene phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">452</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Pathogen phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">9413</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Host phenotype</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">14</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO biological process</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">1453</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO cellular component</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">85</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">GO molecular function</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">152</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Post-translational modification</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">6</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">Physical interaction</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">53</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">WT RNA expression</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">36</td> </tr> <tr style="height: 19.6px;"> <td style="border-width: 1px; width: 84.846%; height: 19.6px;">WT protein expression</td> <td style="border-width: 1px; width: 15.4054%; height: 19.6px;">2</td> </tr> </tbody> </table> <h2>File contents</h2> <ul> <li> <p><strong>phi-base_v5.1.xlsx</strong>: the PHI-base dataset as an Excel spreadsheet. This format follows the layout of the PHI-base 5 website, with sheets corresponding to the sections of gene pages on the website. This format is designed for use by non-technical users.</p> </li> <li> <p><strong>phi-base_v5.1.json</strong>: the PHI-base dataset in JSON format. This is modelled on the export format used by PHI-Canto, the curation tool used by PHI-base. This format is primarily intended for programmatic usage and has additional information (e.g.&nbsp;metadata for curation sessions) that is not included in the spreadsheet format.</p> </li> <li> <p><strong>phi-base.schema.json</strong>: a&nbsp;<a href="https://json-schema.org/">JSON Schema</a> file for the JSON format of the dataset. This is included as documentation for the fields in the JSON file, but can also be used to validate the dataset.</p> </li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Pathogen lifestyle determines host genetic signature of quantitative disease resistance loci in oilseed rape (Brassica napus)

<p>Supplemental datasets associated with publication:&nbsp;Pathogen lifestyle determines host genetic signature of quantitative disease resistance loci in oilseed rape (<em>Brassica napus</em>)</p> <p><strong>Abstract</strong></p> <ul> <li>Crops are affected by several pathogens, but these are rarely studied in parallel to identify common and unique genetic factors controlling diseases. Broad-spectrum quantitative disease resistance (QDR) is desirable for crop breeding as it confers resistance to several pathogen species.</li> <li>Here, we use associative transcriptomics (AT) to identify candidate gene loci associated with <em>Brassica napus</em> constitutive QDR to four contrasting fungal pathogens:&nbsp;<em>Alternaria brassicicola</em>, <em>Botrytis cinerea</em>, <em>Pyrenopeziza</em><em> brassicae</em> and <em>Verticillium longisporum.&nbsp;</em>We did not identify any loci associated with broad-spectrum QDR to fungal pathogens with contrasting lifestyles. Instead, we observed QDR dependent on the lifestyle of the pathogen&mdash;hemibiotrophic and necrotrophic pathogens had distinct QDR responses and associated loci, including some loci associated with early immunity. Furthermore, we identify a genomic deletion associated with resistance to <em>V. longisporum </em>and potentially broad-spectrum QDR.</li> <li>This is the first time AT has been used for several pathosystems simultaneously to identify host genetic loci involved in broad-spectrum QDR. We highlight constitutively expressed candidate loci for broad-spectrum QDR with no antagonistic effects on susceptibility to the other pathogens studies as candidates for crop breeding. In conclusion, this study represents and advancement in our understanding if broad-spectrum QDR in <em>B. napus&nbsp;</em>and is a significant resource for the scientific community. &nbsp;</li> </ul> <p><strong>Description of data files</strong></p> <p><strong>Full dataset for input into AT analysis&nbsp; </strong>Full datasets (infection phenotypes for&nbsp;<em>A. brassicicola, B. cinerea, </em>or&nbsp;<em>V.longisporum,&nbsp;</em>ROS measurements for chitin, flg22, or elf18) and link to original <em>P. brassicae&nbsp;</em>dataset. These datasets were used for input into the Associative Transcriptomics pipeline (Nichols, 2022,&nbsp;<a href="https://github.com/bsnichols/GAGA. https://zenodo.org/badge/latestdoi/512807075">https://github.com/bsnichols/GAGA. https://zenodo.org/badge/latestdoi/512807075</a>).&nbsp;</p> <p><strong>Table S1 </strong>Mean, normalized phenotype data for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). These data were used for association transcriptomic analysis.<strong>&nbsp;</strong></p> <p><strong>Table S2 </strong>Full list of single nucleotide polymorphism (SNP) markers and significance levels from genome-wide association (GWA) analyses for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). Each excel tab contains the analyses for a single trait. The best fit model for GWA analysis is indicated in the tab title. Manhattan plots showing marker-trait association are included for data visualization; x-axis indicates SNP location along the chromosome; the y-axis indicates the -log10(p) (P value). Qqplots are included to demonstrate model fit.</p> <p><strong>Table S3</strong> Full list of gene expression markers (GEMs) and significance levels from GEM analyses for resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae and Verticillium longisporum</em>) and ROS response induced by PAMPS (chitin, flg22, and elf18). Each excel tab contains the analyses for a single trait. Manhattan plots showing marker-trait association are included for data visualization; x-axis indicates GEM location along the chromosome; the y-axis indicates the -log10(p) (P value).&nbsp;</p> <p><strong>Table S4 </strong>184 gene expression markers (GEMs) associated with chitin-induced ROS compared with GEMs associated with resistance to pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and<em> Verticillium longisporum</em>) and ROS response induced by flg22, and elf18. Lists correspond to Venn diagrams in Fig. 2. The first tab includes all 184 GEMs associated with chitin-induced ROS. The subsequent tabs include lists of shared GEMs associated with chitin-induced ROS response and each additional trait (quantitative disease resistance (QDR) to each fungal pathogen or additional PAMP-induced ROS responses). The title of each tab indicates the data included in each comparison and the number of shared GEMs. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible.&nbsp;</p> <p><strong>Table S5</strong> Enrichment analyses to determine if the number of gene expression markers (GEMs) shared between different lists is greater than the number of GEMs that would be expected by chance (e.g., lists of quantitative disease resistance (QDR) GEMs for two fungal pathogens). The representation factor is the number of overlapping GEMs divided by the expected number of overlapping GEMs drawn from two independent groups (traits), considering the total number of GEMs sequenced (53884). A representation factor &gt; 1 indicates more overlap than expected of two groups, a representation factor &lt; 1 indicates less overlap than expected, and a representation factor of 1 indicates that the two groups by the number of genes expected for independent groups of genes.&nbsp;</p> <p><strong>Table S6 R</strong>esults from Weighted Co-expression Gene Network Analysis (WGCNA). The first tab indicates significant modules from WGCNA analysis. Black and magenta modules are associated with antagonistic effects on resistance/susceptibility to all four pathogens. The second tab includes a full list of the GEM markers (Table S3), which are in significant WGCNA modules. The third, fourth and, fifth tabs indicate all significant GEMs in the black module, &nbsp;GO terms associated with GEMs in the black module, and all GO terms associated with the black module, respectively. &nbsp;The sixth, seventh and, eighth tabs indicate all significant GEMs in the magenta module, &nbsp;GO terms associated with GEMs in the magenta module, and all GO terms associated with the magenta module, respectively.</p> <p><strong>Table S7 </strong>Shared gene expression markers (GEMs) associated with resistance to different pathogens (<em>Alternaria brassicicola, Botrytis cinerea, Pyrenopeziza brassicae </em>and <em>Verticillium longisporum</em>). Lists correspond to matrices and Venn diagrams in Fig. 3. The first tab includes all GEMs associated quantitative disease resistance (QDR) to the fungal pathogens. The subsequent tabs include lists of shared GEMs associated with QDR to two or more fungal pathogens. The title of each tab indicates the data included in each comparison and the number of shared GEMs. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible.&nbsp;</p> <p><strong>Table S8 </strong>List of genes in linkage disequilibrium with the top marker for <em>Verticillium longisporum</em> resistance from genome-wide association (GWA) analysis on chromosome A09 (107 genes)(Tab 1) and the homoeologous region on C08 (Tab 2). Their percentage identity and query coverage in <em>Brassica napus</em> reference genotypes Quinta, Tapidor, Westar and Zhongshuang 11 compared to the <em>B. napus</em> pantranscriptome is indicated. Predicted <em>Arabidopsis thaliana</em> orthologs and corresponding descriptions are shown where possible.&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

A Terrylene Bisimide based Universal Host for Aromatic Guests to Derive Contact Surface-Dependent Dispersion Energies

<p>Additional data to report <a href="https://doi.org/10.1002/anie.202318451">https://doi.org/10.1002/anie.202318451</a>:<br><br>&pi;&ndash;&pi; interactions are among the most important intermolecular interactions in supramolecular systems. Here we determine experimentally a universal parameter for their strength that is simply based on the size of the interacting contact surfaces. Toward this goal we designed a new cyclophane based on terrylene bisimide (TBI) &pi;-walls connected by&nbsp;<em>para</em>-xylylene spacer units. With its extended &pi;-surface this cyclophane proved to be an excellent and universal host for the complexation of &pi;-conjugated guests, including small and large polycyclic aromatic hydrocarbons (PAHs) as well as dye molecules. The observed binding constants range up to 10<sup>8</sup> M<sup>&minus;1</sup>&nbsp;and show a linear dependence on the 2D area size of the guest molecules. This correlation can be used for the prediction of binding constants and for the design of new host&ndash;guest systems based on the herewith derived universal Gibbs interaction energy parameter of 0.31 kJ/mol&Aring;<sup>2</sup> in chloroform.</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Neither alpha-synuclein-preformed fibrils derived from patients with GBA1 mutations nor the host murine genotype significantly influence seeding efficacy in the mouse olfactory bulb

<p>Data sets for;</p> <p>Neither alpha-synuclein-preformed fibrils derived from patients with&nbsp;<em>GBA1</em> mutations nor the host murine genotype significantly influence seeding efficacy in the mouse olfactory bulb</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts".

<p>Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts". The data set consists of five folders: &ldquo;Codon_usage_amoebae_and_viruses&rdquo;, "Genome_annotations_amoebae", "Phylogenetic_trees_18S_amoebae", &ldquo;Phylogenomic_trees_amoebae&rdquo;, and "Viral_integration_detection_amoebae". The &ldquo;Codon_usage_amoebae_and_viruses&rdquo; folder contains for each amoeba host the calculated codon usage tables in the subfolder "codon_usage_table_host", the calculated codon usage preferences using different scores in the subfolder "codon_usage_scores_host", and the calculated codon usage preferences of giant viruses versus each host in the subfolder "codon_usage_scores_viruses_vs_host". The giant viruses in the subfolder "codon_usage_scores_viruses_vs_host" are organised by viral family and genus in separate sub-subfolders. The "Genome_annotations_amoebae" folder contains the generated genome annotations in different formats and the manually curated mitochondrial genome annotations for each amoeba host.&nbsp; The "Phylogenetic_trees_18S_amoebae" contains for the eukaryotic phyla <em>Discosea</em>, <em>Heterolobosea</em>, and <em>Tubulinea,&nbsp;</em>the 18S rRNA&nbsp;nucleotide alignments, distance matrices, and computed phylogenetic trees. The folder "Phylogenomic_trees_amoebae" contains for the eukaryotic clades <em>Amoebozoa</em> and <em>Discoba,&nbsp;</em>the protein alignment matrices and computed phylogenomic trees. The folder "Viral_integration_detection_amoebae" contains the MCP databases used (fasta file, alignment file, HMM profile and DIAMOND BLASTX database) and the MCP sequences detected in this study and the blast results of these.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Host-adaptation in Legionellales is 1.9 Gya, coincident with eukaryogenesis

<p>This dataset contains genomes, proteomes and protein alignments mentioned in Hugoson et al (2021). It has been used to analyze the evolution of host-adaptation in the order Legionellales.</p> <p>The data is organized by dataset type, and then by dataset.</p> <p>The four datasets used here are</p> <ul> <li><strong>Gamma105</strong>, comprising 105 <em>Gammaproteobacteria</em> and 5 outgroups;</li> <li><strong>Legio93</strong>, comprising 93 <em>Legionellales</em> and 20 outgroups;</li> <li><strong>Bacteria134</strong>, built on&nbsp;Gamma105, adding 27 genomes from Betts et al. (2018)</li> <li><strong>Bacteria93</strong>, built by removing <em>Legionella</em>, <em>Francisella</em>, <em>Fangia</em> and <em>Piscirickettsia</em> genera from Bacteria134</li> </ul> <p><strong>1_genomes</strong><br> Genomes as downloaded or assembled</p> <ul> <li>1_1_Gamma105</li> <li>1_2_Legio93</li> </ul> <p><strong>2_proteomes</strong><br> Proteomes, as annotated by prokka</p> <ul> <li>2_1_Gamma105</li> <li>2_2_Legio93</li> <li>2_3_Bacteria134</li> </ul> <p><strong>3_alignments</strong></p> <p>In the first three and the fifth folders, the following files are found. All sequence and alignment files are in fasta format:</p> <ul> <li>*_concatenated.fasta: concatenated alignment, trimmed.</li> <li>*.map: map of the files, tab-separated. The first row is a title row. The three first columns give the organism, the marker and the id (as found in the fasta file) for the protein.</li> <li>*_unaligned: non-aligned sequences for each marker.</li> <li>*_aligned: aligned sequences, for each marker. The prefix gives the software used for the alignment.</li> <li>*_trimmed: aligned, trimmed sequences for each marker. The prefix gives the software used to trim the alignment.</li> </ul> <p>&nbsp;</p> <ul> <li><strong>3_1_Gamma105</strong>: Based on the Bact109 set of markers.</li> <li><strong>3_2_Legio93</strong>: Based on the Bact109 set of markers.</li> <li><strong>3_3_Bacteria134</strong>: Based on Gamma105 set and Bact109 set of markers.</li> <li><strong>3_4_Bacteria93</strong>: Based on Bacteria134 (removed fast-evolving genomes).</li> <li><strong>3_5_TB4SS_auto</strong>: Alignment of 12 genes of the T4BSS, automatically detected in all genomes.&nbsp;</li> <li><strong>3_6_TB4SS_manual</strong>: Alignment of 25 genes of the T4BSS, manually curated by collinearity analysis.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data for: Prior exposure of a fungal parasite to cyanobacterial extracts does not impair infection of its Daphnia host

<p>This dataset supports the findings of the study 'Prior exposure of a fungal parasite to cyanobacterial extracts does not impair infection of its <em>Daphnia</em> host', published in Hydrobiologia (https://doi.org/10.1007/s10750-022-04889-7)</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Diversity in the Expressed Genomic Host Response to Myocardial Infarction - Validation Dataset

<p>External validation was performed by separately hierarchically clustering 934 patients with STEMI in an independent cohort[1]&nbsp;into 2 groups (232 and 702 individuals) based on Illumina HT12v4-profiled PBMC expression (median time 21 hour between cardiac catheterization and blood sampling). Probes with most variable expression intensities (SD&ge;0.5, 216 probes, excluding ribosomal genes) were used. From the 20 most differentially expressed genes in the discovery cohort described in the manuscript Toma et al. 2022 [2], 19 were available in the validation cohort.</p> <p>Column names include the Illumina identifyer and the mapped gene name as used in the discovery cohort. Values are log2-transformed, quantile-normalized, batch-corrected values, see also [1] for methodological details.</p> <p>Acknowledgement:</p> <p>This work&nbsp;is supported by LIFE &ndash; Leipzig Research Center for Civilization Diseases, Universit&auml;t Leipzig. LIFE is funded by means of the European Union, by the European Regional Development Fund (ERDF) and by means of the Free State of Saxony within the framework of the excellence initiative.</p> <p>&nbsp;</p> <p>1)&nbsp;Teren A, Kirsten H, Beutner F, Scholz M, Holdt LM, Teupser D, Gutberlet M, Thiery J, Schuler G, Eitel I. Alteration of multiple leukocyte gene expression networks is linked with magnetic resonance markers of prognosis after acute st-elevation myocardial infarction. <em>Scientific Reports</em>. 2017;7:41705</p> <p>2) Toma A, dos Santos&nbsp;C, Burzyńska&nbsp;B, G&oacute;ra&nbsp;M, Kiliszek&nbsp;M, Stickle N, Kirsten H, Kosyakovsky L, Wang B, van Diepen S, Epelman S, Szekely Y, Marshall JC, Billia F, Lawler PR (2022), Diversity in the Expressed Genomic Host Response to Myocardial Infarction, submitted.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Controlled Formation of Dimers and Spatially Isolated Atoms in Bimetallic Au-Ru Catalysts via Carbon-Host Functionalization

<p>Enclosed we report the data in the article:&nbsp;&quot;Controlled Formation of Dimers and Spatially Isolated Atoms in Bimetallic Au-Ru Catalysts via Carbon-Host Functionalization&quot; by&nbsp;P&eacute;rez-Ram&iacute;rez et al.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

PhasAGE Training School 2 - Phase separations and transitions by viral proteins: from viral factories to interference with host cell functions- LECTURE

<p>The Training School 2 &ldquo;Biomolecular condensates in cell function, aging and disease&rdquo; is the <strong>second</strong> edition of a series of PhasAGE training activities.</p> <p>&nbsp;</p> <p>The main goal of this training school is to raise awareness and provide expertise on fundamental aspects of phase separation and formation of <strong>biomolecular condensates</strong>, specifically covering the importance of this process to cellular biology and its contribution to the aging process and age-related diseases.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Data for: Polystyrene nanoplastics differentially influence the outcome of infection by two microparasites of the host Daphnia magna

<p>This dataset supports the findings of the study 'Polystyrene nanoplastics differentially influence the outcome of infection by two microparasites of the host <em>Daphnia magna</em>', published in Philosophical Transactions of the Royal Society B (https://doi.org/10.1098/rstb.2022.0013).</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Identifying and profiling structural similarities between Spike of SARS-CoV-2 and other viral or host proteins with Machaon - Pre-computed features for replication

<p>Machaon&#39;s computed features that were used in the structural comparisons with Spike protein.</p> <p>DATA_PDBS_vir_whole_1-3.zip files are parts of a single folder.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria

<p>Supplementary dataset from &quot;<strong><em>Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria</em></strong>&quot;</p> <p>&nbsp;</p> <p><strong>Supplementary Figure legends</strong></p> <p><strong>Figure S1. Illustration of the transposon insertions around the <em>M. bovis </em>genome. </strong>Sequencing of the input library showed that transposon insertions were evenly distributed around the genome and 27,419 of the permissible 66,931 thymine&ndash;adenine dinucleotide (TA) sites contained an insertion representing an insertion density of ~41%. The outer ring are the genomic coordinates, the blue lines represent transposon insertions and the gray boxes indicate regions of that did not have any insertions. Plot made with Circlize (Gu et al, 2014).</p> <p>&nbsp;</p> <p><strong>Figure S2. Diversity of the output library isolated from lung and thoracic lymph node lesions compared to the input library. </strong>On average, libraries recovered from lung lesions contained 14,456 unique mutants and those recovered from the lymph nodes contained an average of 16,210 unique mutants. Insertion density is represented as a proportion of the TA sites that contained insertions. The numbers on the x-axis refer to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p>&nbsp;</p> <p><strong>Figure S3. Volcano plots showing the distribution of log<sub>2</sub> fold-changes and -log<sub>10</sub> of adjusted p-values for representative lung (A) and lymph node (B) samples. </strong>Adjusted p-values (BH-fdr correction) &lt; 0.000001 cluster at the limits of the plot and precision reflects the number of resampling iterations (10,000).</p> <p>&nbsp;</p> <p><strong>Figure S4. Scatterplot of mean log<sub>2</sub> fold change per gene for all lung samples against all thoracic lymph node samples</strong>. Correlation between mean log<sub>2</sub> fold change among genes between the tissues was calculated with Spearman&#39;s ranked correlation, = 0.878, p-value &lt; 2.2e-16.</p> <p>&nbsp;</p> <p><strong>Figure S5. Fold-changes caused by transposon insertions in <em>RD1<sup>BCG</sup></em> and <em>RD1<sup>MIC</sup> </em>in the lungs and lymph nodes of infected cattle. </strong>Boxplot for log<sub>2 </sub>fold-changes in genes of the RD1<sup>BCG</sup> region. Samples with adjusted p-values (BH-fdr corrected) &lt;0.05 are indicated with purple points. Gene names highlighted in magenta have fewer than 5 TA sites located in the gene; too few to determine the statistical significance of changes in insertion levels with this method.</p> <p>&nbsp;</p> <p><strong>Supplementary Tables </strong></p> <p><strong>Table S1. Sequencing statistics of the input and output transposon libraries. </strong>The numbers in the column labelled &ldquo;filename&rdquo; refers to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p>&nbsp;</p> <p><strong>Table S2. Tissues collected and scored for gross pathology. </strong>Tissues from head and neck lymph nodes (from the right and left sub-mandibular lymph nodes, the right and left medial retropharyngeal lymph nodes), thoracic lymph nodes (the right and left bronchial lymph nodes, the cranial tracheobronchial lymph nodes, the cranial and caudal mediastinal lymph nodes) and from lung lesions, were collected and scored.</p> <p>&nbsp;</p> <p><strong>Table S3. Log<sub>2</sub> fold-changes for insertions across the entire genome of <em>M. bovis</em> AF2122/97. </strong>Cells are coloured according to log<sub>2</sub> fold-change. Refer to the text for the gene groups in individual tabs.</p> <p>&nbsp;</p> <p><strong>Table S4. </strong>Custom transposon sequencing primers and adaptors used in sequencing of the transposon libraries.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record