Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
133
datasets available to search
ShareScore release 0.9.0
Dataset results
133 results for “backbone”
Datasets for Ultra High-Capacity Band and Space Division Multiplexing Backbone EONs
<p>The datasets have been generated for the paper titled "Ultra High-Capacity Band and Space Division Multiplexing Backbone EONs: Multi-core vs. Multi-fiber." </p>
Dataset for simulation of a low-carbon urban energy system using the Backbone model
<p>The dataset contains the input data for cost optimization of an urban energy system. The case study has been described in the article "Impact of power-to-gas on the cost and design of the future low-carbon urban energy system" of Applied Energy.</p> <p>The dataset is in Microsoft Excel format. To make it available for GAMS, one should use e.g. the attached shell script (requires GAMS installation) to convert it to *.gdx file. The generation expansion model is available in the Git repository https://gitlab.vtt.fi/backbone/backbone (under branch projik/planet).</p>
A Repackaged Taxonomic Backbone of Global Biodiversity Information Facility (GBIF)
<p>Publication date:<br> 2022-12-06T07:37:19-06:00</p> <p><br> A Repackaged Taxonomic Backbone of Global Biodiversity Information Facility (GBIF)<br> ---</p> <p>Global Biodiversity Information Facility (GBIF) facilitates access to billions of biodiversity data records. These records include detailed accounts of life on earth.</p> <p>To help records of specific life forms, GBIF provides a taxonomic backbone [1,2]. This backbone contains a long list of names used to describe species and associated hierarchies and taxonomic publications. These lists are sourced from datasets around the world.</p> <p>At time of writing (6 Dec 2022), GBIF publishes a simplified version of their taxonomic backbone at [https://hosted-datasets.gbif.org/datasets/backbone/](https://hosted-datasets.gbif.org/datasets/backbone/) [1].</p> <p>This repository provides script to pre-process https://hosted-datasets.gbif.org/datasets/backbone/current/simple.txt.gz to help facilitate access and improve performance of the creation of search indexes.</p> <p>Pre-process steps currently include:<br> 1. reducing amount of columns<br> 2. reverse sort by id<br> 3. reverse sort by name</p> <p><br> Contents<br> ---</p> <p>README:<br> this file</p> <p>repackage-gbif-backbone.sh:<br> script used to repackage GBIF Simple Backbone.</p> <p>repackage-gbif-backbone.log:<br> log of repackaging of GBIF Simple Backbone.</p> <p>backbone-current-simple.txt.gz:<br> original GBIF backbone archive</p> <p>gbif-backbone-by-name.tsv.gz:<br> two columns, gzipped, tab-separated text file with columns name, and id<br> reverse sorted by name </p> <p>gbif-backbone-by-name.tsv.sha256:<br> sha256 hash of the uncompressed gbif-backbone-by-name.tsv.gz</p> <p>gbif-backbone-by-id.tsv.gz:<br> 20 columns, gzipped, tab-separated text file with first 20 columns of repackaged GBIF backbone file<br> reverse sorted by id</p> <p>gbif-backbone-by-id.tsv.sha256:<br> sha256 hash of the uncompressed gbif-backbone-by-id.tsv.gz</p> <p>References<br> ---</p> <p>[1] Simplied GBIF Backbone Taxonomy. Accessed at https://hosted-datasets.gbif.org/datasets/backbone/ on 2022-12-06.<br> [2] GBIF Secretariat (2021). GBIF Backbone Taxonomy. Checklist dataset https://doi.org/10.15468/39omei accessed via GBIF.org on 2021-08-18.</p> <p><br> Hash URIs<br> ---<br> This publication includes the following content uris:</p> <p>hash://sha256/82d5f2153b4533322692d95eeb18b0f103e1b2297e38bd9ea935b07ba86cd7d5<br> hash://sha256/50c155f66efb2efba0b8b624f8541e81cbe16a701d420a5073791fb993f72919<br> hash://sha256/9cd7d4c91292d86c726210446cd6fe45602505a7c0ea3b7c4f4f481f85f193ad (uncompressed)<br> hash://sha256/f950dde25cce9ba9cce67caa1c68ce0c99cb31fe2dc9658fec85a987d9f31654<br> hash://sha256/f21c6b90f17c6083fcfb4853f3c581dcc2aadd291691fa128392a205321f420b (uncompressed)<br> hash://sha256/5e0a4d1d2d1cccbdcc6b2c9831fafe61c54eb055f2d13ec40d9ac161889b9f89<br> hash://sha256/f6e477133d0585706ee5522963b204200cb3cd198f011cbf62be0fa8519763b5 (uncompressed)<br> </p>
The World Asellidae database and phylogeny: a collaborative backbone resource for comparative studies of subterranean life evolution
<p>Supplementary material for the article "The World Asellidae database and phylogeny: a collaborative backbone resource for comparative studies of subterranean life evolution"</p> <p>- SI Figure 5: The World Asellidae phylogeny with credibility Intervals for the age of the nodes. Node labels of the phylogeny indicate the 95% credibility intervals of the estimated dates.</p> <p>- SI Table 1: Metadata for the 2093 COI sequences used in the study.</p> <p>- SI Table 4: Alignment of the 2093 COI sequences used for the delimitation of MOTUs.</p> <p>- SI Table 5: Alignment of the 424 COI sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 6: Alignment of the 424 16S sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 7: Alignment of the 424 FASTKD4 sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 8: Alignment of the 424 28S sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 9: Metadata for the DNA sequences used for the 4-gene dated phylogeny.</p> <p>- SI Table 11: Data on body size, sexual body size dimorphism, habitat specialization and habitat size used in comparative analyses.</p> <p>- SI Table 12: Metadata for the DNA sequences deposited in NCBI as part of this study.</p>
GBIF Backbone matches - changes
<p>All GBIF occurrence records with changed species matches due to an improved matching service. See http://gbif.blogspot.com/2015/03/improving-gbif-backbone-matching.html for details</p>
Supplementary materials for "Improving diffusion-based protein backbone generation with global-geometry-aware latent encoding"
<h1>Info</h1> <p>This dataset contains the supplementary materials for "Improving diffusion-based protein backbone generation with global-geometry-aware latent encoding". </p> <p>For <strong>source code </strong>and <strong>detailed instructions on usage, </strong>please refer to our <a href="https://github.com/meneshail/TopoDiff/tree/main" target="_blank" rel="noopener">github</a> .</p> <h1>Supplementary data</h1> <h2>weights.tar.gz</h2> <p>The trained model weights used in the paper.</p> <h2>dataset.zip</h2> <p>CATH-60 Dataset used in the paper. In the notebook directory of our <a href="https://github.com/meneshail/TopoDiff/tree/main" target="_blank" rel="noopener">github</a> , we provide an example on encoding and visualize it with our trained encoder.</p> <h2>design.zip</h2> <p>The 21 novel mainly-beta designs selected for experiment validation. Along with the generated backbone, we also provide the prediction results from AlphaFold and ESMFold.</p> <h2>benchmark_sample.zip</h2> <p>Sampled backbones used for all benchmark experiment (All methods and variants included).</p> <h2>evaluation.tar.gz</h2> <p>Precomputed CATH reference data for coverage metric computation. Need to be downloaded for using evaluation scripts. </p>
Text-fig. 9. Phylogenetic tree indicating the number of required character state changes (steps) under parsimony for various positions of Miranthus gen. nov. in a molecular based backbone tree (see material and methods for additional details). in Early Flowers Of Primuloid Ericales From The Late Cretaceous Of Portugal And Their Ecological And Phytogeographic Implications
Text-fig. 9. Phylogenetic tree indicating the number of required character state changes (steps) under parsimony for various positions of Miranthus gen. nov. in a molecular based backbone tree (see material and methods for additional details).
CESNET-QUIC22: A large one-month QUIC network traffic dataset from backbone lines
<p><strong>Please refer to the original data article for further data description: </strong>Jan Luxemburk et al. CESNET-QUIC22: A large one-month QUIC network traffic dataset from backbone lines, Data in Brief, 2023, 108888, ISSN 2352-3409, <a href="https://doi.org/10.1016/j.dib.2023.108888">https://doi.org/10.1016/j.dib.2023.108888</a>. </p><p><strong>We recommend using the</strong> <strong>CESNET DataZoo python library, which facilitates the work with large network traffic datasets. </strong>More information about the DataZoo project can be found in the GitHub repository <a href="https://github.com/CESNET/cesnet-datazoo">https://github.com/CESNET/cesnet-datazoo</a>.</p><p>The QUIC (Quick UDP Internet Connection) protocol has the potential to replace TLS over TCP, which is the standard choice for reliable and secure Internet communication. Due to its design that makes the inspection of QUIC handshakes challenging and its usage in HTTP/3, there is an increasing demand for research in QUIC traffic analysis. This dataset contains one month of QUIC traffic collected in an ISP backbone network, which connects 500 large institutions and serves around half a million people. The data are delivered as enriched flows that can be useful for various network monitoring tasks. The provided server names and packet-level information allow research in the encrypted traffic classification area. Moreover, included QUIC versions and user agents (smartphone, web browser, and operating system identifiers) provide information for large-scale QUIC deployment studies.</p><p><strong>Data capture</strong> The data was captured in the flow monitoring infrastructure of the <a href="https://www.cesnet.cz">CESNET2</a> network. The capturing was done for four weeks between 31.10.2022 and 27.11.2022. The following list provides per-week flow count, capture period, and uncompressed size:</p><ul><li><strong>W-2022-44</strong><ul><li>Uncompressed Size: 19 GB</li><li>Capture Period: 31.10.2022 - 6.11.2022</li><li>Number of flows: 32.6M</li></ul></li><li><strong>W-2022-45</strong><ul><li>Uncompressed Size: 25 GB</li><li>Capture Period: 7.11.2022 - 13.11.2022</li><li>Number of flows: 42.6M</li></ul></li><li><strong>W-2022-46</strong><ul><li>Uncompressed Size: 20 GB</li><li>Capture Period: 14.11.2022 - 20.11.2022</li><li>Number of flows: 33.7M</li></ul></li><li><strong>W-2022-47</strong><ul><li>Uncompressed Size: 25 GB</li><li>Capture Period: 21.11.2022 - 27.11.2022</li><li>Number of flows: 44.1M</li></ul></li><li><strong>CESNET-QUIC22 </strong><ul><li>Uncompressed Size: 89 GB</li><li>Capture Period: 31.10.2022 - 27.11.2022</li><li>Number of flows: 153M</li></ul></li></ul><p> </p><p><strong>Data description</strong> The dataset consists of network flows describing encrypted QUIC communications. Flows were created using <a href="https://github.com/CESNET/ipfixprobe">ipfixprobe</a> flow exporter and are extended with packet metadata sequences, packet histograms, and with fields extracted from the QUIC Initial Packet, which is the first packet of the QUIC connection handshake. The extracted handshake fields are the Server Name Indication (SNI) domain, the used version of the QUIC protocol, and the user agent string that is available in a subset of QUIC communications.</p><p><strong>Packet Sequences</strong> Flows in the dataset are extended with sequences of packet sizes, directions, and inter-packet times. For the packet sizes, we consider payload size after transport headers (UDP headers for the QUIC case). Packet directions are encoded as ±1, <i>+1</i> meaning a packet sent from client to server, and <i>-1</i> a packet from server to client. Inter-packet times depend on the location of communicating hosts, their distance, and on the network conditions on the path. However, it is still possible to extract relevant information that correlates with user interactions and, for example, with the time required for an API/server/database to process the received data and generate the response to be sent in the next packet. Packet metadata sequences have a length of 30, which is the default setting of the used flow exporter. We also derive three fields from each packet sequence: its length, time duration, and the number of roundtrips. The roundtrips are counted as the number of changes in the communication direction (from packet directions data); in other words, each client request and server response pair counts as one roundtrip.</p><p><strong>Flow statistics</strong> Flows also include standard flow statistics, which represent aggregated information about the entire bidirectional flow. The fields are: the number of transmitted bytes and packets in both directions, the duration of flow, and packet histograms. Packet histograms include binned counts of packet sizes and inter-packet times of the entire flow in both directions (more information in the <a href="https://github.com/CESNET/ipfixprobe/tree/master#phists">PHISTS plugin documentation</a> There are eight bins with a logarithmic scale; the intervals are 0-15, 16-31, 32-63, 64-127, 128-255, 256-511, 512-1024, >1024 [ms or B]. The units are milliseconds for inter-packet times and bytes for packet sizes. Moreover, each flow has its end reason - either it was idle, reached the active timeout, or ended due to other reasons. This corresponds with the official <a href="https://www.iana.org/assignments/ipfix/ipfix.xhtml#ipfix-flow-end-reason">IANA IPFIX-specified values</a>. The <i>FLOW_ENDREASON_OTHER</i> field represents the <i>forced end</i> and <i>lack of resources</i> reasons. The <i>end of flow detected</i> reason is not considered because it is not relevant for UDP connections.</p><p><strong>Dataset structure</strong> The dataset flows are delivered in compressed CSV files. CSV files contain one flow per row; data columns are summarized in the provided list below. For each flow data file, there is a JSON file with the number of saved and seen (before sampling) flows per service and total counts of all received (observed on the CESNET2 network), service (belonging to one of the dataset's services), and saved (provided in the dataset) flows. There is also the <i>stats-week.json</i> file aggregating flow counts of a whole week and the <i>stats-dataset.json</i> file aggregating flow counts for the entire dataset. Flow counts before sampling can be used to compute sampling ratios of individual services and to resample the dataset back to the original service distribution. Moreover, various dataset statistics, such as feature distributions and value counts of QUIC versions and user agents, are provided in the <i>dataset-statistics</i> folder. The mapping between services and service providers is provided in the <i>servicemap.csv</i> file, which also includes SNI domains used for ground truth labeling. The following list describes flow data fields in CSV files:</p><ul><li><strong>ID:</strong> Unique identifier</li><li><strong>SRC_IP:</strong> Source IP address</li><li><strong>DST_IP:</strong> Destination IP address</li><li><strong>DST_ASN:</strong> Destination Autonomous System number</li><li><strong>SRC_PORT:</strong> Source port</li><li><strong>DST_PORT:</strong> Destination port</li><li><strong>PROTOCOL:</strong> Transport protocol</li><li><strong>QUIC_VERSION QUIC:</strong> protocol version</li><li><strong>QUIC_SNI:</strong> Server Name Indication domain</li><li><strong>QUIC_USER_AGENT:</strong> User agent string, if available in the QUIC Initial Packet</li><li><strong>TIME_FIRST:</strong> Timestamp of the first packet in format YYYY-MM-DDTHH-MM-SS.ffffff</li><li><strong>TIME_LAST:</strong> Timestamp of the last packet in format YYYY-MM-DDTHH-MM-SS.ffffff</li><li><strong>DURATION:</strong> Duration of the flow in seconds</li><li><strong>BYTES:</strong> Number of transmitted bytes from client to server</li><li><strong>BYTES_REV:</strong> Number of transmitted bytes from server to client</li><li><strong>PACKETS:</strong> Number of packets transmitted from client to server</li><li><strong>PACKETS_REV:</strong> Number of packets transmitted from server to client</li><li><strong>PPI:</strong> Packet metadata sequence in the format: [[inter-packet times], [packet directions], [packet sizes]]</li><li><strong>PPI_LEN:</strong> Number of packets in the PPI sequence</li><li><strong>PPI_DURATION:</strong> Duration of the PPI sequence in seconds</li><li><strong>PPI_ROUNDTRIPS:</strong> Number of roundtrips in the PPI sequence</li><li><strong>PHIST_SRC_SIZES:</strong> Histogram of packet sizes from client to server</li><li><strong>PHIST_DST_SIZES: </strong>Histogram of packet sizes from server to client</li><li><strong>PHIST_SRC_IPT: </strong>Histogram of inter-packet times from client to server</li><li><strong>PHIST_DST_IPT:</strong> Histogram of inter-packet times from server to client</li><li><strong>APP:</strong> Web service label</li><li><strong>CATEGORY:</strong> Service category</li><li><strong>FLOW_ENDREASON_IDLE:</strong> Flow was terminated because it was idle</li><li><strong>FLOW_ENDREASON_ACTIVE:</strong> Flow was terminated because it reached the active timeout</li><li><strong>FLOW_ENDREASON_OTHER:</strong> Flow was terminated for other reasons</li></ul><p> </p><p><strong>Link to other CESNET datasets</strong></p><ul><li><a href="https://www.liberouter.org/technology-v2/tools-services-datasets/datasets/">https://www.liberouter.org/technology-v2/tools-services-datasets/datasets/</a></li><li><a href="https://github.com/CESNET/cesnet-datazoo">https://github.com/CESNET/cesnet-datazoo</a></li></ul><p><strong>Please cite the original data article:</strong></p><blockquote><p>@article{CESNETQUIC22, author = {Jan Luxemburk and Karel Hynek and Tomáš Čejka and Andrej Lukačovič and Pavel Šiška}, title = {CESNET-QUIC22: a large one-month QUIC network traffic dataset from backbone lines}, journal = {Data in Brief}, pages = {108888}, year = {2023}, issn = {2352-3409}, doi = {https://doi.org/10.1016/j.dib.2023.108888}, url = {https://www.sciencedirect.com/science/article/pii/S2352340923000069} }</p></blockquote>
A Repackaged Taxonomic Backbone of Global Biodiversity Information Facility (GBIF) - 2021-11-26
<p>A Repackaged Taxonomic Backbone of Global Biodiversity Information Facility (GBIF)<br> ---</p> <p>Global Biodiversity Information Facility (GBIF) facilitates access to billions of biodiversity data records. These records include detailed accounts of life on earth.</p> <p>To help records of specific life forms, GBIF provides a taxonomic backbone [1,2]. This backbone contains a long list of names used to describe species and associated hierarchies and taxonomic publications. These lists are sourced from datasets around the world.</p> <p>At time of writing (18 Aug 2021), GBIF publishes a simplified version of their taxonomic backbone at [https://hosted-datasets.gbif.org/datasets/backbone/](https://hosted-datasets.gbif.org/datasets/backbone/) [1].</p> <p>This repository provides script to pre-process https://hosted-datasets.gbif.org/datasets/backbone/backbone-current-simple.txt.gz to help facilitate access and improve performance of the creation of search indexes.</p> <p>Pre-process steps currently include:</p> <p>1. reducing amount of columns<br> 2. reverse sort by id<br> 3. reverse sort by name</p> <p><br> Contents<br> ---</p> <p>README:<br> this file</p> <p>repackage-gbif-backbone.sh:<br> script used to repackage GBIF Simple Backbone.</p> <p>backbone-current-simple.txt.gz:<br> original GBIF backbone archive</p> <p>gbif-backbone-by-name.tsv.gz:<br> two columns, gzipped, tab-separated text file with columns name, and id<br> reverse sorted by name</p> <p>gbif-backbone-by-name.tsv.sha256:<br> sha256 hash of the uncompressed gbif-backbone-by-name.tsv.gz</p> <p>gbif-backbone-by-id.tsv.gz:<br> 20 columns, gzipped, tab-separated text file with first 20 columns of repackaged GBIF backbone file<br> reverse sorted by id</p> <p>gbif-backbone-by-id.tsv.sha256:<br> sha256 hash of the uncompressed gbif-backbone-by-id.tsv.gz</p> <p>References<br> ---</p> <p>[1] Simplied GBIF Backbone Taxonomy. Accessed at https://hosted-datasets.gbif.org/datasets/backbone/ on 2021-08-18.<br> [2] GBIF Secretariat (2021). GBIF Backbone Taxonomy. Checklist dataset https://doi.org/10.15468/39omei accessed via GBIF.org on 2021-08-18.</p> <p><br> Hash URIs<br> ---<br> This publication includes the following content uris:</p> <p>repackage-gbif-backbone.sh:<br> hash://sha256/073ac5490252c4ccbbd4f516d391faebe62c9fde9e4d75ae870441a86c382527</p> <p>backbone-current-simple.txt.gz:<br> hash://sha256/15cbfc038e666356af27248935f79e408ed51fd8c0b49a668fed8dbf72591502<br> hash://sha256/1f78788a4a046dcbcf1e36c7658a1e333ca60e7586a372238d58b938d91fde51 (uncompressed)</p> <p>gbif-backbone-by-name.tsv.gz:<br> hash://sha256/6e11ae9961a9498b60d4bdeb489d6c1f5da9c2732310edaecdc79bd287b79ef4<br> hash://sha256/934ce05dbd067abb209168bd1d9389f122d051e1b7374b5d757a12e86f8da9a5 (uncompressed)</p> <p>gbif-backbone-by-id.tsv.gz:<br> hash://sha256/c434c7d3622421b17dadcd119391b32a66edee59f484d4cab924d92fd17713e2<br> hash://sha256/e2cf9116a21966315b0482d391052223e21c8e916ae0c097dfd37bed017b815b (uncompressed)</p>
Research data supporting "Post-polymerisation functionalisation of conjugated polymer backbones and its application in multi-functional emissive nanoparticles"
<p>Research data supporting the publication:</p> <p>Creamer A. et al., "Post-polymerisation functionalisation of conjugated polymer backbones and its application in multi-functional emissive nanoparticles",<em> Nature Communications</em><strong>, 9</strong>:3237 (2018).</p>
Text-fig. 6. Most parsimonious tree obtained after addition of Acaciaephyllum to the data set of Doyle (2008), with modifications discussed in the text, and with relationships of other taxa fixed with a backbone constraint tree based on results of Doyle (2008). Relative parsimony of alternative positions of Acaciaephyllum is indicated as in Text-fig. 2. Gnet = Gnetales. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 6. Most parsimonious tree obtained after addition of Acaciaephyllum to the data set of Doyle (2008), with modifications discussed in the text, and with relationships of other taxa fixed with a backbone constraint tree based on results of Doyle (2008). Relative parsimony of alternative positions of Acaciaephyllum is indicated as in Text-fig. 2. Gnet = Gnetales.
Text-fig. 7. Number of required character state changes under parsimony (steps) for various positions of Mugideiriflora portugallica, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014). in Multiparted, Apocarpous Flowers From The Early Cretaceous Of Eastern North America And Portugal
Text-fig. 7. Number of required character state changes under parsimony (steps) for various positions of Mugideiriflora portugallica, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014).
Text-fig. 8. Number of required character state changes under parsimony (steps) for various positions of Lambertiflora elegans, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014). in Multiparted, Apocarpous Flowers From The Early Cretaceous Of Eastern North America And Portugal
Text-fig. 8. Number of required character state changes under parsimony (steps) for various positions of Lambertiflora elegans, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014).
Text-fig. 9. Number of required character state changes under parsimony (steps) for various positions of Atlantocarpus virginiensis, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014). in Multiparted, Apocarpous Flowers From The Early Cretaceous Of Eastern North America And Portugal
Text-fig. 9. Number of required character state changes under parsimony (steps) for various positions of Atlantocarpus virginiensis, based on the Doyle and Endress character matrix and backbone tree (Doyle and Endress 2000, 2014).
Five-week DoH Dataset collected on ISP backbone lines
<p>This dataset is an additional DoH dataset used for researching DoH traffic data drift phenomena. It contains anonymized packet captures (pcaps) from the following days:</p> <ul> <li>2022-11-28</li> <li>2022-12-05</li> <li>2022-12-12</li> <li>2022-12-19</li> <li>2022-12-26</li> </ul> <p>The traffic was captured on the CESNET2 network and anonymized. The packet capturing and anonymization follow the methodology described in [1]. The list of IP addresses used for DoH recognition is also included within the dataset in <code>doh_resolver_ip.csv</code> file.<br> The structure of the dataset is as follows:</p> <pre><code>. ├── doh_resolver_ip.csv ├── pcap │ ├── 2022-11-28 │ │ ├── DoH-20221128180002.pcapng │ │ └── HTTPS-20221128180002.pcapng │ ├── 2022-12-05 │ │ ├── DoH-20221205180001.pcapng │ │ └── HTTPS-20221205180001.pcapng │ ├── 2022-12-12 │ │ ├── DoH-20221212180001.pcapng │ │ └── HTTPS-20221212180001.pcapng │ ├── 2022-12-19 │ │ ├── DoH-20221219180001.pcapng │ │ └── HTTPS-20221219180001.pcapng │ └── 2022-12-26 │ ├── DoH-20221226180001.pcapng │ └── HTTPS-20221226180001.pcapng └── README.md</code></pre> <p> </p> <p>[1] Jeřábek, K., Hynek, K., Čejka, T., & Ryšavý, O. (2022). Collection of datasets with DNS over HTTPS traffic. Data in Brief, 42, 108310. <a href="https://www.sciencedirect.com/science/article/pii/S2352340922005121">https://www.sciencedirect.com/science/article/pii/S2352340922005121</a></p>
rbLEC - restricted backbone Local Euler Characteristic - from CATH database
<p>-----------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Author: Rodrigo A. Moreira (C) 2023<br> https://orcid.org/0000-0002-7605-8722<br> LICENSE: CC BY-NC-ND 4.0 (https://creativecommons.org/licenses/by-nc-nd/4.0/)</strong></p> <p>----------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>rbLEC - Local Euler Charactersitics - from CATH database</strong></p> <p>----------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>A. rbLEC NETWORK</strong></p> <p> [I] The networks for each PDB[1] structure is defined by the PDB atoms N,CA,C of each residue as nodes of a graph G.<br> [II] An edge of G is set if the distance between two atom in [I] is greater than 2.0 Angstrons.<br> [III] The graph G is defined in the files with extensions ".network_backboneRE_heavy_gt2"</p> <p>Equation (1) [6,7]<br> \begin{equation}<br> \chi = \sum_{k=1}^{N} \kappa_k = \sum_{k=1}^{N} \underbrace{ \left(1 + \sum_{l=1}^{\infty} (-1)^{l} \frac{v_{l-1}}{l+1} \right)_{k}}_{\kappa_k}<br> \end{equation}</p> <p>Equation (2)<br> \begin{equation}<br> LEC = \sum_{m \in R} \kappa_m = \kappa_{N} + \kappa_{CA} + \kappa_{C}<br> \end{equation}</p> <p><strong>B. FILENAME EXTENSIONS</strong></p> <p> <strong> B.1</strong> Basic files</p> <p>".fixed"<br> PDB file after use of pdbfixer[2] in structures from CATH database.</p> <p>".dssp"<br> Output of DSSP[3] software</p> <p>".stride"<br> Output of STRIDE[4] software</p> <p> <strong>B.2</strong> Data files</p> <p>".network_backboneRE_heavy_gt2" - Generate by <strong>D.2</strong> below.<br> Describe the network graph, as described in <strong>A.</strong> above.</p> <p>".knill_curvature" - Generate by <strong>D.1</strong> below.<br> Contain the filtration of kappas for each vertice of the network.</p> <p>".residues_curvature" - Generate by <strong>D.1</strong> below.<br> They are the filtration of LEC, Equation (2) above, for each residue, namely summation of 3 kappas from respective '.knill_curvature', correspoings to PDB atoms N,CA and C, describe in <strong>A.</strong> above.</p> <p>".label" - Generated by <strong>D.3 </strong>below<br> Extra file for easier assesment of structures. They have the same information about LEC as described in respective ".residue_curvature" file extensions, but merge also the information from ".dssp" and ".stride" classes as well as residue name and residue ID for each molecule.<br> Format of columns:<br> cutoff resname resid DSSP_class STRIDE_class LEC</p> <p><strong>C. FOLDERS</strong></p> <p> CATH_FIXED (after uncompress cath_fixed.tar.xz, approximately 13GB)<br> contains the fixed PDBs and LECs from CATH[5] database</p> <p><strong>D. SOFTWARE</strong><br> <strong>D.1</strong> lec.py: compute the kappas in Equation (1) above.<br> Example usage:<br> $ python3 lec.py CATH_FIXED/2x0qA02/2x0qA02<br> It will create the files with extension ".kappas" and ".relec", which reproduces the respectively the files with extension "<strong>.knill_curvature</strong>" and "<strong>.residue_curvature</strong>".</p> <p> <strong> D.2</strong> pdb2network.lua: creates rbLEC network file (number of nodes and edges list) from PDB to be used as input by lec.py.<br> Example usage:<br> $ lua pdb2rbLEC.lua CATH_FIXED/2x0qA02/2x0qA02.fixed<br> Output reproduces the file CATH_FIXED/2x0qA02/2x0qA02.pdb.<strong>network_backboneRE_heavy_gt2</strong></p> <p> <strong>D.3</strong> label.lua: create files with extension '*.label' from files '*.pdb.stride', '*.pdb.dssp' and '*.pdb.network_backboneRE_heavy_gt2.residues_curvature.<br> Example usage:<br> $ lua label.lua CATH_FIXED/2x0qA02/2x0qA02.pdb<br> Output reproduces the file CATH_FIXED/2x0qA02/2x0qA02.pdb.<strong>network_backboneRE_heavy_gt2.residues_curvature.label</strong></p> <p><strong>REFERENCES</strong><br> [1] Herman, H., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T., Weissig, H., Shindyalov, I., & Bourne, P. (2000). The protein data bank. Nucleic acids research, 28, 235–42.<br> [2] Eastman, P., Swails, J., Chodera, J., McGibbon, R., Zhao, Y., Beauchamp, K., Wang, L.P., Simmonett, A., Harrigan, M., Stern, C., & others (2017). OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLoS computational biology, 13(7), e1005659.<br> [3] Kabsch, W., & Sander, C. (1983). Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers: Original Research on Biomolecules, 22(12), 2577–2637.<br> [4] Frishman, D., & Argos, P. (1995). Knowledge-based protein secondary structure assignment. Proteins: Structure, Function, and Bioinformatics, 23(4), 566–579.<br> [5] Knudsen, M., & Wiuf, C. (2010). The CATH database. Human genomics, 4(3), 1–6.<br> [6] Levitt, N. (1992). The Euler characteristic is the unique locally determined numerical homotopy invariant of finite complexes. Discrete & computational geometry, 7, 59–67.<br> [7] Knill, O. (2011). A graph theoretical Gauss-Bonnet-Chern theorem. arXiv preprint arXiv:1111.5395.</p> <p> </p> <p> </p>
Data from: PlaceMyFossils: An integrative approach to analyse and visualize the phylogenetic placement of fossils using backbone trees
Open the record for dataset details and reuse information.
Data from: A phylogenomic backbone for Acoelomorpha inferred from transcriptomic data
Open the record for dataset details and reuse information.
Serial disparity in the carnivoran backbone unveil a complex adaptive role in metameric evolution
Multi-element systems such as the vertebral column of vertebrates represent a major challenge to phenotypic quantification and macroevolutionary analyses. The vertebral column is a metameric structure, composed of serially repeated subunits, and much of what is known so far has been inferred from sparse anatomical samples, providing little insight into local-scale (i.e. vertebra-to-vertebra) variation and its macroevolutionary importance. This limits understanding of how evolutionary constraints and functional adaptation interact during the evolution of multi-element phenotypes. Here, we quantify morphological disparity across all subunits (vertebrae) of the pre-sacral column in the mammalian order Carnivora. We address how vertebral morphology varies among elements, and the extent to which these patterns have been structured by constraints and/or evolutionary adaptation to locomotory capabilities, using 3D geometric morphometrics and multivariate analyses for high-dimensional phenotypes. We find that lumbars and posterior thoracics exhibit high individual disparity but low serial differentiation. These vertebrae are pervasively recruited into locomotory functions, exhibiting high-dimensional ecomorphological signals and patterns of evolution indicative of relaxed constraints. Cervical and anterior thoracic vertebrae have low individual disparity and greater serial differentiation. Individual vertebrae in these regions unexpectedly also show signals of locomotory adaptation that were not generally recognized by previous studies. These are characterized by low-dimensional ecomorphological signals and overall constrained patterns of evolution. Our findings support the hypothesis that the lumbar region is a key innovation that increases ecological versatility of mammalian locomotion. Nevertheless, locomotory adaptation is more widely distributed along the mammalian axial skeleton. This has been masked by local-scale variation and low phenotypic variability in comparison with other skeletal structures such as the skull or limbs. Our analyses demonstrate that the strength of ecomorphological signal does not have a predictable influence on macroevolutionary outcomes even within the same structure, and undermine the traditional view that highly constrained skeletal units are strongly limited in their potential to adapt to new ecological avenues. Our findings emphasize the importance of quantifying local-scale variation in functionally versatile, multi-element phenotypes such as the vertebral column, or indeed, the vertebrate skeleton as a whole.
Data from: Deciphering an extreme morphology: bone microarchitecture of the Hero Shrew backbone (Soricidae: Scutisorex)
<p>Biological structures with extreme morphologies are puzzling because they often lack obvious functions and stymie comparisons to homologous or analogous features with more typical shapes. An example of such an extreme morphotype is the uniquely modified vertebral column of the hero shrew <i>Scutisorex</i>, which features numerous accessory intervertebral articulations and massively expanded transverse processes. The function of these vertebral structures is unknown, and it is difficult to meaningfully compare them to vertebrae from animals with known behavioral patterns and spinal adaptations. Here we use trabecular bone architecture of vertebral centra and quantitative external vertebral morphology to elucidate the forces that may act on the spine of <i>Scutisorex</i> and that of another large shrew with unmodified vertebrae (<i>Crocidura goliath</i>). µCT scans of thoracolumbar columns show that <i>S. thori </i>is structurally intermediate between <i>C. goliath</i> and <i>S. somereni </i>internally and externally, and both <i>Scutisorex</i> species exhibit trabecular bone characteristics indicative of higher <i>in vivo</i> axial compressive loads than <i>C. goliath.</i> Under compressive load, <i>Scutisorex </i>vertebral morphology is adapted to largely restrict bending to the sagittal plane (flexion). Although these findings do not solve the mystery of how <i>Scutisorex</i> uses its byzantine spine <i>in vivo</i>, our work suggests potentially fruitful new avenues of investigation for learning more about the function of this perplexing structure.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.