Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,641
datasets available to search
ShareScore release 0.9.0
Dataset results
1,641 results for “similarity”
Datasets for: Generalizing Monin-Obukhov Similarity Theory (1954) for Complex Atmospheric Turbulence, Stiperski and Calaf 2023, PRL
<p>Scaling variables for the generalized flux-variance scaling relations that include turbulence anisotropy. Dataset is a companion to the manuscript Stiperski, I., Calaf, M., 2023: Generalizing Monin-Obukhov similarity theory (1954) for complex atmospheric turbulence. Physical Review Letters, 130 (12), 124001, https://doi.org/10.1103/PhysRevLett.130.124001</p> <p>The dataset contains the turbulence statistics from 13 datasets: AHATS, Cabauw, CASES-99, METCRAX II campaign (NEAR and RIM towers), T-Rex campaign (Central tower - TRexC, West tower - TRexW) and i-Box measurement network (CCS-VF0 tower - i-Box0, CS-SF1 tower - i-Box1, CS-NF10 tower - i-Box10, CS-NF27 tower - i-Box27, CS-MT21 tower - i-BoxTop, im Hinteren Eis tower - imHint).</p> <p><br>Data are organized in csv files for each datasets and only contain high quality (for applied criteria see the Supplemental Material of the companion paper, https://journals.aps.org/prl/supplemental/10.1103/PhysRevLett.130.124001) data with 30 min averaging for unstable stratification and 1 min for stable stratification. Since the data were used for scaling, there is no reference to time, but the measurement height is provided as an additional variable. </p> <p>Meaning of variables:</p> <p>zeta - z/L where z is height above ground and L is the local Obukhov length</p> <p>SigmaU - $\overline{u'u'}/u_*$ scaled standard deviation of streamwise velocity, where $u_*$ is the local friction velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of spanwise velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of surface-normal velocity</p> <p>SigmaT - $\overline{T'T'}/T_*$ scaled standard deviation of sonic temperature, where $T_*$ is the local temperature scale</p> <p>SigmaEpsU - scaled dissipation rate of the streamwise velocity</p> <p>SigmaEpsW - scaled dissipation rate of the surface-normal velocity </p>
Multilingual news article similarity dataset
<p>This dataset contains the extended version of the authors' earlier work: <a href="../records/6507872">https://zenodo.org/records/6507872,</a> where pairs of news articles drawn from the first half of 2020 are annotated for seven aspects of similarity in the original version as well as an additional FRAME aspect:</p> <ul> <li><strong>GEO</strong>: How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong> How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong> Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong> How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong> Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong> Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> <li><strong>FRAME</strong> Do the articles have similar framing and express similar opinions?</li> </ul>
Environmental Subsidies and Similar Transfers from Europe to the Rest of the World
<p>Environmental subsidies and similar transfers (current, capital, tax abatement, subsidy) for all environmental protection and resource management activities from EU countries to the rest of the world.</p> <p>The original dataset is of Eurostat is plagued with missing data. Our version on the <a href="https://zenodo.org/communities/greendeal_observatory/">Green Deal Data Observatory</a>, though could be further improved, offers a 167% larger congruent data matrix for supervised or unsupervised learning (machine learning, regression analysis) than the <a href="https://ec.europa.eu/eurostat/databrowser/view/ENV_ESST_GG/default/table?lang=en">original dataset</a>: Environmental subsidies and similar transfers from general government, by environmental activity, sector of recipient and ESA category of transfer.</p> <p> </p>
SemEval-2022 Task 8: Multilingual news article similarity
<p>This dataset contains pairs of news articles drawn from the first half of 2020 and annotated for seven aspects of similarity:</p> <ul> <li><strong>GEO</strong>: How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong> How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong> Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong> How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong> Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong> Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> </ul> <p>Further details are provided in</p> <blockquote> <p>Chen et al. (2022). SemEval-2022 Task 8: Multilingual news article similarity. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022). <a href="https://aclanthology.org/2022.semeval-1.155/">https://aclanthology.org/2022.semeval-1.155/</a></p> </blockquote> <p>The data in this repository includes pairs of URLs and annotations. The text of webpages is generally via the Internet Archive in this special collection: https://archive.org/details/2020-multilingual-news-article-similarity . A script to download and process the webpages is available at https://github.com/euagendas/semeval_8_2022_ia_downloader . </p>
Comparative plot about two sequences of integers very similar to each other: A348960 vs A127034
<p>This is the behavior comparison graph between the integer sequences registered in OEIS (The On-Line Encyclopedia of Integer Sequences) with codes: A348960 & A127034 respectively.</p> <p><strong>1)</strong> The sequence A348960 obeys the formula: </p> <p> <span class="math-tex">\(a_{n }=\lfloor(log(\pi n!)\rfloor. \)</span> For any non-negative integer such that <span class="math-tex">\(n\geq0\)</span></p> <p><strong>2) </strong>The sequence A127034 obeys the formula:</p> <p><span class="math-tex">\(a_{n}=\lfloor log(n!)/log(11)\rfloor.\)</span> For any non-negative integer such that <span class="math-tex">\(n\geq0\)</span></p> <p>In this particular case we're conducting the study for the following n-values: i<span class="math-tex">\(1\leq n\leq 60.\)</span> </p> <p> </p>
MiRoR11 - P2 - Annotated corpus for semantic similarity of clinical trial outcomes
<p>Outcome similarity corpus</p> <p>This dataset contains annotations of semantic similarity for pairs of primary and reported outcomes.<br> Tab-separated format is used. The files contain the following columns:<br> filename, sentence pair ID, sentence pair text, primary outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_relations_split contains the dataset splits for 10-fold cross-validation.</p>
Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data
<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed “similarity-based merger models” which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC>0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>
EMBERSim: A Large-Scale Databank for Boosting Similarity Search in Malware Analysis
<p>In recent years there has been a shift from heuristics-based malware detection towards machine learning, which proves to be more robust in the current heavily adversarial threat landscape. While we acknowledge machine learning to be better equipped to mine for patterns in the increasingly high amounts of similar-looking files, we also note a remarkable scarcity of the data available for similarity-targeted research. Moreover, we observe that the focus in the few related works falls on quantifying similarity in malware, often overlooking the clean data. This one-sided quantification is especially dangerous in the context of detection bypass. We propose to address the deficiencies in the space of similarity research on binary files, starting from EMBER — one of the largest malware classification data sets. We enhance EMBER with similarity information as well as malware class tags, to enable further research in the similarity space. Our contribution is threefold: (1) we publish EMBERSim, an augmented version of EMBER, that includes similarity-informed tags; (2) we enrich EMBERSim with automatically determined malware class tags using the open-source tool AVClass on VirusTotal data and (3) we describe and share the implementation for our class scoring technique and leaf similarity method.</p>
The Contributionsof Eye Gaze Fixations and Target-Lure Similarity to Behavioral and fMRI Indices of Pattern Separation and Pattern Completion
Open the record for dataset details and reuse information.
MiRoR11 - P2 - Annotated dataset for spin-related types of statements (statements of similarity and within-group comparisons)
<p>180 abstracts / 2401 sentences annotated for 2 types of spin-related statements: statements of similarity and within-group comparisons.</p>
VSA, Analogy, and Dynamic Similarity
<p>This file is the video of a presentation originally scheduled to be given at the <a href="https://sites.google.com/view/vsaworkshop2020/home">Workshop on Developments in Hyperdimensional Computing and Vector Symbolic Architectures</a>, 16 March 2020, Kirchoff-Institute for Physics at Heidelberg University, Heidelberg, Germany. The workshop meeting was cancelled due to the COVID-19 pandemic and this presentation was given as a webinar on 2020-05-18.</p> <p><strong>Extended Abstract</strong></p> <p>It has been argued that analogy is at the core of cognition [7, 1]. My work in VSA is driven by the goal of building a practical, effective analogical memory/reasoning system. Analogy is commonly construed as structure mapping between a source and target [5], which in turn can be construed as representing the source and target as graphs and finding maximal graph isomorphisms between them. This can also be viewed as a kind of dynamic similarity in that the<br> initially dissimilar source and target are effectively very similar after mapping.</p> <p>Similarity (the angle between vectors) is central to the mechanics of VSA/HDC. Introductory papers (e.g. [8]) necessarily devote space to vector similarityand the effect of the primitive operators (sum, product, permutation) on similarity. Most VSA examples rely on static similarity, where the vector representations are fixed over the time scale of the core computation (which is usually a single-pass, feed-forward computation). This emphasises encoding methods (e.g.<br> [12, 13]) that create vector representations with the similarity structure required by the core computation. Random Indexing [13] is an instance of the vector embedding approach to representation [11] that is widely used in NLP and ML. The important point is that the vector embeddings are developed in advance and then used as static representations (with fixed similarity structure) in the<br> subsequent computation of interest.</p> <p>Human similarity judgments are known to be context-dependent (see [3] for a brief review). It has also been argued that similarity and analogy are based on the same processes [6] and that cognition is so thoroughly context-dependent that representations are created on-the-fly in response to task demands [2]. This seems extreme, but doesn’t necessarily imply that the base representations are context-dependent as long as the cognitive process that compares them is<br> context-dependent, which can be achieved by having dynamic representations that are derived from the static base representations by context-dependent transforms (or any functionally equivalent process).</p> <p>An obvious candidate for a dynamic transformation function in VSA is substitution by binding, because the substitution can be specified as a vector and dynamically generated (see Representing substitution with a computed mapping in [8]). This implies an internal degree of freedom (a register to hold the substitution vector while it evolves) and a recurrent VSA circuit to provide the dynamics to evolve the substitution vector.</p> <p>These essential aspects are present in [4], which finds the maximal subgraph isomorphism between two graphs represented as vectors. This is implemented as a recurrent VSA circuit with a register containing a substitution vector that evolves and settles over the course of the computation. The final state of the substitution vector represents the set of substitutions that transforms the static base representation of each graph into the best subgraph isomorphism to the static base representation of the other graph. This is a useful step along the path to an analogical memory system.</p> <p>Interestingly, the subgraph isomorphism circuit can be interpreted as related to the recently developed Resonator Circuits for factorisation of VSA representations [9], which have internal degrees of freedom for each of the factors to be calculated and a recurrent VSA dynamics that settles on the factorisation. The graph isomorphism circuit can be interpreted as finding a factor (the substitution vector) such that the product of that factor with each of the graphs is the<br> best possible approximation to the other graph. This links the whole enterprise back to statistical modelling, where there is a long history of approximating matrices/tensors as the product of simpler factors [10].</p> <p>References<br> 1. Blokpoel, M., Wareham, T., Haselager, P., van Rooij, I.: Deep Analogical Inference as the Origin of Hypotheses. The Journal of Problem Solving 11(1), 1–24 (2018)<br> 2. Chalmers, D.J., French, R.M., Hofstadter, D.R.: High-level perception, representation, and analogy: A critique of artificial intelligence methodology. Journal of Experimental & Theoretical Artificial Intelligence 4(3), 185–211 (1992)<br> 3. Cheng, Y.: Context-dependent similarity. In: Proceedings of the Sixth Annual Conference on Uncertainty in Artificial Intelligence (UAI’90), pp. 27–30. Cambridge, MA, USA (1990)<br> 4. Gayler, R.W., Levy, S.D.: A distributed basis for analogical mapping. In: Proceedings of the Second International Conference on Analogy (ANALOGY-2009), pp. 165–174. New Bulgarian University, Sofia, Bulgaria (2009)<br> 5. Gentner, D.: Structure-mapping: A theoretical framework for analogy. Cognitive Science 7(2), 155–170 (1983)<br> 6. Gentner, D., Markman, A.B.: Structure mapping in analogy and similarity. American Psychologist 52(1), 45–56 (1997)<br> 7. Gust, H., Krumnack, U., Kühnberger, K.-U., Schwering, A.: Analogical Reasoning: A core of cognition. KI - Künstliche Intelligenz 1(8), 8–12 (2008)<br> 8. Kanerva, P.: Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors. Cognitive Computation 1, 139–159 (2009)<br> 9. Kent, S.J., Frady, E.P., Sommer, F.T., Olshausen, B.A.: Resonator Circuits for factoring high-dimensional vectors. http://arxiv.org/abs/1906.11684 (2019)<br> 10. Kolda, T.G., Bader, B.W.: Tensor decompositions and applications. SIAM Review 51(3), 455–500 (2009)<br> 11. Pennington, J., Socher, R., Manning, C.D.: GloVe: Global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543. Association for Computational Linguistics, Doha, Qatar (2014)<br> 12. Purdy, S.: Encoding data for HTM systems. http://arxiv.org/abs/1602.05925 (2016)<br> 13. Sahlgren, M.: An introduction to random indexing. In: Proceedings of the Methods and Applications of Semantic Indexing Workshop at the 7th International Conference on Terminology and Knowledge Engineering (TKE 2005), Copenhagen, Denmark (2005)</p>
Crowdsourcing Document Similarity Judgements
<p>This is the data obtained from crowdsourcing tasks which ask workers to provide similarity metrics between pairs of documents. Each document, as well as each pair, has a unique ID. We provide crowd workers with the pairs through three different task variations:</p> <ul> <li>Variation 1: We showed workers 5 pairs of documents and, for each, asked them to rate their similarity in a 4-level Likert scale (None, Low, Medium, High), tell us a confidence level of how sure they were (from 0 to 4) and a written reason as to why they chose that similarity level. For quality reasons, two of the 5 pairs were golden-standards, which means we knew their ratings already and checked the workers' responses. They had to give the golden pair with the higher similarity a higher score than the other golden pair, otherwise, their answer would be rejected.</li> <li>Variation 2: We repeated variation 1 but with a slight alteration: instead of a Likert scale for the similarity score, we asked for a Magnitude Estimation, which is any number above 0. It could be 1, 0.0001, 1000, 42, as long as it was coherent, as in a more similar pair had a higher score than a less similar pair and vice-versa;</li> <li>Variation 3: We showed workers 5 rankings. Each ranking had a main document and 3 auxiliary documents to be compared against the main one. They also had to report a confidence score and give a short written reason, just like variation 1. The first ranking is a golden-standard, and we knew the values for the 3 pairs in it (the pairs were the main document paired with each of the 3 auxiliary documents), and they had to give the golden pair with the highest similarity a higher rank than the one with the lower similarity.</li> </ul> <p>The raw results from the tasks are recorded in the JSON file CrowdResults.json. For a description of its contents, please read the file CrowdResults_README.md.</p> <p>These raw annotations from the crowd were then parsed into the three CSVs you see, each corresponding to the aggregated results from one of the task variations.</p> <ul> <li><em>final_scores_likert.csv</em> is the resulting scores for each pair using the variation 1 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>similarity_crowd_simple_maj </em>stores the simple majority result from the crowd's annotations;</li> <li><em>similarity_crowd_simple_mean </em>stores the mean of the crowd's annotations;</li> <li><em>similarity_crowd_simple_median </em>stores the median of the crowd's annotations;</li> </ul> </li> <li><em>final_scores_magnitude.csv</em> is the resulting scores for each pair using the variation 2 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>scaled_similarity_worker</em> is the magnitude score scaled based on worker's behaviours</li> <li><em>scaled_similarity_worker_docset </em>is the magnitude score scaled based both on the worker's behaviour and on the pair</li> </ul> </li> <li><em>final_scores_ranking.csv</em> is the resulting scores for each pair using the variation 3 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>mean_similarity</em> is the mean ranking from that value</li> </ul> </li> </ul> <p>This dataset was built and used as part of the <a href="https://theybuyforyou.eu/">TheyBuyForYou </a>project.</p>
Detection of Functionally Similar Code Clones: Data, Analysis Software, Benchmark
<p>We analysed 2,800 programs in Java and C for which we knew they are functionally similar. We checked if existing clone detection tools are able to find these functional similarities and classified the non-detected differences. We make all used data, the analysis software as well as the resulting benchmark available here.</p>
Biolinks, datasets and algorithms supporting semantic-based distribution and similarity for scientific publications
<p><strong>Background: </strong>Finding articles related to a publication of interest remains a challenge in the Life Sciences domain as the number of scientific publications grows day by day. Publication repositories such as PubMed and Elsevier provides a list of similar articles. There, similarity is commonly calculated based on title, abstract and some keywords assigned to articles. Here we present the datasets and algorithms used in Biolinks. Biolinks uses ontological concepts extracted from publication and makes it possible to calculate a distribution score according to semantic groups as well as a semantic similarity based on either all identified annotations or narrowed to one or more particular semantic groups. Biolinks supports both title and abstract only as well as full-text.</p> <p><strong>Materials: </strong>In a previous work [1], 4,240 articles from the TREC-05 collection [2] were selected. The title-and-abstract for those 4,240 articles were annotated with Unified Medical Language System (UMLS) concepts, such annotations are refer to as our TA-dataset and correspond to the JSON files under the pubmed folder in the JSON-LD.zip file. From those 4,240 articles, full-text was available for only 62. The title-and-abstract annotations for those 62 articles, TAFT-dataset, are located under the pubmed-pmc folder in the JSON-LD.zip file, which also contains the full-text annotations under the folder pmc, FT-dataset. The list corresponding to articles with title-and-abstract is found in the genomics.qrels.large.pubmed.onlyRelevants.titleAndAbstract.tsv file, while those with full-text are recorded in the genomics.qrels.large.pmc.onlyRelevants.fullContent.tsv file.</p> <p>Here we include the annotations on title and abstract as well as those for full-text for all our datasets (profiles.zip). We also provide the global similarity matrices (similarity.zip).</p> <p><strong>Methods:</strong> The TA-dataset was used to calculate the Information Gain (IG) according to the UMLS semantic groups, see IG_umls_groups.PMID.xlsx. A new grouping is proposed for Biolinks, see biolinks_groups.tsv. The IG was calculated for Biolinks groups as well, IG_biolinks_groups.PMID.xlsx, showing a improvement around 5%.</p> <p>In order to assess the similarity metric regarding the cohesion of TREC-05 groups, we used Silhouette Coefficient analyses. An additional dataset Stem-TAFT-dataset was used and compared to TAFT and FT datasets.</p> <p>Biolinks groups were used to calculate a semantic group distribution score for each article in all our datasets. A semantic similarity metric based on PubMed related articles [3] is also provided; the Biolinks groups can be used to narrow the similarity to one or more selected groups. All the corresponding algorithms are open-access and available on GitHub under the license Apache-2.0, a frozen version, biotea-io-parser-master.zip, is provided here. In order to facilitate the analysis of our datasets based on the annotations as well as the distribution and similarity scores, some web-based visualization components were created. All of them open-access and available in GitHub under the license Apache-2.0; frozen versions are provided here, see files biotea-vis-annotation-master.zip, biotea-vis-similarity-master.zip, biotea-vis-tooltip-master.zip and biotea-vis-topicDistribution-master.zip. These components are brought together by biotea-vis-biolinks-master.zip. A demo is provided at http://ljgarcia.github.io/biotea-biolinks/; this demo was built on top of GitHub pages, a frozen version of the gh-pages branch is provided here, see biotea-biolinks-gh-pages.zip.</p> <p><strong>Conclusions: </strong>Biolinks assigns a weight to each semantic group based on the annotations extracted from either title-and-abstract or full-text articles. It also measures similarity for a pair of documents using the semantic information. The distribution and similarity metrics can be narrowed to a subset of the semantic groups, enabling researchers to focus on what is more relevant to them.</p> <p> </p> <p>[1] Garcia Castro, L.J., R. Berlanga, and A. Garcia, <em>In the pursuit of a semantic similarity metric based on UMLS annotations for articles in PubMed Central Open Access.</em> Journal of Biomedical Informatics, 2015. <strong>57</strong>: p. 204-218</p> <p>[2] Text Retrieval Conference 2005 - Genomics Track. <em>TREC-05 Genomics Track ad hoc relevance judgement</em>. 2005 [cited 2016 23rd August]; Available from: http://trec.nist.gov/data/genomics/05/genomics.qrels.large.txt</p> <p>[3] Lin, J. and W.J. Wilbur, <em>PubMed related articles: a probabilistic topic-based model for content similarity.</em> BMC Bioinformatics, 2007. <strong>8</strong>(1): p. 423</p>
KiSSim: Predicting off-targets from structural similarities in the kinome
<p><strong>KiSSim: Predicting off-targets from structural similarities in the kinome</strong></p> <p><strong>Project description.</strong></p> <p>KiSSim (Kinase Structural Similarity) is a novel fingerprint designed specifically for kinase pockets, allowing for similarity studies across the structurally covered kinome. The kinase fingerprint is based on the <a href="https://klifs.net/">KLIFS</a> pocket alignment, which defines 85 pocket residues for all kinase structures. This enables a residue-by-residue comparison without a computationally expensive alignment step.</p> <p>The pocket fingerprint encodes each pocket residue’s spatial and physicochemical properties. The spatial properties describe the residue’s position in relation to the kinase pocket center and important kinase subpockets, i.e. the hinge region, the DFG region, and the front pocket. The physicochemical properties encompass for each residue its size and pharmacophoric features, solvent exposure, and side chain orientation.</p> <p>Some datasets are not part of the `kissim_app` GitHub repository due to their size but can be downloaded from here to the respective kissim_app folders.</p> <p><strong>Data.</strong></p> <ul> <li>`20210902_KLIFS_HUMAN.tar.gz` --- save in `kissim_app/data/external/structures`</li> <li>`complete_SiteAlign.txt.gz` --- save in `kissim_app/data/external/sitealign`</li> </ul> <p><strong>Results.</strong></p> <ul> <li>`results.tar.bz2`--- save as `kissim_app/results`</li> </ul> <p>These are the KiSSim results: fingerprints, feature/fingerprint distances, kinase matrices, and kinase trees for structures in all (`all`), DFG-in (`dfg_in`), and DFG-out (`dfg_out`) conformation. In the case of the DFG-in conformation, we also have KiSSim runs with fingerprint subsets based on only residues that interact with certain ligands in KLIFS IFPs: Erlotinib (`dfg_in_IRE`), Imatinib (`dfg_in_STI`), Bosutinib (`dfg_in_DB8`), and Dopamapimod (`dfg_in_B96`). The folder contains README with a detailed file list.</p> <p><strong>Usage.</strong></p> <p>This dataset can be used to run the notebooks available on <a href="https://github.com/volkamerlab/kissim_app">https://github.com/volkamerlab/kissim_app</a>.</p> <ol> <li>Clone the kissim_app repository.</li> <li>Download the files provided here.</li> <li>If applicable, extract the archive content to the folders as indicated above and run the notebooks.</li> </ol> <pre><code class="language-bash">cd /path/to/your/download tar -xvf results.tar.bz2 -C /path/to/kissim_app/ tar -xvf 20210902_KLIFS_HUMAN.tar.bz2 -C /path/to/kissim_app/data/external/structures/ # In case you want the raw SiteAlign data mv complete_SiteAlign.txt.gz /path/to/kissim_app/data/external/sitealign</code></pre> <p><strong>Citation.</strong></p> <p>These datasets are part of the KiSSim publication: TBA</p>
Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "
<p>Datasets for preprint entitled "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi"</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of <em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi". </p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict. For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22. In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes. </p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset is made up of predicted protein tertiary structures representing the main member of each up-regulated <em>V. inaequalis</em> effector candidate family. Structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). In cases where the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2. Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold (<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>) open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input. </p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi" study. These structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input. </p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>
Evaluation Set - Contributions Similarity in the Open Research Knowledge Graph
<p>This evaluation set has been created for evaluating a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts structured ORKG contribution as input and recommends existing contributions in the ORKG semantically relevant to the given one.</p> <p> </p> <p>The evaluation set is manually annotated based on the <a href="https://www.orkg.org/orkg/featured-comparisons">featured comparisons</a> in the ORKG. In the course of this, it has been distinguished between homogeneous (those who are dissimilar in 2-3 properties) and heterogeneous (otherwise) instances. Multiple annotations have been obtained for the former and exactly one for the latter.</p> <p> </p> <p>It has been also distinguished between "with_response" and "without_response" instances (50 instances for each). The former are those contributions for them the initial version of the contributions similarity service has found similarities and the latter are the opposite case.</p> <p> </p> <p>This evaluation set has been created and applied on a modified version of the contributions similarity service in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>. The modified version of the service has simplified the document representation of contributions that are stored in an ElasticSearch index by omitting redundant terms.</p> <p>The evaluation set has the following schema:</p> <pre><code class="language-json">{ "with_response": [ { "contribution_id": "some_id", "comparison_id": "some_id", "comparison_label": "some_label", "contribution_label": "some_label", "paper": "some_id", "research_field": "some_id", "research_problems": [ "some_id" ], "annotations": [ "some_id of a similar contribution", ... ] }, ... ], "without_response": [ ... ] }</code></pre> <p> </p>
Identifying and profiling structural similarities between Spike of SARS-CoV-2 and other viral or host proteins with Machaon - Pre-computed features for replication
<p>Machaon's computed features that were used in the structural comparisons with Spike protein.</p> <p>DATA_PDBS_vir_whole_1-3.zip files are parts of a single folder.</p> <p> </p>
Self-similarity, density-size dynamics and the sinking speed of marine aggregates.
<p>A collated and referenced data base of observed size and sinking speeds of marine particle aggregates including Reynolds number and estimated excess density.</p>
Spectral irradiance at Lammi Biological Station Research Forest 2015: for assessing scale-wise similarity of curves with a thick pen
<p>This dataset contains records of the solar spectral energy irradiance (W m<sup>-2</sup> nm<sup>-1</sup>) in the understorey of forest stands at Lammi Biological Station, southern Finland (61◦ 3.24’ N, 25◦ 118 2.23’ E) during the spring of 2015. These spectra allow the change in spectral energy irradiance to be followed through the period of canopy leaf flush. Records are the average of recorded spectra from four points recorded at 40-cm above the forest floor using a Maya 2000 Pro array spectrometer. Spectra were recorded from exactly the same location on three dates, 2015-04-25, 2015-05-22, and 2015-06-05, before, during and after leaf flush. Data were recorded from the understorey of a young Betula stand, an old Betula stand, an old mixed Betula stand, a Quercus stand, and a Picea stand, in three positions: shade, semi-shade from leaves, and full sun in a sunfleck. On each occasion control measurements of spectral energy irradiance in full sun of an open field were also recorded at the beginning, middle and end of each measurement period. All measurements were made during the 2 hours either side of solar noon, on clear-sky days. Details of the sampling method and interpretation are given in the paper, Hartikainen et al., (2018) in Ecology and Evolution, which showcases the use of Thick Pen Transform to compare spectra.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.