Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,641
datasets available to search
ShareScore release 0.9.0
Dataset results
1,641 results for “similarity”
Human Family With Sequence Similarity 83 Member B (FAM83B); A Target Enabling Package
<p>FAM83A-H are newly identified oncogenes characterised by a conserved DUF1669 domain. FAM83B can substitute for RAS to promote malignant transformation. Ablation of FAM83B or mutation of Lys230 inhibits malignant phenotypes, implicating FAM83B as potential therapeutic target. As part of this TEP, we solved the first crystal structures from the FAM83 family, including FAM83A and FAM83B. The structures of the DUF1669 domain reveal a phospholipase D-like fold lacking conservation of key catalytic residues. We deorphanise the FAM83 DUF1669 domain as a critical docking scaffold for binding of casein kinase 1 isoforms. Finally, using XChem fragment screening we report chemical fragments that bind to Lys230 in the central pocket of the DUF1669 and form starting points for potential drug development.</p>
Pairwise LexStat similarities for LingPy
<p>This dataset is a PostgreSQL dump of a database generated using <a href="https://github.com/Anaphory/autocode_big">autocode_big</a> from the data in the <a href="https://lexirumah.model-ling.eu/lexirumah">LexiRumah</a> lexical database of eastern Indonesia and Timor-Leste.</p> <p>The dataset contains the forms from the parent dataset, together with automatically generated sound correspondence scorers, pairwise similarity scores, cognate classes, and alignments, all created using <a href="http://lingpy.org/docu/compare/lexstat.html">Lingpy's LexStat algorithm</a>.</p>
(Un)expected similarity of the temporary adhesive systems of marine, brackish, and freshwater flatworms
<p>This repository contains</p> <ul> <li>a container with 441 single *.tiff files from a serial-block-face-imaging experiment of a Macrostomum lignano tail plate. The images were aligned with Dragonfly v. 2021.1 (ORS). This dataset was used to reconstruct the 3D model of a Macrostomum lignano adhesive organ. [Macrostomum_lignano_tailplate_SBFI.tar.gz]</li> <li>The raw Illumina PE150 reads from Macrostomum tuba [Mtub_1_R2_HTY2GBGXC_1_107969_TGTTGATCCTATGTTA.fastq.gz, Mtub_1_R1_HTY2GBGXC_1_107969_TGTTGATCCTATGTTA.fastq.gz]</li> <li>The assembled (Trinity v. 2.11.0 ) transcriptome of Macrostomum tuba including annotated fasta headers using Trinotate (v. 3.2.0)</li> </ul>
Software Similarity Dataset
<p>This dataset contains the post-processed data for software similarity learning. More information is given: <a href="https://github.com/SoftwareUnderstanding/softsim">SoftwareSim_Github</a></p> <p> </p> <p>post_process: All embedded software with autoencoder to make sure each function is the same length (1024 bits), each final is the embedded graph representation of software.</p> <p>final_data: All information obtained by <a href="https://github.com/KnowledgeCaptureAndDiscovery/somef">Somef</a> & <a href="https://github.com/SoftwareUnderstanding/inspect4py">Inspect4py</a> as well as cleaning. Each file represents software in the format given --> Function_Name: [[Called Function], [Function Tokens]]</p> <p>lean_simscore.csv: This file contains software pairs as well as the similarity metrics, format is given:</p> <table> <tbody> <tr> <td>Property</td> <td>Example</td> </tr> <tr> <td>Graph_1</td> <td>kakaobrain_helo_word</td> </tr> <tr> <td>Graph_2</td> <td>mblondel_soft-dtw</td> </tr> <tr> <td>miniLM</td> <td>0.4503</td> </tr> <tr> <td>Sbert</td> <td>0.7204</td> </tr> <tr> <td>TSDAE</td> <td>0.5714</td> </tr> </tbody> </table> <p> </p>
Data for: 'FAS: assessing the similarity between proteins using multi-layered feature architectures'
<p>Raw data and result data for the analyses made for the manuscript:</p> <p>'FAS: assessing the similarity between proteins using multi-layered feature architectures'</p> <p><a href="https://doi.org/10.1093/bioinformatics/btad226">https://doi.org/10.1093/bioinformatics/btad226</a></p> <p>This dataset contains raw data obtained from QFO Orthobench and Gene Ontology database. Analyses were made to showcase the different uses of the FAS algorithm.</p>
Supplemental Data for "Self-similarity of spectral response functions for fractional quantum Hall states"
<p>Scripts and data to supplement the paper "Self-similarity of spectral response functions for fractional quantum Hall states".</p>
Codon similarity data in ATTED-II ver 8.0 (Bra, Mtr)
<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p>
Codon similarity data in ATTED-II ver 8.0 (Ath, Gma, Osa, Sly, Vvi)
<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p> <p> </p>
Supplementary file 1 from: Moliner Cachazo L, Makati K, Chadwick MA, Catford JA, Price BW, Mackay AW, Guiry MD, Murray-Hudson M, Murray-Hudson F (2023) A review of the freshwater diversity in the Okavango Delta and Lake Ngami (Botswana): taxonomic composition, ecology, comparison with similar systems and conservation status. Aquatic Sciences
<p>Dataset with 2,204 freshwater species from the Okavango Delta and Lake Ngami (Botswana), with additional 355 species found in other areas of Botswana that are likely to be present in the study region. The dataset covers the following groups: amphibians, birds, fishes, macroinvertebrates, macrophytes, mammals, reptiles, phytoplankton, and zooplankton. The following information is given for each species: status in the Okavango Delta and Lake Ngami (present/potentially present); conservation status globally, Phylum, Class, Order, Family, Genus, species name, cited synonyms, common name, habitat, presence in high water, presence in low water, ecology, distribution in continental Africa, confirmed locations in the Okavango Delta, site coordinates, references, notes.</p>
Data on specialist and generalist herbivory, environmental variation, and phytochemical similarity from the Atlantic Rainforest of Brazil: 2013-2014
What: These are data on specialist and generalist herbivory from naturally occurring Piper plants as well as environmental data collected in 10 m diameter plots across sites in the Atlantic Rainforest of Brazil. Data also include chemical similarity calculated as the Morisita similarity index for species in a given plot and chemical modules that demonstrate classes of compounds that group together and influence herbivory. Why: Data were collected to understand factors that influence herbivory and tropical forest richness with a particular focus on secondary chemistry metabolomics. Where: Field sites include São Bento de Sapucaí (-22.8758, -45.8581), Parque Nacional de Itatiaia (-22.3698, -44.6285), Parque Estadual de Intervales (-24.3088, -48.2736), and Parque Estadual de Serra do Mar - Núcleo Pincinguaba (-23.6200, -46.7222). When: Data were collected in the field from Dec 2013 – May 2014. How: Field data were collected in 10 m diameter plots centered on randomly located Piper plants. Chemical data come from 1H-NMR metabolomics.
FIG. 5 in The equids represented in cave art and current horses: a proposal to determine morphological differences and similarities
FIG. 5. — Representation of the Magdelian horse figures studied: A-D, "Altamira"; E, F, "Niaux"; G, "Font de Gaume"; H-M, "Lascaux". Designed by Francisco Salado (IAPH), using dashed lines in some recreated parts of the figures.
FIG. 7 in The equids represented in cave art and current horses: a proposal to determine morphological differences and similarities
FIG. 7. — In current breeds of horses, the mane is long and hangs by the neck. In this picture, the "Retuertas" horse could be one of the most ancient breeds. Photo Esteban García-Viñas.
FIG. 6 in The equids represented in cave art and current horses: a proposal to determine morphological differences and similarities
FIG. 6. — Discriminant body proportion analysis of three species of modern horses show three separate groups using: A, absolute variables and B, indexes.
Some Similarities and Differences between the Observed Alfvénic Fluctuations in the Fast Solar Wind and Navier-Stokes Turbulence
<p>Three text data sets are uploaded: wind tunnel data (modane1.txt), solar-wind magnetic-field data (Flat1maginterp.txt) and solar-wind velocity data (Flat15interp3DP.txt)</p>
Data From: Inflorescence and flower development in Orchidantha chinensis T. L. Wu (Lowiaceae; Zingiberales): similarities to inflorescence structure in the Strelitziaceae
<p>The monotypic Lowiaceae remains the least known family in the plant order Zingiberales, yet it holds an important key to unraveling the phylogenetic placement of the families Musaceae, Heliconiaceae, Strelitziaceae, and Lowiaceae. After nine phylogenetic studies the (Lowiaceae, Strelitziaceae) clade is the only stable clade that has emerged in this half of the order. This study was undertaken to verify the unusual inflorescence and flower structure in Orchidantha, and to search for new characters that might be used in future phylogenetic analyses. We describe both inflorescence and flower development in a previously unstudied species, confirm inflorescence morphology in the genus, and compare the structure of the inflorescence in the Lowiaceae with that of the Strelitziaceae, its potential sister group. </p> <p>The inflorescence of Orchidantha is born at the end of a vegetative shoot and is composed of two lateral branches that each bear four bracts and a single flower, before aborting. The fourth bract and its associated flower form the highly reduced flower cluster (florescence) that characterizes this genus. In technical terms Orchidantha has a polytelic synflorescence that lacks a main florescence (it has a truncated polytelic synflorescence) and bears solitary flowers in coflorescences on determinate enriching branches. The enriching branches produce a fixed number of bracts before aborting (i.e., they are special paracladia). Many of these features are shared with the Strelitziaceae.</p> <p>Similarities between the Lowiaceae and Strelitziaceae include inflorescence structure, the presence of a long prolongation of the ovary, and a delay in the formation of the third sepal during flower development, a character that is also shared with the Musaceae. Inflorescence and flower structure is now well established in this small, but important family.</p>
Supporting data and code for "Brown AW, Bohan Brown MM, Onken KL, Beitz DC. Short-term consumption of sucralose, a nonnutritive sweetener, is similar to water with regard to select markers of hunger signaling and short-term glucose homeostasis in women. Nutr Res. 2011 Dec;31(12):882-8. doi: 10.1016/j.nutres.2011.10.004. PMID: 22153513."
<p>Data and code to support the publication, Brown AW, Bohan Brown MM, Onken KL, Beitz DC. Short-term consumption of sucralose, a nonnutritive sweetener, is similar to water with regard to select markers of hunger signaling and short-term glucose homeostasis in women. Nutr Res. 2011 Dec;31(12):882-8. doi: 10.1016/j.nutres.2011.10.004. PMID: 22153513.</p> <ul> <li>Code was updated 2020 NOV 24 to add comments, but otherwise remains unchanged from 2011.</li> <li>Data file was updated to include a data dictionary, but otherwise remains unchanged from 2011.</li> </ul> <p>Treatment identifiers for the four-arm crossover are clarified in the SAS code.</p> <p>Other details are available in the published article.</p>
Long document similarity datasets, Wikipedia excerptions for movies, video games and wine collections
<p>Three corpora in different domains extracted from Wikipedia.</p> <p>For all datasets, the figures and tables have been filtered out, as well as the categories and "see also" sections.</p> <p>The article structure, and particularly the sub-titles and paragraphs are kept in these datasets</p> <p> </p> <p><strong>Wines</strong></p> <p>Wikipedia wines dataset consists of 1635 articles from the wine domain. The extracted dataset consists of a non-trivial mixture of articles, including different wine categories, brands, wineries, grape types, and more. The ground-truth recommendations were crafted by a human sommelier, which annotated 92 source articles with ~10 ground-truth recommendations for each sample. Examples for ground-truth expert-based recommendations are </p> <ul> <li>Dom Pérignon - Moët & Chandon</li> <li>Pinot Meunier - Chardonnay</li> </ul> <p><strong>Movies</strong></p> <p>The Wikipedia movies dataset consists of 100385 articles describing different movies. The movies' articles may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.<br> For this dataset, we have extracted a test set of ground truth annotations for 50 source articles using the "<a href="https://bestsimilar.com/">BestSimilar</a>" database. Each source articles is associated with a list of ${\scriptsize \sim}12$ most similar movies.<br> Examples for ground-truth expert-based recommendations are </p> <ul> <li>Schindler's List - The Pianist</li> <li>Lion King - The Jungle Book</li> </ul> <p><strong>Video games</strong></p> <p>The Wikipedia video games dataset consists of 21,935 articles reviewing video games from all genres and consoles. Each article may consist of a different combination of sections, including summary, gameplay, plot, production, etc. Examples for ground-truth expert-based recommendations are:</p> <ul> <li>Grand Theft Auto - Mafia</li> <li>Burnout Paradise - Forza Horizon 3</li> </ul>
Surveying the communities of users of MATLAB and similar languages (Responses)
<p>This upload includes responses gathered from a survey applied to the users of MATLAB and its clone languages. It covers the participants' programming experience, how they interact with these languages, the importance they give to the reusability of their programs, object-oriented programming and how satisfied they are with these languages.</p>
A dataset used to determine a semantic similarity metric based on UMLS for PMC-OA
<p>We have performed a series of in-silico experiments in order to determine a semantic similarity metric based on UMLS annotations for PubMed Central Open Access. Here we have stored the data used for and obtained from such experiments. We have worked with relevant and partially relevant articles from the TREC-2005 Genomics Track Collection, from now referred as the initial collection, including a total of 4240 unique PubMed articles. From those 4240 articles, only 62 had publicly available; those 62 articles correspond to the full-text collection.</p> <p>Our data comprises flat files using tabs as separators and one Excel sheet. Tab separated values always include a first row with headings:</p> <ul> <li>Stems extracted from title and abstract for articles in the initial collection. Each row contains a stem with its inverse-document-frequency (IDF) within the initial collection. Stems were calculated following the Porter algorithm (available at http://tartarus.org/martin/PorterStemmer/java.txt) <ul> <li>stems.TA.tsv</li> </ul> </li> <li>Article profiles, i.e., terms (either word stems or UMLS concepts) found in the articles with term frequency (TF) and IDF. The first two columns correspond to PubMed Identifier (PMID) and PubMed Central identifier (PMC). PMC identifier was set to 0 whenever full-text was not available. <ul> <li>profiles.TA.tsv: Profiles according word stems in title and abstract for the initial collection</li> <li>profiles.PMID.tsv: Profiles according to UMLS concpets in title and abstract for the initial collection</li> <li>profiles.PMC_TA.tsv: Profiles according to UMLS concepts in title and abstract for the full-text collection</li> <li>profiles.PMC.tsv: Profiles according to UMLS concepts in the full-text for the full-text collection</li> </ul> </li> <li>Similarity matrixes calculated on the article profiles with PubMed Related Article metric (PMRA), BM25, and Cosine. There are matrixes for terms found in title-and-abstract as well as full-text. In a similarity matrix, a reference article (an interest has been already expressed for it) correspond to a row, while the columns correspond to all the other articles for which the similarity was calculated. <ul> <li>Matrixes for our initial collection <ul> <li>similarity.PMRA.TA.profiles.TA.tsv: Similarity matrix for profiles.TA.tsv following the algorithm PMRA. This matrix is considered the baseline for further analyses</li> <li>similarity.PMRA.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm PMRA</li> <li>similarity.BM25_1.2_0.75.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm BM25 with k=1.2 and b=0.75</li> <li>similarity.COSINE.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm Cosine</li> </ul> </li> <li>Matrixes for our full-text collection <ul> <li>similarity.PMRA.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm PMRA</li> <li>similarity.PMRA.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm PMRA</li> <li>similarity.BM25.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm BM25 with k=1.2 and b=0.75</li> <li>similarity.BM25.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm BM25 with k= 1.2 and b= 0.75</li> <li>similarity.COSINE.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm Cosine</li> <li>similarity.COSINE.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm Cosine</li> </ul> </li> </ul> </li> <li>Correlation matrixes for similarities calculated for title-and-abstract taking as reference the similarity values obtained with PMRA for word stems on title-and-abstract. <ul> <li>pearsonCorrelation.PMRA.tsv: Correlation for similarity.PMRA.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.PMRA.tsv: Correlation for similarity.PMRA.profiles.PMID.tsv discriminated by TREC topics</li> <li>pearsonCorrelation.BM25_1.2_0.75.tsv: Correlation for similarity.BM25_1.2_0.75.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.BM25_1.2_0.75.tsv: Correlation for similarity.BM25_1.2_0.75.profiles.PMID.tsv discriminated by TREC topics</li> <li>pearsonCorrelation.COSINE.tsv: Correlation for similarity.COSINE.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.COSINE.tsv: Correlation for similarity.COSINE.profiles.PMID.tsv discriminated by TREC topics</li> </ul> </li> <li>Precision and recall summaries for the similarities calculated based on title-and-abstract. <ul> <li>StatsAllSummary.xlsx: Precision and recall at a global level, i.e., without considering TREC topics. This file includes information for BM25 with multiples values for constants k and b</li> </ul> </li> </ul> <p>Visualization for correlation matrixes as well as scattered plots for full-text based similarity is available at http://ljgarcia.github.io/semsim.benchmark</p>
Similar Sentences from English Wikipedia 20150304
<p>Similar sentences from a dump of enwiki from March 04, 2015 detected using the Wikiduper software (https://github.com/seweissman/wikiduper).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.