Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

61

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

61 results for “sequence similarity”

Learn how ShareScore rates datasets ↗
zenodo48/100

Comparative plot about two sequences of integers very similar to each other: A348960 vs A127034

<p>This is the behavior comparison graph between the integer sequences registered in OEIS (The On-Line Encyclopedia of Integer Sequences) with codes: A348960 &amp; A127034 respectively.</p> <p><strong>1)</strong> The sequence A348960&nbsp;obeys the formula:&nbsp;</p> <p>&nbsp;<span class="math-tex">\(a_{n }=\lfloor(log(\pi n!)\rfloor. \)</span>&nbsp;For any non-negative integer such that&nbsp;<span class="math-tex">\(n\geq0\)</span></p> <p><strong>2) </strong>The sequence A127034 obeys the formula:</p> <p><span class="math-tex">\(a_{n}=\lfloor log(n!)/log(11)\rfloor.\)</span>&nbsp;For any non-negative integer such that&nbsp;<span class="math-tex">\(n\geq0\)</span></p> <p>In this particular case we&#39;re conducting the study for the following n-values: i<span class="math-tex">\(1\leq n\leq 60.\)</span>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Human Family With Sequence Similarity 83 Member B (FAM83B); A Target Enabling Package

<p>FAM83A-H are newly identified oncogenes characterised by a conserved DUF1669 domain. FAM83B can substitute for RAS to promote malignant transformation. Ablation of FAM83B or mutation of Lys230 inhibits malignant phenotypes, implicating FAM83B as potential therapeutic target. As part of this TEP, we solved the first crystal structures from the FAM83 family, including FAM83A and FAM83B. The structures of the DUF1669 domain reveal a phospholipase D-like fold lacking conservation of key catalytic residues. We deorphanise the FAM83 DUF1669 domain as a critical docking scaffold for binding of casein kinase 1 isoforms. Finally, using XChem fragment screening we report chemical fragments that bind to Lys230 in the central pocket of the DUF1669 and form starting points for potential drug development.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

common function paralog pairs and sequence similarity features

<p>This repository contains datasets of selected paralog pairs (Ensembl 111), labeled with various "common function" annotations, including PPI, SL, and GO datasets for both human (<em>Homo sapiens</em>) and budding yeast (<em>Saccharomyces cerevisiae</em>). These paralog pairs are characterized using different sequence similarity features, such as AlphaFold-predicted structures, Protein Language Model embeddings, and similarity searches from various databases.</p> <p>&nbsp;</p> <p>These datasets are used in the following manuscript:&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2024.10.11.617835v1">Evaluating Sequence and Structural Similarity Metrics for Predicting Shared Paralog Functions</a></p> <p>&nbsp;</p> <p>For the corresponding analysis notebooks, see:&nbsp;<a href="https://github.com/cancergenetics/paralog_seq_similarity/tree/main">github.com/cancergenetics/paralog_seq_similarity/</a></p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Sequence Similarity Network (SSN) and Genome Neighbourhood Network (GNN) for Mycobacterium Cytochrome P450 enzymes

<p>This dataset was generated in the context of the Horizon 2020&nbsp;MSCA IF action deCrYPtion (Grant 839116). The aim of this project is to use comparative genomics in order to propose and then test the function of uncharacterised Cytochrome P450 enzymes that are present among Mycobacterium species.</p> <p>More information about this project can be found at:&nbsp;https://cordis.europa.eu/project/id/839116.</p> <p>This dataset contains:</p> <p>- The FASTA sequences files obtained from the UniProt database, for members of the PF00067 protein family (CYP).</p> <p>- A set of reference FASTA sequences, matching the supplementary material from the following publication:&nbsp;Parvez, M.&nbsp;<em>et al.</em>&nbsp;(2016) &lsquo;Molecular evolutionary dynamics of cytochrome P450 monooxygenases across kingdoms: Special focus on mycobacterial P450s&rsquo;,&nbsp;<em>Scientific Reports</em>, 6(1), p. 33099. doi:<a href="https://doi.org/10.1038/srep33099">10.1038/srep33099</a>.</p> <p>- A combined FASTA files of both previously described, that was used for the generation of SSNs</p> <p>- A PNG&nbsp;image&nbsp;produced from the analysis of the&nbsp;Sequence Similarity Networks generated at AST78 (corresponding to 40% identity, defining CYP families)</p> <p>- A PNG&nbsp;image&nbsp;produced from the analysis of the&nbsp;Sequence Similarity Networks generated at AST141&nbsp;(corresponding to 55% identity, defining CYP subfamilies)</p> <p>- A Cytoscape session for&nbsp;the&nbsp;Sequence Similarity Networks from the combined FASTA file&nbsp;generated using the Enzyme Function Initiative web tools (https://efi.igb.illinois.edu), at AST78</p> <p>- A Cytoscape session containing&nbsp;Sequence Similarity Networks and&nbsp;Genome Neighborhood Network from the combined FASTA file&nbsp;&nbsp;generated using the Enzyme Function Initiative web tools (https://efi.igb.illinois.edu), at AST141</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Supplement data for the article "The impact of OTU sequence similarity threshold on diatom-based bioassessment: A case study of the rivers of Mayotte (France, Indian Ocean)", in preparation

<p>These are supplement data for the article &quot;The impact of OTU sequence similarity threshold on diatom-based bioassessment: A case study of the rivers of Mayotte (France, Indian Ocean)&quot;, in preparation</p> <p>The folowing files are available:</p> <ul> <li>Supplement 1. Map of Mayotte with the sampling sites and the rivers.</li> <li>Supplement 2. <em>rbcL</em> primers, reaction mixture, and conditions used for the PCR of the 312-bp <em>rbcL</em> fragment. The information provided is for a single reaction with a final volume of 25&micro;L.</li> <li>Supplement 3. The 20 fastq files containing the demultiplexed DNA reads.</li> <li>Supplement 4. Number of sequence reads for each sample before and after the trimming procedure.</li> <li>Supplement 5. The 20 OTU lists, corresponding to the 20 SSTs, including the number of DNA reads within the 90 samples and their assigned taxonomy.</li> <li>Supplement 6. Sampling site description with sample codes, names of rivers, year, number of raw DNA reads and GPS coordinates.</li> <li>Supplement 7. Values and summary statistics for the environmental variables.</li> <li>Supplement 8. The script used in Mothur for the bioinformatic analysis from trimming to the used OTU lists.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2018View details →
dryad40/100

Optimal sequence similarity thresholds for clustering of molecular operational taxonomic units in DNA metabarcoding studies

<p><span>Clustering approaches are pivotal to handle the many sequence variants obtained in DNA metabarcoding datasets, therefore they have become a key step of metabarcoding analysis pipelines. Clustering often relies on a sequence similarity threshold to gather sequences in Molecular Operational Taxonomic Units (MOTUs), each of which ideally representing a homogeneous taxonomic entity, e.g. a species or a genus. However, the choice of the clustering threshold is rarely justified, and its impact on MOTU over-splitting or over-merging even less tested. Here, we evaluated clustering threshold values for several metabarcoding markers under different criteria: limitation of MOTU over-merging, limitation of MOTU over-splitting, and trade-off between over-merging and over-splitting. We extracted sequences from a public database for nine markers, ranging from generalist markers targeting Bacteria or Eukaryota, to more specific markers targeting a class or a subclass (e.g. Insecta, Oligochaeta). Based on the distributions of pairwise sequence similarities within species and within genera, and on the rates of over-splitting and over-merging across different clustering thresholds, we were able to propose threshold values minimizing the risk of over-splitting, that of over-merging, or offering a trade-off between the two risks. For generalist markers, high similarity thresholds (0.96-0.99) are generally appropriate, while more specific markers require lower values (0.85-0.96). These results do not support the use of a fixed clustering threshold. Instead, we advocate a careful examination of the most appropriate threshold based on the research objectives, the potential costs of over-splitting and over-merging, and the features of the studied markers.</span></p>

opencc-zeroOct 2021View details →
dryad40/100

Optimal sequence similarity thresholds for clustering of molecular operational taxonomic units in DNA metabarcoding studies

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad36/100

Molecular sequencing and morphological identification reveal similar patterns in native bee communities across private and public grasslands of eastern North Dakota

<p>Bees play a key role in the functioning of human-modified and natural ecosystems by pollinating agricultural crops and wild plant communities. Global pollinator conservation efforts need large-scale and long-term monitoring to detect changes in species' demographic patterns and shifts in bee community structure. The objective of this project was to test a molecular sequencing pipeline that would utilize a commonly used locus, produce accurate and precise identifications consistent with morphological identifications, and generate data that are both qualitative and quantitative. We applied this amplicon sequencing pipeline to native bee communities sampled across Conservation Reserve Program (CRP) lands and native grasslands in eastern North Dakota. We found the 28S LSU locus to be more capable of discriminating between species than the 18S SSU rRNA locus, and in some cases even resolved instances of cryptic species or morphologically ambiguous species complexes. Overall, we found the amplicon sequencing method to be a qualitatively accurate representation of the sampled bee community richness and species identity, especially when a well-curated database of known 28S LSU sequences is available. Both morphological identification and molecular sequencing revealed similar patterns in native bee community structure across CRP lands and native prairie. Additionally, a genetic algorithm approach to compute taxon-specific correction factors using a small subset of the most concordant samples demonstrated that a high level of quantitative accuracy could be possible if the specimens are fresh and processed soon after collection. Here we provide a first step to a molecular pipeline for identifying insect pollinator communities. This tool should prove useful for future national monitoring efforts as use of molecular tools becomes more affordable and as numbers of 28S LSU sequences for pollinator species increase in publicly-available databases.</p>

opencc-zeroDec 2019View details →
zenodo36/100

Registration of multi-view echocardiography sequences using a subspace similarity measure

<p>Data employed for the validation of the PCA-based similarity metric proposed in &quot;<em>Registration of multi-view echocardiography sequences using a subspace similarity measure.</em>&quot; Peressutti <em>et al.</em> (under review).</p> <p>Data consists of echocardiography sequences of the Left ventricle of four volunteers (vol_A-vol_D) from different acoustic windows (aw_1-aw_5). The image format is metadata.<br /> For each subject, the ground-truth rigid transformations that aligns each sequence to all the others is provided. Such transformations&nbsp;are provided by optically tracking the position of the ultrasound imaging probe. &nbsp;</p> <p>For more detailed information regarding the datasets, please refer to the paper.</p> <p>Python code for testing the proposed method can be downloaded from (https://github.com/devisperessutti/Python.git), while MATLAB code can be downloaded from&nbsp;&nbsp;(https://github.com/gomezalberto/Matlab.git).</p>

opencc-zeroSep 2015View details →
zenodo36/100

NSRC-Search: Efficient searching for similar protein sequences of non-standard amino acid composition

<p>This research was funded by the National Science Centre in Poland (grant number 2021/41/N/ST6/01919)</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Do we similarly assess diversity with microscopy and High Throughput Sequencing? Case of microalgae in lakes.

<p>These are the repository files of the paper &quot;Do we similarly assess diversity with microscopy and High Throughput Sequencing? Case of microalgae in lakes&quot; published in Organisms Diversity and Evolution</p> <p>In these files are given:</p> <p>- the sampling sites descrptions (coordinates, lake names)</p> <p>- species relative abundances obtained with light microscopy for each sampling site</p> <p>- OTUs amounts and relative abundances obtained High-Throughput Sequencing for each sampling site</p> <p>- code correspondence between lake codes and FastQ files codes</p> <p>- FastQ files of the sampling sites</p>

opencc-by-4.0Jan 2018View details →
dryad36/100

Molecular sequencing and morphological identification reveal similar patterns in native bee communities across public and private grasslands of eastern North Dakota

Open the record for dataset details and reuse information.

publicJan 2020View details →
zenodo32/100

Heme oxygenase pfam14518 sequence similarity network

<p>Sequence similarity network generated from pfam14518 using the EFI-GNT webtool. Each node is representative of sequences that are 80% identical.&nbsp;&nbsp;</p>

opencc-by-4.0Nov 2021View details →
dryad32/100

Amplicon_sorter: a tool for reference-free amplicon sorting based on sequence similarity and for building consensus sequences

<p>Oxford Nanopore Technologies (ONT) is a third-generation sequencing technology that is gaining popularity in ecological research for its portable and low-cost sequencing possibilities. Although the technology excels at long-read sequencing, it can also be applied to sequence amplicons. The downside of ONT is the low quality of the raw reads. Hence, generating a high-quality consensus sequence is still a challenge. We present Amplicon_sorter, a tool for reference-free sorting of ONT sequenced amplicons based on their similarity in sequence and length and for building solid consensus sequences.</p>

opencc-zeroMay 2022View details →
zenodo32/100

Supplementary material 3 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265

Common OTU lists and Statistical analysis : Explanation note: This file contains detected OTUs in both methods and biodiversity analysis (GLM and t-test)

opencc-by-4.0Feb 2017View details →
zenodo32/100

Supplementary material 2 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265

Biodiversity workflow in R : Explanation note: Bundle of files for biodiversity analysis in R. All necessary input files and a commented script of R-commands are provided.

opencc-by-4.0Feb 2017View details →
zenodo32/100

Supplementary material 4 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265

Master data sheet : Explanation note: Spreadsheet file containing information about read abundances of operational taxonomic units (OTUs) and sample metadata. Here, data were prepared for subsequent biodiversity analysis in R.

opencc-by-4.0Feb 2017View details →
zenodo32/100

Supplementary material 1 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265

Bioinformatics pipeline : Explanation note: This file provides all steps and commands necessary for quality filtering and demultiplexing of raw paired fastq sequences.

opencc-by-4.0Feb 2017View details →
zenodo32/100

Fig. 3 in Do we similarly assess diversity with microscopy and high-throughput sequencing? Case of microalgae in lakes

Fig. 3 Correlation between the samples positions obtained on the first axes of the PCA based on microscopy and HTS diatom composition of the samples. There is a highly significant correlation (p &lt;0.001, R 2 = 31%) between both axes

opennotspecifiedFeb 2018View details →
zenodo32/100

Fig. 4 in Do we similarly assess diversity with microscopy and high-throughput sequencing? Case of microalgae in lakes

Fig. 4 Comparison of the diatom assemblages heterogeneity inside each lake, obtained with HTS and microscopy. Inside lake assemblage heterogeneity is the sum of the Bray-Curtis distances between the three samples of a lake. Correlation is significant (Pearson correlation p &lt;0.001) and follows a linear model (p &lt;0.001, R 2 = 50.8%) (see black line)

opennotspecifiedFeb 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record