Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

61

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

61 results for “sequence similarity”

Learn how ShareScore rates datasets ↗
zenodo28/100

Figure 1 from: Siddique AB, Khokon AM, Unterseher M (2017) What do we learn from cultures in the omics age? High-throughput sequencing and cultivation of leaf-inhabiting endophytes from beech (Fagus sylvatica L.) revealed complementary community composition but similar correlations with local habitat conditions. MycoKeys 20: 1-16. https://doi.org/10.3897/mycokeys.20.11265

Figure 1 - Diversity indexes and accumulation curves for a Illumina and b cultivation data of fungal leaf-inhabiting endophytes of beech. Except of the accumulation curves of cultivation data, both methods revealed a clear and partly significant trend of higher fungal diversity at the valley site.

opencc-by-4.0Feb 2017View details →
zenodo28/100

Fig. 1 in Do we similarly assess diversity with microscopy and high-throughput sequencing? Case of microalgae in lakes

Fig. 1 Location of the sampled lakes in the French Northern Alps

opennotspecifiedFeb 2018View details →
dryad28/100

Data from: Sequencing of the needle transcriptome from Norway spruce (Picea abies Karst L.) reveals lower substitution rates, but similar selective constraints in gymnosperms and angiosperms

BACKGROUND: A detailed knowledge about spatial and temporal gene expression is important for understanding both the function of genes and their evolution. For the vast majority of species, transcriptomes are still largely uncharacterized and even in those where substantial information is available it is often in the form of partially sequenced transcriptomes. With the development of next generation sequencing, a single experiment can now simultaneously identify the transcribed part of a species genome and estimate levels of gene expression. RESULTS: mRNA from actively growing needles of Norway spruce (Picea abies) was sequenced using next generation sequencing technology. In total, close to 70 million fragments with a length of 76 bp were sequenced resulting in 5 Gbp of raw data. A de novo assembly of these reads, together with publicly available expressed sequence tag (EST) data from Norway spruce, was used to create a reference transcriptome. Of the 38,419 PUTs (putative unique transcripts) longer than 150 bp in this reference assembly, 83.5% show similarity to ESTs from other spruce species and of the remaining PUTs, 3,704 show similarity to protein sequences from other plant species, leaving 4,167 PUTs with limited similarity to currently available plant proteins. By predicting coding frames and comparing not only the Norway spruce PUTs, but also PUTs from the close relatives Picea glauca and Picea sitchensis to both Pinus taeda and Taxus mairei, we obtained estimates of synonymous and non-synonymous divergence among conifer species. In addition, we detected close to 15,000 SNPs of high quality and estimated gene expression differences between samples collected under dark and light conditions. CONCLUSIONS: Our study yielded a large number of single nucleotide polymorphisms as well as estimates of gene expression on transcriptome scale. In agreement with a recent study we find that the synonymous substitution rate per year (0.6 x 10-09 and 1.1 x 10-09) is an order of magnitude smaller than values reported for angiosperm herbs. However, if one takes generation time into account, most of this difference disappears. The estimates of the dN/dS ratio (non-synonymous over synonymous divergence) reported here are in general much lower than 1 and only a few genes showed a ratio larger than 1.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Genetic barcoding of dark-spored myxomycetes (Amoebozoa)—Identification, evaluation and application of a sequence similarity threshold for species differentiation in NGS studies

Unicellular, eukaryotic organisms (protists) play a key role in soil food webs as major predators of microorganisms. However, due to the polyphyletic nature of protists, no single universal barcode can be established for this group, and the structure of many protistean communities remains unresolved. Plasmodial slime moulds (Myxogastria or Myxomycetes) stand out among protists by their formation of fruit bodies, which allow for a morphological species concept. By Sanger sequencing of a large collection of morphospecies, this study presents the largest database to date of dark-spored myxomycetes and evaluate a partial 18S SSU gene marker for species annotation. We identify and discuss the use of an intraspecific sequence similarity threshold of 99.1% for species differentiation (OTU picking) in environmental PCR studies (ePCR) and estimate a hidden diversity of putative species, exceeding those of described morphospecies by 99%. When applying the identified threshold to an ePCR data set (including sequences from both NGS and cloning), we find 64 OTUs of which 21.9% had a direct match (>99.1% similarity) to the database and the remaining had on average 90.2 ± 0.8% similarity to their best match, thus thought to represent undiscovered diversity of dark-spored myxomycetes.

opencc-zeroDec 2016View details →
zenodo28/100

Metadata supporting the AFDB90v4 annotated sequence similarity network

<p>Driven by the development and upscaling of fast genome sequencing and assembly pipelines, the number of protein-coding sequences deposited in public protein sequence databases is increasing exponentially. Recently, the dramatic success of deep learning-based approaches applied to protein structure prediction has done the same for protein structures. We are now entering a new era in protein sequence and structure annotation, with hundreds of millions of predicted protein structures made available through the AlphaFold database. These models cover most of the catalogued natural proteins, including those difficult to annotate for function or putative biological role based on standard, homology-based approaches. In this work, we quantified how much of such &quot;dark matter&quot; of the natural protein universe was structurally illuminated by AlphaFold2 at a high predicted accuracy and modelled this diversity as an interactive sequence similarity network that can be navigated at <a href="https://uniprot3d.org/atlas/AFDB90v4">https://uniprot3d.org/atlas/AFDB90v4</a>. The dataset deposited here corresponds to the metadata generated, and that makes the base of the similarity network constructed and its interpretation. These files are either generated or processed using the code available at <a href="https://github.com/ProteinUniverseAtlas/AFDB90v4">https://github.com/ProteinUniverseAtlas/AFDB90v4</a>.</p> <p>This repository further contains the detailed, individual sequence similarity networks (in CLANS format) generated for the 3 example protein (super)families described in the text.</p> <p>The full content of this repository includes:</p> <ol> <li> <p><strong>AFDBv4_pLDDT_diggestion_UniRef50_2023-02-01.csv</strong>: table listing all uniref50 clusters in UniProt, including information on structural representatives from AFDB. Each column provides different annotations, including functional brightness, median and best pLDDT, brightness and structural representatives, etc.</p> </li> <li> <p><strong>AFDBv4_DUF_dark_diggestion_UniRef50_2023-02-06.csv</strong>: table listing all uniref50 clusters in UniProt and whether they include proteins mapped to known domains of unknown function (DUF).</p> </li> <li> <p><strong>AFDBv4_90.fasta</strong>: fasta file with the sequences of all UniRef50 clusters selected, and used for the all-against-all mmseqs searches that make the base of the network.</p> </li> <li> <p><strong>AFDB90v4_data.csv</strong>: the subset of file (1) that corresponds to the AFDB90v4 dataset, including columns such as functional brightness, median and best pLDDT, brightness and structural representatives, etc.</p> </li> <li> <p><strong>AFDB90v4_data_with_graph_labels.csv</strong>: table listing each individual uniref50 cluster included in the AFDB90v4 dataset, together with their mapping to communities, and connected components.</p> </li> <li> <p><strong>AFDB90v4_cc_data.csv</strong>: table of uniref50 clusters in connected components, including their annotations, and the columns in file (5).</p> </li> <li> <p><strong>AFDB90v4_cc_data_uniprot_community_taxonomy_map.csv</strong>: mapping of each uniprotAC entry to their corresponding component, community and taxonomy.</p> </li> <li> <p><strong>AFDB90v4_subgraphs_summary.csv</strong>: table summarising the properties of individual connected components, including the average brightness, the number of members, the number of unique protein sequences, the median length, and the number of communities.</p> </li> <li> <p><strong>communities_summary.csv</strong>: table summarising the properties of individual communities, including average brightness, the number of members, the number of unique protein sequences, the median length, the most common superkingdom represented, the average structure outlier score, etc.</p> </li> <li> <p><strong>communities_edge_list-coordinates.csv</strong>: the coordinates of each community in the graphical representation. Singleton communities or singleton UniRef50 clusters are not included.</p> </li> <li> <p><strong>communities_edge_list_no_duplicates.csv</strong>: list of edges making the graph.</p> </li> <li> <p><strong>node_class.json</strong>: map between each Uniref50 cluster and its corresponding component and community.</p> </li> <li> <p><strong>subgraphs.tar.gz</strong>: tar file containing gml files for each individual connected component.</p> </li> <li> <p><strong>AFDB90v4_outlier_scores.tsv</strong>: table containing the outlier scores for each community representative.</p> </li> <li> <p><strong>AFDB90v4_dark_galaxies_summary.csv</strong>: table containing the summary of all dark connected components, including average brightness, median length, representatives, number of communities, etc.</p> </li> <li> <p><strong>AFDB90v4_uniprot_naming_assessment_counts.csv</strong>: table listing the per-component semantic diversity scores, as well as the major source of the titles of the proteins included and their count.</p> </li> <li> <p><strong>uniprot_naming_assessment.tar.gz</strong>: tar file containing the per community assessment of predicted protein names in UniProt as of February 2023.</p> </li> <li> <p><strong>CLANS_files.tar.gz</strong>: stores the 3 sequence similarity networks, in CLANS format, constructed for the analysis of the sequence diversity and sequence similarities of the proteins in components, 27, 159 and 3314. These CLANS files make the base of panel A in all figures 3 and 4 and extended data figure 5.</p> </li> </ol> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
dryad28/100

Data from: Sequencing of the needle transcriptome from Norway spruce (Picea abies Karst L.) reveals lower substitution rates, but similar selective constraints in gymnosperms and angiosperms

Open the record for dataset details and reuse information.

publicNov 2012View details →
dryad28/100

Data from: Genetic barcoding of dark-spored myxomycetes (Amoebozoa)—Identification, evaluation and application of a sequence similarity threshold for species differentiation in NGS studies

Open the record for dataset details and reuse information.

publicOct 2017View details →
dryad28/100

Australian long-finned pilot whales (Globicephala melas) emit stereotypical, variable, biphonic, multi-component, and sequenced vocalisations, similar to those recorded in the northern hemisphere

Open the record for dataset details and reuse information.

publicSep 2020View details →
geo24/100

Sequence variation within similar cis-elements promotes context-specific functions of two Drosophila GA-binding transcription factors (protein binding microarray data set)

GEO Series GSE110653. Drosophila melanogaster. 2 samples. Type: Genome binding/occupancy profiling by array.

openGEO-OpenFeb 2018View details →
geo24/100

miRNA expression in family with sequence similarity 222 member B (Fam222B) knockdown B16 cells

GEO Series GSE289508. Mus musculus. 2 samples. Type: Non-coding RNA profiling by array.

openGEO-OpenAug 2025View details →
geo24/100

Single cell RNA and TCR sequencing reveals substantive similarities between synovial fluid and synovial tissue T-cells in inflammatory arthritis.

GEO Series GSE250242. Homo sapiens. 21 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2024View details →
geo24/100

The bZIP mutant CEBPB(V285A) has sequence specific DNA binding propensities similar to CREB1

GEO Series GSE112020. synthetic construct; Mus musculus. 31 samples. Type: Other.

openGEO-OpenMar 2019View details →
geo24/100

Sequence variation within similar cis-elements promotes context-specific functions of two Drosophila GA-binding transcription factors

GEO Series GSE110654. Drosophila melanogaster. 39 samples. Type: Genome binding/occupancy profiling by array; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →
geo24/100

The bZIP mutant CEBPB(V285A) has sequence specific DNA binding propensities similar to CREB1 [CEBPMut_5mCG]

GEO Series GSE112019. synthetic construct; Mus musculus. 10 samples. Type: Other.

openGEO-OpenMar 2019View details →
geo24/100

Illumina SBS sequencing and DNBSEQ perform similarly for single-cell transcriptomics

GEO Series GSE277740. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2025View details →
geo24/100

Similarity Regression predicts evolution of transcription factor sequence specificity

GEO Series GSE121420. synthetic construct. 682 samples. Type: Other.

openGEO-OpenMar 2019View details →
geo24/100

Variations in the expression pattern and sequence similarity of the key genes involved in the metabolism of taxoids among three Taxus species

GEO Series GSE121523. Taxus x media; Taxus mairei; Taxus cuspidata. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2018View details →
geo20/100

Sequence variation within similar cis-elements promotes context-specific functions of two Drosophila GA-binding transcription factors (ChIP-seq data set)

GEO Series GSE107059. Drosophila melanogaster. 37 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →
zenodo20/100

FIGURE. Phylogram of Tolypocladium generated from Maximum likelihood analysis of ITS, SSU and LSU sequence data. Purpureocillium lilacinum (CBS 284.36) was selected as an outgroup taxon. The tree topology of the ML analysis was similar to the BI. Maximum likelihood bootstrap values greater than 75 and Bayesian posterior probabilities over 0.90 were indicated above the nodes. The scale bar indicates 0.006 changes. The new species was in blue. in Yunnan-Guizhou Plateau: a mycological hotspot

FIGURE. Phylogram of Tolypocladium generated from Maximum likelihood analysis of ITS, SSU and LSU sequence data. Purpureocillium lilacinum (CBS 284.36) was selected as an outgroup taxon. The tree topology of the ML analysis was similar to the BI. Maximum likelihood bootstrap values greater than 75 and Bayesian posterior probabilities over 0.90 were indicated above the nodes. The scale bar indicates 0.006 changes. The new species was in blue.

opennotspecifiedOct 2021View details →
geo16/100

aCGH Detects Copy Number Variation with Similar Resolution to PacBio Sequencing Approaches

GEO Series GSE141976. Haplochromis burtoni; Neolamprologus brichardi; Pundamilia nyererei; Oreochromis niloticus; Maylandia zebra. 4 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenDec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record