Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

59

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

59 results for “sequence databases”

Learn how ShareScore rates datasets ↗
zenodo36/100

Nucleotide sequence database of Copper-containing membrane monooxygenases genes for analysing primer pairs targeting the ammonia monooxygenase subunit A gene of complete ammonia oxidising Nitrospira

<p>Nucleotide sequences of 487 Cu-mmo genes, including amoA comammox clade A and clade B, amoA ammonia oxidizing bacteria as well as other Cu-mmo genes.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Analysis of public single-cell sequencing database of COVID lung samples

<p>Lung endothelial cells from three published scRNA-seq datasets (GSE122960, GSE149878, GSE171668) of healthy subjects and COVID-19 patients were collected for further integrative analyses. The endothelial cells were classified into three sub-groups according to their distinguished expression of IL7R, DKK2, and EDNRB. For differential analysis of gene expression, counts per million of aggregated UMIs in each group were adopted in Wilcoxon rank-sum test.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Literature consistency of bioinformatics sequence databases is effective for assessing record quality

<p>Bioinformatics sequence databases such as Genbank or UniProt contain hundreds of millions of records of genomic data. These records are derived from direct submissions from individual laboratories, as well as from bulk submissions from large-scale sequencing centres; their diversity and scale means that they suffer from a range of data quality issues including errors, discrepancies, redundancies, ambiguities, incompleteness and inconsistencies with the published literature. In this work, we seek to investigate and analyze the data quality of sequence databases from the perspective of a curator, who must detect anomalous and suspicious records. Specifically, we emphasize the detection of inconsistent records with respect to the literature. Focusing on GenBank, we propose a set of 24 quality indicators, which are based on treating a record as a query into the published literature, and then use query quality predictors. We then carry out an analysis that shows that the proposed quality indicators and the quality of the records have a mutual relationship, in which one depends on the other. We propose to represent record literature consistency as a vector of these quality indicators. By reducing the dimensionality of this representation for visualization purposes using principal component analysis, we show that records which have been reported as inconsistent with the literature fall roughly in the same area, and therefore share similar characteristics. By manually analyzing records not previously known to be erroneous that fall in the same area than records know to be inconsistent, we show that one record out of four is inconsistent with respect to the literature. This high density of inconsistent record opens the way towards the development of automatic methods for the detection of faulty records. We conclude that literature inconsistency is a meaningful strategy for identifying suspicious records.</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

The databases of linear-recursive and "core" sequences

<p>These two files are <em>linres</em> and <em>core</em> databases of integer sequences related to my paper where I apply the algorithm <em>Diofantos</em> on them. They are also included in the related source code repository with the DOI:</p> <div> <pre>https://doi.org/10.5281/zenodo.13692310<br><br>I created this repository separate from the program code to make it easier to find.<br><br><em>linres</em> : linear_database_newbl.csv<br><em>core</em> : cores_test.csv</pre> </div>

openmit-licenseSep 2024View details →
zenodo36/100

Data mining antibody sequences for database searching in bottom-up proteomics

<p>Mass spectrometry (MS)-based proteomics is a powerful method for identifying and quantifying antibodies. Among the various MS approaches, bottom-up proteomics is especially effective for analyzing thousands of antibodies in complex mixtures. In this method, proteins are enzymatically digested into smaller peptides, typically using the protease trypsin, which are then analyzed via mass spectrometry. These peptides are matched to sequences in standard databases like UniProt or NCBI-RefSeq for identification.</p> <p>However, a major limitation of this approach is the absence of comprehensive disease-specific antibody databases. Current databases, such as UniProt, include only a fraction of the antibody sequences present in the human body. For instance, as of January 2024, UniProt contains just 38,800 immunoglobulin sequences, far short of the billions of antibodies the human immune system can produce. As a result, relying on such limited databases can lead to under-detection of antibodies, particularly those associated with specific diseases. Expanding antibody databases with disease-specific sequences is crucial for improving the accuracy of MS-based proteomics in identifying antibodies relevant to human health.</p> <p>Recently, through next-generation sequencing of antibody gene repertoires, it has become possible to obtain billions of antibody sequences (in amino acid format) by annotating, translating, and numbering antibody gene sequences. These large numbers of sequences are now available in public databases such as the&nbsp;<a href="https://opig.stats.ox.ac.uk/webapps/oas/" rel="nofollow">Observed Antibody Space</a>. We hypothesize that using these theoretical antibody sequences as new databases for bottom-up proteomics could address the current lack of antibody coverage in standard databases.</p> <p>We developed a workflow to create disease-specific antibody peptide databases for bottom-up proteomics. The workflow details are available on <a href="https://github.com/trinhxt/SDU_Immunoinformatics">GitHub</a>. The database and metadata files generated by this workflow are stored in this Zenodo dataset, and they are used in DAT-DB &mdash; a web application that allows researchers to obtain FASTA files of disease-specific antibody peptides for direct use in bottom-up proteomics (see <a href="https://trinhxt.shinyapps.io/DAT-DB/">Demo version</a>).</p> <p>Each database file in this dataset is in <em>.duckdb</em> format and contains tables with 10 columns: Sequence, Filename, Patient, BSource, BType, Isotype, N_patient, N_antibody, Length_aa, and CDR3. The "<strong>Sequence</strong>" column contains tryptic peptides. "<strong>Filename</strong>" is the file where the data was collected. "<strong>Patient</strong>" refers to the patient number as listed in <em>metadata2.csv</em>. "<strong>BSource</strong>" refers to the B-cells' source, and "<strong>BType</strong>" refers to the type of B-cells. "<strong>Isotype</strong>" specifies the antibody isotype (IgA, IgD, IgE, IgG, IgM, or Bulk). "<strong>N_patient</strong>" indicates the number of patients having this peptide, and "<strong>N_antibody</strong>" specifies the number of antibodies containing this peptide. "<strong>Length_aa</strong>" indicates the number of amino acids in the peptide, while "<strong>CDR3</strong>" shows whether the peptide is found in the CDR3 region.</p> <p>The file <em>metadata1.csv</em> contains information about each database file, while <em>metadata2.csv</em> provides details about the sources of the collected antibodies.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Database for mi-faser: Functional sequencing read annotation for high precision microbiome analysis

<p><strong>[Database for mi-faser]</strong></p> <p><strong>mi-faser:&nbsp;</strong><em>microbiome - functional annotation of sequencing reads</em></p> <p>A super-fast ( &lt; 20min/10GB of reads ) and accurate ( &gt; 90% precision ) method for annotation of molecular functionality encoded in sequencing read data without the need for assembly or gene finding.</p> <p>Web Service:&nbsp;http://services.bromberglab.org/mifaser/|<br> Repository:&nbsp;https://bitbucket.org/bromberglab/mifaser_base/</p>

opennosl3.0Nov 2017View details →
dryad32/100

Data from: Exploitation of a turbot (Scophthalmus maximus L.) immune-related expressed sequence tag (EST) database for microsatellite screening and validation

In this study, we identified and characterized 160 microsatellite loci from an expressed sequence tag (EST) database generated from immune-related organs of turbot (Scophthalmus maximus). A final set of 83 new polymorphic microsatellites were validated after the analysis of 40 individuals from Atlantic origin including both wild and farmed individuals. The allele number and the expected heterozygosity ranged from 2 to 18 and from 0.021 to 0.951, respectively. Evidences of null alleles at moderate-high frequencies were detected at six loci using population data. None of the analyzed loci showed deviations from Mendelian segregation after analysis of five full-sib families including ~92 individuals/family. The markers are used to consolidate the turbot genetic map and, since they are mostly EST-derived, they will be very useful for comparative genomic studies within flatfishes and with model fish species. Using an in silico approach, we detected significant homologies of microsatellite sequences with the EST databases of the flatfish species with highest genomic resources (Senegalese sole, Atlantic halibut, bastard halibut) at 31% of these turbot markers. The conservation of these microsatellites within Pleuronectiformes will pave the way for anchoring genetic maps of different species and identifying genomic regions related to productive traits.

opencc-zeroDec 2011View details →
dryad32/100

Data from: A public database of memory and naive B-cell receptor sequences

The vast diversity of B-cell receptors (BCR) and secreted antibodies enables the recognition of, and response to, a wide range of epitopes, but this diversity has also limited our understanding of humoral immunity. We present a public database of more than 37 million unique BCR sequences from three healthy adult donors that is many fold deeper than any existing resource, together with a set of online tools designed to facilitate the visualization and analysis of the annotated data. We estimate the clonal diversity of the naive and memory B-cell repertoires of healthy individuals, and provide a set of examples that illustrate the utility of the database, including several views of the basic properties of immunoglobulin heavy chain sequences, such as rearrangement length, subunit usage, and somatic hypermutation positions and dynamics.

opencc-zeroDec 2015View details →
zenodo32/100

16S rRNA sequence database

<p>&nbsp;This 16S rRNA database was curated for honeybee gut microbiome.</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7). in An integrative taxonomic revision of the genus Triphosa Stephens, 1829 (Geometridae: Larentiinae) in the Middle East and Central Asia, with description of two new species

APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7).

opennotspecifiedMay 2019View details →
zenodo32/100

SQL database with detailed characteristic measurement results of photovoltaic modules that were subjected to accelerated aging sequences

<p>All measurement results are organized in an optimized database, which forms the information base for setting up models for climate sensitive ageing and degradation processes/mechanisms. The database is structured around the modules (module = device under test), see general scheme given in <a href="https://doi.org/10.1002/pip.3090">https://doi.org/10.1002/pip.3090</a>. Modules are logically connected via their specific ageing module groups with strictly associated ageing actions and instances. As stated above, a set of three identical modules is stored together in each specific ageing action (= accelerated ageing test as described in detail in Table&nbsp;<a title="Link to table" href="https://onlinelibrary.wiley.com/doi/10.1002/pip.3090#pip3090-tbl-0001">1</a> of <a href="https://doi.org/10.1002/pip.3090">https://doi.org/10.1002/pip.3090</a>) in order to increase the statistical reliability. Those triples are logically grouped in the database with the corresponding acquired measurement results being canonicalized and stored in the database as well. For future applications (modelling), all measurement information is kept as complete as possible; aggregation is avoided.</p> <a href="https://onlinelibrary.wiley.com/cms/asset/d5235350-b448-4094-b022-3b572d4c9317/pip3090-fig-0001-m.jpg" target="_blank" rel="noopener"></a>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Virus sequences associated with the rodent virome database for the years 2014 and 2016/2017.

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

In-house sequence database of S.pn recombinant protein vaccine target-related genes

<p>In-house sequence database of S.pn recombinant protein vaccine target-related genes.</p>

opencc-by-4.0Jul 2021View details →
dryad32/100

Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences

<p>A fasta-formatted database of 36,866,870 predicted proteins representing 4,351 unique species from 117 phyla.</p>

opencc-zeroFeb 2023View details →
zenodo32/100

ML DNA sequencing database

<p>The machine learning technique is utilized with the transmission fingerprint database for the MoS<sub>2</sub> nanochannel-assisted identification and classification of DNA nucleotides.</p>

opencc-by-4.0Jun 2023View details →
dryad32/100

Data from: A public database of memory and naive B-cell receptor sequences

Open the record for dataset details and reuse information.

publicAug 2017View details →
dryad32/100

Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad32/100

Data from: Exploitation of a turbot (Scophthalmus maximus L.) immune-related expressed sequence tag (EST) database for microsatellite screening and validation

Open the record for dataset details and reuse information.

publicJan 2012View details →
zenodo28/100

Evaluation SIHUMI dataset: A sectioning and database enrichment approach for improved peptide spectrum matching in large, genome-guided protein sequence databases

<p>Dataset for&nbsp;evaluation of&nbsp;database sectioning method for generating an enriched database for mass-spectrometry-based proteomic approaches&nbsp;using large databases. Our evaluation demonstrates that this method helps to&nbsp;increase the sensitivity of PSMs while maintaining acceptable FDR statistics. This dataset was acquired from the protein samples containing proteins from eight microorganisms (<em>Anaerostipes caccae, Bacteroides thetaiotaomicron, Bifidobacterium longum, Blautia producta, Clostridium butyricum, Clostridium ramosum, Escherichia coli, Lactobacillus plantarum</em>) that were grown in a bioreactor. The MS-data for this dataset was acquired by Dr. Hettich&#39;s group at Oak Ridge National Laboratory, and it was made available by Dr. Nico Jehmlich from&nbsp;Helmholtz Center for Environmental Research through&nbsp;the 3rd International Metaproteome Symposium (<a href="https://www.ufz.de/index.php?en=44639">https://www.ufz.de/index.php?en=44639</a>).&nbsp;</p>

opencc-by-4.0Apr 2020View details →
dryad28/100

Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants

Background: Simple Sequence Repeats (SSRs) are widely used in population genetic studies but their classical development is costly and time-consuming. The ever-increasing available DNA datasets generated by high-throughput techniques offer an inexpensive alternative for SSRs discovery. Expressed Sequence Tags (ESTs) have been widely used as SSR source for plants of economic relevance but their application to non-model species is still modest. Methods: Here, we explored the use of publicly available ESTs (GenBank at the National Center for Biotechnology Information-NCBI) for SSRs development in non-model plants, focusing on genera listed by the International Union for the Conservation of Nature (IUCN). We also search two model genera with fully annotated genomes for EST-SSRs, Arabidopsis and Oryza, and used them as controls for genome distribution analyses. Overall, we downloaded 16 031 555 sequences for 258 plant genera which were mined for SSRsand their primers with the help of QDD1. Genome distribution analyses in Oryza and Arabidopsis were done by blasting the sequences with SSR against the Oryza sativa and Arabidopsis thaliana reference genomes implemented in the Basal Local Alignment Tool (BLAST) of the NCBI website. Finally, we performed an empirical test to determine the performance of our EST-SSRs in a few individuals from four species of two eudicot genera, Trifolium and Centaurea. Results: We explored a total of 14 498 726 EST sequences from the dbEST database (NCBI) in 257 plant genera from the IUCN Red List. We identify a very large number (17 102) of ready-to-test EST-SSRs in most plant genera (193) at no cost. Overall, dinucleotide and trinucleotide repeats were the prevalent types but the abundance of the various types of repeat differed between taxonomic groups. Control genomes revealed that trinucleotide repeats were mostly located in coding regions while dinucleotide repeats were largely associated with untranslated regions. Our results from the empirical test revealed considerable amplification success and transferability between congenerics. Conclusions: The present work represents the first large-scale study developing SSRs by utilizing publicly accessible EST databases in threatened plants. Here we provide a very large number of ready-to-test EST-SSR (17 102) for 193 genera. The cross-species transferability suggests that the number of possible target species would be large. Since trinucleotide repeats are abundant and mainly linked to exons they might be useful in evolutionary and conservation studies. Altogether, our study highly supports the use of EST databases as an extremely affordable and fast alternative for SSR developing in threatened plants.

opencc-zeroDec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record