Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
59
datasets available to search
ShareScore release 0.9.0
Dataset results
59 results for “Sequence databases”
Consortium of Long Read Sequencing Database (CoLoRSdb)
<p>Visit the <a href="https://colorsdb.org">CoLoRSdb website</a> for additional information.</p>
CircRNAs sequencing database of Gastric carcinoma cells (SGC-7901) and 5-fluorouracil-resistant cells (SGC-7901-5-fu)
<p>The circRNAs of two different gastric cancer cell lines were sequenced in this database. The sequencing results were compared and annotated with the database as analysis background data, and the screening conditions for differential expression of circRNA in the two cells line were defined as fold changes (FC) ≥ 2 and P <0.05.</p>
swissprot90_2022_03 COMER2 sequence profile database
<p>swissprot90_2022_03 COMER2 sequence profile database</p>
Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing - Klebsiella spp. Isolates for RASE Databases and Paired Isolates
Open the record for dataset details and reuse information.
Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing - E. coli Isolates for RASE Databases and Paired Isolates
Open the record for dataset details and reuse information.
Candidate antimicrobial peptide sequences mined from the UniProtKB/Swiss-Prot database using AMPlify
<p>Here we provide the datasets used in the study of AMP mining from the UniProtKB/Swiss-Prot database (2022_02 release), as well as the predicted AMP sequences by AMPlify v2.0.0. </p>
Reference sequence database for eDNA metabarcoding of San Francisco estuary fishes and invertebrates
<p>Environmental DNA (eDNA) methods complement traditional monitoring and can be configured to detect multiple species simultaneously. One such approach, eDNA metabarcoding, uses high-throughput DNA sequencing to indirectly detect many different organisms, spanning broad taxonomic boundaries, from water samples. We are optimizing a non-invasive, low cost eDNA metabarcoding protocol to be used in conjunction with existing monitoring programs. One resource that is currently lacking for metabarcoding studies in general, including those in the San Francisco Estuary (SFE), is a comprehensive database of DNA barcode reference sequences. Without this foundational data, many species go undetected or misidentified in metabarcoding studies. To meet this need, we generated a custom barcode sequence database for the SFE by DNA sequencing and mining of public DNA seqeunce data for estuarine and freshwater species of interest to monitoring programs and ecological studies. Here we present custom reference sequence databases for three barcodes: Cytochrome C Oxidase I (COI), 12S MiFish and 16S.</p>
Reference sequence database for eDNA metabarcoding of San Francisco estuary fishes and invertebrates
Open the record for dataset details and reuse information.
Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants
Open the record for dataset details and reuse information.
CircRNAs sequencing database of Gastric carcinoma cells (SGC-7901) and 5-fluorouracil-resistant cells (SGC-7901-5-fu)
Open the record for dataset details and reuse information.
Characterization of eQTLs associated with androstenone by RNA sequencing in porcine testis (RNASeq Database)
GEO Series GSE119474. Sus scrofa. 32 samples. Type: Expression profiling by high throughput sequencing.
YM500: An integrative small RNA sequencing (smRNA-seq) database for microRNA research
GEO Series GSE39841. Homo sapiens. 34 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Characterization of eQTLs associated with androstenone by RNA sequencing in porcine testis - RNASeq Database
GEO Series GSE114225. Sus scrofa. 56 samples. Type: Expression profiling by high throughput sequencing.
RNA-sequencing of 69 cell lines in the Human Protein Atlas database
GEO Series GSE240542. Homo sapiens. 141 samples. Type: Expression profiling by high throughput sequencing.
Characterization of eQTLs associated with androstenone by RNA sequencing in porcine testis - SNP Genotype Database
GEO Series GSE114518. Sus scrofa. 48 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.
Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing - RASE Databases for E. coli and Klebsiella spp
<p>RASE databases used for the prediction of antibiotic phenotype in the paper titled "Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing". A database for each of <em>Escherichia coli </em>and <em>Klebsiella spp.</em> are available, including the original metadata and genetic trees used to create the databases.</p>
MRI Guided Radiotherapy and Radiobiological Data: the ISRAR Database (Irm Sequences for Radiobiological Adaptative Radiotherapy)
ClinicalTrials.gov study NCT06041555. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Single-cell sequencing analysis and weighted co-expression network analysis based on public databases identified that TNC is a novel biomarker for keloid
GEO Series GSE190626. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Database: results of the genotyping of the population of recombinant lines TU/MUSICA (TUM) by means of Genotyping by Sequencing
<p>The database contains the files:</p> <p>- imagen with a seed per recombinant line as well as seed of the two parents</p> <p><strong><em> Figure1_TUM RIL population.jpg</em></strong></p> <p>- imagen of image of the resulting genetic map</p> <p><strong><em> Figure 3.pdf</em></strong></p> <p>- file with the GBS results (95678 SNPs)</p> <p><em><strong>all.filter_ordered.vcf</strong></em></p> <p>- file with the SNPs included in the genetic map published (842 loci/ 175 lines)</p> <p><strong><em>GENETIC MAP_TUM.xlsx</em></strong></p> <p>Genotyping was performed according to Garcia Fernández et al (2021)</p> <p><strong>Reference</strong>: García-Fernández, C., Campa, A. & Ferreira, J.J. Dissecting the genetic control of seed coat color in a RIL population of common bean (<em>Phaseolus vulgaris</em> L.). <em>Theor Appl Genet</em> <strong>134, </strong>3687–3698 (2021). https://doi.org/10.1007/s00122-021-03922-y</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.