Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

96

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

96 results for “nucleotide sequences”

Learn how ShareScore rates datasets ↗
zenodo44/100

Control T-cell receptor (TCR) alpha and beta chain nucleotide and amino acid sequences from human and mouse

<p>A dataset of pooled T-cell receptor (TCR) sequences for TCR alpha and beta chains of human and mouse.</p> <p>Sequences are obtained from various samples of healthy individuals/mice using our conventional protocols:&nbsp;see for example [Britanova et al &quot;Dynamics of individual T cell repertoires: from cord blood to centenarians&quot;&nbsp;The Journal of Immunology 2016] and [Izraelson et al. &quot;Comparative analysis of murine T‐cell receptor repertoires.&quot;&nbsp;Immunology 2018].</p> <p>The sequences are stored as gzipped clonotype tables in VDJtools format,&nbsp;see [https://vdjtools-doc.readthedocs.io/en/master/input.html#vdjtools-format].</p> <p>This control dataset can be used as a proxy for a generative VDJ rearrangement model to estimate the expected frequency distribution of TCRs and check for enrichment of rare TCR clonotypes and groups of similar TCR sequences. For the implementation of the enrichment analysis, please see CalcDegreeStats routine from VDJtools software, see [https://vdjtools-doc.readthedocs.io/en/master/annotate.html#calcdegreestats].</p> <p>Files named &quot;human.tra.strict.txt.gz&quot;, etc are pools of random/naive TCR clonotypes containing unique V/J/CDR3 nucleotide sequence combinations observed in data. The pools.zip file is used for TCR motif inference in VDJdb database [https://github.com/antigenomics/vdjdb-motifs], it contains human.tra.aa.txt, etc files that contain random/naive TCR clonotypes grouped by CDR3 amino acid sequence with the most frequent representative V and J.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Alignments and ML trees of cassava brown streak virus and Ugandana cassava brown streak virus polyprotein nucleotide sequences

<p>Alignments of full and nearly full polyprotein-length nucleotide sequences from GenBank for the two ipomoviruses that cause cassava brown streak disease, in fasta format.&nbsp; Separate alignments for 67 cassava brown streak virus sequences and 81 Ugandan cassava brown streak virus sequences are provided, as well as a combined alignment of 148 sequences.&nbsp; Alignments were created with MUSCLE and then modified by eye in AliView.</p> <p>Also, two tree files (in nexus) format are supplied, resulting from a maximum likelihood analysis with IQTree on each of the two single-species datasets. Support for nodes with aLRT and 100 actual bootstrap replicates are provided (aLRT/bootstrap).</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Fig. 3 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data

Fig. 3. Bayesian consensus tree for the anabantoids, channids and catfishes (silurids, bagrids, clariids) obtained using partial Cytochrome b sequences with cyprinids as outgroup. The heteronchocleidids genera present on the anabantoids and channids are shown with their geographical areas. Values shown at each node refer to Bayesian posterior probabilities. (*refer to Table 3 for names used in GenBank).

opencc-by-4.0Aug 2011View details →
zenodo40/100

Fig. 2. Bayesian consensus tree generated from partial 28S in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data

Fig. 2. Bayesian consensus tree generated from partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Values shown at each node refer to Bayesian (BI) posterior probabilities/maximum likelihood (ML) percentages of the bootstrap values with 100 replicates. Bootstrap values lower than 50 are given as dashes (-).

opencc-by-4.0Aug 2011View details →
zenodo40/100

Fig. 1 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data

Fig. 1. Neighbour joining (NJ) tree constructed by PAUP* using partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Percentages of the bootstrap values for neighbour joining (NJ)/maximum parsimony (MP) (NJ &amp; MP=1,000 replicates) are shown along the branches. Bootstrap values lower than 50 are given as dashes (-).

opencc-by-4.0Aug 2011View details →
zenodo40/100

MGBC-26640: nucleotide sequences for gene annotations

<p>Nucleotide sequences of annotated genes from the 26,640 high-quality, non-redundant genomes of the MGBC.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S in A herpetological survey of western Zambia

Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S rRNA of Phrynobatrachus natalensis. Numbers above branches are non-parametric bootstrap support values. Specimen vouchers or GenBank accession numbers are shown in parentheses. Colored polygons highlight the clades comprising specimens from this study. (*) Nearest sample from type locality of Phrynobatrachus natalensis; (**) Haplotype groups A and B in Zimkus and Schick (2010).

opencc-by-4.0Aug 2019View details →
zenodo40/100

Fig. 1 in Polymerase chain reaction and gyrA nucleotide sequence analysis of Wolbachia endosymbionts (Rickettsiales: Anaplasmataceae) in various species of Culicidae, Cimex lectularius (Hemiptera: Cimicidae) and Dirofilaria immitis (Rhabditida: Onchocercidae)

Fig. 1. Phylogenetic tree based on Maximum Likelihood depicting the grouping of Wolbachia from various hosts based on analysis of the gyrA gene. The numerical value displayed on branches is the bootstrap value (1,000 replicates), and branches with values below 50% are collapsed. The tree illustrates that gyrA sequences distinguish Wolbachia subtypes based on host taxonomy, demonstrating that this gene may contribute to Wolbachia strain typing projects and future phylogenetic analysis.

opencc-by-4.0Jan 2021View details →
zenodo40/100

RefSeq bacterial protein coding (nucleotide) sequences

<p><strong>Bacteria_Nucleotide.fas.gz</strong></p><p>151,835,459 protein coding (nucleotide) sequences extracted from 44,831 randomly selected bacterial genomes from NCBI's RefSeq (release 220). Sequences are named by their accession number, followed by "|" and their PGAP predicted function ("protein" tag). For example, the first sequence is named:</p><blockquote><p>WP_125174066.1|iron ABC transporter permease</p></blockquote><p>The process of creating the file involved the following steps.<br><strong>Step 1.</strong> Download 318,613 faa and fna files associated with a bacterial assembly in RefSeq. The following query was used:<br><i>esearch -db assembly -query '"Bacteria"[Organism] AND "latest refseq"[properties] AND "refseq has annotation"[properties]' | esummary | xtract -pattern DocumentSummary -element FtpPath_RefSeq</i><br><strong>Step 2.</strong> Verify all protein coding sequences match the expected protein sequence lengths within three codons, otherwise skip the assembly.<br><strong>Step 3.</strong> Remove all redundant protein coding or protein sequences in a genome. Only exact duplicates were removed, but they were removed from both nucleotides and proteins. Hence, a duplicated amino acid sequence would be discarded along with its coding sequence even if the coding sequence was unique. This was done to keep the two sets of sequences consistent.<br><strong>Step 4.</strong> Name sequences by their accession and PGAP predicted function, separated by a "|" character. The PGAP predicted function is generally uniform, although there are subtle difference between some taxon specific predictions. The predicted function is reasonably dependable but certainly not perfect.<br><strong>Step 5.</strong> Discard any sequences without a predicted function ("hypothetical protein"). These were discarded under the assumption that the protein's function would be required for downstream uses of the sequences.<br><strong>Step 6.</strong> Append protein and protein coding (nucleotide) sequences from randomly ordered assemblies to separate gzipped FASTA formatted files until the Zenodo file size limit was met for either file. Hence, there are many exact duplicate sequences in the set, but none for the sequences from each genome.</p><p>The final sets of sequences are intended to provide large sets of matched protein coding (nucleotide) and protein (amino acid) sequences with consistent labels. The FASTA descriptions in both files are identical. Note, the protein coding sequences do not exactly translate into the protein sequences because of slight differences in length (typically inclusion/exclusion of the first or last codon), as well as use of different translation tables depending on the organism.</p><p>See <i>Related works</i> for the companion file of protein (amino acid) sequences (DOI: 10.5281/zenodo.10030000).</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula

<p>This genomic dataset provides highly variable single-nucleotide polymorphism&nbsp;(SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>

opencc-by-4.0Jun 2020View details →
zenodo36/100

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models

<p>This is the consensus sequence repository to the manuscript &quot;Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models&quot;.<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>

openJun 2023View details →
dryad36/100

Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan

<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Nucleotide sequence database of Copper-containing membrane monooxygenases genes for analysing primer pairs targeting the ammonia monooxygenase subunit A gene of complete ammonia oxidising Nitrospira

<p>Nucleotide sequences of 487 Cu-mmo genes, including amoA comammox clade A and clade B, amoA ammonia oxidizing bacteria as well as other Cu-mmo genes.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Single Nucleotide Polymorphisms (SNPs) identified from the whole genome sequences of hilsa shad (Tenualosa ilisha) of the Bay of Bengal

<p>The data file contains 792,939 isolated SNPs identified by discoSnp++ v2.3.x (Uricaru et al., 2015) from the whole genome sequence of T. ilisha of the Bay of Bengal. The central sequence of length 2k-1 is seen in upper case, while the flanking sequences are seen in lower case. SNP_higher/lower: one of the two alleles. id: id of the SNP (each SNP has a unique id).</p> <p>FOR SNPs:</p> <p>P_i:pos_Alt1/Alt2: Information about a ith SNP (If more than a unique SNP is found, the following format is used: P_1:pos_Alt1/Alt2,P_2:pos_Alt1/Alt2,...</p> <p>pos: position of the SNP with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>Alt1: One of the two alleles</p> <p>Alt2: the other</p> <p>FOR INDELs:</p> <p>P_1:pos_size_repeatSize</p> <p>pos: predicted position of the indel with respect to the starting position of the bubble, i.e. the starting of the upper case sequence.</p> <p>size: predicted size of the indel</p> <p>repeatSize: Size of the longest sequence both prefix of the indel and prefix of the sequence located just after the insertion.</p> <p>high/low: sequence complexity. If the sequence if of low complexity (e.g. ATATATATATATATAT) this variable would be low</p> <p>nb_pol: number of polymorphism.</p> <p>left_unitig_length: size of the full left extension.</p> <p>right_unitig_length: size of the right extension.</p> <p>left_contig_length: size of the full left extension.</p> <p>right_contig_length: size of the right extension.</p> <p>C1: number of reads mapping the central upper case sequence from the first read set.</p> <p>C2: number of reads mapping the central upper case sequence from the second read set.</p> <p>Q1 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the first read set.</p> <p>Q2 [if reads were given in fastq]: average phred quality of the central nucleotide from the mapped reads from the second read set.</p> <p>G1: Genotype of the variant in the first read set.</p> <p>G2: Genotype of the variant in the second read set.</p> <p>rank: ranks the predictions according to their read coverage in each condition favoring SNPs that are discriminant between conditions.</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Towards a rapid sequencing-based molecular surveillance and mosaicism investigation of Toxoplasma gondii (nucleotide alignment dataset)

<p>This dataset includes the nucleotide alignment of eight Toxoplasma gondii genome loci (Sag1 / Chromossome VIII, Gra6 / Chromossome X, PK1 / Chromossome VI, Sag3 / Chromossome XII, L363 / Chromossome VIIb, CB21-4 / Chromossome III, M102 / Chromossome VIIa, Sag2&nbsp;/ Chromossome VIII). Each alignment includes sequences from T. gondii reference strains (retrieved from ToxoDB) as well as sequences from multiple clinical strains (obtained by Sanger /&nbsp;Next-generation sequencing) of the collection of the&nbsp;National Reference Laboratory of Parasitic and Fungal Infections, Department of Infectious Diseases, National Institute of Health Dr. Ricardo Jorge, Portugal.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

S and Z loci nucleotide sequences of Lolium multiflorum cultivar Rabiosa

<p>Nucleotide sequences of two scaffolds spanning the&nbsp;<em>S</em>-locus and two scaffolds the&nbsp;<em>Z</em>-locus in&nbsp;<em>Lolium multiflorum</em>&nbsp;cultivar Rabiosa</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Microsatellite data of bank voles and <em>Ixodes ricinus</em> plus TBEV nucleotide sequences

Open the record for dataset details and reuse information.

publicOct 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record