Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
704
datasets available to search
ShareScore release 0.9.0
Dataset results
704 results for “Nucleotides”
Supplementary material for: "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"
<p>Supplementary material for the publication "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"</p> <p>The realated preprint can be found at Research Square (<a href="https://doi.org/10.21203/rs.3.rs-861830/v1">https://doi.org/10.21203/rs.3.rs-861830/v1</a>)</p> <p>Supplementary file 1: Supplementary results, tables and figures.</p> <p>Supplementary file 2: MultiQC report.</p> <p>Supplementary file 3: Observer concordance of the visual filtering step.</p> <p>Supplementary file 4: Snakemake workflow and scripts.</p> <p> </p>
MGBC-26640: nucleotide sequences for gene annotations
<p>Nucleotide sequences of annotated genes from the 26,640 high-quality, non-redundant genomes of the MGBC.</p>
High nucleotide similarity of three Copia lineage LTR retrotransposons among plant genomes
<p>Transposable elements (TEs) are mobile genetic elements found in the majority of eukaryotic genomes. TEs deeply impact the structure and evolution of chromosomes and can induce mutations affecting coding genes. In plants, the major group of TEs is Long Terminal Repeats retrotransposons (LTR-RT). They are classified into superfamilies (<em>Gypsy</em>, <em>Copia</em>) and sub-classified into lineages. Horizontal transfer (HT), defined as the nonsexual transmission of genetic material between species, is a process allowing LTR-RTs to invade a new genome. Although this phenomenon was considered rare, recent studies demonstrate numerous transfers of LTR-RTs, suggesting that HT may be more frequent than initially estimated.</p> <p>This study aims to determine which LTR-RT lineages are shared with high similarity among 69 reference plant genomes. We identified and classified 88,450 LTR-RTs and determined 143 cases (involving 94 elements) of high similarities between pairs of genomes. Most of them involved three <em>Copia</em> lineages (<em>Oryco/Ivana</em>, <em>Retrofit/Ale</em> and <em>Tork/Tar/Ikeros</em>). A detailed analysis of three cases of high similarities involving <em>Tork/Tar/Ikeros</em> group shows a patchy distribution of the elements and phylogenetic incongruities, indicating they originated from potential HTs. Overall, our results suggest that <em>Copia</em> LTR-RTs share outstanding similarity between very distant species and may probably be more involved in HT mechanisms.</p>
Data set for "A Magnesium Binding Site And The Anomeric Effect Regulate The Abiotic Redox Chemistry Of Nicotinamide Nucleotides"
<p>Data associated with Sebastianelli L, Kaur H, Chen Z, Krishnamurthy R, Mansy SS (2024) A magnesium binding site and the anomeric effect regulate the abiotic redox chemistry of nicotinamide nucleotides. Chem Eur J. 30, e202400411. DOI: 10.1002/chem.202400411 [<a href="https://chemistry-europe.onlinelibrary.wiley.com/doi/abs/10.1002/chem.202400411">link</a>]</p>
◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West) in Bumps on the back: An unusual morphology in phylogenetically distinct Peridinium aff. cinctum (= Peridinium tuberosum; Peridiniales, Dinophyceae)
◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West)
◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae) in Morphological and molecular variability of Peridinium volzii Lemmerm. (Peridiniaceae, Dinophyceae) and its relevance for infraspecific taxonomy
◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae)
Anchored phylogenomics unravels the evolution of spider flies (Diptera, Acroceridae) and reveals discordance between nucleotides and amino acids
<p>Supplementary Material accompanying the manuscript titled "Anchored phylogenomics unravels the evolution of spider flies (Acroceridae) and reveals discordance between nucleotides and amino acids", including Supplementary Figures, Tables and Datasets.</p>
Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group. in Distribution of extracellular enzyme-producing bacteria in the digestive tracts of 4 brackish water fish species
Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group.
Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S in A herpetological survey of western Zambia
Fig. 5. Maximum likelihood phylogenetic tree inferred from nucleotide sequence data from mitochondrial 16S rRNA of Phrynobatrachus natalensis. Numbers above branches are non-parametric bootstrap support values. Specimen vouchers or GenBank accession numbers are shown in parentheses. Colored polygons highlight the clades comprising specimens from this study. (*) Nearest sample from type locality of Phrynobatrachus natalensis; (**) Haplotype groups A and B in Zimkus and Schick (2010).
Unravelling the Dependence of Aptamer–Target Binding Orientation and Strength on Nucleotide Modification Size, Position, and Number: A Computational Case Study of TBA
<p>Data deposition for "Unravelling the Dependence of Aptamer–Target Binding Orientation and Strength on Nucleotide Modification Size, Position, and Number: A Computational Case Study of TBA" published in Nucleic Acids Research. Data includes custom paramaterized nucleobase residues for use in AMBER molecular dyanmics simulations. Initial structures and setup code along with starting structures for generation of simulations in AMBER (pmemd.cuda). A in-house "MD-Generator" script to create nessessary files with our select simulation settings. Lastly, also included are CPPTRAJ analysis commands and scripts with visualizion scripts for use in Pymol 2.5 and above. </p>
Fig. 1 in Polymerase chain reaction and gyrA nucleotide sequence analysis of Wolbachia endosymbionts (Rickettsiales: Anaplasmataceae) in various species of Culicidae, Cimex lectularius (Hemiptera: Cimicidae) and Dirofilaria immitis (Rhabditida: Onchocercidae)
Fig. 1. Phylogenetic tree based on Maximum Likelihood depicting the grouping of Wolbachia from various hosts based on analysis of the gyrA gene. The numerical value displayed on branches is the bootstrap value (1,000 replicates), and branches with values below 50% are collapsed. The tree illustrates that gyrA sequences distinguish Wolbachia subtypes based on host taxonomy, demonstrating that this gene may contribute to Wolbachia strain typing projects and future phylogenetic analysis.
Single-nucleotide polymorphisms in isolates Pyricularia oryzae isolates from rice and other hosts
<p>We investigated the presence of inter-lineage hybrids in a set of 886 isolates of the rice-infecting lineage of the rice blast fungus (<em>Pyricularia oryzae</em>).</p> <p>More details in https://doi.org/10.1101/2020.06.02.129296</p> <p>List of files:</p> <p>concatenation_5190SCOs_179isolates.fasta.gz: concatenated sequences of 5190 single copy orthologs from 179 isolates (used in Figures A and B, and computations of polymorphism and divergence in Appendix 1).</p> <p>concatenation_3686SNPs_415isolates.fasta.gz: concatenated polymorphisms identified in 415 isolates at 3686 genomic positions (used in Figure C, Appendix 1).</p> <p>concatenation_503889SNPs_96isolates.fasta.gz: concatenated polymorphisms identified in 96 isolates at 503889 genomic positions (used in Figures D and E, Appendix 1).</p> <p>orthology_table_with_70-15_genemodels.txt: Orthology table, with one orthogroup per line, one species per column, 16445 orthogroups, 179 isolates. The MGG column represents gene models from the 70-15 reference genome MG8, determined subsequently to orthology analysis by blastx of 70-15 sequences against isolate VT0027, which was the closest isolate in a neighbor-net network.</p>
Suppl. files to: Cellular and molecular targets of nucleotide-tagged trithiola-to-bridged arene ruthenium complexes in the protozoan para-sites Toxoplasma gondii and Trypanosoma brucei
<p>These are supplementary files for the manuscript entitled:</p> <p>Cellular and molecular targets of nucleotide-tagged trithiolato-bridged arene ruthenium complexes in the protozoan parasites <em>Toxoplasma gondii</em>and <em>Trypanosoma brucei</em></p> <p>submitted to International Journal of Molecular Sciences</p> <p>by: <strong>Nicoleta Anghel<sup>1¥</sup>, Joachim Müller<sup>1¥*</sup>, Mauro Serricchio<sup> 2</sup>, Jennifer Jelk <sup>2</sup>, Peter Bütikofer<sup>2</sup>, Ghalia Boubaker<sup>1</sup>, Dennis Imhof<sup>1</sup>, Jessica Ramseier<sup>1</sup>, Oksana Desiatkina<sup>3</sup>, Emilia Păunescu<sup>3</sup>, Sophie Braga-Lagache<sup>4</sup>, Manfred Heller<sup>4</sup>, Julien Furrer<sup>3</sup>, Andrew Hemphill<sup>1*</sup></strong></p>
DATASET: Replication-competent HIV-1 in human alveolar macrophages and monocytes despite nucleotide pools with elevated dUTP
<p>Complete data set to support content of the study: Cui, et al (2022) <strong>Replication-competent HIV-1 in human alveolar macrophages and monocytes despite nucleotide pools with elevated dUTP</strong></p>
Dataset and fitting methods for Kinetic Proofreading can Enhance Single Nucleotide Discrimination in a Non-enzymatic DNA Strand Displacement Network
<p>This upload contains raw experimental data and the fitting methods used for the article "Kinetic Proofreading can Enhance Single Nucleotide Discrimination in a Non-enzymatic DNA Strand Displacement Network".</p>
RefSeq bacterial protein coding (nucleotide) sequences
<p><strong>Bacteria_Nucleotide.fas.gz</strong></p><p>151,835,459 protein coding (nucleotide) sequences extracted from 44,831 randomly selected bacterial genomes from NCBI's RefSeq (release 220). Sequences are named by their accession number, followed by "|" and their PGAP predicted function ("protein" tag). For example, the first sequence is named:</p><blockquote><p>WP_125174066.1|iron ABC transporter permease</p></blockquote><p>The process of creating the file involved the following steps.<br><strong>Step 1.</strong> Download 318,613 faa and fna files associated with a bacterial assembly in RefSeq. The following query was used:<br><i>esearch -db assembly -query '"Bacteria"[Organism] AND "latest refseq"[properties] AND "refseq has annotation"[properties]' | esummary | xtract -pattern DocumentSummary -element FtpPath_RefSeq</i><br><strong>Step 2.</strong> Verify all protein coding sequences match the expected protein sequence lengths within three codons, otherwise skip the assembly.<br><strong>Step 3.</strong> Remove all redundant protein coding or protein sequences in a genome. Only exact duplicates were removed, but they were removed from both nucleotides and proteins. Hence, a duplicated amino acid sequence would be discarded along with its coding sequence even if the coding sequence was unique. This was done to keep the two sets of sequences consistent.<br><strong>Step 4.</strong> Name sequences by their accession and PGAP predicted function, separated by a "|" character. The PGAP predicted function is generally uniform, although there are subtle difference between some taxon specific predictions. The predicted function is reasonably dependable but certainly not perfect.<br><strong>Step 5.</strong> Discard any sequences without a predicted function ("hypothetical protein"). These were discarded under the assumption that the protein's function would be required for downstream uses of the sequences.<br><strong>Step 6.</strong> Append protein and protein coding (nucleotide) sequences from randomly ordered assemblies to separate gzipped FASTA formatted files until the Zenodo file size limit was met for either file. Hence, there are many exact duplicate sequences in the set, but none for the sequences from each genome.</p><p>The final sets of sequences are intended to provide large sets of matched protein coding (nucleotide) and protein (amino acid) sequences with consistent labels. The FASTA descriptions in both files are identical. Note, the protein coding sequences do not exactly translate into the protein sequences because of slight differences in length (typically inclusion/exclusion of the first or last codon), as well as use of different translation tables depending on the organism.</p><p>See <i>Related works</i> for the companion file of protein (amino acid) sequences (DOI: 10.5281/zenodo.10030000).</p>
Single nucleotide polymorphism (SNPs) data for Scurria scurra, Scurria variabilis, Scurria ceciliana and Scurria araucana
Open the record for dataset details and reuse information.
Assessing kinship detection: Single nucleotide polymorphism array density and estimator comparison in white-tailed deer
Open the record for dataset details and reuse information.
<i>k</i>-mer-based diversity scales with population size proxies more than nucleotide diversity in a meta-analysis of 98 plant species
Open the record for dataset details and reuse information.
Data from: FSHB transcription is regulated by a novel 5’ distal enhancer containing a fertility-associated single nucleotide polymorphism
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.