Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
934
datasets available to search
ShareScore release 0.7.1
Dataset results
934 results for “Amino acids”
RefSeq bacterial protein (amino acid) sequences
<p><strong>Bacteria_Protein.fas.gz</strong></p><p>151,835,459 protein (amino acid) sequences extracted from 44,831 randomly selected bacterial genomes from NCBI's RefSeq (release 220). Sequences are named by their accession number, followed by "|" and their PGAP predicted function ("protein" tag). For example, the first sequence is named:</p><blockquote><p>WP_125174066.1|iron ABC transporter permease</p></blockquote><p>The process of creating the file involved the following steps.<br><strong>Step 1.</strong> Download 318,613 faa and fna files associated with a bacterial assembly in RefSeq. The following query was used:<br><i>esearch -db assembly -query '"Bacteria"[Organism] AND "latest refseq"[properties] AND "refseq has annotation"[properties]' | esummary | xtract -pattern DocumentSummary -element FtpPath_RefSeq</i><br><strong>Step 2.</strong> Verify all protein coding sequences match the expected protein sequence lengths within three codons, otherwise skip the assembly.<br><strong>Step 3.</strong> Remove all redundant protein coding or protein sequences in a genome. Only exact duplicates were removed, but they were removed from both nucleotides and proteins. Hence, a duplicated amino acid sequence would be discarded along with its coding sequence even if the coding sequence was unique. This was done to keep the two sets of sequences consistent.<br><strong>Step 4.</strong> Name sequences by their accession and PGAP predicted function, separated by a "|" character. The PGAP predicted function is generally uniform, although there are subtle difference between some taxon specific predictions. The predicted function is reasonably dependable but certainly not perfect.<br><strong>Step 5.</strong> Discard any sequences without a predicted function ("hypothetical protein"). These were discarded under the assumption that the protein's function would be required for downstream uses of the sequences.<br><strong>Step 6.</strong> Append protein and protein coding (nucleotide) sequences from randomly ordered assemblies to separate gzipped FASTA formatted files until the Zenodo file size limit was met for either file. Hence, there are many exact duplicate sequences in the set, but none for the sequences from each genome.</p><p>The final sets of sequences are intended to provide large sets of matched protein coding (nucleotide) and protein (amino acid) sequences with consistent labels. The FASTA descriptions in both files are identical. Note, the protein coding sequences do not exactly translate into the protein sequences because of slight differences in length (typically inclusion/exclusion of the first or last codon), as well as use of different translation tables depending on the organism.</p><p>See <i>Related works</i> for the companion file of protein coding (nucleotide) sequences (DOI: 10.5281/zenodo.10031801).</p>
Oral Amino Acid Nutrition to Improve Glucose Excursions in PCOS
ClinicalTrials.gov study NCT03717935. IPD Sharing: NO. Countries: 1. Publications: 0.
The Dayhoff Exchange Score: A new metric to quantify site saturation in amino acid datasets prior to phylogenetic analysis
Open the record for dataset details and reuse information.
Data for: Genetic control of grain amino acid composition in a UK soft wheat mapping population
Open the record for dataset details and reuse information.
Negative linkage disequilibrium between amino acid changing variants reveals interference among deleterious mutations in the human genome
Open the record for dataset details and reuse information.
Climate Change Across Seasons Experiment (CCASE) at the Hubbard Brook Experimental Forest; concentrations of foliar metabolites: polyamines, amino acids, chlorophyll, carotenoids, soluble proteins, soluble elements, sugars, and total nitrogen and carbon in red maple (Acer rubrum) trees.
Foliage was collected in 2015 and 2017 from red maple trees at the Climate Change Across Seasons Experiment (CCASE) as part of the Hubbard Brook Ecosystem Study (HBES). Analyses of foliar metabolites include polyamines, amino acids, chlorophylls, carotenoids, soluble proteins, soluble inorganic elements, sugars, and total nitrogen and carbon. There are six (11 x 14m) plots in total in this study; two control (plots 1 and 2), two warmed 5 degrees (°) Celsius (C) above ambient throughout the growing season (plots 3 and 4), and two warmed 5 °C in the growing season, with snow removal during the winter to induce soil freezing and then warmed with buried heating cables to create a subsequent thaw (plots 5 and 6). Each soil freeze/thaw cycle includes 72 hours of soil freezing followed by 72 hours of thaw. Four kilometers (km) of heating cable are buried in the soil to warm these four plots. Together, these treatments led to warmer growing season soil temperatures and an increased frequency of soil freeze-thaw cycles (FTCs) in winter. Our goal was to determine how these changes in soil temperature affect foliar nitrogen (N) and carbon metabolism of red maple trees. These data were gathered as a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
The Structural Basis of the Genetic Code: Amino Acid Recognition by Aminoacyl-tRNA Synthetases
<p>Data sets for the characterization of amino acid recognition in aminoacyl-tRNA synthetases:</p> <ul> <li>multiple-sequence alignment files in FASTA format</li> <li>Excel tables to infer original sequence positions from renumbered positions</li> </ul> <p>Accompanying the paper: <a href="https://www.biorxiv.org/content/10.1101/606459v1">https://www.biorxiv.org/content/10.1101/606459v1</a></p>
Isotope analyses of amino acids in fungi and fungal feeding Diptera larvae allow differentiating ectomycorrhizal and saprotrophic fungi-based food chains
1- Both ectomycorrhizal (ECM) and saprotrophic fungi are fundamental to carbon and nutrient dynamics in forest ecosystems; however, the relative importance of these different fungal functional groups for higher trophic levels of the soil food web is virtually unknown. 2- To explore differences between fungal functional groups and their importance for higher trophic levels, we analysed isotopic composition of nitrogen and carbon in amino acids (AAs) and bulk tissue of leaf litter, fungi, and fungal-feeding Diptera larvae. 3- By accounting for isotopic variability of utilized substrates, compound-specific isotope analyses of nitrogen in AAs yielded more realistic results for the trophic position of fungi than bulk isotope analyses, with converging trophic positions of saprotrophic and ECM fungi. 4- Saprotrophic and ECM fungi possessed different AA δ<sup>13</sup>C signatures separating fungal functional groups and their consumers in fingerprinting approaches, thereby allowing to trace energy fluxes from these basal resources to higher trophic levels. 5- A pronounced isotopic fractionation even in essential/source AAs of fungal-feeding Diptera larvae necessitates further studies on tissue-/compound-specific isotopic differences in fungi and on potential supplementation by gut microorganisms. 6- The results highlight the potential of compound-specific isotope analysis of amino acids to identify and integrate contributions of different fungal functional groups to higher trophic levels in soil food webs.
Data from: Trophic plasticity in a common reef-building coral: Insights from δ13C analysis of essential amino acids
1. Reef-building corals are mixotrophic organisms that can obtain nutrition from endosymbiotic microalgae (autotrophy) and particle capture (heterotrophy). Heterotrophic nutrition is highly beneficial to many corals, particularly in times of stress. Yet the extent to which different coral species rely on heterotrophic nutrition remains largely unknown because it is challenging to quantify. 2. We developed a quantitative approach to investigate coral nutrition using carbon isotope (δ13C) analysis of six essential amino acids (AAESS) in a common Indo-Pacific coral (Pocillopora meandrina) from the fore reef habitat of Palmyra Atoll. We sampled particulate organic matter (POM) and zooplankton as the dominant heterotrophic food sources in addition to the coral host and endosymbionts. We also measured bulk tissue carbon (δ13C) and nitrogen (δ15N) isotope values of each sample type. 3. Patterns among δ13C values of individual AAESS provided complete separation between the autotrophic (endosymbionts) and heterotrophic nutritional sources. In contrast, bulk tissue δ13C and δ15N values were highly variable across the putative food sources and among the coral and endosymbiont fractions, preventing accurate estimates of coral nutrition on Palmyra. 4. We used linear discriminant analysis to quantify differences among patterns of AAESS δ13C values, or 'fingerprints', of the food resources available to corals. This allowed for the development of a quantitative continuum of coral nutrition that can identify the relative contribution of autotrophic and heterotopic nutrition to individual colonies. Our approach revealed exceptional variation in conspecific colonies at scales of meters to kilometers. On average, 41% of AAESS in P. meandrina on Palmyra are acquired via heterotrophy but some colonies appear capable of obtaining the majority of AAESS from one source or the other. 5. The use of AAESS δ13C fingerprinting analysis offers a significant improvement on the current methods for quantitatively assessing coral trophic ecology. We anticipate that this approach will facilitate studies of coral nutrition in the field, which are essential for comparing coral trophic ecology across taxa and multiple spatial scales. Such information will be critical for understanding the role of heterotrophic nutrition in coral resistance and/or resilience to ongoing environmental change.
Reinvestigating the Photoprotection Properties of a Mycosporine Amino Acid Motif
<p>With the growing concern regarding commercially available ultraviolet (UV) filters damaging the environment, there is an urgent need to discover new UV filters. A family of molecules called mycosporines and mycosporine-like amino acids (referred to as MAAs collectively) are synthesized by cyanobacteria, fungi and algae and act as the natural UV filters for these organisms. Mycosporines are formed of a cyclohexenone core structure while mycosporine-like amino acids are formed of a cyclohexenimine core structure. To better understand the photoprotection properties of MAAs, we implement a bottom-up approach by first studying a simple analog of an MAA, 3-aminocyclohex-2-en-1-one (ACyO). Previous experimental studies on ACyO using transient electronic absorption spectroscopy (TEAS) suggest that upon photoexcitation, ACyO becomes trapped in the minimum of an S1 state, which persists for extended time delays (>2.5 ns). However, these studies were unable to establish the extent of electronic ground state recovery of ACyO within 2.5 ns due to experimental constraints. In the present studies, we have implemented transient vibrational absorption spectroscopy (as well as complementary TEAS) with Fourier transform infrared spectroscopy and density functional theory to establish the extent of electronic ground state recovery of ACyO within this time window. We show that by 1.8 ns, there is >75% electronic ground state recovery of ACyO, with the remaining percentage likely persisting in the electronic excited state. Long-term irradiation studies on ACyO have shown that a small percentage degrades after 2 h of irradiation, plausibly due to some of the aforementioned trapped ACyO going on to form a photoproduct. Collectively, these studies imply that a base building block of MAAs already displays characteristics of an effective UV filter.</p>
Tracking long-distance migration of marine fishes using compound-specific stable isotope analysis of amino acids
The long-distance migrations by marine fishes are difficult to track by field observation. Here, we propose a new method to track such migrations using stable nitrogen isotopic composition at the base of the food web (<i>δ</i><sup>15</sup>N<sub>Base</sub>), which allows for direct comparison of isotope ratios between proxy organisms of the isoscape and the target migratory animal. We initially constructed a <i>δ</i><sup>15</sup>N<sub>Base</sub> isoscape in the North Pacific by bulk and compound-specific isotope analyses of copepods (<i>n</i> = 360 and 24, respectively). We then determined retrospective <i>δ</i><sup>15</sup>N<sub>Base</sub> values of spawning chum salmon (<i>Oncorhynchus keta</i>) from their vertebral centra (10 sections from each of two salmon), and estimated their migration routes by using a state-space model. Our isotope tracking method successfully reproduced a known chum salmon migration route between the Okhotsk and Bering seas, and our findings suggest the presence of a migration route to the Bering Sea Shelf during a later growth stage.
Quantifying capital vs. income breeding: new promise with stable isotope measurements of individual amino acids
<p>1. Capital breeders accumulate nutrients prior to egg development, then use these stores to support offspring. In contrast, income breeders rely on local nutrients consumed contemporaneously with offspring development. Understanding such nutrient allocations is critical to assessing wildlife reliance on different habitats.</p> <p>2. Despite the contrast between these strategies, it remains challenging to trace nutrients from endogenous stores or exogenous food intake into offspring. Here, we tested a new solution to this problem.</p> <p>3. Using tissue samples collected opportunistically from wild emperor penguins (<i>Aptenodytes forsteri</i>), which exemplify capital breeding, we hypothesized that the stable carbon (δ<sup>13</sup>C) and nitrogen (δ<sup>15</sup>N) isotope values of individual amino acids (AA) in endogenous stores (e.g., muscle) and in egg yolk and albumen reflect the nutrient sourcing that distinguishes capital versus income breeding. Unlike other methods, this approach does not require untested assumptions or diet sampling.</p> <p>4. We found that over half of essential AA had δ<sup>13</sup>C values that did not differ between muscle and yolk or albumen, suggesting that most of these AA were directly routed from muscle into eggs. In contrast, almost all non-essential AA differed in δ<sup>13</sup>C values between muscle and yolk or between muscle and albumen, suggesting <i>de novo</i> synthesis. Over half of AA that have labile nitrogen atoms (i.e., "trophic" AA) had higher δ<sup>15</sup>N values in yolk and albumen than in muscle, suggesting that they were transaminated during their routing into egg tissue. This effect was smaller for AA with less labile nitrogen atoms (i.e., "source" AA).</p> <p>5. Our results indicate that the δ<sup>15</sup>N offset between trophic-source AA (Δ<sup>15</sup>N<sub>trophic-source</sub>) may provide an index of the extent of capital breeding. The value of emperor penguin Δ<sup>15</sup>N<sub>Pro-Phe</sub> was higher in yolk and albumen than in muscle, reflecting the mobilization of endogenous stores; in comparison, the value of Δ<sup>15</sup>N<sub>Pro-Phe</sub> was similar across muscle and egg tissue in previously-published data for income-breeding herring gulls (<i>Larus argentatus smithsonianus</i>). Our results provide a quantitative basis for using AA δ<sup>13</sup>C and δ<sup>15</sup>N, and isotopic offsets among AA (e.g., Δ<sup>15</sup>N<sub>Pro-Phe</sub>), to explore the allocation of endogenous versus exogenous nutrients across the capital versus income spectrum of avian reproduction.</p>
Amino acids (AA) all genes for: Beyond Drosophila: resolving the rapid radiation of schizophoran flies with phylotranscriptomics
<p><b>Background:</b></p> <p>The largest radiation of animal life since the end Cretaceous extinction event 66 million years ago is that of schizophoran flies: a third of fly diversity including <i>Drosophila </i>lab fruit flies, house flies, and many other well and poorly known true flies. Rapid diversification has hindered previous attempts to elucidate the phylogenetic relationships among major schizophoran clades. A robust phylogenetic hypothesis for the major lineages containing these 55,000 described species would be critical to understand the processes that contributed to the diversity of these agriculturally, medically, and forensically important flies. We use protein encoding sequence data from transcriptomes, including 3,145 genes from 70 species, representing all superfamilies, to improve the resolution of this previously intractable phylogenetic challenge.</p> <p><b>Results:</b></p> <p>Our results support a paraphyletic acalyptrate grade including a monophyletic Calyptratae and the monophyly of half of the acalyptrate superfamilies. The primary branching framework of Schizophora is well supported for the first time, revealing the primarily parasitic Pipunculidae and Sciomyzoidea s.l. as successive sister groups to the remaining Schizophora. Ephydroidea, <i>Drosophila</i>'s superfamily, is the sister group of Calyptratae. Sphaeroceroidea has modest support as the sister to all non-sciomyzoid Schizophora. We define two novel lineages corroborated by morphological traits, the Modified Oviscapt Clade containing Tephritoidea, Nerioidea, and other families, and the Cleft Pedicel Clade containg Calyptratae, Ephydroidea, and other families. Support values remain low among a challenging subset of lineages, including Diopsidae. The placement of these families remained uncertain in both concatenated maximum likelihood and multi-species coalescent approaches Rogue taxon removal was effective in increasing support values compared with strategies that maximize gene coverage or minimize missing data.</p> <p><b>Conclusions:</b></p> <p>Dividing most acalyptrate fly groups into four major lineages is supported consistently across analyses. Understanding the fundamental branching patterns of schizophoran flies provides a foundation for future comparative research on the genetics, ecology, and biocontrol.</p>
Data from: Conformational dynamics in TRPV1 channels reported by an encoded coumarin amino acid
TRPV1 channels support the detection of noxious and nociceptive input. Currently available functional and structural data suggest that TRPV1 channels have two gates within their permeation pathway: one formed by a ′bundle-crossing′ at the intracellular entrance and a second constriction at the selectivity filter. To describe conformational changes associated with channel gating, the fluorescent non-canonical amino acid coumarin-tyrosine was genetically encoded at Y671, a residue proximal to the selectivity filter. Total internal reflection fluorescence microscopy was performed to image the conformational dynamics of the channels in live cells. Photon counts and optical fluctuations from coumarin encoded within TRPV1 tetramers correlates with channel activation by capsaicin, providing an optical marker of conformational dynamics at the selectivity filter. In agreement with the fluorescence data, molecular dynamics simulations display alternating solvent exposure of Y671 in the closed and open states. Overall, the data point to a dynamic selectivity filter that may serve as a gate for permeation.
An unusual amino acid substitution within hummingbird cytochrome c oxidase alters a key proton-conducting channel
<p>Hummingbirds in flight exhibit the highest metabolic rate of all vertebrates. The bioenergetic requirements associated with sustained hovering flight raise the possibility of unique amino acid substitutions that would enhance aerobic metabolism. Here, we have identified a non-conservative substitution within the mitochondria-encoded cytochrome <i>c</i> oxidase subunit I (COI) that is fixed within hummingbirds, yet exceedingly rare among other vertebrates. This unusual change is also rare among metazoans, but can be identified in several clades with diverse life histories. We performed atomistic molecular dynamics simulations using bovine and hummingbird COI models, thereby bypassing experimental limitations imposed by the inability to modify mtDNA in a site-specific manner. Intriguingly, our findings suggest that COI amino acid position 153 (bovine numbering system) provides control over the hydration and activity of a key proton channel in COX. We discuss potential phenotypic outcomes linked to this intriguing alteration encoded by the hummingbird mitochondrial genome.</p>
Amino acid sequences of annotated genes in Pelargonium zonale
Open the record for dataset details and reuse information.
Dataset of: Deconvolving feeding niches and strategies of abyssal holothurians from their stable isotope, amino acid, and fatty acid composition
Open the record for dataset details and reuse information.
Data from: programming co-assembled peptide nanofiber morphology via anionic amino acid type: insights from molecular dynamics simulations
<p>Co-assembling peptides can be crafted into supramolecular biomaterials for use in biotechnological applications, such as cell culture scaffolds, drug delivery, biosensors, and tissue engineering. Peptide co-assembly refers to the spontaneous organization of two different peptides into a supramolecular architecture. Here we use molecular dynamics simulations to quantify the effect of anionic amino acid type on co-assembly dynamics and nanofiber structure in binary CATCH(+/-) peptide systems. CATCH peptide sequences follow a general pattern: CQCFCFCFCQC, where all C's are either a positively charged or a negatively charged amino acid. Specifically, we investigate the effect of substituting aspartic acid residues for the glutamic acid residues in the established CATCH(6E-) molecule, while keeping CATCH(6K+) unchanged. Our results show that structures consisting of CATCH(6K+) and CATCH(6D-) form flatter β-sheets, have stronger interactions between charged residues on opposing β-sheet faces, and have slower co-assembly kinetics than structures consisting of CATCH(6K+) and CATCH(6E-). Knowledge of the effect of sidechain type on assembly dynamics and fibrillar structure can help guide the development of advanced biomaterials and grant insight into sequence-to-structure relationships.</p>
Bulk Carbon and Amino Acid nitrogen isotope data from Baltic cod (Gadus morhua) and European flounder (Platichthys flesus) muscle tissue samples from the western and central Baltic Sea
<p><span>Eutrophication, increased temperatures and stratification can lead <span>to massive, filamentous, N<sub>2</sub>-fixing cyanobacterial (FNC) blooms in coastal ecosystems with largely unresolved consequences for the mass and energy supply in pelagic and benthic food webs. Mesozooplankton adapt to not top-down controlled FNC blooms by switching diets from phytoplankton to microzooplankton, resulting in a directly quantifiable increase in its trophic position (TP) from 2.0 (herbivore) to as high as 3.0 (carnivore). If this process in mesozooplankton, we call trophic lengthening, was transferred up to higher trophic levels of a food web, a large loss of energy could result in massive declines of fish biomass. </span></span><span>We used compound-specific nitrogen stable isotope data of amino acids (CSIA) to estimate and compare </span><span>the nitrogen (N) sources and TPs of cod and flounder (mesopredators) from areas</span><span> </span><span>with influence of FNC blooms (central Baltic Sea) and without it (western Baltic Sea)</span><span>. We tested if FNC-caused </span><span>trophic lengthening in mesozooplankton is carried over to fish.</span><span> The TP of cod from the western Baltic, feeding mainly on decapods, was equal to the global mean value (4.1, secondary carnivore). Only cod from the central Baltic, mainly feeding on zooplanktivorous pelagics, had a higher TP (4.8, near-tertiary carnivore), indicating a strong carry-over effect of </span><span>FNC-</span><span>caused trophic lengthening from mesozooplankton. In contrast, the TP of molluscivorous flounder (3.2 ± 0.2 in both areas), associated with the benthic food web, was unaffected by trophic lengthening. This suggests that FNC blooms cause a large loss of energy in zooplanktivorous but not in molluscivorous mesopredators. If FNC blooms continue to detour energy at the base of the pelagic food web, the TP of cod will not return to global mean values and the fish stock not recover. Monitoring the TP of key species can identify fundamental changes in ecosystems and provide useful information for resource management.</span></p>
NSRC-Search: Efficient searching for similar protein sequences of non-standard amino acid composition
<p>This research was funded by the National Science Centre in Poland (grant number 2021/41/N/ST6/01919)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.