Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequences”
SMAdd-seq: Probing chromatin accessibility with small molecule DNA intercalation and nanopore sequencing
<p>Studies of in vivo chromatin organization have relied on the accessibility of the underlying DNA to nucleases or methyltransferases, which is limited by their requirement for purified nuclei and enzymatic treatment. Here, we introduce a nanopore-based sequencing technique called Small-Molecule Adduct sequencing (SMAdd-seq), where we profile chromatin accessibility by treating nuclei or intact cells with a small molecule, angelicin. Angelicin reacts with thymine bases in linker DNA not bound to core nucleosomes after UV light exposure, thereby labeling accessible DNA regions. By applying SMAdd-seq in Saccharomyces cerevisiae, we demonstrate that angelicin-modified DNA can be detected by its distinct nanopore current signals. To systematically identify angelicin modifications and analyze chromatin structure, we developed a neural network model, NEural network for mapping MOdifications in nanopore long-reads (NEMO). NEMO accurately called expected nucleosome occupancy patterns near transcription start sites at both bulk and single-molecule levels. We observe heterogeneity in chromatin structure and identify clusters of single-molecule reads with varying configurations at specific yeast loci. Furthermore, SMAdd-seq performs equivalently on purified yeast nuclei and intact cells, indicating the promise of this method for in vivo chromatin labeling on long single molecules to measure native chromatin dynamics and heterogeneity.</p>
Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"
Open the record for dataset details and reuse information.
ProTInSeq: transposon insertion tracking by ultra-deep DNA sequencing applied to identify small and large translated ORFs
<p>ProTInSeq is a novel -omics technique designed to characterize proteomes by using DNA ultra-deep sequencing. The technique is based on transposons engineered to have a positive or negative protein selection marker expressed when the transposon is inserted in-frame into a protein-coding gene. In the genome-reduced bacterium Mycoplasma pneumoniae, ProTInSeq identifies 80% of known expressed proteins, as well as 5 new open reading frames (ORFs; >100 amino acids); and 153 novel small ORF-encoded proteins (SEPs; ≤100 aa) that represent up to 18% of this bacterium’s proteome. ProTInSeq can be used to detect translational noise, for protein quantification and to provide insight into functional protein aspects such as relative half-life, stability, and membrane topology. Herein, we describe a methodology that can be easily implemented in any living system and allows the deep understanding of proteomes and more importantly the identification of small proteins by DNA ultra-sequencing.</p> <p>We include the following files:</p> <p>- processed_inscalling.zip: output obtain after running FASTQINS transposon calling tool over the datasets found at <a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee</a>. This include every genome position in <em>M. pneumoniae</em>, the number of times an insertion has been mapped to that position and the total read count value.</p> <p>- separated_library_metrics.zip: insertion and read count processed from processed_inscalling files associated to every ORF and intergenic region in <em>M. penumoniae. </em>Columns include frame measured (0 - whole gene, 1 - in-frame, 2 and 3 for following positions) and metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant).</p> <p>- allmetrics.xlsx: merged table with the combination of results from separated_library_metrics.zip tab-delimited files.</p> <p>- selective_metrics_allannotations.xlsx: this table includes all the available information about the 30,112 sequences that could encode for a coding sequence in <em>M. pneumoniae</em>. For each identifier (column B), we include coordinates information and nucleotide and amino acid length information (columns C-H). Column I includes the gene name when the entry is found annotated in <em>M. pneumoniae</em>. Localization and function are described in columns J and K. Column L includes the operon number in which the annotation would be expressed. We also included transcription-related information average expression (column M; as log2(gene read count/gene length) and estimated average RNA copies per cell (column N) considering 4 RNA sequencing samples covering different growth times (6, 24 and 48 hours, ArrayExpress identifier E‐MTAB‐6203). Column O accounts for the number of mass spectrometry experiments detecting that entry (to a maximum of 116) and column P accounts for the total number of unique tryptic peptides detected. This comes <a href="https://paperpile.com/c/BImj5N/eljz">[5]</a>, available for 12,426 sequences that present an amino acid length ≥19 (from 116 mass spectrometry experiments, ID PRIDE: PXD008243). Columns Q to T recapitulate protein copies per cell under different conditions (overall, extracting with urea, extracting with SDS and mean, respectively). Column U includes half-lives of the proteins. Columns V and W describe the reference density of insertion and essentiality assigned in previous studies. Column X-AA includes the predicted RanSEPs score, ribosome binding site presence, homology into seven groups: 0—no hits passed the thresholds defined; 1—conserved with an annotated function; 2—conserved as an annotated SEP in NCBI but no associated function; 3—conserved in a different species but target and homologous sequence not found in NCBI; 4—sequence is completely or partially (> 75%) repeated ≥ 3 times in the reference genome; 5—potential pseudogene; and 6—to depict those annotations that are found in the reference NCBI annotation file. , and function expected by homology, respectively. Columns AB to AD cover the output provided by Phobius, including the number of transmembrane segments, presence of signal peptide and transmembrane topology predicted by TM-HMM. Column AE includes the complex information where 1 implies that entry is functional as a monomer, 2 as dimer, and so on. Finally, columns AF-AH will be 1 if the protein is a Lon protease target, a lipoprotein, and/or a truncated gene or pseudogene, respectively, 0 otherwise. Following columns include for every sample presenting selective insertion rates in-frame using the following identifiers separated by underscores: marker (BarnB, Cm or Ery), type (control-AC or selection-BD, antibiotic concentration, sample replicate, frame measured, metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant). Last columns combine the number of samples each annotation has been identified. Notice for barnase library the results need to be interpreted considering it is a negative selection marker inverting the 0 and 1 meaning.</p>
DNA sequence data generated using non-invasive feather and eggshell samples from the Grenada Dove for two gene regions: Cyt b and ND2
<p>As an island endemic with a decreasing population, the Critically Endangered Grenada Dove <em>Leptotila wellsi</em> is threatened by accelerated loss of genetic diversity resulting from ongoing habitat fragmentation. Small, threatened populations are difficult to sample directly but advances in molecular methods mean that non-invasive samples can be used. We performed the first assessment of genetic diversity of populations of Grenada Dove by a) assessing mtDNA genetic diversity in the only two areas of occupancy on Grenada, b) defining the number of haplotypes present at each site and c) evaluating evidence of isolation between sites. We used non-invasively collected samples from two locations: Mt Hartman (n=18) and Perseverance (n=12). DNA extraction and PCR were used to amplify 1,751 bps of mtDNA from two mitochondrial markers: NADH dehydrogenase 2 (<em>ND2</em>) and Cytochrome b (<em>Cyt b</em>). Haplotype diversity (<em>h</em>) of 0.4, a nucleotide diversity (π) of 0.00023 and two unique haplotypes were identified within the <em>ND2</em> sequences; a single haplotype was identified within the <em>Cyt b </em>sequences. Of the two haplotypes identified; the most common haplotype (haplotype A = 73.9%) was observed at both sites and the other (haplotype B = 26.1%) was unique to Perseverance. Our results show low mitochondrial genetic diversity and clear evidence for genetically isolated populations. The Grenada Dove needs urgent conservation action, including habitat protection and potential augmentation of gene flow by translocation in order to increase genetic resilience and diversity with the ultimate aim of securing the long-term survival of this Critically Endangered species. </p>
Multimodal Epigenetic Sequencing Analysis (MESA) of Cell-free DNA for Non-invasive Colorectal Cancer Detection
<p>Processed data (feature-by-sample matrices) of non-disruptive bisulfite-free methylation sequencing for cfDNA samples from 4 clinical cohorts (Cohort 1, Cohort 2, Cohort 3, and cfTAPS dataset). Codes used to repeat the results in our paper can be found https://rpubs.com/LiYumei/926228 and https://github.com/ChaorongC/MESA. </p>
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China
<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>
Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays
<p>Data repository accompanying manuscript titled of "Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays."</p> <p>Prasetyo, A. P., Murray, J. M., Kurniawan, M. F. A. K., Sales, N. G., McDevitt, A. D., & Mariani, S. (2023). Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays. Conservation Letters, 16, e12971. https://doi.org/10.1111/conl.12971</p>
Oxford Nanopore sequencing for comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity
<p><span>The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here, we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. </span></span></p>
DNA large fragment deleting by compact, sequence-motif-free and specific TaqTth-hpRNA assisted with the microhomology-mediated end joining pathway
<p><span>A DNA editing tool TaqTth-hpRNA was developed in this study, composed of a compact recombinant TaqTth nuclease (832 aa) and a simple hairpin-RNA guiding probe (hpRNA). <em>In vitro</em> biochemical studies showed the TaqTth-hpRNA efficiently cleaves artificially synthesized ssDNA without stringent sequence motif like PAM. It can also cleave the genomic DNA of <em>E. coli</em> with ~80% efficiency. The TaqTth-hpRNA cleavage of genomic DNA in mammalian cells generated products with large fragment deletions mediated by the microhomology-mediated end joining (MMEJ) pathway. In addition, the cleavage was sensitive to mismatches in targeted regions, which was applied to specific damage of the <em>APP<sup>lon</sup></em> mutation in Alzheimer’s disease without disrupting the <em>APP<sup>wt</sup></em> locus. It is worth mentioning that the <em>APP<sup>lon</sup></em> sequence has only one base difference from that of <em>APP<sup>wt</sup></em>. The characteristics of small size, no PAM requirement, high specificity, and large deletion products make the TaqTth-hpRNA a potential therapeutic strategy for treating autosomal dominant disorders in the future.</span></p>
A revised classification of Glossopetalon (Crossosomataceae) based on restriction site-associated DNA sequencing
Glossopetalon inhabits arid regions in the American west and northern Mexico on limestone substrates. The genus comprises four species: G. clokeyi ; G. pungens ; G. texense ; and G. spinescens . Three of the species are narrow endemics. The fourth, G. spinescens , is a widespread species with six recognized varieties. All six varieties are intricately branched shrubs that have been difficult to identify due to a lack of clearly delineating morphological characters. Characters typically used to differentiate the varieties of G. spinescens, such as stem coloration, leaf blade size, and presence of stipules, are highly variable within and among populations. A custom protocol of double digest restriction-site associated DNA sequencing (ddRAD) was used to resolve the phylogeny of Glossopetalon and address if population genetic data analyses (such as STRUCTURE, SVDquartets, and phylogenetic networks) support the recognition of six varieties of G. spinescens . Glossopetalon was fully supported as monophyletic and G. pungens was resolved sister to the remaining taxa in the genus. The varieties of G. spinescens were resolved as two distinct lineages corresponding to their biogeography – one to the northwest (lineage 1) and one to southeast (lineage 2). Glossopetalon clokeyi was resolved at the base of lineage 1 and G. texense was embedded within lineage 2 sister to var. spinescens . Taxonomic changes include the recognition of G. texense and G. clokeyi as varieties of G. spinescens and description of a unique population from northern Arizona as a new variety – G. spinescens var. goodwinii .
Aligned DNA sequence matrix for phylogenetic analyses in the article "Dos nuevas especies del grupo Pristimantis boulengeri (Anura: Strabomantidae) de la cuenca alta del río Napo, Ecuador" by Bejarano, et al.
<p>Matrix in nexus format that include sequences of 16S (1-1295), RAG1 (1296-1922), 12S (1923-3306), and COI (3307-3984) for 97 specimens belonging to the genus <em>Pristimantis</em>, in addition to specimens of <em>Strabomantis</em> and <em>Niceforonia</em> as outgroups.</p>
Targeted sequencing of T-DNA borders in OCP1xOGC transgenic lines of Camelina
<p>Background: Genetic engineering of crop plants has been successful in transferring traits into elite lines beyond what can be achieved with breeding techniques. Introduction of transgenes originating from other species has conferred resistance to biotic and abiotic stresses, increased efficiency, and modified developmental programs. The next challenge is now to combine multiple transgenes into elite varieties via gene stacking to combine traits. Generating stable homozygous lines with multiple transgenes requires selection of segregating generations which is time consuming and labor intensive, especially if the crop is polyploid. Insertion site effects and transgene copy number are important metrics for commercialization and trait efficiency.</p> <p>Results: We have developed a simple method to identify the sites of transgene insertions using T-DNA-specific primers and high-throughput sequencing that enables identification of multiple insertion sites in the T1 generation of any crop transformed via <em>Agrobacterium</em>. We present an example using the allohexaploid oil-seed plant <em>Camelina sativa</em> to determine insertion site location of two transgenes.</p> <p>Conclusion: This new methodology enables the early selection of desirable transgene location and copy number to generate homozygous lines within two generations.</p>
Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing
<p>Studies in vertebrate genomics require sampling from a broad range of tissue types, taxa, and localities. Recent advancements in long-read and long-range genome sequencing have made it possible to produce high-quality chromosome-level genome assemblies for almost any organism. However, adequate tissue preservation for the requisite ultra-high molecular weight DNA (uHMW DNA) remains a major challenge. Here we present a comparative study of preservation methods for field and laboratory tissue sampling, across vertebrate classes and different tissue types. We find that no single method is best for all cases. Instead, the optimal storage and extraction methods vary by taxa, by tissue, and by down-stream application. Therefore, we provide sample preservation guidelines that ensure sufficient DNA integrity and amount required for use with long-read and long-range sequencing technologies across vertebrates. Our best practices generate the uHMW DNA needed for the high-quality reference genomes for Phase 1 of the Vertebrate Genomes Project (VGP), whose ultimate mission is to generate chromosome-level reference genome assemblies of all ~70,000 extant vertebrate species.</p>
Satyrium longicauda (Orchidaceae) species complex: Morphometric characters and DNA sequences
<p><strong>Morphology.xlsx</strong></p> <p>Morphological data obtained from 1802 individuals of the <em>Satyrium longicauda</em> (Orchidaceae) complex. Each entry has fifteen values representing both vegetative and reproductive characters. The full dataset or part of it has been used to perform univariate and multivariate analyses.<br> Each data point was obtained by measuring fresh material either with a ruler or a pair of calipers or by counting the number of different structures.</p> <p><strong>TS Alignment_130Samples.nex</strong></p> <p>Dataset that contains 130 ITS sequences including 14 gaps codified. This dataset has been used to perform phylogenetic analyses using both parsimony and Bayesian inference. All sequences are publically available on Genbank.</p>
A Monomeric Mycobacteriophage Immunity Repressor Utilizes Two Domains to Recognize an Asymmetric DNA Sequence
<p>Supporting source data consisting of simulation topologies and trajectories in pdb and xtc file formats, respectively.</p>
First large-scale quantification study of DNA preservation in insects from natural history collections using genome-wide sequencing
<p>Insect declines are a global issue with significant ecological and economic ramifications. Yet we have a poor understanding of the genomic impact these losses can have. Genome-wide data from historical specimens has the potential to provide baselines of population genetic measures to study population change, with natural history collections representing large repositories of such specimens. However, an initial challenge in conducting historical DNA data analyses, is to understand how molecular preservation varies between specimens. Here, we highlight how Next Generation Sequencing methods developed for studying archaeological samples can be applied to determine DNA preservation from only a single leg taken from entomological museum specimens, some of which are more than a century old. An analysis of genome-wide data from a set of 113 red-tailed bumblebee (Bombus lapidarius) specimens, from five British museum collections, was used to quantify DNA preservation over time. Additionally, to improve our analysis and further enable future research we generated a novel assembly of the red-tailed bumblebee genome. Our approach shows that museum entomological specimens are comprised of short DNA fragments with mean lengths below 100 base pairs (BP), suggesting a rapid and large-scale post-mortem reduction in DNA fragment size. After this initial decline, however, we find a relatively consistent rate of DNA decay in our dataset, and estimate a mean reduction in fragment length of 1.9bp per decade. The proportion of quality filtered reads mapping our assembled reference genome was around 50 %, and decreased by 1.1 % per decade. We demonstrate that historical insects have significant potential to act as sources of DNA to create valuable genetic baselines. The relatively consistent rate of DNA degradation, both across collections and through time, mean that population level analyses - for example for conservation or evolutionary studies - are entirely feasible, as long as the degraded nature of DNA is accounted for. </p>
DNA metabarcoding sequence data for diet analysis of caribou
<p>Woodland caribou (<em>Rangifer tarandus caribou</em>) are threatened in Canada due to the drastic decline in population size caused primarily by human-induced landscape changes that decrease habitat and increase predation risk. Conservation efforts have largely focused on reducing predators and protecting critical habitat, whereas research on dietary niches and the role of potential food constraints in lichen-poor environments is limited. To improve our understanding of dietary niche variability, we used a next-generation sequencing approach with metabarcoding of DNA extracted from faecal pellets of woodland caribou located on Lake Superior in lichen-rich (mainland) and lichen-poor (island) environments. Amplicon sequencing of fungal ITS2 region revealed lichen-associated fungi as predominant in samples from both populations, but amplification at the chloroplast <em>trnL </em>region, which was only successful on island samples, revealed primary consumption of yew based on relative read abundance (<em>Taxus spp.</em>; 83.68%) with dogwood (<em>Cornus spp</em>.; 9.67%) and maple (<em>Acer spp.</em>; 4.10%) also prevalent. These results suggest that conservation efforts for caribou need to consider the availability of food resources beyond lichen to ensure successful outcomes. More broadly, we provide a reliable methodology for assessing ungulate diet from archived faecal pellets that could reveal important dietary shifts over time in response to climate change.</p>
Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data
<p>This upload include data and scripts supporting the results described in the manuscript of <em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>. </em>Detailed description of the contents can be found in README.txt.</p>
Raw DNA sequence data of an individual known as "whitequark" (part 1)
<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with 350bp insert. 30-40× coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.