Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequence”
Base-resolution analyses of sequence and parent-of-origin dependent DNA methylation
GEO Series GSE33722. Homo sapiens; Mus musculus. 17 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
SpDamID: Marking DNA Bound by Protein Complexes Identifies Notch-Dimer Responsive Enhancers [next-generation sequencing]
GEO Series GSE70387. Mus musculus. 25 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.
The transcriptional coactivator Eya1 exerts transcriptional repressive activity by interacting with REST corepressors and DNA recognition sequences to maintain nephron progenitor identity [ChIP-seq]
GEO Series GSE202955. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Deep sequencing of MYC DNA-binding sites in Burkitt's lymphoma
GEO Series GSE30726. Homo sapiens. 22 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by array.
Unraveling the whole genome DNA methylation profile of zebrafish kidney marrow through Oxford Nanopore sequencing
GEO Series GSE232842. Danio rerio. 12 samples. Type: Methylation profiling by high throughput sequencing.
Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing
GEO Series GSE15353. Homo sapiens. 13 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Efficient small DNA fragment sequencing and miRNA, small RNA or csRNA-seq libraries using AVITI
GEO Series GSE267848. Homo sapiens; Bison bison; Bos taurus. 28 samples. Type: Other; Non-coding RNA profiling by high throughput sequencing.
Whole genome methylation sequencing for daughter fertility DNA mehylation biomarker
GEO Series GSE211926. Bos taurus. 12 samples. Type: Methylation profiling by high throughput sequencing.
DNA copy number, including telomeres and mitochondria, assayed using next-generation sequencing
GEO Series GSE21159. Homo sapiens. 3 samples. Type: Genome variation profiling by high throughput sequencing.
Sequence-dependent activity and compartmentalization of foreign DNA in a eukaryotic nucleus [ChIP-Seq]
GEO Series GSE217016. Saccharomyces cerevisiae. 54 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Methylated DNA Immunoprecipitation Sequencing of WT and βRapKO mice
GEO Series GSE95775. Mus musculus. 4 samples. Type: Methylation profiling by high throughput sequencing.
Data from: From environmental DNA sequences to ecological conclusions: how strong is the influence of methodological choices?
Aim: Environmental DNA (eDNA) is increasingly used for analysing and modelling all-inclusive biodiversity patterns. However, the reliability of eDNA-based diversity estimates is commonly compromised by arbitrary decisions for curating the data from molecular artefacts. Here, we test the sensitivity of common ecological analyses to these curation steps, and identify the crucial ones to draw sound ecological conclusions. Location : Valloire, French Alps. Taxon: Vascular plants and Fungi. Methods: Using soil eDNA metabarcoding data for plants and fungi from twenty plots sampled along a 1000-m elevation gradient, we tested how the conclusions from three types of ecological analyses: (i) the spatial partitioning of diversity, (ii) the diversity-environment relationship, and (iii) the distance-decay relationship, are robust to data curation steps. Since eDNA metabarcoding data also comprise erroneous sequences with low frequencies, diversity estimates were further calculated using abundance-based Hill numbers, which penalize rare sequences through a scaling parameter, namely the order of diversity q (Richness with q=0, Shannon diversity with q~1, Simpson diversity with q=2). Results: We showed that results from different ecological analyses had varying degrees of sensitivity to data curation strategies and that the use of Shannon and Simpson diversities led to more reliable results. We demonstrated that MOTU clustering, removal of PCR errors and of cross-sample contaminations had major impacts on ecological analyses. Main conclusions: In the Era of Big Data, eDNA metabarcoding is going to be one of the major tools to describe, model and predict biodiversity in space and time. However, ignoring crucial data curation steps will impede the robustness of several ecological conclusions. Here, we propose a roadmap of crucial curation steps for different types of ecological analyses.
Data from: Scaling up DNA barcoding - primer sets for simple and cost efficient arthropod systematics by multiplex PCR and Illumina amplicon sequencing
1. The simplicity and cost efficiency of Illumina amplicon sequencing has greatly contributed to the advancement of DNA barcoding and metabarcoding applications. However, current amplicon sequencing based barcoding approaches are usually restricted to short, single-locus fragments, limiting their taxonomic and phylogenetic resolution. 2. Here, we establish a cost efficient and simple multiplex PCR protocol for arthropod systematics by Illumina amplicon sequencing. We introduce primer sets, including several new, generic primers, to reliably amplify nine loci across a wide range of arthropods. Using a diverse collection of arthropod species from 19 orders, we test loci for amplification efficiency and estimate the effect of cross-species amplification bias on taxon recovery from bulk community samples. We then explore the taxonomic and phylogenetic utility of the primer sets, focusing on a dataset of spiders that includes both deep and recent divergences. 3. The set of loci provides good phylogenetic support across a wide taxonomic spectrum, making it a useful addition to COI for resolving lineages within a comparative context. All loci recover sequences for the majority of arthropod taxa in separate PCRs. However, cross-species amplification bias in some primers prevents an exhaustive taxon recovery from bulk community samples. 4. Our protocol makes it possible to generate multilocus datasets for large numbers of arthropod taxa for a fraction of the price and workload of Sanger sequencing. This opens up the possibility for parallel phylogenetic and taxonomic analysis of large collections of arthropods, but also enables rapid exploratory analyses of target lineages. Primers for metabarcoding applications should be carefully evaluated for their performance in bulk community samples and chosen to minimize cross-species amplification bias.
Full sequencing dataset of payload segments for "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency" (Part 2, FASTA file)
<p><span lang="EN-US">This is the full sequencing dataset of payload segments for the manuscript "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency", which is submitted for review.</span></p> <p><span lang="EN-US">We successfully encoded the first color film "Becky Sharp" (54.79 MB) into ~10.08 million DNA strands, achieving a sustained readout throughput of 459 kbit/s.</span></p> <p><span lang="EN-US">This full sequencing dataset contains all base-called data (in BIN and FASTA) of the payload segments. It can be used for recovery of the 28-layer color film data with the <br>DNA-LC software. </span></p> <p><span lang="EN-US">Due to size limitations, the dataset is divided into two parts. Part 1 contains all the files in BIN format, and Part 2 contains the corresponding files in FASTA format. This link is for Part 2.</span></p>
Full sequencing dataset of payload segments for "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency" (Part 1, BIN file)
<p><span lang="EN-US">This is the full sequencing dataset of payload segments for the manuscript "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency", which is submitted for review.</span></p> <p><span lang="EN-US">We successfully encoded the first color film "Becky Sharp" (54.79 MB) into ~10.08 million DNA strands, achieving a sustained readout throughput of 459 kbit/s.</span></p> <p><span lang="EN-US">This full sequencing dataset contains all base-called data (in BIN and FASTA) of the payload segments. It can be used for recovery of the 28-layer color film data with the<br>DNA-LC software. </span></p> <p><span lang="EN-US">Due to size limitations, the dataset is divided into two parts. Part 1 contains all the files in BIN format, and Part 2 contains the corresponding files in FASTA format. This link is for Part 1.</span></p>
ALV-like ERV DNA sequences
<p>These sequences correspond to ERVs extracted from genomic and wgs data from published genomes.</p> <p>Naming is as follows: in Alectoris-magna2-1, first is the host species, then followed by the contig number (the contigs with matches were asigned a natural (1,2,3) number from lowest to highest ID number) and ending in a second number to identify individual insertions if there were more than one in the same contig, corresponding the lower numbers to the lowest start position in the contig. This means Alectoris-magna1 and Alectoris-magna2-1 are in different contigs from the same species; while Alectoris-magna2-1 and Alectoris-magna2-2 are two different ERVs in the same contig.</p> <p>These sequences are associated with paper "Comparison of Endogenous Alpharetroviruses (ALV-like) across Galliform Species: New Distant Proviruses" (<a href="https://doi.org/10.3390/microorganisms12010086">https://doi.org/10.3390/ microorganisms12010086)</a></p> <p>Further details about genes and regions present in each ERV, as well as genome location, ORF integrity and other information can be found in the research paper above.</p>
Targeted DNA Sequencing of Circulating Tumor Cell-Enriched Samples from Stage II/III NSCLC Patients
<p>Targeted DNA sequencing of circulating tumor cell (CTC) enriched blood samples from stage II/III non-small cell lung cancer (NSCLC) patients prior to undergoing chemoradiation and immunotherapy treatment. The CTC samples were enriched using a label-free microfluidic platform, the Labyrinth. Following enrichment, the samples were fixed in 80% methanol or ethanol and stored at -80C or -20C, respectively. DNA was extracted from the samples using Qiagen's DNA/RNA AllPrep kit and submitted to the Univeristy of Michigan's Advanced Genomics Core (AGC) to perform targeted DNA sequencing using Illumina's TruSight Oncology (TSO) 500 assay. The TSO 500 assay consists of a 523 gene panel of specific oncogenes. Following sequecing, the pair-end fastq files were locally processed through the TSO 500 analysis pipeline (version: ruo-2.2.0.12) to produce the CombinedVariantOutput files per patient. Each file contains a list of all variants detected within the patient. For each variant, the following information is provided: gene name, chromosome, genomic position, reference call, alternative call, allele frequency, depth, p-dot notation, c-dot notation, consequence of variant, and the number of affected exons.</p>
Data from: Phylogenetic Systematics and Evolution of Primate-Derived Pneumocystis Based on Mitochondrial or Nuclear DNA Sequence Comparison
[No abstract entered]
Figure 1 from: Kurina O, Mantič M, Ševčík J (2017) A remarkable new genus of Keroplatidae (Insecta, Diptera) from the Afrotropical region, with DNA sequence data. African Invertebrates 58(1): 93-105. https://doi.org/10.3897/afrinvertebr.58.12655
Figure 1 - The sampling locality of Kibaleana apicospinosa sp. n. in southern Uganda.
Figure 2 from: Kurina O, Mantič M, Ševčík J (2017) A remarkable new genus of Keroplatidae (Insecta, Diptera) from the Afrotropical region, with DNA sequence data. African Invertebrates 58(1): 93-105. https://doi.org/10.3897/afrinvertebr.58.12655
Figure 2 - Malaise trapping at Kibale National Park in southern Uganda (Photo by O. Kurina).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.