Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,696

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,696 results for “DNA sequence”

Learn how ShareScore rates datasets ↗
geo24/100

Base-resolution analyses of sequence and parent-of-origin dependent DNA methylation

GEO Series GSE33722. Homo sapiens; Mus musculus. 17 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2012View details →
geo24/100

SpDamID: Marking DNA Bound by Protein Complexes Identifies Notch-Dimer Responsive Enhancers [next-generation sequencing]

GEO Series GSE70387. Mus musculus. 25 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenAug 2015View details →
geo24/100

The transcriptional coactivator Eya1 exerts transcriptional repressive activity by interacting with REST corepressors and DNA recognition sequences to maintain nephron progenitor identity [ChIP-seq]

GEO Series GSE202955. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenOct 2022View details →
geo24/100

Deep sequencing of MYC DNA-binding sites in Burkitt's lymphoma

GEO Series GSE30726. Homo sapiens. 22 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by array.

openGEO-OpenNov 2011View details →
geo24/100

Unraveling the whole genome DNA methylation profile of zebrafish kidney marrow through Oxford Nanopore sequencing

GEO Series GSE232842. Danio rerio. 12 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenMay 2023View details →
geo24/100

Genome-wide profiles of STAT1 DNA association using chromatin immunoprecipitation and massively parallel sequencing

GEO Series GSE15353. Homo sapiens. 13 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMar 2009View details →
geo24/100

Efficient small DNA fragment sequencing and miRNA, small RNA or csRNA-seq libraries using AVITI

GEO Series GSE267848. Homo sapiens; Bison bison; Bos taurus. 28 samples. Type: Other; Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenJun 2024View details →
geo24/100

Whole genome methylation sequencing for daughter fertility DNA mehylation biomarker

GEO Series GSE211926. Bos taurus. 12 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenAug 2023View details →
geo24/100

DNA copy number, including telomeres and mitochondria, assayed using next-generation sequencing

GEO Series GSE21159. Homo sapiens. 3 samples. Type: Genome variation profiling by high throughput sequencing.

openGEO-OpenApr 2010View details →
geo24/100

Sequence-dependent activity and compartmentalization of foreign DNA in a eukaryotic nucleus [ChIP-Seq]

GEO Series GSE217016. Saccharomyces cerevisiae. 54 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2025View details →
geo24/100

Methylated DNA Immunoprecipitation Sequencing of WT and βRapKO mice

GEO Series GSE95775. Mus musculus. 4 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenApr 2017View details →
dryad24/100

Data from: From environmental DNA sequences to ecological conclusions: how strong is the influence of methodological choices?

Aim: Environmental DNA (eDNA) is increasingly used for analysing and modelling all-inclusive biodiversity patterns. However, the reliability of eDNA-based diversity estimates is commonly compromised by arbitrary decisions for curating the data from molecular artefacts. Here, we test the sensitivity of common ecological analyses to these curation steps, and identify the crucial ones to draw sound ecological conclusions. Location : Valloire, French Alps. Taxon: Vascular plants and Fungi. Methods: Using soil eDNA metabarcoding data for plants and fungi from twenty plots sampled along a 1000-m elevation gradient, we tested how the conclusions from three types of ecological analyses: (i) the spatial partitioning of diversity, (ii) the diversity-environment relationship, and (iii) the distance-decay relationship, are robust to data curation steps. Since eDNA metabarcoding data also comprise erroneous sequences with low frequencies, diversity estimates were further calculated using abundance-based Hill numbers, which penalize rare sequences through a scaling parameter, namely the order of diversity q (Richness with q=0, Shannon diversity with q~1, Simpson diversity with q=2). Results: We showed that results from different ecological analyses had varying degrees of sensitivity to data curation strategies and that the use of Shannon and Simpson diversities led to more reliable results. We demonstrated that MOTU clustering, removal of PCR errors and of cross-sample contaminations had major impacts on ecological analyses. Main conclusions: In the Era of Big Data, eDNA metabarcoding is going to be one of the major tools to describe, model and predict biodiversity in space and time. However, ignoring crucial data curation steps will impede the robustness of several ecological conclusions. Here, we propose a roadmap of crucial curation steps for different types of ecological analyses.

opencc-zeroJul 2020View details →
dryad24/100

Data from: Scaling up DNA barcoding - primer sets for simple and cost efficient arthropod systematics by multiplex PCR and Illumina amplicon sequencing

1. The simplicity and cost efficiency of Illumina amplicon sequencing has greatly contributed to the advancement of DNA barcoding and metabarcoding applications. However, current amplicon sequencing based barcoding approaches are usually restricted to short, single-locus fragments, limiting their taxonomic and phylogenetic resolution. 2. Here, we establish a cost efficient and simple multiplex PCR protocol for arthropod systematics by Illumina amplicon sequencing. We introduce primer sets, including several new, generic primers, to reliably amplify nine loci across a wide range of arthropods. Using a diverse collection of arthropod species from 19 orders, we test loci for amplification efficiency and estimate the effect of cross-species amplification bias on taxon recovery from bulk community samples. We then explore the taxonomic and phylogenetic utility of the primer sets, focusing on a dataset of spiders that includes both deep and recent divergences. 3. The set of loci provides good phylogenetic support across a wide taxonomic spectrum, making it a useful addition to COI for resolving lineages within a comparative context. All loci recover sequences for the majority of arthropod taxa in separate PCRs. However, cross-species amplification bias in some primers prevents an exhaustive taxon recovery from bulk community samples. 4. Our protocol makes it possible to generate multilocus datasets for large numbers of arthropod taxa for a fraction of the price and workload of Sanger sequencing. This opens up the possibility for parallel phylogenetic and taxonomic analysis of large collections of arthropods, but also enables rapid exploratory analyses of target lineages. Primers for metabarcoding applications should be carefully evaluated for their performance in bulk community samples and chosen to minimize cross-species amplification bias.

opencc-zeroDec 2017View details →
zenodo24/100

Full sequencing dataset of payload segments for "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency" (Part 2, FASTA file)

<p><span lang="EN-US">This is the full sequencing dataset of payload segments for the manuscript "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency", which is submitted for review.</span></p> <p><span lang="EN-US">We successfully encoded the first color film "Becky Sharp" (54.79 MB) into ~10.08 million DNA strands, achieving a sustained readout throughput of 459 kbit/s.</span></p> <p><span lang="EN-US">This full sequencing dataset contains all base-called data (in BIN and FASTA) of the payload segments. It can be used for recovery of the 28-layer color film data with the&nbsp;<br>DNA-LC software.&nbsp;</span></p> <p><span lang="EN-US">Due to size limitations, the dataset is divided into two parts. Part 1 contains all the files in BIN format, and Part 2 contains the corresponding files in FASTA format. This link is for Part 2.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo24/100

Full sequencing dataset of payload segments for "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency" (Part 1, BIN file)

<p><span lang="EN-US">This is the full sequencing dataset of payload segments for the manuscript "A Base Multilayer DNA Storage Architecture Enables Real-Time Decoding to Overcome Latency", which is submitted for review.</span></p> <p><span lang="EN-US">We successfully encoded the first color film "Becky Sharp" (54.79 MB) into ~10.08 million DNA strands, achieving a sustained readout throughput of 459 kbit/s.</span></p> <p><span lang="EN-US">This full sequencing dataset contains all base-called data (in BIN and FASTA) of the payload segments. It can be used for recovery of the 28-layer color film data with the<br>DNA-LC software.&nbsp;</span></p> <p><span lang="EN-US">Due to size limitations, the dataset is divided into two parts. Part 1 contains all the files in BIN format, and Part 2 contains the corresponding files in FASTA format. This link is for Part 1.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo24/100

ALV-like ERV DNA sequences

<p>These sequences correspond to ERVs extracted from genomic and wgs data from published genomes.</p> <p>Naming is as follows: in Alectoris-magna2-1, first is the host species, then followed by the contig number (the contigs with matches were asigned a natural (1,2,3) number from lowest to highest ID number) and ending in a second number to identify individual insertions if there were more than one in the same contig, corresponding the lower numbers to the lowest start position in the contig. This means Alectoris-magna1 and Alectoris-magna2-1 are in different contigs from the same species; while Alectoris-magna2-1 and Alectoris-magna2-2 are two different ERVs in the same contig.</p> <p>These sequences are associated with paper "Comparison of Endogenous Alpharetroviruses (ALV-like) across Galliform Species: New Distant Proviruses" (<a href="https://doi.org/10.3390/microorganisms12010086">https://doi.org/10.3390/ microorganisms12010086)</a></p> <p>Further details about genes and regions present in each ERV, as well as genome location, ORF integrity and other information can be found in the research paper above.</p>

opencc-by-4.0Feb 2024View details →
zenodo24/100

Targeted DNA Sequencing of Circulating Tumor Cell-Enriched Samples from Stage II/III NSCLC Patients

<p>Targeted DNA sequencing of circulating tumor cell (CTC) enriched blood samples from stage II/III non-small cell lung cancer (NSCLC) patients prior to undergoing chemoradiation and immunotherapy treatment. The CTC samples were enriched using a label-free microfluidic platform, the Labyrinth. Following enrichment, the samples were fixed in 80% methanol or ethanol and stored at -80C or -20C, respectively. DNA was extracted from the samples using Qiagen's DNA/RNA AllPrep kit and submitted to the Univeristy of Michigan's Advanced Genomics Core (AGC) to perform targeted DNA sequencing using Illumina's TruSight Oncology (TSO) 500 assay. The TSO 500 assay consists of a 523 gene panel of specific oncogenes. Following sequecing, the pair-end fastq files were locally processed through the TSO 500 analysis pipeline (version: ruo-2.2.0.12) to produce the CombinedVariantOutput files per patient. Each file contains a list of all variants detected within the patient. For each variant, the following information is provided: gene name, chromosome, genomic position, reference call, alternative call, allele frequency, depth, p-dot notation, c-dot notation, consequence of variant, and the number of affected exons.</p>

restrictedcc-by-4.0Apr 2024View details →
dryad24/100

Data from: Phylogenetic Systematics and Evolution of Primate-Derived Pneumocystis Based on Mitochondrial or Nuclear DNA Sequence Comparison

[No abstract entered]

opencc-zeroDec 2008View details →
zenodo24/100

Figure 1 from: Kurina O, Mantič M, Ševčík J (2017) A remarkable new genus of Keroplatidae (Insecta, Diptera) from the Afrotropical region, with DNA sequence data. African Invertebrates 58(1): 93-105. https://doi.org/10.3897/afrinvertebr.58.12655

Figure 1 - The sampling locality of Kibaleana apicospinosa sp. n. in southern Uganda.

opencc-by-4.0May 2017View details →
zenodo24/100

Figure 2 from: Kurina O, Mantič M, Ševčík J (2017) A remarkable new genus of Keroplatidae (Insecta, Diptera) from the Afrotropical region, with DNA sequence data. African Invertebrates 58(1): 93-105. https://doi.org/10.3897/afrinvertebr.58.12655

Figure 2 - Malaise trapping at Kibale National Park in southern Uganda (Photo by O. Kurina).

opencc-by-4.0May 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record