Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequences”
Simultaneous profiling of histone modifications and DNA methylation via nanopore sequencing
<p>Datasets that contain a minimum of nanopore reads sufficient for hidden Markov model training and for evaluating the performance of our computational tool - nanoHiMe at simultaneously calling CpG and/or adenine methylation on individual nanopore reads.<em> Ecoli</em>_PCR_amplicons_100k.tgz, <em>Ecoli</em>_PCR_MSssI_100k.tar.gz and <em>Ecoli</em>_PCR_pA-Hia5_100k.tar.gz are used for training new parameters of the emission distributions of individual <em>k</em>-mers from DNA template without modification, with fully methylated CpGs, and with partially methylated adenines, respectively. nanoHiMe_H3K27me3.fast5.tgz are the nanopore sequencing reads from H3K27me3 nanoHiMe-seq experiments in GM12878 cells and used for evaluating the performance of nanoHiMe at jointly calling CpG and adenine methylation.</p>
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
DNA sequence and bioacoustic data of the Guibemantis liber complex from Madagascar (Amphibia, Mantellidae)
<p>Data from a taxonomic revision of the Guibemantis liber complex from Madagascar. The following data are included:</p> <p>- Advertisement call recordings of Guibemantis liber, G. razoky and G. razandry from different localities in wav format. See associated publication for metadata.</p> <p>- A table in Excel format with all DNA sequences used, metadata of the respective samples and specimens, and Genbank accession numbers.</p>
Supplementary data for: DNA sequences are as useful as protein sequences for inferring deep phylogenies
<p>Inference of deep phylogenies has almost exclusively used protein rather than DNA sequences, based on the perception that protein sequences are less prone to homoplasy and saturation or to issues of compositional heterogeneity than DNA sequences. Here we analyze a model of codon evolution under an idealized genetic code and demonstrate that those perceptions may be misconceptions. We conduct a simulation study to assess the utility of protein versus DNA sequences for inferring deep phylogenies, with protein-coding data generated under models of heterogeneous substitution processes across sites in the sequence and among lineages on the tree, and then analyzed using nucleotide, amino acid, and codon models. Analysis of DNA sequences under nucleotide-substitution models (possibly with the third codon positions excluded) recovered the correct tree at least as often as analysis of the corresponding protein sequences under modern amino acid models. We also applied the different data-analysis strategies to an empirical dataset to infer the metazoan phylogeny. Our results from both simulated and real data suggest that DNA sequences may be as useful as proteins for inferring deep phylogenies and should not be excluded from such analyses. Analysis of DNA data under nucleotide models has a major computational advantage over protein-data analysis, potentially making it feasible to use advanced models that account for among-site and among-lineage heterogeneity in the nucleotide-substitution process in inference of deep phylogenies.</p>
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
Open the record for dataset details and reuse information.
Data from: Restriction site-associated DNA sequencing reveals local adaptation despite high levels of gene flow in Sardinella lemuru (Bleeker, 1853) along the northern coast of Mindanao, Philippines
Open the record for dataset details and reuse information.
Data from: Microhaplotypes provide increased power from short-read DNA sequences for relationship inference
Open the record for dataset details and reuse information.
Optimal sequence similarity thresholds for clustering of molecular operational taxonomic units in DNA metabarcoding studies
Open the record for dataset details and reuse information.
Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens
Open the record for dataset details and reuse information.
DNA sequencing of drive and standard Teleopsis dalmanni males
Open the record for dataset details and reuse information.
Supplementary data for: DNA sequences are as useful as protein sequences for inferring deep phylogenies
Open the record for dataset details and reuse information.
Gaps in DNA sequence libraries for Macaronesian marine macroinvertebrates imply decades till completion and robust monitoring
Open the record for dataset details and reuse information.
A novel method to assess the integrity of frozen archival DNA samples: Alpha-diversity ratios of short and long-read 16S rRNA gene sequences
Open the record for dataset details and reuse information.
Aligned DNA sequence matrix for phylogenetic analyses in the article "New species of fossorial salamanders of the genus Oedipina (Plethodontidae) from the northwestern Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "New species of fossorial salamanders of the genus Oedipina (Plethodontidae) from the northwestern Ecuador". The matrix is in NEXUS format.</p> <p>Gene partitions are arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>16S = 4- 789 1761- 1891 ;<br> ND1-codonPos1 = 790-1759\3;<br> ND1-codonPos2 = 791-1760\3;<br> ND1-codonPos3 = 792-1758\3;<br> CytB-codonPos1 = 1893-2274\3;<br> CytB-codonPos2 = 1894-2275\3;<br> CytB-codonPos3 = 1892-2276\3;</p>
Fig. 12 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 12. Known distribution of the species of the T. opinatus subgroup.
Data from: Accumulation curves of environmental DNA sequences predict coastal fish diversity in the Coral Triangle
Environmental DNA (eDNA) has the potential to provide more comprehensive biodiversity assessments particularly for vertebrates in species-rich regions. Yet, this method requires the completeness of a reference database, i.e. a list of DNA sequences attached to each species, which is never met. As an alternative, a diversity of Operational Taxonomic Units (OTUs) can be extracted from eDNA metabarcoding. However, the extent to which the diversity of OTUs provided by a limited eDNA sampling effort can predict regional species diversity is unknown. Here, by modelling OTU accumulation curves of eDNA seawater samples across the Coral Triangle, we obtained an asymptote reaching 1,531 fish OTUs while 1,611 fish species are recorded in the region. Besides, we also accurately predict (R² = 0.92) the distribution of species richness among fish families from OTU-based asymptotes. Thus, the multi-model framework of OTU accumulation curves extends the use of eDNA metabarcoding in ecology, biogeography and conservation.
Data from: Intraspecific DNA contamination distorts subtle population structure in a marine fish: decontamination of herring samples before restriction-site associated (RAD) sequencing and its effects on population genetic statistics
Wild specimens are often collected in challenging field conditions, where samples may be contaminated with the DNA of conspecific individuals. This contamination can result in false genotype calls, which are difficult to detect, but may also cause inaccurate estimates of heterozygosity, allele frequencies, and genetic differentiation. Marine broadcast spawners are especially problematic, because population genetic differentiation is low and samples are often collected in bulk and sometimes from active spawning aggregations. Here, we used contaminated and clean Pacific herring (Clupea pallasi) samples to test (i) the efficacy of bleach decontamination, (ii) the effect of decontamination on RAD genotypes, and (iii) the consequences of contaminated samples on population genetic analyses. We collected fin tissue samples from actively spawning (and thus contaminated) wild herring and non-spawning (uncontaminated) herring. Samples were soaked for 10 minutes in bleach or left untreated, and extracted DNA was used to prepare DNA libraries using a restriction-site associated DNA (RAD) approach. Our results demonstrate that intraspecific DNA contamination affects patterns of individual and population variability, causes an excess of heterozygotes, and biases estimates of population structure. Bleach decontamination was effective at removing intraspecific DNA contamination and compatible with RAD sequencing, producing high-quality sequences, reproducible genotypes, and low levels of missing data. Although sperm contamination may be specific to broadcast spawners, intraspecific contamination of samples may be common and difficult to detect from high-throughput sequencing data, and can impact downstream analyses.
Data from: Concealed by darkness: interactions between predatory bats and nocturnally migrating songbirds illuminated by DNA sequencing
Recently, several species of aerial-hawking bats have been found to prey on migrating songbirds, but details on this behaviour and its relevance for bird migration are still unclear. We sequenced avian DNA in feather-containing scats of the bird-feeding bat Nyctalus lasiopterus from Spain collected during bird migration seasons. We found very high prey diversity, with 31 bird species from eight families of Passeriformes, almost all of which were nocturnally flying sub-Saharan migrants. Moreover, species using tree hollows or nest boxes in the study area during migration periods were not present in the bats' diet, indicating that birds are solely captured on the wing during night-time passage. Additional to a generalist feeding strategy, we found that bats selected medium-sized bird species, thereby assumingly optimizing their energetic cost-benefit balance and injury risk. Surprisingly, bats preyed upon birds half their own body mass. This shows that the 5% prey to predator body mass ratio traditionally assumed for aerial hunting bats does not apply to this hunting strategy or even underestimates these animals' behavioural and mechanical abilities. Considering the bats' generalist feeding strategy and their large prey size range, we suggest that nocturnal bat predation may have influenced the evolution of bird migration strategies and behaviour.
Aligned and trimmed 16S and COI DNA sequences of Oceaniidae (Hydrozoa)
<p>Aligned and trimmed 16S and COI sequences of Oceaniidae (Hydrozoa) used for the study "The polyps of <em>Oceania armata</em> identified by DNA barcoding (Cnidaria, Hydrozoa)"</p> <p>Format is Fasta, files are text files</p>
Application of high-throughput sequencing (HTS) metabarcoding to diatom biomonitoring: Do DNA extraction methods matter?
<p>The 8 benthic samples from Mainland France (stream Edian, stream Aire, lake Geneva), Sweden (stream Dåmman, Agricultural stream, lake Båtkåjaure) and Mayotte (stream Dapani, stream Majimbini) were collected by scraping material from the surface of stones, following the French standard (AFNOR 2007) used in routine biomonitoring programs.DNA was extracted from each sample (2 replicates) using five DNA extraction methods, followed by the amplification of a short rbcL DNA barcode (312bp) specific to diatoms. PCR products were then sequenced in one random direction using the Ion Torrent™ Personal Genome Machine® (PGM) System according to the manufacturer’s instructions. The data file contains one fastq file per library sequenced with the raw DNA reads, as provided by the sequencing platform (demultiplexing performed by the sequencing platform). An excel file is also provided to make the link between the fastq file number and the sample information (sample origin, DNA extraction method used, number of raw reads).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.