Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18,449
datasets available to search
ShareScore release 0.7.1
Dataset results
18,449 results for “DNA”
Multi-locus DNA metabarcoding of western spotted skunk diet in the McKenzie River Ranger District of the Willamette National Forest from 2017-2019
There are increasing concerns about the declining population trends of small mammalian carnivores around the world. Their conservation and management is often challenging due to limited knowledge about their ecology and natural history. To address one of these deficiencies for western spotted skunks (Spilogale gracilis), we investigated their diet in the Oregon Cascades of the Pacific Northwest during 2017 –2019. We collected 130 spotted skunk scats opportunistically and with detection dog teams and identified prey items using DNA metabarcoding and mechanical sorting. Western spotted skunk diet consisted of invertebrates such as wasps, millipedes, and gastropods, vertebrates such as small mammals, amphibians, and birds, and plants such as Gaultheria, Rubus, and Vaccinium. Diet also consisted of items such as black-tailed deer that were likely scavenged. Comparison in diet by season revealed that spotted skunks consumed more insects during the dry season (June –August), particularly wasps (75% of scats in the dry season), and marginally more mammals during the wet season(September –May). We observed similar diet in areas with no record of human disturbance and areas with a history of logging at most spatial scales, but scats collected in areas with older forest within a skunk’s home range (1 km buffer) were more likely to contain insects. Western spotted skunks provide food web linkages between aquatic, terrestrial, and arboreal systems and serve functional roles of seed dispersal and scavenging. Due to their diverse diet and prey-switching, western spotted skunks may dampen the effects of irruptions of prey, such as wasps during dry springs and summers. By studying the natural history of western spotted skunks in the Pacific Northwest forests while they are still abundant, we provide key information necessary to achieve the conservation goal of keeping this common species common.
DNA Origami Raw AFM Data - NanoLocz: Image analysis platform for AFM, high-speed AFM and localization AFM
<p>The data file is in the original ARIS data format as captured on a Cypher VRS1250 AFM (Oxford Instruments)<br><br><br></p>
Arctic-boreal bryophyte dynamics since the last glacial from ancient DNA metabarcoding
<p>A total of 26 lake-sediment cores collected from 26 study sites spanning the glacial and interglacial transition are used in this study. These sites are distributed across Siberia, Beringia, and Alaska regions, with a gradient of vegetation types dominated by tundra in the northern region and transitioning to boreal forest in the southern extents. DNA samples from the sediment core were analysed with a standard sedimentary ancient DNA metabarcoding pipeline (see additional description), which resulted in a raw dataset of all DNA plant sequences, which were then filtered for Bryophytes (Bryophyte DNA dataset). The Bryophyte DNA dataset contains 120 unique ASV. Samples in the Bryophyte DNA dataset are then grouped into 1000-year time slices and are subsequently resampled to a base count of 500 read counts for each time slice. After that, a Bryophyte trait datastet is assigned to the Bryophyte DNA dataset. </p> <p> </p> <h3>Input files</h3> <ul> <li><strong>Excel file with all data used in the R-Script:</strong> "Bryophytes_data.xlsx"</li> <li><strong>WorldClim 2.0 dataset with mean temperatures of Warmest Quarter</strong> (https://www.worldclim.org/; Fick, S.E. and R.J. Hijmans, 2017. WorldClim 2: new 1km spatial resolution climate surfaces for global land areas. <a href="https://rmets.onlinelibrary.wiley.com/doi/abs/10.1002/joc.5086">International Journal of Climatology 37 (12): 4302-4315</a>): "wc2.1_30s_bio_10.tif"</li> </ul> <h3>R script</h3> <ul> <li><strong>R-Script:</strong> "2025-01-14_R-Script_ arctic_boreal_bryophyte_dynamics_DNA_metabarcoding.R"</li> </ul> <h3>R outputs</h3> <ul> <li><strong>resampled Bryophyte metabarcoding percentage dataset with ASV:</strong> "2025-01-14_bryophyta_resampled_percentages_mean_100runs_sequences.csv"</li> <li><strong>resampled Bryophyte metabarcoding percentage dataset with unique scientific names: </strong>"2025-01-14_bryophyta_resampled_percentages_mean_100runs_scientific_names.csv"</li> <li><strong>GBIF taxa occurrences with WorldClim temperature data: </strong>"2025-01-14_gbif_taxa_occurrences_seqtypes_climate.csv"</li> </ul> <p> </p>
Conformations and cryo-force spectroscopy of spray-deposited single-strand DNA on gold: Lifting atomic coordinates
<p>Here we provide the atomic coordinates and the topology file concerning the lifting process of a single stranded DNA molecule previously adsorbed on gold. In order to visualize it you will need a visualization software. Using VMD, you would only need to do in a terminal:</p> <p>vmd -e visualize.vmd </p> <p>and that is it. If you find this useful, please cite the corresponding paper:<br> Nature Communications 10, 685 (2019) [DOI: https://doi.org/10.1038/s41467-019-08531-4 ]</p>
Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study (Supplementary Data)
<p>The zip file contains supplementary data for the publication - Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study, accepted for publication in Environmental Health Perspectives (DOI: 10.1289/EHP6174).</p> <p>The description of the files are noted below:</p> <p><strong>1. Readme File for SAPALDIA Noise and Air Pollution EWAS Single Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_SingleExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_SingleExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p> </p> <p><strong>General footnote for all files:</strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from single exposure epigenome-wide linear mixed models, with random intercept at the level of participant. Each model was adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator (for Lden models) and leukocyte composition. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites.</p> <p>Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p> <p> </p> <p><strong>2. Readme File for SAPALDIA Noise and Air Pollution EWAS Multi Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_MultiExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_MultiExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p><strong>General table footnotes: </strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from multi-exposure epigenome-wide linear mixed models, with random intercept at the level of participant, and were adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator and leukocyte composition. Multi-exposure models included all five exposures (Aircraft, railway, road traffic Lden and respective truncation indicators, NO<sub>2</sub> and PM<sub>2.5</sub>) at the same time. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites. Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p>
Phylogeny of "Philoceanus complex" seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences
<p>Data from "Phylogeny of “<em>Philoceanus </em>complex” seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences". See the file index.html for details. Data includes NEXUS files for sequences, tree files output by MrBayes and PAUP, and host-parasite association files for TreeMap.</p>
Detecting local variations across metazoan communities in backreef depressions of Reunion Island (Mascarene Archipelago) through environmental DNA survey
<p>The back-reef depressions, or lagoons, of Reunion Island (western Indian Ocean) host a high abundance of organisms living amongst the coral reefs and are critical sites for artisanal fishing, tourism, and shoreline stability for the island. Over time, increasing degradation of Reunionese reefs has been observed due to overexploitation, beach erosion and eutrophication. Efforts to mitigate the impact of these pressures on aquatic organisms include biodiversity surveys primarily performed through visual censuses that can be logistically complex and may unintentionally overlook organisms. Surveys integrating environmental DNA (eDNA) collections have provided rapid biodiversity assessments, while helping to circumvent some limitations of visual surveys. The present study describes the results of an exploratory eDNA survey, which aims to characterize metazoan communities of four Reunionese lagoons located along the west coast of the island. As eDNA surveys first require deliberate study design and optimization for each new context, we sought to establish a modernized workflow implementing specialized equipment to collect and preserve samples to facilitate future studies in these lagoons. During the austral summer of 2023, samples were pumped directly from surface and bottom depths at each site through self-preserving filters which were then processed for DNA metabarcoding using regions of the 12S ribosomal RNA (12S), small ribosomal subunit 18S (18S) and Cytochrome Oxidase I (COI) genes. The survey detected high species richness that varied by site, and in a single collection period, recovered the presence of 60 teleost families and numerous invertebrate taxa, including members of the coral faunal community that are less studied in Reunion. Distinct biological communities were observed at each site, and within a single lagoon, suggesting that these differences are due to site-specific factors (e.g., environmental variables, geographic distance, etc.). Although continued protocol optimization is needed, the present findings demonstrate the successful application of an eDNA-based survey for biodiversity assessment within Reunionese lagoons.</p>
Inter-Chemical Correlation results for the study: HHEARx2017-1740 (Mitochondrial DNA biomarkers of prenatal metal mixture exposure: intergenerational inheritance and infant growth)
Title: Mitochondrial DNA biomarkers of prenatal metal mixture exposure: intergenerational inheritance and infant growth <br>Species: Homo sapiens <br>Number of samples: 1423 <br>Number of named analytes: 20 <br>Datasource url: https://hheardatacenter.mssm.edu/PublicFile/ViewPublicFile?projectid=19 <br>
Data from 'Tracability of Forest Reproductive Material with the quality label 'Plant van Hier': A DNA database with genetic profiles of native autochthonous tree and shrub species of Flanders, Belgium'
<h2>Background</h2> <p>Indigenous trees and shrubs play an important role in multifunctional forest management. They form a significant part of the biodiversity in our forests. Forest reproductive material (FRM) of autochthonous Flemish origin is sold under the quality label ‘Plant van Hier’, a certification mark of the Agency for Nature and Forests. To ensure the provenance of the seedlings, we developed a DNA-database of genetic profiles of potential parent trees, using species-specific genetic markers. This database enables the traceability of FRM of the ‘Plant van Hier’ label throughout the entire production chain; from seed harvesting and cultivation to planting by the end user.</p> <p>This database contains the genetic profiles of almost all possible parent trees present within 27 Flemish autochthonous seed orchards of eight ecologically important tree and shrub species: <em>Carpinus betulus</em>, <em>Corylus avellana</em>, <em>Frangula alnus</em>, <em>Populus tremula</em>, <em>Sorbus aucuparia</em>, <em>Tilia cordata</em>, <em>Tilia platyphyllos,</em> and <em>Ulmus laevis</em>. The profiles were established using microsatellite markers (11 to 24 markers per species). New genetic markers were developed for <em>Carpinus betulus</em> and <em>Ulmus laevis</em>. PCR products were run on an ABI 3500 Genetic Analyser (Thermo Fisher Scientific).</p> <h2>Files</h2> <p>The files will be updated when new genotypes are added to the seed orchards. The current data files contain data from genotypes collected in the period 2018-2023. </p> <h3>Species_genotypes</h3> <p>These files contain the genetic fingerprints of the parent trees of autochthonous Flemish seed orchards. Missing data is indicated as ‘MD’. For <em>Carpinus betulus</em>, an octoploid species, the allelic phenotype is given instead of the genotype as the number of times that an allele occurs on a specific locus is not known.</p> <p>The next metadata is additionally given:<br>- Species: the Latin name of the species<br>- Seed_orchard: the name of the seed orchard in which the genotypes are located<br>- Code_seed_orchard: the code of the seed orchard in which the genotypes are located as given in the Register of Flemish Forest Reproductive Material (‘Register bosbouwkundig uitgangsmateriaal’; inbo.be)<br>- Genotype: the fieldname given to the genotype<br>- Origin: the location where the genotype was collected in Flanders, Belgium. Genotypes were collected from natural stands which are assumed to have an autochthonous origin. When the specific location is unknown, the location ‘Flanders’ is given. <br>- Year_sampled: the year in which the genotypes were sampled in the respective seed orchard for genetic analysis.</p> <h3>Species_binsets</h3> <p>These files contain the binsets and allele names that are used to score the alleles of the genotypes in the programme Geneious Prime 2019.3.2 (<a href="https://www.geneious.com">https://www.geneious.com</a>). For <em>Tilia platyphyllos </em>and <em>Tilia cordata</em>, the same binsets were used.</p>
Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing
<p><strong>OV2295 Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz: Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz: Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number </li> <li>minor_cn: HMM predicted minor copy number </li> <li>major_cn: HMM predicted major copy number </li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz: Table of SNVs per clone for OV2295 samples. Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het: is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format. Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: ‘SA922’: ‘OV2295(R2)’, ‘SA921’: ‘TOV2295(R)’, ‘SA1090’: ‘OV2295’,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>
DNA methylation dynamics during stress-response in woodland strawberry (Fragaria vesca)
<p><strong>Genome sequence and annotation of Fragaria vesca cv. Reine des Vallées</strong></p> <p>In order to generate a reference genome for Fragaria vesca cv. Reine des Vallées, we used MinIon long-read sequencing data to substitute the <em>F. vesca</em> genome v.4.0.a2 genome. The detailed method used to obtain these results were the following:</p> <p><em>Genome sequencing and assembly NIL Fb2</em></p> <p>Genomic DNA from strawberry plants was extracted by a Hexadecyltrimethylammonium bromide (Cetrimonium bromide, CTAB) modified protocol (Healey, Furtado, Cooper, & Henry, 2014) and purified with Agencourt AMPure XP beads (cat# A63880). Long-read sequencing was performed for the genome assembly; Genomic DNA by Ligation (Oxford Nanopore, cat# SQK-LSK109) library was prepared as described by the manufacturer and sequenced on a MinION for 72 h (Oxford Nanopore).</p> <p><em>Reference genome polishing</em></p> <p>Reads obtained from nanopore were filtered with Filtlong v0.2.1 (<a href="https://github.com/rrwick/Filtlong">https://github.com/rrwick/Filtlong</a>) using --min_mean_q 80 and --min_length 200. Cleaned reads were then aligned to the most recent version of the <em>F. vesca</em> genome v4.0.a2, downloaded from the Genome Database for Rosaceae (GDR) (<a href="https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2">https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2</a>), using minimap2 v2.21 (H. Li, 2018) with parameters -aLx map-ont --MD -Y. The generated BAM file was then sorted and indexed with samtools v1.11 (H. Li et al., 2009). We used mosdepth v0.3.1 (Pedersen & Quinlan, 2018) to verify that coverage on chromosomic scaffolds was over 50 X. Sniffles v1.0.12a (Sedlazeck et al., 2018) with parameters -s 10 -r 1000 -q 20 --genotype -l 30 -d 1000 was used to detect structural variations larger than 30 bp. The VCF files obtained from Sniffles was sorted and filtered with BCFtools v1.14 (Danecek et al., 2021) to keep only structural variants (SV) with smaller than 200,00 bp (we observed that larger SV were most of the time false positive caused by misalignments in regions with gaps or Ns), supported by 10 or more reads and with allelic frequencies above 0.8 (we were interested in homozygous changes). The complete filtering command used is “bcftools view -q 0.8 -Oz -i '(SVTYPE = "DUP" || SVTYPE = "INS" || SVTYPE = "DEL" || SVTYPE = "TRA" || SVTYPE = "INV" || SVTYPE = "INVDUP") && %FILTER = "PASS" && FMT/DV>9 && SVLEN>29 && SVLEN<200000' “</p> <p>From the VCF listing all the structural variants that we detected in our <em>F. vesca </em>accession, we generated a substituted genome version based on the reference <em>F. vesca</em> genome v.4.0.a2. The reference genome was first indexed with samtools faidx v1.11(Danecek et al., 2021) and a sequence dictionary was generated with Picard CreateSequenceDictionary v2.25.6 (<a href="https://broadinstitute.github.io/picard">https://broadinstitute.github.io/picard</a>). The VCF containing the SV produced from our Nanopore sequencing was also indexed with gatk (Van der Auwera GA & O'Connor BD, 2020) IndexFeatureFile v4.2.0.0 (<a href="https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile">https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile</a>). FastaAlternateReferenceMaker v4.2.0.0 (<a href="https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker">https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker</a>) was then run with the reference genome and the VCF file to generate a substituted genome representative of our <em>Fragaria</em> accession.</p> <p>As substituting our genome with the detected structural variants changes genomic coordinates, we also corrected the public GFF genome annotation of <em>F. vesca</em> (Y, Pi, Gao, Liu, & Kang, 2019) using liftoff v1.6.1 (Shumate & Salzberg, 2021). Liftoff also detects and annotates duplications within the substituted genome.</p> <p>Transposable elements annotation was carried out using the EDTA transposable element annotation pipeline v. 1.9.6 (S. Ou et al., 2019) on the substituted genome using default parameters<em>.</em></p> <p><strong>Differentially methylated regions</strong></p> <p>The file Stress_vs_control_DMRs.zip file contains the DMRs that were called using the reads submitted to ENA (ERP135585) and obtained as follows:</p> <p>First, bedGraph files from wgbs pipeline were pre-filtered for a minimum coverage of 5 reads using awk command. These output files were then used as input for the EpiDiverse/dmr bioinformatics analysis pipeline for non-model plant species to define DMRs (Nunn <em>et al</em>., 2021) with default parameters (minimum coverage threshold 5; maximum q-value 0.05; minimum differential methylation level 10%; 10 as minimum number of Cs; Minimum distance (bp) between Cs that are not to be considered as part of the same DMR is 146 bp). The pipeline uses metilene v.0.2.6.1 (<a href="https://www.bioinf.uni-leipzig.de/Software/metilene/">https://www.bioinf.uni-leipzig.de/Software/metilene/</a>) for pairwise comparison between groups and R-packages ggplot2 v.3.3.5 and gplots v.3.1.1, for visualization results (Fig. S1). Based on our <em>F. vesca</em> genome transcript annotation and methylation data (overlapped regions with DNA methylation cytosines and DMRs), we detected the methylated genes, promoters, 3’ UTRs, 5’UTR and transposable elements in strawberry. Global DNA methylation and DMR plots were performed with R-package ggplot2. Gene analyses by methylation patterns and analysis of per-family TE DNA methylation profiles were performed with deepTools v.3.5.0 (Ramírez <em>et al</em>., 2014). DMRs comparison between treatments were done by the Venn diagram v.1.7.0 R-package.</p> <p>We produced several genome browsers tracks with DMRs that we integrated in our local instance of JBrowse available at the following url: <a href="https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub">https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub</a></p>
Minimal dataset to test multiplexed DNA imaging (Hi-M) software pipelines
<p>This is a dataset of nuclei (DAPI), and 3 multiplexed DNA imaging cycles to test and validate processing software packages, such as pyHiM (https://github.com/marcnol/pyHiM). This dataset was acquired in a nc14 Drosophila embryo.</p> <p>File contents:</p> <p>scan_001_RT27_001_ROI_converted_decon_ch00.tif barcode 27, fiducial <br> scan_001_RT27_001_ROI_converted_decon_ch01.tif barcode 27<br> scan_001_RT29_001_ROI_converted_decon_ch00.tif barcode 29, fiducial <br> scan_001_RT29_001_ROI_converted_decon_ch01.tif barcode 29 <br> scan_001_RT37_001_ROI_converted_decon_ch00.tif barcode 37, fiducial <br> scan_001_RT37_001_ROI_converted_decon_ch01.tif barcode 37 <br> scan_006_DAPI_001_ROI_converted_decon_ch00.tif DAPI <br> scan_006_DAPI_001_ROI_converted_decon_ch01.tif DAPI, fiducial <br> scan_006_DAPI_001_ROI_converted_decon_ch02.tif RNA</p> <p> </p> <p>To test this dataset please refer to <a href="https://github.com/marcnol/pyHiM">pyHiM documentation page</a>.</p>
Slimfield: Escherichia coli DNA repair proteins (RecA-mGFP and RecB-sfGFP)
<p>Imaging modality / instrument: <em>Brightfield</em> + <em>Slimfield</em></p> <p>Image format:<em> OME TIFF (16 bit) + MicroManager metadata files</em></p> <p>Microscope settings:</p> <p><em>488 nm triggered excitation; split detection, cropped to GFP or RFP/GFP (left/right) channels; 3 ms/frame laser exposure; Photometrics Prime95b CMOS</em></p> <p>Samples and acquisitions:</p> <p>Fluorescent fusions in live E.coli cells. MMC = mitomycin C</p> <table> <tbody> <tr> <td> <p>No. fields of view</p> </td> <td> <p>MMC-</p> </td> <td> <p>MMC+ (0.5 ug/ml 3h)</p> </td> </tr> <tr> <td> <p>RecA-mGFP</p> </td> <td> <p>7</p> </td> <td> <p>15</p> </td> </tr> <tr> <td> <p>RecB-sfGFP</p> </td> <td> <p>17</p> </td> <td> <p>21</p> </td> </tr> <tr> <td> <p>MG1655 control</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> </tr> </tbody> </table> <p>Approx. size before/after compression: 21 GB / 7 GB</p>
Multiplexed DNA-FISH imaging dataset, drosophila embryos, nuclear cycles 11-14
<p>Multiplexed DNA-FISH imaging dataset from Drosophila embryos at nuclear cycles 11-14.</p> <p>Examples on how to load and use this dataset can be found at this <a href="https://github.com/NollmannLab/Goetz_etal">GitHub repository</a>.</p> <p><strong>Data processing details</strong></p> <p>Barcodes were segmented using a neural network (<a href="https://github.com/stardist/stardist"><em>stardist</em></a>) specifically trained for the detection of 3D diffraction limited spots produced by our microscope. To extract the position of the barcode with sub-pixel accuracy, a subsequent 3D Gaussian fit of the regions segmented by <em>stardist</em> was performed with Big-FISH (<a href="https://github.com/fish-quant/big-fish">https://github.com/fish-quant/big-fish</a>). Barcode localizations with intensities lower than 1.5 times that of the background were filtered out.</p> <p>Nuclei were segmented from projected DAPI images using <em><a href="https://github.com/stardist/stardist">stardist</a> </em>with a neural network trained for detection of nuclei from <em>Drosophila</em> embryos under our imaging conditions. Barcodes were then attributed to single nuclei by using the XY coordinates of the barcodes and the DAPI masks of the nuclei. Finally, pairwise distance matrices were calculated for each single nucleus.</p> <p><strong>Processed data in Figures</strong></p> <p>This new version of the dataset contains the raw data for each of the figures in the manuscript:</p> <p><strong>Associated publication</strong></p> <p><strong>Multiple parameters shape the 3D chromatin structure of single nuclei at the doc locus in </strong><em>Drosophila</em>.</p> <p>Markus Götz, Olivier Messina, Sergio Espinola, Jean-Bernard Fiche, Marcelo Nollmann</p> <p>Nature Communications (2022).</p>
Begomovirus DNA-B Movement Protein and Nuclear Shuttle Protein Ref-Seq Datasets
<p>Multiple Sequence Alignment of proteins encoded on begomovirus DNA-B (movement protein and nuclear shuttle protein; n=131). One isolate per species based on ICTV reference list.</p> <p> </p>
Arctic specimens in the NHMO DNA bank Vascular plants collection 2022
<p>All Arctic specimens in the NHMO DNA bank Vascular plants collection as of August 2022. See Bjorå et al. 2023 "Collections of Arctic<br> plants, lichens and fungi in the Natural History Museum, University of Oslo, Norway" for further details.</p>
Arctic specimens in the NHMO DNA bank Fungi & Lichens collection 2022
<p>All Arctic specimens in the NHMO DNA bank Fungi & Lichens collection as of August 2022. See Bjorå et al. 2023 "Collections of Arctic plants, lichens and fungi in the Natural History Museum, University of Oslo, Norway" for further details.</p>
Graphic Illustration of Neal Platt's Talk: Targeted sequencing of pathogen DNA from museum specimens
<p><a href="https://lib.ku.edu/people/courtney-foat" target="_blank" rel="noopener">Courtney Foat</a>, Advisor for Strategic Initiatives & Organizational Engagement at the University of Kansas, graphically recorded this invited talk by Neal Platt at an NSF-supported Workshop: Digital Collections Data and Tracking Disease.</p>
Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'
<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions. </p> <p> </p>
AlphaFold2 models of human DNA polymerase epsilon catalytic subunit A
<p>Models of the structure of human DNA polymerase epsilon catalytic subunit A built with the program AlphaFold2. Files are in mmCIF format. Models include:</p> <p>1) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_NterminalLobe_4m8o.cif">DPOE1_HUMAN_AF2_NterminalLobe_4m8o.cif </a>: AlphaFold2 model built with template PDB:4M8O (yeast DNA polymerase epsilon N-terminal lobe with bound DNA)</p> <p>2) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_FullLength_6wjv.cif">DPOE1_HUMAN_AF2_FullLength_6wjv.cif </a>: AlphaFold2 model built with template PDB:6WJV (yeast DNA polymerase epsilon full-length protein, N and C terminal lobes, without DNA).</p> <p>3) <a href="https://zenodo.org/api/files/c8652c0b-1bc6-44d5-b11e-e8e107dba67e/DPOE1_HUMAN_AF2_NterminalLobe_4m8o_withDNA.cif">DPOE1_HUMAN_AF2_NterminalLobe_4m8o_withDNA.cif</a> : AlphaFold2 model built with template 4M8O (File #1 above) with DNA added from PDB entry 4M8O and subjected to energy minimization with the program AMBER using force field ff14SB.</p> <p>The models were used to estimate the change in free energy of mutations found in patients with ovarian cancer, colon cancer, and endometrial cancer. Paper to be submitted Dec 2022.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.