Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad36/100

List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)

<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>

opencc-zeroDec 2020View details →
dryad36/100

Measuring phylogenetic information of incomplete sequence data

<p>Widely used approaches for extracting phylogenetic information from aligned sets of molecular sequences rely upon probabilistic models of nucleotide substitution or amino-acid replacement. The phylogenetic information that can be extracted depends on the number of columns in the sequence alignment and will be decreased when the alignment contains gaps due to insertion or deletion events. Motivated by the measurement of information loss, we suggest assessment of the Effective Sequence Length (ESL) of an aligned data set. The ESL can differ from the actual number of columns in a sequence alignment because of the presence of alignment gaps. Furthermore, the estimation of phylogenetic information is affected by model misspecification. Inevitably, the actual process of molecular evolution differs from the probabilistic models employed to describe this process. This disparity means the amount of phylogenetic information in an actual sequence alignment will differ from the amount in a simulated data set, which motivated us to develop a new test for model adequacy. Via theory and empirical data analysis, we show how to disentangle the effects of gaps and model misspecification. By comparing the Fisher information of actual and simulated sequences, we identify which alignment sites and tree branches are most affected by gaps and model misspecification.</p>

opencc-zeroSep 2021View details →
zenodo36/100

Input Data for "Protein Function Prediction for newly sequenced organisms"

<p>The input sequence files in FASTA format and the detailed list of all organisms excluded when testing each specific bacterium.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Extended data table 1: Taqman array card results showing all individual target hits with Ct values and whether validated by conventional microbiology and/or microbial sequencing.

<p>Data from study, protocol published at&nbsp;10.5281/zenodo.5081880</p> <p><strong>Extended Data Table 1: TAC results showing all individual target hits with Ct values and whether validated by conventional microbiology and/or microbial sequencing.&nbsp;</strong>(BAL:&nbsp;bronchoalveolar lavage.) *Not included in validation numbers as duplicate at sub-species or genus level detection, **MecA was not included in validation numbers. TAC hits which&nbsp;did not pass the internal quality control standards required for reporting are indicated by (not reported). Samples which did not undergo sequencing indicated by&nbsp;<em>ND</em>.&nbsp;(✓) indicates low confidence hits from sequencing.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

AliSim: Ultrafast and Realistic Sequence Alignment Simulator for Phylogenetics - Supplementary Data

<p>This supplementary data&nbsp;contains&nbsp;testing scripts, input/output data for validating and benchmarking AliSim.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Rock-magnetic, paleomagnetic and multimethod paleointensity data from an upper Miocene lava flow sequence from São Vicente (Cape Verde)

<p>The folder &ldquo;Praia Grande Paleomagnetism Data.zip&rdquo; contains paleomagnetic thermal and alternating field demagnetisation data obtained on upper Miocene volcanic rocks from S&atilde;o Vicente (Cape Verde). Measurements were performed in the paleomagnetic laboratory of the University of Burgos (Spain). Data are in .txt format with the extension .rs3. Columns are separated by empty spaces. Data can be visualised and analysed with the Remasoft software (Chadima and Hrouda, 2006).</p> <p>The folder &ldquo;Praia Grande rock magnetism Data.zip&rdquo; contains data in .txt format of IRM acquisition curves (extension .irm), hysteresis curves (extension .hys), backfield curves (extension .coe) and thermomagnetic magnetisation versus temperature curves (extension .rmp) obtained on upper Miocene volcanic rocks from S&atilde;o Vicente (Cape Verde). Measurements were performed in the paleomagnetic laboratory of the University of Burgos (Spain). Columns are separated by tabs. Data can be visualised and analysed with the RockMagAnalyzer 1.0 software (Leonhardt, 2006).</p> <p>The folder &ldquo;Praia Grande Thellier-Coe Data.zip&rdquo;contains paleointensity determination data obtained with the Thellier-Coe method on upper Miocene volcanic rocks from S&atilde;o Vicente (Cape Verde). Measurements were performed in the paleomagnetic laboratory of the University of Burgos (Spain). Data are in .txt format with the extension .tdt separated by tabs. Data can be visualised and analysed with the ThellierTool software (Leonhardt et al.,2004).</p> <p>The folder &ldquo;Praia Grande Multispecimen.zip&rdquo; contains paleointensity determination data on upper Miocene volcanic rocks from S&atilde;o Vicente (Cape Verde) obtained with the multispecimen method (Biggin and Poidras, 2006; Dekkers and B&ouml;hnel, 2006; Fabian and Leonhardt, 2010) at Laboratorio Interinstitucional de Magnetismo Natural, Instituto de Geof&iacute;sica, Unidad Michoac&aacute;n, UNAM, Mexico. Two kinds of files can be found: Five .txt files (M0.txt, M1.txt, M2.txt, M3.txt, M4.txt) and five .jr6 files (M0.jr6, M1.jr6, M2.jr6, M3.jr6, M4.jr6). Both types of files are in .txt format. The .txt files directly provide the measurement data generated during the multispecimen experiments. The first columns of the .jr6 files display the following information: column 1: specimen name; column 2: experimental step (explanation below); columns 3, 4 and 5: three magnetisation components M(x), M(y) and M(z); column 6: exponent of the magnetisation components to obtain the magnetisation value in A/m.&nbsp;In the multispecimen experiments, the specimen-name (e.g.&nbsp; PRG1-3) can be divided in two parts, the first three characters indicate the sample (flow), the last character the specimen number (1 to 7). Eight different experimental steps (column 2) can be distinguished: NRM, A10, A20, A30, A40, A50, A60 and A70.&nbsp; They correspond to the measurement of the NRM and of steps in which fields of 10, 20, 30, 40, 50, 60 and 70 mT where respectively applied. Files M0 (.txt and .jr6) include only NRM measurements, files M1 include measurements of specimens heated with an applied field parallel to their NRM, files M2 include measurements of specimens heated with an applied field antiparallel to their NRM, files M3 include measurements of specimens heated in zero field and cooled down in an applied field parallel to their NRM, and files M4 include again include measurements of specimens heated with an applied field parallel to their NRM.</p> <p><strong>REFERENCES</strong></p> <p>Biggin, A., Poidras, T., 2006. First-order symmetry of weak-field partial thermoremanence in multi-domain ferromagnetic grains. 1. Experimental evidence and physical implications. Earth Planet. Sci. Lett. 245, 438&ndash;453. doi:10.1016/j.epsl.2006.02.035</p> <p>Chadima, M. and Hrouda, F., 2006. Remasoft 3.0 a user friendly paleomagnetic data browser and analyzer. <em>Travaux G&eacute;ophysiques</em>, XXVII, 20-21.</p> <p>Dekkers, M.J., B&ouml;hnel, H.N., 2006. Reliable absolute palaeointensities independent of magnetic domain state. Earth Planet. Sci. Lett. 248, 507&ndash;516. doi:10.1016/j.epsl.2006.05.040</p> <p>Fabian, K., Leonhardt, R., 2010. Multiple-specimen absolute paleointensity determination: An optimal protocol including pTRM normalization, domain-state correction, and alteration test. Earth Planet. Sci. Lett. 297, 84&ndash;94. doi:10.1016/j.epsl.2010.06.006</p> <p>Leonhardt, R., 2006. Analyzing rock magnetic measurements; The RockMagAnalyzer 1.0 software.<em>Computers and Geosciences</em>, 32, 1420-1431.</p> <p>Leonhardt, R., Heunemann, C. and Kr&aacute;sa, D., 2004. Analyzing absolute paleointensity determinations: Acceptance criteria and the software ThellierTool4.0. <em>Geochem. Geophys. Geosyst.</em>, Vol. 5, no. 12, doi.: 10.1029/2004GC000807.</p> <p>Monster, M.W.L., de Groot, L. V., Dekkers, M.J., 2015. MSP-Tool: A VBA-Based Software Tool for the Analysis of Multispecimen Paleointensity Data. Front. Earth Sci. 3, 1&ndash;9. https://doi.org/10.3389/feart.2015.00086</p>

opencc-by-4.0Oct 2021View details →
dryad36/100

Emergence and radiation of distemper viruses in terrestrial and marine mammals - Input files, bash and R codes for analysing PDV and CDV sequence data

<p><span>Canine distemper virus (CDV) and phocine distemper virus (PDV) are major pathogens to terrestrial and marine mammals. Yet little is known about the timing and geographical origin of distemper viruses and to what extent it was influenced by environmental change and human activities. To address this, we i) performed the first comprehensive time-calibrated phylogenetic analysis of the two distemper viruses; ii) mapped distemper antibody and virus detection data from marine mammals collected between 1972-2018; iii) and compiled historical reports on distemper dating back to the 18<sup>th</sup> century. We find that CDV and PDV diverged in the early 17<sup>th</sup> century. Modern CDV strains last shared a common ancestor in the 19<sup>th</sup> century with a marked radiation during the 1930s-50s. Modern PDV strains are of more recent origin, diverging in the 1970s-80s. Based on the compiled information on distemper distribution, the diverse host range of CDV and basal phylogenetic placement of terrestrial morbilliviruses, we hypothesize a terrestrial CDV-like ancestor giving rise to PDV in the North Atlantic. Moreover, given the estimated timing of distemper origin and radiation, we hypothesize a prominent role of environmental change such as the Little Ice Age, and human activities like globalisation and war in distemper virus evolution. </span></p>

opencc-zeroOct 2021View details →
zenodo36/100

epicPCR sequencing data, October 20th

<p>Samples:</p> <p>Rhodo100dilWWMocksBC10e8SM16S<br> Rhodo100dilWWMocksBC10e8SM16S<br> Rhodo100dilWWMocksBC10e8SM18S<br> Rhodo100dilWWMocksBC10e8SM18S<br> Rhodo10dilWWMocksBC10e8SM16S<br> Rhodo10dilWWMocksBC10e8SM16S<br> Rhodo10dilWWMocksBC10e8SM18S<br> Rhodo10dilWWMocksBC10e8SM18S<br> RhodoMocksBC10e7MC16S<br> RhodoMocksBC10e7MC16S<br> RhodoMocksBC10e7MC18S<br> RhodoMocksBC10e7MC18S<br> RhodoMocksBC10e8SM16S<br> RhodoMocksBC10e8SM16S<br> RhodoMocksBC10e8SM18S<br> RhodoMocksBC10e8SM18S<br> RhodoNoMocksBC10e7MC16S<br> RhodoNoMocksBC10e7MC16S<br> RhodoNoMocksBC10e7MC18S<br> RhodoNoMocksBC10e7MC18S<br> RhodoWWMocksBC10e8SM16S<br> RhodoWWMocksBC10e8SM16S<br> RhodoWWMocksBC10e8SM18S<br> RhodoWWMocksBC10e8SM18S<br> WWMocksBC10e7MC16S<br> WWMocksBC10e7MC16S<br> WWMocksBC10e7MC18S<br> WWMocksBC10e7MC18S<br> WWMocksBC10e8SM16S<br> WWMocksBC10e8SM16S<br> WWMocksBC10e8SM18S<br> WWMocksBC10e8SM18S<br> WWNoMocksBC10e7MC16S<br> WWNoMocksBC10e7MC16S<br> WWNoMocksBC10e7MC18S<br> WWNoMocksBC10e7MC18S</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

TIL1383I TCR mutation sequencing and SPR binding data

<p>Sequence and mutation sequence data for the TIL1383I TCR and surface plasmon resonance data and analysis files for TIL1383I binding to tyrosinase/HLA-A2.</p>

opencc-by-4.0Oct 2022View details →
dryad36/100

Illumina next generation ddRAD sequencing SNP data from: Contrasting genetic diversity and structure between endemic and widespread damselfishes are related to differing adaptive strategies

<p class="MsoNormal"><strong><u><span>Aim:</span></u></strong><span> Discerning when, where, and how processes of isolation lead to differing biogeography is especially complex for marine species with similar ecological niches and within the same geographic location. We assessed population genetics of congeneric and ecologically similar damselfishes within their overlapping distributions and across potential barriers to geneflow.</span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Taxon:</span></u></strong><span> <em>Dascyllus marginatus </em>(endemic) and <em>Dascyllus abudafur </em>(widespread)<em>.</em></span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Location:</span></u></strong><span> Coral reefs from the Red Sea, Djibouti, Yemen, Oman, and Madagascar. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Methods:</span></u></strong><span> We used RADseq derived SNPs to investigate key differences in population genetics between both species and discuss barriers shaping genetic differentiation (neutral vs. selective) and biogeography. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Results:</span></u></strong><strong><span> </span></strong><em><span>Dascyllus marginatus </span></em><span>inhabited the Red Sea, the coasts of Yemen (including Socotra), and the Gulf of Oman. <em>Dascyllus abudafur</em> species was present from the Red Sea to Madagascar but was absent from Yemen and Oman. Populations of <em>D. marginatus </em>had an order of magnitude higher genetic differentiation compared to <em>D. abudafur</em>, as well as several outlier loci (suggesting selective pressure), which were absent in <em>D. abudafur</em> despite equal sampling locations. In both species, specimens from the Red Sea and Djibouti formed one genetic cluster separated from all other locations.  </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Main conclusions:</span></u></strong><span> The stronger genetic structure at smaller geographic scale of the endemic species seems associated to faster adaptation to environmental differences; whereas the widespread species only experienced reduced geneflow and neutral differentiation at much larger geographic scales. Restrictive transitions (between the Gulf of Aqaba and the Red Sea or the Red Sea and the Gulf of Aden) did not affect the genetic architecture of either species, while the environmental shift within the Red Sea (at 22°N/20°N) affected the endemic but not the widespread species. Samples from continental Yemen revealed that a genetic break in the Gulf of Aden likely reflects historical colonization processes and not contemporary environmental regimes.</span></p>

opencc-zeroOct 2022View details →
dryad36/100

The topological nature of tag jumping in environmental DNA metabarcoding studies (sequencing raw data)

<p>Metabarcoding of environmental DNA constitutes a state-of-the-art tool for environmental studies. One fundamental principle implicit in most metabarcoding studies is that individual sample amplicons can still be identified after being pooled with others – based on their unique combinations of tags – during the so-called demultiplexing step that follows sequencing. Nevertheless, it has been recognized that tags can sometimes be changed (i.e. tag jumping), which ultimately leads to sample crosstalk. Here, using four DNA metabarcoding datasets derived from the analysis of soils and sediments, we show that tag jumping follows very specific and systematic patterns. Specifically, we find a strong correlation between the number of reads in blank samples and their topological position in the tag matrix (described by vertical and horizontal vectors). This observed spatial pattern of artefactual sequences could be explained by polymerase activity, which leads to the exchange of the 3' tag of single stranded tagged sequences through the formation of heteroduplexes with mixed barcodes. Importantly, tag jumping substantially distorted our datasets – despite our use of methods suggested to minimize this error. We developed a topologic model to estimate the noise based on the counts in our blanks, which suggested that 40-80% of the taxa in our soil and sedimentary samples were likely false positives introduced through tag jumping. We highlight that the amount of false positive detections caused by tag jumping strongly biased our community analyses. </p>

opencc-zeroNov 2022View details →
zenodo36/100

Saved model and preprocessed data for "CRMnet:a deep learning model for predicting gene expression from large regulatory sequence datasets"

<p>Saved TUNet model and preprocessed training data&nbsp;for &quot;CRMnet: a deep learning model for predicting&nbsp;gene expression from large regulatory&nbsp;sequence datasets&quot;</p> <p>To load the trained model:</p> <pre><code class="language-python">import tensorflow as tf tf.keras.models.load_model("path to the model folder")</code></pre> <p>for more information please find our repository:&nbsp;https://github.com/jiayuwen/CRMnet</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Genome-wide RAD sequencing data suggest predominant role of vicariance in Sino-Japanese disjunction of the monotypic genus Conandron (Gesneriaceae)

<p>Disjunct distribution is a key issue in biogeography and ecology, but it is often difficult to determine relative roles of dispersal vs. vicariance in disjunctions. We studied phylogeographic pattern of the monotypic <em>Conandron</em> <em>ramondioides</em> (Gesneriaceae), which shows Sino-Japanese disjunctions, with ddRAD sequencing based on a comprehensive sampling of 11 populations from mainland China, Taiwan Island, and Japan. We found a very high degree of genetic differentiation among these three regions, with very limited gene flow and a clear Isolation by Distance pattern. Mainland China and Japan clades diverged first from a widespread ancestral population in the middle Miocene, followed by a later divergence between mainland China and Taiwan Island clades in the early Pliocene. Three current groups have survived in various glacial refugia during the Last Glacial Maximum (LGM), and experienced contraction and/or bottlenecks since their divergence during Quaternary glacial cycles, with strong niche divergence between mainland China + Japan and Taiwan Island ranges. Thus, we verified a predominant role of vicariance in the current disjunction of the monotypic genus <em>Conandron</em>. The sharp phylogenetic separation, ecological niche divergences among these three groups and the great number of private alleles in all populations sampled indicate a considerable time of independent evolution and suggest the need for a taxonomic survey to detect potentially overlooked taxa.</p>

opencc-zeroDec 2022View details →
dryad36/100

Flow cytometry YFP and CFP data and deep sequencing data of populations evolving in galactose

<p><span>Copy-number and point mutations form the basis for most evolutionary novelty through the process of gene duplication and divergence. While a plethora of genomic sequence data reveals the long-term fate of diverging coding sequences and their cis-regulatory elements, little is known about the early dynamics around the duplication event itself. In microorganisms, selection for increased gene expression often drives the expansion of gene copy-number mutations, which serves as a crude adaptation, prior to divergence through refining point mutations. Using a simple synthetic genetic system that allows us to distinguish copy-number and point mutations, we study their early and transient adaptive dynamics in real-time in <em>Escherichia</em> <em>coli</em>. We find two qualitatively different routes of adaptation depending on the level of functional improvement selected for: In conditions of high gene expression demand, the two types of mutations occur as a combination. Under</span><span> low gene expression demand, negative epistasis between the two types of mutations renders them mutually exclusive. Thus, owing to their higher frequency, adaptation is dominated by copy-number mutations. Ultimately, due to high rates of reversal and pleiotropic cost, copy-number mutations may not only serve as a crude and transient adaptation but also <a>constrain</a></span><span> sequence divergence over evolutionary time scales.</span></p>

opencc-zeroDec 2022View details →
dryad36/100

Data from: Phylogenomics of superrosids and core rosids based on nuclear sequences and synteny

<p class="MsoListParagraph">Superrosids form one of the largest clades of angiosperms, including 18 orders (Vitales, Saxifragales and core rosids) which exhibits remarkable morphological and ecological diversity. However, phylogenetic relationships within superrosids remain unclear.</p> <p class="MsoListParagraph">To resolve the phylogeny of superrosids, we screened 122 single copy nuclear genes from 37 species, representing all 18 orders.</p> <p class="MsoListParagraph">Vitales was revealed as sister to all other superrosids. Within core rosids, the fabids should be restricted only to the nitrogen-fixing clade, while Picramniales, the CM clade, Huerteales, Oxalidales, Sapindales, Malvales and Brassicales composed an "expanded" malvids. The COM clade (sensu APG IV) did not form a monophyletic group. Crossosomatales, Geraniales, Myrtales and Zygophyllales did not belong to either malvids or fabids. The difficult phylogeny of superrosids is likely due to the combined effects of ancient reticulation and incomplete lineage sorting.</p> <p class="MsoListParagraph">To provide broader genomic representation of Saxifragales, we constructed a high-quality chromosome-level genome assembly for <em>Tiarella polyphylla</em> (Saxifragaceae). Whole genome microsynteny analysis of superrosids showed that Saxifragales shared more synteny clusters with core rosids than Vitales, which also indicated that Saxifragales has a closer relationship with core rosids.</p> <p class="MsoListParagraph">Our findings contribute to a better understanding of the phylogeny and evolution of angiosperms.</p>

opencc-zeroJan 2023View details →
dryad36/100

The phylogeny and global biogeography of Primulaceae based on high-throughput DNA sequence data

<p>The angiosperm family Primulaceae is morphologically diverse and distributed nearly worldwide. However, phylogenetic uncertainty has limited the ability to identify major morphological and biogeographic transitions. We used target capture sequencing with the Angiosperms353 kit for over 300 species across Ericales, tree-based sequence curation, and multiple phylogenetic approaches to investigate the phylogenetics of the major clades of Primulaceae and their relationship to other Ericales. The study included 150 samples of Primulaceae comprising nearly all recognized genera of the family, with a particular focus on the most diverse subfamily, Myrsinoideae, for which previous phylogenetic knowledge was poor. We used fossil and secondary calibrations to generate dated phylogenetic trees and conducted broad-scale biogeographic analyses as well as ancestral state reconstructions of plant habit.</p>

opencc-zeroDec 2022View details →
dryad36/100

Data from: Higher evolutionary dynamics of gene copy number for Drosophila glue genes located near short repeat sequences

<p><strong>Background</strong></p> <p>During evolution, genes can experience duplications, losses, inversions and gene conversions. Why certain genes are more dynamic than others is poorly understood. Here we examine how several <em>Sgs</em> genes encoding glue proteins, which make up a bioadhesive that sticks the animal during metamorphosis, have evolved in <em>Drosophila</em> species.</p> <p><strong>Results</strong></p> <p>We examined high-quality genome assemblies of 24 <em>Drosophila</em> species to study the evolutionary dynamics of four glue genes that are present in <em>D. melanogaster</em> and are part of the same gene family <em>–</em> <em>Sgs1, Sgs3, Sgs7 and Sgs8 –</em> across approximately 30 millions of years. We annotated a total of 102 <em>Sgs</em> genes and grouped them into 4 subfamilies. We present here a new nomenclature for these <em>Sgs</em> genes based on protein sequence conservation, genomic location and presence/absence of internal repeats. Two types of glue genes were uncovered. The first category (<em>Sgs1, Sgs3x, Sgs3e</em>) showed a few gene losses but no duplication, no local inversion and no gene conversion. The second group (<em>Sgs3b, Sgs7, Sgs8</em>) exhibited multiple events of gene losses, gene duplications, local inversions and gene conversions. Our data suggest that the presence of short "new glue" genes near the genes of the latter group may have accelerated their dynamics.</p> <p><strong>Conclusions</strong></p> <p>Our comparative analysis suggests that the evolutionary dynamics of glue genes is influenced by genomic context. Our molecular, phylogenetic and comparative analysis of the four glue genes <em>Sgs1, Sgs3, Sgs7</em> and <em>Sgs8 </em>provides the foundation for investigating the role of the various glue genes during <em>Drosophila</em> life.</p>

opencc-zeroJan 2023View details →
dryad36/100

MinION sequencing data of mtDNA from BH10 cells

<p><span>Mitochondrial DNA (mtDNA) recombination in animals has remained enigmatic because of its uniparental inheritance and subsequent homoplasmic state, which excludes the biological need for genetic recombination, as well as limits tools to study it. However, molecular recombination is an important genome maintenance mechanism for all organisms, most notably being required for double-strand break repair. To demonstrate the existence of mtDNA recombination, we have taken advantage of a cell model with two different types of mitochondrial genomes and impaired ability to turn over broken mtDNA. The resulting excess of linear DNA fragments caused increased formation of cruciform mtDNA, appearance of heterodimeric mtDNA complexes and recombinant mtDNA genomes, detectable by Southern blot. Combining our observations with previously published work, we propose that the mitochondrial replisome can catalyze microhomology-mediated recombination of linear mtDNA ends, thus rendering a specialized mitochondrial recombinase unnecessary. The error-proneness of this system is likely to contribute to the formation of pathological mtDNA rearrangements.</span></p>

opencc-zeroFeb 2023View details →
dryad36/100

Perianth evolution and implications for generic delimitation in the Eucalypts (Myrtaceae): DNA sequences, morphological data

<p><em>Eucalyptus</em> was traditionally defined by the operculate perianth—hence the generic name (Latin, meaning "well-covered"). But after previous phylogenetic analysis placed <em>Angophora</em>, which has free sepals and petals, as sister to the bloodwood eucalypts, the latter were segregated into a new genus, <em>Corymbia</em>. We made a targeted capture of 101 low-copy nuclear exons from 392 samples representing 329 species-level taxa. The phylogeny was estimated using maximum likelihood (IQtree and RAxML) and the multi-species coalescent (Astral). We tested alternative relationships between four genera within Eucalypteae (<em>Arillastrum</em>, <em>Angophora</em>, <em>Eucalyptus</em>, <em>Corymbia</em>) at each of two nodes critical to generic delimitation using Shimodaira's Approximately Unbiased (AU) test. Monophyly of <em>Arillastrum</em> + (<em>Corymbia</em> + <em>Angophora</em>) relative to <em>Eucalyptus</em> sensu stricto was supported whereas monophyly of <em>Corymbia</em> relative to <em>Angophora</em> was decisively rejected. These results indicate that either <em>Eucalyptus</em> should be expanded to include all four genera or <em>Corymbia</em> should be split into two. All of the alternative relationships among the four currently recognised genera imply homoplasy in perianth evolution, specifically with respect to origins of the bud cap (operculum or calyptra), which has been traditionally used to define <em>Eucalyptus</em>. Inferred evolutionary transitions in perianth traits are generally congruent with divergences between major clades with a single exception: expression of separate sepals and petals in <em>Angophora</em>, which is nested within the operculate genus <em>Corymbia</em>, appears prima facie to be a reversal to the plesiomorphic perianth structure. Strictly, this is not a reversal because the petals of <em>Angophora</em> and <em>Corymbia</em> have a novel compound keel-and-limb structure that is absent in the outgroups. This structure is evident in early development, irrespective of whether the petals remain free or later become part of an operculum. Many of the currently recognised infrageneric taxa down to sectional level (and below in some cases) are well-supported by the sequence data and definable by morphological traits. Inclusion of <em>Angophora</em> within <em>Eucalyptus</em> was formally proposed two decades ago but did not gain acceptance. Here instead, we formally raise <em>Corymbia</em> subg. <em>Blakella</em> to genus rank and make the relevant new combinations.</p>

opencc-zeroFeb 2023View details →
dryad36/100

Raw sequence data and OTU tables of soil microorganisms obtained across a summit in the Lesotho highlands

<p>Mountain regions represent unique environments characterized by strong topographical diversity which drive climatic and environmental variability within these environments. These regions thus provide an opportunity to explore the relationships between various environmental factors and soil microorganisms. In this study, we investigated the impact of micro-topographical (i.e., north/south-facing slope aspects and flat plateau between them) variations on microbial diversity and community structures across a Lesotho mountain summit.</p> <p>Raw sequenced data were generated using the Illumina MiSeq platform on DNA extracted from soil samples collected across the plateau, north- and south-facing slopes. This data was then used for taxonomic classification of the bacterial and fungal OTUs for the determination of the alpha- and beta-diversity across the slopes. These analyses revealed that a relatively greater bacterial and fungal diversity could be observed for the north-facing slope compared to the south-facing slope and plateau. While there was no difference in group variance of bacterial and fungal community structures across the plateau, north- and south-facing slopes.</p> <p>Multiple comparison analyses were conducted to determine the impact of various abiotic and geographical factors on bacterial and fungal diversity and community structures. These analyses indicated that the slope aspect significantly affects bacterial and fungal community structures at this location. These results provide an original insight into soil microbial diversity in the Lesotho highlands and offer an opportunity to investigate the response of soil microorganisms to changes in environmental and climatic factors in highly variable mountain environments such as the Lesotho highlands.</p>

opencc-zeroFeb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record