Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Data from: Oral samples as non-invasive proxies for assessing the composition of the rumen microbial community
Microbial community analysis was carried out on ruminal digesta obtained directly via rumen fistula and buccal fluid, regurgitated digesta (bolus) and faeces of dairy cattle to assess if non-invasive samples could be used as proxies for ruminal digesta. Samples were collected from five cows receiving grass silage based diets containing no additional lipid or four different lipid supplements in a 5 x 5 Latin square design. Extracted DNA was analysed by qPCR and by sequencing 16S and 18S rRNA genes or the fungal ITS1 amplicons. Faeces contained few protozoa, and bacterial, fungal and archaeal communities were substantially different to ruminal digesta. Buccal and bolus samples gave much more similar profiles to ruminal digesta, although fewer archaea were detected in buccal and bolus samples. Bolus samples overall were most similar to ruminal samples. The differences between both buccal and bolus samples and ruminal digesta were consistent across all treatments. It can be concluded that either proxy sample type could be used as a predictor of the rumen microbial community, thereby enabling more convenient large-scale animal sampling for phenotyping and possible use in future animal breeding programs aimed at selecting cattle with a lower environmental footprint.
Data from: Revisiting metazoan phylogeny with genomic sampling of all phyla
Proper biological interpretation of a phylogeny can sometimes hinge on the placement of key taxa – or fail when such key taxa are not sampled. In this light, we here present the first attempt to investigate (though not conclusively resolve) animal relationships using genome-scale data from all phyla. Results from the site-heterogeneous CAT+GTR model recapitulate many established major clades, and strongly confirm some recent discoveries, such as a monophyletic Lophophorata, and a sister group relationship between Gnathifera and Chaetognatha, raising continued questions on the nature of the spiralian ancestor. We also explore matrix construction with an eye towards testing specific relationships; this approach uniquely recovers support for Panarthropoda, and shows that Lophotrochozoa (a subclade of Spiralia) can be constructed in strongly conflicting ways using different taxon- and/or orthologue sets. Dayhoff-6 recoding sacrifices information, but can also reveal surprising outcomes, e.g., full support for a clade of Lophophorata and Entoprocta+Cycliophora, a clade of Placozoa+Cnidaria, and raising support for Ctenophora as sister group to the remaining Metazoa, in a manner dependent on the gene and/or taxon sampling of the matrix in question. Future work should test the hypothesis that the few remaining uncertainties in animal phylogeny might reflect violations of the various stationarity assumptions used in contemporary inference methods.
Data from: Long-tailed macaques (Macaca fascicularis) can use simple heuristics but fail at drawing statistical inferences from populations to samples
Human infants, apes, and capuchin monkeys engage in intuitive statistics: they generate predictions from populations of objects to samples based on proportional information. This suggests that statistical reasoning might depend on some core knowledge that humans share with other primate species. To aid the reconstruction of the evolution of this capacity, we investigated whether intuitive statistical reasoning is also present in a species of Old World monkey. In a series of 4 experiments, 11 long-tailed macaques were offered different pairs of populations containing varying proportions of preferred vs. neutral food items. One population always contained a higher proportion of preferred items than the other. An experimenter simultaneously drew one item out of each population, hid them in her fists and presented them to the monkeys to choose. Although some individuals performed well across most experiments, our results imply that long-tailed macaques as a group did not make statistical inferences from populations of food items to samples but rather relied on heuristics. These findings suggest that there may have been convergent evolution of this ability in New World monkeys and apes (including humans).
Data from: A comparison of single-sample estimators of effective population sizes from genetic marker data
In molecular ecology and conservation genetics studies, the important parameter of effective population size (Ne) is increasingly estimated from a single sample of individuals taken at random from a population and genotyped at a number of marker loci. Several estimators are developed, based on the information of linkage disequilibrium (LD), heterozygote excess (HE), molecular coancestry (MC) and sibship frequency (SF) in marker data. The most popular is the LD estimator, because it is more accurate than HE and MC estimators and is simpler to calculate than SF estimator. However, little is known about the accuracy of LD estimator relative to that of SF and about the robustness of all single-sample estimators when some simplifying assumptions (e.g. random mating, no linkage, no genotyping errors) are violated. This study fills the gaps and uses extensive simulations to compare the biases and accuracies of the four estimators for different population properties (e.g. bottlenecks, nonrandom mating, haplodiploid), marker properties (e.g. linkage, polymorphisms) and sample properties (e.g. numbers of individuals and markers) and to compare the robustness of the four estimators when marker data are imperfect (with allelic dropouts). Extensive simulations show that SF estimator is more accurate, has a much wider application scope (e.g. suitable to nonrandom mating such as selfing, haplodiploid species, dominant markers) and is more robust (e.g. to the presence of linkage and genotyping errors of markers) than the other estimators. An empirical data set from a Yellowstone grizzly bear population was analysed to demonstrate the use of the SF estimator in practice.
Data from: Nitrogen and chlorine co-doped carbon dots as probe for sensing and imaging in biological samples
A facile one step hydrothermal synthesis approach was proposed to prepare nitrogen and chlorine co-doped carbon dots using l-ornithine hydrochloride as the sole precursor. The configuration and component of carbon dots were characterized by TEM, XPS, and FTIR. The obtained CDs (Orn-CDs) with a mean diameter of 2.1 nm were well monodispersed in aqueous solutions. The as-prepared CDs exhibited a bright blue fluorescence with a high yield of 60%, good photostability and low cytotoxicity. The emission of Orn-CDs could be selectively and effectively suppressed by Fe3+. Thus, a quantitative assay of Fe3+ was realized by this nanoprobe with a detection limit of 95.6 nmol L-1 in the range of 0.3-50 µmol L-1. Furthermore, ascorbic acid could recover the fluorescence of Orn-CDs suppressed by Fe3+, owing to the transformation of Fe3+ to Fe2+ by ascorbic acid. The limit of detection for ascorbic acid was 137 nmol L-1 in the range of 0.5-10 µmol L-1. In addition, the established method was successfully applied for Fe3+ and ascorbic acid sensing in human serum and urine specimans and for imaging of Fe3+ in living cells. With merits of low economic cost, easy to scale up, without additional functionalized and sample pretreatment, Orn-CDs based sensing platform showed its potential to be used for biomedical related study.
Data from: Interpreting ELISA analyses from wild animal samples: some recurrent issues and solutions
1. Many studies in disease and immunological ecology rely on the use of assays that quantify the amount of specific antibodies (immunoglobulin) in samples. Enzyme-Linked Immuno Sorbent Assays (ELISAs) are increasingly used in ecology due to their availability for a broad array of antigens and the limited amount of sampling material they require. Two recurrent methodological issues are nevertheless faced by researchers: (i) the limited availability of immunological assays and reagents developed for non-model species, and (ii) the statistical determination of the cut-off threshold used to distinguish individual samples that are likely to have or not to have antibodies against a specific antigen. 2. Here, we outline two solutions to deal with these issues. First, we show that implementing two assays with differing detection methods can help validate the use of reagents, such as antibodies, in species different from their intended target. We illustrate this by comparing the quantification of specific vaccinal antibodies against Newcastle Disease Virus (NDV) using two ELISA approaches in four seabird species (Cory's shearwater, European shag, European storm petrel, and Southern rockhopper penguin). 3. Second, we provide a simple way to determine from the distribution of ELISA values whether the assayed samples are likely to be made of a single group of individuals (likely negative) or of two groups of individuals (negative and positive). We illustrate the use of this approach with two independent datasets: NDV antibody levels following vaccination and anti-Borrelia antibody levels following natural exposure. 4. The practical implementation of these methodological approaches could provide a way to efficiently apply ELISAs and other immune-based assays to address questions in the growing fields of ecological immunology and disease ecology.
Data from: A lateral flow immunochromatographic strip test for rapid detection of hexoestrol in fish samples
A lateral flow immunochromatographic test strip was developed for on-site rapid and sensitive detection of Hexoestrol (HES) residues in fish samples with colloidal gold labeled the anti-HES monoclonal antibody (mAb). The strip is composed of a sample pad, a conjugate reagent pad, an absorbent pad, and a test membrane containing a control line and a test line. The sensitivity (half inhibitory concentration, IC50) of the strip in the detection of fish extract samples was confirmed to be 1.86 μg/kg, and the limit detection (LOD) value was 0.62 μg/kg. For intra-assay and inter-assay reproducibility, recoveries of HES spiked samples were ranged from 86.3% to 92.3% and 85.8% to 93.4%, coefficients of variation were 2.91-4.64% and 4.24-5.17% respectively. High-performance liquid chromatography (HPLC) was employed to confirm the performance of the strip. The strip test only took less than 10 minutes, and thus provides a repaid method for on-site detection of HES residues.
Data from: Sex, size and timing: sampling design for reliable population genetics analyses using microsatellite data
1. Population genetics is used in a wide variety of fields such as ecology and biodiversity conservation. How estimated genetic characteristics of natural populations can be influenced by the sampling design has been a long-standing concern. Multiple simulation and empirical studies illustrated the influence of both sample size and polymorphism of markers. However, our review of studies on butterfly population genetics indicates no consensus on sample size for the estimation of genetic diversity or differentiation. Furthermore, other aspects of sampling design (sex ratio and timing of sampling) were not addressed and their potential impact on genetic parameter estimates rarely explored. 2. Using a large empirical dataset (with spatial and temporal replicates) collected on a butterfly species, Boloria aquilonaris, as well as simulated datasets reflecting (1) three scenarios of migration-genetic drift equilibrium and (2) one scenario of parameter stabilization after 100,000 generations, we quantified the impacts of three aspects of genetic sampling design (namely sample size, sex ratio, and timing of sampling) on the estimation of allele frequencies and its potential downstream impact on the estimation of genetic parameters. 3. With empirical data, we found that sample size and timing of sampling strongly affected the accuracy of allele frequencies and the downstream analyses, while sex ratio did not. Our results were consistent across spatial and temporal replicates. Also, with simulated data, we showed that the genetic sampling design had limited effect in systems where dispersal outweighs genetic drift, while it can have major consequences on our understanding of the genetic diversity and population differentiation in systems dominated by genetic drift (such as most study systems with conservation concerns). 4. We advocate for careful consideration of all aspects of the sampling design in population genetics studies, i.e. a sufficient number of samples, while ensuring similar sex ratio among sampling locations and collecting with timing appropriate to the question under study. This is particularly important when the study aims at species conservation.
Data from: Spatiotemporal sampling patterns in the 230 million year fossil record of terrestrial crocodylomorphs and their impact on diversity
The 24 extant crocodylian species are the remnants of a once much more diverse and widespread clade. Crocodylomorpha has an approximately 230 million year evolutionary history, punctuated by a series of radiations and extinctions. However, the group's fossil record is biased. Previous studies have reconstructed temporal patterns in subsampled crocodylomorph palaeobiodiversity, but have not explicitly examined variation in spatial sampling, nor the quality of this record. We compiled a dataset of all taxonomically diagnosable non‐marine crocodylomorph species (393). Based on the number of phylogenetic characters that can be scored for all published fossils of each species, we calculated a completeness value for each taxon. Mean average species completeness (56%) is largely consistent within subgroups and for different body size classes, suggesting no significant biases across the crocodylomorph tree. In general, average completeness values are highest in the Mesozoic, with an overall trend of decreasing completeness through time. Many extant taxa are identified in the fossil record from very incomplete remains, but this might be because their provenance closely matches the species' present‐day distribution, rather than through autapomorphies. Our understanding of nearly all crocodylomorph macroevolutionary 'events' is essentially driven by regional patterns, with no global sampling signal. Palaeotropical sampling is especially poor for most of the group's history. Spatiotemporal sampling bias impedes our understanding of several Mesozoic radiations, whereas molecular divergence times for Crocodylia are generally in close agreement with the fossil record. However, the latter might merely be fortuitous, i.e. divergences happened to occur during our ephemeral spatiotemporal sampling windows.
Data from: How many more? Sample size determination in studies of morphological integration and evolvability
The variational properties of living organisms are an important component of current evolutionary theory. As a consequence, researchers working on the field of multivariate evolution have increasingly used integration and evolvability statistics as a way of capturing the potentially complex patterns of trait association and their effects over evolutionary trajectories. Little attention has been paid, however, to the cascading effects that inaccurate estimates of trait covariance have on these widely used evolutionary statistics. Here, we analyze the relationship between sampling effort and inaccuracy in evolvability and integration statistics calculated from 10-trait matrices with varying patterns of covariation and magnitudes of integration. We then extrapolate our initial approach to different numbers of traits and different magnitudes of integration and estimate general equations relating the inaccuracy of the statistics of interest to sampling effort. We validate our equations using a dataset of cranial traits, and use them to make sample size recommendations. Our results suggest that highly inaccurate estimates of evolvability and integration statistics resulting from small sample sizes are likely common in the literature, given the sampling effort necessary to properly estimate them. We also show that patterns of covariation have no effect on the sampling properties of these statistics, but overall magnitudes of integration interact with sample size and lead to varying degrees of bias, imprecision, and inaccuracy. Finally, we provide R functions that can be used to calculate recommended sample sizes or to simply estimate the level of inaccuracy that should be expected in these statistics, given a sampling design.
Operando electronic conductivity data - Handbook protocol results (other samples)
<p>Operando electronic conductivity data obtained from the microwave cavity perturbation technique using the Handbook protocol. Supplement to DOI: <a href="https://doi.org/10.5281/zenodo.5008960">10.5281/zenodo.5008960</a></p> <p>Results for other samples studied with the Handbook protocol.</p> <p>Contents:</p> <ul> <li>instrument_data.zip contains raw instrument (MCPT and GC) data for the whole dataset</li> <li>individual zip files contain the <em>schema</em>, <em>datagram</em>, <em>run protocol</em>,<em> </em>and <em>parameter file</em> for each sample and reproduction, as well as an plot.png file and results.json file from dg2png / dg2json analysis</li> </ul> <p>Note that calibration files for the instrument are included at <a href="https://doi.org/10.5281/zenodo.5894835">DOI: 10.5281/zenodo.5894835</a>.</p> <p>Sample to archive matrix:</p> <ul> <li>19760 silica gel: 19760-01.zip</li> <li>30649 LaMnO<sub>3</sub>: 30649-01.zip</li> <li>30650 PrMnO<sub>3</sub>: 30650-01.zip</li> <li>30869 Sm<sub>0.95</sub>MnO<sub>3</sub>: 30869-01.zip</li> <li>31034 V<sub>2</sub>O<sub>5</sub>: 31034-01.zip</li> <li>31652 MoVTeNbOx: 31652-01.zip</li> </ul>
Data from: The influence sampling design on species tree inference: a new relationship for the New World chickadees (Aves: Poecile)
In this study, we explore the long-standing issue of how many loci are needed to infer accurate phylogenetic relationships, and whether loci with particular attributes (i.e., parsimony informativeness, variability, gene tree resolution) outperform others. To do so, we use an empirical dataset consisting of the seven species of chickadees (Aves: Paridae), an analytically tractable, recently diverged group, and well studied ecologically but lacking a nuclear phylogeny. We estimate relationships using 40 nuclear loci and mitochondrial DNA using four coalescent-based species tree inference methods (BEST, *BEAST, STEM, STELLS). Collectively, our analyses contrast with previous studies and support a sister relationship between the Black-capped and Carolina Chickadee, two superficially similar species that hybridize along a long zone of contact. Gene flow is a potential source of conflict between nuclear and mitochondrial gene trees, yet, we find a significant, albeit low, signal of gene flow. Our results suggest that relatively few loci with high information content may be sufficient for estimating an accurate species tree, but that substantially more loci are necessary for accurate parameter estimation. We provide an empirical reference point for researchers designing sampling protocols with the purpose of inferring phylogenies and population parameters of closely related taxa.
Data from: A non-lethal sampling method to obtain, generate and assemble whole-blood transcriptomes from small, wild mammals
The acquisition of tissue samples from wild populations is a constant challenge in conservation biology, especially for endangered species and protected species where nonlethal sampling is the only option. Whole blood has been suggested as a nonlethal sample type that contains a high percentage of bodywide and genomewide transcripts and therefore can be used to assess the transcriptional status of an individual, and to infer a high percentage of the genome. However, only limited quantities of blood can be nonlethally sampled from small species and it is not known if enough genetic material is contained in only a few drops of blood, which represents the upper limit of sample collection for some small species. In this study, we developed a nonlethal sampling method, the laboratory protocols and a bioinformatic pipeline to sequence and assemble the whole blood transcriptome, using Illumina RNA-Seq, from wild greater mouse-eared bats (Myotis myotis). For optimal results, both ribosomal and globin RNAs must be removed before library construction. Treatment of DNase is recommended but not required enabling the use of smaller amounts of starting RNA. A large proportion of protein-coding genes (61%) in the genome were expressed in the blood transcriptome, comparable to brain (65%), kidney (63%) and liver (58%) transcriptomes, and up to 99% of the mitogenome (excluding D-loop) was recovered in the RNA-Seq data. In conclusion, this nonlethal blood sampling method provides an opportunity for a genomewide transcriptomic study of small, endangered or critically protected species, without sacrificing any individuals.
Data from: Sampling diverse characters improves phylogenies: craniodental and postcranial characters of vertebrates often imply different trees
Morphological cladograms of vertebrates are often inferred from greater numbers of characters describing the skull and teeth than from postcranial characters. This is either because the skull is believed to yield characters with a stronger phylogenetic signal (i.e., contain less homoplasy), because morphological variation therein is more readily atomized, or because craniodental material is more widely available (particularly in the palaeontological case). An analysis of 85 vertebrate datasets published between 2000 and 2013 confirms that craniodental characters are significantly more numerous than postcranial characters, but finds no evidence that levels of homoplasy differ in the two partitions. However, a new partition test based on tree-to-tree distances (as measured by Robinson Foulds metric) rather than tree length reveals that relationships inferred from the partitions are significantly different about one time in three, much more often than expected. Such differences may reflect divergent selective pressures in different body regions, resulting in different localized patterns of homoplasy. Most systematists attempt to sample characters broadly across body regions, but this is not always possible. We conclude that trees inferred largely from either craniodental or postcranial characters in isolation may differ significantly from those that would result from a more holistic approach. We urge the latter.
Data from: Biofilm morphotypes and population structure among Staphylococcus epidermidis from commensal and clinical samples
Bacterial species comprise related genotypes that can display divergent phenotypes with important clinical implications. Staphylococcus epidermidis is a common cause of nosocomial infections and, critical to its pathogenesis, is its ability to adhere and form biofilms on surfaces, thereby moderating the effect of the host's immune response and antibiotics. Commensal S. epidermidis populations are thought to differ from those associated with disease in factors involved in adhesion and biofilm accumulation. We quantified the differences in biofilm formation in 98 S. epidermidis isolates from various sources, and investigated population structure based on ribosomal multilocus typing (rMLST) and the presence/absence of genes involved in adhesion and biofilm formation. All isolates were able to adhere and form biofilms in in vitro growth assays and confocal microscopy allowed classification into 5 biofilm morphotypes based on their thickness, biovolume and roughness. Phylogenetic reconstruction grouped isolates into three separate clades, with the isolates in the main disease associated clade displaying diversity in morphotype. Of the biofilm morphology characteristics, only biofilm thickness had a significant association with clade distribution. The distribution of some known adhesion-associated genes (aap and sesE) among isolates showed a significant association with the species clonal frame, with the exception of. These data challenge the assumption that biofilm-associated genes, such as those on the ica operon, are genetic markers for less invasive S. epidermidis isolates, and suggest that phenotypic characteristics, such as adhesion and biofilm formation, are not fixed by clonal descent but are influenced by the presence of various genes that are mobile among lineages.
Data from: Influence of preexisting preference for color on sampling and tracking behavior in bumble bees
Animals reduce uncertainty in their lifetime by using information to guide decision making. Information available can be inherited from the past or gathered from the present. Therefore, animals must balance inherited biases with new information that may be in conflict with those potential biases. In our study, we set up color pairings such that an arbitrarily chosen focal color, human-orange, would result in an inherent bias in comparison to three other colors tested resulting in equal, medium, and strong preference differences. We chose color pairings through a series of preferences tests across 8 colonies of bumblebees. We subsequently used these pairings with rewards that varied in quality (good or bad states) and consistency (steady and fluctuating) in order to investigate how inherited biases affect the foraging choices of bumblebees when new information is gathered. We found that the pre-existing color biases within our bees were only maintained when the reward associated with those colors was steady, even if paired with mediocre sugar concentrations. When maintained, we observed that other aspects of bee choice also reflected this bias, including increased sampling for the preferred color and an increased likelihood of choosing that color in a subsequent choice. Thus, environmental change and reward differences interact with the level of pre-existing bias to determine whether inherited information is more heavily weighted than newly gathered information, and even a strong pre-existing bias can be quickly erased with experience under some conditions.
Data from: State-space reduction and equivalence class sampling for a molecular self-assembly model
Direct simulation of a model with a large state space will generate enormous volumes of data, much of which is not relevant to the questions under study. In this paper, we consider a molecular self-assembly model as a typical example of a large state-space model, and present a method for selectively retrieving 'target information' from this model. This method partitions the state space into equivalence classes, as identified by an appropriate equivalence relation. The set of equivalence classes H, which serves as a reduced state space, contains none of the superfluous information of the original model. After construction and characterization of a Markov chain with state space H, the target information is efficiently retrieved via Markov chain Monte Carlo sampling. This approach represents a new breed of simulation techniques which are highly optimized for studying molecular self-assembly and, moreover, serves as a valuable guideline for analysis of other large state-space models.
Data from: Time for a rethink: time sub-sampling methods in disparity-through-time analyses
Disparity-through-time analyses can be used to determine how morphological diversity changes in response to mass extinctions, and to investigate the drivers of morphological change. These analyses are routinely applied to palaeobiological datasets, yet although there is much discussion about how to best calculate disparity, there has been little consideration of how taxa should be sub-sampled through time. Standard practice is to group taxa into discrete time bins, often based on stratigraphic periods. However, this can introduce biases when bins are of unequal size, and implicitly assumes a punctuated model of evolution. In addition, many time bins may have few or no taxa, meaning that disparity cannot be calculated for the bin and making it harder to complete downstream analyses. Here we describe a different method to complement the disparity-through-time tool-kit: time-slicing. This method uses a time-calibrated phylogenetic tree to sample disparity-through-time at any fixed point in time rather than binning taxa. It uses all available data (tips, nodes and branches) to increase the power of the analyses, specifies the implied model of evolution (punctuated or gradual), and is implemented in R. We test the time-slicing method on four example datasets and compare its performance in common disparity-through-time analyses. We find that the way you time sub-sample your taxa can change your interpretations of the results of disparity-through-time analyses. We advise using multiple methods for time sub-sampling taxa, rather than just time binning, to gain a better understanding disparity-through-time.
Denoising Autoencoders for Phenotype Stratification (DAPS) Sample Trained Simulated Patient Data
<p>DAPS Trained data</p>
Sample data
<p>Sample data</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.