Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"
<p>Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"</p>
Petrophysical data for 29 samples from the Chicxulub impact crater.
<p>Note: ɸ-porosity, ρ<sub>b</sub>-bulk density, ρ<sub>g</sub>-grain density, k-permeability, F-formation factor, m-cementation exponent, τ<sup>2</sup>-tortuosity, C<sub>s</sub>-surface conductivity, Vp-acoustic velocity of compressional waves. Uncertainty for porosity, density, permeability, velocity and conductivity is 5%. Uncertainty for formation factor, cementation exponent and tortuosity is 8%). Lith <sup>1 </sup>and Unit <sup>1</sup> after Morgan et al. (2017), Unit <sup>2</sup> after de Graaf et al. (2021, UIM-upper impact melt rock unit, LIMB-lower impact melt rock-bearing unit)) and Kaskes et al. (2021).</p> <p> </p> <p>Morgan, J. V., Gulick, S. P. S., Bralower, T. J., Chenot, E., Christeson, G. L., Claeys, P., et al. (2016). The formation of peak rings in large impact craters. Science, 354(6314), 878–882. <a href="https://doi.org/10.1126/science.aah6561">https://doi.org/10.1126/science.aah6561</a></p> <p>de Graaff, S. J., Kaskes, P., Déhais, T., Goderis, S., Vinciane, D., Ross, C. H., et al. (2021). New insights into the formation and emplacement of impact melt rocks within the Chicxulub impact structure, following the 2016 IODP-ICDP Expedition 364. Geological Society of America Bulletin. <a href="https://doi.org/doi:">https://doi.org/doi:</a> <a href="https://doi.org/10.1130/B35795.1">https://doi.org/10.1130/B35795.1</a></p> <p>Kaskes, P., de Graaff, S. J., Feignon, J. G., Déhais, T., Goderis, S., Ferrière, L., et al. (2021). Formation of the crater suevite sequence from the Chicxulub peak ring: A petrographic, geochemical, and sedimentological characterization. Geological Society of America Bulletin. <a href="https://doi.org/https://doi.org/10.1130/B36020.1">https://doi.org/https://doi.org/10.1130/B36020.1</a></p>
Data from: When less is more and more is less: the impact of sampling effort on species delineation
Taxonomy is the very first step of most biodiversity studies, but how confident can we be in the taxonomic-systematic exercise? One may hypothesise that the more material, the better the taxonomic delineation, because the more accurate the description of morphological variability. As rarefaction curves assess the degree of knowledge on taxonomic diversity through sampling effort, we aim to test the impact of sampling effort on species delineation by subsampling a given assemblage. To do so, we use an abundant and morphologically diverse conodont fossil record. Within the assemblage, we first recognize four well established morphospecies but about 80% of the specimens share diagnostic characters of these morphospecies. We quantify these diagnostic characters on the sample using geometric morphometrics, and assess the number of morphometric groups, i.e. morphospecies, using ordination and cluster analyses. Then we gradually subsample the assemblage in two ways (randomly and by mimicking taxonomist work) and redo the 'ordination + clustering' protocol to appraise the evolution of the number of clusters related to sampling effort. We observe the number of delineated morphospecies decreasing when increasing the number of specimens, whatever the subsampling method, resulting mostly in less morphospecies than expected. Such rather counter-intuitive influence of sampling effort on species delineation highlights the complexity of taxonomical work. This indicates that new morphotaxa should not be erected based on small samples, and encourages researchers to largely illustrate, measure, and quantitatively compare their material to better constrain the morphological variability of a clade, and so to better characterize and delineate morphospecies. --
Data from: Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range - wild grapevine sampling locations, Maxent input files, morphological and microsatellite data
<p><span>This dataset contains raw data described in the paper: "Rahimi O., Ohana-Levi N., Brauner H., Inbar N., Hübner S. and Drori E. (2021) "Demographic and ecogeographic factors limit wild grapevine spread at the southern edge of its distribution range", accepted for publication in "Ecology and Evolution".</span></p> <p><span>The spatial distribution of plants is constrained by demographic and eco-geographic factors that determine the range and abundance of the species. In this study, we performed genetic and morphological analyzes based on SSR and OIV datasets. In addition, according to the spatial distribution model performed by Maxent software we found that distance to water sources, Normalized difference vegetation index, and precipitation are the main environmental factors constraining <i>V.v. sylvestris</i> distribution at its southern distribution range. All raw data used for this study can be found in this deposit which contains a table with grapevine locations, Maxent input files, morphological and microsatellite data. </span></p>
DCSsim (simulated) and DCSsub (sub-sampled) ChIP-seq data with different FRIP.
<p>These data are the results from three independent runs of DCSsim and DCSsub for TF, sharp and broad mark signals in 50:50 regulation scenarios for four (sim) and three (sub) different FRIP ranges.</p> <p>Simulated data from DCSsim: simulated_ChIP-seq_data.zip</p> <p>Set13: TF 50:50 x0.5 background<br> Set14: TF 50:50 x2 background<br> Set15: TF 50:50 x3 background<br> Set25: TF 50:50 x1 background</p> <p>Set16: Sharp mark 50:50 x0.5 background<br> Set17: Sharp mark 50:50 x2 background<br> Set18: Sharp mark 50:50 x3 background<br> Set26: Sharp mark 50:50 x1 background</p> <p>Set19: Broad mark 50:50 x0.5 background<br> Set20: Broad mark 50:50 x2 background<br> Set21: Broad mark 50:50 x3 background<br> Set27: Broad mark 50:50 x1 background</p> <p><br> Sub-sampled data from DCSsub: sub-sampled_ChIP-seq_data.zip</p> <p>Set1: PU1-ChIP-seq 50:50<br> Set2: STAT6-ChIP-seq 50:50<br> Set8: C/EBPa-ChIP-seq 50:50</p> <p>Set4: H3K4me3-ChIP-seq 50:50<br> Set9: H3K27ac-ChIP-seq 50:50<br> Set15: H3K9ac-ChIP-seq 50:50</p> <p>Set6: H3K27me3-ChIP-seq 50:50<br> Set10: H3K36me3-ChIP-seq 50:50<br> Set11: H3K79me2-ChIP-seq 50:50</p>
Data from: A genotyping-in-thousands by sequencing panel to inform invasive deer management using non-invasive fecal and hair samples
<p>Studies in ecology, evolution, and conservation often rely on non-invasive samples, making it challenging to generate large amounts of high-quality genetic data for many elusive and at-risk species. We developed and optimized a Genotyping-in-Thousands by sequencing (GT-seq) panel using non-invasive samples to inform the management of invasive Sitka black-tailed deer (<em>Odocoileus hemionus sitkensis</em>) in Haida Gwaii (Canada). We validated our panel using paired high-quality tissue and non-invasive fecal and hair samples to simultaneously distinguish individuals, identify sex and reconstruct kinship among deer sampled across the archipelago, then provided a proof-of-concept application using field-collected feces on SGang Gwaay, an island of high ecological and cultural value. Genotyping success across 244 loci was high (90.3%) and comparable to that of high-quality tissue samples genotyped using restriction-site associated DNA sequencing (92.4%), while genotyping discordance between paired high-quality tissue and non-invasive samples was low (0.50%). The panel will be used to inform future invasive species operations (culls or eradications) in Haida Gwaii by providing individual and population information to inform management. More broadly, our GT-seq workflow that includes quality control analyses for targeted SNP selection and a modified protocol may be of wider utility for other studies and systems where non-invasive genetic sampling is employed.</p>
Sample Input Data and Supporting Files for the SELECT Model of Urbanization
<p>Sample Input Data and Supporting Files for the SELECT Model of Urbanization</p> <p>Code available at: https://github.com/IMMM-SFA/select</p>
Data from: Between the Cape Fold Mountains and the deep blue sea: comparative phylogeography of selected codistributed ectotherms reveals asynchronous cladogenesis. Sampling locations and MaxEnt input files
<p><span>We compare the phylogeographic structure of thirteen codistributed ectotherms including four reptiles (a snake, a legless skink and two tortoise species) and nine invertebrates (six freshwater crabs and three velvet worm species) to test the presence of congruent evolutionary histories. </span><span>Phylogenies were estimated and dated using maximum likelihood and Bayesian methods with combined mitochondrial and nuclear DNA sequence datasets. </span><span>All taxa demonstrated a marked east/west phylogeographic division, separated by the Cape Fold Mountain range.</span><span> <span>Phylogeographic concordance factors were calculated to assess the degree of evolutionary congruence among the study species and </span></span><span>generally supported a shared pattern of diversification along the east/west longitudinal axis</span><span>. Testing simultaneous divergence between the eastern and western phylogeographic regions indicated </span><span>pseudo-congruent evolutionary histories among the study taxa, with at least three separate divergence events throughout the Mio/Plio/Pleistocene epochs.</span><span> <span>Climatic refugia were identified for each species using climatic niche modeling, </span></span><span>demonstrating taxon-specific responses to climatic fluctuations. Climate and the Cape Fold Mountain barrier explained the highest proportion of genetic diversity in all taxa, while climate was the most significant individual abiotic variable. </span><span>This study highlights the complex interactions between the Cape Fold Mountains and past climatic oscillations during the Mio/Plio/Pleistocene. The congruent east/west phylogeographic division observed in all taxa lends support to the conclusion that the longitudinal climatic gradient within the Greater Cape Floristic Region, mediated in part by the barrier to dispersal posed by the Cape Fold Mountains, plays a major role in lineage diversification and population differentiation.</span></p>
Orignal data for publication: Probing traps in the persistent phosphor SrAl2O4:Eu2+,Dy3+,B3+ - A wavelength, temperature and sample dependent thermoluminescence investigation
<p>Original Data for publication (part of University of Geneva):</p> <p>Probing traps in the persistent phosphor SrAl2O4:Eu2+,Dy3+,B3+ - A wavelength, temperature and sample dependent thermoluminescence investigation<br> Jakob Bierwagen, Teresa Delgado, Guillaume Jiranek, Songhak Yoon, Nando Gartmann, Bernhard Walfort, Markus Pollnau , Hans Hagemann,<br> J. Lumin. 222 (2020) 117113.</p>
Data for: Development of a method for the measurement of human scent samples using comprehensive two-dimensional gas chromatography with mass detection
<p>This dataset was used for development of a method for the measurement of human scent samples using comprehensive two-dimensional gas chromatography with mass detection [<a href="https://doi.org/10.1016/j.forsciint.2016.09.011">https://doi.org/10.1016/j.forsciint.2016.09.011</a>].</p> <p>The dataset contains chromatograms of a model mixture of human scent and chromatograms of the human scent of one volunteer measured on different column setups. Each sample was processed in ChromaToF(version 4.72.0.0) by LECO corp. The processing step was executed at the signal-to-noise (SN) ratio levels 100, 300, and 500 (human scent samples chromatograms).</p>
Additional Figures for winning models for sample in A Comparative L-dwarf Sample Exploring the Interplay Between Atmospheric Assumptions and Data Properties
<p>Additional Figures for winning models for sample in <em>A Comparative L-dwarf Sample Exploring the Interplay Between Atmospheric Assumptions and Data Properties (<a href="https://arxiv.org/pdf/2209.02754.pdf">https://arxiv.org/pdf/2209.02754.pdf</a>).</em></p> <p>Model naming key: NC = cloud-free, d2_89 = power-law deck cloud</p> <p>SDSS J1416+1348A: Winning model: power-law deck cloud</p> <p>Spectral Type Comparison J1526+2043 Winning model: Cloud-free</p> <p>Temperature Comparisons</p> <p>J1539-0520 Winning model: Power-law deck cloud and cloud-free tied.</p> <p>J0539-0059 Winning model: Power-law deck cloud and cloud-free tied. </p> <p><br> </p>
Quantitative 3D OPT and LSFM datasets of pancreata from mice with streptozotocin-induced diabetes: Sample data sets
<p><span>Mouse models for streptozotocin (STZ) induced diabetes probably represent the most widely used systems for preclinical diabetes research, owing to the compound's toxic effect on pancreatic ß-cells. However, a comprehensive view of pancreatic β-cell mass distribution subject to STZ administration is lacking. Previous assessments have largely relied on the extrapolation of stereological sections, which provide limited 3D-spatial and quantitative information. This data descriptor presents multiple ex vivo tomographic optical image data sets of the full β-cell mass distribution in mice subject to single high and multiple low doses of STZ administration, and in glycaemia recovered mice. The data further include information about structural features, such as individual islet β-cell volumes, spatial coordinates, and shape as well as signal intensities for both insulin and GLUT2. Together, they provide the most comprehensive anatomical record of the effects of STZ administration on the islet of Langerhans in mice. As such, this data descriptor may serve as reference material to facilitate the planning, use and (re)interpretation of this widely used disease model.</span></p>
Data samples for Flow-matching -- efficient coarse-graining molecular dynamics without forces
<p>CG samples generated during the training and validation processes in the flow-matching project. Accompanying the preprint "Flow-matching -- efficient coarse-graining molecular dynamics without forces": https://arxiv.org/abs/2203.11167. Detailed descriptions can be found in the preprint as well as the included README.</p>
Sample data for tephritid metabarcoding workshop
<p>Raw data files and illumina sample sheets for tephritid metabarcoding workshop hosted at Agribio in September 2022</p>
Data from: Resolution, conflict and rate shifts: Insights from a densely sampled plastome phylogeny for Rhododendron (Ericaceae)
<p><strong>Background and Aims</strong> <em>Rhododendron </em>is a species-rich and taxonomically challenging genus due to recent adaptive radiation and frequent hybridization. A well-resolved phylogenetic tree would help to understand the diverse history of <em>Rhododendron </em>in the Himalaya–Hengduan Mountains where the genus is most diverse.</p> <p><strong>Methods </strong>We reconstructed the phylogeny based on plastid genomes with broad taxon sampling, covering 161 species representing all eight subgenera and all 12 sections, including ~45 % of the <em>Rhododendron </em>species native to the Himalaya–Hengduan Mountains. We compared this phylogeny with nuclear phylogenies to elucidate reticulate evolutionary events and clarify relationships at all levels within the genus. We also estimated the timing and diversification history of <em>Rhododendron</em>, especially the two species-rich subgenera <em>Rhododendron</em> and <em>Hymenanthes </em>that comprise >90 % of <em>Rhododendron </em>species in the Himalaya–Hengduan Mountains.</p> <p><strong>Key Results </strong>The full plastid dataset produced a well-resolved and supported phylogeny of <em>Rhododendron</em>. We identified 13 clades that were almost always monophyletic across all published phylogenies. The conflicts between nuclear and plastid phylogenies strongly suggested that reticulation events may have occurred in the deep lineage history of the genus. Within <em>Rhododendron</em>, subgenus <em>Therorhodion </em>diverged first at 56 Mya, then a burst of diversification occurred from 23.8 to 17.6 Mya, generating ten lineages among the component 12 clades of core <em>Rhododendron</em>. Diversification in subgenus <em>Rhododendron </em>accelerated c. 16.6 Mya and then became fairly continuous. Conversely, <em>Hymenanthes </em>diversification was slow at first, then accelerated very rapidly around 5 Mya. In the Himalaya–Hengduan Mountains, subgenus <em>Rhododendron </em>contained one major clade adapted to high altitudes and another to low altitudes, whereas most clades in <em>Hymenanthes </em>contained both low- and high-altitude species, indicating greater ecological plasticity during its diversification.</p> <p><strong>Conclusions </strong>The 13 clades proposed here may help to identify specific ancient hybridization events. This study will help to establish a stable and reliable taxonomic framework for <em>Rhododendron</em>, and provides insight into what drove its diversification and ecological adaption. Denser sampling of taxa, examining both organelle and nuclear genomes, is needed to better understand the divergence and diversification history of <em>Rhododendron</em>.</p>
Genome-wide association results from Phase 1 data comparing uveitis-JIA cases to non-uveitis JIA samples
<p>Summary-level GWAS results for Phase 1 data affiliated with the manuscript "An amino acid motif in HLA-DRβ1 distinguishes patients with uveitis in juvenile idiopathic arthritis."</p> <p>Columns are:</p> <p> 1. CHR: chromosome</p> <p> 2. SNP: SNP identifier</p> <p> 3. BP: basepair position (hg19)</p> <p> 4. A1: minor allele and tested allele</p> <p> 5. A2: the other allele (major allele)</p> <p> 6. FRQ: frequency of the A1 allele</p> <p> 7. INFO: imputation info score</p> <p> 8. EFFECT: beta/effect size of the SNP</p> <p> 9. SE: standard error of the SNP</p> <p> 10. P: p-value at that SNP</p> <p> </p>
Multimodal data set for the investigation of the early stage of plasticity in a polycrystalline titanium sample
<p>This dataset is the result of several experiments to study the early stage of plasticity in a polycrystalline sample of a commercially pure alpha phase grade 2 titanium (CP-Ti family). The study aimed to achieve three critical goals: first, the acquisition of a 3D representation of the microstructure; second, the conduction of in situ measurements capturing grain-scale plasticity dynamics during a controlled tensile test; and third, a rigorous comparison of these experimental observations against the predictions derived from a microstructure-sensitive crystal plasticity simulation. This simulation was conducted on a digital twin of the titanium sample, aiming to assess the predictive accuracy of the model at the local scale. The data set was assembled from the different sources using the Pymicro package. </p>
Data belonging to "Successful invasion: camera trap distance sampling reveals higher density for invasive raccoon dog compared to native mesopredators"
<p>Data files (comma separated text files) containing the camera data (CameraData) containing the information on camera trap placements in the various sites and their operation time in days and aperture, the distance sampling data (DistanceData) containing the information on the species and distance detected for each 1s time interval in front of each camera, and the trigger data (TriggerData) containing the time stamps for the pictures taken of each species with each camera, collected in the years 2020 and 2021 in southern Finland. The repository further contains an R script "distanceSamplingScript" which uses the reposited above-described files for analysis reported in the publication "Successful invasion: camera trap distance sampling reveals higher density for invasive raccoon dog compared to native mesopredators" https://doi.org/10.1007/s10530-024-03323-4. The R script has been confirmed to run in R version 4.3.3 using packages "activity" vs 1.3.4 and "Distance" vs 1.0.9</p>
EBSD data, Entia Dome Amphibolite samples (2017 sample collection, unoriented)
<p><strong>Description:</strong> This dataset contains EBSD data (.ctf file format) collected on amphibolite (± clinopyroxene, ± garnet-bearing) samples from the Entia dome, central Australia.</p> <p><strong>Data collection Methods (in brief):</strong> EBSD data was collected in 2020 at the University of Minnesota Characterization Facility (multi-user facility), using a JEOL 6500 scanning electron microscope. EBSD data was collected on an Oxford Instruments Symmetry detector; data collection and processing was performed using the AZtec software. </p>
Costless correction of chain based nested sampling parameter estimaion in gravitational wave data and beyond (supplementary material)
<p>These are the nested sampling inference products used to obtain the results for <span><a href="https://arxiv.org/abs/2404.16428">arXiv:2404.16428</a>. The chains are given as pickled dataframes as there are over 1000 runs provided here. The script for reproducing the plots in the paper is also given. </span></p> <p><span>Folders:</span></p> <p><span>outdir_simulated_BBH - contains 200 runs on the same simulated signal. 'samples' contains the full sets of weighted samples from each run and 'full_dfs' contains the dataframes with the weighted samples AND the phantom points from the run. The samples in 'full_dfs' with chain_no=0 are the weighted samples, and all others are phantoms. </span></p> <p><span>outdir_test - contains the single run on the above simulated signal which was performed with num_repeats=25*ndims=100. </span></p> <p><span>outdir_coverage - contains 1000 runs on different simulated signals, with parameters drawn from the prior. The drawn values are saved as '{}_params.npy' for each run.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.