Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.7.1
Dataset results
86 results for “Mixture models”
Fluorescent Confocal Laser Scanning Microscopy of White Blood Cells, Cancer Cell Line MCF7, and Mixtures of these Cells: A Model System for Circulating Tumor Cell Biomarker Evaluation V.1
<p>This is a confocal laser scanning microscopy data set of white blood cells (leukocytes), the cancer cell line MCF7, and mixtures of these cells acquired on a Zeiss LSM 780 microscope in the University of Colorado Anschutz Medical Campus Advanced Light Microscopy Core. Cells are fluorescently labeled for DNA with DAPI (Sigma D9542), lipids with Bodipy 495/503 (Thermo Fisher D3922), the filament protein cytokeratin (CK) with pan-cytokertain-alexa555 antibodies (Cell Signaling Technologies 3478S) and the surface membrane antigen CD45 with CD45-alexa647 antibodies (Biolegend 304020). Bodipy was excited with a continuous wave (CW) 488 nm laser, alexa555 was excited with CW 561 nm laser, and alexa647 was excited with a CW 633 nm laser. The acquiring instrument does not have a CW 405 nm source so DAPI was excited by two photon process using a Coherent Cameleon ultrafast pulsed laser tuned to 765 nm. The objective used was a Zeiss Plan-Apochromat 20x, 0.8 NA, air.</p> <p>The data consists of 4 channel 8x8 mosaic z-stacks. The Zeiss software performed stitching of the mosaics. These stitched data images are included and marked with _Stitched at the end. Those interested in performing the stitching themselves can do this with the raw data files (without the _Stitched). The jpeg images are processed from the stitched LSM images. The LSM files contain additional meta data on the experiment including power levels and acquisition settings.</p> <p>The _Stiched .lsm files will load in ImageJ (tested with V.1.49) as 4 channel 3 stack images.</p> <p>This data is a model system for evaluating the DNA/Lipids/CK/CD45 biomarker panel to identify circulating tumor cells (CTCs). The D- population of the model is the WBCs and the D+ population is the MCF7 cancer cell line. The amount of separation the biomarker panel plus analysis algorithm can produce between these populations (D+/D-) is an estimate the sensitivity and specificity of the biomarker panel plus algorithm to CTCs.</p> <p>Experiments generating the data were performed over the course of 15 days. Peripheral blood samples were collected from the Gynecological Tissue and Fluid Bank (COMIRB 07-0935 / COMIRB 05-1081) from consenting patients undergoing surgery at the University of Colorado Hospital. Blood samples were used the same day they were collected. Blood samples were collected from 3 patients with benign conditions, labeled WBBN#, and 3 patients with ovarian cancer, labeled WBCA#. We do not expect there to be any difference in the isolated white blood cells samples prepared from the cancer and benign patients. Samples were stored at room temperature until white blood cells were isolated. Mixed samples were prepared by passaging a MCF7 flask and mixing it with isolated white blood cells before fixation. A schedule showing the time duration between collection, processing and imaging is included as “experimental schedule.gif”.</p> <p>The MCF7 cancer cell line was a kind gift from Dr. Heide Ford. Genomic DNA was isolated from the MCF7 cell line after the experiment and sent for cell line authentication. The gDNA was a match to MCF7. The authentication report and data are included in this submission.</p> <p>CD45 antibodies were exhausted on day 7. New antibody was purchased and received on day 8. The day 7 images only has labels for DAPI and Bodipy. The samples prepared with the old antibodies on days 4 and 7 were relabeled and imaged with the new antibodies on days 14 and 15. This labeling was also done to confirm the pan-CK antibodies remained good since they are dim in the MCF7 cells imaged on days 12 and 13. The pan-CK on days 14 and 15 looks the same as it did on days 5 and 7 confirming the antibodies are good.</p> <p>Four of the filters containing cells were not sufficiently flat to be acquired with a 3 slice z-stack so a 5 slice z-stack was used. These files have been zipped to compress them under the 2 GB limit permitted by zenodo.org</p> <p>Further information on how these samples were prepared, processed, and analyzed can be found in our associated 2016 SPIE Photonics West BIOS conference proceeding titled, “Quantitative image cytometry measurements of lipids, DNA, CD45 and cytokeratin for circulating tumor cell identification in a model system”, http://dx.doi.org/10.1117/12.2222317.</p> <p>This work was supported by funding provided to the University of Colorado Cancer Center by the American Cancer Society and awarded as Institutional Research Grant Number 57-001-53, by funding provided by the Defense Advanced Research Projects Agency under grant number N66001-10-4035, and by funding provided by NIH/NCATS Colorado CTSI Grant Number TL1 TR001081. The University of Colorado Anschutz Medical Campus Advanced Light Microscopy Core is also supported in part by NIH/NCATS Colorado CTSI Grant Number UL1 TR001082. The funders had no role in the study design, data collection, analysis, or decision to publish.</p>
Stiffness Moduli Modelling and Prediction in Four-Point Bending of Asphalt Mixtures: A Machine Learning-Based Framework within Weave-UNISONO 2021 project, NCN project No 2021/03/Y/ST8/00079, and GACR project GA22-04047K
<div><strong>Summary:</strong></div> <div>Two selected mixtures were thoroughly investigated in an experimental trial carried out by means of a four-point bending test (4PBT) apparatus. The mixtures were prepared using spilite aggregate, a conventional 50/70 penetration grade bitumen, and limestone filler. Their stiffness moduli (SM) were determined while samples were exposed to 11 loading frequencies (from 0.1 to 50 Hz) and 4 testing temperatures (from 0 to 30 °C). Observations were recorded and used to develop a machine learning (ML) model. The main scope was the prediction of the stiffness moduli based on the volumetric properties and testing conditions of the corresponding mixtures, which would provide the advantage of reducing the laboratory efforts required to determine them.</div> <div> </div> <div><strong>The dataset includes:</strong></div> <div>Characteristics of bituminous binder, CSV raw data</div> <div> <ul> <li>bituminous binder.csv</li> </ul> </div> <div>Grading curves of tested asphalt mixtures</div> <ul> <li>AML16 Grading curves.csv</li> <li>AMP22 Grading curves.csv</li> </ul> <div>Volumetric characterizations of AML16 and AMP22 mixtures</div> <ul> <li>AML16 Volumetric characterizations.csv</li> <li>AMP22 Volumetric characterizations.csv</li> </ul> <div>Outcomes of the 4PBT experimental trial carried out on AML16 and AMP22 mixtures</div> <ul> <li>AML16 Stiffness Modulus 4PB.csv</li> <li>AMP22 Stiffness Modulus 4PB.csv</li> </ul>
Example code and data for ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework
<p>This repository contains an R script (grouse_example.R) and data (grouse_data.csv) used to reproduce the grouse abundance analysis described in Kellner, K. F., et al. (2021) ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework. Methods in Ecology and Evolution. The R script requires installation of the ubms R package, which can be obtained from CRAN (https://cran.r-project.org/package=ubms).</p> <p>The repository also contains an additional example occupancy analysis (occupancy_example.R) using the crossbill dataset included with the unmarked R package.</p>
Metrics for two-sample tests: results on Mixture of Gaussians and Correlated Gaussians models
<p>The repository includes version 1.0 (v1.0) of the code and results corresponding to the GitHub repository <a href="https://github.com/TwoSampleTests/GenerativeModelsMetrics">GenerativeModelsMetrics</a>.</p> <p>Publishing information and arXiv identifier will be added after publication of the main manuscript related to the data.</p>
Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - MiSeq - U5-H1-M1-M2-M3-M4-M5 - FASTQ
<p>mtDNA mixture model of 2 mtDNA sequences belonging to haplogroups U5 and H1. Run on Illumina MiSeq with 3 different polymerases (Clontech, Herculase, NEB Taq), and different DNA extraction protocols - Paired-end Fastq files</p> <p>M1 = Mixture 1:2 i.e. 50%</p> <p>M2 = Mixture 1:10 i.e. 10%</p> <p>M3 = Mixture 1:50 i.e. 2%</p> <p>M4 = Mixture 1:100 i.e. 1%</p> <p>M5 = Mixture 1:200 i.e. 0.5%</p>
Fig. 2 in Assessing The Abundance Of Caucasian Salamander, Mertensiella Caucasica (Caudata, Salamandridae), With N-Mixture Model In Northeastern Anatolia
Fig. 2. The average abundance of Caucasian salamanders from the East Black Sea Region, Turkey. X-axis shows the sampling plots number in each city; Y-axis shows the estimated population size.
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Abstract Phylogenetic reconstruction using concatenated loci ("phylogenomics" or "supermatrix phylogeny") is a powerful tool for solving evolutionary splits that are poorly resolved in single gene/protein trees (SGTs). However, recent phylogenomic attempts to resolve the eukaryote root have yielded conflicting results, along with claims of various artefacts hidden in the data. We have investigated these conflicts using two new methods for assessing phylogenetic conflict. ConJak uses whole marker (gene or protein) jackknifing to assess deviation from a central mean for each individual sequence, while ConWin uses a sliding window to screen for incongruent protein fragments (mosaics). Both methods allow selective masking of individual sequences or sequence fragments in order to minimize missing data, an important consideration for resolving deep splits with limited data. Analyses focused on a set of 76 eukaryotic proteins of bacterial-ancestry previously used in various combinations to assess the branching order among the three major divisions of eukaryotes: Amorphea (mainly animals, fungi and Amoebozoa), Diaphoretickes (most other well-known eukaryotes and nearly all algae) and Excavata, represented here by Discoba (Jakobida, Heterolobosea, and Euglenozoa). ConJak analyses found strong outliers to be concentrated in under-sampled lineages, while ConWin analyses of Discoba, the most under-sampled of the major lineages, detected potentially incongruent fragments scattered throughout. Phylogenetic analyses of the full data using an LG-gamma model support a Discoba sister scenario (neozoan-excavate root), which rises to 99-100% bootstrap support with data masked according to either protocol. However, analyses with two site-specific (CAT) mixture models yielded widely inconsistent results and a striking sensitivity to missing data. The neozoan-excavate root places Amorphea and Diaphoretickes as more closely related to each other than either is to Discoba, a fundamental relationship that should remain unaffected by additional taxa.
Roadside turfgrass seed mixtures: models and figures
<div> <div> <div> <div> <p>Roadsides in urban areas are often seeded with turfgrass mixtures to provide ground cover and reduce weed abundance. Designing mixtures to withstand exposure to biotic and abiotic stress is challenging. Research from managed and natural ecosystems have shown that increasing plant species richness and diversity can increase groundcover and suppress weed cover, but it is unclear whether such relationships hold in roadside environments. Our objective was to determine the effect of seeded turfgrass species richness on ground cover and weed suppression alongside roadsides in diverse regions in Minnesota. We tested six turfgrass species in monocultures, two-way mixtures, some three-way mixtures, and a single six-way mixture at seven sites seeded in the fall of 2018, and seven sites seeded in the fall of 2019. Seeded turfgrass, weed, and bare soil coverage was measured at each site over two growing seasons. There was a positive relationship between turfgrass species richness and turfgrass cover, and this interaction effect increased over time. We found that increasing turfgrass species richness reduced bare soil coverage. Turfgrass cover was also more consistent across research sites (i.e., greater spatial stability) with increasing species richness. Our results show that positive relationships between plant species richness and groundcover hold in highly disturbed and managed roadside environments. These findings can improve the design of seed mixtures for roadsides and in other ecological contexts where vegetative cover is important.</p> </div> </div> </div> </div>
Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models
<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model. We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate. We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>
Speech and noise mixtures used in Modelling Auditory Processing and Organisation
<p>Speech and noise signals used in Cooke, M (1991) Modelling Auditory Processing and Organisation, Ph. D. Thesis, Department of Computer Science, University of Sheffield</p>
Machine Learning Models for Surface Wave Dispersion Curve Inversion using Mixture Density Networks
<p>Machine learning (ML) approach for dispersion curve inversion using mixture density networks (MDN) based on Keil and Wassermann (2023).</p> <p>The ML approach presented here allows the simultaneous estimation of layer numbers, layer depth and a complete probability distribution of the S-wave velocity structure in the upper 100 m. This is achieved by a two-step ML approach, where 1) a regular NN classifies the number of layers within the upper 100 m of the subsurface and 2) individual trained mixture density networks output the depth estimates together with a fully probabilistic solution of the S-wave velocity structure. We trained the model to distinguish structures with 2 - 7 subsurface layers.</p> <p>The trained classification NN and the individual MDNs are located in the folder ./trained_models.<br> With the jupyter notebook Prediction.ipynb the dispersion curve inversion can be performed using the already trained ML models.<br> With the jupyter notebooks Training-MDN.ipynb and Training-classification.ipynb the models can be trained on new data.<br> The code for the set-up of the MDN is based on Earp et al. (2020).</p> <p> </p> <p>More details and updates on the code can be found on: <a href="https://github.com/SabrinaKeil/MDN_Inversion">https://github.com/SabrinaKeil/MDN_Inversion</a> </p>
HETEAC – The Hybrid End-To-End Aerosol Classification model for EarthCARE: Look-Up Table (LUT) for aerosol mixtures
<p>The dataset contains the look-up table (LUT) of EarthCARE’s Hybrid End-To-End Aerosol Classification (HETEAC) model. The LUT contains optical and radiative parameters for four pure aerosol components (fine mode weakly absorbing, fine mode strongly absorbing, coarse mode spherical and coarse mode non-spherical) and their mixtures. In total, 314 aerosol mixtures are considered. The LUT returns the mixing state of an aerosol mixture based on the lidar ratio and the particle linear depolarization ratio at 355 nm. The mixing state is expressed in terms of relative volume contribution of the four pure aerosol components. Additionally, the LUT returns the effective radius, the asymmetry parameter, the single scattering albedo (at 355, 532, 550, 670, 865, 1064, 1650 and 2210 nm) and the Angstrom exponent (at 28 wavelength combinations) of the aerosol mixture. The lidar ratio and the particle linear depolarization ratio is also provided at 532, 550, 670, 865, 1064, 1650 and 2210 nm.</p> <p>The datafile contains two top-level groups: the HeaderData, which contains the header variables, and the ScienceData with the variables. The latter contains two groups, the AerosolComponents, which includes the aerosol-component-related optical and microphysical variables, and the LookUpTable, which contains the HETEAC LUT variables.</p> <p>The variables included in the datafile are listed below. For each variable, a full description is provided in the long_name attribute.</p> <ul> <li>HeaderData <ul> <li>angstrom_exponent_header</li> </ul> </li> <li>ScienceData <ul> <li>AerosolComponents <ul> <li>backscatter</li> <li>effective_radius</li> <li>extinction</li> <li>logarithmic_width</li> <li>mode_radius_number</li> <li>mode_radius_volume</li> <li>particle_linear_depolarization_ratio</li> <li>refractive_index_imaginary</li> <li>refractive_index_real</li> <li>scattering</li> </ul> </li> <li>LookUpTable <ul> <li>angstrom_exponent</li> <li>asymmetry_parameter</li> <li>effective_radius</li> <li>lidar_ratio</li> <li>particle_linear_depolarization_ratio</li> <li>relative_volume_contribution</li> <li>single_scattering_albedo</li> </ul> </li> <li>radiation_wavelength</li> </ul> </li> </ul> <p>Contact</p> <p>For any further clarifications or expression of interest with respect to the EarthCARE LUT, please contact Ulla Wandinger (ulla.wandinger@tropos.de) and/or Athena Augusta Floutsi (floutsi@tropos.de).</p> <ul> </ul>
Hyperspectral Mixture Models in the CHIME Mission Implementation for Topsoil Texture Retrieval
<p>This dataset provides the steps of the image analysis techniques used to soil texture classes retrieval related to the paper 'Hyperspectral Mixture Models in the CHIME Mission Implementation for Topsoil Texture Retrieval' in wich the principles of the spectral mixture analyses are used.</p>
Experimental Results for "A Unified Perspective on Natural Gradient Variational Inference with Gaussian Mixture Models"
<p>This package contains the raw data / logs (fetched from WandB) for the experiments of the following publication:</p> <p>O. Arenz, P. Dahlinger, Z. Ye, M. Volpp, and G. Neumann. A unified perspective on natural gradient variational inference with gaussian mixture models. Transactions on Machine Learning Research, 2023. URL: <a href="https://openreview.net/forum?id=tLBjsX4tjs">https://openreview.net/forum?id=tLBjsX4tjs</a>.</p> <p> </p>
Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models
Open the record for dataset details and reuse information.
Evidence of absence regression: a binomial N-mixture model for estimating fatalities at wind power facilities
Open the record for dataset details and reuse information.
Roadside turfgrass seed mixtures: models and figures
Open the record for dataset details and reuse information.
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Open the record for dataset details and reuse information.
Monitoring animal populations with cameras using open, multistate, N-mixture models
Open the record for dataset details and reuse information.
Estimating spawning Green Sturgeon (Acipenser medirostris Ayres, 1854) abundance using side scan sonar and N-mixture models
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.