Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
990
datasets available to search
ShareScore release 0.9.0
Dataset results
990 results for “quantification”
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data B
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part B.</p>
CUSP - UBC Workshop: Analytics: Characterization and quantification / enumeration of particles in the environment and in tissue
<p>The CUSP-UBC Workshop was held online on 28th January 2022.</p> <p>69 participants took part from across Europe and British Columbia to share experiences, exchange knowledge, and to discuss challenges and solutions as part of a great collaboration between the two clusters.</p> <p><strong>Acknowledgements:</strong></p> <p><strong>Co-Organisation and Cluster Presentations:</strong></p> <p>Lesley Tobin (CUSP Working Group 6 Communication and Dissemination, PlasticsFatE) <a href="mailto:lesley.tobin@optimat.co.uk"> </a><a href="mailto:lesley.tobin@optimat.co.uk">lesley.tobin@optimat.co.uk</a></p> <p>Mahdi Takaffoli (Coordinator, Cluster for Microplastics, Health and the Environment, The University of British Columbia) <a href="mailto:mahdi.takaffoli@ubc.ca">mahdi.takaffoli@ubc.ca</a></p> <p><strong>Presenters:</strong></p> <p><strong>Florian Meirer </strong>(Associate Professor, Inorganic Chemistry and Catalysis research group at Utrecht University; Polyrisk & Aurora)</p> <p>“Characterizing Nanoplastics with Force Microscopy – An Update”) <a href="mailto:F.Meirer@uu.nl">F.Meirer@uu.nl</a></p> <p><strong>Anna Costa</strong> (Environmental Nanotechnology and Nano-Safety group of CNR-ISTEC; PlasticsFatE) “Strategies for MP/NP simulated samples-laboratory tests” <a href="mailto:anna.costa@istec.cnr.it">anna.costa@istec.cnr.it</a></p> <p><strong>Tao Huan</strong> (Assistant Professor, Chemistry, The University of British Columbia)</p> <p>“Pilot Study of the Impact of Microplastics on Cell Liability and Potential Application of Metabolomics in Understanding the Biological Mechanisms” <a href="mailto:thuan@chem.ubc.ca">thuan@chem.ubc.ca</a></p> <p><strong> Edward Grant</strong> (Professor, UBC Chemistry) “The challenge of representative microplastic analysis” edgrant@chem.ubc.ca</p> <p>Thank you to Michelle Epstein, Doctor of Allergy and Clinical Immunology, MedUni Vienna, for such a useful, stimulating idea, and to everyone who took part despite the unsocial hours!</p> <p> </p>
Probabilistic simulation of big climate data for robust quantification of changes in compound hazard events
<p>Data, code and supplementary Figures for paper "Probabilistic simulation of big climate data for robust quantification of changes in compound hazard events".</p>
Dataset - Controlled release experiment to investigate uncertainties in UAV-based emission quantification for methane point sources
<p>This dataset was created by Randulph Morales (randulph.morales@empa.ch) and was used for Morales et al. (2021) AMT publication (amt-2021-314). A short description of the files is written in <strong>readme.txt</strong></p> <p>The dataset contains:</p> <ul> <li>QCLAS methane measurement</li> <li>Active AirCore methane measurement</li> <li>Meteorology files</li> </ul>
Monte-Carlo-simulated MR spectra of the rat brain and their quantification results obtained with QUEST and QUEST-MM jMRUI algorithms
<p>MC-simulated spectra of the rat brain along with the corresponding basis set (metabolite signals simulated using NMRScopeB from jMRUI) and the QUEST-MM, QUEST, QUEST(Met+Back) and QUEST(Met+MM) quantification results are stored in the MC_results folder.</p> <p>Results of an in-vivo rat brain MRS signal (SPECIAL, dead time t0 = 0.134 ms, TE = 2.8 ms, at 9.4 T) quantification performed with the QUEST-MM jMRUI algorithm (origin for MC-simulation) are stored in the folder rat_results.</p> <p>All files can be loaded to jMRUI software version 5 and later.</p>
Dataset from Remote analysis of Sputum Smears for Mycobacterium Tuberculosis Quantification using Digital Crowdsourcing
<p>Worldwide, TB is one of the top 10 causes of death and the leading cause from a single infectious agent. Although the development and roll out of Xpert MTB/RIF has recently become a major breakthrough in the field of TB diagnosis, smear microscopy remains the most widely used method for TB diagnosis, especially in low- and middle-income countries.</p> <p>This is a minimal dataset to reproduce our research that tests the feasibility of a crowdsourced approach to tuberculosis image analysis. In particular, we investigated whether anonymous volunteers with no prior experience would be able to count acid-fast bacilli in digitized images of sputum smears by playing an online game. Following this approach 1790 people identified the acid-fast bacilli present in 60 digitized images, the best overall performance was obtained with a specific number of combined analysis from different players and the performance was evaluated with the F1 score, sensitivity and positive predictive value, reaching values of 0.933, 0.968 and 0.91, respectively.</p> <p>The dataset includes 24 digitized images of sputum smears and the corresponding gameplays clicks. </p>
Data for A Locally Activatable Sensor for Robust Quantification of Organellar Glutathione
<p>Supporting data to paper A Locally Activatable Sensor for Robust Quantification of Organellar Glutathione,</p> <p>including NMR, MS, microscopy, etc</p>
Dataset for 'Experimental Quantification of Gas Dispersion in 3D-Printed Logpile Structures Using a Noninvasive Infrared Transmission Technique'
<p>This dataset contains the infrared images of tracer flow that were taken in the investigations of transverse dispersion in 3D-printed logpile structures. Accompanying the files (which are labelled according to the convention of the camera software) is a Python script which can be used to link the images to the operating conditions at which they were obtained. Documentation of this script can be found in the file at the very top. <br> This dataset was used as basis for the journal article 'Experimental Quantification of Gas Dispersion in 3D-Printed Logpile Structures Using a Noninvasive Infrared Transmission Technique', published in ACS Engineering Au under DOI:<a href="https://doi.org/10.1021/acsengineeringau.1c00040">10.1021/acsengineeringau.1c00040</a>. This paper can also be found in this repository at https://zenodo.org/record/6517082</p> <p> </p>
Dataset and code for paper "An automated quantification tool for angiogenic sprouting from endothelial spheroids"
<p>This repository contains raw data and code for the manuscript with DOI: 10.3389/fphar.2022.883083</p>
SupportingDataset Identification and Quantification of Within-Burst Dynamics in Singly-Labeled Single-Molecule Fluorescence Lifetime Experiments
<p>The Jupyter notebooks and resulting files used to demonstrate divisor-based mpH<sup>2</sup>MM. The analysis is demonstrated with both simulations and analyses of alpha-synuclein.</p> <ol> <li> <p><em><strong>Notebooks.zip</strong></em>:* Zip file containing the Jupyter notebooks for producing, analyzing and visualizing the simulated photon trajectories. Note: this folder contains all code needed to reproduce simulations. All other files related to the simulations are produced by one of the notebooks in this trajectory. However, as simulations can take a long time, the various results files are included in this repository so that notebooks can be run from intermediate steps.</p> <ol> <li> <p><strong>1-PIFE-pybromo-sims.ipynb</strong> : The code for producing simulated diffusion trajectories and photon-HDF5 files of two-state systems undergoing transition dynamics (the results of this notebook are stored in the sub-folder <em>PyBroMo_photonHDF5</em>)</p> </li> <li> <p><strong>2-PIFE-mpH2MM-sim-[lifetime components].ipynb </strong>: Notebooks performing divisor-based mpH<sup>2</sup>MM on simulated datasets for a given combination of lifetime states. (these notebooks store files that are contained in the sub-folder <em>H2MMresults</em>)</p> </li> <li> <p><strong>3-PIFE-mpH2MM-compiled-plots.ipynb</strong> : Jupyter notebook for producing figures comparing all results globally</p> </li> <li> <p><strong>532nm_IRF_19-10-2021.csb</strong>: the file containing the experimental IRF used in the simulations</p> </li> </ol> </li> <li> <p><em><strong>PyBroMo_photonHDF5.zip</strong></em>:* Zip file containing the simulated results of <em>1-PIFE-pybromo-sims</em> notebook as photon-HDF5 files (1 file per transition rate/lifetime combination)</p> </li> <li> <p><strong>PIFE-sim-dynamicmix_[lifetime components]_result.hdf5</strong>: special HDF5 files containing the results of each notebook in <em>Notebooks</em>, which are used by <em>3-PIFE-mpH2MM-compiled-plots</em></p> </li> <li> <p><strong>PIFE-mpH2MM-alpha-syn-vFinal.ipynb</strong>: divisor-based mpH<sup>2</sup>MM analysis of alpha-synuclein smPIFE data</p> </li> <li> <p><strong>H2MM-Lifetime_example.ipynb</strong>: A demonstration of divisor-based mpH<sup>2</sup>MM using nsALEX-smFRET data. This method could potentially demonstrate states differentiated in lifetimes independently of potential changes in E & S.</p> </li> <li> <p><strong>Template_ltH2MM.ipynb</strong>: An easy-to-follow implementation of divisor-based mpH<sup>2</sup>MM demonstrated on a single alpha-synuclein experimental data acquisition file. This can be used for learning how to implement and analyze single dye fluorescence lifetime data with mpH<sup>2</sup>MM</p> </li> </ol> <p> </p> <p>* For running these notebooks, generally, all files in <em>PyBroMo_photonHDF5.zip</em> should be placed into a single directory (i.e., the files in <em>Notebooks</em>.<em>zip</em> should be placed into the same directory as the files in <em>PyBroMo_photonHDF5</em>.<em>zip</em>) as the notebooks are set to read in files from their current directory.</p>
Number of chamber measurement locations for accurate quantification of landscape-scale greenhouse gas fluxes: Importance of land use, seasonality, and greenhouse gas type
<p>Contains all raw data measured in the Schwingbach Earth Observatory (SEO) from Spring, Summer and Autumn 2020. Data was measured with an on-site LGR laser from the GHG emissions, and with 100cm³ soil cores for the soil characteristics. Details can be found in the corresponding manuscript "Number of chamber measurement locations for accurate quantification of landscape-scale greenhouse gas fluxes: Importance of land use, seasonality, and greenhouse gas type"</p>
Dataset for "Quantification of Muscle Fiber Malformations Using Edge Detection to Investigate Chronic Wound Healing"
<p>Primary images, spreadsheets and files for "Quantification of Muscle Fiber Malformations Using Edge Detection to Investigate Chronic Wound Healing"</p>
Quantification of recyclability indicators for three case studies
<p>Supplementary data table belonging to the original research publication entitled <em>Early-stage assessment of minor metal recyclability</em>.</p>
R code and associated data for: A review of riverine ecosystem service quantification: research gaps and recommendations
<p>This publication contains the R code and associated data used in the Journal of Applied Ecology publication entitled "A review of riverine ecosystem service quantification: research gaps and recommendations". </p>
CompanyKG Dataset V2.0: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
<p><strong>CompanyKG</strong> is a heterogeneous graph consisting of 1,169,931 nodes and 50,815,503 undirected edges, with each node representing a real-world company and each edge signifying a relationship between the connected pair of companies.</p> <p><strong>Edges</strong>: We model 15 different inter-company relations as undirected edges, each of which corresponds to a unique edge type. These edge types capture various forms of similarity between connected company pairs. Associated with each edge of a certain type, we calculate a real-numbered weight as an approximation of the similarity level of that type. It is important to note that the constructed edges do not represent an exhaustive list of all possible edges due to incomplete information. Consequently, this leads to a sparse and occasionally skewed distribution of edges for individual relation/edge types. Such characteristics pose additional challenges for downstream learning tasks. Please refer to our paper for a detailed definition of edge types and weight calculations.</p> <p><strong>Nodes</strong>: The graph includes all companies connected by edges defined previously. Each node represents a company and is associated with a descriptive text, such as "<em>Klarna is a fintech company that provides support for direct and post-purchase payments</em> ...". To comply with privacy and confidentiality requirements, we encoded the text into numerical embeddings using four different pre-trained text embedding models: <a href="https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v2">mSBERT</a> (multilingual Sentence BERT), <a href="https://platform.openai.com/docs/guides/embeddings/what-are-embeddings">ADA2</a>, <a href="https://github.com/princeton-nlp/SimCSE">SimCSE</a> (fine-tuned on the raw company descriptions) and <a href="https://github.com/EQTPartners/pause">PAUSE</a>.</p> <p><strong>Evaluation Tasks</strong>. The primary goal of CompanyKG is to develop algorithms and models for quantifying the similarity between pairs of companies. In order to evaluate the effectiveness of these methods, we have carefully curated three evaluation tasks:</p> <ul> <li><strong>Similarity Prediction (SP)</strong>. To assess the accuracy of pairwise company similarity, we constructed the SP evaluation set comprising 3,219 pairs of companies that are labeled either as positive (similar, denoted by "1") or negative (dissimilar, denoted by "0"). Of these pairs, 1,522 are positive and 1,697 are negative.</li> <li><strong>Competitor Retrieval (CR)</strong>. Each sample contains one <em>target company</em> and one of its direct competitors. It contains 76 distinct target companies, each of which has 5.3 competitors annotated in average. For a given target company A with <em>N</em> direct competitors in this CR evaluation set, we expect a competent method to retrieve all <em>N</em> competitors when searching for similar companies to A. </li> <li><strong>Similarity Ranking (SR)</strong> is designed to assess the ability of any method to rank <em>candidate companies</em> (numbered 0 and 1) based on their similarity to a <em>query company</em>. Paid human annotators, with backgrounds in engineering, science, and investment, were tasked with determining which candidate company is more similar to the query company. It resulted in an evaluation set comprising 1,856 rigorously labeled ranking questions. We retained 20% (368 samples) of this set as a validation set for model development. </li> <li><strong>Edge Prediction (EP)</strong> evaluates a model's ability to predict future or missing relationships between companies, providing forward-looking insights for investment professionals. The EP dataset, derived (and sampled) from new edges collected between April 6, 2023, and May 25, 2024, includes 40,000 samples, with edges not present in the pre-existing CompanyKG (a snapshot up until April 5, 2023).</li> </ul> <p><strong>Background and Motivation</strong></p> <p>In the investment industry, it is often essential to identify similar companies for a variety of purposes, such as market/competitor mapping and Mergers & Acquisitions (M&A). Identifying comparable companies is a critical task, as it can inform investment decisions, help identify potential synergies, and reveal areas for growth and improvement. The accurate quantification of inter-company similarity, also referred to as <strong>company similarity quantification</strong>, is the cornerstone to successfully executing such tasks. However, company similarity quantification is often a challenging and time-consuming process, given the vast amount of data available on each company, and the complex and diversified relationships among them.</p> <p>While there is no universally agreed definition of company similarity, researchers and practitioners in PE industry have adopted various criteria to measure similarity, typically reflecting the companies' operations and relationships. These criteria can embody one or more dimensions such as industry sectors, employee profiles, keywords/tags, customers' review, financial performance, co-appearance in news, and so on. Investment professionals usually begin with a limited number of companies of interest (a.k.a. seed companies) and require an algorithmic approach to expand their search to a larger list of companies for potential investment. </p> <p>In recent years, transformer-based Language Models (LMs) have become the preferred method for encoding textual company descriptions into vector-space embeddings. Then companies that are similar to the seed companies can be searched in the embedding space using distance metrics like cosine similarity. The rapid advancements in Large LMs (LLMs), such as GPT-3/4 and LLaMA, have significantly enhanced the performance of general-purpose conversational models. These models, such as ChatGPT, can be employed to answer questions related to similar company discovery and quantification in a Q&A format.</p> <p>However, graph is still the most natural choice for representing and learning diverse company relations due to its ability to model complex relationships between a large number of entities. By representing companies as nodes and their relationships as edges, we can form a <strong>Knowledge Graph (KG)</strong>. Utilizing this KG allows us to efficiently capture and analyze the network structure of the business landscape. Moreover, KG-based approaches allow us to leverage powerful tools from network science, graph theory, and graph-based machine learning, such as Graph Neural Networks (GNNs), to extract insights and patterns to facilitate similar company analysis. While there are various company datasets (mostly commercial/proprietary and non-relational) and graph datasets available (mostly for single link/node/graph-level predictions), there is a scarcity of datasets and benchmarks that combine both to create a large-scale KG dataset expressing rich pairwise company relations.</p> <p><strong>Source Code and Tutorial:<br></strong><a href="https://github.com/llcresearch/CompanyKG2"><strong>https://github.com/llcresearch/CompanyKG2</strong></a></p> <p><strong>Paper: to be published<br></strong></p>
Manual quantification of peroxisome counts in yeast from 2-channel fluorescence Z-stacks
<p>This dataset contains fluorescence microscopy imaging data from various strains of <em>Saccharomyces cerevisiae</em>. The images were used to test software called <em>perox-per-cell,</em> which automatically quantifies peroxisome features in yeast cells based on microscopy data. There are 44 imaging instances in the dataset, each consisting of two Z-stacks, one capturing signal from calcofluor white to identify cell boundaries (blue channel), and one capturing signal from GFP tagged with peroxisome targeting sequence 1 (PTS1) to locate peroxisomes (green channel). These raw microscopy imaging sets are provided as ZVI files in <strong>Zstacks.zip</strong>.</p> <p>We compared <em>perox-per-cell</em>'s automatically-generated peroxisome counts to those derived manually by two individuals. For manual counting, images were deconvolved with theoretically generated point spread functions using Axiovision software V4.9.1 SP2 followed by the generation of maximum intensity Z-projections (MIP) of both blue and green channels. All the deconvolved MIP images from WT and mutant strains were blinded and labelled as ‘1-44’, and their grey levels were set to ‘best fit’ in the Axiovision software prior to providing them to two individuals who manually counted peroxisomes in cells using the ‘measure events’ tool in Axiovision. The maximum intensity projection images used for manual counting are provided as ZVI files in <strong>MaxIntensityProjections.zip</strong>.</p> <p>Each individual's manual counts are included in this dataset within the CSV file <strong>ManualPeroxisomeCounts.csv</strong>. Please note that cell IDs in this file are only indicative of the order in which each individual counted peroxisomes, they do not indicate a specific cell within an image. For example, "Cell5" in Image 3 that was processed by manual counter 1 may not be the same cell as "Cell5" in Image3 processed by manual counter 2. These two entries have the same cell ID only because for both manual counters, they were the 5th cell counted.</p> <p>For our software test, we used wild-type (WT) yeast strains as well as several mutant strains with known peroxisomal defects. The strains used for each image are indicated in the<em> </em><strong>ImageAndStrainTable.csv</strong> file.</p> <p>Experimental details: <em>Saccharomyces cerevisiae</em> cells were grown in synthetic defined medium (SD: 6.7 g/L Yeast nitrogen base without amino acids + 0.79 g/L CSM) with 2% Dextrose in flask cultures shaken at 250 rpm at 30 °C until log phase after which they were pelleted and resuspended in 50 µg/ml calcofluor white stain (Sigma, Cat No. 18909) for 5-10 min followed by imaging at room temperature. 3D images consisting of 26 XY images with a Z-slice spacing of 0.204 µm (total Z-stack thickness 5.1 µm) were acquired at 100× magnification using a fluorescence microscope (Axioskop 2 MOT plus, Carl Zeiss, Inc.) equipped with a Plan Apochromat 100×/1.4 Oil DIC objective, an Axio Cam HRm camera and an HBO 100 Mercury lamp. Identical exposure times (50 ms) were used to acquire the green channel images whereas the exposure time for blue channel was adjusted for individual images based on the intensity of calcofluor staining. </p> <p> </p> <p> </p>
Nissl_4, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains some images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_3, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_2, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p>
Nissl_1, Raw images for Machine learning for histological annotation and quantification of cortical layers.
<p>This dataset contains images (TIFF image data) of <strong>brain juvenile rats Wistar Han (P14)</strong> scanned immunostained slides by using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel.</p> <p>These raw images are part of another Zenodo dataset <span><a href="https://doi.org/10.5281/zenodo.11544829" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.11544829</a></span>, that contains the QuPath projects that uses this dataset and 5 others (from Nissl_1 to Nissl_6).</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.