Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Model Counting”
Evaluation of Mask R-CNN Model for Counting Reproductive Structures of Six Plant Species 1895-2018
Phenology––the timing of life-history events––is a key trait for understanding responses of organisms to climate. The digitization and online mobilization of herbarium specimens is rapidly advancing our understanding of plant phenological response to climate and climatic change. The current common practice of manually harvesting data from individual specimens greatly restricts our ability to scale data collection to entire collections. Recent investigations have demonstrated that machine-learning models can facilitate data collection from herbarium specimens. However, present attempts have focused largely on simplistic binary coding of reproductive phenology (e.g., flowering or not). Here, we use crowd-sourced phenological data of numbers of buds, flowers, and fruits of more than 3000 specimens of six common wildflower species of the eastern United States (Anemone canadensis, A. hepatica, A. quinquefolia, Trillium erectum, T. grandiflorum, and T. undulatum} to train a model using Mask R-CNN to segment and count phenological features. A single global model was able to automate the binary coding of reproductive stage with greater than 90% accuracy. Segmenting and counting features were also successful, but accuracy varied with phenological stage and taxon. Counting buds was significantly more accurate than flowers or fruits. Moreover, botanical experts provided more reliable data than either crowd-sourcers or our Mask R-CNN model, highlighting the importance of high-quality human training data. Finally, we also demonstrated the transferability of our model to automated phenophase detection and counting of the three Trillium species, which have large and conspicuously-shaped reproductive organs. These results highlight the promise of our two-phase crowd-sourcing and machine-learning pipeline to segment and count reproductive features of herbarium specimens, providing high-quality data with which to study responses of plants to ongoing climatic change.
CMIP6 variable counts per model
<p>The number of variables (y-axis) published for the historical simulation by each model (as represented in the DKRZ Earth System Grid Federation (ESGF) index node March 2022) is shown in blue columns against the model rank, where models are ranked in order of decreasing variable count. Also shown, in orange, is the number of variables which are included by all models up to the given rank.</p> <p>Data provided by Martin Juckes, image created by Beth Dingley</p>
Detection of ultra-weak laser pulses by free-running single-photon detectors: modeling dead time and dark counts effects
<p>In quantum communication systems, the precise estimation of the detector´s response to the incoming light is necessary to avoid security breaches. The typical working regime uses a free-running single-photon avalanche diode in combination with attenuated laser pulses at telecom wavelength for encoding information. We demonstrate the validity of an analytical model for this regime which considers the effects of dark counts and dead time on the measured count rate. For the purpose of gaining a better understanding of these effects, the photon detections were separated from the dark counts via a software-induced gating mechanism. The model was verified by experimental data for mean photon numbers covering three orders of magnitude as well as for laser repetition frequencies below and above the inverse dead time. Consequently, our model would be of interest for predicting the detector response not only in the field of quantum communications, but also in any other quantum physics experiment where high detection rates are needed.</p>
Model Counting and Uniform Sampling Instances
<p>These instances mainly consist of the formulas that have been used in the evaluation of recent model counting techniques. A significant set of benchmarks involving sampling set, i.e., they are meant for projected model counting.<br> <br> The specification for reading such files can be found at <a href="https://github.com/meelgroup/approxmc">https://github.com/meelgroup/approxmc</a><br> <br> Here is list of some of the papers that have reported results on these instances:</p> <p>1. BIRD: Engineering an Efficient CNF-XOR SAT Solver and its Applications to Approximate Model Counting<br> Mate Soos and Kuldeep S. Meel<br> Proceedings of AAAI Conference on Artificial Intelligence (AAAI), 2019.</p> <p>2. Accelerating Approximate Techniques for Counting and Sampling Models Through Refined CNF-XOR Solving<br> Mate Soos, Stephan Gocht, and Kuldeep S. Meel<br> Proceedings of International Conference on Computer-Aided Verification (CAV), 2020.<br> </p>
Code and supplementary plots for "Flexible distributed lag models for count data using mgcv"
<p>R code and supplementary plots accompanying the paper: "Flexible distributed lag models for count data using mgcv".</p>
A Parameterized Model for Differential Galaxy Counts at Any Wavelength
<p>Collected literature values of Schechter functions, resulting parameter values of fits to those data as a function of wavelength, and Python scripts to compute simulated differential galaxy counts based on those parameterisations.</p> <p> </p> <p>Future revisions of galaxy_counts.py can be found at https://github.com/Onoddil/macauff/blob/main/macauff/ galaxy counts.py, and future revisions of literature and parameterisations can be found at https://onoddil.github.io/galaxy evolution/galaxy counts.html.</p>
Ruthwell Cross 3D Model High Resolution (112M poly count)
<p>This is a very High-Resolution 3D model the Ruthwell Cross, developed as part of the ongoing Visionary Cross project. This model has a 112M poly count.</p> <p>This record contains:</p> <ul> <li>An xml record: <a href="http://zenodo.org/record/1490878/_Ruthwell_3DModel_112M_Metadata.xml">_Ruthwell_3DModel_112M_Metadata.xml</a></li> <li> A 3D Model of the Ruthwell cross: <a href="https://zenodo.org/record/1490878/cross_ColorMapped_112M.ply">cross_ColorMapped_112M.ply</a> </li> <li>A 2D thumbnail: <a href="https://zenodo.org/record/1490878/Ruthwell_Cross00.png">Ruthwell_Cross00.png</a></li> </ul> <pre>The DOI for this version of the record is <a href="https://zenodo.org/record/1490878">10.5281/zenodo.1490878</a>. The DOI <a href="https://10.5281/zenodo.1233638">10.5281/zenodo.1233638</a> always points to the latest version of this record.</pre> <p>This file in .ply format is best viewed using 3DHOP, but can also be viewed using 3D rendering software like MeshLab.</p> <p>For individual panels and other related material go to The Visionary Cross community at: https://zenodo.org/communities/the_visionary_cross/</p>
Counting animals in aerial images with a density map estimation model
<p>Animal abundance estimation is increasingly based on drone or aerial survey photography. Manual post-processing has been used extensively, however, volumes of such data are increasing, necessitating some level of automation, either for complete counting or as a labour-saving tool. Any automated processing can be challenging when using such tools on species that nest in close formation such as <em>Pygoscelis</em> penguins. We present here a customized CNN-based density map estimation method for counting of penguins from low-resolution aerial photography. Our model, an indirect regression algorithm, performed significantly better in terms of counting accuracy than standard detection algorithm (Faster RCNN) when counting small objects from low-resolution images and gave an error rate of only 0.8 percent. Density map estimation methods as demonstrated here can vastly improve our ability to count animals in tight aggregations, and demonstrably improve monitoring efforts from aerial imagery. </p>
Supporting Material for "Lunar Surface Model Age Derivation: Comparisons Between Automatic and Human Crater Counting Using LRO-NAC And Kaguya TC Images"
<p>Supporting Material for "Lunar Surface Model Age Derivation: Comparisons Between Automatic and Human Crater Counting Using LRO-NAC And Kaguya TC Images"</p> <p>Contents of this material</p> <ul> <li>Supplemental Text S1 and Text S2.</li> <li>Figures S1, S2, S2, S4, S5.</li> <li>Tables S1, S2</li> </ul> <p>For any questions email JHF (john.h.fairweaher@gmail.com).</p>
Counting animals in aerial images with a density map estimation model
Open the record for dataset details and reuse information.
Model Age Derivation of Large Martian Impact Craters, using automatic crater counting methods / Dataset
<ul> <li>"counting_area" folder: shapefiles of the mapped ejecta layers considered in this study</li> <li>"scc" folder: .scc files readable on CraterStats listing the size and location of craters detected by our CDA and recognized as primaries by the ASCI. The counting area considered for each crater slightly vary from the area indicated in the shapefile due to the removal of Thiessen polygons associated to secondary craters by the ASCI.</li> </ul> <p>The ASCI code and toolbox implementable to ESRI ArcGIS (10.6) is discoverable here: https://github.com/curtin-crater-detection/secondary-crater-removal<br> </p>
Weekly travel times by direction and sample count for each link in Development of AIS Model of Texas Gulf Intracoastal Waterway Travel Times
<p>Excel spreadsheet containing all of the travel times and sample counts for each link (by direction).</p>
Dataset for On the Sparsity of XORs in Approximate Model Counting (SAT-20 Paper)
<p>The artifact consists of the necessary data to reproduce the results reported in the SAT-20 Paper titled "On the Sparsity of XORs in Approximate Model Counting". <br> <br> In particular, the artifact consists of the binaries, the log files generated by our computing cluster, and scripts to generate tables and the plots used in the paper. </p>
Complex ecological phenotypes on phylogenetic trees: a Markov process model for comparative analysis of multivariate count data
The evolutionary dynamics of complex ecological traits – including multistate representations of diet, habitat, and behavior – remain poorly understood. Reconstructing the tempo, mode, and historical sequence of transitions involving such traits poses many challenges for comparative biologists, owing to their multidimensional nature. Continuous-time Markov chains (CTMC) are commonly used to model ecological niche evolution on phylogenetic trees but are limited by the assumption that taxa are monomorphic and that states are univariate categorical variables. A necessary first step in the analysis of many complex traits is therefore to categorize species into a pre-determined number of univariate ecological states, but this procedure can lead to distortion and loss of information. This approach also confounds interpretation of state assignments with effects of sampling variation because it does not directly incorporate empirical observations for individual species into the statistical inference model. In this study, we develop a Dirichlet-multinomial framework to model resource use evolution on phylogenetic trees. Our approach is expressly designed to model ecological traits that are multidimensional and to account for uncertainty in state assignments of terminal taxa arising from effects of sampling variation. The method uses multivariate count data for individual species to simultaneously infer the number of ecological states, the proportional utilization of different resources by different states, and the phylogenetic distribution of ecological states among living species and their ancestors. The method is general and may be applied to any data expressible as a set of observational counts from different categories.
Complex ecological phenotypes on phylogenetic trees: a Markov process model for comparative analysis of multivariate count data
Open the record for dataset details and reuse information.
Protea repens whole transcriptome count data for control and drought treatment for 8 populations, climatic data for the 8 populations and phenotypic data collected, and data used for linear mixed models for climate gene expression/trait correlation testing
Open the record for dataset details and reuse information.
Avian point-counts from Rhode Island and Connecticut used to test species distribution models
<p>Spatial-biases are a common feature of presence-absence data from citizen scientists. Spatial thinning can mitigate errors in species distribution models (SDMs) that use these data. When detections or non-detections are rare, however, SDMs may suffer from class imbalance or low sample size of the minority (i.e. rarer) class. Poor predictions can result, the severity of which may vary by modeling technique. To explore the consequences of spatial bias and class imbalance in presence-absence data, we used eBird citizen science data for 102 bird species from the northeastern USA to compare spatial thinning, class balancing, and majority-only thinning (i.e., retaining all samples of the minority class). We created SDMs using two parametric or semi-parametric techniques (generalized linear models and generalized additive models) and two machine-learning techniques (random forest and boosted regression trees). We tested the predictive abilities of these SDMs using an independent and systematically collected reference dataset with a combination of discrimination (area under the receiver operator characteristic curve; true skill statistic; area under the precision-recall curve) and calibration (Brier score; Cohen's kappa) metrics. We found large variation in SDM performance depending on thinning and balancing decisions. Across all species, there was no single best approach, with the optimal choice of thinning and/or balancing depending on modeling technique, performance metric, and the baseline sample prevalence of species in the data. Spatially thinning all the data was often a poor approach, especially for species with baseline sample prevalence < 0.1. For most of these rare species, balancing classes improved model discrimination between presence and absence classes, but hindered model calibration. Baseline sample prevalence, sample size, modeling approach, and the intended application of SDM output – whether discrimination or calibration – should guide decisions about how to thin or balance data, given the considerable influence of these methodological choices on SDM performance. For prognostic applications requiring good model calibration (vis-à-vis discrimination), the match between sample prevalence and true species prevalence may be the overriding feature and warrants further investigation.</p>
Data from: Estimating density for species conservation: comparing camera trap spatial count models to genetic spatial capture-recapture models
Density estimation is integral to the effective conservation and management of wildlife. Camera traps in conjunction with spatial capture-recapture (SCR) models have been used to accurately and precisely estimate densities of "marked" wildlife populations comprising identifiable individuals. The emergence of spatial count (SC) models holds promise for cost-effective density estimation of "unmarked" wildlife populations when individuals are not identifiable. We evaluated model agreement, precision, and survey costs, between i) a fully marked approach using SCR models fit using non-invasive genetic data, and ii) an unmarked approach using SC models fit using camera trap data, for a recovering population of the mesocarnivore fisher (Pekania pennanti). The SCR density estimates ranged from 2.95 to 3.42 (2.18–5.19 95% BCI) fishers 100 km−2. The SC density estimates were influenced by their priors, ranging from 0.95 (0.65–2.95 95% BCI) fishers 100 km−2 for the uninformative model to 3.60 (2.01–7.55 95% BCI) fishers 100 km−2 for the model informed by prior knowledge of a 16 km2 fisher home range. We caution against using strongly informative priors but instead recommend using a range of unweighted prior knowledge. Thin detection data was problematic for both SCR and SC models, potentially producing biased low estimates. The total cost of the genetic survey ($47 610) was two-thirds of the camera trap survey ($77 080), or comparable ($75 746) if genetic sampling effort was increased to include sex and trap-behaviour covariates in SCR models. Density estimation of unmarked populations continues to be a series of trade-offs but as methods improve and integrate, so will our estimates.
Model Counting Competition 2020: Submitted Solvers
<p>The dataset contains the submissions that have been evaluated in the Model Counting Competition 2020 on the tracks:</p><p>- Track 1 (Model Counting)<br>- Track 2 (Weighted Model Counting)<br>- Track 3 (Projected Model Counting)</p><p><br>Details will be made public in the publication "The Model Counting Competition 2020" by Fichte, Hecher, and Hamiti. Due to numerous names, we omit authors of the solvers from the meta information of this dataset. We refer to the aforementioned report for details.</p><p>We also refer to the competition website: https://mccompetition.org/</p><p> </p><p>Changelist:</p><p>2023-10-21 (v2): We updated wrappers to take the more recent competition format (MC2021 format), which is DIMACS CNF compatible.</p>
Model Counting Competition 2020: Competition Instances
<p>Instances for the Model Counting Competition 2020</p> <p>- Track 1 (Model Counting)<br>- Track 2 (Weighted Model Counting)<br>- Track 3 (Projected Model Counting)</p> <p><br>The even instances (trackX_000.mcc2020_cnf, trackX_000.mcc2020_cnf, ...) were made public for all participants during the testing phase of the solvers, whereas the private instances (trackX_001.mcc2020_cnf, trackX_003.mcc2020_cnf, ...) were used for the final evaluation and disclosed after the submission. For details, we refer to: <a href="https://mccompetition.org/2020/mc_description">https://mccompetition.org/2020/mc_description</a>.<br>Instances originate from various publicly available data sets and submissions made after a call for benchmarks. The full instances from which we selected is available on Zenodo under Model Counting Competition 2020: Full Instance Set.<br>For details, see the MCC2020 report (Fichte, Hecher, Hamiti: The Model Counting Competition 2020).</p> <p>2023-10-22 (v3): Updated instances to the most recent competition format (MC2021).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.