Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
295
datasets available to search
ShareScore release 0.9.0
Dataset results
295 results for “approximation”
Estimating total species richness: fitting rarefaction by asymptotic approximation
Open the record for dataset details and reuse information.
Identifying the best approximating model in Bayesian phylogenetics: Bayes factors, cross-validation or wAIC?
Open the record for dataset details and reuse information.
All simulation results, figures and code regarding the manuscript: Calibrating models of cancer invasion: parameter estimation using Approximate Bayesian Computation and gradient matching
Open the record for dataset details and reuse information.
High-Resolution Orthorectified Imagery from Approximately 1990, Niwot Ridge LTER Project Area, Colorado
Citation: Manley, W.F., Parrish, E.G., and Lestak, L.R., 2009, High-Resolution Orthorectified Imagery and Digital Elevation Models for Study of Environmental Change at Niwot Ridge and Green Lakes Valley, Colorado: Niwot Ridge LTER, INSTAAR, University of Colorado at Boulder, digital media. This image is a mosaic of orthorectified aerial photography from 1988 and 1990 for the Niwot Ridge Long Term Ecological Research (LTER) project area at 0.6 m resolution. The image also covers the Green Lakes Valley portion of the Boulder Creek Critical Zone Observatory (CZO). The mosaic has the qualities of a photograph and the functionality of a map layer for use in Geographic Information Systems (GIS) or remote sensing software. The mosaic is derived from approx. 1:40,000 scale, color infrared (CIR) photographs acquired by the United States Geological Survery (USGS) National Aerial Photography Program (NAPP). The aerial photos were obtained as 1800 dpi digital scans from the USGS EROS Data Center (EDC) and then fully orthorectified in a Leica Photogrammetry Suite (LPS) bundle blockfile using an air-photo camera model, a Digital Elevation Model (DEM), and known focal length and fiducial coordinates from a calibration report. Individual photo frames were mosaiced with cutlines and clipped to the Niwot project extent area. The photography was registered to 2008 orthocorrected Denver Region Council of Governments (DRCOG) aerial photography. Horizontal accuracy is 0.9 m (RMSE, relative to the 2008 reference imagery, based on 9 independent check points). The mosaic covers an area of 98 km2 and is available in GeoTIFF format, in a UTM zone 13 projection and NAD83 horizontal datum, with FGDC-compliant metadata. The mosaic is available through an unrestricted public license, and can be obtained by request (see Distributor contact information below). Other datasets available in this series includes orthorectified aerial photograph mosaics (for 1953, 1972, 1985, 1999, 2000, 2002, 2004, 200
Source Index Map Layer for High-Resolution Orthorectified Imagery from Approximately 1990, Niwot Ridge LTER Project Area, Colorado
Citation Manley, W.F., Parrish, E.G., and Lestak, L.R., 2009, High-Resolution Orthorectified Imagery and Digital Elevation Models for Study of Environmental Change at Niwot Ridge and Green Lakes Valley, Colorado: Niwot Ridge LTER, INSTAAR, University of Colorado at Boulder, digital media. This vector shapefile is a source index map layer for the mosaic of orthorectified aerial photography from 1988 and 1990 for the Niwot Ridge Long Term Ecological Research (LTER) project. The index also covers the Green Lakes Valley portion of the Boulder Creek Critical Zone Observatory (CZO). The index polygons are attributed with source photo date and photo year. The mosaic is derived from approx. 1:40,000 scale, color infrared (CIR) photographs acquired by the United States Geological Survery (USGS) National Aerial Photography Program (NAPP). Other datasets available in this series includes orthorectified aerial photograph mosaics (for 1953, 1972, 1985, approximately 1990, 1999, 2000, 2002, 2004, 2006 and 2008), digital elevation models (DEM's), and accessory map layers. Together, the DEM's and imagery will be of interest to students, research scientists, and others for observation and analysis of natural features and ecosystems. NOTE: This EML metadata file does not contain important geospatial data processing information. Before using any NWT LTER geospatial data read the arcgis metadata XML file in either ISO or FGDC compliant format, using ArcGIS software (ArcCatalog > description), or by viewing the .xml file provided with the geospatial dataset.
GAP model parameter files for "Combining phonon accuracy with high transferability in Gaussian approximation potential models"
<p>GAP model parameter files to accompany the publication</p> <p>"Combining phonon accuracy with high transferability in Gaussian approximation potential models"</p> <p>by Janine George, Geoffroy Hautier, Albert P. Bartók, Gábor Csányi, and Volker L. Deringer</p> <p>All models are defined by the main parameter file "gp_iter6C.xml" and the associated file "gp_iter6C.xml.sparseX.GAP_00001", where "GAP_0000" is a placeholder for the unique identifier of the potential (also given in the XML header), and the trailing "1" indicates that only one set of descriptor (here, SOAP) parameters is given.</p> <p>The directories in this dataset follow the figures in the publication for which the respective potentials have been first used.</p> <p>Fig_2/only_random/M_1000<br> Fig_2/only_random/M_3000<br> Fig_2/only_random/M_5000<br> Fig_2/only_random/M_7000<br> Fig_2/only_random/M_9000</p> <p>Fig_2/only_individual/M_1000<br> Fig_2/only_individual/M_3000<br> Fig_2/only_individual/M_5000<br> Fig_2/only_individual/M_7000<br> Fig_2/only_individual/M_9000</p> <p>Fig_2/combined/M_1000<br> Fig_2/combined/M_3000<br> Fig_2/combined/M_5000<br> Fig_2/combined/M_7000<br> Fig_2/combined/M_9000</p> <p>Fig_4/only_individual/f_0.1000<br> Fig_4/only_individual/f_0.0100<br> Fig_4/only_individual/f_0.0010<br> Fig_4/only_individual/f_0.0001</p> <p>Fig_4/combined/f_0.1000<br> Fig_4/combined/f_0.0100<br> Fig_4/combined/f_0.0010<br> Fig_4/combined/f_0.0001</p> <p>Fig_5/GAP-18_plus_SCs_f0.01 <br> Fig_5/GAP-18_plus_SCs_f0.001</p> <p> </p>
Stability properties of a projector-splitting scheme for dynamical low rank approximation of random parabolic equations - Numerical tests
<p>Code used to obtain numerical results for the work titled "Stability properties of a projector-splitting scheme for dynamical low rank approximation of random parabolic equations".</p> <p>Associated publication (pre-print) is available at https://arxiv.org/abs/2006.05211</p>
Supplementary materials for "Energetics of a fluid under the Boussinesq approximation"
<p>Supplementary materials for "Energetics of a fluid under the Boussinesq approximation"</p> <p>MARUYAMA, Kiyoshi. (2017, June 13). Energetics of a fluid under the Boussinesq approximation. Zenodo. http://doi.org/10.5281/zenodo.3757240</p>
Approximation of a marine ecosystem model by artificial neural networks designed using a genetic algorithm
<p>Data from the Paper: Approximation of a marine ecosystem model by artificial neural networks designed using a genetic algorithm.</p> <p>Abstract: </p> <p>Marine ecosystem models are important to identify the processes that affects for example the global carbon cycle. Computation of an annually periodic solution (i.e., a steady annual cycle) for these models requires a high computational effort. To reduce this effort, we approximated an exemplary marine ecosystem model by different artificial neural networks. We used a fully connected network, then applied the sparse evolutionary training (SET) procedure, and finally applied a genetic algorithm (GA) to optimize both the network topology. With all three approaches, a direct approximation of the steady annual cycle was not sufficiently accurate. However, using the mass-corrected prediction of the ANN as initial concentration for additional model runs, the results were in very good agreement. In this way, we achieved a runtime reduction by about 15 \%. The result from the SET algorithm were comparable to those of the full network. Further application of the GA may lead to an even higher reduction.</p> <p>Content:</p> <p>Database sqlite <a href="https://zenodo.org/api/files/669d208b-7304-4d31-a6cc-3ad78e3544e9/ANN_Database.db">ANN_Database.db</a></p> <p>zip-files with data: </p> <p><a href="https://zenodo.org/api/files/669d208b-7304-4d31-a6cc-3ad78e3544e9/ANN-Data.zip">ANN-Data.zip</a> structure and weights of used networks</p> <p><a href="https://zenodo.org/api/files/669d208b-7304-4d31-a6cc-3ad78e3544e9/ANN-Results.zip">ANN-Results.zip</a> results obtained with networks</p> <p><a href="https://zenodo.org/api/files/669d208b-7304-4d31-a6cc-3ad78e3544e9/Reference-Results.zip">Reference-Results.zip</a> reference results and training data</p> <p> </p> <p> </p> <p> </p> <p> </p>
The limits of normal approximation for adult height - old version
<p>Supplementary tables from the article "The limits of normal approximation for adult height".</p>
Coefficients for Global Minimax Approximations and Bounds for the Gaussian Q-Function by Sums of Exponentials
<p>This is a supplementary dataset for the publication:</p> <p>I. M. Tanash and T. Riihonen, "Global Minimax Approximations and Bounds for the Gaussian Q-Function by Sums of Exponentials," in <em>IEEE Transactions on Communications</em>, vol. 68, no. 10, pp. 6514-6524, Oct. 2020, doi: 10.1109/TCOMM.2020.3006902.</p> <p>The dataset contains the sets of the optimized coefficients for the novel minimax approximations and bounds of the Gaussian Q-function, its first four integer powers and for the case of average symbol error probability (SEP) in optimal detection of 4-QAM that is actually a polynomial of the Q-function. The proposed approximations and bounds have the form of a weighted sum of exponential functions. The corresponding optimized coefficients are found up to twenty-five exponential terms with the right boundary of the finite interval on the x-axis (x_K+1) ranging from 1 to 10 in steps of 0.1 for the relative error.</p> <p>The Matlab function (func_extract_coef.m) extracts the required set of optimal coefficients from the provided dataset according to the selected error type, variation, number of terms and the right end-point in case of relative error. See help func_extract_coef for more information.</p> <p>A Matlab script (Example.m) is also provided as an example to illustrate the use of the provided Matlab function in extracting the required coefficients from the dataset, to calculate and plot the corresponding relative error which is shown by figure Example.jpg.</p> <p> </p>
Experiment Dataset for L. Zhu, G. Casale, I. Perez, Fluid approximation of closed queueing networks with discriminatory processor sharing, Performance Evaluation (2020): 102094
<p>This dataset provides the results of the validation experiments for transient, steady-state and response time distribution<br> analysis published in L. Zhu, G. Casale, I. Perez, Fluid approximation of closed queueing networks with discriminatory processor sharing, Performance Evaluation (2020): 102094.</p>
Distributed Distance Approximation
Full video presentation of the paper: Distributed Distance Approximation.<br><br>Appears in Session 4 of the 24th International Conference on Principles of Distributed Systems OPODIS 2020<br><a href="https://opodis2020.unistra.fr">https://opodis2020.unistra.fr</a>
Estimating on the fly: the approximate number system in rufous hummingbirds (Selasphorus rufus)
<p></p><p>When presented with resources that differ in quantity, many animals use a numerosity system to discriminate between them. One taxonomically widespread system is the approximate number system. This is a numerosity system that allows the rapid evaluation of the number of objects in a group and which is regulated by Weber's Law. Here we investigated whether wild, free-living rufous hummingbirds (Selasphorus rufus) possess an approximate number system. The hummingbirds were presented with two experiments. In the first we investigated whether hummingbirds spontaneously chose an array containing more flowers than an alternate array. In the second we asked whether the hummingbirds could learn to use numerosity as a cue to which of two arrays contained the better reward. The birds did not spontaneously prefer an array containing more flowers. After minimal training, however, they learned to choose the more numerous array and could differentiate between arrays of five and seven flowers. These data support the presence of an approximate number system in the rufous hummingbird. It seems plausible that having such a system would enable much more efficient foraging in this species.</p><p></p>
Measurement of the B-band Galaxy Luminosity Function with Approximate Bayesian Computation
<p>In our recent paper Tortorelli et al. 2020 (arXiv: 2001.07727), we presented the first measurement in literature of galaxy population properties using the Approximate Bayesian Computation (ABC), specifically the B-band Galaxy Luminosity function. The novel methodology we developed that combines forward-modeling and ABC offers excellent prospects to robustly measure galaxy population properties as a function of redshift in large wide-field galaxy surveys, such as LSST and NGRST. The methodology has the advantage of overcoming part of the limitations that arise in wide-field galaxy studies, such as incorrect redshift estimation or separation of galaxies into red and blue. We forward model wide-field broad-band galaxy surveys using the fast image simulator UFig. We use ABC to constrain the galaxy population model parameters of the simulations and match data from the CFHTLS. We define a number of distance metrics between the simulated and the survey data. By exploring the parameter space of the galaxy population model through ABC to find the set of parameters that minimize these distance metrics, we obtain constraints on the LFs of blue and red galaxies as a function of redshift. We compare our results to other measurements, finding good agreement at all redshifts, for both blue and red galaxies. We also compare the redshift distribution we obtain applying the same cuts as the VIPERS survey on our image simulations, finding good agreement with survey data.</p>
Data from: Reconstructing the demographic history of orang-utans using approximate Bayesian computation
Investigating how different evolutionary forces have shaped patterns of DNA variation within and among species requires detailed knowledge of their demographic history. Orang-utans, whose distribution is currently restricted to the Southeast Asian islands of Borneo (Pongo pygmaeus) and Sumatra (Pongo abelii), have likely experienced a complex demographic history, influenced by recurrent changes in climate and sea levels, volcanic activities and anthropogenic pressures. Using the most extensive sample set of wild orang-utans to date, we employed an approximate Bayesian computation (ABC) approach to test the fit of 12 different demographic scenarios to the observed patterns of variation in autosomal, X-chromosomal, mitochondrial and Y-chromosomal markers. In the best-fitting model, Sumatran orang-utans exhibit a deep split of populations north and south of Lake Toba, probably caused by multiple eruptions of the Toba volcano. In addition, we found signals for a strong decline in all Sumatran populations ~24 ka, probably associated with hunting by human colonizers. In contrast, Bornean orang-utans experienced a severe bottleneck ~135 ka, followed by a population expansion and substructuring starting ~82 ka, which we link to an expansion from a glacial refugium. Therefore, we showed that orang-utans went through drastic changes in population size and connectedness, caused by the recurrent contraction and expansion of rainforest habitat during Pleistocene glaciations, and probably also by the impact of hunting by early humans. Our findings also emphasize the fact that important aspects of the evolutionary past of species with complex demographic histories might remain obscured when applying overly simplified models.
Data from: Approximate Bayesian computation for modular inference problems with many parameters: the example of migration rates
We propose a two-step procedure for estimating multiple migration rates in an approximate Bayesian computation (ABC) framework, accounting for global nuisance parameters. The approach is not limited to migration, but generally of interest for inference problems with multiple parameters and a modular structure (e.g. independent sets of demes or loci). We condition on a known, but complex demographic model of a spatially subdivided population, motivated by the reintroduction of Alpine ibex (Capra ibex) into Switzerland. In the first step, the global parameters ancestral mutation rate and male mating skew have been estimated for the whole population in Aeschbacher et al. (Genetics 2012; 192: 1027). In the second step, we estimate in this study the migration rates independently for clusters of demes putatively connected by migration. For large clusters (many migration rates), ABC faces the problem of too many summary statistics. We therefore assess by simulation if estimation per pair of demes is a valid alternative. We find that the trade-off between reduced dimensionality for the pairwise estimation on the one hand and lower accuracy due to the assumption of pairwise independence on the other depends on the number of migration rates to be inferred: the accuracy of the pairwise approach increases with the number of parameters, relative to the joint estimation approach. To distinguish between low and zero migration, we perform ABC-type model comparison between a model with migration and one without. Applying the approach to microsatellite data from Alpine ibex, we find no evidence for substantial gene flow via migration, except for one pair of demes in one direction.
Data from: Approximate Bayesian computation analysis of EST-associated microsatellites indicates that the broadleaved evergreen tree Castanopsis sieboldii survived the Last Glacial Maximum in multiple refugia in Japan
Climatic changes have played major roles in plants' evolutionary history. Glacial oscillations have been particularly important, but some of their effects on plants' populations are poorly understood, including the numbers and locations of refugia in Asian warm temperate zones. In the present study, we investigated the demographic history of the broadleaved evergreen tree species Castanopsis sieboldii (Fagaceae) during the last glacial period in Japan. We used approximate Bayesian computation (ABC) for model comparison and parameter estimation for the demographic modelling using 27 EST associated microsatellites. We also performed the species distribution modelling (SDM). The results strongly support a demographic scenario that the Ryukyu Islands and the western parts in the main islands (Kyushu and western Shikoku) were derived from separate refugia and the eastern parts in the main islands and the Japan Sea groups were diverged from the western parts prior to the coldest stage of the Last Glacial Maximum (LGM). Our data indicate that multiple refugia survived at least one in the Ryukyu Islands, and the other three regions of the western and eastern parts and around the Japan Sea of the main islands of Japan during the LGM. The SDM analysis also suggests the potential habitats under LGM climate conditions were mainly located along the Pacific Ocean side of coastal region. Our ABC-based study helps efforts resolve the demographic history of a dominant species in warm temperate broadleaved forests during and after the last glacial period, which provides a basic model for future phylogeographical studies using this approach.
Data from ``Developments in Stochastic Coupled Cluster Theory: The initiator approximation and application to the Uniform Electron Gas''
<p>We describe further details of the Stochastic Coupled Cluster method and a diagnostic of such calculations, the shoulder height, akin to the plateau found in Full Configuration Interaction Quantum Monte Carlo. We describe an initiator modification to Stochastic Coupled Cluster Theory and show that initiator calculations can be extrapolated to the unbiased limit. We apply this method to the 3D 14-electron uniform electron gas and present complete basis set limit values of the CCSD and previously unattainable CCSDT correlation energies for up to $r_s=2$, showing a requirement to include triple excitations to accurately calculate energies at high densities.</p>
FIGURE 1. Chiridota heheva new species. Approximately 4 in Chiridota heheva, new species, from Western Atlantic deepsea cold seeps and anthropogenic habitats (Echinodermata: Holothuroidea: Apodida)
FIGURE 1. Chiridota heheva new species. Approximately 4 individuals in situ near whitish bacterial mats (?) at Florida Escarpment seep site, eastern Gulf of Mexico, 3,270 meters. Alvin Dive 1343. Approximate diameter of body 5 mm. Photo, S. Golubic.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.