Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
295
datasets available to search
ShareScore release 0.9.0
Dataset results
295 results for “approximation”
Revisiting the historical scenario of a disease dissemination using genetic data and Approximate Bayesian Computation methodology: the case of Pseudocercospora fijiensis invasion in Africa
<p class="MsoNormal"><span>The reconstruction of geographic and demographic scenarios of dissemination for invasive pathogens of crops is a key step towards improving the management of emerging infectious diseases. Nowadays, the reconstruction of biological invasions typically uses the information of both genetic and historical information to test for different hypotheses of colonization. The Approximate Bayesian Computation framework and its recent Random Forest development (ABC-RF) have been successfully used in evolutionary biology to decipher multiple histories of biological invasions. Yet, for some organisms, typically plant pathogens, historical data may not be reliable notably because of the difficulty to identify the organism and the delay between the introduction and the first mention. We investigated the history of the invasion of Africa by the fungal pathogen of banana, <em>Pseudocercospora fijiensis</em>, by testing the historical hypothesis against other plausible hypotheses. We analysed the genetic structure of eight populations from six eastern and western African countries, using 20 microsatellite markers, and tested competing scenarios of population foundation using the ABC-RF methodology. We do find evidence for an invasion front consistent with the historical hypothesis, but also for the existence of another front never mentioned in historical records. We question the historical introduction point of the disease on the continent. Crucially, our results illustrate that even if ABC-RF inferences may sometimes fail to infer a single, well-supported scenario of invasion, they can be helpful in rejecting unlikely scenarios, which can prove much useful to shed light on disease dissemination routes.</span></p>
Does Taylor approximation really works? (Portuguese version)
<p>Didactic video that shows how Taylor approximation can be used to approximate the exponential function. It is illustrated that the approximation becomes more accurate as the polynomial degree increases.</p>
Dataset: Evaluating approximate asymptotic distributions for fast neutrino flavor conversions in a periodic 1D box
<p>In this data package we include two txt files and two zip files.</p> <p>The file Gaussian.txt contains 4 columns. The first column shows the indices of the file numbers contained in the file Gaussian_100.zip. The second, third, and fourth columns list the corresponding <span class="math-tex">\(I_{\bar\nu}\)</span>, <span class="math-tex">\(F_\nu\)</span>, and <span class="math-tex">\(F_{\bar\nu}\)</span>, respectively. </p> <p>The file Maxent.txt contains 5 columns. The first column shows the indices of the file numbers contained in the file Maxent_100.zip while the second column lists the corresponding indices in the file Gaussian.txt. The third, fourth, and fifth columns list the corresponding <span class="math-tex">\(I_{\bar\nu}\)</span>, <span class="math-tex">\(F_\nu\)</span>, and <span class="math-tex">\(F_{\bar\nu}\)</span>, respectively. </p> <p>Files Gaussian_100 and Maxent_100 give the asymptotic simulated outcomes of fast neutrino flavor conversions with the initial Gaussian and maximum-entropy distributions, respectively. Each data file xxxx.dat lists four columns for <span class="math-tex">\(v_z\)</span>, <span class="math-tex">\(P_{ee}(v_z)\)</span>, <span class="math-tex">\(g_\nu(v_z)\)</span>, and <span class="math-tex">\(g_{\bar\nu}(v_z)\)</span>, respectively. We take 100 <span class="math-tex">\(v_z\)</span> grids for all data.</p>
Cophylogeny reconstruction allowing for multiple associations through approximate Bayesian computation
Open the record for dataset details and reuse information.
Revisiting the historical scenario of a disease dissemination using genetic data and Approximate Bayesian Computation methodology: the case of Pseudocercospora fijiensis invasion in Africa
Open the record for dataset details and reuse information.
Data from: Assessing uncertainties and approximations in solar heating of the climate system
Open the record for dataset details and reuse information.
Bird observations collected by volunteers along census tracks in the Parker River National Wildlife Refuge, intervals of observations are approximately twice a month.
This file contains bird observations collected by volunteers along census tracks in the Parker River National Wildlife Refuge, Massachusetts. Intervals of observations are approximately twice a month.
All simulation results, figures and code regarding the manuscript: Calibrating models of cancer invasion: parameter estimation using Approximate Bayesian Computation and gradient matching
<p>We present two different methods to estimate parameters within a partial differential equation (PDE) model of cancer invasion. The model describes the spatio-temporal evolution of three variables -- tumour cell density, extracellular matrix density and matrix degrading enzyme concentration -- in a one-dimensional tissue domain. The first method is a likelihood-free approach associated with Approximate Bayesian Computation (ABC); the second is a two-stage gradient matching method based on smoothing the data with a Generalized Additive Model (GAM) and matching gradients from the GAM to those from the model. Both methods performed well on simulated data. To increase realism, additionally we tested the gradient matching scheme with simulated measurement error and found that the ability to estimate some model parameters deteriorated rapidly as measurement error increased.</p>
List of specimens, collection numbers, localities, and GenBank accessions of sequences. The neotype of Scinax x‑signatus is underlined. New sequences produced for this study are in bold. Abbreviations are as follow. Countries: ARG = Argentina, BOL = Bolivia, BRA = Brazil, GUF = French Guiana, GUY = Guyana, MTQ = Martinique, PER = Peru, SUR = Suriname; Brazilian states: AP = Amapá, BA = Bahia, CE = Ceará, ES = Espírito Santo, MA = Maranhão, MG = Minas Gerais, PE = Pernambuco, RJ = Rio de Janeiro, RS = Rio Grande do Sul, SP = São Paulo. An asterisk (*) indicates approximate coordinates taken from Google Earth. in A neotype for Hyla x-signata Spix, 1824 (Amphibia, Anura, Hylidae)
List of specimens, collection numbers, localities, and GenBank accessions of sequences. The neotype of Scinax x‑signatus is underlined. New sequences produced for this study are in bold. Abbreviations are as follow. Countries: ARG = Argentina, BOL = Bolivia, BRA = Brazil, GUF = French Guiana, GUY = Guyana, MTQ = Martinique, PER = Peru, SUR = Suriname; Brazilian states: AP = Amapá, BA = Bahia, CE = Ceará, ES = Espírito Santo, MA = Maranhão, MG = Minas Gerais, PE = Pernambuco, RJ = Rio de Janeiro, RS = Rio Grande do Sul, SP = São Paulo. An asterisk (*) indicates approximate coordinates taken from Google Earth.
Dataset for On the Sparsity of XORs in Approximate Model Counting (SAT-20 Paper)
<p>The artifact consists of the necessary data to reproduce the results reported in the SAT-20 Paper titled "On the Sparsity of XORs in Approximate Model Counting". <br> <br> In particular, the artifact consists of the binaries, the log files generated by our computing cluster, and scripts to generate tables and the plots used in the paper. </p>
Supplement to "Estimation of the total magnetization direction of approximately spherical bodies"
<p>Supplementary data and source code to "Estimation of the total magnetization direction of approximately spherical bodies".</p> <p>Article submitted for publication to Nonlinear Processes in Geophysics.</p> <p>The original submission and open peer-review can be viewed at http://dx.doi.org/10.5194/npgd-1-1465-2014</p> <p>This is an archive of the git version controlled repository containing: the manuscript LaTeX source files; IPython notebooks (www.ipython.org) containing source code to perform the applications to synthetic and real data; the total field magnetic anomaly data used in the real data application.</p> <p>See the original repository at https://github.com/pinga-lab/Total-magnetization-of-spherical-bodies</p> <p>The source code that implements the proposed methodology is included in version 0.3 of the open-source software Fatiando a Terra (www.fatiando.org). See module fatiando.gravmag.magdir.</p>
Accurate and efficient representation of intramolecular energy in ab initio generation of crystal structures. Part I: Adaptive local approximate models
<p>The global search stage of Crystal Structure Prediction (CSP) methods requires a fine balance between accuracy and computational cost, particularly for the study of large flexible molecules. A major improvement in the accuracy and cost of the intramolecular energy function used in the CrystalPredictor II (Habgood, M., Sugden, I. J., Kazantsev, A. V., Adjiman, C. S. & Pantelides, C. C. (2015).<em> J Chem Theory Comput</em> <strong>11</strong>, 1957-1969) program is presented, where the most efficient use of computational effort is ensured via the use of adaptive Local Approximate Model (LAM) placement. The entire search space of relevant molecule’s conformations is initially evaluated using a coarse, low accuracy grid. Additional LAM points are then placed at appropriate points determined via an automated process, aiming to minimise the computational effort expended in high energy regions whilst maximising the accuracy in low energy regions. As the size, complexity, and flexibility of molecules increase, the reduction in computational cost becomes marked. This improvement is illustrated with energy calculations for benzoic acid and the ROY molecule, and a CSP study of molecule XXVI from the sixth blind test (Reilly <em>et al.</em>, (2016).<em> Acta Cryst. B, accepted</em>.), which is challenging due its size and flexibility. Its known experimental form is successfully predicted as the global minimum. The computational cost of the study is tractable without the need to make unphysical simplifying assumptions. </p>
Contents for: Influence of Ionization on the Polytropic Index of the Solar Atmosphere within Local Thermodynamic Equilibrium Approximation
<p>This is the data used for the creation of the manuscript figures. Explanation about the data is in the "README.txt" file. Version 2 includes compressed archive file containing Fortran module files with subroutines and input data files for the calculations of the study.</p>
Supplementary material for "A Constrained Spectral Approximation of Subgrid-Scale Orography on Unstructured Grids"
<p>Supplementary material for:</p> <ul> <li>Chew, R.; Dolaptchiev, S.; Wedel, M.-S.; Achatz, U. <br>A Constrained Spectral Approximation of Subgrid-Scale Orography on Unstructured Grids</li> </ul> <p>The <em>results_datasets.tar.gz</em> archive contains the simulation datasets for the results presented in Fig. 4-9, 11-16, and B1.</p> <p>The <em>Fig_10-wind_direction_study.tar.gz</em> archive contains the simulation datasets for the results presented in Fig. 10.</p> <p>Input parameters necessary to reproduce these simulation runs are included as metadata (attributes) to the datasets. Otherwise, the input parameter scripts can be found in the <code>inputs</code> subpackage of the source code.</p> <p>Furthermore, the following results can be generated by the corresponding scripts:</p> <table> <tbody> <tr> <td><strong>Figures</strong></td> <td><strong>Scripts</strong></td> </tr> <tr> <td>Fig. 1-3</td> <td><code>runs.idealised_isosceles</code></td> </tr> <tr> <td>Fig. D1</td> <td><code>runs.taper_test</code></td> </tr> <tr> <td>Fig. E1</td> <td><code>runs.idealised_delaunay</code></td> </tr> </tbody> </table> <p> </p> <p>Refer to the software repository URL for details on downloading the source code.</p> <p> </p>
Evolutionary Approximated Audio Signals
<p>This dataset provides supplementary material for publication [1] and contains ten target audio tracks from AAM [2] and ten single chords which are approximated using evolutionary algorithms after [3]. The best mixes for experiments with different configurations of audio features, distance measures, normalization, and window size are provided.</p>
Geoclaw inputs and processed outputs used in building tsunami ML surrogates for nearshore and onshore approximation
<p>This dataset contains some input files needed for GeoClaw tsunami simulation and post-processed outputs used for training the nearshore and onshore surrogates discussed in the article - Advancing nearshore and onshore tsunami hazard approximation with machine learning surrogates available as preprint (https://doi.org/10.5194/nhess-2024-72) and project repo - https://github.com/naveenragur/tsunami-surrogates</p> <p>The GeoClaw simulation requires the following files: </p> <p><strong><a href="../api/records/10817116/draft/files/geoclaw_dtopo_files.tar.gz/content" target="_blank" rel="noopener noreferrer">geoclaw_dtopo_files.tar.gz</a></strong><strong>: </strong>input bathymetry and topography elevation in asc format from multiple sources of datasets.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_his.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_his.tar.gz</a></strong>: tsunami displacement(dtopo) input files in tt3 format for historic earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeA.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeA.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeB.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeB.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p>The machine learning training and testing requires:</p> <p><strong><a href="../api/records/10817116/draft/files/procesed_tsunami_data.tar/content" target="_blank" rel="noopener noreferrer">procesed_tsunami_data.tar</a>: </strong>the post-processed outputs for the three test location i.e. (1) the waveforms( time series recorded for the water level at the offshore and the nearshore gauge locations and (2) the maximum inundation depth recorded at the fixed grids onshore.</p> <p><a href="https://zenodo.org/uploads/14902165"><strong>tsunami-surrogates-nhess-2024-72.tar.gz</strong></a><a href="https://zenodo.org/uploads/14902165"> </a>: Model code, scripts and notebooks used in the study</p>
The exact and approximate solutions of phase velocities and reflection coefficients
<p>These are the scripts of the exact and approximate solutions of phase velocities and reflection coefficients in the paper "Approximate equations of PP-, PS1- and PS2-wave reflection coefficients in fluid-filled monoclinic media".</p>
Approximate Entropy of Spiking Series Reveals Different Dynamical States in Cortical Assemblies
<p>Self-organized criticality theory proved that information transmission and computational performances of neural networks are optimal in critical state. By using recordings of the spontaneous activity originated by dissociated neuronal assemblies coupled to Micro-Electrode Arrays (MEAs), we tested this hypothesis using Approximate Entropy (ApEn) as a measure of complexity and information transfer. We analysed 60 min of electrophysiological activity of three neuronal cultures exhibiting either sub-critical, critical or super-critical behaviour. The firing patterns on each electrode was studied in terms of the inter-spike interval (ISI), whose complexity was quantified using ApEn. We assessed that in critical state the local complexity (measured in terms of ApEn) is larger than in sub- and super-critical conditions. Our estimations were stable when considering epochs as short as 5 min. These preliminary results indicate that ApEn has the potential of being a reliable and stable index to monitor local information transmission in a neuronal network during maturation. Thus, ApEn applied on ISI time series appears to be potentially useful to reflect the overall complex behaviour of the neural network, even monitoring a single specific location.</p>
tgEDMD: Approximation of the Kolmogorov Operator in Tensor Train Format
<p>Data sets required to re-produce numerical examples in</p> <p>Lücke, M. and Nüske, F. <em>tgEDMD: Approximation of the Kolmogorov Operator in Tensor Train Format</em>, arxiv 2111.09606 (2021)</p> <p><strong>Lemon Slice Example:</strong></p> <p>- Simulation_LS_Full.npy: Complete set of ten independent simulations, at time spacing 10^{-3}, each comprising 300,000 steps.</p> <p>- Simulation_LS_delta_100.npy: Downsampled set of ten independent simulations, at time spacing 10^{-1}, each comprising 3,000 steps.</p> <p><strong>Deca Alanine Example:</strong></p> <p>- Dih_Traj_*.npy: Trajectories of sixteen backbone dihedral angles for 50,000 steps each, at 10ps time spacing.</p> <p>- Dih_Jac_Traj_*.npy: Trajectories of Jacobian matrices for sixteen backbone dihedral angles. Derivatives are taken with respect to the Euclidean coordinates of 26 atoms required for the computation of the dihedrals, and evaluated for 50,000 steps each, at 10ps time spacing.</p> <p>- Timescales_MSM.npy: Implied timescales computed by MSM analysis of the same data set. Contains the first 499 timescales computed using seven different MSM lag times.</p>
Results of "Comparison of ice dynamics using full-Stokes and Blatter-Pattyn approximation: application to the Northeast Greenland Ice Stream"
<p>This archive provides ice sheet model output produced as part of the publication:</p> <p>Rückamp, M., Kleiner, T., and Humbert, A.: Comparison of ice dynamics using full-Stokes and Blatter–Pattyn approximation: application to the Northeast Greenland Ice Stream, The Cryosphere, 16, 1675–1696, https://doi.org/10.5194/tc-16-1675-2022, 2022.</p> <p>About the data:<br> The archive contains 2 zipped tar-files for the two 2 different regions 'icestream' and 'outlet'. Each tar-file contains ascii files for the different conducted scenarios. An exemplary filename looks like:<br> FSvsBPlike_icestream_res12800m_E1_SlidingM1_isFS_P1P1GLS_BCstrong.txt</p> <p>'FSvsBPlike' indicates the general study<br> 'icestream' indicates the region (either icestream or outlet)<br> 'res12800m' indicates the horizontal resolution (ranges from 12800 to 100 meter)<br> 'E1' indicates the enhancement factor used (available values E=0.1, 1, 3 and 6)<br> 'SlidingM1' indicates the employed sliding exponent (available values m=1 and 3)<br> 'isFS' indicates that the full-Stokes equation was employed (either isFS or isBPlike)<br> 'P1P1GLS' describes the employed finite element discretization (either P1P1GLS or P2P1)<br> 'BCstrong' indicates that the strong basal boundary implementation was used (either BCstrong or BCweak)<br> <br> An exemplary file structure looks as follows:<br> % Creation date: 03-May-2022<br> % Institution: AWI (Bremerhaven-Germany), BADW (Munich-Germany)<br> % Correspondence: Martin Rueckamp (martin.rueckamp@badw.de), Angelika Humbert (angelika.humbert@awi.de)<br> % Source: COMSOL version 5.6 (www.comsol.com)<br> % Citation: Rückamp et al. (2022): https://tc.copernicus.org/preprints/tc-2021-193/<br> % Projection: EPSG:3413 (proj4 string: +init=epsg:3413)<br> x(m) y(m) z_s(m) z_b(m) vx_s(m/a) vy_s(m/a) vz_s(m/a) p_s(Pa) vx_b(m/a) vy_b(m/a) vz_b(m/a) p_b(Pa)<br> 301099.83 -1256559.78 1840.83 205.40 -0.031 0.028 -5.456 -157729.374 -2.312 2.076 0.012 14684584.911<br> ...<br> <br> The fields x, y contain x- and y coordinates of the grid nodes in EPSG:3413 projection. The fields z_s and z_b are the ice surface and ice base, respectively. The fields vx, vy, and vz are the ice velocities in horizontal (x,y) and vertical (z) direction. The field p is the pressure. The subscripts _s and _b indicate the surface and ice base, respectively.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.