Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Cellpose training data and scripts from "Machine learning for histological annotation and quantification of cortical layers"
<p>This Workflow contains all the material necessary to reproduce the cells detection, thanks to the QuPath performed in the paper</p> <p> "<strong>Machine learning for histological annotation and quantification of cortical layers</strong>"</p> <p>Inside this workflow and dataset, you will find the following folders</p> <ol> <li><strong>QuPath Training Project</strong>: A QuPath 0.5.0 project containing all the manual annotations (ground truths) used to train the cellpose model, as well as the script to start the training</li> <li><strong>Training Images</strong> and <strong>Demo Images</strong>: The raw whole slide scanner images needed by the above QuPath project</li> <li><strong>Model</strong>: The fodler containing the trained cellpose model</li> <li><strong>cellpose-training Folder</strong>: The exported raw and ground truth images that the above cellpose model was trained on</li> <li><strong>Scripts</strong>: The QuPath scripts, also located in their respective QuPath projects, that were created for this whole workflow</li> <li><strong>QC</strong>: A Jupyter notebook, based on ZeroCostDL4Mic that computes quality metrics in order to assess the performance of the trained cellpose model. The folder also contains the resulting metrics.</li> </ol> <p>Installation and Use</p> <p>If you are going to use the QuPath projects, you need a local QuPath Installation https://qupath.github.io/ that is configured to run the QuPath Cellpose Extension https://github.com/BIOP/qupath-extension-cellpose as well as a working Cellpose installation https://github.com/MouseLand/cellpose</p> <p>Instructions for installation are available from the links above.</p> <p>After that, you should be able to open the QuPath project, navigate to the "Automate > Project scripts" menu and locate the script you wish to run.</p> <p><br>1. train a cell segmentation algorithm in the context of the rat brain Layer <br>Boundaries project </p> <p>2. trigger cell segmentation from a QuPath project in a semi-automated pipeline</p>
Data sets of multi-body models for conventional, articulated, and equidistant-axles trains and single span bridges
<p>Information about characteristics of multi-body models of conventional, articulated, and equidistant-axles trains.</p> <p>Information about bridge data set, used for vehicle-bridge interaction investigations.</p>
BRAIN Journal-High Performance Data mining by Genetic Neural Network-Figure 6. Training Accuracy with prograess generation
<p>There are many features will reduce the efficiency of the algorithm and its complexity.<br> Among the methods for selecting the appropriate features, the algorithm is a GA.<br> One of the important parameters for testing methods is accuracy rate on progress generation.<br> In fact accuracy is reverse error in algorithm results. As reader can compare the results of our paper<br> with another works. Figure 4 show that accuracy present for Training step. We achieve to best<br> answers of 800 generation to after generation.</p>
Figure 2.Flow Chart for Data preprocessing & Training-Comparative study of Financial Time Series Prediction by Artificial Neural Network with Gradient Descent Learning
<p>Methodology<br> This paper develops an ANN based comparative predictive model for NASDAQ stock<br> prediction. The first ANN model is developed with Multi-Layer Feed forward Network<br> Architecture & the second model is developed with Recurrent Neural Network Architecture. In this<br> paper gradient descent based back propagation learning algorithm is used for the supervised<br> learning of the predictive network.</p>
Unipen data set of on-line (vectorial) handwriting - train_r01_v07
<p>/*****************************************************************************\<br> * *<br> * *<br> * This is the first UNIPEN distribution of the iUF *<br> * *<br> * This distribution comprises NIST train_r01_v07 *<br> * *<br> * http://www.unipen.org/ *<br> * *<br> * Source code: C/Linux at *<br> * http://www.sourcefiles.org/Scientific/Other_Sciences/uptools3.tar.gz *<br> * *<br> * *<br> * The International Unipen Foundation, December 1999 *<br> * *<br> * *<br> *******************************************************************************<br> * *<br> * *<br> * DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CDROM: *<br> * *<br> * *<br> * 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH *<br> * PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL *<br> * PURPOSES. *<br> * *<br> * Copyright 1999, International Unipen Foundation - All rights reserved *<br> * *<br> * 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY *<br> * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE *<br> * DISCLAIMED. *<br> * *<br> * 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, *<br> * INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS *<br> * DATA. *<br> * *<br> * 4) THE CONDITIONS OF USE REQUIRE PROPER REFERENCE TO THIS DATABASE *<br> * AS DESCRIBED IN ACCOMPANYING DOCUMENT 'unipen-conditions-of-use.html' *<br> * *<br> \*****************************************************************************/</p> <p>Contents of the CDROM:<br> ----------------------</p> <p>1) This file, called CDROM-README<br> 2) The nist distribution, of which part of the directory tree is listed here.</p> <p>train_r01_v07<br> include<br> abm apb app atu bbd ced gmd ibm kai lou pap pri sta uqb<br> aga apc art bba cea cee hpb imp kar mot par rim syn val<br> anj apd ata bbb ceb cef hpp imt lav nic pcl scr tos<br> apa ape att bbc cec dar huj int lex not phi sie ugi</p> <p> data<br> 1a<br> aga apb art ceb gmd imp pri tos val<br> apa app cea ced ibm lou syn uqb<br> 1b 1c 1d 2 3 4 5 6 7 8</p> <p>All files on the the CDROM were tested on UNIPEN integrity using uplib.<br> The description of the contents is given below:</p> <p><br> Description of the contents:<br> ----------------------------</p> <p>For a description and examples of the UNIPEN format, see http://www.unipen.org/</p> <p>The UNIPEN files contained in this release are organized in 10 categories, listed<br> below. The number of .SEGMENTS and number of files for each category are given:</p> <p> cat nsegm nfiles<br> 1a 15953 634 isolated digits<br> 1b 28069 1423 isolated upper case<br> 1c 61351 2145 isolated lower case<br> 1d 17286 1222 isolated symbols (punctuations etc.)<br> 2 122628 2735 isolated characters, mixed case<br> 3 67352 1949 isolated characters in the context of words or texts<br> 4 0 0 isolated printed words, not mixed with digits and symbols<br> 5 0 0 isolated printed words, full character set<br> 6 75529 3298 isolated cursive or mixed-style words (without digits and symbols)<br> 7 85213 3393 isolated words, any style, full character set<br> 8 14544 4563 text: (minimally two words of) free text, full character set</p> <p>In each directory representing a category, e.g., data/1a, a number of<br> sub-directories are contained. The name of a subdirectory is a<br> three-letter word identifying the contributor of the data.</p> <p>Consider for example the UNIPEN files contributed by 'aga' of category<br> 1a (isolated digits). The files containing .SEGMENT entries are contained<br> in the 'data' directory:<br> data/1a/aga</p> <p>Most files in this distribution contain one or more .INCLUDE statements.<br> The corresponding files are found in the 'include' directory, in this case:<br> include/aga<br> Some files (such as the 'imp' contributions) use nested .INCLUDE statements.<br> The software contained in the uptools3 distribution contains code to find<br> files to be included based on an environment variable.</p> <p><br> Distribution of categories per contributor:<br> -------------------------------------------</p> <p> 1a | 1b | 1c | 1d | 2 | 3 | 6 | 7 | 8<br> --------------------------------------------------------------------------------------------------------------<br> abm | | | | | | | 628 4 | 646 4 | 7 3 |<br> aga | 405 14 | 1115 14 | 1063 14 | 221 14 | 2804 14 | | | | 605 14 |<br> anj | | | | | | | 1435 6 | 1435 6 | |<br> apa | 692 74 | 2236 247 | 7414 391 | 1953 268 | 12295 527 | 12295 527 | | | 527 527 |<br> apb | 2033 138 | 3450 466 | 8869 434 | 946 233 | 15298 590 | 15298 590 | | | 590 590 |<br> apc | | | | | | | 1724 441 | 1798 444 | 444 444 |<br> apd | | | | | | | 1958 453 | 2448 507 | 507 507 |<br> ape | | | | | | | 1384 286 | 1848 322 | 322 322 |<br> app | 1046 115 | 3010 353 |10370 556 | 2886 400 | 17312 745 | 17312 745 | | | 745 745 |<br> art | 170 6 | 1042 6 | 2301 6 | 202 6 | 3715 6 | 3715 6 | 687 6 | 933 6 | 186 6 |<br> att | | | | | | | 932 29 | 2253 29 | 819 30 |<br> atu | | | | | | | | | 92 92 |<br> bba | | | | | | | | | 63 63 |<br> bbb | | | | | | | | | 51 51 |<br> bbc | | | | | | | | | 61 61 |<br> bbd | | | | | | | | | 858 858 |<br> cea | 7 3 | 57 6 | 1402 6 | 35 6 | 1501 6 | 1501 6 | 311 6 | 345 6 | 38 6 |<br> ceb | 16 2 | 30 4 | 488 4 | 8 3 | 542 4 | 542 4 | 116 4 | 129 4 | 22 4 |<br> cec | | | | | | | 4880 35 | 5625 35 | 604 35 |<br> ced | 1369 42 | 2691 42 | 2619 43 | 1077 43 | 7756 43 | 7756 43 | | | 1100 43 |<br> cee | | | | | | | 3977 29 | 3978 29 | |<br> dar | | | | | | | 277 2 | 316 2 | 36 2 |<br> gmd | 1145 3 | | 2921 3 | 832 3 | 4898 3 | | | | |<br> hpb | | | | | | | 1524 7 | 2292 7 | 1832 23 |<br> hpp | | | | | | | 8323 32 | 10820 32 | 2591 29 |<br> huj | | | | | | | 104 1 | 104 1 | |<br> ibm | 1571 22 | 4264 22 | 4354 22 | 1994 22 | 12183 22 | | 1196 9 | 1196 9 | |<br> imp | 257 50 | 645 50 | 656 50 | 851 50 | 2409 50 | | 1119 22 | 1119 22 | |<br> imt | | | | | | | 242 1 | 242 1 | |<br> int | | | | | | | 2012 4 | 2012 4 | |<br> kai | | 1961 28 | 8663 46 | 1585 22 | 12209 57 | 8933 28 | 1013 28 | 1663 28 | |<br> kar | | | | | | | 1809 33 | 1860 33 | |<br> lav | | | 1324 9 | | 1324 9 | | 213 5 | 213 5 | |<br> lex | | | | | | | 5660 13 | 7235 13 | 1937 13 |<br> lou | 7 1 | 11 1 | 15 1 | 2 1 | 35 1 | | 1538 7 | 1599 7 | |<br> mot | | | 2701 8 | | 2701 8 | | | | |<br> nic | | | | | | | 6813 66 | 6813 66 | |<br> not | | | | | | | 1452 8 | 1452 8 | |<br> pap | | | | | | | 2203 39 | 2213 41 | |<br> par | | | | | | | 496 8 | 512 8 | |<br> pcl | | | | | | | 616 21 | 616 21 | |<br> phi | | | | | | | 2506 12 | 2506 12 | 91 4 |<br> pri | 78 15 | 212 15 | 191 15 | 230 15 | 711 15 | | 106 3 | 110 3 | 49 18 |<br> rim | | | | | | | 277 21 | 277 21 | |<br> scr | | | | | | | | | 211 44 |<br> sie | | | 377 377 | | 377 377 | | 1593 1593 | 1593 1593 | |<br> sta | | | | | | | 15808 61 | 16415 61 | 156 29 |<br> syn | 4554 17 | 637 8 | 589 8 | 415 8 | 6195 17 | | | | |<br> tos | 543 108 | 1432 108 | 1381 108 | 1660 108 | 4985 108 | | | | |<br> ugi | | | | | | | 597 3 | 597 3 | |<br> uqb | 598 4 | 1514 4 | | 1327 4 | 3439 4 | | | | |<br> val | 1462 20 | 3762 49 | 3653 44 | 1062 16 | 9939 129 | | | | |<br> --------------------------------------------------------------------------------------------------------------<br> | | | | | | | | | |<br> tot |15953 634 |28069 1423|61351 2145|17286 1222|122628 2735|67352 1949 | 75529 3298 | 85213 3393 |14544 4563|<br> --------------------------------------------------------------------------------------------------------------<br> 1a | 1b | 1c | 1d | 2 | 3 | 6 | 7 | 8</p>
SILVA_123 Eukaryota taxonomic training data formatted for DADA2; plus version with (uncurrated) additional sequences
<p>This is the Silva123 taxonomy reformatted to work with eukaryote sequences for dada2:</p> <p><a href="https://zenodo.org/api/files/4f873b5b-fe02-4344-8c8b-c723c3caba67/SILVA_123_dada2.fasta">SILVA_123_dada2.fasta </a></p> <p>It has been created using a script available <a href="https://github.com/derele/AA_Hyena/blob/master/R/convert_silva_taxonomy.r">here</a>.</p> <p>To increase coverage (Silva has a low coverage for eukaryotes). Uncurrated additional sequences similar to ASVs found in the intestine of hyenas (BLAST) have been added. The script used for taxonomic annotation of these sequences is available <a href="https://github.com/derele/AA_Hyena/blob/master/scripts/blast2alltax_outfmt11.pl">here</a>. The file containing these additional (uncurrated!) sequences is:</p> <p><a href="https://zenodo.org/api/files/4f873b5b-fe02-4344-8c8b-c723c3caba67/SILVA_123_dada2_exp.fasta?versionId=78a7526a-72e0-4707-ab5c-297671b5ff13">SILVA_123_dada2_exp.fasta </a></p> <p>This expanded file has been used for taxonomic annotation in: </p> <p><a href="https://doi.org/10.3389/fcimb.2017.00262">Heitlinger, E., Ferreira, S., Thierer, D., Hofer, H., & East, M. L. (2017). The intestinal eukaryotic and bacterial biome of spotted hyenas: the impact of social status and age on diversity and composition. <em>Frontiers in cellular and infection microbiology</em>, <em>7</em>, 262.</a></p>
Trained neural network data for synchrotron radiative transfer in the Stokes basis, power law model, computed by rimphony, for consumption by neurosynchro
<p>This archive contains data representing a trained-up neural network suitable for use with the <a href="https://github.com/pkgw/neurosynchro/">neurosynchro</a> package. The network generates coefficients that can be used for numerical radiative transfer of synchrotron emission in the Stokes basis with a package such as <a href="https://github.com/jadexter/grtrans/">grtrans</a>.</p> <p>In this particular dataset, networks were trained on a training set of coefficients generated by <a href="https://github.com/pkgw/rimphony/">rimphony</a> that is available as <a href="https://doi.org/10.5281/zenodo.1341154">DOI:10.5281/zenodo.1341154</a>. The data were generated using a model of a power law electron distribution isotropic in pitch angle. The input parameters, which were sampled randomly in a three-dimensional space, were:</p> <ul> <li><em>s</em>, the harmonic number, dimensionless, sampled logarithmically between 5 and 50,000,000.</li> <li><em>theta</em>, the angle between the ray path and the local magnetic field, measured in radians, sampled linearly between 0.001 and π/2 (namely, 1.5707963267948966).</li> <li><em>p</em>, the power-law index of the energetic electrons, dimensionless, sampled linearly between 1.5 and 7.</li> </ul> <p>The training set was computed on Harvard’s Odyssey cluster using Git commit <a href="https://github.com/pkgw/rimphony/commit/772161ebda0217b8c1ccb8ce3801ad9dc3701a4f">772161</a> of rimphony. A total of about 5,000 CPU hours were used, with 500 processes running for about 10 hours each, yielding about 22 million numbers. Training the networks took about 3 hours on an 8-core laptop.</p> <p>For the purposes of <em>neurosynchro</em>, the formats of the files in this package should be regarded as internal implementation details. The <a href="https://pypi.org/project/neurosynchro/">neurosynchro</a> Python package will load up the files in this archive and use them to predict synchrotron coefficients. For specifics, see <a href="https://neurosynchro.readthedocs.io/en/stable/">the neurosynchro documentation</a>.</p>
Data for Galaxy CLIP-Seq Training Material
<p>The eCLIP data provided here is a subset of the eCLIP data of RBFOX2 from a study published by Nostrand et al. (2016, http://dx.doi.org/10.1038/nmeth.3810). The dataset contains the first biological replicate of RBFOX2 CLIP-seq and the input control experiment (*fastq files). The data was changed and downsampled to reduce data processing time, thus the datasets does not correspond to the original data pulled from Nostrand et al. (2016, http://dx.doi.org/10.1038/nmeth.3810). Also included is a text file (.txt) encompassing the chromosome sizes of hg19 and hg38 obtained from UCSC (http://hgdownload.cse.ucsc.edu/goldenPath/hg19/bigZips/hg19.chrom.sizes, http://hgdownload.cse.ucsc.edu/goldenPath/hg38/bigZips/hg38.chrom.sizes) and a genome annotation for hg19 (.gtf) taken from Ensembl (http://ftp.ensemblorg.ebi.ac.uk/pub/release-74/gtf/homo_sapiens/) and for hg38 taken from the Galaxy libraries (https://usegalaxy.eu/library/list#folders/F30cab321d898d2fb/datasets/9ba790aa79c9cf23). The data is used for a galaxy training course about CLIP-Seq data analysis. </p>
Training data for 'Somatic variant calling' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates identification of somatic and germline variants from tumor and normal sample pairs.</p>
Data - The Benefits of Neurofeedback Training for Alpha Enhancement and Cognitive Performance - a Single-Blind, Sham-Feedback Study Using a Low-Prized EEG Device
<p>This data set includes the minimal data set, which was used to obtain the results in Naas, Rodrigues, Knirsch, & Sonderegger (2019, doi: http://dx.doi.org/10.1101/527598).</p>
Training data for 'Genome annotation with Apollo' tutorial (Galaxy Training Material)
<p>Published scaffolds from the Apis mellifera assembly Amel_4.5 and Official Gene Set 3.2.</p> <p>Source: <a href="http://hymenopteragenome.org/beebase/?q=download_sequences">http://hymenopteragenome.org/beebase/?q=download_sequences</a></p>
Logistics Transport Label Data - 'Lean Training Data Generation for Planar Object Detection Models in Unsteady Logistics Contexts'
<p>Example dataset described in ICMLA2019 Paper 'Lean Training Data Generation for Planar Object Detection Models in Unsteady Logistics Contexts' (Dörr, Brandt, Meyer, Pouls).</p>
Neural-network-based molecular dynamics simulations reveal that proton transport in water is doubly gated by sequential hydrogen-bond exchange: Neural network potentials training data
<h1>Neural network potentials of an excess proton in bulk water, training data</h1> <p>This dataset contains 2188 configurations labeled at two hybrid DFT levels (revPBE0-D3 and B3LYP-D3).</p> <p>The configurations are given as a single XYZ file: configurations.xyz</p> <p>The box dimensions are written in box.txt</p> <p>The energies for all configurations at a given level of theory are written in energies_LEVEL.txt (one configuration per line)</p> <p>The atomic forces for each configuration at a given level of theory are gathered in a XYZ file: forces_LEVEL.xyz</p> <p>The relative displacements of the Wannier centroids, with respect to the closest oxygen atom, for each configuration at a given level of theory, are in the following XYZ file: wannier-centroids-displacements_LEVEL.xyz</p>
Training data set for: Graph Neural Network based elastic deformation emulators for magmatic reservoirs of complex geometries
<h2>Overview</h2> <p>This is a synthetic volcano deformation dataset accompanying the publication of <em><strong>Graph Neural Network based elastic deformation emulators for magmatic reservoirs of complex geometries</strong></em>,<em><strong> </strong></em>on the journal <em>Volcanica</em>. Synthetic, quasi-static deformation is computed for magma chambers of various geometries, parameterized as spheroids or superpositions of spherical harmonics. Surface deformation is computed using the boundary element method (BEM) of Nikkhoo & Walter (2015). Please reference our paper for details of computational methods.</p> <p>The dataset contains 50,000 realizations of magma chamber geometries/orientations/centroid depths and associated deformation fields. Surface deformation fields are sampled at discrete locations, with a uniform random distribution within [Lh x Lh], and a distribution that concentrates near the chamber (at radial distances, r = 10^(-3 <em> random number) * </em>Lh/2). Note this dataset contains only a small fraction of the total dataset. In total, 824,393 realizations of magma chambers were used to train our emulators. For accessing the complete training data set, please contact the authors. </p> <p>Each .mat file contains the deformation field associated with a single chamber geometry. Use visData.m to visualize chamber geometry and associated surface displacement. Each file contains two MATLAB structures, "input" and "output". </p> <h2>Naming of each zip file</h2> <p>The numbers after the underscore, N:M, indicate that this file contains N of the M total chamber realizations for this particular setup. </p> <p><a href="../api/records/13800065/draft/files/sph_20AspRatios_1e4:151211.zip.zip/content" target="_blank" rel="noopener noreferrer">sph_20AspRatios_1e4:151211.zip</a>: deformation corresponding to spheroidal magma chambers parameterized by aspect ratios. </p> <p><a href="../api/records/13800065/draft/files/sh_complex_1e4:152283.zip/content" target="_blank" rel="noopener noreferrer">sh_complex_1e4:152283.zip</a>: deformation corresponding to chamber geometry produced by superposition of spherical harmonic modes. </p> <p><a href="../api/records/13800065/draft/files/sh_mode_approx_1e4:138380.zip/content" target="_blank" rel="noopener noreferrer">sh_mode_approx_1e4:138380.zip</a>: deformation corresponding to chamber geometries corresponding to individual spherical harmonic modes, combined with a spherical mode (the spherical mode prevents chamber surfaces from having zero radii locally)</p> <p><a href="../api/records/13800065/draft/files/sh_spheroid_approx1e4:202272.zip/content" target="_blank" rel="noopener noreferrer">sh_spheroid_approx1e4:202272.zip</a>: deformation corresponding to chambers approximating spheroids, but parameterized by spherical harmonics.</p> <p><a href="../api/records/13800065/draft/files/sh_spheroid_perturb_1e4:180247.zip/content" target="_blank" rel="noopener noreferrer">sh_spheroid_perturb_1e4:180247.zip</a>: same as above, but with additional random perturbations parameterized in spherical harmonics.</p> <h2>Variables in each file</h2> <p><strong>Input</strong> contains the following fields:</p> <p><strong>dp2mu</strong>: pressure change to shear modulus ratio.</p> <p><strong>dx</strong>, <strong>dy</strong>, <strong>dz</strong>: the coordinates of chamber centroid [meters]</p> <p><strong>mu: </strong>dimensionless crustal shear modulus (always set to 1)</p> <p><strong>nu</strong>: crustal Poisson's ratio (always set to 0.25)</p> <p><strong>Ns</strong>: number of points on the surface where displacements are computed</p> <p><strong>Lh</strong>, <strong>Lv</strong>: horizontal and vertical dimensions of the model domain [meters]. Lh is determined such that at the edge of the model domain, the displacement magnitude is below 10 percent of the maximum. Lv = Lh/2 + abs(dz)</p> <p>for the spheroids -----------------------------------------------------------------------------------------------------------</p> <p>the input files contain</p> <p><strong>asp</strong>: aspect ratio of chamber (length of the semi-major axis divided by that of the semi-minor axis)</p> <p><strong>ra</strong>, <strong>rb</strong>: semi-major, -minor, axis length [meters]</p> <p><strong>thetax</strong>, <strong>thetay</strong>, <strong>thetaz</strong>: counterclockwise rotation angles with regard to x, y, z axis [degrees]. thetax = [0, 90] degrees, thetay = 0 degrees, thetaz = 360 degrees.</p> <p>for the general geometries--------------------------------------------------------------------------------------------------</p> <p>the input files contain</p> <p><strong>ls</strong>, <strong>ms</strong>, <strong>fs</strong>: degree, order, coefficients of spherical harmonic modes. Spherical harmonics are sampled up to degree 5. fs is a complex vector of coefficients such that the resulting shape is real. </p> <p><strong>normF</strong>: normalization factor applied to the shape parameterized by ls, ms, fs, such that the shape as a maximum radius of unity.</p> <p><strong>rmax</strong>: scale factor to scale the spherical harmonics parameterized shape to real dimensions [meters].</p> <p>=============================================================================================</p> <p>Output contains the following fields,</p> <p><strong>X</strong>, <strong>Y</strong>, <strong>Z</strong>: coordinates of points where displacement vectors are computed [meters]</p> <p><strong>Ux</strong>, <strong>Uy</strong>, <strong>Uz</strong>: displacements in x, y, z directions [meters]</p> <p><strong>P</strong>, <strong>T</strong>: coordinates [meters] of vertices for the triangular mesh used in BEM calculation, and the connectivity matrix </p> <p><strong>C</strong>: coordinates [meters] of the center of each triangular element</p> <p><strong>that</strong>, <strong>dhat</strong>, <strong>nhat</strong>: unit vectors for orthogonal coordinate systems local to each triangular element. that ("t-hat") extends from vertex one to vertex two, nhat is outward normal, and dhat = cross (nhat, that).</p> <p>Reference:</p> <p>1. Nikkhoo, M., & Walter, T. R. (2015). Triangular dislocation: an analytical, artefact-free solution. <em>Geophysical Journal International</em>, <em>201</em>(2), 1119-1141.</p>
Training data for ARPESNet
<p><strong>Datasets used for training the ARPESNet autoencoder.</strong></p> <p>Datasets:</p> <p>The training and test data are stored in the .zip files. Unzipping these will provide the files created by saving PyTorch tensors. These can be loaded with the python code:</p> <pre><code>import torch data = torch.load("filename.pt")</code></pre> <p> where filename can be any of the files in this repository.</p> <p><strong>train_data.zip & test_data.zip: </strong>training and testing data, respectively. Collection of 256x256 ARPES spectra (images) obtained by randomly slicing 3D of 28 (train) and 18 (test) high resolution angle scans, covering 19 material systems: Au(110), Au(111), Bi2Se3(111), CoO2, on Au(111), CrSBr(001), Graphene on Ir(111), Graphene on Ru(0001), single-layer MoS2 on Au(111), single-layer NbSe2 on bilayer graphene, NdTe3(010), Pd(100), Pd(111), Pt(111), Rb-doped Bi2Se3(111), Ru(0001), P-δ-layer on Si(001), single-layer TaS2 on Au(111), single-layer WS2 on Ag(111) and single-layer WS2on Au(111).</p> <p>Each file corresponds to one material system, for which 500 images were generated, resulting in a tensor of shape 500x256x256 each.</p> <p><strong>test_imgs.pt: </strong>6 ARPES spectra used for visual inspection and performance test of the ARPESNet autoencoder.</p> <p><strong>cluster_centers.pt: </strong>ARPES spectra extracted by slicing an angle scan obtained measuring a Bi2Se3 crystal. These are used to generate simulated nanoARPES maps for testing clustering performance.</p> <p>dataset_info.csv: tabluar data describing the single datasets, their use in test or training and appropriate citation to the source publication wherre the data was first published.</p> <p>This repository contains the data related to the publication</p> <p>Steinn Ýmir Ágústsson, Mohammad Ahsanul Haque, Thi Tam Truong, Marco Bianchi, Nikita Klyuchnikov, Davide Mottin, Panagiotis Karras, Philip Hofmann; <strong>An autoencoder for compressing angle-resolved photoemission spectroscopy data</strong>. <em>Mach. Learn.: Sci. Technol.</em> <strong>6</strong> 015019 (2025) DOI: <a href="https://doi.org/10.1088/2632-2153/ada8f2" target="_blank" rel="noopener">10.1088/2632-2153/ada8f2</a></p> <p>Please cite the paper above in case of re-use of these data in a scientific publication.</p>
Training data for "PepINVENT: Generative peptide design beyond the natural amino acids"
<p>The zipped file contains the training and the validation data used to train the PepINVENT model.</p>
Data for Manuscript: Instrumental Validity of the Motion Detection Accuracy of a Smartphone Based Training Game
<p><strong>Background: </strong>In the project TRIMOTEP we developed a low-cost augmented reality training game. Aim of the training game ist to support patients after total hip replacement in their rehabilitation. The project was funded by the Austrian Research Promotion Agency (FFG, grant number 862050). As hardware the training game uses a headset, an android smartphone and a step board. The goal of the training game is to dodge animals and objects while performing exercises. A current version of the training game can be downloaded here: https://trimotep.fh-joanneum.at/exer-game-ar_walker/ . The training game is based on Google ARCore and uses a movement detection approach to recognise different exercises. To detect movements ARCore uses the smartphone inbuilt inertial measurement unit and the front camera (https://developers.google.com/ar/discover). In order to investigate the possibilities of the training game, it is necessary to examine the accuracy of movement detection in more detail.</p> <p><strong>Data: </strong>To investigate the accuracy, comparative measurements were carried out with 30 healthy subjects. During the measurements, the subjects motion was recorded simultaneously with the training game and an optoelectronic motion capture system (Vicon). Two trials were recorded with each subject.</p> <p>First Trial: subjects followed a protocol</p> <p>Second Trial: subjects played the training game for one minute</p> <p>The training game measures the movement of the smartphone (and therefore of the headset and the head). The optoelectronic motion capture system uses a marker set consisting of four markers. Those markers are labeled HMD_F, HMD_B, HMD_R, HMD_L. Markers HMD_R and HMD_L as well as HMD_B and HMD_F form an axis in a karthesian coordinate system. This coordinate system is rotated by 8 degrees compared to the training game along the transversal axis.</p> <p><strong>Structure of the Data Set:</strong> The data set includes an excel sheet with general data of the subjects and a figure showing the tilt between the two coordinate systems. Further one folder contains the measurement data of the training game as json files. Another folder contains the measurement data of the optoelectronic motion capturing system as csv files.</p> <p> </p> <p>For further information or help to process the data please contact:</p> <p>Bernhard Guggenberger, bernhard.guggenberger2@fh-joanneum.at</p>
Accurate annotation of protein coding sequences with IDTAXA - Training Data
<p>Training data used to test IDTAXA, HMMER, and BLAST performance of classification of amino acid and nucleotide sequences.</p>
Training and test data set of thickness distribution of turbidites
<p>This is a data set of thickness distribution used for training and test of the inverse model of turbidites. Details were described in https://esurf.copernicus.org/preprints/esurf-2020-93/esurf-2020-93.pdf</p>
Training data for the component contribution method
<p>Data files containing measured Gibbs free energies of reaction and formation (or electrostatic potential). This data is used by the component-contribution method to train the regression model that is used to estimate Gibbs free energies of many other biochemical reactions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.