Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,956
datasets available to search
ShareScore release 0.9.0
Dataset results
1,956 results for “test data”
UTCI-Test-Data
<p>Independent data for testing the accuracy of approximations to values of the Universal Thermal Climate Index (UTCI) calculated by the UTCI-Fiala model, cf. Fig. 12 in https://doi.org/10.1007/s00484-011-0454-1</p> <p>File name: UTCI-Test-Data.txt<br> Purpose: Providing independent test data for assessing the accuracy of UTCI approximations <br> This TAB-delimited file tabulates values of UTCI and the Offset (= UTCI - Ta) in °C calculated by the UTCI-Fiala model in comparison to UTCI approximated by the table look-up approach and the regression polynomial, respectively, for 1000 conditions with different input values of:<br> Ta: air temperature in°C (range: -50 °C to +42 °C)<br> Tr-Ta: difference between mean radiant temperature (Tr) and air temperature in °C (-17 °C to +50 °C)<br> va: wind speed in m/s measured 10 m above ground level (1.0 m/s to 30 m/s)<br> rH: relative humidity in % (5% to 100%)<br> pa: water vapour pressure in kPa (0 kPa to 3.4 kPa)<br> Calculated variables:<br> Offset: Offset (= UTCI - Ta) in °C calculated by the UTCI-Fiala model (range -62 to 13 °C)<br> UTCI: UTCI in °C calculated by the UTCI-Fiala model (range -111 to 43 °C)<br> UTCI_Table: UTCI in °C approximated by the look-up approach using the data table ESM4 from the above referred publication<br> UTCI_polynomial: UTCI in °C approximated by the polynomial regression function from ESM3 of the above referred publication</p> <p>Headers in line 34, first data row in line 35</p>
Performance Data of an Ice-Melting Probe from Field Tests in two Different Ice Environments
<p>This dataset was acquired at field tests of the steerable ice-melting probe "EnEx-IceMole" (Dachwald et al., 2014). A field test in summer 2014 was used to test the melting probe's system, before the probe was shipped to Antarctica, where, in international cooperation with the MIDGE project, the objective of a sampling mission in the southern hemisphere summer 2014/2015 was to return a clean englacial sample from the subglacial brine reservoir supplying the Blood Falls at Taylor Glacier (Badgeley et al., 2017, German et al., 2021).</p> <p>The standardized log-files generated by the IceMole during melting operation include more than 100 operational parameters, housekeeping information, and error states, which are reported to the base station in intervals of 4 s. Occasional packet loss in data transmission resulted in a sparse number of increased sampling intervals, which where compensated for by linear interpolation during post processing. The presented dataset is based on a subset of this data: The penetration distance is calculated based on the ice screw drive encoder signal, providing the rate of rotation, and the screw's thread pitch. The melting speed is calculated from the same data, assuming the rate of rotation to be constant over one sampling interval. The contact force is calculated from the longitudinal screw force, which es measured by strain gauges. The used heating power is calculated from binary states of all heating elements, which can only be either switched on or off. Temperatures are measured at each heating element and averaged for three zones (melting head, side-wall heaters and back-plate heaters).</p>
Testing and Demonstration Data for DataRig Software
<p>This repository holds the testing and demonstration data for <a href="https://github.com/mscaudill/datarig">DataRig</a>, an opensource software program for downloading datasets from data repositories utilizing RESTful APIs. This repository contains 5 sample datasets.</p> <p> </p> <p><strong>annotations_001.txt</strong></p> <p>This data set is a tab-separated text file containing 6 columns that start on line number 7. The column headers are; </p> <p> 'Number' 'Start Time' 'End Time' 'Time From Start' 'Channel' 'Annotation'</p> <p>There are 13 rows of data under each of these column headers representing the start and end times of annotated events from an eeg recording file in this repository called recording_001.edf. The events describe the behavior of a mouse in 5 sec increments with each behavior being one of 'exploring', 'grooming' or 'rest'.</p> <p> </p> <p><strong>recording_001.edf</strong></p> <p>A European Data Format file consisting of 4 channels of EEG data lasting approximately 1 hour. The times in the annotations_001.txt file are referenced against this file.</p> <p> </p> <p><strong>sample_arr.npy</strong></p> <p>A numpy array of shape (4, 250) with values sequentially running from 0 to 1000.</p> <p> </p> <p><strong>sample_excel.xls</strong></p> <p>An excel file with a single column of 10 numbers from 0-9 sequentially.</p> <p> </p> <p><strong>sample_text.txt</strong></p> <p>A text file with 4 rows containing 250 values per row. The values in the file run from 0 to 1000 sequentially.</p>
Archaeological shovel test data for Little Sapelo Island, Mary Hammock, Patterson Island, and Pumpkin Hammock; Spring 2007, Summer 2007 and 2008
The purpose of this study was to determine the human occupational history of Little Sapelo Island, Mary Hammock, Patterson Island, and Pumpkin Hammock in McIntosh County, GA. During the Spring of 2007 and Summers of 2007 and 2008, a systematic archaeological shovel test survey was performed at a closely-spaced interval of 20 meters. The shovel test grid was laid out using a Trimble GPS, in accordance with the Universal Trans Mercator coordinate system Zone 17, and the North American Datum of 1927 (with a typical accuracy between 1 and 6 m). Teams of two people would go to these locations, dig holes (typically 50 cm in diameter and 1 m deep), and put the soil through wire hardware cloth with a mesh size of 0.6 cm (i.e., 0.25 inch). All cultural material (except shell) were collected and put in labeled plastic bags. Only some shell samples were kept for future analysis. Ceramics were classified to determine surface treatment/ decoration, temper, and rim form, with the ultimate goal of determining their period of use. All of the data related to these shovel tests were entered into a spreadsheet, and that spreadsheet was turned into an ArcGIS point shapefile for analysis. These data are available for professional archaeologists only. Please contact the investigator by E-mail if you would like a copy of these data.
Test data for running snakePipes : ATAC-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ATAC-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for fruit fly (<strong>dm6</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Test data for running snakePipes : mRNA-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the mRNA-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for mouse (<strong>GRCm38</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Språkbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. "Korp-the corpus infrastructure of Språkbanken." <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dannélls, and Nina Tahmasebi. "Exploring the Quality of the Digital Historical Newspaper Archive KubHist." <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p> </p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>
Test Data from a Study on Latin Vocabulary Acquisition (Cicero)
<p>The dataset contains test results from an intervention study with intermediate learners in two high schools in Berlin. In total, 58 students participated in three groups (= classes). The intervention materials and tests are published as well.</p> <p>The study was the first to collect empirical data on what German students actually know about Latin vocabulary and how they handle their vocabulary knowledge. One of the main goals of the research project is to establish a broad understanding of vocabulary knowledge in Latin lessons in Germany, which aims at a versatile education of (cross-linguistically helpful) vocabulary competence.</p>
Test Data from a Study on Latin Vocabulary Acquisition (Ovid)
<p>The dataset contains test results from an intervention study with intermediate learners in two high schools in Berlin (2018-2019). In total, 60 students participated in three groups (= classes). The intervention materials and tests are published as well.</p> <p>A key question of the still ongoing research project is: How can vocabulary competence in a historical language such as Latin be acquired and deepened by using corpus-based, i.e. context-based, methods? This question is based on a broad understanding of vocabulary that refers back to theories of the mental lexicon.</p>
Tierpsy Tracker test data
<p>Test data for Tierpsy Tracker, the MultiWorm Tracker developed at André Brown's Behavioural Genomics lab, at the MRC London Institute of Medical Sciences.</p> <p><a href="https://github.com/Tierpsy/tierpsy-tracker">https://github.com/Tierpsy/tierpsy-tracker</a></p> <p>Original paper:</p> <p>Javer, A., Currie, M., Lee, C.W. <em>et al.</em> An open-source platform for analyzing and sharing worm-behavior data. <em>Nat Methods</em> <strong>15, </strong>645–646 (2018). https://doi.org/10.1038/s41592-018-0112-1</p>
Test data for `rpackageutils`
<p>Test data for use with `rpackageutils`, common utilities used in R modeling software packages available here: <a href="https://github.com/JGCRI/rpackageutils">https://github.com/JGCRI/rpackageutils</a></p>
Voice Conversion Challenge 2020 Listening Test Data
<pre>Voice conversion (VC) is a technique to transform a speaker identity included in a source speech waveform into a different one while preserving linguistic information of the source speech waveform. In 2016, we have launched the Voice Conversion Challenge (VCC) 2016 [1][2] at Interspeech 2016. The objective of the 2016 challenge was to better understand different VC techniques built on a freely-available common dataset to look at a common goal, and to share views about unsolved problems and challenges faced by the current VC techniques. The VCC 2016 focused on the most basic VC task, that is, the construction of VC models that automatically transform the voice identity of a source speaker into that of a target speaker using a parallel clean training database where source and target speakers read out the same set of utterances in a professional recording studio. 17 research groups had participated in the 2016 challenge. The challenge was successful and it established new standard evaluation methodology and protocols for bench-marking the performance of VC systems. In 2018, we have launched the second edition of VCC, the VCC 2018 [3]. In the second edition, we revised three aspects of the challenge. First, we educed the amount of speech data used for the construction of participant's VC systems to half. This is based on feedback from participants in the previous challenge and this is also essential for practical applications. Second, we introduced a more challenging task refereed to a Spoke task in addition to a similar task to the 1st edition, which we call a Hub task. In the Spoke task, participants need to build their VC systems using a non-parallel database in which source and target speakers read out different sets of utterances. We then evaluate both parallel and non-parallel voice conversion systems via the same large-scale crowdsourcing listening test. Third, we also attempted to bridge the gap between the ASV and VC communities. Since new VC systems developed for the VCC 2018 may be strong candidates for enhancing the ASVspoof 2015 database, we also asses spoofing performance of the VC systems based on anti-spoofing scores. In 2020, we launched the third edition of VCC, the VCC 2020 [4][5]. In this third edition, we constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. The dataset for intra-lingual VC consists of a smaller parallel corpus and a larger nonparallel corpus, where both of them are of the same language. The dataset for cross-lingual VC consists of a corpus of the source speakers speaking in the source language and another corpus of the target speakers speaking in the target language. As a more challenging task than the previous ones, we focused on cross-lingual VC, in which the speaker identity is transformed between two speakers uttering different languages, which requires handling completely nonparallel training over different languages. As for listening test, we subcontracted the crowd-sourced perceptual evaluation with English and Japanese listeners to Lionbridge TechnologiesInc. and Koto Ltd., respectively. Given the extremely large costs required for the perceptual evaluation, we selected 5 utterances (E30001, E30002, E30003,E30004, E30005) only from each speaker of each team. To evaluate the speaker similarity of the cross-lingual task, we used audio in both the English language and in the target speaker’s L2language as reference. For each source-target speaker pair, we selected three English recordings and two L2 language recordings as the natural reference for the converted five utterances. </pre> <p>This data repository includes the audio files used for the crowd-sourced perceptual evaluation and raw listening test scores. </p> <pre>[1] Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "The Voice Conversion Challenge 2016" in Proc. of Interspeech, San Francisco. [2] Mirjam Wester, Zhizheng Wu, Junichi Yamagishi "Analysis of the Voice Conversion Challenge 2016 Evaluation Results" in Proc. of Interspeech 2016. [3] Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, Zhenhua Ling, "The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods", Proc Speaker Odyssey 2018, June 2018. [4] Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda. "Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion" Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 80-98, DOI: 10.21437/VCC_BC.2020-14. [5] Rohan Kumar Das, Tomi Kinnunen, Wen-Chin Huang, Zhenhua Ling, Junichi Yamagishi, Yi Zhao, Xiaohai Tian, and Tomoki Toda. "Predictions of subjective ratings and spoofing assessments of voice conversion challenge 2020 submissions." Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020, 99-120, DOI: 10.21437/VCC_BC.2020-15. </pre>
1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques
<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong> University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\'c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included 'README' file contains all the instructions.</p> <p>The 'images' directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. The file names for the direct binarized output are of the format '1QIsaa_col<columnnr>.pbm', for example, '1QIsaa_col15.pbm'. And, for the cleaned version, the format is '1QIsaa_col<columnnr>_cleaned.pbm', for example, '1QIsaa_col15_cleaned.pbm'. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The 'features' directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, '1QIsaa_col15_cleaned.hinge' and '1QIsaa_col15_cleaned.adjoined'. They are also arranged in separate directories for ease of use.</p> <p>The 'plots' directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The 'README_plot' file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick's' identify' tool, the original images are in grayscale (.jpg) from Brill collection, in '8-bit Gray 256c'. These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns' images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface's degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1. </strong>L. Schomaker & M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. & Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> <br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali <m.a.dhali(at)rug.nl><br> Lambert Schomaker <l.r.b.schomaker(at)rug.nl><br> Mladen Popović <m.popovic(at)rug.nl></p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., & Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>
Listening test results for sound field synthesis localization experiment -- head movement data
<p>This data set contains recorded head movements listeners did during several localisation tasks in the context of sound field synthesis. This is an add-on to the actual localisation results provided by [1].</p> <p>[1] Wierstorf, H. (2016). Listening test results for sound field synthesis localization experiment [Data set]. Zenodo. http://doi.org/10.5281/zenodo.55439</p>
Model outputs from the study "A scalable framework for soil property mapping tested across a highly diverse tropical data-scarce region"
<p>Model outputs from the study "A scalable framework for soil property mapping tested across a highly diverse tropical data-scarce region". The study is published as open access and can be found at the following link: <a href="https://www.sciencedirect.com/science/article/pii/S2950289625000326">https://www.sciencedirect.com/science/article/pii/S2950289625000326</a></p> <p> </p> <p>The file "SWAT_USERSOIL.csv" was included to facilitate the assimilation of the soil mapping data into the Soil & Water Assessment Tool (SWAT, https://swat.tamu.edu/) for hydrological modeling. </p> <p> </p> <p>Regarding the raster files, please note:</p> <p>a) All values in these datasets have been multiplied by 10,000 to optimize file sizes.</p> <p>b) Files are named using the variable acronym, followed by the corresponding soil layer. For outputs derived from pedotransfer functions (PTFs), the PTF reference is appended after the variable acronym.</p> <p>c) Available data decrease with increasing soil layer number. This occurs because not all locations (grid cells) have the same soil depth or number of soil layers.</p> <p> </p> <p>If you have any questions about the dataset or its use, please don't hesitate to contact us.</p> <p> </p> <p> </p>
scelda test data
<p>20210323_4i4color_149x25</p> <p><strong>m1a_features.csv</strong></p> <p>No imagenet features.</p> <p>columns 0-1 spatial x, y coordinates</p> <p>columns 2-16 morphological features: area, area_bbox, area_convex, area_filled, axis_major_length, axis_minor_length, eccentricity, equivalent_diameter_area, euler_number, extent, feret_diameter_max, orientation, perimeter, perimeter_crofton, solidity</p> <p>columns 17-24 protein mean intensity: GFAP, ELAVL2, LMN1b, MBP, LMN1b_5, GFAP_5, MBP_5, ELAVL2_5</p> <p>cells were segmented using LMN1b stain using cellpose2 </p> <p><strong>m1a_celltypes.csv</strong></p> <p>cell types determined using CCA in Seurat</p>
Data set of a thermal response test on a planar trench collector
<p>This data set provides three different temperature measurements (PT100, fiber optic measurement, thermistors), as well as the determination of the volumetric water content and the bulk electrical conductivity of the subsurface. During the published period, a thermal response test was carried out. On April 28th, at 15:25, the fluid circulation began, and heat injection started at 15:38 on April 28th and continued until May 3rd. The test was conducted with a constant volume flow of 1.00 m³/h and a constant heat injection rate of 0.88 kW.</p>
Wind tunnel test data for the evaluation of the aerodynamic coefficients of an antenna mast with ancillaries.
<p>This dataset comprises measured data and results from static wind tunnel tests conducted in April 2024 at the Giovanni Solari Wind Tunnel Facility (GS-WinDyn). The tests aim to assess the drag, lift, and moment coefficients <span>of an antenna mast designed as a triangular lattice tower, equipped with both linear and discrete ancillary components.</span> The wind tunnel experiments are carried out under both smooth and turbulent flow conditions using a scaled 3D model of the antenna mast. Five ancillary configurations, based on predominant patterns observed, are tested. Drag forces, lift forces and moments are measured using two six-component force balances attached to the ends of the model, while downstream three-component velocity data is captured by a Cobra probe. For each configuration, aerodynamic coefficients are determined for angles of attack ranging from 0° to 360°, with increments of up to 10°. The dataset provides the measured data and the obtained aerodynamic coefficients and it has significant reuse potential in several applications: comparison with experimental wind tunnel data, validation of analytical and numerical CFD models with similar configurations, estimation of wind loads due to ancillary structures, and characterization of wake effects.</p>
A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena subject to imposed damage
<p>A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena (HBTA), a full-scale steel bridge subject to imposed damage, has been established. The data set includes organized dynamic response and load measurement data of the bridge under different structural state conditions, where the structural state conditions range from an undamaged (reference) state to known damage states. Furthermore, the data set includes acceleration and strain data from the response monitoring and acceleration data from the load monitoring, where a modal vibration shaker is used as an excitation source. The data is collected in one h5-file (hierarchical data format version 5) with a sampling rate of 100 Hz. Signal processing and resampling of the data has been performed according to the description provided in the references below. The data set is now published in this open-access data repository and can be accessed and downloaded freely. As such, the data set provides an important benchmark to the scientific community within bridge damage detection and SHM.</p>
Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing
<p><strong>Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing</strong></p> <p><strong>Latest version of files: repository_v2.0.zip, Behavior Data_v2.0.xlsx and MRI IDs Testing&Replication Cohort.xlsx (please ignore repository.zip)</strong></p> <p>Open data repository Knab et al. Prediction of stroke outcome in mice based on non-invasvive MRI and behavioral testing</p> <p>Open code and documentation of prediction models available via <a href="https://github.com/major-s/mouse-mcao-outcome-predictor">https://github.com/major-s/mouse-mcao-outcome-predictor</a></p> <p><strong>Content:</strong></p> <p>README.txt</p> <p>This information</p> <p><strong>dat</strong></p> <p>Contains MRI data in NIFTI format and secondary data from atlas registration. For documentation of atlas registration files see https://pubmed.ncbi.nlm.nih.gov/28829217/<br>Files used for the manuscript:<br>t2.nii: t2 weighted image acquired 24 h post stroke<br>masklesion.nii: manually delineated lesion<br>x_masklesion.nii: lesion in atlas space<br>ix_ANO.nii: Allen brain atlas in native space (i.e. matching t2.nii)<br>Lesion volume was calculated by volume of voxels unequal 0 in x_masklesion.nii<br>Overlap of regions defined by ix_ANO.nii with masklesion.nii were used for calculating percent damage in each atlas region</p> <p><strong>prediction_models</strong></p> <p>Contains separated training and test data as xlsx and csv files with lesion volumes in cubic mm of the Allen brain atlas space, percent damage per atlas region and behavioral data. The training data was used as input for training prediction models in MATLAB, the results were created using the test data.<br>The files have following sturcture:<br>Column 1: animal ID<br>Columns 2-537: MRI regions (column title corresponds to the region number as used in the Allen common coordinate framework)<br>Column 538: lesion volume<br>Column 539: initial performance (subacute deficit) = mean performance/deficit on days 2-6<br>Column 540: mean performance/deficit on days 2-6 = initial performance (subacute deficit) - this column equals column 539 but has different header which was used to train the residual from initial deficit<br>Column 541: residual performance/deficit<br>Column 542: test or training group<br>Consecutive rows contain data for each animal specified by the animal id</p> <p>The repository also contains all trained models, prediction results for the test data and tables with resulting median absolute error (MedAE) and 5th, 25th, 75th and 95 absolute error quantiles for each model.<br>The model files end with '_models.mat' and contain 50 independently trained models each. Each model version is specified by number 1-50.<br>The result files end with '_test_results.mat' or '_test_results.xlsx', files with MedAE and quantiles end with '_test_errors.xlsx' or '_test_errors.csv. The common part of filenames specifies the used paradigm<br>Folder 'subacute deficit prediction' contains:<br> - initial_performance_from_lesion_volume: prediction of subacute deficit using lesion volume<br> - initial_performance_from_segmented_mri: prediction of subacute deficit using segmented mri<br>Folder 'long-term outcome prediction' contains:<br> - lesion_volume: prediction of residual deficit using lesion volume<br> - segmented_mri: prediction of residual deficit using segmented_mri<br> - initial_performance: prediction of residual deficit using subacute deficit<br>Folder 'mri_inc_oob_imp' contains models trained using increasing number of mri segments sorted according to the out-of-bag importance. The number of used segments is given in the file name. The models, results and errors are separated in subfolders.</p> <p>Files with equal file name and different extension always contain the same data</p> <p><strong>templates</strong><br>Allen atlas, template, brain mask, hemisphere masks, tissue probability masks in NIFTI format including annotations of region IDs and parameter.m file for use in MATLAB toolbox ANTx2<br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.