Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
111
datasets available to search
ShareScore release 0.9.0
Dataset results
111 results for “test collection”
1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques
<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong> University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\'c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included 'README' file contains all the instructions.</p> <p>The 'images' directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. The file names for the direct binarized output are of the format '1QIsaa_col<columnnr>.pbm', for example, '1QIsaa_col15.pbm'. And, for the cleaned version, the format is '1QIsaa_col<columnnr>_cleaned.pbm', for example, '1QIsaa_col15_cleaned.pbm'. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The 'features' directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, '1QIsaa_col15_cleaned.hinge' and '1QIsaa_col15_cleaned.adjoined'. They are also arranged in separate directories for ease of use.</p> <p>The 'plots' directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The 'README_plot' file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick's' identify' tool, the original images are in grayscale (.jpg) from Brill collection, in '8-bit Gray 256c'. These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns' images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface's degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1. </strong>L. Schomaker & M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. & Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> <br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali <m.a.dhali(at)rug.nl><br> Lambert Schomaker <l.r.b.schomaker(at)rug.nl><br> Mladen Popović <m.popovic(at)rug.nl></p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., & Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>
Resource status collected during AWOPS tests
<p>The dataset is about resource status collected during the validation of the Automated Work Planning Services (AWOPS). The dataset includes:</p> <ul> <li>5 csv files regarding crews composition involved at JEA renovation works from March 7<sup>th</sup> to April 11<sup>th</sup>;</li> <li>9 json files regarding crews efforts in the same period.</li> </ul>
Building status collected during AWOPS tests
<p>The dataset is about building status collected during the validation of the Automated Work Planning Services (AWOPS). The dataset includes:</p> <ul> <li>1 csv file regarding renovation works' activities carried out from March 7<sup>th</sup> to April 11<sup>th</sup>;</li> <li>9 json files regarding progress data collected during the following days: <ul> <li>March 14<sup>th</sup> 2022 (day 2);</li> <li>March 17<sup>th</sup> 2022 (day 5);</li> <li>March 21<sup>st</sup> 2022 (day 9);</li> <li>March 24<sup>th</sup> 2022 (day 12);</li> <li>March 28<sup>th</sup> 2022 (day 16);</li> <li>March 31<sup>st</sup> 2022 (day 19);</li> <li>April 4<sup>th</sup> 2022 (day 23);</li> <li>April 7<sup>th</sup> 2022 (day 26);</li> <li>April 11<sup>th</sup> 2022(day 30).</li> </ul> </li> </ul>
Bangla Information Retrieval Test Collection | Revisiting Anwesha
<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, no Gold standard dataset existed for Bangla IR until recently (https://zenodo.org/record/6583149). Our work expands the existing Gold standard dataset by creating 100 query document relevance pairs across a new test collection of 1000 documents. The corpus contains news articles from Ebela, Zee News and Anandabazar Patrika, Vikaspedia and various Bangla travel blogs. The definition of the complexity level of a query is described below:</p> <p>Complexity Level 1: The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2: The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3: The query is a generalised phrase capturing the overall story or the document’s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>
ERT data collected at the Corona volcano (Lanzarote, Canary Islands) during the European Space Agency (ESA) testing campaign PANGAEA-X 2017
<p>This dataset contains the ERT (Electrical Resistivity Tomography) data collected between 22 and 23 November 2017 at the Corona volcano (Lanzarote, Canary Islands, Fig. 1) for the detection of lava tubes and the stratigraphic investigation of planetary volcanic analogues. This geophysical survey was carried out within the European Space Agency (ESA) testing campaign PANGAEA-X 2017 (Bessone et al., 2018), aimed at integrating astronaut training-data collection, documentation, analogue field geology procedures with remote sensing and in situ geophysical methods. </p> <p>Two ERT profiles were acquired in NE-SW and NNE-SSW orientations (Fig. 1). These were located roughly orthogonal to the Corona lava tube system and as far as possible on top of the main lava tube axes. The longer profile, profile D, is 470 m in length and was obtained using 48 electrodes spaced 10 m apart. The profile orientation is from SW to NE (electrode 1 to 48). The profile was acquired to detect lava tubes in test site D (sub-area south) where the exact location of a lava tube was known thanks to a LiDAR TLS (Terrestrial Laser Scan) subsurface survey (Santagata et al., 2018). A shorter profile, profile E, is 235 m long and was obtained using 48 electrodes 5 m apart. The profile orientation is from SSW to NNE (electrode 1 to 48). This profile was acquired in test site E (sub-area north) to provide a more detailed investigation of the potential existence of inaccessible sections of the tube whose location could be indicated by the evidence of closely-spaced aligned collapse structures.</p> <p>Each profile was collected using measure sequences compounded by 276 Wenner-Schlumberger array quadrupoles which ensure high vertical resolution and signal amplitude and 328 dipole-dipole array quadrupoles which provide enhanced lateral resolution. A fully automatic multi-electrode resistivity meter SYSCAL Jr Switch-48 by IRIS Instruments (400 V max output voltage, 1200 mA max output current, 100 W max output power, <a href="http://www.iris-instruments.com/syscal-juniorsw.html">http://www.iris-instruments.com/syscal-juniorsw.html</a>), was used for data collection.</p> <p>At most of the measurement points, it was necessary to drill the basalt using a hand drilling machine in order to place the tips of the electrodes into the ground at a depth of approximately 40 cm. The electrodes also needed to kept moist to reduce contact resistance between the electrode and the ground. A large amount of water (up to 2 liters per point) was needed for profile D, situated in an area above the lava tubes with very porous dry soil cover.</p> <p>The dataset is presented as a spreadsheet format which has the "space" as separator and the ".txt" extension. The structure of such a file is the following one:</p> <p>#, El array, Spa1/4, Rho, Dev, M, Sp, Vp, In, Time, Spa5/12, M1/20</p> <p>- #: Data point number</p> <p>- El array: Electrode array</p> <p>- Spa. 1/4: four spacing parameters (corresponding to the electrode array – in m)</p> <p>- Rho: resistivity value (in Ohm.m)</p> <p>- Dev: standard deviation (quality factor, in %)</p> <p>- M: global chargeability value (induced polarization parameter (in mV/V – "=0" if only-resistivity data))</p> <p>- Sp: spontaneous polarization (measured just before the injection, in mV)</p> <p>- Vp: measured primary voltage (in mV)</p> <p>- In: injected current intensity (in mA)</p> <p>- Time: injection time (pulse duration, in s)</p> <p>- Spa. 5/8: other spacing parameters (in m)</p> <p>- Spa. 9/12: electrode elevation (in m)</p> <p>- M1/M20: partial chargeability values (induced polarization window (in mV/V – "=0" if only-resistivity data))</p> <p> </p> <p>Acknowledgements</p> <p>The authors are grateful to ESA and all PANGAEA-X 2017 staff, particularly Loredana Bessone, Matthias Maurer, Herve Stevenin and Igor Drozdovskiy for their participation in data collection during some of the experiments and to the MilesBeyond Team, particularly Francesco Maria Sauro for his logistical support. Regional and local remote sensing data were obtained by the Spanish Instituto Geográfico Nacional (https://www.ign.es) and Gobierno de Canarias (https://www.grafcan.es, <a href="https://opendata.sitcan.es/">https://opendata.sitcan.es</a>).</p> <p> </p> <p>References</p> <p>Bessone, L., et al., 2018, Testing technologies and operational concepts for field geology exploration of the Moon and beyond: the ESA PANGAEA-X campaign, Geophysical Research Abstract, #EGU2018-4013.</p> <p>Santagata, T., Sauro, F., Massironi, M., Pozzobon, R., Del Vecchio, U., Lazzaroni, M., Damiano, N., Tonello, M., Tomasi, I., Martínez-Frìas, J. and Mateo Medero, E., 2018. Subsurface laser scanning and photogrammetry in the Corona Lava Tube System, Lanzarote, Spain, EGU General Assembly 2018, pp. EGU2018-5290.</p>
Test Collection Reliability: A Study of Bias and Robustness to Statistical Assumptions via Stochastic Simulation
<p>This archive contains the simulated collections, their diagnosis data, and the estimates of accuracy. For the full code and description, please refer to https://github.com/julian-urbano/irj2015-reliability</p>
Toward Estimating the Rank Correlation between the Test Collection Results and the True System Performance
<p>This archive contains the simulated collections and the estimated correlation coefficients. For the full code and description, please refer to https://github.com/julian-urbano/sigir2016-correlation</p>
Bangla Information Retrieval Test Collection
<p>There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, there is no Gold standard dataset available to test the effectiveness of Bangla IR. So, we have created a document collection containing 182 short stories, novels, and essays written by Rabindranath Tagore11 and 1000 newspaper articles published in 2013 crawled from the Bangla newspaper Prothom Alo12. The collection contains 100 newspaper articles each from one of the ten categories: বাংলাদেশ/ Bānlādēśa(EN: `Bangladesh'), খেলা/ khēlā(EN: `sports'), বিজ্ঞান ও প্রযুক্তি/ bijñāna ō prayukti(EN: `technology'), বিনোদন/ binōdana(EN: `entertainment'), আন্তর্জাতিক/ āntarjātika(EN: `international'), অর্থনীতি/ arthanīti(EN: `economy'), জীবনযাপন/ jībanayāpana(EN: `life-style'), মতামত/ matāmata(EN: `opinion'), শিক্ষা/ śikṣā(EN: `education') and আমরা/ āmarā(EN:`we-are'). There are 94 queries in the dataset, 26 queries belonging to complexity levels 1 and 2, 19 queries in complexity level 3 and 23 queries in complexity level 4. The definition of the complexity level of a query is described below:</p> <p>Complexity Level 1: The query contains exact words, phrases or sentence from the document.</p> <p>Complexity Level 2: The query is not present as it is in the document. There is a slight deviation.</p> <p>Complexity Level 3: The query is a generalised phrase capturing the overall story or the document’s theme.</p> <p>Complexity Level 4: It is a general query not related to any specific document.</p>
Experimental data collected during first-phase testing of the Ground CO2 Mapper
<p>The various Excel files included in this dataset report data from tests performed to assess the technical capabilities of the Ground CO2 Mapper, a newly developed tool that can be used to help reduce uncertainty in the mapping of geological or anthropogenic CO2 leakage from the ground surface. These files include data from a number of laboratory experiments as well as tests performed at a controlled release site and a natural site where geological CO2 is released over a large area. This dataset was used to create the various figures presented in the article "Development and testing of a rapid, sensitive, high-resolution tool to improve mapping of CO<sub>2</sub> leakage at the ground surface" by Graziani, Beaubien, Ciotoli and Bigi to be published in Applied Geochemistry.</p>
Fig. 2 in Application of a universal parasite diagnostic test to biological specimens collected from animals
Fig. 2. Cluster dendrogram showing parasite species detected in mammalian hosts. Sequences detected in each specimen using nUPDx were clustered alongside parasite-derived reference sequences of known identity obtained from GenBank. These reference sequences are labelled on the dendrogram branch tips (where appropriate). A peripheral color-coded heat map ring indicates the host animal from which the parasite-derived sequence was detected. Gray branches and blocks on the heat map reflect the position of reference sequences within the tree. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
Fig. 1 in Application of a universal parasite diagnostic test to biological specimens collected from animals
Fig. 1. Schematic describing the nested UPDx protocol employed in this study. This schematic provides a summary of the nested UPDx (nUPDx) protocol originally described by Flaherty et al. (Flaherty et al., 2021). Briefly, the DNA extract is subjected to a restriction digestion using the PstI restriction enzyme, and the digest product is subjected to PCR1 using primers 5'TTGATCCTGC- CAGTAGTCATATGC'3 (outer forward) and 5'GGTGTGTA- CAAAGGGCAGGGAC'3 (outer reverse). The resultant ~2 kb amplicon is digested using the restriction enzymes BamHI and BsoBI. The digest product is then subjected to PCR2 using internal primers 5'CCGGAGAGGGAGCCTGAGA'3 (inner forward) and 5'GAGCTGGAATTACCGCGG'3 (inner reverse) originally described by Flaherty et al. (Flaherty et al., 2018, 2021). The amplicon of PCR2 (~200 base pairs) is finally subjected to Illumina amplicon sequencing.
Fig. 3 in Application of a universal parasite diagnostic test to biological specimens collected from animals
Fig. 3. Cluster dendrogram showing parasite species detected in avian and reptilian hosts. Sequences detected in each specimen using nUPDx were clustered in this dendrogram alongside parasite-derived reference sequences of known identity obtained from GenBank. These reference sequences are labelled on the dendrogram branch tips (where appropriate). A peripheral color-coded heat map ring indicates the host animal from which the parasite-derived sequence was detected. Gray branches and blocks on the heat map reflect the position of reference sequences within the tree. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)
How to collect dried blood spot samples for hepatitis C testing (Spanish: Cómo obtener muestras de gota de sangre seca para el cribado de la hepatitis C)
<p>Short video in Spanish describing how to collect, store and ship to the laboratory, dried blood spot (DBS) samples for hepatitis C virus (HCV) testing.</p> <p>With proper training of the staff involved, DBS samples can be collected outside the healthcare setting, thus facilitating access to diagnosis of hepatitis C by the most vulnerable populations who attend different centers in the community.<br> With our experience in detecting HCV RNA from DBS samples since 2015, we have produced an explanatory video and a booklet with step-by-step instructions on how to obtain good quality DBS samples for laboratory HCV testing.<br> DBS samples not only allow us to improve the diagnosis rate of viremic HCV infection, but also to monitor the elimination of hepatitis C. Within the following website you can see the publications of different studies that we have carried out using DBS, which have also helped us to characterize the HCV epidemic at the local level.</p> <p>https://www.researchgate.net/project/Development-and-assessment-of-alternative-testing-strategies-for-the-detection-of-active-hepatitis-C-virus-infection-among-vulnerable-groups-at-risk-and-micro-elimination/update/61011c48647f3906fc8c31b8</p> <p> </p>
Data from: Use of the lung flute ECO to assist in sputum collection for tuberculosis testing: a randomized crossover trial
<p>The Lung Flute ECO, a self-powered, low cost, oscillatory positive expiratory pressure (OPEP) device, assisted people with presumptive tuberculosis to produce an adequate sputum volume for diagnostic testing and was well-tolerated.</p>
Oscillatory Flow Testing Data Collected at Field Site for Research in Fractured Sedimentary Rock (FSR)^2
<p>This dataset contains raw and processed pressure data collected in 2019 during oscillatory flow testing experiments at the Field Site for Research in Fractured Sedimentary Rock (FSR)^2 near Madison, WI. The included ReadMe file describes the data and code used in data processing. The companion processing and analysis code is included as a separate upload (doi:10.5281/zenodo.6584777)</p>
3D cranial landmark coordinates from the Athens Collection. A modern Greek population reference dataset for testing 3D-ID software
<p>The present dataset comprises the landmark coordinates of 158 intact crania (80 males and 78 females) of adult individuals from the Athens Collection. The 3D coordinates of up to 34 landmarks have been extracted from high quality textured 3D models produced with photogrammetry. The dataset aims to evaluate the correct classification performance of 3D-ID software. Hence, the dataset contains the landmark 3D coordinates both in Meshlab's PickedPoints files (.pp) but also in 3D-ID's input text format (.3did). The dataset is accompanied by certain GNU Octave scripts and functions used for data conversion and integrity check. For more details see the Dataset Description pdf.</p>
Assessing the prevalence of Female Genital Schistosomiasis and comparing the acceptability and performance of health worker-collected and self-collected cervical-vaginal swabs using PCR testing among women in North-Western Tanzania: the ShWAB study
<p>Female genital schistosomiasis (FGS) is a severe neglected disease, caused by infection with <em>Schistosoma haematobium</em>. The WHO has prioritized the improvement of diagnostics for FGS and previous studies have explored the PCR-based detection of <em>Schistosoma</em> DNA on genital specimens, with encouraging results. We aimed to determine the prevalence of FGS among women living in an endemic district in North-western Tanzania, applying and preliminary comparing self-collected and operator-collected cervical-vaginal swabs followed by PCR, and to assess the acceptability of these sampling procedures.</p>
Study to Collect Sera for Immunogenicity Testing in Children Vaccinated With Fluzone®
ClinicalTrials.gov study NCT00836953. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Use of the lung flute ECO to assist in sputum collection for tuberculosis testing: a randomized crossover trial
Open the record for dataset details and reuse information.
Retrospective data collected on SATURN, a public domain self-administered cognitive screening test
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.