Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,038
datasets available to search
ShareScore release 0.7.1
Dataset results
8,038 results for “validation”
DCASE 2024 Task 9: Language-Queried Audio Source Separation | Validation Set
<p>This is the <strong>validation set for Task 9, Language-Queried Audio Source Separation (LASS), in DCASE 2024 Challenge</strong>. </p> <p>This validation split is meant to be used for Task 9 at the scientific challenge DCASE 2024. This split is not meant to be used for training LASS methods. This split is meant to be used for evaluating LASS methods during the model development stage.</p> <p>This validation set consists of 1000 audio files sourced from Freesound [1], uploaded between April and October 2023. Each audio file has been manually annotated with three captions. In the annotation guidance, we instructed annotators to describe the content of audio clips using 5-20 words (similar to the caption style in Clotho [3] and AudioCaps [4] datasets). The tags of each audio file were verified and revised according to the FSD50K [2] sound event categories. Each audio file has been chunked into a 10-second clip and downsampled to 16kHz.</p> <p><strong>== Details ==</strong></p> <p>The audio files in the archives:</p> <ul> <li>lass_validation.zip</li> </ul> <p>and the associated metadata (including tags and captions) in the JSON file:</p> <ul> <li>lass_validation.json</li> </ul> <p>Participants will evaluate their LASS models using synthetic mixture data in the development stage. Specifically, given an audio clip A1 and its corresponding caption C, we select an additional audio clip, A2, to serve as background noise, thereby creating a mixed audio, A3. We anticipate that the LASS system, given A3 and C as inputs, will be able to separate the A1 source. We use the revised tags information to ensure that the two audio clips used in each mix do not share overlapping sound source classes. Three thousand synthetic audio mixtures with signal-to-noise ratios (SNR) ranging from -15dB to 15dB will be generated for the validation of LASS model development. These synthetic mixtures can be generated based on the provided CSV file:</p> <ul> <li>lass_synthetic_validation.csv</li> </ul> <p>The evaluation tool can be found at: https://github.com/Audio-AGI/dcase2024_task9_baseline/blob/main/dcase_evaluator.py</p> <p><strong>== References ==</strong></p> <p>[1] Fonseca E, Pons Puig J, Favory X, et al. Freesound datasets: a platform for the creation of open audio datasets. International Society for Music Information Retrieval (ISMIR), 2017.</p> <p>[2] Fonseca E, Favory X, Pons J, et al. FSD50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 30: 829-852.</p> <p>[3] Drossos K, Lipping S, Virtanen T. Clotho: An audio captioning dataset. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2020: 736-740.</p> <p>[4] Kim C D, Kim B, Lee H, et al. AudioCaps: Generating captions for audios in the wild. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2019: 119-132.</p>
Exemplary validation of SMOS Level 3 version 339 Descending vs ERA5-Land v20190904 vs ISMN 20230110 global for EGU24
Exemplary QA4SM validation using FRMs for EGU24: SMOS Level 3 version 339 Descending vs ERA5-Land v20190904 vs ISMN 20230110 global. URL: https://qa4sm.eu/ui/validation-result/7f27889e-95e4-4e67-9a85-2e472cbb935c. Produced on QA4SM (https://qa4sm.eu)
Clermont validation set
<p>This is the validation set presented in [1]. </p> <p>It was used to test the validitity of EzClermont. It includes 125 <em><em><a href="http://doi.org/10.1601/nm.3093" target="_blank" rel="noopener">E. coli</a> </em></em>genomes where the isolates have had Clermont phylotypes experimentally determined. </p> <div> <div>It is presented here as a benchmark set for other tools implementing in silico prediction of Clermont phylotypes. </div> <div> </div> <div>[1]: Waters NR, Abram F, Brennan F, Holmes A, Pritchard L. Easy phylotyping of <em>Escherichia coli via</em> the EzClermont web app and command-line tool. Access Microbiol. 2020 Jun 19;2(9):acmi000143. doi: 10.1099/acmi.0.000143. PMID: 33195978; PMCID: PMC7656184.</div> </div>
Data Sets for Evaluation of the Psychometric Properties and Validity of the German Version of the Process Model of Emotion Regulation Scale (PMERQ)
<p>Data files relate to an investigation of the psychometric properties of the German Version of the Process Model of Emotion Regulation Scale (PMERQ). Data set 1 (pmerq_1) contains information regarding the age, gender, ethnicity, and educational status of participants. In addition, responses to the 45 items of the initial translation of the 10-scale PMERQ are included. Data set 2 (pmerq_2) contains identical sociodemographic variables and responses to the 45 items of the revised translation of the 10-scale PMERQ. In addition, data set 2 contains responses to the 16-item German Interpersonal Emotion Regulation Questionnaire (IERQ), the 10-item German Emotion Regulation Questionnaire (ERQ), , the German version of the 10-item Big Five Inventory-10 (BFI-10), the 4-item German version of the Patient Health Questionnaire-4 (PHQ-4), the German version of the Satisfaction with Life Scale (SWLS), and the 17-item German Social Desirability Scale-17 (SES-17). Data set 2 (pmerq_2) contains identical sociodemographic variables and responses to the 45 items of the readability-improved translation of the 10-scale PMERQ. In addition, data set 2 contains responses to the German version of the Satisfaction with Life Scale (SWLS) and the German version of the 9-items UCLA Loneliness Scale (UCLA).</p>
INTELLIMAN_WP2_Application Requirements and Integration_T2.4_Fresh food handling use case analysis, integration and validation_Apple 6D pose estimation dataset_v0
<p>The dataset contains the data generated for the training of the 6D pose estimation neural network<br>DOPE related to the publication:<br>M. Costanzo, M. De Simone, S. Federico, C. Natale and S. Pirozzi, "Enhanced 6D Pose Estimation for<br>Robotic Fruit Picking," 2023 9th International Conference on Control, Decision and Information<br>Technologies (CoDIT), Rome, Italy, 2023, pp. 901-906, doi: 10.1109/CoDIT58514.2023.10284072.</p>
Methodology of diachronic analysis of old prints and its validation by tracing the changing meaning of key concepts in the intellectual debate of 16th century Italy
<p>This dataset in corpus.zip file contains OCR texts for 16th century Italian books available from BNF (folder gallica) and archive.org (folder internetarchive). Source URL is given in 1st line of each text file. The pages within each document may come in order different than in the source, however, the order of lines in each page is preserved. The archive is protected with password, which will be made public immediately after the publication of the related research paper under the same title.</p> <p>Samples of three documents along with images of their initial pages are provided at the current time.</p>
Validation of the UK'37 paleotemperature proxy in the South Brazilian Bight from core-top sediments
<p>Alkenone unsaturation index, water depth temperature, and nutrient concentration from the South Brazilian Bight (SBB). Temperature and nutrient content were retrieved from the World Ocean Atlas 2018. Residual analyses were performed using six paleoequations against observed temperature from WOA18. The PCA analysis was performed using alkenone unsaturation index, temperature, and nutrient data from each data point.</p>
In situ data for GLEAM4 validation
<p>In situ validation data from eddy covariance stations used in the validation of GLEAM4, as described in "GLEAM4: global land evaporation dataset at 0.1° resolution from 1980 to near present" by Miralles et al. (2025) (preprint doi: <a href="https://doi.org/10.21203/rs.3.rs-5488631/v1">https://doi.org/10.21203/rs.3.rs-5488631/v1</a>). This data is used in <a href="https://github.com/h-cel/GLEAM4">https://github.com/h-cel/GLEAM4</a> repository. </p> <p>Data sources for the in situ data</p> <ul> <li>FLXUNET LA Thuile, FLUXNET2015 and FLUXNET-CH4: <a href="https://fluxnet.org/data/">https://fluxnet.org/data/</a></li> <li>AmerifFlux: <a href="https://ameriflux.lbl.gov/data/aboutdata/">https://ameriflux.lbl.gov/data/aboutdata/</a></li> <li>ICOS: <a href="https://www.icos-cp.eu/data-products">https://www.icos-cp.eu/data-products </a></li> <li>EFDC: <a href="https://www.europe-fluxdata.eu/home/data/data-policy">https://www.europe-fluxdata.eu/home/data/data-policy </a></li> </ul> <p>Data sources for the model data:</p> <ul> <li>GLEAM4 (v4.2a with data assimilation) and GLEAM v3.8 data: <a href="https://www.gleam.eu/">https://www.gleam.eu/</a> </li> <li>ERA5-Land data from the Climate Data Store: <a href="https://cds.climate.copernicus.eu/datasets/reanalysis-era5-land?tab=download">https://cds.climate.copernicus.eu/datasets/reanalysis-era5-land?tab=download</a></li> <li>FLUXCOM data: <a href="https://www.fluxcom.org/EF-Download/" target="_blank" rel="noopener">https://www.fluxcom.org/EF-Download/</a></li> </ul>
Dataset for "Validation of SSDE calculation in a modern CT scanner and correlation with effective dose"
<p>Size-Specific Dose Estimate (SSDE) is a size-adjusted dosimetric index that addresses the limitations of the Computed Tomography Dose Index (CTDIvol). This research aims to verify the SSDE generated by a modern Computed Tomography (CT) scanner and examine its relationship with effective dose (E). Sixty CT scans were performed on anthropomorphic phantoms, including models representing pediatric and obese patients, and then analyzed. SSDE values from the CT scanner were compared with those calculated independently using a Python-based method and Radimetrics, a dose monitoring software.</p> <p>The published dataset contains all the CT images and an Excel file with the main parameters given by the CT scanner, alsongside the ones calculated with Python and Radimetrics. </p>
Validation of MUISCA for Blockchain Interoperability: A Case Study in a Healthcare Environment
<p>This repository contains the questions asked to the 31 experts who participated in the study “Validation of MUISCA for Blockchain Interoperability: A Case Study in a Healthcare Environment”. It also contains the spreadsheet with the recorded answers.</p>
Resolution, validation and divergence of heterozygous haplotypes from pooled long read sequencing of the diamondback moth (Lepidoptera: Plutellidae)
<p>This data sets includes the full set of intermediate genome assemblies produced during our analyses. We make these available for researchers who may be interested in the variation between assembly results and variation within the study organism (<em>Plutella xylostella</em>) prior to removal during subsequent genome processing.</p>
Validation of Work Values Instrument in Final Year University Students
<p><em>This is the raw dataset when we conducted the adaptation of the work values instrument in final year students. The number of participants in this study was 316 students, comprised of final year students from various majors who were selected by quota sampling. The raw dataset has been disseminated in the undergraduate thesis defense, so the publication date stated above is referring to it. The raw data set has never been published elsewhere. </em></p>
Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: validation cohort meta data and parsed TCR repertoire data
<p>Meta data corresponding the the validation cohort for the paper, "Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities" by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include: </p> <p>(1) SNP genotypes for the two SNPs which overlap with the discovery cohort<br> - (nicaragua_snp_genotypes_ints.tsv) -- SNP genotypes as integers<br> - (nicaragua_snp_genotypes_strings.tsv) -- SNP genotypes as allele strings <br> (2) the ancestry PCs for each individual in the validation cohort (nicaragua_snp_ancestry_PCA.tsv)<br> (3) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (4) a file including IMGT genes used for parsing TCRA repertoire data (human_vj_alleles_alpha.tsv)<br> (5) Parsed TCRA repertoire data (nicaragua_parsed_TCRA.tgz)<br> (6) Parsed TCRB repertoire data (nicaragua_parsed_TCRB.tgz) </p> <p><strong>Corresponding raw validation cohort TCR repertoire data is available here:</strong> https://www. ncbi.nlm.nih.gov/bioproject/PRJNA762269 (The BioProject database, accession number: PRJNA762269)</p> <p><strong>Software tools designed to work with these data are available here:</strong> https://github.com/phbradley/tcr-gwas</p>
Multi-contrast MRI and histology datasets used to train and validate MRH networks to generate virtual mouse brain histology
<p><span>H MRI maps brain structure and function non-invasively through versatile contrasts that exploit inhomogeneity in tissue micro-environments. Inferring histopathological information from MRI findings, however, remains challenging due to absence of direct links between MRI signals and cellular structures. Here, we provided deep convolutional neural networks, called MRH-Nets, developed using co-registered multi-contrast MRI and histological data of the mouse brain, can estimate histological staining intensity directly from MRI signals at each voxel. The results provide three-dimensional maps of axons and myelin with tissue contrasts that closely mimics target histology and enhanced sensitivity and specificity compared to conventional MRI markers. </span><span> </span>The dataset contains multi-contrast MRI and histology used for the training and testing and the acquisition parameters. The datasets have been carefully registered to mouse brain images from the Allen Mouse Brain Atlas (https://mouse.brain-map.org). The source codes for MRH-Nets can be found at <a href="https://github.com/liangzifei/MRH-Net">https://github.com/liangzifei/MRH-Net</a>.</p>
Multi-stakeholder Validation of Entrustable Professional Activities in Family Medicine - Survey Responses
<p>Survey Responses from the 2021 Multi-Stakeholder Validation of Entrustable Professional Activities in Family Medicine medical education research project.</p>
OTUs with valid matched taxid on gg_13_15 99_otu tree
<p>A mapping file between OTUs and Taxids on the 99_otus tree of gg_13_5 data package, for reproducibility of the WGSUniFrac project.</p>
Content Validity and Reliability of the Italian Language Version of the US National Cancer Institute's Patient-Reported Outcomes version of the Common Terminology Criteria for Adverse Events (PRO-CTCAE®)
<p><strong>ABSTRACT </strong></p> <p><strong>Introduction: </strong>The US National Cancer Institute’s (NCI) Patient-Reported Outcomes version of the Common Terminology Criteria for Adverse Events (PRO-CTCAE<sup>®</sup>) is a library of 78 symptom terms and 124 items enabling patient reporting of symptomatic adverse events in cancer trials. This multicenter study used mixed methods to develop an Italian language version of this widely accepted measure, and evaluate content validity and reliability in a diverse sample of Italian-speaking patients.</p> <p><strong>Methods: </strong>All PRO-CTCAE items were translated in accordance with international guidelines. Subsequently, the content validity of the PRO-CTCAE-Italian was examined and iteratively refined through cognitive debriefing interviews. Participants (n=96; 52% male; median age 64 years; 26% older adults; 18% lower educational attainment) completed a PRO-CTCAE survey and participated in a semi-structured interview to determine if the translation captured the concepts of the original English language PRO-CTCAE, and to evaluate comprehension, clarity and ease of judgement. Test-retest reliability of the finalized measure was evaluated in a second sample (n=135).</p> <p><strong>Results: </strong>Four rounds of cognitive debriefing interviews were conducted. The majority of PRO-CTCAE symptom terms, attributes and associated response choices were well-understood, and respondents found the items easy to judge. To improve comprehension and clarity, the symptom terms for nausea and pain were rephrased and retested in subsequent interview rounds. Test-retest reliability was excellent for 41/49 items (84%); the median intraclass correlation coefficient was 0.83 (range 0.64-0.94).</p> <p><strong>Discussion: </strong>Results support the semantic, conceptual and pragmatic equivalence of PRO-CTCAE-Italian to the original English version, and provide preliminary evidence of content validity and reliability.</p>
Validation of the Arabic version of the mental toughness questionnaire
<p>The objective of this study was to assess the construct validity of the Arabic version of mental toughness MT based on the theoretical conception of the MTQ48 questionnaire <a href="#_ENREF_7">Clough et al. (2002)</a>. The 48 items consist of six components 6C :(1)Challenge CH, (2) Commitment CO, (3)Emotion Control EC, (4)Life Control LC, (5)Confidence in abilities CA,(6)Interpersonal confidence IC ,and the global score MT. The sample consisted of 853 Tunisian participants (444 males and 409 females; 409 athletes and 444 non-athletes), aged 14-27 years (<em>M=20.38 SD=4.12</em>). Cronbach's alpha suggests that over 48 items has adequate internal consistency (<em>α=.72)</em>. The EFA revealed a good sampling quality (df = 1128; p < .001). The CFA approved a good model fit (<em>χ²=1146.33; df =1065; CFI=.93; SRMR=.063; RMSEA=.009</em>). In conclusion, our results allowed us to propose a valid Arabic measure of the 6C of MT. The results confirm the factorial validity of the MTQ48 and indicate that the Arabic version of the questionnaire has a robust measure of the psychometric properties of mental toughness. Finally, stakeholders in the Arab region should be able to benefit from the MTQ48 questionnaire, such as and reliable measurement and assessment instruments.</p>
Modeling and validating a SuperDARN radar's Poynting flux profile
<p>Numerical ray trace modelingl output corresponding to Radio Science manuscript "Modeling and validating a SuperDARN radar's Poynting flux profile" by G. W. Perry.</p>
Vasculature Segmentation Validation Dataset - Part IV - Biological Difference Dataset 2/2 (Development)
<p>Dataset to allow exploration of data enhancement, segmentation, and validation for: https://www.biorxiv.org/content/10.1101/2020.07.21.213843v1 and associated future publications<br> </p> <p><strong>Dataset description</strong>:</p> <p><strong>Development</strong>: Example data of the head vasculature of developing zebrafish.</p> <p><br> <strong>Data acquisition</strong>:<br> Experiments were performed according to the rules and guidelines of institutional and UK Home Office regulations under the Home Office Project Licence 70/8588 held by TC. Maintenance of adult zebrafish Tg(kdrl:HRAS-mCherry)s916 (Chi et al., 2008) was performed as described in standard husbandry protocols (Westerfield, 1993). Embryos, obtained from controlled mating, were kept in E3 medium buffer with methylene blue and staged as previously described (Kimmel et al., 1995).</p> <p>Embryos were embedded in 2% LMP-agarose with 0.01% Tricaine in E3 (MS-222, Sigma). Data were acquired using a Zeiss Z.1 light sheet microscope, Plan-Apochromat 20x/1.0 Corr nd=1.38 objective, dual-side illumination with online fusion, activated Pivot Scan, image acquisition chamber incubation at 28°C, with a scientific complementary metal-oxide semiconductor (sCMOS) detection unit. Data properties can be summarised as: 16bit image depth, voxel dimensions in x, y and z of 1920 x 1920 x 400-600, respectively, giving a voxel size of 0.33 x 0.33 x 0.5 µm). </p> <p><strong>Contact</strong>: kugler.elisabeth[at]gmail.com</p> <p><strong>Useful code links</strong>: </p> <ul> <li>https://github.com/ElisabethKugler/ZFVascularQuantification</li> <li>https://github.com/ElisabethKugler/Matlab3D-ImageAnalysis<br> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.