Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Models and Predictions for "The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction"
<p><strong>Models and Predictions</strong></p> <p>This dataset contains the trained XGBoost and EA-LSTM models and the models' predictions for the paper <a href="https://github.com/gauchm/ealstm_regional_modeling"><em>The Proper Care and Feeding of CAMELS: How Limited Training Data Affects Streamflow Prediction</em></a>.</p> <p>For each input sequence length (10, 30, 100, 270*, 365*) and each combination of model (XGBoost, EA-LSTM), training years (3, 6, 9), number of basins (13, 26, 53, 265, 531), and seed (111-888), there are five folders. Each corresponds to a random basin sample (for 531 basins there's only one folder, since it's all basins).<br> In each folder, there are three files:</p> <ul> <li><span class="math-tex">\(\texttt{model.pkl}\)</span> (XGBoost) or <em><span class="math-tex">\(\texttt{model_epoch30.pt}\)</span></em> (EA-LSTM), which stores the pickled trained model</li> <li><em><span class="math-tex">\(\texttt{xgboost_seedNNN.p}\)</span></em> or <em><span class="math-tex">\(\texttt{ealstm_seedNNN.p}\)</span></em>, which stores a pickled dictionary that maps each basin to the DataFrame of predicted and actual daily streamflow.</li> <li><span class="math-tex">\(\texttt{attributes.db}\)</span>, which stores static catchment attributes needed for inference.</li> </ul> <p>In addition to each folder, there is a SLURM submission script called <em><span class="math-tex">\(\texttt{<foldername>.sbatch}\)</span></em> that was used to create and evaluate the model in the folder.</p> <p> </p> <p>* sequence lengths 270 and 365 only contain data for EA-LSTM.</p>
Platynereis EM training data
<p>Training data for Convolutional Neural Networks used in the publication Whole-body integration of gene expression and single-cell morphology. We provide training data for segmenting structures in the SerialBlockface Electron Microscopy data-set containing a complete 6 day old Platynereis dumerilii larva, in particular for:<br> - cell membranes: 9 training blocks @ resolution 20x20x25 nm. Based on initial training data provided by https://ariadne.ai/.<br> - cilia: 3 training and 2 validation blocks @ resolution 20x20x25 nm.<br> - cuticle: 5 training blocks @ resolution 40x40x50 nm.<br> - nuclei: 12 training blocks @ resolution 80x80x100 nm. Based on initial training data provided by https://ariadne.ai/.<br> <br> For details on how to use this data for training, see https://github.com/platybrowser/platybrowser-backend/tree/master/segmentation.</p>
Training Data of Quantitative Online NMR Spectroscopy for Artificial Neural Networks
<p>Data set of low-field NMR spectra of continuous synthesis of nitro-4’-methyldiphenylamine (MNDPA). <sup>1</sup>H spectra (43 MHz) were recorded as single scans.</p> <p> Two different approaches for the generation of artificial neural networks training data for the prediction of reactant concentrations were used: (<em>i</em>) Training data based on combinations of measured pure component spectra and (<em>ii</em>) Training data based on a spectral model.</p> <p><strong>Synthetic low-field NMR spectra</strong></p> <p>First 4 columns in MAT-files represent component areas of each reactant within the synthetic mixture spectrum.</p> <p><em>X<sub>i</sub></em> (“pure component spectra dataset”)</p> <p><em>X<sub>ii</sub></em> (“spectral model dataset”)</p> <p><strong>Experimental low-field NMR spectra from MNDPA-Synthesis</strong></p> <p>This data set represents low-field NMR-spectra recorded during continuous synthesis of nitro-4’-methyldiphenylamine (MNDPA). Reference values from high-field NMR results are included.</p>
Example computer vision classification training data derived from British Library 19th Century Books Image collection
<p>Example computer vision classification training data derived from British Library 19th Century Books Image collection</p> <p>This dataset provides training data for image classification for use in a computer vision workshop. The images are derived from '<a href="https://doi.org/10.21250/db17">Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG'</a> from the year '1839'.</p> <p>Currently, included are four folders containing a variety of images derived from the BL books corpus.</p> <ul> <li>'cv_workshop_exercise_data' include images of: 'building', 'people', 'coat of arms'</li> <li>'humancats' contains images of humans and images of cats</li> </ul> <p>The 'fashion' and 'portraits' folders both contain images of people organised into 'female' and 'male'. These labels were annotated by a single annotator and these categories may themselves not be meaningful. They are included in the workshop data as a point of discussion about how we should label data both in general and when working with historical data. </p> <p>This data is intended primarily as an educational resource.</p>
deepCR Original Training Data
<p>deepCR original training dataset.</p>
Training data for "Apply modeling on population or community data to see effect of year, habitat or site on species abundance"
<p>Datasets for the "Apply modeling on population or community data to see effect of year, habitat or site on species abundance" Galaxy for ecology tutorial</p>
M. tuberculosis bioinformatics training data
<p>Training data based on for tutorial on M. tuberculosis variant calling and annotation based on inferred ancestral reference genome of <em>M. tuberculosis</em> combined with NCBI NC_000962.3 annotation in Genbank format and reads from "Use of whole-genome sequencing to distinguish relapse from reinfection in a completed tuberculosis clinical trial".</p>
Composite 2D video of raw and processed video footage from an outdoor camerawork training session for qualitative data collection
<p>The video clip, in traditional 2D format, contains a short excerpt (1:43 minutes) from an outdoor camerawork training session. The raw and processed footage from the different cameras is composited in a single frame. The original composite file has 32 channels of audio so that one can switch between different microphones and combinations of microphones. This is indicated in the video clip but is not available in this file. All participants are playing particular roles in the training session, and each carries a camera. In preparation for the real data collection with a guide, one person is pretending to be a nature guide. She carries a GoPro camera on a gimbal. There is an instructor, who is carrying a single lens 360° camera on a raised extension pole with a separate ambisonic microphone. Two others are filming with a prosumer camcorder and a single lens 360° camera on a lowered extension pole respectively. And a fifth person is filming with a stereoscopic 360° camera and an independent ambisonic microphone on a monopod. In a nutshell, this is a typical team filming arrangement, in which the team needs to attentively yet silently coordinate their joint camerawork. Languages: Danish and English</p>
A live screen capture of the AVA360VR prototype being used to annotate camerawork training video data
<p>In this 2D video clip, we see a live screen capture of the AVA360VR prototype being used to annotate camerawork training data. This clip was recorded in January 2018 with an alpha version of the prototype. The functions shown do not necessarily reflect those in the final software tool.</p> <p><em>AVA360VR </em>(Annotate, Visualise, Analyse 360° video in VR) is a VR software tool developed by the BigSoftVideo team at Aalborg University. The aim is to support ‘inhabiting’ 360-degree video data – that is, to explore complex spatial video and audio recordings of a scene in which social interaction took place through a tangible interface in virtual reality.</p> <p> </p> <p> </p> <p> </p>
FASTQE (training data)
<p>Data is from <a href="https://qubeshub.org/publications/1092/2">https://qubeshub.org/publications/1092/2</a></p>
University of Denver Collections as Data - HTR Train and Validation Set JCRS_2020_5_27
<p><a href="https://zenodo.org/api/files/333ecb88-1f48-4ffd-b39e-5e70b800c276/HTR_Train_Set_JCRS_2020_5_27.zip">HTR_Train_Set_JCRS_2020_5_27.zip</a> <br> Description</p>
RDP taxonomic training data formatted for DADA2 (RDP trainset 18/release 11.5)
<p>These DADA2-formatted training fasta files were derived from the Ribosomal Database Project's Training Set 18 and the 11.5 release of the RDP database.</p> <p>These fastas were generated by the following commands (using the dada2 R package version 1.19.1):</p> <blockquote> <p>## The RDP trainset data was downloaded from: https://sourceforge.net/projects/rdp-classifier/files/RDP_Classifier_TrainingData/<br> path <- "~/Desktop/RDP/RDPClassifier_16S_trainsetNo18_rawtrainingdata"<br> dada2:::makeTaxonomyFasta_RDP(file.path(path, "trainset18_062020.fa"), <br> file.path(path, "trainset18_db_taxid.txt"), <br> "~/tax/rdp_train_set_18.fa.gz")<br> ## Download "ten_16s.100.fa" from Robert Edgar's taxonomy testing page: https://drive5.com/taxxi/doc/fasta_index.html<br> dada2:::tax.check("~/tax/rdp_train_set_18.fa.gz", "~/Desktop/ten_16s.100.fa")</p> <p>## This function creates the dada2 assignSpecies fasta file for the RDP from the RDP's _Bacteria_unaligned.fa file.<br> dada2:::makeSpeciesFasta_RDP("~/Desktop/RDP/current_Bacteria_unaligned.fa", "~/tax/rdp_species_assignment_18.fa.gz")<br> dada2:::tax.check("~/tax/rdp_species_assignment_18.fa.gz", "~/Desktop/ten_16s.100.fa", mode="species")</p> </blockquote>
Data from: Predicting classifier performance with limited training data: applications to computer-aided diagnosis in breast and prostate cancer
Clinical trials increasingly employ medical imaging data in conjunction with supervised classifiers, where the latter require large amounts of training data to accurately model the system. Yet, a classifier selected at the start of the trial based on smaller and more accessible datasets may yield inaccurate and unstable classification performance. In this paper, we aim to address two common concerns in classifier selection for clinical trials: (1) predicting expected classifier performance for large datasets based on error rates calculated from smaller datasets and (2) the selection of appropriate classifiers based on expected performance for larger datasets. We present a framework for comparative evaluation of classifiers using only limited amounts of training data by using random repeated sampling (RRS) in conjunction with a cross-validation sampling strategy. Extrapolated error rates are subsequently validated via comparison with leave-one-out cross-validation performed on a larger dataset. The ability to predict error rates as dataset size increases is demonstrated on both synthetic data as well as three different computational imaging tasks: detecting cancerous image regions in prostate histopathology, differentiating high and low grade cancer in breast histopathology, and detecting cancerous metavoxels in prostate magnetic resonance spectroscopy. For each task, the relationships between 3 distinct classifiers (k-nearest neighbor, naive Bayes, Support Vector Machine) are explored. Further quantitative evaluation in terms of interquartile range (IQR) suggests that our approach consistently yields error rates with lower variability (mean IQRs of 0.0070, 0.0127, and 0.0140) than a traditional RRS approach (mean IQRs of 0.0297, 0.0779, and 0.305) that does not employ cross-validation sampling for all three datasets.
Data from: Effect of canine oxytocin receptor gene polymorphism on the successful training of drug detection dogs
Drug detection dogs can be trained to locate various prohibited drugs with targeted odors, and they play an important role in interdiction of drug smuggling in human society. Recent studies provide the interesting hypothesis that the oxytocin system serves as a biological basis for co-evolution between dogs and humans. Here, we offer the new possibility that genetic variation of the canine oxytocin receptor (OXTR) gene may regulate the success of a dog's training to become a drug detection dog. A total of 340 Labrador Retriever dogs that were trained to be drug detection dogs in Japan were analyzed. We genotyped an exonic SNP (rs8679682) in the OXTR gene and compared the training success rate of dogs with different genotypes. We also asked dog trainers in the training facility to evaluate subjective personality assessment scores for each dog, and examined how each dog's training success was related to those scores. A significant effect of the OXTR genotype on the success of the dogs' training was found, with a higher proportion of dogs carrying the C allele (T/C and C/C genotypes) being successful candidates than dogs carrying the T/T genotype. Dog personality scores of Training Focus (Factor 1) were positively correlated with an increased likelihood that a dog would successfully complete training. Although the molecular mechanism of the OXTR gene and its functional pathway related to dog behavior remains unknown, our findings suggest that canine OXTR gene variants may regulate individual differences between dogs in their responsiveness to training for drug detection.
Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains
With the ability to observe the activity from large numbers of neurons simultaneously using modern recording technologies, the chance to identify sub-networks involved in coordinated processing increases. Sequences of synchronous spike events (SSEs) constitute one type of such coordinated spiking that propagates activity in a temporally precise manner. The synfire chain was proposed as one potential model for such network processing. Previous work introduced a method for visualization of SSEs in massively parallel spike trains, based on an intersection matrix that contains in each entry the degree of overlap of active neurons in two corresponding time bins. Repeated SSEs are reflected in the matrix as diagonal structures of high overlap values. The method as such, however, leaves the task of identifying these diagonal structures to visual inspection rather than to a quantitative analysis. Here we present ASSET (Analysis of Sequences of Synchronous EvenTs), an improved, fully automated method which determines diagonal structures in the intersection matrix by a robust mathematical procedure. The method consists of a sequence of steps that i) assess which entries in the matrix potentially belong to a diagonal structure, ii) cluster these entries into individual diagonal structures and iii) determine the neurons composing the associated SSEs. We employ parallel point processes generated by stochastic simulations as test data to demonstrate the performance of the method under a wide range of realistic scenarios, including different types of non-stationarity of the spiking activity and different correlation structures. Finally, the ability of the method to discover SSEs is demonstrated on complex data from large network simulations with embedded synfire chains. Thus, ASSET represents an effective and efficient tool to analyze massively parallel spike data for temporal sequences of synchronous activity.
Data from: Adherence to trained standards after a faculty development workshop on "Teaching With Simulated Patients"
Background: Nowadays, faculty development programs to improve teaching quality are considered to be very important by medical educators from all over the world. However, the assessment of the impact of such programs rarely exceeds tests of participants' knowledge gain or self-assessments of their teaching behavior. It remains unclear what exactly is expected of the attending faculty and how the transfer to practice may be measured more comprehensively and accurately. Method: This study evaluates how specific teaching standards were applied after a workshop (10 teaching units) focusing on teaching communication skills with simulated patients. Trained observers used a validated checklist to observe 60 teaching sessions (held by 60 different teachers) of a communication skills course integrating simulated patients. Additionally, we assessed the amount of time that had passed since their participation in the workshop and asked them to rate the importance of communication and social skills in medical education. Results: The observations showed that more than two thirds of teaching standards were met by at least 75% of teachers. Fulfillment of standards was significantly connected to teachers' rating of the importance of communication and social skills (tb=-.21, p=.03). In addition, the results suggest a slight decrease in the amount of fulfilled standards over time (r=-.14, p=.15). Conclusions: Teachers' adherence to basic teaching standards was already satisfying after a one-day workshop. More complex issues need to be re-addressed in further faculty development courses with a special focus on teachers' attitude towards teaching. In future, continuing evaluations of the transfer of knowledge and skills from faculty development courses into practice, preferably including pre-tests or control groups, are needed.
Data from: An integrated iterative annotation technique for easing neural network training in medical image analysis
Neural networks promise to bring robust, quantitative analysis to medical fields. However, their adoption is limited by the technicalities of training these networks and the required volume and quality of human-generated annotations. To address this gap in the field of pathology, we have created an intuitive interface for data annotation and the display of neural network predictions within a commonly used digital pathology whole-slide viewer. This strategy used a 'human-in-the-loop' to reduce the annotation burden. We demonstrate that segmentation of human and mouse renal micro compartments is repeatedly improved when humans interact with automatically generated annotations throughout the training process. Finally, to show the adaptability of this technique to other medical imaging fields, we demonstrate its ability to iteratively segment human prostate glands from radiology imaging data.
Data from: Handover training for medical students – a controlled educational trial of a pilot curriculum
Background: Handovers are a critical point of patient care and a significant source of adverse events. The WHO patient safety curriculum provides some structure for handover teaching; in Europe, there is no standardized curriculum for undergraduate handover training. To address this, the Aachen Interdisciplinary Training Centre for Medical Education, developed and established a pilot curriculum for handover training in the context of the EU-funded PATIENT-project Objective: To develop and implement a handover curriculum for medical students and to assess its effect on students' awareness, confidence and knowledge regarding patient safety and handover in multiple settings. Methods: The pilot handover training curriculum was designed following Kern´s principles of curriculum development and was integrated into a curricular course led by departments for anesthesiology and intensive care (AI) at the University Hospital. A controlled educational research study was conducted with 4th year medical students (n=147) who either received the standard existing curriculum (no teaching of handover, n=78) or the pilot handover training (n=69). Paper-based questionnaires regarding attitude, confidence and knowledge towards handover and patient safety were used for pre- and post-assessment. The pilot curriculum consisted of 3 units (1-2 hours each) integrated into a 4-week course of AI. Multiple types of handover (end-of-shift, operating room/post anesthesia recovery unit/ICU, telephone, discharge) were addressed. Results: Students showed a significant increase in knowledge (p<0.01) and self-confidence for the use of standardized handover tools (p<0.01) and accurate handover performance (p<0.01) among the pilot group. Discussion/Conclusion: We developed and implemented a pilot curriculum for undergraduate handover training. Students displayed a significant increase in knowledge and self-confidence for the use of standardized handover tools and accurate handover performance. An evaluation of the curriculum by other faculties is needed. Further studies should evaluate whether the observed effect of a specific handover strategy is associated with a patient benefit.
Data from: Lower education level is a risk factor for peritonitis and technique failure but not a risk for overall mortality in peritoneal dialysis under comprehensive training system
Background: Lower education level could be a risk factor for higher peritoneal dialysis (PD)-associated peritonitis, potentially resulting in technique failure. This study evaluated the influence of lower education level on the development of peritonitis, technique failure, and overall mortality. Methods: Patients over 18 years of age who started PD at Seoul National University Hospital between 2000 and 2012 with information on the academic background were enrolled. Patients were divided into three groups: middle school or lower (academic year ? 9, n = 102), high school (9 < academic year ? 12, n = 229), and higher than high school (academic year > 12, n = 324). Outcomes were analyzed using Cox proportional hazards models and competing risk regression. Results: A total of 655 incident PD patients (60.9% male, age 48.4 ± 14.1 years) were analyzed. During follow-up for 41 (interquartile range, 20-65) months, 255 patients (38.9%) experienced more than one episode of peritonitis, 138 patients (21.1%) underwent technique failure, and 78 patients (11.9%) died. After adjustment, middle school or lower education group was an independent risk factor for peritonitis (adjusted hazard ratio [HR], 1.61; 95% confidence interval [CI], 1.10-2.36; P = 0.015) and technique failure (adjusted HR, 1.87; 95% CI, 1.10-3.18; P = 0.038), compared with higher than high school education group. However, lower education was not associated with increased mortality either by as-treated (adjusted HR, 1.11; 95% CI, 0.53-2.33; P = 0.788) or intent-to-treat analysis (P = 0.726). Conclusions: Although lower education was a significant risk factor for peritonitis and technique failure, it was not associated with increased mortality in PD patients. Comprehensive training and multidisciplinary education may overcome the lower education level in PD.
Data from: Using matrix and tensor factorizations for the single-trial analysis of population spike trains
Advances in neuronal recording techniques are leading to ever larger numbers of simultaneously monitored neurons. This poses the important analytical challenge of how to capture compactly all sensory information that neural population codes carry in their spatial dimension (differences in stimulus tuning across neurons at different locations), in their temporal dimension (temporal neural response variations), or in their combination (temporally coordinated neural population firing). Here we investigate the utility of tensor factorizations of population spike trains along space and time. These factorizations decompose a dataset of single-trial population spike trains into spatial firing patterns (combinations of neurons firing together), temporal firing patterns (temporal activation of these groups of neurons) and trial-dependent activation coefficients (strength of recruitment of such neural patterns on each trial). We validated various factorization methods on simulated data and on populations of ganglion cells simultaneously recorded in the salamander retina. We found that single-trial tensor space-by-time decompositions provided low-dimensional data-robust representations of spike trains that capture efficiently both their spatial and temporal information about sensory stimuli. Tensor decompositions with orthogonality constraints were the most efficient in extracting sensory information, whereas non-negative tensor decompositions worked well even on non-independent and overlapping spike patterns, and retrieved informative firing patterns expressed by the same population in response to novel stimuli. Our method showed that populations of retinal ganglion cells carried information in their spike timing on the ten-milliseconds-scale about spatial details of natural images. This information could not be recovered from the spike counts of these cells. First-spike latencies carried the majority of information provided by the whole spike train about fine-scale image features, and supplied almost as much information about coarse natural image features as firing rates. Together, these results highlight the importance of spike timing, and particularly of first-spike latencies, in retinal coding.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.