Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
Chemical Name Extraction Based on Automatic Training Data Generation
<p>The automation of extracting chemical names from text has significant value to biomedical and life science research. A major barrier in this task is the difficulty of getting a sizable and good quality data to train a reliable entity extraction model. Another difficulty is the selection of informative features of chemical names, since comprehensive domain knowledge on chemistry nomenclature is required. Leveraging random text generation techniques, we explore the idea of automatically creating training sets for the task of chemical name extraction. Assuming the availability of an incomplete list of chemical names, called a dictionary, we are able to generate well-controlled, random, yet realistic chemical-like training documents. We statistically analyze the construction of chemical names based on the incomplete dictionary, and propose a series of new features, without relying on any domain knowledge. Compared to state-of-the-art models learned from manually labeled data and domain knowledge, our solution shows better or comparable results in annotating real-world data with less human effort. Moreover, we report an interesting observation about the language for chemical names. That is, both the structural and semantic components of chemical names follow a Zipfian distribution, which resembles many natural languages.</p>
Supplementary material 7 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
This document contains an annotated set of data quality checks that participants report they use when evaluating and cleaning datasets. These items outline how participants are judging if the data suits their purpose.
Supplementary material 3 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
The informed consent request and workshop survey questions given to participants after the workshop each day for 4 consecutive days.
Supplementary material 2 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
This document shows just the questions we asked the applicants who applied to participate in this Georeferencing for Research Use workshop. We used a Google Form to deliver these questions and collect responses. It is both an application and serves as our pre-workshop survey.
Supplementary material 8 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
Summary of desired future workshop topics that were listed by participants on the last day of the workshop.
Supplementary material 6 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
Summary of topics to be covered in an ideal workshop as identified by workshop applicants in the workshop call for participation. We incorporated as many as possible that also fit our scope.
Supplementary material 5 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
Questions we asked in the Georeferencing for Research Follow Up Survey done 3 months after the workshop.
Supplementary material 4 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
Three months after the workshop, participants were surveyed to assess what workshop-related knowledge and materials were being used and disseminated to others. This document summarized data collected in this particular survey.
Supplementary material 1 from: Seltmann K, Lafia S, Paul D, James S, Bloom D, Rios N, Ellis S, Farrell U, Utrup J, Yost M, Davis E, Emery R, Motz G, Kimmig J, Shirey V, Sandall E, Park D, Tyrrell C, Thackurdeen R, Collins M, O'Leary V, Prestridge H, Evelyn C, Nyberg B (2018) Georeferencing for Research Use (GRU): An integrated geospatial training paradigm for biocollections researchers and data providers. Research Ideas and Outcomes 4: e32449. https://doi.org/10.3897/rio.4.e32449
Darwin Core Archive file downloaded from the iDigBio portal for use in the Georeferencing for Research Use workshop. Total 25,429 records, accessed on 2016-08-29. Collections contributing to the record set are listed in the archive records.citation.txt file. Dataset GUID: a69d1541-4726-465d-84ad-50c7ed556eee
HDNNP training data set for Cu2S
<p>High-dimensional neural network potential (HDNNP) training data set for copper sulfide (Cu<sub>2</sub>S). Unpack archive and see README.pdf for details.</p>
HDNNP training data set for H2O
<p>High-dimensional neural network potential (HDNNP) training data set for water. Unpack archive and see README.pdf for details.</p>
ATAC-Seq (training data)
<p>Training dataset for a Galaxy ATAC-seq tutorial.</p>
Data for the training and testing of ccAFv2
<p><span>Single-cell transcriptomics has unveiled a vast landscape of cellular heterogeneity in which the cell cycle is a significant component. We trained a high-resolution cell cycle classifier (ccAFv2) using single cell RNA-seq (scRNA-seq) characterized human neural stem cells. The features of this classifier are that it classifies six cell cycle states (G1, Late G1, S, S/G2, G2/M, and M/Early G1) and a quiescent-like G0 state, and it incorporates a tunable parameter to filter out less certain classifications. The ccAFv2 classifier performed better than or equivalent to other state-of-the-art methods even while classifying more cell cycle states, including G0. We showcased the versatility of ccAFv2 by successfully applying it to classify cells, nuclei, and spatial transcriptomics data in humans and mice, using various normalization methods and gene identifiers. We provide methods to regress the cell cycle expression patterns out of single cell or nuclei data to uncover underlying biological signals. The classifier can be used either as an R package integrated with Seurat (</span><span><a href="https://github.com/plaisier-lab/ccafv2_R"><span>https://github.com/plaisier-lab/ccafv2_R</span></a></span><span>) or a PyPI package integrated with scanpy (</span><span><a href="https://pypi.org/project/ccAF/"><span>https://pypi.org/project/ccAF/</span></a></span><span>). We proved that ccAFv2 has enhanced accuracy, flexibility, and adaptability across various experimental conditions, establishing ccAFv2 as a powerful tool for dissecting complex biological systems, unraveling cellular heterogeneity, and deciphering the molecular mechanisms by which proliferation and quiescence affect cellular processes.</span></p>
Deep learning for the occurrence of tipping points: training data
<p>This data accompanies the manuscript by Chengzuo Zhuge et al. “Deep learning for the occurrence of tipping points” and the Github repository <a href="https://github.com/zhugchzo/dl_occurrence_tipping">https://github.com/zhugchzo/dl_occurrence_tipping</a>. It contains the model time series data that are used to train the deep learning algorithm. The directory <br>increased_bifurcation contains 150k time series (50k Fold, Hopf, Transcritical respectively) with parameter increasing and the directory decreased_bifurcation contains 150k time series (50k Fold, Hopf, Transcritical respectively) with parameter decreasing. The directory pitchfork contains 100k time series (50k supercritical and subcritical pitchfork respectively) with parameter increasing. Both directories contain files labels.csv and groups.csv which provide numbers corresponding to the labels (The tipping points) and groups (Training, Validation, Test) for each time series respectively.</p>
The NIMROD training data for the STFNO model
Open the record for dataset details and reuse information.
Sub Volume of Training Data to Train and Test 3DUnet
<p>download test for colab </p>
Koster et al. - Balance training in older adults enhances feedback control after perturbations - data & code
<p>Data & code for: </p> <p><strong><u>Balance training in older adults enhances feedback control after perturbations</u></strong></p> <p><em><u>R A J Koster, L Alizadehsaravi, W Muijres, S M Bruijn, N Dominici, J H van Dieen</u></em></p> <p> </p> <p>This folder contains two subfolders: <em>Code</em> & <em>Data</em>. The folder <em>Code</em> contains the MATLAB scripts used to analyse the data contained within the folder <em>Data</em>.</p> <p><strong>Scripts</strong>/</p> <p>· <strong>RK_Analysis_Kinematics.m</strong>: This is the main script used to analyse the kinematics. Running it will compute the kinematic parameters investigated in the paper and contrast them between conditions. Visualisations of these as well as the statistical results will be stored in the newly created folder <em>/Data/Processed/Figures/Kinematics/</em></p> <p>· <strong>RK_Analysis_Synergies.m</strong>: This is the main script used to analyse the EMG activity. Running it will compute the synergies from the EMGs and contrast them between conditions. Visualisations of these as well as the statistical results will be stored in the newly created folder <em>/Data/Processed/Figures/Synergies/</em></p> <p>· <strong>Subfunctions/</strong>: This folder contains all scripts used within the two main scripts. These scripts are subdivided into 3 folders (<em>General/, Kinematics analysis/, Synergy analysis/</em>) based on which part of the analysis they belong to.</p> <p>o <strong>General/spmi1d/</strong>: The external toolbox used to perform the statistics.</p> <p>o <strong>Kinematics analysis/=VU 3D model=/</strong>: the VU 3D model used to compute kinematic parameters from the trajectories of individual body segments.</p> <p><strong>Data/</strong></p> <p>· <strong>EMG/</strong>: Contains the raw EMG data for each participant & recording.</p> <p>· <strong>Excel sheets/</strong>: Contains participant & recording information sheets. This is used to determine on which leg the participant was standing.</p> <p>· <strong>PT/</strong>: Contains information about the rotating platform the participants were standing on for every recording. This is used to determine perturbation onset and direction.</p> <p>· <strong>Trajectory/</strong>: Contains the trajectories of the participant’s body segments during the recordings.</p> <p>· <strong>Processed/</strong>: (This folder is not there initially, but it is added by the 2 main codes) Will contain the kinematic parameters and synergy weights & activation patterns for every condition and participant after they are computed. This folder is also where the results from the analysis will be stored.</p> <p>o <strong>Figures/</strong>: For kinematics & synergies separately, will contain figures displaying the parameters and the statistical results. The ‘<em>Stats – X – Y.tiff</em>’ figure files display the ANOVA and post hoc test results. If statistical differences are found, their p-values are reported in the title of the individual plots (Note: for subthreshold results p-values are undefined in such analyses). Additional individual muscle analysis is included in the synergy analysis.</p>
De-identified data on outcomes of a caregiver-led versus therapist-led training programme
Open the record for dataset details and reuse information.
Training data for muTOV using a piecewise polytope model
Open the record for dataset details and reuse information.
Data for training BNN model
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.