Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
558
datasets available to search
ShareScore release 0.9.0
Dataset results
558 results for “Training Data”
CNN models and training, validation and test datasets for "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"
<p>Convolutional neural network (CNN) models and their respective training, validation and test datasets used in manuscript:</p> <p>Tuomo Hartonen, Teemu Kivioja and Jussi Taipale, "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"</p>
Data for "Training data composition affects performance of protein structure analysis algorithms" by A. Derry, K. A. Carpenter, & R. B. Altman
<p><strong>Description</strong></p> <p>This repository contains all data used in "Training data composition affects performance of protein structure analysis algorithms", published in the Pacific Symposium on Biocomputing 2022 by A. Derry, K. A. Carpenter, & R. B. Altman. </p> <p>The data consists of the following files:</p> <ul> <li>ema_zenodo_data.tar.gz: train, validation, and test splits for Estimation of Model Accuracy task, in LMDB format</li> <li>design_zenodo_data.tar.gz: train, validation, and test splits for Protein Sequence Design task, in JSON format</li> <li>enz_cat_res_zenodo_data.tar.gz: train, validation, and test splits for Catalytic Residue and Enzyme Prediction task, in TF record format</li> </ul> <p>Details on dataset construction can be found in our paper and dataloaders can be found in our <a href="https://github.com/awfderry/ml-structure-bias">Github repo</a>.</p> <p><strong>Reference</strong></p> <p>A. Derry*, K. A. Carpenter*, & R. B. Altman, "Training data composition affects performance of protein structure analysis algorithms", 2021.</p> <p><strong>Dataset References</strong></p> <p>Datasets used were derived from the following works:</p> <p>Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K., & Moult, J. (2019). Critical assessment of methods of protein structure prediction (CASP)—Round XIII. In <em>Proteins: Structure, Function and Bioinformatics</em> (Vol. 87, Issue 12, pp. 1011–1020). https://doi.org/10.1002/prot.25823</p> <p>Ingraham, J., Garg, V. K., Barzilay, R., & Jaakkola, T. (2019). <em>Generative Models for Graph-Based Protein Design</em>. https://openreview.net/pdf?id=SJgxrLLKOE</p> <p>Furnham, N., Holliday, G. L., de Beer, T. A. P., Jacobsen, J. O. B., Pearson, W. R., & Thornton, J. M. (2014). The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. <em>Nucleic Acids Research</em>, <em>42 </em>(Database issue), D485–D489.</p>
2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)
<p>This dataset is part of the <a href="https://www.osgeo.org/foundation-news/2021-osgeo-un-committee-educational-challenge/?fbclid=IwAR0UvwkPO2pay7C0tJawb63eewjBGfeL9TIQpYUFccza9OIo6HAolmHXLWE">2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)</a>,</p> <p>Exercise 1: Supervised Change Detection: Monitoring deglaciation in Huascaran, Peru.</p>
Data Set on Content Excerpts from Relevant Literature for a Scoping Review of Evacuation Training Methods in Buildings
<p>This Excel-file contains a set of data from a scoping review on methods for fire evacuation training in buildings. The review follows the PRISMA approach (Transparent Reporting of Systematic Reviews and Meta-Analyses) and systematically identifies 73 sources among scientific literature published between 1997 and 2022. The dataset contains information excerpted through a custom template on the employed training methods and technology, study information, participants, and contents of the discussion of the 73 sources of evidence that were identified in the systematic review process.</p>
Discover the Data Archiving Guide (DAG) - a training event for new(ish) staff members [Workshop recording]
<p>The CESSDA Data Archiving Guide (https://dag.cessda.eu/) is a new resource developed by CESSDA and is designed to provide employees at data archives and repositories with an understanding of the work a data archive performs. The information in the DAG was collected by experts from CESSDA social science data archives reflecting the procedures and policies at their local archives. While the context of these archives varies — in size, the underlying technical architecture or in the specific services provided to researchers — the DAG focuses on common ground and is a useful tool for professionals new to data archiving or those who are knowledgeable in one domain and now seek to broaden their expertise.<br> <br> The full-day workshop was targeted mainly for new employees in data archives; people generally interested in the DAG were welcome as well.<br> <br> This workshop focused specifically on the Chapters Pre-Ingest, Ingest and FAIR with an additional excursion into the glossary to deepen participant's knowledge in a playful way.</p> <p>The video is also available on the <a href="https://www.youtube.com/watch?v=yzPzVK5UZKE">CESSDA Training YouTube Channel</a>.</p>
Data Set on the Literature Screening for a Scoping Review of Evacuation Training Methods in Buildings
<p>This data set contains all retrieved literature records of a scoping review on fire evacuation training methods in buildings together with the reasoning for their in- or exclusion in the review.</p>
Environmental data and fractional abundance of iso and branched GDGT data used to train the BIGMaC algorithm
<p>Location, environmental data -depth (m), elevation, distance to land (km), Mean Annual Air Temperature (C), and pH-, as well as fractional abundance of isoprenoid and branched GDGTs for unpublished samples used for the training of the Branched and Isoprenoid GDGT Machine learning Classification (BIGMaC) algorithm (Martínes-Sosa, et al., in prep).</p>
Immersive haptic simulation for training nurses in emergency medical procedures - Data collected and statistical analysis
<p>Data collected during the evaluation presented in "Haptic simulation for emergency procedures in nursing training" paper.</p> <table> <caption>HR ALL</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.857</td> <td>29</td> <td>0.008</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-8.089</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>7.567</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.962</td> <td>29</td> <td>0.006</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR FIRST MANN</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>1.665</td> <td>14</td> <td>0.118</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-7.104</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>6.498</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-1.461</td> <td>14</td> <td>0.166</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR FIRST VR</caption> <thead> <tr> <th>Measure 1</th> <th> </th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.341</td> <td>14</td> <td>0.035</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-4.612</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>4.482</td> <td>14</td> <td>< .001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.688</td> <td>14</td> <td>0.018</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Paired samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>HR BETWEEN GROUPS</caption> <thead> <tr> <th> </th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-1.958</td> <td>28</td> <td>0.060</td> </tr> <tr> <td>Mann post HR</td> <td>-1.902</td> <td>28</td> <td>0.068</td> </tr> <tr> <td>VR pre HR</td> <td>-4.013</td> <td>28</td> <td>< .001</td> </tr> <tr> <td>VR post HR</td> <td>-2.344</td> <td>28</td> <td>0.026</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em> Independent samples student's t-test.</td> </tr> </tbody> </table> <p> </p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the mannequin.</caption> <thead> <tr> <th>First variable</th> <th>μ</th> <th>σ</th> <th>Second variable</th> <th>μ</th> <th>σ</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>128.333</td> <td>10.715</td> <td>SBP pre-simulator</td> <td>134.533</td> <td>11.819</td> <td>-1.870</td> <td>14</td> <td>0.083</td> </tr> <tr> <td>SBP post-mannequin</td> <td>125.600</td> <td>11.648</td> <td>SBP post-simulator</td> <td>131.467</td> <td>14.643</td> <td>-2.094</td> <td>14</td> <td>0.055</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>80.133</td> <td>5.527</td> <td>DBP pre-simulator</td> <td>81.533</td> <td>9.039</td> <td>-0.623</td> <td>14</td> <td>0.544</td> </tr> <tr> <td>DBP post-mannequin</td> <td>78.667</td> <td>6.956</td> <td>DBP post-simulator</td> <td>81.400</td> <td>8.475</td> <td>-2.073</td> <td>14</td> <td>0.057</td> </tr> <tr> <td>HR pre-mannequin</td> <td>92.133</td> <td>14.837</td> <td>HR pre-simulator</td> <td>75.733</td> <td>9.9625</td> <td>6.498</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>87.400</td> <td>9.132</td> <td>HR post-simulator</td> <td>91.400</td> <td>14.217</td> <td>-1.461</td> <td>29</td> <td>0.166</td> </tr> </tbody> </table> <p>SBP = Systolic blood pressure. DBP = Diastolic blood pressure. HR = Heart Rate.</p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the ParaVR simulator.</caption> <thead> <tr> <th>First variable</th> <th>μ</th> <th>σ</th> <th>Second variable</th> <th>μ</th> <th>σ</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>119.067</td> <td>12.898</td> <td>SBP pre-simulator</td> <td>130.600</td> <td>12.188</td> <td>-3.799</td> <td>14</td> <td>0.002</td> </tr> <tr> <td>SBP post-mannequin</td> <td>117.533</td> <td>13.410</td> <td>SBP post-simulator</td> <td>128.200</td> <td>13.385</td> <td>-4.022</td> <td>14</td> <td>0.001</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>76.533</td> <td>8.943</td> <td>DBP pre-simulator</td> <td>80.200</td> <td>6.899</td> <td>-1.815</td> <td>14</td> <td>0.091</td> </tr> <tr> <td>DBP post-mannequin</td> <td>74.333</td> <td>8.541</td> <td>DBP post-simulator</td> <td>79.133</td> <td>7.864</td> <td>-2.003</td> <td>14</td> <td>0.065</td> </tr> <tr> <td>HR pre-mannequin</td> <td>102.067</td> <td>12.876</td> <td>HR pre-simulator</td> <td>91.533</td> <td>11.825</td> <td>4.482</td> <td>29</td> <td>< .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>95.867</td> <td>14.623</td> <td>HR post-simulator</td> <td>103.667</td> <td>14.450</td> <td>-2.688</td> <td>29</td> <td>0.018</td> </tr> </tbody> </table>
Datasets used to train the models in "Deep learning for denoising High-Rate Global Navigation Satellite System data."
<p>Datasets used to train the models in "Deep learning for denoising High-Rate Global Navigation Satellite System data." Additional information can be found at https://github.com/amtseismo/hrgnss_denoising.</p>
DeliCS Training+Validation Data - SPI-TGAS-MRF+GRE
<p>This data set consists of raw MRI k-space data from 12 healthy volunteers. The data were acquired on a 3T Premier MRI scanner (GE Healthcare, Waukesha, WI) with a 48-channel head receiver-coil. The raw data was saved as numpy-arrays to remove any potentially identifying meta-data, and to work in the reconstruction pipeline presented in [1]. </p> <p>SPI-TGAS-MRF (files named <strong>raw_mrf.npy</strong>):</p> <p>The acquisition consists of an initial adiabatic inversion pulse followed by a 500 TR long readout train (TI/TE/TR = 20/0.7/12ms) with varying flip angles (10 to 75 degrees) and a rotating 3D center-out spiral trajectory. 48 repeats of the TR train are used for a 6 min acquisition. Details available in [2]. The data shape is: (2000, 48, 24000) = (data along spiral readout, number of receive channels, number of spirals across 500 TR's and 48 repeats)</p> <p>GRE (files named <strong>raw_gre.npy</strong>):</p> <p>A 20 second, low resolution (6.9 mm isotropic) gradient echo (GRE) pre-scan with a large FOV of 440x440x440mm^3. The data shape is: (64, 48, 4096) = (data along readout, number of receive channels, number of phase encode lines (64x64))</p> <p>Noise estimation (files named <strong>noise.npy</strong>):</p> <p>Data from a noise scan acquired using all receive channels to calculate the noise coherence matrix. The data shape is: (48, 4096) = (number of receive channels, noise measurement points)</p> <p>To run the processing pipeline presented in [1], please follow the instructions on <a href="http://github.com/SetsompopLab/deli-cs">https://github.com/SetsompopLab/deli-cs</a> and download Zenodo datasets <a href="https://doi.org/10.5281/zenodo.7734431">10.5281/zenodo.7734431</a> and <a href="http://doi.org/10.5281/zenodo.7703200">10.5281/zenodo.7703200</a> with testing data and meta data needed for the reconstruction pipeline.</p> <p> </p> <p>[1] Iyer S, Schauman S, Sandino C, et al. Deep Learning Initialized Compressed Sensing (Deli-CS) in Volumetric Spatio-Temporal Subspace Reconstruction. <em>BioRxiv: </em><a href="https://www.biorxiv.org/content/10.1101/2023.03.28.534431v1">https://www.biorxiv.org/content/10.1101/2023.03.28.534431v1</a></p> <p>[2] Cao, X, Liao, C, Iyer, SS, et al. Optimized multi-axis spiral projection MR fingerprinting with subspace reconstruction for rapid whole-brain high-isotropic-resolution quantitative imaging. <em>Magn Reson Med</em>. 2022; 88: 133- 150. doi:<a href="https://doi.org/10.1002/mrm.29194">10.1002/mrm.29194</a></p>
Data Management Training Clearinghouse Metadata and Collection Statistics Report
<p>This collection contains a snapshot of the learning resource metadata from ESIP's <a href="https://dmtclearinghouse.esipfed.org">Data management Training Clearinghouse</a> (DMTC) associated with the closeout (March 30, 2023) of the Institute of Museum and Library Services funded (Award Number: <a href="https://imls.gov/grants/awarded/lg-70-18-0092-18">LG-70-18-0092-18</a>) <em>Development of an Enhanced and Expanded Data Management Training Clearinghouse project.</em> The shared metadata are a snapshot associated with the final reporting date for the project, and the associated data report is also based upon the same data snapshot on the same date.</p> <p>The materials included in the collection consist of the following:</p> <ul> <li><strong>esip-dev-02.edacnm.org.json.zip</strong> - a zip archive containing the metadata for 587 published learning resources as of March 30, 2023. These metadata include all publicly available metadata elements for the published learning resources with the exception of the metadata elements containing individual email addresses (submitter and contact) to reduce the exposure of these data.</li> <li><strong>statistics.pdf</strong> - an automatically generated report summarizing information about the collection of materials in the DMTC Clearinghouse, including both published and unpublished learning resources. This report includes the numbers of published and unpublished resources through time; the number of learning resources within subject categories and detailed subject categories, the dates items assigned to each category were first added to the Clearinghouse, and the most recent data that items were added to that category; the distribution of learning resources across target audiences; and the frequency of keywords within the learning resource collection. This report is based on the metadata for published resourced included in this collection, <strong>and</strong> preliminary metadata for unpublished learning resources that are not included in the shared dataset. </li> </ul> <p>The metadata fields consist of the following:</p> <table> <thead> <tr> <th scope="col">Fieldname</th> <th scope="col">Description</th> </tr> </thead> <tbody> <tr> <td>abstract_data</td> <td>A brief synopsis or abstract about the learning resource</td> </tr> <tr> <td>abstract_format</td> <td>Declaration for how the abstract description will be represented.</td> </tr> <tr> <td>access_conditions</td> <td>Conditions upon which the resource can be accessed beyond cost, e.g., login required.</td> </tr> <tr> <td>access_cost</td> <td>Yes or No choice stating whether othere is a fee for access to or use of the resource.</td> </tr> <tr> <td>accessibililty_features_name</td> <td>Content features of the resource, such as accessible media, alternatives and supported enhancements for accessibility.</td> </tr> <tr> <td>accessibililty_summary</td> <td>A human-readable summary of specific accessibility features or deficiencies.</td> </tr> <tr> <td>author_names</td> <td>List of authors for a resource derived from the given/first and family/last names of the personal author fields by the system</td> </tr> <tr> <td>author_org<br> - name<br> - name_identifier<br> - name_identifier_type</td> <td> <p><br> - Name of organization authoring the learning resource.<br> - The unique identifier for the organization authoring the resource.<br> - The identifier scheme associated with the unique identifier for the organization authoring the resource.</p> </td> </tr> <tr> <td> <p>authors<br> - givenName<br> - familyName<br> - name_identifier<br> - name_identifier_type</p> </td> <td> <p><br> - Given or first name of person(s) authoring the resource.<br> - Last or family name of person(s) authoring the resource.<br> - The unique identifier for the person(s) authoring the resource.<br> - The identifier scheme associated with the unique identifier for the person(s) authoring the resource, e.g., ORCID.</p> </td> </tr> <tr> <td>citation</td> <td>Preferred Form of Citation.</td> </tr> <tr> <td>completion_time</td> <td>Intended Time to Complete</td> </tr> <tr> <td> <p>contact<br> - name<br> - org<br> - email</p> </td> <td> <p><br> - Name of person(s) who has/have been asserted as the contact(s) for the resource in case of questions or follow-up by resource user.<br> - Name of organization that has/have been asserted as the contact(s) for the resource in case of questions or follow-up by resource user.<br> - (excluded) Contact email address.</p> </td> </tr> <tr> <td>contributor_orgs<br> - name<br> - name_identifier<br> - name_identifier_type<br> - type</td> <td>- Name of organization that is a secondary contributor to the learningresource. A contributor can also be an individual person.<br> - The unique identifier for the organization contributing to the resource.<br> - The identifier scheme associated with the unique identifier for the organization contributing to the resource.<br> - Type of contribution to the resource made by an organization.</td> </tr> <tr> <td>contributors<br> - familyName<br> - givenName<br> - name_identifier<br> - name_identifier_type</td> <td> <p>- Last or family name of person(s) contributing to the resource.<br> - Given or first name of person(s) contributing to the resource.<br> - The unique identifier for the person(s) contributing to the resource.<br> - The identifier scheme associated with the unique identifier for the person(s) contributing to the resource, e.g., ORCID.</p> </td> </tr> <tr> <td> <p>contributors.type</p> </td> <td> <p>Type of contribution to the resource made by a person.</p> </td> </tr> <tr> <td>created</td> <td>The date on which the metadata record was first saved as part of the input workflow.</td> </tr> <tr> <td>creator</td> <td>The name of the person creating the MD record for a resource.</td> </tr> <tr> <td>credential_status</td> <td>Declaration of whether a credential is offered for comopletion of the resource.</td> </tr> <tr> <td> <p>ed_frameworks<br> - name<br> - description<br> - nodes.name</p> </td> <td>- The name of the educational framework to which the resource is aligned, if any. An educational framework is a structured description of educational concepts such as a shared curriculum, syllabus or set of learning objectives, or a vocabulary for describing some other aspect of education such as educational levels or reading ability.<br> - A description of one or more subcategories of an educational framework to which a resource is associated.<br> - The name of a subcategory of an educational framework to which a resource is associated.</td> </tr> <tr> <td>expertise_level</td> <td>The skill level targeted for the topic being taught.</td> </tr> <tr> <td>id</td> <td>Unique identifier for the MD record generated by the system in UUID format.</td> </tr> <tr> <td>keywords</td> <td>Important phrases or words used to describe the resource.</td> </tr> <tr> <td>language_primary</td> <td>Original language in which the learning resource being described is published or made available.</td> </tr> <tr> <td>languages_secondary</td> <td>Additional languages in which the resource is tranlated or made available, if any.</td> </tr> <tr> <td>license</td> <td>A license for use of that applies to the resource, typically indicated by URL.</td> </tr> <tr> <td>locator_data</td> <td>The identifier for the learning resource used as part of a citation, if available.</td> </tr> <tr> <td>locator_type</td> <td>Designation of citation locatorr type, e.g., DOI, ARK, Handle.</td> </tr> <tr> <td>lr_outcomes</td> <td>Descriptions of what knowledge, skills or abilities students should learn from the resource.</td> </tr> <tr> <td>lr_type</td> <td>A characteristic that describes the predominant type or kind of learning resource.</td> </tr> <tr> <td>media_type</td> <td>Media type of resource.</td> </tr> <tr> <td>modification_date</td> <td>System generated date and time when MD record is modified.</td> </tr> <tr> <td>notes</td> <td>MD Record Input Notes</td> </tr> <tr> <td>pub_status</td> <td>Status of metadata record within the system, i.e., in-process, in-review, pre-pub-review, deprecate-request, deprecated or published.</td> </tr> <tr> <td>published</td> <td>Date of first broadcast / publication.</td> </tr> <tr> <td>publisher</td> <td>The organization credited with publishing or broadcasting the resource.</td> </tr> <tr> <td>purpose</td> <td>The purpose of the resource in the context of education; e.g., instruction, professional education, assessment.</td> </tr> <tr> <td>rating</td> <td>The aggregation of input from all user assessments evaluating users' reaction to the learning resource following Kirkpatrick's model of training evaluation.</td> </tr> <tr> <td>ratings</td> <td>Inputs from users assessing each user's reaction to the learning resource following Kirkpatrick's model of training evaluation.</td> </tr> <tr> <td>resource_modification_date</td> <td>Date in which the resource has last been modified from the original published or broadcast version.</td> </tr> <tr> <td>status</td> <td>System generated publication status of the resource w/in the registry as a yes for published or no for not published.</td> </tr> <tr> <td>subject</td> <td>Subject domain(s) toward which the resource is targeted. There may be more than one value for this field.</td> </tr> <tr> <td>submitter_email</td> <td>(excluded) Email address of person who submitted the resource.</td> </tr> <tr> <td>submitter_name</td> <td>Submission Contact Person</td> </tr> <tr> <td>target_audience</td> <td>Audience(s) for which the resource is intended.</td> </tr> <tr> <td>title</td> <td>The name of the resource.</td> </tr> <tr> <td>url</td> <td>URL that resolves to a downloadable version of the learning resource or to a landing page for the resource that contains important contextual information including the direct resolvable link to the resource, if applicable.</td> </tr> <tr> <td>usage_info</td> <td>Descriptive information about using the resource, not addressed by the License information field.</td> </tr> <tr> <td>version</td> <td>The specific version of the resource, if declared.</td> </tr> </tbody> </table> <p> </p>
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)
<p>Datasets from Main study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)
<p>Datasets from Pilot study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
VoroIF-GNN training, validation, and testing data
<p>Data used to train, validate, and test the VoroIF-GNN method described in the paper "VoroIF-GNN: Voronoi tessellation-derived protein-protein interface assessment using a graph neural network".</p>
Training Data for "Binning of metagenomic sequencing data" tutorial
<p><strong>Metagenomics is the study of genetic material recovered directly from environmental samples, such as soil, water, or gut contents, without the need for isolation or cultivation of individual organisms. Metagenomics binning is a process used to classify DNA sequences obtained from metagenomic sequencing into discrete groups, or bins, based on their similarity to each other</strong>. The goal of metagenomics binning is to assign the DNA sequences to the organisms or taxonomic groups that they originate from, allowing for a better understanding of the diversity and functions of the microbial communities present in the sample. This is typically achieved through computational methods that use sequence similarity, composition, and other features to group the sequences into bins.</p> <p>There are two main types of metagenomics binning: <strong>reference-based</strong> and <strong>de novo</strong>.</p> <ul> <li><strong>reference-based binning</strong> involves aligning the sequences to a database of known genomes or reference sequences</li> <li><strong>de novo binning</strong> involves clustering the sequences based on similarity without prior knowledge of the organisms or reference sequences present in the sample.</li> </ul> <p>Both methods have their strengths and limitations, and researchers often use a combination of approaches to improve the accuracy of their binning results. Metagenomics binning is an important tool for understanding the functional potential of microbial communities in various environments and has applications in fields such as biotechnology, environmental science, and human health.</p> <p>In this tutorial, we will learn how to run metagenomic binning tools and evaluate the quality of the results. In order to do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a> and MetaBAT2 algorithm. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>
Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value < 0.0001).</p>
Training data for the "Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints"
<p>There are two ZIP-files consisting of small histological image tiles that have been used to detect and quantify distinct tissue textures and lymphocyte proportions from H&E-stained clear cell renal cell carcinoma (KIRC) digital tissue sections of the Cancer Genome Atlas (TCGA) image archive and the Helsinki dataset.</p> <p>The <strong>tissue_classification </strong>file contains 300x300px tissue texture image tiles (n=52,713) representing renal cancer (“cancer”; n=13,057, 24.8%); normal renal (“normal”; n=8,652, 16.4%); stromal (“stroma”; n= 5,460, 10.4%) including smooth muscle, fibrous stroma and blood vessels; red blood cells (“blood”; n=996, 1.9%); empty background (“empty”; n=16,026, 30.4%); and other textures including necrotic, torn and adipose tissue (“other”; n=8,522, 16.2%). Image tiles have been randomly selected from the TCGA-KIRC WSI and the Helsinki datasets.</p> <p>The <strong>binary_lymphocytes </strong>file contains mostly 256x256px-sized but also smaller image tiles of Low (n=20,092, 80.1%) or High (n=5,003, 19.9%) lymphocyte density (n=25,095). Image tiles have been randomly selected from the TCGA-KIRC WSI dataset.</p> <p>All accuracy of all annotations have been double-checked. However, the classification between multiple tissue textures or lymphocyte density can be sometimes ambiguous.</p> <p>The deep learning model parameters trained with the ResNet-18 infrastructure for (1) lymphocyte and (2) texture classification are named as (1) <strong>resnet18_binary_lymphocytes.pth</strong> and (2) <strong>resnet18_tissue_classification.pth</strong>. Codes and instructions to use these are found in <a href="https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis">https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis</a>.</p> <p> </p> <p>If you use either work, please cite the publication by Brummer O et al (1) AND the TCGA Research Network (2):<br><strong>(1) </strong><strong>Brummer, O., Pölönen, P., Mustjoki, S. <em>et al.</em> Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints. <em>Br J Cancer</em> 129, 683–695 (2023). </strong><a href="https://doi.org/10.1038/s41416-023-02329-4">https://doi.org/10.1038/s41416-023-02329-4</a></p> <p><strong>(2) The results shown here are in whole or part based upon data generated by the TCGA Research Network: </strong><strong><a href="https://www.cancer.gov/tcga">https://www.cancer.gov/tcga</a></strong><strong>.</strong></p>
Neural net training and test data set
<p>In this zip file you will find several folders as well as a readme file explaining the dataset. The neural net file can be directly used in cellpose to segment <strong>widefield</strong> <strong>20x objective</strong> imaging data of neutrophils. The neural net could be used for other types of data, but might perform below expectations.</p>
U-T training data and test data for Sigsbee2A m odel
<p>Here are the training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal “Journal of Geophysical Research: Solid Earth”, named “Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density”: SigsbeeA model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>
U-T training and test data for Saltblock model
<p>Here are the training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal “Journal of Geophysical Research: Solid Earth”, named “Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density”: Saltblock model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.