Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo40/100

CNN models and training, validation and test datasets for "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"

<p>Convolutional neural network (CNN) models and their respective training, validation and test datasets used in manuscript:</p> <p>Tuomo Hartonen, Teemu Kivioja and Jussi Taipale, &quot;PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data&quot;</p>

opencc-by-4.0Mar 2021View details →
zenodo40/100

Data for "Training data composition affects performance of protein structure analysis algorithms" by A. Derry, K. A. Carpenter, & R. B. Altman

<p><strong>Description</strong></p> <p>This repository contains all data used in&nbsp;&quot;Training data composition affects performance of protein structure analysis algorithms&quot;, published in the Pacific Symposium on Biocomputing 2022 by A. Derry, K. A. Carpenter, &amp; R. B. Altman.&nbsp;</p> <p>The data consists of the following files:</p> <ul> <li>ema_zenodo_data.tar.gz: train, validation, and test&nbsp;splits for Estimation of Model Accuracy task, in LMDB format</li> <li>design_zenodo_data.tar.gz: train, validation, and test&nbsp;splits for Protein Sequence Design&nbsp;task, in JSON format</li> <li>enz_cat_res_zenodo_data.tar.gz:&nbsp;train, validation, and test&nbsp;splits for Catalytic Residue and Enzyme Prediction task, in TF record format</li> </ul> <p>Details on dataset construction can be found in our paper and dataloaders can be found in our&nbsp;<a href="https://github.com/awfderry/ml-structure-bias">Github repo</a>.</p> <p><strong>Reference</strong></p> <p>A. Derry*, K. A. Carpenter*, &amp; R. B. Altman, &quot;Training data composition affects performance of protein structure analysis algorithms&quot;, 2021.</p> <p><strong>Dataset References</strong></p> <p>Datasets used were derived from the following works:</p> <p>Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K., &amp; Moult, J. (2019). Critical assessment of methods of protein structure prediction (CASP)&mdash;Round XIII. In <em>Proteins: Structure, Function and Bioinformatics</em> (Vol. 87, Issue 12, pp. 1011&ndash;1020). https://doi.org/10.1002/prot.25823</p> <p>Ingraham, J., Garg, V. K., Barzilay, R., &amp; Jaakkola, T. (2019). <em>Generative Models for Graph-Based Protein Design</em>. https://openreview.net/pdf?id=SJgxrLLKOE</p> <p>Furnham, N., Holliday, G. L., de Beer, T. A. P., Jacobsen, J. O. B., Pearson, W. R., &amp; Thornton, J. M. (2014). The Catalytic Site Atlas 2.0: cataloging catalytic sites and residues identified in enzymes. <em>Nucleic Acids Research</em>, <em>42&nbsp;</em>(Database issue), D485&ndash;D489.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)

<p>This dataset is part of the&nbsp;<a href="https://www.osgeo.org/foundation-news/2021-osgeo-un-committee-educational-challenge/?fbclid=IwAR0UvwkPO2pay7C0tJawb63eewjBGfeL9TIQpYUFccza9OIo6HAolmHXLWE">2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)</a>,</p> <p>Exercise 1:&nbsp;Supervised Change Detection: Monitoring deglaciation in Huascaran, Peru.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Data Set on Content Excerpts from Relevant Literature for a Scoping Review of Evacuation Training Methods in Buildings

<p>This Excel-file contains a set of data&nbsp;from a scoping review on methods for fire evacuation training in buildings. The review&nbsp;follows the PRISMA approach (Transparent Reporting of Systematic Reviews and Meta-Analyses) and systematically identifies 73 sources among scientific literature published between 1997 and 2022. The dataset contains information excerpted through a custom template&nbsp;on the employed training methods and technology, study information, participants, and contents of the discussion of the 73 sources of evidence that were identified in the systematic review process.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Discover the Data Archiving Guide (DAG) - a training event for new(ish) staff members [Workshop recording]

<p>The CESSDA Data Archiving Guide (https://dag.cessda.eu/) is a new resource developed by CESSDA and is designed to provide employees at data archives and repositories with an understanding of the work a data archive performs. The information in the DAG was collected by experts from CESSDA social science data archives reflecting the procedures and policies at their local archives. While the context of these archives varies &mdash; in size, the underlying technical architecture or in the specific services provided to researchers &mdash; the DAG focuses on common ground and is a useful tool for professionals new to data archiving or those who are knowledgeable in one domain and now seek to broaden their expertise.<br> <br> The full-day workshop was targeted mainly for new employees in data archives; people generally interested in the DAG were welcome as well.<br> <br> This workshop focused specifically on the Chapters Pre-Ingest, Ingest and FAIR with an additional excursion into the glossary to deepen participant&#39;s knowledge in a playful way.</p> <p>The video is also available on the <a href="https://www.youtube.com/watch?v=yzPzVK5UZKE">CESSDA Training YouTube Channel</a>.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Data Set on the Literature Screening for a Scoping Review of Evacuation Training Methods in Buildings

<p>This data set contains all retrieved literature records of a scoping review on fire evacuation training methods in buildings together with the reasoning for their in- or exclusion in the review.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Environmental data and fractional abundance of iso and branched GDGT data used to train the BIGMaC algorithm

<p>Location, environmental data -depth (m), elevation, distance to land (km), Mean Annual Air Temperature (C), and pH-, as well as fractional abundance of isoprenoid and branched GDGTs for unpublished samples used for the training of the Branched and Isoprenoid GDGT Machine learning Classification (BIGMaC) algorithm (Mart&iacute;nes-Sosa, et al., in prep).</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Immersive haptic simulation for training nurses in emergency medical procedures - Data collected and statistical analysis

<p>Data collected during the evaluation presented in &quot;Haptic simulation for emergency procedures in nursing training&quot; paper.</p> <table> <caption>HR ALL</caption> <thead> <tr> <th>Measure 1</th> <th>&nbsp;</th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.857</td> <td>29</td> <td>0.008</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-8.089</td> <td>29</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>7.567</td> <td>29</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.962</td> <td>29</td> <td>0.006</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em>&nbsp; Paired samples student&#39;s t-test.</td> </tr> </tbody> </table> <p>&nbsp;</p> <table> <caption>HR FIRST MANN</caption> <thead> <tr> <th>Measure 1</th> <th>&nbsp;</th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>1.665</td> <td>14</td> <td>0.118</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-7.104</td> <td>14</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>6.498</td> <td>14</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-1.461</td> <td>14</td> <td>0.166</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em>&nbsp; Paired samples student&#39;s t-test.</td> </tr> </tbody> </table> <p>&nbsp;</p> <table> <caption>HR FIRST VR</caption> <thead> <tr> <th>Measure 1</th> <th>&nbsp;</th> <th>Measure 2</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-</td> <td>Mann post HR</td> <td>2.341</td> <td>14</td> <td>0.035</td> </tr> <tr> <td>VR pre HR</td> <td>-</td> <td>VR post HR</td> <td>-4.612</td> <td>14</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann pre HR</td> <td>-</td> <td>VR pre HR</td> <td>4.482</td> <td>14</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>Mann post HR</td> <td>-</td> <td>VR post HR</td> <td>-2.688</td> <td>14</td> <td>0.018</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em>&nbsp; Paired samples student&#39;s t-test.</td> </tr> </tbody> </table> <p>&nbsp;</p> <table> <caption>HR BETWEEN GROUPS</caption> <thead> <tr> <th>&nbsp;</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>Mann pre HR</td> <td>-1.958</td> <td>28</td> <td>0.060</td> </tr> <tr> <td>Mann post HR</td> <td>-1.902</td> <td>28</td> <td>0.068</td> </tr> <tr> <td>VR pre HR</td> <td>-4.013</td> <td>28</td> <td>&lt;&nbsp;.001</td> </tr> <tr> <td>VR post HR</td> <td>-2.344</td> <td>28</td> <td>0.026</td> </tr> <tr> </tr> </tbody> <tbody> <tr> <td><em>Note.</em>&nbsp; Independent samples student&#39;s t-test.</td> </tr> </tbody> </table> <p>&nbsp;</p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the mannequin.</caption> <thead> <tr> <th>First variable</th> <th>&mu;</th> <th>&sigma;</th> <th>Second variable</th> <th>&mu;</th> <th>&sigma;</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>128.333</td> <td>10.715</td> <td>SBP pre-simulator</td> <td>134.533</td> <td>11.819</td> <td>-1.870</td> <td>14</td> <td>0.083</td> </tr> <tr> <td>SBP post-mannequin</td> <td>125.600</td> <td>11.648</td> <td>SBP post-simulator</td> <td>131.467</td> <td>14.643</td> <td>-2.094</td> <td>14</td> <td>0.055</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>80.133</td> <td>5.527</td> <td>DBP pre-simulator</td> <td>81.533</td> <td>9.039</td> <td>-0.623</td> <td>14</td> <td>0.544</td> </tr> <tr> <td>DBP post-mannequin</td> <td>78.667</td> <td>6.956</td> <td>DBP post-simulator</td> <td>81.400</td> <td>8.475</td> <td>-2.073</td> <td>14</td> <td>0.057</td> </tr> <tr> <td>HR pre-mannequin</td> <td>92.133</td> <td>14.837</td> <td>HR pre-simulator</td> <td>75.733</td> <td>9.9625</td> <td>6.498</td> <td>29</td> <td>&lt; .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>87.400</td> <td>9.132</td> <td>HR post-simulator</td> <td>91.400</td> <td>14.217</td> <td>-1.461</td> <td>29</td> <td>0.166</td> </tr> </tbody> </table> <p>SBP = Systolic blood pressure. DBP = Diastolic blood pressure. HR = Heart Rate.</p> <table> <caption>Physiological T-Test results for the participants who started the experiment performing the procedure in the ParaVR simulator.</caption> <thead> <tr> <th>First variable</th> <th>&mu;</th> <th>&sigma;</th> <th>Second variable</th> <th>&mu;</th> <th>&sigma;</th> <th>t</th> <th>df</th> <th>p</th> </tr> </thead> <tbody> <tr> <td>SBP pre-mannequin</td> <td>119.067</td> <td>12.898</td> <td>SBP pre-simulator</td> <td>130.600</td> <td>12.188</td> <td>-3.799</td> <td>14</td> <td>0.002</td> </tr> <tr> <td>SBP post-mannequin</td> <td>117.533</td> <td>13.410</td> <td>SBP post-simulator</td> <td>128.200</td> <td>13.385</td> <td>-4.022</td> <td>14</td> <td>0.001</td> </tr> <tr> <td>DBP pre-mannequin</td> <td>76.533</td> <td>8.943</td> <td>DBP pre-simulator</td> <td>80.200</td> <td>6.899</td> <td>-1.815</td> <td>14</td> <td>0.091</td> </tr> <tr> <td>DBP post-mannequin</td> <td>74.333</td> <td>8.541</td> <td>DBP post-simulator</td> <td>79.133</td> <td>7.864</td> <td>-2.003</td> <td>14</td> <td>0.065</td> </tr> <tr> <td>HR pre-mannequin</td> <td>102.067</td> <td>12.876</td> <td>HR pre-simulator</td> <td>91.533</td> <td>11.825</td> <td>4.482</td> <td>29</td> <td>&lt; .001</td> </tr> <tr> <td>HR post-mannequin</td> <td>95.867</td> <td>14.623</td> <td>HR post-simulator</td> <td>103.667</td> <td>14.450</td> <td>-2.688</td> <td>29</td> <td>0.018</td> </tr> </tbody> </table>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Datasets used to train the models in "Deep learning for denoising High-Rate Global Navigation Satellite System data."

<p>Datasets used to train the models in &quot;Deep learning for denoising High-Rate Global Navigation Satellite System data.&quot;&nbsp; Additional information can be found at&nbsp;https://github.com/amtseismo/hrgnss_denoising.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

DeliCS Training+Validation Data - SPI-TGAS-MRF+GRE

<p>This data set consists of raw MRI k-space data from 12 healthy volunteers.&nbsp;The data were acquired on a 3T Premier MRI scanner (GE Healthcare, Waukesha, WI) with a 48-channel head receiver-coil. The raw data was saved as numpy-arrays to remove any potentially identifying meta-data, and to work in the reconstruction pipeline presented in [1].&nbsp;</p> <p>SPI-TGAS-MRF (files named <strong>raw_mrf.npy</strong>):</p> <p>The acquisition consists of an initial adiabatic inversion pulse followed by a 500 TR long readout train (TI/TE/TR = 20/0.7/12ms) with varying flip angles (10 to 75 degrees) and a rotating 3D center-out spiral trajectory. 48 repeats of the TR train are used for a 6 min acquisition. Details available in [2]. The data shape is: (2000, 48, 24000) = (data along spiral readout, number of receive channels, number of spirals across 500 TR&#39;s and 48 repeats)</p> <p>GRE&nbsp;(files named <strong>raw_gre.npy</strong>):</p> <p>A 20 second, low resolution (6.9 mm isotropic) gradient echo (GRE)&nbsp;pre-scan with a large FOV of 440x440x440mm^3. The data shape is: (64, 48, 4096) = (data along readout, number of receive channels, number of phase encode lines (64x64))</p> <p>Noise estimation (files named <strong>noise.npy</strong>):</p> <p>Data from a noise scan acquired using all receive channels to calculate the noise coherence matrix. The data shape is: (48, 4096) = (number of receive channels, noise measurement points)</p> <p>To run the processing pipeline presented in [1], please follow the instructions on <a href="http://github.com/SetsompopLab/deli-cs">https://github.com/SetsompopLab/deli-cs</a>&nbsp; and download Zenodo datasets <a href="https://doi.org/10.5281/zenodo.7734431">10.5281/zenodo.7734431</a> and&nbsp;<a href="http://doi.org/10.5281/zenodo.7703200">10.5281/zenodo.7703200</a>&nbsp;with testing data and meta data needed for the reconstruction pipeline.</p> <p>&nbsp;</p> <p>[1] Iyer S, Schauman S, Sandino C, et al.&nbsp;Deep Learning Initialized Compressed Sensing (Deli-CS) in Volumetric Spatio-Temporal Subspace Reconstruction.&nbsp;<em>BioRxiv:&nbsp;</em><a href="https://www.biorxiv.org/content/10.1101/2023.03.28.534431v1">https://www.biorxiv.org/content/10.1101/2023.03.28.534431v1</a></p> <p>[2]&nbsp;Cao, X,&nbsp;&nbsp;Liao, C,&nbsp;&nbsp;Iyer, SS, et al.&nbsp;&nbsp;Optimized multi-axis spiral projection MR fingerprinting with subspace reconstruction for rapid whole-brain high-isotropic-resolution quantitative imaging.&nbsp;<em>Magn Reson Med</em>.&nbsp;2022;&nbsp;88:&nbsp;133-&nbsp;150. doi:<a href="https://doi.org/10.1002/mrm.29194">10.1002/mrm.29194</a></p>

openbsd-licenseMar 2023View details →
zenodo40/100

Data Management Training Clearinghouse Metadata and Collection Statistics Report

<p>This collection contains a snapshot of the learning resource metadata from ESIP&#39;s <a href="https://dmtclearinghouse.esipfed.org">Data management Training Clearinghouse</a> (DMTC) associated with the closeout (March 30, 2023) of the Institute of Museum and Library Services funded (Award Number: <a href="https://imls.gov/grants/awarded/lg-70-18-0092-18">LG-70-18-0092-18</a>) <em>Development of an Enhanced and Expanded Data Management Training Clearinghouse project.</em> The shared metadata are a snapshot associated with the final reporting date for the project, and the associated data report is also based upon the same data snapshot on the same date.</p> <p>The materials included in the collection consist of the following:</p> <ul> <li><strong>esip-dev-02.edacnm.org.json.zip</strong> - a zip archive containing the metadata for 587 published learning resources as of March 30, 2023. These metadata include all publicly available metadata elements for the published learning resources with the exception of the metadata elements containing individual email addresses (submitter and contact) to reduce the exposure of these data.</li> <li><strong>statistics.pdf</strong> - an automatically generated report summarizing information about the collection of materials in the DMTC Clearinghouse, including both published and unpublished learning resources. This report includes the numbers of published and unpublished resources through time; the number of learning resources within subject categories and detailed subject categories, the dates items assigned to each category were first added to the Clearinghouse, and the most recent data that items were added to that category; the distribution of learning resources across target audiences; and the frequency of keywords within the learning resource collection. This report is based on the metadata for published resourced included in this collection, <strong>and</strong> preliminary metadata for unpublished learning resources that are not included in the shared dataset.&nbsp;</li> </ul> <p>The metadata fields consist of the following:</p> <table> <thead> <tr> <th scope="col">Fieldname</th> <th scope="col">Description</th> </tr> </thead> <tbody> <tr> <td>abstract_data</td> <td>A brief synopsis or abstract about the learning resource</td> </tr> <tr> <td>abstract_format</td> <td>Declaration for how the abstract description will be represented.</td> </tr> <tr> <td>access_conditions</td> <td>Conditions upon which the resource can be accessed beyond cost, e.g., login required.</td> </tr> <tr> <td>access_cost</td> <td>Yes or No choice stating whether othere is a fee for access to or use of the resource.</td> </tr> <tr> <td>accessibililty_features_name</td> <td>Content features of the resource, such as accessible media, alternatives and supported enhancements for accessibility.</td> </tr> <tr> <td>accessibililty_summary</td> <td>A human-readable summary of specific accessibility features or deficiencies.</td> </tr> <tr> <td>author_names</td> <td>List of authors for a resource derived from the given/first and family/last names of the personal author fields by the system</td> </tr> <tr> <td>author_org<br> - name<br> - name_identifier<br> - name_identifier_type</td> <td> <p><br> - Name of organization authoring the learning resource.<br> - The unique identifier for the organization authoring the resource.<br> - The identifier scheme associated with the unique identifier for the organization authoring the resource.</p> </td> </tr> <tr> <td> <p>authors<br> - givenName<br> - familyName<br> - name_identifier<br> - name_identifier_type</p> </td> <td> <p><br> - Given or first name of person(s) authoring the resource.<br> - Last or family name of person(s) authoring the resource.<br> - The unique identifier for the person(s) authoring the resource.<br> - The identifier scheme associated with the unique identifier for the person(s) authoring the resource, e.g., ORCID.</p> </td> </tr> <tr> <td>citation</td> <td>Preferred Form of Citation.</td> </tr> <tr> <td>completion_time</td> <td>Intended Time to Complete</td> </tr> <tr> <td> <p>contact<br> - name<br> - org<br> - email</p> </td> <td> <p><br> - Name of person(s) who has/have been asserted as the contact(s) for the resource in case of questions or follow-up by resource user.<br> - Name of organization that has/have been asserted as the contact(s) for the resource in case of questions or follow-up by resource user.<br> - (excluded) Contact email address.</p> </td> </tr> <tr> <td>contributor_orgs<br> - name<br> - name_identifier<br> - name_identifier_type<br> - type</td> <td>- Name of organization that is a secondary contributor to the learningresource.&nbsp; A contributor can also be an individual person.<br> - The unique identifier for the organization contributing to the resource.<br> - The identifier scheme associated with the unique identifier for the organization contributing to the resource.<br> - Type of contribution to the resource made by an organization.</td> </tr> <tr> <td>contributors<br> - familyName<br> - givenName<br> - name_identifier<br> - name_identifier_type</td> <td> <p>- Last or family name of person(s) contributing to the resource.<br> - Given or first name of person(s) contributing to the resource.<br> - The unique identifier for the person(s) contributing to the resource.<br> - The identifier scheme associated with the unique identifier for the person(s) contributing to the resource, e.g., ORCID.</p> </td> </tr> <tr> <td> <p>contributors.type</p> </td> <td> <p>Type of contribution to the resource made by a person.</p> </td> </tr> <tr> <td>created</td> <td>The date on which the metadata record was first saved as part of the input workflow.</td> </tr> <tr> <td>creator</td> <td>The name of the person creating the MD record for a resource.</td> </tr> <tr> <td>credential_status</td> <td>Declaration of whether a credential is offered for comopletion of the resource.</td> </tr> <tr> <td> <p>ed_frameworks<br> - name<br> - description<br> - nodes.name</p> </td> <td>- The name of the educational framework to which the resource is aligned, if any.&nbsp; An educational framework is a structured description of educational concepts such as a shared curriculum, syllabus or set of learning objectives, or a vocabulary for describing some other aspect of education such as educational levels or reading ability.<br> - A description of one or more subcategories of an educational framework to which a resource is associated.<br> - The name of a subcategory of an educational framework to which a resource is associated.</td> </tr> <tr> <td>expertise_level</td> <td>The skill level targeted for the topic being taught.</td> </tr> <tr> <td>id</td> <td>Unique identifier for the MD record generated by the system in UUID format.</td> </tr> <tr> <td>keywords</td> <td>Important phrases or words used to describe the resource.</td> </tr> <tr> <td>language_primary</td> <td>Original language in which the learning resource being described is published or made available.</td> </tr> <tr> <td>languages_secondary</td> <td>Additional languages in which the resource is tranlated or made available, if any.</td> </tr> <tr> <td>license</td> <td>A license for use of that applies to the resource, typically indicated by URL.</td> </tr> <tr> <td>locator_data</td> <td>The identifier for the learning resource used as part of a citation, if available.</td> </tr> <tr> <td>locator_type</td> <td>Designation of citation locatorr type, e.g., DOI, ARK, Handle.</td> </tr> <tr> <td>lr_outcomes</td> <td>Descriptions of what knowledge, skills or abilities students should learn from the resource.</td> </tr> <tr> <td>lr_type</td> <td>A characteristic that describes the predominant type or kind of learning resource.</td> </tr> <tr> <td>media_type</td> <td>Media type of resource.</td> </tr> <tr> <td>modification_date</td> <td>System generated date and time when MD record is modified.</td> </tr> <tr> <td>notes</td> <td>MD Record Input Notes</td> </tr> <tr> <td>pub_status</td> <td>Status of metadata record within the system, i.e., in-process, in-review, pre-pub-review, deprecate-request, deprecated or published.</td> </tr> <tr> <td>published</td> <td>Date of first broadcast / publication.</td> </tr> <tr> <td>publisher</td> <td>The organization credited with publishing or broadcasting the resource.</td> </tr> <tr> <td>purpose</td> <td>The purpose of the resource in the context of education; e.g., instruction, professional education, assessment.</td> </tr> <tr> <td>rating</td> <td>The aggregation of input from all user assessments evaluating&nbsp; users&#39; reaction to the learning resource following Kirkpatrick&#39;s model of training evaluation.</td> </tr> <tr> <td>ratings</td> <td>Inputs from users assessing each user&#39;s reaction to the learning resource following Kirkpatrick&#39;s model of training evaluation.</td> </tr> <tr> <td>resource_modification_date</td> <td>Date in which the resource has last been modified from the original published or broadcast version.</td> </tr> <tr> <td>status</td> <td>System generated publication status of the resource w/in the registry as a yes for published or no for not published.</td> </tr> <tr> <td>subject</td> <td>Subject domain(s) toward which the resource is&nbsp; targeted. There may be more than one value for this field.</td> </tr> <tr> <td>submitter_email</td> <td>(excluded) Email address of person who submitted the resource.</td> </tr> <tr> <td>submitter_name</td> <td>Submission Contact Person</td> </tr> <tr> <td>target_audience</td> <td>Audience(s) for which the resource is intended.</td> </tr> <tr> <td>title</td> <td>The name of the resource.</td> </tr> <tr> <td>url</td> <td>URL that resolves to a downloadable version of the learning resource or to a landing page for the resource that contains important contextual information including the direct resolvable link to the resource, if applicable.</td> </tr> <tr> <td>usage_info</td> <td>Descriptive information about using the resource, not addressed by the License information field.</td> </tr> <tr> <td>version</td> <td>The specific version of the resource, if declared.</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)

<p>Datasets from Main&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)

<p>Datasets from Pilot&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

VoroIF-GNN training, validation, and testing data

<p>Data used to train, validate, and test the VoroIF-GNN method described in the paper &quot;VoroIF-GNN: Voronoi tessellation-derived protein-protein interface assessment using a graph neural network&quot;.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Training Data for "Binning of metagenomic sequencing data" tutorial

<p><strong>Metagenomics is the study of genetic material recovered directly from environmental samples, such as soil, water, or gut contents, without the need for isolation or cultivation of individual organisms. Metagenomics binning is a process used to classify DNA sequences obtained from metagenomic sequencing into discrete groups, or bins, based on their similarity to each other</strong>. The goal of metagenomics binning is to assign the DNA sequences to the organisms or taxonomic groups that they originate from, allowing for a better understanding of the diversity and functions of the microbial communities present in the sample. This is typically achieved through computational methods that use sequence similarity, composition, and other features to group the sequences into bins.</p> <p>There are two main types of metagenomics binning: <strong>reference-based</strong>&nbsp;and <strong>de novo</strong>.</p> <ul> <li><strong>reference-based binning</strong>&nbsp;involves aligning the sequences to a database of known genomes or reference sequences</li> <li><strong>de novo binning</strong>&nbsp;involves clustering the sequences based on similarity without prior knowledge of the organisms or reference sequences present in the sample.</li> </ul> <p>Both methods have their strengths and limitations, and researchers often use a combination of approaches to improve the accuracy of their binning results. Metagenomics binning is an important tool for understanding the functional potential of microbial communities in various environments and has applications in fields such as biotechnology, environmental science, and human health.</p> <p>In this tutorial, we will learn how to run metagenomic binning tools and evaluate the quality of the results. In order to do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a>&nbsp;and MetaBAT2 algorithm. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Training data for 'Genome annotation with Funannotate' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with funannotate.</p> <p>Genome was assembled following the GTN Flye assembly tutorial, then masked with RepeatMasker.</p> <p>RNASeq data: SRR8534859 reads were mapped to the genome using STAR (toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.8a+galaxy0), then the bam was downsampled (10% with toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_DownsampleSam/2.18.2.1) to reduce the size of the dataset. Fastq files were then extracted from the resulting bam file (toolshed.g2.bx.psu.edu/repos/devteam/picard/picard_SamToFastq/2.18.2.1).</p> <p>SwissProt_subset.fasta is a subset of SwissProt proteins that are known to have some similarity with the genome (found using Diamond against the genome, then extracting sequences matching with e-value &lt; 0.0001).</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Training data for the "Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints"

<p>There are two ZIP-files consisting of small histological image tiles that have been used to detect and quantify distinct tissue textures and lymphocyte proportions from&nbsp;H&amp;E-stained clear cell renal cell carcinoma (KIRC)&nbsp;digital tissue sections of the Cancer Genome Atlas (TCGA) image archive and the Helsinki&nbsp;dataset.</p> <p>The <strong>tissue_classification </strong>file contains 300x300px tissue texture image tiles (n=52,713) representing renal cancer (&ldquo;cancer&rdquo;; n=13,057, 24.8%); normal renal (&ldquo;normal&rdquo;; n=8,652, 16.4%); stromal (&ldquo;stroma&rdquo;; n= 5,460, 10.4%) including smooth muscle, fibrous stroma and blood vessels; red blood cells (&ldquo;blood&rdquo;; n=996, 1.9%); empty background (&ldquo;empty&rdquo;; n=16,026, 30.4%); and other textures including necrotic, torn and adipose tissue (&ldquo;other&rdquo;; n=8,522, 16.2%). Image tiles have been randomly selected from the TCGA-KIRC WSI and the Helsinki datasets.</p> <p>The <strong>binary_lymphocytes </strong>file contains mostly 256x256px-sized but also smaller image tiles of Low (n=20,092, 80.1%) or High (n=5,003, 19.9%) lymphocyte density (n=25,095). Image tiles have been randomly selected from the TCGA-KIRC WSI dataset.</p> <p>All accuracy of all annotations have been double-checked. However, the classification between multiple tissue textures or lymphocyte density can be sometimes ambiguous.</p> <p>The deep learning model parameters&nbsp;trained with the ResNet-18 infrastructure for (1) lymphocyte and (2) texture classification are named as (1)&nbsp;<strong>resnet18_binary_lymphocytes.pth</strong>&nbsp;and (2)&nbsp;<strong>resnet18_tissue_classification.pth</strong>. Codes and instructions to use these are found in&nbsp;<a href="https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis">https://github.com/vahvero/RCC_textures_and_lymphocytes_publication_image_analysis</a>.</p> <p>&nbsp;</p> <p>If you use either work, please cite the publication by Brummer O et al (1) AND the TCGA Research Network (2):<br><strong>(1) </strong><strong>Brummer, O., P&ouml;l&ouml;nen, P., Mustjoki, S.&nbsp;<em>et al.</em>&nbsp;Computational textural mapping harmonises sampling variation and reveals multidimensional histopathological fingerprints.&nbsp;<em>Br J Cancer</em>&nbsp;129, 683&ndash;695 (2023). </strong><a href="https://doi.org/10.1038/s41416-023-02329-4">https://doi.org/10.1038/s41416-023-02329-4</a></p> <p><strong>(2) The results shown here are in whole or part based upon data generated by the TCGA Research Network: </strong><strong><a href="https://www.cancer.gov/tcga">https://www.cancer.gov/tcga</a></strong><strong>.</strong></p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Neural net training and test data set

<p>In this zip file you will find several folders as well&nbsp;as a readme file explaining the dataset. The neural net file can be directly used in cellpose to segment&nbsp;<strong>widefield</strong>&nbsp;<strong>20x objective</strong>&nbsp;imaging data of neutrophils. The neural net could be used for other types of data, but might perform below expectations.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

U-T training data and test data for Sigsbee2A m odel

<p>Here are the&nbsp;training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal &ldquo;Journal of Geophysical Research: Solid Earth&rdquo;, named &ldquo;Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density&rdquo;:&nbsp; SigsbeeA model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

U-T training and test data for Saltblock model

<p>Here are the&nbsp; training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal &ldquo;Journal of Geophysical Research: Solid Earth&rdquo;, named &ldquo;Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density&rdquo;: Saltblock model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record