Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,688
datasets available to search
ShareScore release 0.7.1
Dataset results
3,688 results for “Computer”
arXiv abstracts and titles from 1,469 single-authored papers (100 unique authors) in computer science
<p>This dataset is meant to be used for experiments of Authorship Analysis. The dataset consists of abstracts of single-author papers from arXiv crawled using the arXiv's API by querying a list of computer-science-related keywords ("deep learning", "machine learning", "information retrieval", "computer science", "data mining", "support vector", "logistic regression", "artificial intelligence", "supervised learning"'). The corpus somehow follows a power-law distribution, with few prolific authors and many authors accounting for very few papers each: we retained authors with at least 10 papers, resulting in a total of 1,469 documents from 100 authors. The most prolific authors (Peter D. Turney and Subhash Kak) have 34 abstracts to their names, the 10 most prolific authors have written 22 or more articles, while 50% of the authors have no more than 12 abstracts to their names. In order to divide the corpus into a training set and a test set we perform a stratified split, with the production of each author being split into a training set (70%) and a test set (30%). We use these documents as examples of "scientific communication", characterised by a precise and compact style, with an abundance of technical terminology.</p>
COG-BCI database: A multi-session and multi-task EEG cognitive dataset for passive brain-computer interfaces
<p>Brain-Computer Interfaces, and especially passive Brain-Computer Interfaces (pBCI), with their ability to estimate and detect mental states, are receiving increasing attention from both the scientific and the research and development communities. Many pBCIs aim to increase the safety of complex work environments such as in the aeronautical domain. Therefore, mental workload, vigilance and decision-making are some of the most commonly examined aspects of cognition within this field of research. A large proportion of pBCIs involve a component of machine learning and signal processing as the data that are collected need to be transformed into a reliable estimate of the users’ current mental state (e.g. mental workload). Improving this component is a major challenge for researchers, requiring large quantities of data. While data sharing is common for the active BCI community, open pBCI datasets are scarcer and generally incomplete with regards to the information they report. This is particularly true for datasets encompassing several tasks or sessions, which are of importance for tackling the challenges of transfer learning. Testing new pipelines, feature extraction algorithms and classifiers are central issues for future advances in research within this domain, as well as for algorithm benchmark and research reproducibility.The COG-BCI database presented here is comprised of the recordings of 29 participants over 3 individual sessions with 4 different tasks designed to elicit different cognitive states. This results in a total of over 100 hours of open electrophysiological (EEG) and electrocardiogram (ECG) data. The project was validated by the local ethical committee of the University of Toulouse (CER number 2021-342). The dataset was validated on a subjective, behavioral and physiological level (i.e. cardiac and cerebral activity), to ensure its usefulness to the pBCI community. This body of work represents a large effort to promote the use of pBCIs, as well as the use of open science.</p> <p> </p> <p><strong>The data are in the Brain Imaging Data Structure (BIDS) format. For more information, please read the COG-BCI_info.pdf file.</strong></p> <p><strong>Please note that version 4 corrected an electrode name mismatch, for which we sincerely apologize. The answers to the RSME and KSS questionnaires are provided in two separate .txt files.</strong></p>
Computed surface and chemical potentials, expansion coefficients, structures, models and results for the PMFPredictor Toolkit
<p>The PMFPredictor toolkit enables the prediction of the potentials of mean force describing the interaction between a surface and a small molecule in aqueous solution, which would otherwise be obtained from lengthy metadynamics simulations. This repository contains files to enable the operation of the toolkit, with source code available at https://github.com/ijrouse/PMFPredictor-Toolkit and corresponding to release v0.5-alpha.</p> <p>In PMFPredictor-Repository.zip we provide supplementary data necessary for the operation of the PMFPredictor Toolkit including:</p> <ul> <li>Structures of surfaces ("Structures/Surfaces") and chemicals ("Structures/Chemicals") in a united tabulated (.csv) format, listing x/y/z co-ordinates, atom IDs, mass (in amu), charge (in elementary units), Lennard Jones 6-12 parameters: sigma (in nm) and epsilon (in kJ/mol).</li> <li>Interaction potentials of surfaces ("SurfacePotentials") and chemicals ("ChemicalPotentials") with probe atoms and molecules in tabulated format with distances relative to reference points in nm and energies in kJ/mol. Also included in these folders are the potentials with the molecular probes in individual files.</li> <li>Hypergeometric expansion coefficients of the interaction potentials ("Datasets/SurfacePotentialCoefficientsNoise-1-oct12.csv" and "Datasets/ChemicalPotentialCoefficients-oct10.csv") in tabulated form, corresponding to potentials with units of nm for distance and kJ/mol for energy. Descriptions of the headers are provided in DatasetHeaderDescription.txt, included in the archive.</li> <li>Trained TensorFlow models for the prediction of potentials of mean force from HG interaction coefficients, suitable for loading via the Keras backend.</li> <li>PMFs generated for a range of surfaces and chemicals as output from the trained model, in both text format and figures showing comparisons to training PMFs where available. PMFs are supplied as tabulated data with comma separated values of distance in nm and interaction energies in kJ/mol.</li> <li>Adsorption energies in kJ/mol evaluated at T=300K extracted from all PMFs and compared to the values obtained from known PMFs where available.</li> </ul> <p>The surface_pmfpredictor.zip archive contains PMFs selected for the operation of the UnitedAtom software package for the calculation of protein-nanoparticle interactions. This data is included in the main repository file and provided separately to avoid the download of unnecessary data if only the final PMFs are required. As with the main set, these are provided in tabulated form with distance [nm], energy [kJ/mol] pairs. This repository also contains the sets of figures illustrating these PMFs for each surface. Both archives contain further information on the contents, including descriptions of the surfaces and chemicals for which PMFs are computed. We also supply the training data used to build the model in a separate archive, PMFPredictor-TrainingData.zip, along with a text file containing descriptions of all headers in this file. This training data is quite large when uncompressed, c.a. 7 Gb, hence its exclusion from the main archive.</p> <p>If you use results from this repository please cite the following paper in addition to the repository itself:</p> <p>I. Rouse, V. Lobaskin, Machine-learning based prediction of small molecule -- surface interaction potentials, arXiv:2211.07999<br> https://arxiv.org/abs/2211.07999</p>
Source code and simulation results for computing resonance expansions of quadratic quantities with regularized quasinormal modes
<p>This data publication supplements the article "Resonance expansion of quadratic quantities with regularized quasinormal modes" [1]. Tabulated data related to the figures in the manuscript is provided along with the Matlab scripts used to generate the results. The Riesz projection software package RPExpand [2] has been extended to support quasi normal modes (QNMs) and, in particular, the proposed method for quadratic quantities. A current version is contained in the directory <code>Code</code>. Furthermore, the input files required for scattering and resonance simulations with the finite element method (FEM) solver JCMsuite [3] are contained.</p> <p><strong>Requirements</strong></p> <ul> <li>JCMsuite (version 5.2.1 or newer)</li> <li>MATLAB (tested with version R2019b)</li> </ul> <p>In order to run the scripts you must replace the corresponding place holders in the files by a path to your installation of JCMsuite. Free trial licenses are available, please refer to the homepage of <a href="https://jcmwave.com/">JCMwave</a>. </p> <p><strong>References</strong></p> <p>[1] Fridtjof Betz, Felix Binkowski, Martin Hammerschmidt, Lin Zschiedrich, Sven Burger: Resonance expansion of quadratic quantities with regularized quasinormal modes, Physica Status Solidi A <strong>220</strong>, 2370013 (2023)</p> <p>[2] Fridtjof Betz, Felix Binkowski, Sven Burger, RPExpand: Software for Riesz projection expansion of resonance phenomena, SoftwareX <strong>15</strong>, 100763 (2021), https://doi.org/10.1016/j.softx.2021.100763</p> <p>[3] Jan Pomplun, Sven Burger, Lin Zschiedrich, Frank Schmidt, Adaptive finite element method for simulation of optical nano structures, Physica Status Solidi B <strong>244</strong>, 3419 (2007), http://dx.doi.org/10.1002/pssb.200743192</p>
Residual dynamics resolves recurrent contributions to neural computation
<p>This repository contains all the neural and simulated datasets used in the above publication. All the data are stored as .mat files.</p> <p>This repository is organized into three main folders. For ease of use, we recommend extracting the contents of each of the three folders into a single folder titled ‘data’.</p> <ol> <li>neuraldata –Raw and preprocessed neural data. Organized into the following folders: <ol> <li>array_TDR_dotsTask_data – contains the non-processed neural data of each session for both monkeys</li> <li>array_dotsTask_beh_xx – contains the behavioral data of each session for both monkeys (xx = monkey)</li> <li>array_reppasdotsTask_datasegmented_binsize=xxms – contains the pre-processed neural data (session wise) of both monkeys for a specified bin size (xx = bin size)<br> </li> </ol> </li> <li>simulations – Data from the simulated models. Organized into the following folders: <ol> <li>toymodels – Simulated data from all models of decisions/movement. Also includes simulations of the augmented line attractor and rotational dynamics models.</li> <li>twoarearnn – Simulated data from the two types of (feedback and nofeedback) two-area RNN models, along with the impulse response simulations.<br> </li> </ol> </li> <li>analyses – contains processed data files from various stages of the analysis pipeline. Only the most important files are listed below: <ol> <li>xx_aligned_reppasdotsTask_binsize=45ms.mat – session-aligned neural data for a specified monkey (xx = monkey)</li> <li>xx_ndim=8_lag=3_alpha=200_50_allconfigsresdynresults.mat – contains the residual dynamics for a specified monkey.</li> <li>xx_aligned_hankelagdimCV_reppasdotsTask_binsize=45ms.mat and xx_aligned_smoothnessCV_reppasdotsTask_binsize=45ms.mat – results of cross-validation of the residual dynamics pipeline for a specified monkey (xx = monkey)</li> </ol> </li> </ol>
Soybean dependence on biotic pollination decreases with latitude - Data and Computer code
<p>Release of Datasets and R scripts needed to reproduce the analyses and figures published in the article <em>'Soybean dependence on biotic pollination decreases with latitude'</em>, published in Agriculture, Ecosystems & Environment, Volume 347, 1 May 2023, 108376. <a href="https://doi.org/10.1016/j.agee.2023.108376">https://doi.org/10.1016/j.agee.2023.108376</a></p> <p><strong>Highlights</strong></p> <ul> <li>In the absence of pollinators, soybean yield decreases between 0 and ~50%.</li> <li>Variation in pollinator dependence (PD) was found to be structured latitudinally.</li> <li>PD decreases at high latitudes due to an apparently higher incidence of autogamy.</li> <li>Temperature and photoperiod could play an important role in determining PD.</li> <li>Changes in cleistogamy and androsterility might explain the reported trends.</li> </ul> <p><strong>Abstract</strong></p> <p>Identifying large-scale patterns of variation in pollinator dependence (PD) in crops is important from both basic and applied perspectives. Evidence from wild plants indicates that this variation can be structured latitudinally. Individuals from populations at high latitudes may be more selfed and less dependent on pollinators due to higher environmental instability and overall lower temperatures, environmental conditions that may affect pollinator availability. However, whether this pattern is similarly present in crops remains unknown. Soybean (Glycine max), one of the most important crops globally, is partially self-pollinated and autogamous, exhibiting large variation in the extent of PD (from a 0 to ~50% decrease in yield in the absence of animal pollination). We examined latitudinal variation in soybean's PD using data from 28 independent studies distributed along a wide latitudinal gradient (4-43 degrees). We estimated PD by comparing yields between open pollinated and pollinator-excluded plants. In the absence of pollinators, soybean yield was found to decrease by an average of ~30%. However, PD decreases abruptly at high latitudes, suggesting a relative increase in autogamous seed production. Pollinator supplementation does not seem to increase seed production at any latitude. We propose that latitudinal variation in PD in soybean may be driven by temperature and photoperiod affecting the expression of cleistogamy and androsterility. Therefore, an adaptive mating response to an unpredictable pollinator environment apparently common in wild plants can also be imprinted in highly domesticated and genetically-modified crops.</p> <p><strong>Content</strong></p> <p>The dataset consists of two files</p> <p>1 - <a href="https://github.com/NERC-CEH/Soybean-dependence-on-biotic-pollination-decreases-with-latitude/blob/main/%5Bdata%5D%20Cunha%20et%20al.%20MS_soybean.xlsx">[data] Cunha et al. MS_soybean.xlsx</a> is an excel file with two sheets, <strong>data</strong> and <strong>data_map</strong>. These sheets contain the data used in the models defined in the R script <a href="https://github.com/NERC-CEH/Soybean-dependence-on-biotic-pollination-decreases-with-latitude/blob/main/%5BR%20script%5D%20Cunha%20et%20al.%20MS_soybean.R">[R script] Cunha et al. MS_soybean.R</a>.</p> <ul> <li> <p>1.1 The <strong>data</strong> sheet contains the variables:</p> <ul> <li>Value = log_ratios</li> <li>Lat = latitude in decimal degrees</li> <li>Variable = yield component</li> <li>Treatment = treatment type for comparing pollinator dependence</li> <li>Reference_Data_owner = study ID where the data was obtained</li> <li>Site = site within the study where each field experiment was performed</li> </ul> </li> <li> <p>1.2 The <strong>data_map</strong> sheet contains information used for plotting the geographical distribution of the used studies:</p> <ul> <li>Reference_Data_owner = study ID where the data was obtained</li> <li>Country = country where the study was performed</li> <li>Province = province where the study was performed</li> <li>Locality/Farm = locality where the study was performed</li> <li>Lat = latitude in decimal degrees</li> <li>Long = longitude in decimal degrees</li> </ul> </li> </ul> <p>2 - <a href="https://github.com/NERC-CEH/Soybean-dependence-on-biotic-pollination-decreases-with-latitude/blob/main/%5Bdata%5D%20Cunha%20et%20al.%20MS_soybean%20%5Bdate_photoperiod%5D.csv">[data] Cunha et al. MS_soybean [date_photoperiod].csv</a> is a comma-separated file that contains the information used in the R script <a href="https://github.com/NERC-CEH/Soybean-dependence-on-biotic-pollination-decreases-with-latitude/blob/main/%5BR%20script%5D%20Cunha%20et%20al.%20AGEE%20-%20gee_temp_ts_extract.R">[R script] Cunha et al. AGEE - gee_temp_ts_extract.R</a> and produces Figure S2.</p> <ul> <li> <p>2.1 The dataset contains the following variables:</p> <ul> <li>study_ID = study ID number where the data was obtained</li> <li>study_ref = study ID where the data was obtained</li> <li>latitude = latitude in decimal degrees</li> <li>longitude = longitude in decimal degrees</li> <li>date1 = date of the sowing or flowering when the experiment was done</li> <li>date2 = a second date, when available, of the sowing or flowering when the experiment was done</li> <li>event = if the date was related to the sowing of seeds or flowering of soybean.</li> </ul> </li> </ul>
Controllable temporal dynamics of titanium oxide memristor for analog time-based neuromorphic computing: Dataset
<p>Dataset used to produce graphs related to the temporal behavior of the Pt/TiO/Au memristors.</p>
Datasets for testing the computational thinking of children with mental disabilities by cCT-test
<p>This study aims to evaluate the impact of web technologies on the development of computational thinking of students with mental disabilities. The experiment involved 14 students aged 8-12. For 8 weeks children were trained in computational thinking and computer science. Assessment of computational thinking was performed with cCT-test by El-Hamamsi et al. before and after the experiment (El-Hamamsy, L., Zapata-Cáceres, M., Barroso, E. M., Mondada, F., Zufferey, J. D., & Bruno, B. (2022). The competent Computational Thinking Test: Development and Validation of an Unplugged Computational Thinking Test for Upper Primary School. Journal of Educational Computing Research, 07356331221081753 <a href="https://doi.org/10.1177/07356331221081753">https://doi.org/10.1177/07356331221081753</a>).</p> <p>After conducting computer science lessons using web technologies the respondents showed a higher level of computational thinking (M=15,7, SD=3,69), compared to the results of preliminary testing (M=5,93, SD=2,3). Web technologies can significantly increase the effectiveness of inclusive pedagogy, which establishes the importance of integrating web technologies into the teaching system in inclusive classes of general education schools.</p> <p>Examples of tasks are located at the link <a href="https://miro.com/app/board/uXjVPz8p6LE=/?share_link_id=864046044746">https://miro.com/app/board/uXjVPz8p6LE=/?share_link_id=864046044746</a> </p> <p>https://wordwall.net/resource/48631685</p> <p>https://wordwall.net/resource/48650842</p> <p>https://wordwall.net/resource/48655774</p> <p>https://wordwall.net/resource/48661305</p> <p> </p>
Computer Vision Datasets for Visual Blockage Assessment at Culverts
<p>Blockage of culverts caused by transported debris is a major factor in causing flash floods in urban areas. Traditional hydraulic models have been unsuccessful in solving this problem due to a lack of data on peak flood hydraulics and the complex behavior of debris at culverts. To address this problem, a new approach of developing intelligent video analytics (IVA) algorithms is being proposed, which uses computer vision algorithms to extract information about visual blockage. This approach is expected to help in timely and safe maintenance operations and reduce the risk of culverts being blocked. To support the development of computer vision solutions, two datasets have been created: the Synthetic Images of Culvert (SIC) and the Visual Hydraulics Lab Dataset (VHD).</p> <ul> <li>The Synthetic Images of Culvert (SIC) dataset consists of synthetic images of culverts that were generated using a 3D computer application built on the Unity3D gaming engine. The application was designed to simulate various blockage scenarios by allowing users to place different types of debris materials in the scene in various orientations and locations. These blockage scenarios were captured as images using batch capture functionality. The dataset offers diversity in terms of the type of debris (urban, vegetative, mixed), culvert types (pipe, single circular, double circular, single box, double box, triple box), camera viewpoints, time of day, and water levels. However, it has some limitations, such as a single natural background and unrealistic effects and animations.</li> <li>The Visual Hydraulics-Lab Dataset (VHD) is a dataset of simulated images of culverts that were captured during controlled hydrology lab experiments. The experiments involved a series of tests using scaled physical models of culverts under different flooding conditions. The experiments were recorded using two high definition (HD) cameras and images of culverts in both blocked and unblocked conditions were extracted. The VHD dataset includes a variety of images with different culvert configurations (single circular, double circular, single box, double box), blockage types (urban, vegetative, mixed), simulated lighting conditions, camera viewpoints, and flood levels controlled by inlet water discharge. The limitations of the dataset include reflections from the water surface and flume walls, an identical background and scaling, and clear water.</li> </ul>
Supplemental data for "Computational screening of chemically active metal center in coordinated dipyridyl tetrazine network"
<p>Atomic coordinates of structures used in N. Ud Din, D. Le, T. S. Rahman "Computational screening of chemically active metal center in coordinated dipyridyl tetrazine network", J. Phys.: Condens. Matter .(2023). DOI: 10.1088/1361-648X/acb8f3</p>
Series expansions for computing rhumb areas
<p>This data accompanies the paper</p> <blockquote>C. F. F. Karney,<br> <a href="https://arxiv.org/abs/2303.03219">The area of rhumb polygons</a>,<br> Technical Report, SRI International, March 2023.<br> <a href="https://arxiv.org/abs/2303.03219">https://arxiv.org/abs/2303.03219</a></blockquote> <p>Files:</p> <ul> <li><code>rhumbarea.mac</code>: the Maxima code used to obtain the series expansion for the rhumb area. Instructions for using this code are included in the file.</li> <li><code>auxvals40.mac</code>: the 40th-order series for auxiliary latitudes as Maxima code (required by rbumbarea.mac).</li> <li><code>rhumbvals40.mac</code>: the 40th-order series for the rhumb area as Maxima code.</li> <li><code>rhumbvals6.m</code>, <code>rhumbvals16.m</code>, and <code>rhumbvals40.m</code>: the 6th-, 16th-, and 40th-order series as Octave/MATLAB code. The 40th-order series are given as floating-point numbers; the others are as exact fractions.</li> </ul>
A computationally efficient statistically downscaled 100 m resolution Greenland product from the regional climate model MAR: accompanying dataset
<p>Dataset containing surface temperature and surface mass balance datasets generated from the MAR regional climate model over Greenland over two test areas using statistical downscaling tools from 6 km to 100m. The abstract of the accompanying submitted paper follows: </p> <p> </p> <p>The Greenland Ice Sheet (GrIS) has been contributing directly to sea level rise and this contribution is projected to accelerate over next decades. A crucial tool for studying the evolution surface mass loss (e.g., surface mass balance, SMB) consists of regional climate models (RCMs) which can provide current estimates and future projections of sea level rise associated with such losses. However, one of the main limitations of RCMs is the relatively coarse horizontal spatial resolution at which outputs are currently generated. Here, we report results concerning the statistical downscaling of the SMB modeled by the Modèle Atmosphérique Régional (MAR) RCM from the original spatial resolution of 6 km to 100 m building on the relationship between elevation and mass losses in Greenland. To this goal, we developed a geospatial framework that allows the parallelization of the downscaling process, a crucial aspect to increase the computational efficiency of the algorithm. The results obtained in the case of the SMB, assessed through the comparison of the modeled outputs with in-situ SMB measurements, show a considerable improvement in the case of the downscaled product with respect to the original, coarse output. In the case of the downscaled MAR product, the coefficient of determination (R<sup>2</sup>) increases from 0.868 for the original MAR output to 0.935 for the downscaled product. Moreover, the value of the slope and intercept of the linear regression fitting modeled and measured SMB values shifts from 0.865 for the original MAR to 1.015 for the downscaled product in the case of the intercept and from the value -235mm (original) to -57 mm (downscaled) in the case of the slope, considerably improving upon results previously published in the literature.</p>
Raw data for the computation of ExPaNDS facilities maturity wrt FAIR data catalogues
<p>Raw data for the computation of ExPaNDS facilities maturity wrt FAIR data catalogues.</p> <p>The method is explained in the <a href="https://doi.org/10.5281/zenodo.4146819">report on status, gap analysis and roadmap towards harmonised and federated metadata catalogues for EU national Photon and Neutron RIs</a>.</p>
Dataset for "A computational fluid dynamics—Population balance equation approach for evaporating cough droplets transport"
<p>Dataset for figures and tables of the article "A computational fluid dynamics—Population balance equation approach for evaporating cough droplets transport" submitted to "International Journal of Multiphase Flow".</p>
A Modular Quantum Compilation Framework for Distributed Quantum Computing
<p>This repository contains the data used for the plots in "<em>A Modular Quantum Compilation Framework for Distributed Quantum Computing</em>" by D. Ferrari, S. Carretta and M. Amoretti.</p> <p>Data is located in the <em>'data'</em> directory in <em>.csv</em> format, a python script to generate the plots can be found in the main directory. The script was tested with <strong>python3.10</strong> and needs <strong>matplotlib</strong>, <strong>pandas</strong> and <strong>seaborn</strong> packages. Plots are saved as <em>.pdf</em> files in the <em>'figures'</em> directory.</p>
childPoeDE: A corpus of German Children's Poems for Computational and Experimental Studies - Metadata
<p>The childPoeDE corpus is a collection of 1082 German poems for children created within the CHYLSA project. The poems were taken from anthologies published between 1991 and 2019. This publication includes the poem-level metadata for each poem with information about the author, the poem's length, data on case, punctuation, layout, rhyme, type-token ratio (TTR and MATTR) and lexical density. It also includes token-level metadata, namely word length and position, POS tags in different levels of granularity as well as data on onomatopoeia and sonority. Furthermore, this publication provides a word frequency table and a Python script which was used to extract some of the metadata from the texts (poemtool.py). The childPoeDE corpus does not contain all poems from the anthologies. A list of the poems that have been omitted for different reasons (length, language, typography, ...) can be accessed as well.</p> <p>Read more about the childPoeDE corpus in our data paper published in the Journal of Open Humanities Data: <a href="https://doi.org/10.5334/johd.102">The ChildPoeDE Corpus: 1082 German Children’s Poems for Computational and Experimental Studies on Poetry Reception</a>.</p> <p>DFG Schwerpunktprogramm SPP 2207 “Computational Literary Studies“<br> Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/402743989">https://gepris.dfg.de/gepris/projekt/402743989</a></li> <li><a href="https://dfg-spp-cls.github.io/">https://dfg-spp-cls.github.io<em>/</em></a></li> </ol> <p>Subproject: „CHYLSA (Children’s and Youth Literature Sentiment Analysis)“</p> <p>Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/424250469">https://gepris.dfg.de/gepris/projekt/424250469</a></li> <li><a href="https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/">https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/</a></li> </ol>
Physiological Signals During Motor Imagery Brain-Computer Interface Training Using Virtual Reality and Haptics
<p><strong>Participant demographics:</strong></p> <p>The sample is consisted by 20 healthy volunteers with a mean age of 24.79 years (SD = 3.54 years). The cohort was 68% male and 32% female. In terms of education, 16% had attended only high school, while 32% had a bachelor's degree, 42% a master's degree, and 11% a doctorate. All participants signed an informed consent before participating in the study in accordance with the 1964 Declaration of Helsinki.</p> <p><strong>Experiment Description:</strong></p> <p>The experiment consisted in having the subjects perform motor imagery of a bimanual rowing task with two individual paddles, one in each hand, under five experimental conditions. Four of these conditions used NeuRow (<a href="https://link.springer.com/chapter/10.1007/978-3-030-27950-9_1"><strong>Vourvopoulos et al. (2016-2019</strong>))</a>—a VR environment that renders virtual arms from a first-person perspective—while the other conditions used abstract feedback based on the BCI-Graz paradigm<a href="https://ieeexplore.ieee.org/abstract/document/1214714"> (<strong>Pfurtscheller et al. (2003))</strong></a>. All six conditions and their acronyms are described below:</p> <ol> <li><strong>Motor Imagery(MI)</strong>: The standard motor imagery training, with a fixation cross and directional arrows on a black background guiding the subjects through the experiment.</li> <li><strong>Motor Imagery/Motor Observation (MIMO):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a monitor.</li> <li><strong>Motor Imagery/Motor Observation with Haptics (MIMOHP): </strong>A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a monitor. Hand controllers also provided haptic feedback through vibrotactile stimulation.</li> <li><strong>Motor Imagery/Motor Observation with VR HMD (MIMOVR):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a VR HMD.</li> <li><strong>Motor Imagery/Motor Observation with VR HMD and Haptics (MIMOVRHP):</strong> A motor imagery training paradigm using NeuRow, with a fixation cross and directional arrows overlaid on the VR environment, which was displayed through a VR HMD. Hand controllers also provided haptic feedback through vibrotactile stimulation.</li> <li><strong>Motor Execution (ME):</strong> A fixation cross and directional arrows were displayed on a black background through a monitor (same as in MI), and guided the subjects through the experiment by having them tap their fingers accordingly. Data from this condition was available only after S07, so only 10 subjects<br> have performed ME.</li> </ol> <p>Finally, this experiment followed a within-subject design, in a randomized order of the conditions to minimize any order effects, while MI and ME conditions acted as control.</p> <p><strong>Equipment:</strong></p> <p>A wireless EEG amplifier (LiveAmp; Brain Products GmbH, Gilching, Germany) was used, with 32 active electrodes(+3 ACC) with a sampling rate of 500Hz. In addition, <strong>ECG, PPG</strong> and <strong>Respiration</strong> signals have been recorded synchronously in a bipolar montage, and connected to the EEG amplifier’s AUX input through the Brain Products BIP2AUX adapter.</p> <p>Visual feedback was provided through a monitor in all conditions except in MIMOVR and MIMOVRHP, in which an Oculus Rift CV1 headset (Reality Labs, formerly Facebook, Inc., CA, USA) was used instead. Haptic feedback was provided through the Oculus Rift hand controllers.<br> </p> <p><strong>Channel Indices:</strong></p> <p><strong>EEG</strong>: 1-32<br> <strong>PPG</strong> (AUX1): 33<br> <strong>Resp</strong>. (AUX2): 34<br> <strong>ECG</strong> (AUX3): 35<br> <strong>ACC</strong>: 36-38</p> <p> </p> <p><strong>Event codes:</strong></p> <table> <tbody> <tr> <td><strong>Code</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>S01</td> <td>Experiment Start</td> </tr> <tr> <td>S02</td> <td>Baseline Start</td> </tr> <tr> <td>S03</td> <td>Baseline Stop</td> </tr> <tr> <td>S04</td> <td>Start Of Trial</td> </tr> <tr> <td>S05</td> <td>Cross On Screen</td> </tr> <tr> <td>S07</td> <td>class1, Left hand </td> </tr> <tr> <td>S08</td> <td>class2, Right hand </td> </tr> <tr> <td>S09</td> <td>Feedback Continuous</td> </tr> <tr> <td>S10</td> <td>End of Trial</td> </tr> <tr> <td>S11</td> <td>End Of Session</td> </tr> <tr> <td>S12</td> <td>Experiment Stop</td> </tr> </tbody> </table> <p> </p> <p><strong>Directory tree:</strong></p> <p>ROOT<br> |<br> +--- USER #<br> | +---SESSION #<br> | | +---TASK #<br> | | | +---MI<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMO<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOHP<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOVR<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---MIMOHPVR<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk<br> | | | +---ME<br> | | | | .eeg<br> | | | | .vhdr<br> | | | | .vmrk</p> <p> </p> <p><strong>Note: </strong>The first three datasets are from pilot sessions: sub-p01 to p03. From sub-01 to 19, subjects 10 and 11 have been removed due to the lack of markers. Subject sub-13, task MIMOVRHP is missing.</p> <p> </p>
AI-derived annotations for the NLST and NSCLC-Radiomics computed tomography imaging collections
<p>Public imaging datasets are critical for the development and evaluation of automated tools in cancer imaging. Unfortunately, many of the available datasets do not provide annotations of tumors or organs-at-risk, crucial for the assessment of these tools. This is due to the fact that annotation of medical images is time consuming and requires domain expertise. It has been demonstrated that artificial intelligence (AI) based annotation tools can achieve acceptable performance and thus can be used to automate the annotation of large datasets. As part of the effort to enrich the public data available within NCI Imaging Data Commons (IDC) (<a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov/</a>) [1], we introduce this dataset that consists of such AI-generated annotations for two publicly available medical imaging collections of Computed Tomography (CT) images of the chest. For detailed information concerning this dataset, please refer to our publication <a href="https://www.nature.com/articles/s41597-023-02864-y">here</a> [2]. </p> <p>We use publicly available pre-trained AI tools to enhance CT lung cancer collections that are unlabeled or partially labeled. The first tool is the nnU-Net deep learning framework [3] for volumetric segmentation of organs, where we use a pretrained model (Task D18 using the SegTHOR dataset) for labeling volumetric regions in the image corresponding to the heart, trachea, aorta and esophagus. These are the major organs-at-risk for radiation therapy for lung cancer. We further enhance these annotations by computing 3D shape radiomics features using the pyradiomics package [4]. The second tool is a pretrained model for per-slice automatic labeling of anatomic landmarks and imaged body part regions in axial CT volumes [5].</p> <p>We focus on enhancing two publicly available collections, the Non-small Cell Lung Cancer Radiomics (NSCLC-Radiomics collection) [6,7], and the National Lung Screening Trial (NLST collection) [8,9]. The CT data for these collections are available both in The Cancer Imaging Archive (TCIA) [10] and in NCI Imaging Data Commons (IDC). Further, the NSLSC-Radiomics collection includes expert-generated manual annotations of several chest organs, allowing us to quantify performance of the AI tools in that subset of data.</p> <p>IDC is relying on the DICOM standard to achieve FAIR [10] sharing of data and interoperability. Generated annotations are saved as DICOM Segmentation objects (volumetric segmentations of regions of interest) created using the <em>dcmqi</em> [12], and DICOM Structured Report (SR) objects (per-slice annotations of the body part imaged, anatomical landmarks and radiomics features) created using <em>dcmqi </em>and <em>highdicom</em> [13]. 3D shape radiomics features and corresponding DICOM SR objects are also provided for the manual segmentations available in the NSCLC-Radiomics collection. </p> <p>The dataset is available in IDC, and is accompanied by our publication <a href="https://www.nature.com/articles/s41597-023-02864-y">here</a> [2]. This pre-print details how the data were generated, and how the resulting DICOM objects can be interpreted and used in tools. Additionally, for further information about how to interact with and explore the dataset, please refer to our <a href="https://github.com/ImagingDataCommons/nnU-Net-BPR-annotations/">repository</a> and accompanying <a href="https://github.com/ImagingDataCommons/nnU-Net-BPR-annotations/blob/main/usage_notebooks/scientific_data_paper_usage_notes.ipynb">Google Colaboratory notebook</a>. </p> <p>The annotations are organized as follows. For NSCLC-Radiomics, three nnU-Net models were evaluated ('2d-tta', '3d_lowres-tta' and '3d_fullres-tta'). Within each folder, the PatientID and the StudyInstanceUID are subdirectories, and within this the DICOM Segmentation object and the DICOM SR for the 3D shape features are stored. A separate directory for the DICOM SR body part regression regions ('sr_regions') and landmarks ('sr_landmarks') are also provided with the same folder structure as above. Lastly, the DICOM SR for the existing manual annotations are provided in the 'sr_gt' directory. For NSCLC-Radiomics, each patient has a single StudyInstanceUID. The DICOM Segmentation and SR objects are named according to the SeriesInstanceUID of the original CT files. </p> <ul> <li>nsclc <ul> <li>2d-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>3d_lowres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>3d_fullres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_regions <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_regions_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_landmarks <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_landmarks_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_gt <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>For NLST, the '3d_fullres-tta' model was evaluated. The data is organized the same as above, where within each folder the PatientID and the StudyInstanceUID are subdirectories. For the NLST collection, it is possible that some patients have more than one StudyInstanceUID subdirectory. A separate directory for the DICOM SR body par regions ('sr_regions') and landmarks ('sr_landmarks') are also provided. The DICOM Segmentation and SR objects are named according to the SeriesInstanceUID of the original CT files. </p> <ul> <li>nlst <ul> <li>3d_fullres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_regions <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_regions_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_landmarks <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_landmarks_SR.dcm </li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>The query used for NSCLC-Radiomics is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/NSCLC_Radiomics_query.txt">here</a>, and a list of corresponding SeriesInstanceUIDs (along with PatientIDs and StudyInstanceUIDs) is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/zenodo_nsclc_radiomics_series_analyzed.csv">here</a>. The query used for NLST is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/NLST_query.txt">here</a>, and a list of corresponding SeriesInstanceUIDs (along with PatientIDs and StudyInstanceUIDs) is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/zenodo_nlst_series_analyzed.csv">here</a>. The two csv files that describe the series analyzed, <em>nsclc_series_analyzed.csv</em> and <em>nlst_series_analyzed.csv</em>, are also available as uploads to this repository. </p> <p><em>Version updates: </em></p> <p><em>Version 2: For the regions SR and landmarks SR, changed to use a distinct TrackingUniqueIdentifier for each MeasurementGroup. Also instead of using TargetRegion, changed to use FindingSite. Additionally for the landmarks SR, the TopographicalModifier was made a child of FindingSite instead of a sibling.</em></p> <p><em>Version 3: Added the two csv files that describe which series were analyzed </em></p> <p><em>Version 4: Modified the landmarks SR as the TopographicalModifier for the Kidney landmark (bottom) does not describe the landmark correctly. The Kidney landmark is the "first slice where both kidneys can be seen well." Instead, removed the use of the TopographicalModifier for that landmark. For the features SR, modified the units code for the Flatness and Elongation, as we incorrectly used mm units instead of no units. </em></p>
Experimental Demonstration of In-Memory Computing in a Ferrofluid System
<p>Measurements that demonstrate the feasibility of in-memory computing using an amorphous ferrofluid system (in a liquid aggregation state). This work poses the basis for the exploitation of a colloid as both an in-memory computing device and as a full-electric liquid computer thanks to its fluidity and the reported complex dynamics, via probing read-out and programming ports.</p>
CaRCC Research Computing and Data (RCD) Workforce Survey Data 2021 - Part 2
<p>Data sets to accompany "Compensation of Academic Research Computing and Data Professionals"</p> <p>Paper citation: Christina Maimone, Carrie Brown, Kimberly Grasch, Chris Reidy, and Ashley Stauffer. 2023. Compensation of Academic Research Computing and Data Professionals. In Proceedings of Practice and Experience in Advanced Research Computing (PEARC23). ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3569951.3593599</p> <p>Survey questions: Maimone, Christina, Yockel, Scott, Middelkoop, Timothy, Alameda, Jay, Stauffer, Ashley, & Neeser, Amy. (2021). CaRCC Research Computing and Data (RCD) Workforce 2021 Census Questions (1.0). Zenodo. https://doi.org/10.5281/zenodo.5914431</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.