Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
431
datasets available to search
ShareScore release 0.9.0
Dataset results
431 results for “Decoding”
Data from: Decoding and encoding models reveal the role of mental simulation in the brain representation of meaning
<p>How the brain representation of conceptual knowledge vary as a function of processing goals, strategies and task-factors remains a key unresolved question in cognitive neuroscience. Here we asked how the brain representation of semantic categories is shaped by the depth of processing during mental simulation. Participants were presented with visual words during functional magnetic resonance imaging (fMRI). During shallow processing, participants had to read the items. During deep processing, they had to mentally simulate the features associated with the words. Multivariate classification, informational connectivity and encoding models were used to reveal how the depth of processing determines the brain representation of word meaning. Decoding accuracy in putative substrates of the semantic network was enhanced when the depth processing was high, and the brain representations were more generalizable in semantic space relative to shallow processing contexts. This pattern was observed even in association areas in inferior frontal and parietal cortex. Deep information processing during mental simulation also increased the informational connectivity within key substrates of the semantic network. To further examine the properties of the words encoded in brain activity, we compared computer vision models - associated with the image referents of the words - and word embedding. Computer vision models explained more variance of the brain responses across multiple areas of the semantic network. These results indicate that the brain representation of word meaning is highly malleable by the depth of processing imposed by the task, relies on access to visual representations and is highly distributed, including prefrontal areas previously implicated in semantic control.</p>
Decoding Movement Goals from the Fronto-Parietal Reach Network
<p>Here we provide fMRI data used for MVPA in the following project (for details see data description file):</p> <p>Gertz H, Lingnau A and Fiehler K (2017). Decoding Movement Goals from the Fronto-Parietal Reach Network. <em>Front. Hum. Neurosci.</em> <strong>11</strong>:84. doi: 10.3389/fnhum.2017.00084</p> <p>During reach planning, fronto-parietal brain areas need to transform sensory information into a motor code. It is debated whether these areas maintain a sensory representation of the visual cue or a motor representation of the upcoming movement goal. Here, we present results from a delayed pro-/anti-reach task which allowed for dissociating the position of the visual cue from the reach goal. In this task, the visual cue was combined with a context rule (pro vs. anti) to infer the movement goal. Different levels of movement goal specification during the delay were obtained by presenting the context rule either before the delay together with the visual cue (specified movement goal) or after the delay (underspecified movement goal). By applying fMRI multivoxel pattern analysis (MVPA) we demonstrate movement goal encoding in the left dorsal premotor cortex (PMd) and bilateral superior parietal lobule (SPL) when the reach goal is specified. This suggests that fronto-parietal reach regions maintain a prospective motor code during reach planning. When the reach goal is underspecified, only area PMd but not SPL represents the visual cue position indicating an incomplete state of sensorimotor integration. Moreover, this result suggests a potential role of PMd in movement goal selection.</p>
Data supporting "Local Clustering Decoder: a fast and adaptive hardware decoder for the surface code"
<p>The data consists of a CSV file containing the raw performance data collected from running our decoder on a Xilinx Virtex Ultrascale+ VU19P FPGA. The Stim circuits that were used to create the samples are provided in a ZIP file. Our internal fork of Stim with support for leakage is needed to sample noise from the circuits.</p>
Dataset containing raw simulation data for a paper on decoding bosonic quantum LDPC codes
<p>This is the dataset from evaluations used in code related to a paper on analog information decoding of bosonic quantum LDPC codes, available on Github (https://github.com/cda-tum/mqt-qecc/). For more information we refer to the Github repository and the paper.</p>
Video-EEG Encoding-Decoding Dataset KU Leuven
<p><strong> If using this dataset, please cite the following paper and the current Zenodo repository.</strong></p> <p>This dataset is described in detail in the following paper:</p> <p><a href="https://iopscience.iop.org/article/10.1088/1741-2552/ad2333/meta">[1] Yao, Y., Stebner, A., Tuytelaars, T., Geirnaert, S., & Bertrand, A. (2024). Identifying temporal correlations between natural single-shot videos and EEG signals. <em>Journal of Neural Engineering</em>, <em>21</em>(1), 016018. doi:10.1088/1741-2552/ad2333</a></p> <p>The associated code is available at: <a href="https://github.com/YYao-42/Identifying-Temporal-Correlations-Between-Natural-Single-shot-Videos-and-EEG-Signals?tab=readme-ov-file">https://github.com/YYao-42/Identifying-Temporal-Correlations-Between-Natural-Single-shot-Videos-and-EEG-Signals?tab=readme-ov-file</a></p> <h2><strong>Introduction</strong></h2> <p>The research work leading to this dataset was conducted at the Department of Electrical Engineering (ESAT), KU Leuven.</p> <p>This dataset contains electroencephalogram (EEG) data collected from 19 young participants with normal or corrected-to-normal eyesight when they were watching a series of carefully selected YouTube videos. The videos were muted to avoid the confounds introduced by audio. For synchronization, a square box was encoded outside of the original frames and flashed every 30 seconds in the top right corner of the screen. A photosensor, detecting the light changes from this flashing box, was affixed to that region using black tape to ensure that the box did not distract participants. The EEG data was recorded using a BioSemi ActiveTwo system at a sample rate of 2048 Hz. Participants wore a 64-channel EEG cap, and 4 electrooculogram (EOG) sensors were positioned around the eyes to track eye movements.</p> <p>The dataset includes a total of <strong>(19 subjects x 63 min + 9 subjects x 24 min)</strong> of data. Further details can be found in the following section.</p> <h2><strong>Content</strong></h2> <ul> <li>YouTube Videos: Due to copyright constraints, the dataset includes links to the original YouTube videos along with precise timestamps for the segments used in the experiments. The features proposed in [1] (<em>Object Flow</em>) have been extracted and can be downloaded here: <a href="https://drive.google.com/file/d/1J1tYrxVizrl1xP-W1imvlA_v-DPzZ2Qh/view?usp=sharing">https://drive.google.com/file/d/1J1tYrxVizrl1xP-W1imvlA_v-DPzZ2Qh/view?usp=sharing</a>.</li> <li>Raw EEG Data: Organized by subject ID, the dataset contains EEG segments corresponding to the presented videos. Both EEGLAB .set files (containing metadata) and .fdt files (containing raw data) are provided, which can also be read by popular EEG analysis Python packages such as MNE. <ul> <li>The naming convention links each EEG segment to its corresponding video. E.g., the EEG segment 01_eeg corresponds to video 01_Dance_1, 03_eeg corresponds to video 03_Acrob_1, Mr_eeg corresponds to video Mr_Bean, etc.</li> <li>The raw data have 68 channels. The first 64 channels are EEG data, and the last 4 channels are EOG data. The position coordinates of the standard BioSemi headcaps can be downloaded here: <a href="https://www.biosemi.com/download/Cap_coords_all.xls">https://www.biosemi.com/download/Cap_coords_all.xls</a>.</li> <li>Due to minor synchronization ambiguities, different clocks in the PC and EEG recorder, and missing or extra video frames during video playback (rarely occurred), the length of the EEG data may not perfectly match the corresponding video data. The difference, typically within a few milliseconds, can be resolved by truncating the modality with the excess samples.</li> </ul> </li> <li>Signal Quality Information: A supplementary .txt file detailing potential bad channels. Users can opt to create their own criteria for identifying and handling bad channels.</li> </ul> <p>The dataset is divided into two subsets: Single-shot and MrBean, based on the characteristics of the video stimuli.</p> <h3><strong>Single-shot Dataset</strong></h3> <p>The stimuli of this dataset consist of 13 single-shot videos (63 min in total), each depicting a single individual engaging in various activities such as dancing, mime, acrobatics, and magic shows. All the participants watched this video collection.</p> <table> <tbody> <tr> <th>Video ID</th> <th>Link</th> <th>Start time (s)</th> <th>End time (s)</th> </tr> </tbody> <tbody> <tr> <td>01_Dance_1</td> <td><a href="https://youtu.be/uOUVE5rGmhM">https://youtu.be/uOUVE5rGmhM</a></td> <td>8.54</td> <td>231.20</td> </tr> <tr> <td>03_Acrob_1</td> <td><a href="https://youtu.be/DjihbYg6F2Y">https://youtu.be/DjihbYg6F2Y</a></td> <td>4.24</td> <td>231.91</td> </tr> <tr> <td>04_Magic_1</td> <td><a href="https://youtu.be/CvzMqIQLiXE">https://youtu.be/CvzMqIQLiXE</a></td> <td>3.68</td> <td>348.17</td> </tr> <tr> <td>05_Dance_2</td> <td><a href="https://youtu.be/f4DZp0OEkK4">https://youtu.be/f4DZp0OEkK4</a></td> <td>5.05</td> <td>227.99</td> </tr> <tr> <td>06_Mime_2</td> <td><a href="https://youtu.be/u9wJUTnBdrs">https://youtu.be/u9wJUTnBdrs</a></td> <td>5.79</td> <td>347.05</td> </tr> <tr> <td>07_Acrob_2</td> <td><a href="https://youtu.be/kRqdxGPLajs">https://youtu.be/kRqdxGPLajs</a></td> <td>183.61</td> <td>519.27</td> </tr> <tr> <td>08_Magic_2</td> <td><a href="https://youtu.be/FUv-Q6EgEFI">https://youtu.be/FUv-Q6EgEFI</a></td> <td>3.36</td> <td>270.62</td> </tr> <tr> <td>09_Dance_3</td> <td><a href="https://youtu.be/LXO-jKksQkM">https://youtu.be/LXO-jKksQkM</a></td> <td>5.61</td> <td>294.17</td> </tr> <tr> <td>12_Magic_3</td> <td><a href="https://youtu.be/S84AoWdTq3E">https://youtu.be/S84AoWdTq3E</a></td> <td>1.76</td> <td>426.36</td> </tr> <tr> <td>13_Dance_4</td> <td><a href="https://youtu.be/0wc60tA1klw">https://youtu.be/0wc60tA1klw</a></td> <td>14.28</td> <td>217.18</td> </tr> <tr> <td>14_Mime_3</td> <td><a href="https://youtu.be/0Ala3ypPM3M">https://youtu.be/0Ala3ypPM3M</a></td> <td>21.87</td> <td>386.84</td> </tr> <tr> <td>15_Dance_5</td> <td><a href="https://youtu.be/mg6-SnUl0A0">https://youtu.be/mg6-SnUl0A0</a></td> <td>15.14</td> <td>233.85</td> </tr> <tr> <td>16_Mime_6</td> <td><a href="https://youtu.be/8V7rhAJF6Gc">https://youtu.be/8V7rhAJF6Gc</a></td> <td>31.64</td> <td>388.61</td> </tr> </tbody> </table> <h3><strong>MrBean Dataset</strong></h3> <p>Additionally, 9 participants watched an extra 24-minute clip from the first episode of Mr. Bean, where multiple (moving) objects may exist and interact, and the camera viewpoint may change. The subject IDs and the signal quality files are inherited from the single-shot dataset.</p> <table> <tbody> <tr> <th>Video ID</th> <th>Link</th> <th>Start time (s)</th> <th>End time (s)</th> </tr> </tbody> <tbody> <tr> <td>Mr_Bean</td> <td><a href="https://www.youtube.com/watch?v=7Im2I6STbms">https://www.youtube.com/watch?v=7Im2I6STbms</a></td> <td>39.77</td> <td>1495.00</td> </tr> </tbody> </table> <h2><strong>Acknowledgement</strong></h2> <p>This research is funded by the Research Foundation - Flanders (FWO) project No G081722N, junior postdoctoral fellowship fundamental research of the FWO (for S. Geirnaert, No. 1242524N), the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation program (grant agreement No 802895), the Flemish Government (AI Research Program), and the PDM mandate from KU Leuven (for S. Geirnaert, No PDMT1/22/009).</p> <p>We also thank the participants for their time and effort in the experiments.</p> <h2><strong>Contact Information</strong></h2> <p>Executive researcher: Yuanyuan Yao, <a href="mailto:yuanyuan.yao@kuleuven.be">yuanyuan.yao@kuleuven.be</a></p> <p>Led by: Prof. Alexander Bertrand, <a href="mailto:alexander.bertrand@kuleuven.be">alexander.bertrand@kuleuven.be</a></p> <p> </p>
Latency and energy characterization of 5G LDPC FEC Decoding on CPU and GPU
<p>CloudRIC is a system that meets specific reliability targets in 5G FEC processing while sharing pools of heterogeneous processors among DUs, which leads to more cost- and energy-efficient vRANs. The details of the solution are presented in <a title="CloudRIC: Open Radio Access Network (O-RAN) Virtualization with Shared Heterogeneous Computing" href="https://doi.org/10.1145/3636534.3649381">CloudRIC: Open Radio Access Network (O-RAN) Virtualization with Shared Heterogeneous Computing</a>. These repository provides a dataset, analyzed therein, with experiments carried out with different 5G LDPC decoding processors: (i) Intel FlexRAN library and two open-source alternative libraries on an Intel Xeon Gold 6240R CPU, and (ii) a proprietary driver on an NVIDIA GPU V100.</p> <p>See README file for a description of the dataset.</p>
Dataset for publication "Decoding Niobium Carbide MXene Dual Functional Photoactive Cathode in Photoenhanced Hybrid Zinc-Ion Capacitor"
<p>This is a dataset for publication "Decoding Niobium Carbide MXene Dual-Functional Photoactive Cathode in Photoenhanced Hybrid Zinc-Ion Capacitor". It contains all relevant data from the aforementioned publication. The details about how to read these data is contained within the readme file.</p>
Raman Signal Denoising using Fully Convolutional Encoder Decoder
<p>Test dataset used in the manuscript 'Raman Signal Denoising using Fully Convolutional Encoder Decoder'.</p>
Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments
<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2) A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p> </p> <p>Manuscript title Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>
Decoding Physical and Cognitive Impacts of Particulate Matter Concentrations at Ultra-fine Scales
<p>Data, plots, and software to accompany (unpublished) paper: Decoding Physical and Cognitive Impacts of Particulate Matter Concentrations at Ultra-fine Scales. This work uses an ultra-fine, holistic environmental and biometric sensing paradigm to generate empirical particulate matter models estimated by biometric variables.</p> <p>GitHub repository: <a href="https://github.com/mi3nts/DUEDARE">https://github.com/mi3nts/DUEDARE</a></p>
Image dataset to train a deep learning model to decode Leetspeak obfuscated characters
<p>The dataset contains an image database (18,981 images) that could be used to train a deep learning model to accurately detect characters. We have successfully used it to create a model that identifies characters encoded using LeetSpeak. The original dataset can be found in the Mondragon Unibertsitatea Repository -- https://gitlab.danz.eus/datasharing/ski4spam</p> <p>The training dataset consists of:</p> <p>- Alphabetic letters (a-z) written using different fonts and styles (regular, cursive, bold, cursive+bold)</p> <p>- Handwritten letters: English handwriting from the Chars74k dataset [2] which is available at http://www.ee.surrey.ac.uk/CVSSP/demos/chars74k/.</p>
Two light sensors decode moonlight versus sunlight to adjust a plastic circadian/circalunidian clock to moon phase
<p>Many species synchronize their physiology and behavior to specific hours. It is commonly assumed that sunlight acts as the main entrainment signal for ~24h clocks. However, the moon provides similarly regular time information. Consistently, a growing number of studies have reported correlations between diel behavior and lunidian cycles. Yet, mechanistic insight into the possible influences of the moon on ~24hr timers remains scarce.</p> <div> <div> <div class="msocomtxt"> <p class="MsoNormal"><span>We have explored the marine bristleworm </span><em><span>Platynereis dumerilii</span></em><span> to investigate the role of moonlight in the timing of daily behavior. We uncover that moonlight, besides its role in monthly timing, also schedules the exact hour of nocturnal swarming onset to the nights' darkest times. Our work reveals that extended moonlight impacts on a plastic clock that exhibits <24h (moonlit) or >24h (no moon) periodicity. Abundance, light sensitivity, and genetic requirement indicate that the <em>Platynereis </em>light receptor molecule r-Opsin1 serves as a receptor that senses moonrise, whereas the cryptochrome protein L-Cry<em> </em>is required to discriminate the proper valence of nocturnal light as either moon- or sunlight. Comparative experiments in <em>Drosophila </em>suggest that cryptochrome's principal requirement for light valence interpretation is conserved. Its exact biochemical properties differ, however, between species with dissimilar timing ecology.</span></p> <p class="MsoNormal"><span>Our work advances the molecular understanding of lunar impact on fundamental rhythmic processes, including those of marine mass spawners endangered by anthropogenic change.</span></p> </div> </div> </div>
Datasets for "Versatile Domain Mapping of Scanning Electron Nanobeam Diffraction datasets utilising Variational Auto Encoders and decoder-assisted latent clustering"
<p>20210925_152115_data.hdf5 is the P2 sample raw data.</p> <p>FinalMap-weights.hdf5 is the weights for the P2 model used for clustering</p> <p>SimulatedDSA-data.hdf5 is the simulated data set raw data.</p> <p>FullyTrainedModel.hdf5 is the weights for the Simulated Dataset model used for clustering</p> <p><br> </p>
Decoding the metabolic response of Escherichia coli for sensing trace heavy metals in water
<p>As: Raman spectra from E. coli lysate sample after exposing to As in DI water</p> <p>Cr: Raman spectra from E. coli lysate sample after exposing to Cr in DI water</p> <p>As_TapWater: Raman spectra from E. coli lysate sample after exposing to As in tap water</p> <p>WasteWater_FineTune_Dataset: Raman spectra from E. coli lysate sample after exposing to As in waste water</p> <p>WasteWater 'Unknow' Dataset: Raman spectra from E. coli lysate sample after exposing to waste water</p>
Decoding diabetes biomarkers and related molecular mechanisms using machine learning, text mining, and gene expression analysis
<p>The molecular basis of diabetes mellitus is yet to be fully elucidated. We aimed to identify the most frequently reported and differential expressed genes (DEGs) in diabetes using bioinformatics approaches. Text mining was used to screen 40,225 article abstracts from diabetes literature. These studies highlighted 5939 diabetes-related genes spread across 22 human chromosomes, with 112 genes mentioned in more than 50 studies. Among these genes, HNF4A, PPARA, VEGFA, TCF7L2, HLA- DRB1, PPARG, NOS3, KCNJ11, PRKAA2, and HNF1A were mentioned in more than 200 articles. These genes are correlated with the regulation of glycogen and polysaccharide, adipogenesis, AGE/RAGE, and macrophage differentiation. Three datasets (44 patients and 57 controls) were subjected to gene expression analysis. The analysis revealed 135 significant DEGs, of which CEACAM6, ENPP4, HDAC5, HPCAL1, PARVG, STYXL1, VPS28, ZBTB33, ZFP37 and CCDC58 were the top ten DEGs. These genes were enriched in aerobic respiration, T-Cell antigen receptor pathway, Tricarboxylic acid metabolic process, vitamin D receptor pathway, Toll-like receptor signaling, and endoplasmic reticulum (ER) unfolded protein response. The results of text mining and gene expression analyses used as attribute values for ML analysis . The "Decision tree", "Extra-tree regressor" and "Random forest" algorithms were used in ML analysis to identify unique markers that could be used as diabetes diagnosis tools. These algorithms produced prediction models with accuracy ranges from 0.6364 to 0.88 and overall confidence interval (CI) of 95%. There were 39 biomarkers that could distinguish diabetic and non-diabetic patients, 12 of which were repeated multiple times. The majority of these genes are associated with stress response, signalling regulation, locomotion, cell motility, growth, and muscle adaptation. ML algorithms highlighted the use of the HLA-DQB1 gene as a biomarker for diabetes early detection. Our data mining and gene expression analysis have provided useful information about potential biomarkers in diabetes.</p>
Data supporting "A real-time, scalable, fast and resource-efficient decoder for a quantum computer"
<p>Data includes the circuits (stim_circuits.zip) used to create samples to benchmark CC decoder across different noise rates and code sizes. The resulting accuracy and cycle data is in fpga_accuracy_data.csv. The memory footprint (in KB) of the algorithm for different code sizes is in fpga_memory_data.csv.</p> <p>Weights of syndromes for different noise rates for both phenomenological and circuit-level noise at distance d=23 and d=21 are in noise_rate_sampling_full_d23.csv and noise_rate_sampling_full_d21.csv respectively.</p>
Simulation results for "Localized statistics decoding: A parallel decoding algorithm for quantum low-density parity-check codes"
<p>This dataset contains simulations results presented in the paper "Localized statistics decoding: A parallel decoding algorithm for quantum low-density parity-check codes".</p> <p>The files are in `csv` file format, with data easily processable using the python library `sinter`.</p>
Data from: Reduced palatability, fast flight, and tails: Decoding the defence arsenal of Eudaminae skipper butterflies in a Neotropical locality
<p>Prey often rely on multiple defences against predators, such as flight speed, attack deflection from vital body parts, or unpleasant taste, but our understanding on how often and why they are co-exhibited remains limited. Eudaminae skipper butterflies use fast flight and mechanical defences (hindwing tails), but whether they use other defences like unpalatability (consumption deterrence), and how these defences interact, has not been assessed.</p> <p>We tested the palatability of 12 abundant Eudaminae species in Peru, using training and feeding experiments with domestic chicks. Further, we approximated the difficulty of capture explained by flight speed and quantified by wing loading. We performed phylogenetic regressions to find any association between multiple defences, body size, and habitat preference.</p> <p>We found a broad range of palatability in Eudaminae, within and among species. Contrary to current understanding, palatability was negatively correlated with wing loading, suggesting that faster butterflies tend to have lower palatability.</p> <p>The relative length of hind wing tails did not explain the level of butterfly palatability, showing that attack deflection and consumption deterrence are not mutually exclusive. Habitat preference (open or forested environments) did not explain the level of palatability either, although butterflies with high wing loading tended to occupy semi-closed or closed habitats.</p> <p>Finally, the level of unpalatability in Eudaminae is size dependent. Larger butterflies are less palatable, perhaps because of higher detectability/preference by predators. Altogether, our findings shed light on the contexts favouring the prevalence of single vs. multiple defensive strategies in prey.</p>
Deep in the Bowel: Highly Interpretable Neural Encoder-Decoder Predicts Gut Metabolites from Gut Microbiome: Supplemental Data
<p>This upload contains the supplemental data to the manuscript titled, "Deep in the Bowel: Highly Interpretable Neural Encoder-Decoder Predicts Gut Metabolites from Gut Microbiome". It contains the transformed data used to train the model (*-grouped-clr.csv), the latent feature space (latent_z.txt), and the differential abundance results (*DA*.csv).</p>
Decoding the Epigenetics and Chromatin Loop Dynamics of Androgen Receptor-Mediated Transcription
<p>This repository stores the datasets for the "Decoding the Dynamic Regulation and Chromatin Architecture of Androgen Receptor-mediated Gene Expression" paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.