Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,139

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,139 results for “recognition”

Learn how ShareScore rates datasets ↗
zenodo36/100

Molecular basis for the increased affinity of an RNA recognition motif with re-engineered specificity: A molecular dynamics and enhanced sampling simulations study.

<p>This repository contains the representative structures of the 20 clusters obtained, which constitute the &ldquo;MD-adapted structure ensemble&rdquo;: i.e., sets of atomic coordinates&nbsp;that capture the flexibility and the pre-miR20b (<a href="https://zenodo.org/api/files/ee12021f-4398-465a-9ff6-ddb7be32765f/ensemble_MD_2n7x.pdb?versionId=310a80f6-aa64-445d-8641-45faf9f1ac03">ensemble_MD_2n7x.pdb</a>)&nbsp; and Rbfox/pre-miR20b (<a href="https://zenodo.org/api/files/ee12021f-4398-465a-9ff6-ddb7be32765f/ensemble_MD_2n82.pdb?versionId=82d7afcb-8a92-4daa-8150-789dbd7b2474">ensemble_MD_2n82.pdb</a>) conformers suggested by MD simulations while still retaining the highest possible level of agreement with the primary NMR data.</p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

Dataset from 'Billino, J., van Belle, G., Rossion, B., & Schwarzer, G. (2018). The nature of individual face recognition in preschool children: Insights from a gaze-contingent paradigm. Cognitive Development, 47, 168-180. DOI: 10.1016/j.cogdev.2018.06.007

<p>The folder contains a data file and a description file providing column labels.</p> <p>For further questions, please contact:<br> jutta.billino[at]psychol.uni-giessen.de</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

To bee or not to bee: An annotated dataset for beehive sound recognition

<p><strong>-- Dataset documentation --</strong></p> <p><br> <strong>1- Introduction</strong></p> <p>The present dataset was developed in the context of our work in [1] that focus on the automatic recognition of beehive sounds. The problem is posed as the classification of sound segments in two classes: Bee and noBee. The novelty of the explored approach and the need for annotated data, dictated the construction of such dataset.</p> <p><strong>2- Description</strong></p> <p><strong>2.1- Audio recordings:</strong></p> <p>The annotated dataset was developed based on a selected set of recordings acquired in the context of two different projects: the Open Source Beehive (OSBH) project [2] and the NU-Hive project [3]. Both projects main goal is to develop a beehive monitoring system capable of identifying and predict certain events and states of the hive that are of interest to the beekeeper. Among many different variables that can be measured and that help the recognition of different states of the hive, the analysis and use of the sound the bees produce is a big focus for both projects.</p> <p>The recordings from the OSBH project were acquired through a citizen science initiative which asked people from the general public to record the sound from their beehives together with the registering of the hive state at the moment. Because of the amateur and collaborative nature of this project, the recordings from the OSBH project present great diversity due to the very different conditions in which the signals were acquired: different recording devices used, different environments where the hives were placed, and even different position for the microphones inside the hive. This variety of settings makes this dataset a very interesting tool to help evaluate and challenge the methods developed.</p> <p>The NU-Hive project is a comprehensive effort of data acquisition, concerning not only sound, but a vast amount of variables that will allow the study of bees behaviors and other unknown aspects. The selected recordings are taken from 2 hives and labeled regarding two states: queen bee is present, and queen bee not present. Contrary to the OSBH project recordings, the recordings from&nbsp;the NU-Hive project are from a much more controlled and homogeneous environment. Here the occurring external sounds are mainly traffic, car honks and birds.</p> <p><strong>The annotated dataset:</strong></p> <p>For each selected recording, time segments are labeled as Bee or noBee depending on the perceived source of the sound signal being from bees or external to the hive.</p> <p>The whole annotated dataset consists of 78 recordings of varying lengths which make up for a total duration of approximately 12 hours of which 25% is annotated as noBee events.</p> <p>About 60% of the recordings are from the NU-Hive dataset and represent 2 hives, the remaining are recordings from the OSBH dataset and 6 different hives. The recorded hives are from 3 main locations: North America, Australia and Europe.</p> <p>&nbsp;</p> <p><strong>2- Annotation procedure<a href="http://localhost:8888/notebooks/Dropbox/QMUL/BEESzzzz/Data/Annotations/readme.ipynb#2--Annotation-procedure">&para;</a></strong></p> <p>The annotation procedure consists in hearing the selected recordings and marking the beginning and the end of every sound that could not be recognized as a beehive sound. The recognition of external sounds is based primarily on the perceived heard sounds, but a visual aid is also used by visualizing the log-mel frequency spectrum of the signal. All the above are functionalities offered by the Sonic Visualiser software, which was used by two volunteers that are neither bee-specialists nor specially trained in sound annotation tasks.</p> <p>By marking these pairs of moments corresponding to the beginning and end of external sound periods, we are able to get the whole recording labeled into&nbsp;Bee and noBee intervals. Thus in the resulting Bee intervals only pure beehive sounds, (no external sounds) should be perceived for the entirety of the segment. The noBee intervals refer to periods where an external sound can be perceived (superimposed to the bee sounds).</p> <p>&nbsp;</p> <p><strong>File Structure:</strong></p> <p>Each audio file is coupled with its corresponding annotation file, identified by the same name and extension <em>.lab</em>.<br> For convenience, all the annotations are collected in a single master label file named <em>beeAnnotations.mlf</em></p> <p>The <em>.lab</em>&nbsp;files consist of :&nbsp;</p> <ul> <li>First row identifies the audio file to which the annotations refer to.</li> <li>Each line after that describes an interval with starting time point, end time point and label. The time points are expressed in seconds.</li> </ul> <p>Below is an example of such an annotation file:&nbsp;</p> <pre><code>Hive3_20_07_2017_QueenBee_H3_audio_15_30_00 0 78.45 bee 78.46 78.95 nobee 78.96 103.92 bee 103.93 112.48 nobee 112.49 152.48 bee . </code></pre> <p>This dataset is licensed under a Creative Commons Attribution 4.0 International License.<br> When using this dataset, please cite [1]:</p> <p>[1] I. Nolasco and E. Benetos, &ldquo;To bee or not to bee: Investigating machine learning approaches to beehive sound recognition,&rdquo; in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2018, submitted.</p> <p>[2] &ldquo;Open Source Beehives Project,&rdquo; https://www.osbeehives.com/.</p> <p>[3] S. Cecchi, A. Terenzi, S. Orcioni, P. Riolo, S. Ruschioni, and N. Isidoro, &ldquo;A preliminary study of sounds emitted by honey bees in a beehive,&rdquo; in Audio Engineering Society Convention 144, 2018.</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

Molecular basis for the increased affinity of an RNA recognition motif with re-engineered specificity: A molecular dynamics and enhanced sampling simulations study.-PART 8

<p>Simulations of the miR20b&nbsp;RNA with the Case vdW modification to amber force field and&nbsp;OPC water molecules.</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

ICDAR2013 – Handwritten Digit and Digit String Recognition Competition

<p>The CVL Single Digit dataset consists of 7000 single digits (700 digits per class) written by approximately 60 different writers. The validation set has the same size but different writers. The validation set may be used for parameter estimation and validation but not for supervised training. The CVL Digit Strings dataset uses 10 different digit strings from a total of about 120 writers resulting in 1262 training images. The digits from the CVL Single Digit dataset were extracted from these strings.</p> <p>This database may be used for non-commercial research purpose only. If you publish material based on this database, we request you to include a reference to:</p> <p>Markus Diem, Stefan Fiel, Angelika Garz, Manuel Keglevic, Florian Kleber and Robert Sablatnig, <em>ICDAR 2013 Competition on Handwritten Digit Recognition (HDRC 2013)</em>, In Proc. of the 12th Int. Conference on Document Analysis and Recognition (ICDAR) 2013, pp. 1454-1459, 2013.</p>

opencc-by-nc-4.0Nov 2018View details →
zenodo36/100

Fig. 1. – 50 in Enlarging the monotypic Monocarpieae (Annonaceae, Malmeoideae): recognition of a second genus from Vietnam informed by morphology and molecular phylogenetics

Fig. 1. – 50% majority-rule consensus phylogram derived from Bayesian inference of combined seven plastid DNA regions. Bayesian posterior probabilities (PP) indicated on the right; maximum likelihood bootstrap (BS) percentages in the middle; parsimony symmetric resampling (SR) percentages on the left [** denotes BS/SR &lt;50%]. DEN. = Dendrokingstonieae; MAL. = Malmeeae; MIL. = Miliuseae; MON. = Monocarpieae; PIP. = Piptostigmateae. Scale bar unit = substitutions per site.

opencc-by-4.0Nov 2018View details →
zenodo36/100

Example of using of two-stage graphic primitives recognition system

<p>Example of using of two-stage graphic primitives recognition system contains one video file with demonstration of result of work.</p>

opencc-by-4.0Apr 2019View details →
zenodo36/100

Dataset for Evaluating Pedalling Techniques Recognition Using Gesture Data

<p>With the help of a dedicated measurement system (see reference for the details of the system),&nbsp;the pedalling gestures and the piano sound can be synchronously recorded at an audio sampling rate and a high resolution.&nbsp;The measurement system was deployed on the sustain pedal of a Yamaha baby grand piano situated in the studios at Queen Mary University of London. Ten well known passages of Chopin&#39;s piano music were selected to form this dataset. Therefore the dataset consists of:</p> <p>- <strong>audio-data.zip</strong>:&nbsp;piano sound of the ten passages&nbsp;recorded at 44.1kHz, each&nbsp;saved as &quot;<strong>PASSAGE.wav</strong>&quot;.</p> <p>- <strong>pedal-data.zip</strong>: associated gesture data recorded at 22.05kHz, each saved as &quot;<strong>PASSAGE.npy</strong>&quot;. The gesture data correspond&nbsp;to the movement trajectory of the sustain pedal.</p> <p>- <strong>pedal-label.zip</strong>: label the &quot;continuous&quot; gesture data by &quot;discrete&quot; pedalling techniques at every 0.02 second.&nbsp;Label 0-4 represents none, 1/4, 1/2, 3/4 and full pedalling technique, respectively.&nbsp;Labels for gesture data&nbsp;&quot;<strong>PASSAGE.npy</strong>&quot; are saved in &quot;<strong>PASSAGE-label.npy</strong>&quot;.</p> <p>- <strong>passage.zip</strong>: music scores of the ten passages, each saved as &quot;<strong>PASSAGE.pdf</strong>&quot;. They were annotated with pedalling techniques by the experimenter in advance and then performed by a pianist, who was asked to follow the annotated scores. The resulting audio recording and gesture data formed the above &quot;<strong>PASSAGE.wav</strong>&quot; and&nbsp;the &quot;<strong>PASSAGE.npy</strong>&quot;. The annotated score guided the labelling process and formed the &quot;<strong>PASSAGE-label.npy</strong>&quot;.</p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

ICDAR 2019 Competition on Table Detection and Recognition (cTDaR)

<p>The aim of this competition is to evaluate the performance of state of the art methods for table detection (TRACK A) and table recognition (TRACK B). For the first track, document images containing one or several tables are provided. For TRACK B two subtracks exist: the first subtrack (B.1) provides the table region. Thus, only the table structure recognition must be performed. The second subtrack (B.2) provides no a-priori information. This means, the table region and table structure detection has to be done. The Ground Truth is provided in a similar format as for the ICDAR 2013 competition (see [2]):</p> <p>&lt;?xml version=&quot;1.0&quot; encoding=&quot;UTF-8&quot;?&gt;</p> <p>&lt;<strong>document</strong> filename=&#39;filename.jpg&#39;&gt;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&lt;<strong>table</strong> id=&#39;Table_1540517170416_3&#39;&gt;</p> <p><strong>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&lt;Coords points=&quot;180,160 4354,160 4354,3287 180,3287&quot;/&gt;</strong></p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&lt;<strong>cell</strong> id=&#39;TableCell_1540517477147_58&#39; <strong>start-row</strong>=&#39;0&#39; <strong>start-col</strong>=&#39;0&#39; <strong>end-row</strong>=&#39;1&#39; <strong>end-col</strong>=&#39;2&#39;&gt;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&lt;<strong>Coords</strong> <strong>points</strong>=&quot;180,160 177,456 614,456 615,163&quot;/&gt;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&lt;/cell&gt;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;...</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&lt;/table&gt;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;...</p> <p>&lt;/document&gt;</p> <p>&nbsp;</p> <p>The difference to Gobel et al. [2] is the Coords tag which defines a table/cell as a polygon specified by a list of coordinates. For B.1 the table and its coordinates is given together with the input image.</p> <p>Important Note:</p> <p>For the modern dataset, the convex hull of the content describes a cell region. For the historical dataset, it is requested that the output region of a cell is the cell boundary. This is necessary due to the characteristics of handwritten text, which is often overlapping with different cells.</p> <p>See also: http://sac.founderit.com/tasks.html</p> <p>The evaluation tool is available at github: https://github.com/cndplab-founder/ctdar_measurement_tool</p>

opencc-by-4.0Apr 2019View details →
zenodo36/100

T-REC Song Recognition Dataset

<p>This csv file (tab delimited) contains the track names, artist(s), total number of responses,&nbsp;measured recognition (user study), computed recognition (T-REC) and measured recognition on specific demographics i.e. male, female and age groups 18-24, 25-34, 35-44, 45-54, 55-65 for 100 music tracks used&nbsp;in the paper &quot;Data-driven song recognition estimation using collective memory dynamics models&quot; accepted for publication in the ISMIR 2019 conference.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Supervised Molecular Dynamics Movies from: Deciphering the molecular recognition mechanism of multidrug resistance Staphylococcus aureus NorA efflux pump using a Supervised Molecular Dynamics approach

<p>Molecular Recognition pathway of Supervised molecular dynamic simulations of MdfA-CLM NorA-CPX and&nbsp;&nbsp;NorA-CPX.</p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Deciphering the molecular recognition mechanism of multidrug resistance Staphylococcus aureus NorA efflux pump using a Supervised Molecular Dynamics approach.

<p><strong>Legend of Movie-S1</strong></p> <p>The Movie is composed by four synchronized and animated panels that show different aspects of the SuMD simulation. The time evolution is reported in nanosecond. In the first panel (upper left), the molecular representation of the system is shown. The MdfA backbone is represented by the new cartoon style (cyan). The CLM is shown in yellow and by a transparent surface. The protein residues within 3 &Aring; from the ligand are made explicit by a stick representation.</p> <p>In the second panel (upper-right), the CM-distance between the protein and the ligand centers of mass is reported.</p> <p>In the third panel (lower left), the MMGBSA energy profile is reported.</p> <p>In the fourth panel (lower-right) cumulative electrostatic interactions are reported for the 15 MdfA residues most contacted by CLM during the whole simulation.</p> <p>&nbsp;</p> <p><strong>Legend of Video-S2</strong></p> <p>The Movie shows the SuMD trajectory of CLM on MdfA compared to the CLM crystallographic pose. MdfA is represented in cyan new cartoon transparency. The crystallographic pose is showed in yellow while the experimental one in light green. At 16.69 ns a RMSD value of 1.77 &Aring; is highlighted.</p> <p>&nbsp;</p> <p><strong>Legend of Video-S3</strong></p> <p>The Movie is composed by four synchronized and animated panels that show different aspects of the SuMD simulation. The time evolution is reported in nanosecond. In the first panel (upper left), the system is shown. The NorA backbone is represented by the new cartoon style (red) and the protein residues within 3 &Aring; of CPX are showed in stick. CPX is rendered by a green stick.</p> <p>In the second panel (upper-right), the distance between the centre of mass of the ligand and the protein during the trajectory is reported.</p> <p>In the third panel (lower left), the MMGBSA energy profile is reported. In the fourth panel (lower-right) cumulative electrostatic interactions are reported for the 15 NorA residues most contacted by CPX during the whole SuMD trajectory.</p> <p>It is important to note that the following video has a duration that is half of the simulation of SuMD. However, this straid does not alter the description of the trajectory performed by the ligand.</p> <p>&nbsp;</p> <p><strong>Legend of Video-S4</strong></p> <p>The Movie depicts the clustering analysis of CPX during the whole SuMD simulation. The NorA protein is shown in red new cartoon transparency. CPX is rendered by a light-green stick and by a transparent surface. The spheres are shown in 7 different colours, according to the different clusters. Each sphere dimension is in according to the cluster dimensions. After a first recognition site, the ligand conformations are clustered in different sites of the NorA channel. It is important to note that the following video has a duration that is half of the simulation of SuMD. However, this straid does not alter the description of the trajectory performed by the ligand.</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Occupancy Sensing and Activity Recognition with Cameras and Wireless Sensors

<p>This dataset contains human activity data from a&nbsp;wireless sensing system, which includes a Doppler motion sensor and a wireless network.&nbsp;The Doppler sensor is a low-cost dual Doppler sensor modified from a commercial-off-the-shelf range-controlled radar, which operates at 5.8 GHz with two directional antennas. The wireless network uses four IEEE 802.15.4 radio nodes (CC2531 from TI) to create a mesh network to measure the RSS between each pair of radio nodes operating on the 16 frequency channels at 2.4 GHz.</p> <p>For the activity experiment, we recruited human subjects to perform 42 trials of four activities &nbsp;(each one with two minutes duration): (1) &nbsp;walking in a room (10 trials), (2) sitting in a chair (10 trials), (3) lying on a bed (12 trials), and (4) body turning on a bed (10 trials).&nbsp;For the walking activity, the human subjects walk along different paths at different locations in the room. For the lying on bed activity, we ask human subjects to breathe normally on bed with three orientations facing upwards, right and left. Finally, for the turning on bed case, human subjects turn their bodies from one side to the other on bed with random time intervals. We also recorded two-minute data of the empty room case before and after each human subject trial. Note that each data file name has its&nbsp;corresponding activity&nbsp;in it, so it is pretty self-explanatory.&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Supplementary Data for "DNA hairpin base-flipping dynamics drives APOBEC3A recognition and selectivity"

<p>Contains CSV-formatted files with RMSD and Sugar Pucker averages and standard deviations for 3- and 4-nt hairpin loop simulations as described in the manuscript "DNA hairpin base-flipping dynamics drives APOBEC3A recognition and selectivity".</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

System Fingerprint Recognition for Deepfake Audio (SFR) - Compressed Set

<div>The rapid progress of deep speech synthesis models&nbsp;has posed significant threats to society such as malicious manip</div> <div>ulation of content. This has led to an increase in studies aimed&nbsp;at detecting so-called &ldquo;deepfake audio&rdquo;. However, existing works</div> <div>focus on the binary detection of real audio and fake audio. In&nbsp;real-world scenarios such as model copyright protection and</div> <div>digital evidence forensics, it is needed to know what tool or&nbsp;model generated the deepfake audio to explain the decision. This</div> <div>motivates us to ask: &lsquo;Can we recognize the system fingerprints&nbsp;of deepfake audio?&rsquo; In this paper, we present the first deepfake</div> <div>audio dataset for System Fingerprint Recognition (SFR) and&nbsp;conduct an initial investigation. We collected the dataset from</div> <div>the speech synthesis systems of seven Chinese vendors that use&nbsp;the latest state-of-the-art deep learning technologies, including</div> <div>both clean and compressed sets. In addition, we provide extensive benchmarks and research findings to facilitate the further development of system fingerprint recognition methods. The dataset is publicly available.&nbsp;</div> <div>&nbsp;</div> <div>The subsets 01, 02, and 03 represent the training set, development set, and test set, respectively.</div> <div>&nbsp;</div> <div> <div>This data set is licensed with a CC BY-NC-ND 4.0 license.</div> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

System Fingerprint Recognition for Deepfake Audio (SFR) - Clean Set

<div>The rapid progress of deep speech synthesis models&nbsp;has posed significant threats to society such as malicious manip</div> <div>ulation of content. This has led to an increase in studies aimed&nbsp;at detecting so-called &ldquo;deepfake audio&rdquo;. However, existing works</div> <div>focus on the binary detection of real audio and fake audio. In&nbsp;real-world scenarios such as model copyright protection and</div> <div>digital evidence forensics, it is needed to know what tool or&nbsp;model generated the deepfake audio to explain the decision. This</div> <div>motivates us to ask: &lsquo;Can we recognize the system fingerprints&nbsp;of deepfake audio?&rsquo; In this paper, we present the first deepfake</div> <div>audio dataset for System Fingerprint Recognition (SFR) and&nbsp;conduct an initial investigation. We collected the dataset from</div> <div>the speech synthesis systems of seven Chinese vendors that use&nbsp;the latest state-of-the-art deep learning technologies, including</div> <div>both clean and compressed sets. In addition, we provide extensive benchmarks and research findings to facilitate the further development of system fingerprint recognition methods. The dataset is publicly available.&nbsp;</div> <div>&nbsp;</div> <div>The subsets 01, 02, and 03 represent the training set, development set, and test set, respectively.</div> <div>&nbsp;</div> <div> <div>This data set is licensed with a CC BY-NC-ND 4.0 license.</div> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

RIS Based Hand Gesture Recognition Dataset

<h1><strong>RIS Based Hand Gesture Recognition Dataset</strong></h1> <h2><strong>Overview</strong></h2> <div>This dataset contains images for gesture recognition, divided into two main sets: dataset0608 and data_synthetic_variab. The data was collected using a wooden hand.&nbsp; &nbsp;</div> <div>&nbsp;</div> <h3>dataset0608</h3> <div>This dataset consists of two modes: ris_random and ris_optimized. The main difference between the two subfolders is the configuration of the RIS (random or optimized).</div> <div>&nbsp;</div> <div>This dataset consists of four subfolders: ris_random, ris_random2, ris_optimized, and ris_optimized2. The main difference between the subfolders is the format of the data:</div> <div>- ris_random and ris_optimized: Data is stored in individual files for each frame, named as 'frame_{i}{posture}{n_med}'&nbsp;</div> <div>- ris_random2 and ris_optimized2: Data has already been processed and combined into single files for all frames using the compact_files_frames.txt function, named as 'all_frames_{posture}_{n_med}'&nbsp;</div> <div>&nbsp;</div> <div>For each gestures = {close, two, open}, we have n_med values from 0 to 114 and 10 frames. Therefore, the ris_random and ris_optimized folders contain 10 frames &times; 115 measurements &times; 3 gestures = 3450 files, while the ris_random2 and ris_optimized2 folders contain 1 &times; 115 measurements &times; 3 gestures = 345 files.</div> <div>&nbsp;</div> <h3>data_synthetic_variab</h3> <div>This dataset consists of two modes: ris_random and ris_optimized. The main difference between the two subfolders is the configuration of the RIS (random or optimized).&nbsp;</div> <div>&nbsp;</div> <div>This dataset consists of four subfolders: ris_random, ris_random2, ris_optimized, and ris_optimized2. The main difference between the subfolders is the format of the data:</div> <div>- ris_random and ris_optimized: Data is stored in individual files for each frame, named as 'frame_{i}{posture}{n_med}'&nbsp;</div> <div>- ris_random2 and ris_optimized2: Data has already been processed and combined into single files for all frames using the compact_files_frames.txt function, named as 'all_frames_{posture}_{n_med}'&nbsp;</div> <div>&nbsp;</div> <div>For each gestures = {close, two, open},&nbsp; we have n_med values from 0 to 8 and 10 frames. This dataset provides additional synthetic data with variations in hand position to increase the dataset's diversity. Each gesture is represented by 8 different ways, where the hand position was slightly modified between each sample. These real data were used as a basis for generating synthetic data. By using the functions in the files "multiply_files.txt" and "add_gaussian_noise.txt," the dataset was expanded and made more realistic by adding Gaussian noise to the images.</div> <div>&nbsp;</div> <div>Therefore, the ris_random and ris_optimized folders contain 10 frames &times; 8 measurements &times; 3 gestures = 240 files, while the ris_random2 and ris_optimized2 folders contain 1 &times; 8 measurements &times; 3 gestures = 24 files.</div> <div>&nbsp;</div> <h3>Functions</h3> <div>* **add_gaussian_noise.txt:** This script adds Gaussian noise to the images to simulate real-world conditions and improve the robustness of the model.</div> <div>* **compact_files_frames.txt:** This script combines multiple frames into a single image, which can be useful for certain types of analysis.</div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction

<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models.&nbsp;</p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p>&nbsp; &nbsp; English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br>&nbsp; &nbsp; German: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br>&nbsp; &nbsp; Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br>&nbsp; &nbsp; Greek: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br>&nbsp; &nbsp; Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br>&nbsp; &nbsp; Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p>&nbsp;</p> <p><strong>Dataset Building&nbsp;</strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model&nbsp;</li> <li>Translation into the targeted language</li> <li>Word Alignment&nbsp;</li> <li>Data Augmentation&nbsp;</li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website:&nbsp;</span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

DAGHAR: A Benchmark for Domain Adaptation and Generalization in Smartphone-Based Human Activity Recognition

<p>DAGHAR benchmark is a curated dataset collection designed for domain adaptation and domain generalization studies in HAR tasks, using inertial sensors such as accelerometers and gyroscopes, from "A benchmark for domain adaptation and generalization in smartphone-based human activity recognition" work.&nbsp;It features raw inertial sensor data sourced exclusively from smartphones. Six public datasets were selected and standardized in terms of accelerometer units of measurement, sampling rate, gravity component, activity labels, user partitioning, and time window size. This standardization process allows for creating a comprehensive benchmark for evaluating the generalization capabilities of HAR models in cross-dataset scenarios.</p> <p>The benchmark is based on the following datasets:</p> <ul> <li><strong>Ku-HAR</strong>, from "Sikder, N. and Nahid, A.A., 2021. KU-HAR: An open dataset for heterogeneous human activity recognition. Pattern Recognition Letters, 146, pp.46-54", avaliable at <a href="https://data.mendeley.com/datasets/45f952y38r/5">Mendeley</a>. Distributed under CC BY 4.0.</li> <li><strong>MotionSense</strong>, from "Malekzadeh, M., Clegg, R.G., Cavallaro, A. and Haddadi, H., 2019, April. Mobile sensor data anonymization. In Proceedings of the international conference on internet of things design and implementation (pp. 49-58)", available at <a href="https://www.kaggle.com/datasets/malekzadeh/motionsense-dataset" target="_blank" rel="noopener">Kaggle</a>. Distributed under Open Data Commons Open Database License (ODbL) v1.0.</li> <li><strong>RealWorld</strong>, from "Sztyler, T. and Stuckenschmidt, H., 2016, March. On-body localization of wearable devices: An investigation of position-aware activity recognition. In 2016 IEEE international conference on pervasive computing and communications (PerCom) (pp. 1-9). IEEE", available at <a href="https://www.uni-mannheim.de/dws/research/projects/activity-recognition/dataset/dataset-realworld/" target="_blank" rel="noopener">this link</a>. We obtained explicitly permission to distribute a copy of the preprocessed data from the original authors.</li> <li><strong>UCI-HAR</strong>, from "Reyes-Ortiz, J.L., Oneto, L., Sam&agrave;, A., Parra, X. and Anguita, D., 2016. Transition-aware human activity recognition using smartphones. Neurocomputing, 171, pp.754-767", available at <a href="https://archive.ics.uci.edu/dataset/240/human+activity+recognition+using+smartphones">UCI Repository</a>. Distributed under CC BY 4.0.</li> <li><strong>WISDM</strong>, from "Weiss, G.M., Yoneda, K. and Hayajneh, T., 2019. Smartphone and smartwatch-based biometrics using activities of daily living. Ieee Access, 7, pp.133190-133202", available at <a href="https://archive.ics.uci.edu/dataset/507/wisdm+smartphone+and+smartwatch+activity+and+biometrics+dataset">UCI repository</a>. Distributed under CC BY 4.0.</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Annotated Dataset for Named Entity Recognition and Relation Extraction in French Building Technical Specifications (BTS)

<p>This dataset contains 233 raw requirements extracted from French Building Technical Specifications (BTS), referred to as "<a href="https://www.aglo.ai/cctp/#:~:text=Le%20CCTP%20(Cahier%20des%20Clauses%20Techniques%20Particuli%C3%A8res)%20est%20un%20document,code%20du%20march%C3%A9%20public%20donc."><em>Cahier des Clauses Techniques Particuli&egrave;res (CCTP)</em></a>", specifically focused on carpentry ("<em>lot menuiserie</em>") in public French construction projects. The requirements have been collected from 72 CCTP documents, resulting in a total of 19,725 sentences and 651,948 words.</p> <p>The dataset has been annotated using <a title="Open-source text annotation tool" href="https://github.com/doccano/doccano">Doccano </a>for Named Entity Recognition (NER) and Relation Extraction (RE). The annotations involve identifying entities and the relationships between them within the domain of building requirements. This dataset is intended for research on Natural Language Processing (NLP) models for Requirements Engineering (RE) in the Architecture, Engineering, and Construction (AEC) sector. Potential applications include requirements extraction, compliance analysis, and knowledge management in construction.</p> <p>The dataset includes the following components:</p> <ol> <li><strong>CCTP Documents</strong>: The original CCTP files from which the raw requirements were extracted.</li> <li><strong>Annotated Dataset</strong>: A JSONLines file containing the annotated dataset, including labels for Named Entity Recognition (NER) and Relation Extraction (RE).</li> </ol> <p>Key features of the dataset:</p> <ul> <li>Language: French</li> <li>Number of requirements: 233</li> <li>Number of sentences: 19,725</li> <li>Number of words: 651,948</li> <li>Annotation tasks: Named Entity Recognition (NER) and Relation Extraction (RE)</li> </ul> <p>This dataset is relevant for NLP research focused on structured information extraction from domain-specific texts in the construction industry.</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record