Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6 results for “machine perception”

Learn how ShareScore rates datasets ↗
zenodo40/100

Molecular similarity perception based on machine-learning models

<p>Molecular similarity is an particularly important notion for chemical legislation, specifically in the evaluation process for orphan drugs (i.e., drugs for rare diseases). A new molecule needs to be dissimilar from any other existing drug for a given disease to be assigned the financially advantageous status of orphan drug. Currently, there are many ways to define whether two molecules are similar or dissimilar. Thus far, the European Medicines Agency has used experts majority voting on discretional judgments of similarity when assessing new drugs for rare diseases. The decision of individual expert whether two compounds are similar is inherently subjective, depending on factors such as gender, age, state of mind, and previous experiences. It is therefore desirable, in this context, to benefit from an objective measure of similarity. To answer this need, we report a new dataset of molecular similarity assessments, that includes complex and difficult similarity scenarios. As a result, we propose new and improved models for similarity-prediction procedures, including 3D properties. These models are publicly available: <a href="https://chematlas.chimie.unistra.fr/ReadySim/">https://chematlas.chimie.unistra.fr/ReadySim/</a>.</p> <p>Software, 3D structures and pictures are available in the git related to this deposit:&nbsp;<a href="https://github.com/enricogandini/paper_similarity_prediction.git">https://github.com/enricogandini/paper_similarity_prediction.git</a></p> <p>The deposit contains two files.</p> <ul> <li>original_training_set.csv: this is one of the dataset published initially in [doi: 10.1186/1758-2946-6-5].</li> <li>new_dataset.csv: result from a new survey organized in 2020</li> </ul> <p>The columns are the following:</p> <ul> <li>id_pair: unique identifier of the compound pair</li> <li>curated_smiles_molecule_a: first compound of the pair</li> <li>curated_smiles_molecule_b: second compound of the pair</li> <li>tanimoto_cdk_Extended: ECFP similarity measure</li> <li>TanimotoCombo: ComboScore similarity measure</li> <li>pchembl_distance: difference of activity of the compound pair</li> <li>target_name: protein to which the compound pair is binding</li> <li>simil_2D: similar based on ECFP (0 or 1)</li> <li>simil_3D: similar based on ComboScore&nbsp;(0 or 1)</li> <li>dissimil_2D: dissimilar based on ECFP&nbsp;(0 or 1)</li> <li>dissimil_3D: dissimilar based on ComboScore&nbsp;(0 or 1)</li> <li>pair_type: pairs are classified based on ECFP and ComboScore as similar or dissimilar in 2D and 3D - Sim2DSim3D, Sim2DDis3D, Dis2D,Sim3D, Dis2DSim3D</li> <li>n_answers: number of answers from experts</li> <li>n_similar: number of answers labeling the pair as similar compounds</li> <li>frac_similar: n_similar/n_answers</li> </ul>

openmit-licenseApr 2022View details →
zenodo40/100

Supplementary Material for the paper "Developers' Perception Matters: Machine Learning to Detect Developer-sensitive Smells"

<p>Supplementary material for the paper &quot;Daniel Oliveira; Wesley K. G. Assun&ccedil;&atilde;o; Alessandro Garcia; Baldoino Fonseca; and M&aacute;rcio Ribeiro. Developers&#39; Perception Matters: Machine Learning to Detect Developer-sensitive Smells. In: Empirical Software Engineering. 2022. Springer.&quot;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

BRAIN Journal-Brain-Like Artificial Intelligence for Automation-Figure 6. Overview of Brain-Inspired Architecture for Machine Perception

<p>Figure 6 gives an overview about the developed architecture for human-like machine perception which bases on insights about the working mechanisms of the human perceptual system. The central element of the model is the so-called &ldquo;neuro-symbolic network&rdquo;, which processes data coming from different sensor sources and additionally considers information coming from<br> &ldquo;higher-level&rdquo; sources referred to as memory, knowledge, and focus of attention . Within the neuro-symbolic network, so called &ldquo;neuro-symbolic information processing&rdquo; takes place based on information exchange of &ldquo;neuro-symbols&rdquo;. The focus in this article will be on the description of the<br> functioning of neuro-symbols and the neuro-symbolic network. Details about the other modules and functional aspects of the model can amongst others be found in.</p>

opencc-by-4.0Oct 2013View details →
zenodo32/100

A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots

<p>This&nbsp;dataset is a part of the supplementary materials to the 2017 RAL&nbsp;<a href="https://ieeexplore.ieee.org/document/7358076">article</a>&nbsp;with the same title.</p> <blockquote> <p>A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots</p> <p>IEEE Robotics and Automation Letters</p> <p>Alessandro Giusti, Jerome Guzzi, Dan Ciresan, Fang Lin He, Juan Pablo Rodriguez, Flavio Fontana, Matthias Faessler, Christian Forster, Jurgen Schmidhuber, Gianni A. Di Caro, Davide Scaramuzza, Luca Gambardella</p> </blockquote> <p>You can find more information on&nbsp;the&nbsp;<a href="http://bit.ly/perceivingtrails">project web page</a>&nbsp;(alessandrog@idsia.ch).</p> <p><strong>Dataset</strong></p> <p>Folders 001..010 contain the dataset used to train the networks. Folder 000 contains preliminary test data. Folders 011..014 contain data for testing the system.</p> <ul> <li>000 and 003 were shot with an handheld cellphone.</li> <li>001 and 002 were shot with 3 GOPRO Hero 3 cameras, fixed on the head with straps.</li> <li>004..014 were shot with 3 Bluefox cameras, fixed on a rigid helm (the same model and with the same lens as the camera mounted on the quadcopter).</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo12/100

DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception

<p>DEMoS (Database of Elicited Mood in Speech), is a corpus of induced emotional speech in Italian. DEMoS encompasses 9,365 emotional and 332 neutral samples produced by 68 native speakers (23 females, 45 males) in seven emotional states: the 'big six' anger, sadness, happiness, fear, surprise, disgust, and the secondary emotion guilt. To get more realistic productions, instead of acted speech, DEMoS contains emotional speech elicited by combinations of Mood Induction Procedures (MIP). Three elicitation methods are presented, made up by the combination of at least three MIPs, and considering six different MIPs in total. To select samples 'typical' of each emotion, evaluation strategies based on self- and external assessment were applied. The selected part of the corpus encompasses 1,564 prototypical samples produced by 59 speakers (21 females, 38 male). DEMoS has been published in the Journal Language, Resousrces, and Evalaution.</p> <p>&nbsp;</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Maximilian Schmitt, and Bj&ouml;rn Schuller (2019), <em>DEMoS: An Italian emotional speech corpus. Elicitation methods, machine learning, and perception</em>, Language, Resources, and Evaluation, Feb 2019. <a href="http://em.rdcu.be/wf/click?upn=lMZy1lernSJ7apc5DgYM8eCoqdGxOfRWEudjYRrxU-2BI-3D_Ru5N6PJ4ngeR7K-2Fncs2CW1jGAzl4dMvrVh77-2BVH-2B9g5urNss1KItQNXvWL1jiHKvcYDtUVs2c78DX20PMDTauCGehGiQvHdgrAknGggtu7pHINBqVKjp16-2BTn63kNrm22m52e-2FPV-2FidpRe8A-2FplLxPMV-2FjTR-2FLLIK8Wqe7u0-2BLSZ9w-2BWYtrAXRYn2lvPcjGTP1La8yiTxBuJKbHJpnNeFb6LmBIiNMmGRSZPIY0leXhyj4k07rx5cETF6n34aIQHP-2FwcafanNMN4BoA9QKhXGgFxvRgZQidsQ-2BCDbbTBL0PPjM3CgitSGk66qut9E3pd">https://rdcu.be/bn7oI</a></p> <p>&nbsp;</p> <p><strong>How to access DEMoS</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1v6GaCVyNcib5v802t2uXHYOioqIkoBQ-/view?usp=share_link</p>

restrictedFeb 2019View details →
zenodo12/100

Light and Perception: Study of the Effects of Daylight and Artificial Light on Affect, Mood and Sleepiness under a Sky-Lighting Machine - Data-sets A-B

<p>These data-sets contain information relative to the Paper entitled&nbsp; &ldquo;Light and Perception: Study of the Effects of Daylight and Artificial Light on Affect, Mood and Sleepiness under a Sky-Lighting Machine&rdquo;</p> <p>&nbsp;</p> <p>Data-set A contains Affect, Sleepiness and Mood data divided by participant ID (N = 28) and session number (N=9), summarised by Task (Reading and Game) and Light Condition (Daylight DL; Artificial Light AL). We coded AL as 0 and DL as 1; Reading as 0 and Game as 1.&nbsp;</p> <p>&nbsp;</p> <p>Data-set A includes as factors Gender, Order and Time of Day. Plus these subjective parameters, used as&nbsp; predictors in the models: perceived Color of Light and perceived Level of Light (Brightness).&nbsp;</p> <p>&nbsp;</p> <p>These photometric values were included as predictors in the model: Mean Illuminance, referred to as Room Illuminance; Mean Correlated Color Temperature (CCT); Variation, defined as the Standard Deviation (SD) of the Illuminance values; mEDI means calculated with CIE S 026/E:2018 on the spectral values of the two median Illuminance measurements taken over the time-span of thirty minutes; Illuminance measured by the actigraphs, referred to as Personal Illuminance. Personal Illuminance values were asymmetrically distributed and were logged.&nbsp;</p> <p>&nbsp;</p> <p>Data-set B contains relevant Health Symptoms data divided by participant ID (N = 28) and session number (N=9), summarised by Task (Reading and Game) and Light Condition (Daylight DL; Artificial Light AL).&nbsp;</p> <p>Data-set B includes as factors Gender, Order and Time of Day.&nbsp;</p> <p>&nbsp;</p> <p>These photometric values were included as predictors in the model relative to data-set B: Mean Illuminance, referred to as Room Illuminance in the paper; Variation, defined as the Standard Deviation (SD) of the Illuminance values; mEDI means calculated with CIE S 026/E:2018 on the spectral values of the last two median Illuminance measurements of the session.&nbsp;</p> <p>&nbsp;</p> <p>The file Light-Perception_Exposure-DL-Data-set includes photometric measurements each 30 minutes in Daylight conditions.&nbsp;</p> <p>&nbsp;</p>

restrictedOct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record