Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,970

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,970 results for “concepts”

Learn how ShareScore rates datasets ↗
OpenNeuro48/100

Model-based fMRI reveals co-existing specific and generalized concept representations

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

Alignment between type of landmark in different sources and the concept in the spatial reference objects ontology

<p>The five datasets represent a manually alignment between the landmark type of five different datasets archived <a href="https://doi.org/10.5281/zenodo.6480986">here</a> and a common vocabulary extracted from an application ontology defined for mountain rescue purposes, named&nbsp;<a href="https://hamac.ign.fr/owa/redir.aspx?C=cjlWje9SCaYsVOTLbxbOoIBLZUCS56nVb248cRSMTEDSENDFzybaCA..&amp;URL=http%3a%2f%2fchoucas.ign.fr%2fdoc%2fontologies%2foor.owl%2f">Ontology of landmarks</a>&nbsp;(OOR).</p> <p>Each file represents the alignment for features belonging to a data source with the same OOR ontology.</p> <p>For example, the type &laquo;bivouac&raquo; from camptocamp.org source is aligned with the uri <a href="http://purl.org/choucas.ign.fr/oor#abri">http://purl.org/choucas.ign.fr/oor#abri</a> of the corresponding class &laquo;Shelter&nbsp;&raquo; in the ontology of landmark. The alignments models can be considered as a ground truth data.</p> <p>The alignments results are obtained using an ontology application named <a href="http://choucas.ign.fr/doc/ontologies/index-fr.html">OOR</a>. These specific results are obtained using the version of OOR V1.0.1 which is an improved version and contains new concepts compared to the first release 1.0.0. The new version of OOR (i.e. 1.0.1) will be released by the end of May 31 2022. The new link will be added here.</p> <p>This archive is released for transparency and reproducibility purposes.</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

MUHAI Benchmark : Task 3 (Understanding Complex Concepts)

<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 3 Understanding complex concepts</strong></p> <p>&nbsp;</p> <p>This dataset helps investigating whether&nbsp;symbolic reasoning can help statistical models tro understand complex concepts. Complex concepts are expressed in the form of Image Schemas (i.e.,&nbsp;mental templates that&nbsp;summarise human&nbsp;experiences in the form of patterns of object relations and actions).<br> The submission includes the the ImageSchema dataset with ground truth labels :<br> 1.&nbsp;A question to be asked<br> 2. the type of Image schema (class)<br> 3.&nbsp;the type of phrasing (literal , metaphoric, a distracting sentence)<br> 4. the type of questioning (one referring to the image schema by name, and another describing its content)&nbsp;<br> 5. a question indicating whether it is a yes or no answer<br> 6. the image schema&nbsp;variables identified<br> <br> Each sample in the datasets starts with a question about the presence of the given schema in the following sentence, and follows with a single sentence to be classified as either &quot;yes&quot;&nbsp;or &quot;no&quot;&nbsp;(presence or absence of a schema).<br> <br> This&nbsp;can be used by a system (eg a language model, a symbolic system, a neuro-symbolic approach) to&nbsp;identify image schemas. The&nbsp;file &quot;language-models.csv&quot; includes&nbsp;the results of&nbsp;two language models that were tested (T0pp, GPT-3).</p> <p>Metrics used to evaluate:<br> 1.&nbsp;Accuracy&nbsp;: no. correct&nbsp;predictions&nbsp; / no. of total sentences&nbsp;(TP + TN / P + N)<br> 2. Precision: no. correct&nbsp;image schema predictions&nbsp; /&nbsp;total correct image schema&nbsp;predictions (TP / TP + FP)<br> 3. Recall :&nbsp; &nbsp;: no. correct&nbsp;image schema predictions&nbsp; /&nbsp;total predicted image schema (TP / TP + FN)<br> 4. F1 : harmonic mean of Precision and Recall</p> <p>Code :&nbsp;https://github.com/kmitd/muhai-EPL&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Dissecting the FAIR Guiding Principles - Key Categories, Core Concepts, Focus Elements, and Harmonized Indicators

<p>A comprehensive workbook created to facilitate and document&nbsp;the process of decomposing the FAIR Guiding Principles and mapping them to key categories, requirements, core concepts, focus elements, and harmonized&nbsp;indicators. It also contains&nbsp;a complete list of the indicators.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Dataset from the Survey on Industry 5.0 Concepts and Enabling Technologies, Towards an Enhanced Conservation Practice

<p>This database contains all the responses from the participants in the survey: Industry 5.0 Concepts and Enabling Technologies, Towards an Enhanced Conservation Practice.</p> <p>The main purpose of this survey was to explore how the Architecture, Engineering, Construction, Management, Operation, and Conservation (AECMO&amp;C)<br>industry can adapt and better prepare to embrace the innovative principles and enabling technologies of Industry 5.0. This could ultimately result in<br>enhanced conservation practices for built cultural heritage.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset and Benchmarks

<p><strong>Abstract:</strong> Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited. Supply chains are inherently graph-like, making them ideal for GNN methodologies, which can optimize and solve complex problems. The barriers include a lack of proper conceptual foundations, familiarity with graph applications in SCM, and real-world benchmark datasets for GNN-based supply chain research. To address this, we discuss and connect supply chains with graph structures for effective GNN application, providing detailed formulations, examples, mathematical definitions, and task guidelines. Additionally, we present a multi-perspective real-world benchmark dataset from a leading FMCG company in Bangladesh, focusing on supply chain planning. We discuss various supply chain tasks using GNNs and benchmark several state-of-the-art models on homogeneous and heterogeneous graphs across six supply chain analytics tasks. Our analysis shows that GNN-based models consistently outperform statistical ML and other deep learning models by around 10-30% in regression, 10-30% in classification and detection tasks, and 15-40% in anomaly detection tasks on designated metrics. With this work, we lay the groundwork for solving supply chain problems using GNNs, supported by conceptual discussions, methodological insights, and a comprehensive dataset.</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Virtual Reality Dataset used for Proof of Concept in the Validation of the Conflict Detection and Resolution Use Case (ARTIMATION)

<p>This dataset contains the <strong>dataset </strong>used in the Virtual Reality POC for the validation of the Conflict Detection and Resolution (CD&amp;R) use case.</p> <p>This dataset represent a extract of different (using K-means) candidate solution, either good or bad ones.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Example of Force Digital Calibration Certificate used in ComTraForce 18SIB08 project to demonstrate Digital Twin concept

<p>Force Digital Calibration Certificate (DCC) was developed in the frameworks of 18SIB08&nbsp;ComTraForce project. It was used to demonstrate the way of data connection between the physical object (force transducer) and virtual object (Finite Element model) within&nbsp;the developed Digital Twin&nbsp;concept. The developed at PTB v3.1.2 xsd schema was used to convert analog calibration certificate to machine readable XML&nbsp;format. The DCC covers static and continuous calibration processes. Note that the current Force DCC is not a Good Practice example. Please follow further developments of force DCC Good Practice example at&nbsp;https://gitlab.com/ptb/dcc.</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Data of "Towards a More Reliable Forecast of Ice Supersaturation: Concept of a One-Moment Ice Cloud Scheme that Avoids Saturation Adjustment"

<p>These are the data used for generating the figures in the ACP article &quot;Towards a More Reliable Forecast of Ice Supersaturation: Concept of a One-Moment Ice Cloud Scheme that Avoids Saturation Adjustment&quot; by Sperber and Gierens.</p> <p>The data sets labeled&nbsp;&quot;Box&quot; have been generated by the stochastic box model, &quot;adj&quot; refers to the parameterisation using saturation adjustment and data labeled&nbsp;&quot;par&quot; originate from&nbsp;the newly developed parameterisation.</p> <p>The label &quot;const&quot; followed by a number refers to simulations with a constant updraught of the speed specified by the number in cm/s. The label &quot;cos&quot; refers to the simulations in which&nbsp;the updraught velocity follows a cosine function in time.</p> <p>&quot;a10&quot; labels simulations with less&nbsp;initial clear sky humidity fluctuations of plus/minus 10% instead of plus/minus 25%. &quot;al0028&quot; labels simulations with a higher deposition rate of 0.0028 1/s instead of 0.0003 1/s. &quot;step10&quot; labels simulations with a longer time step of 10 minutes instead of 1 minute.</p> <p>&quot;Box_const2_rh1.txt&quot; contains data from a simulation similar to &quot;Box_const2.txt&quot; but with an initial mean relative humidity of 100% instead of 110%. &quot;Box_het.txt&quot; contains data from a simulation including heterogeneous nucleation. &quot;Box_slow_nuc.txt&quot; contains data from a simulation where the deposition rate increases over time from zero after&nbsp;nucleation in every air parcel. &quot;Box_upvar.txt&quot; contains data from a simulation, where the updraught velocity in every air parcel varies randomly between 1 cm/s and 3 cm/s and the deposition rate inside the air parcel depends on the updraught velocity at the time of nucleation.</p> <p>&nbsp;</p> <p>The columns in the &quot;Box&quot; files represent from left to right:</p> <p>1. Time since the simulation start in s</p> <p>2. Cloud fraction</p> <p>3. Mean relative humidity across all air parcels</p> <p>4. Mean specific humidity across all air parcels</p> <p>5. Mean specific ice content across all air parcels</p> <p>6. Mean relative humidity across all cloudy air parcels</p> <p>7. Mean relative humidity across all clear air parcels</p> <p>8. Mean equilibrium supersaturation</p> <p>9. Mean threshold relative humidity for homogeneous nucleation</p> <p>10. Mean deposition rate across all cloudy air parcels</p> <p>11. Mean updraught velocity</p> <p>&nbsp;</p> <p>The columns in the &quot;adj&quot; files represent from left to right:</p> <p>1. Time since the simulation start in s</p> <p>2. Cloud fraction</p> <p>3. Mean relative humidity</p> <p>4. Mean specific humidity</p> <p>5. Mean specific ice content</p> <p>6. In-cloud Humidity</p> <p>7. Clear sky humidity</p> <p>&nbsp;</p> <p>The columns in the &quot;par&quot; files represent from left to right:</p> <p>1. Time since the simulation start in s</p> <p>2. Cloud fraction</p> <p>3. Mean relative humidity</p> <p>4. Mean specific humidity</p> <p>5. Mean specific ice content</p> <p>6. In-cloud Humidity</p> <p>7. Clear sky humidity</p> <p>8. Obsolete</p> <p>9. Equilibrium supersaturation</p>

opencc-by-4.0Oct 2023View details →
edi48/100

Does the River Continuum Concept apply on a tropical island? Longitudinal variation in a Puerto Rican stream

We examined whether a tropical stream in Puerto Rico matched predictions of the River Continuum Concept (RCC) for macroinvertebrate functional feeding groups (FFGs). Sampling sites for macroinvertebrates, basal resources, and fishes ranged from headwaters to within 2.5 km of the fourth-order estuary. In a comparison to a model temperate system where RCC predictions generally held, we used catchment area as a measure of stream size in order to examine truncated RCC predictions (i.e., cut off to correspond to the largest stream size sampled in Puerto Rico). Despite dominance of generalist freshwater shrimps, which use more than one feeding mode, RCC predictions held for scrapers, shredders, and predators. Collector-filterers showed a trend opposite that predicted by the RCC, but patterns in basal resources suggest that this is consistent with the central RCC theme: longitudinal distributions of FFGs follow longitudinal patterns in basal resources. Alternatively, the filterer pattern may be explained by fish predation affecting distributions of filter-feeding shrimp. Our results indicate that the RCC generally applies to running waters on tropical islands. However, additional theoretical and field studies across a broad array of stream types should examine whether the RCC needs to be refined to reflect the potential influence of top-down trophic controls on FFG distributions. Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.

openCC0Nov 2023View details →
zenodo44/100

Supplementary material for "Song et al., Modelling Simul. Mater. Sci. Eng., 2021: Data-mining of dislocation microstructures: concepts for coarse-graining of internal energies"

<p>This zip archive contains supplementary material in the form of datasets and jupyter notebooks that are used in the following publication:</p> <ul> <li>authors: Hengxu Song, Nina Gunkelmann, Giacomo Po, and Stefan Sandfeld</li> <li>journal: Modelling Simul. Mater. Sci. Eng.</li> <li>year: 2021</li> <li>title: Data-mining of dislocation microstructures: concepts for coarse-graining of internal energies</li> </ul>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Node2Vec model - Czech Wikidata (knowledge graph / concepts / l80 / rw40)

<p>Node2Vec&nbsp;embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 80</li> <li>number of random walks = 40</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Node2Vec model - Czech Wikidata (knowledge graph / concepts / l40 / rw10)

<p>Node2Vec&nbsp;embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 40</li> <li>number of random walks = 10</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Le concept d'article de données

<p>Réalisée pour l'<a href="https://scienceouverte.univ-lorraine.fr/donnees-de-la-recherche-ul/donnees-de-la-recherche/">atelier de la donnée ADOC Lorraine</a>, cette illustration synthétise le concept d'article de données (angl. <i>data paper</i>). Créant des liens entre les articles de recherche et les données déposées dans un entrepôt adapté, l'article de données permet aux chercheurs et aux chercheuses d'améliorer la qualité de leur recherche, en accord avec les principes de la science ouverte et du partage de données.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Embodied Spatial Navigation Training in Mild Cognitive Impairment: A Proof-of-Concept Trial

<p>Raw data of included cognitive test and VR data Starting Grant Ricerca Finalizzata, code: SG-2018-12368175</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Bibliographic Dataset for the Systematic Literature Review on Industry 5.0 Concepts and Enabling Technologies, Towards an Enhanced Conservation Practice

<p>This database contains all the bibliographic information found after applying the Search Strategy used for the Industry 5.0 Concepts and Enabling Technologies, Towards an Enhanced Conservation Practice: Systematic Literature Review. The following electronic databases were searched:</p> <ul> <li>Scopus.</li> </ul> <p>A total of 907 records were found. The search was conducted on 16/02/2024.</p> <p>The information is presented in .ris, .bib, and .csv format.</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

2023 HYDRAS Proof of concept experiment with soybean genotypes

<h2>Description</h2> <p>This data sets contains metadata and data of the <strong>2023_POC experiment in the HYDRAS facility</strong>.</p> <p>&nbsp;The 2023 proof-of-concept study in hydras tested all standard field-phenotyping measurement types available in the infrastructure with 3 contrasting varieties of soybean + a control and a drought treatment using rain-out shelters. Start date: 23/5/2023, End dat: 4/10/2023. Location: Melle, Belgium.&nbsp;</p> <p>Data set contains: UAV sensor data, Electrical Resistivity Tomography data, soil point sensor data (water content, water potential and temperature), weather data, yield data and associated experimental information.&nbsp;</p> <h2>Content</h2> <ul> <li>2023_POC_metadata.xlsx contains all information about the experiment (goal, design, sensors, treatments, biological material, ...) and about the associated data files.</li> <li>Subfolder SPATIAL_INFO contains all spatial information about the experimental layout (location of field, plots, sensors, transects, ...) in GEOJSON files</li> <li>Other subfolders contain the actual data from various sources as .csv files.&nbsp;</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Supplement to "Proof of concept for Bayesian inference of dynamic rating curve uncertainty" (v3)

<div>This deposit contains part of the updated supplement to &ldquo;Proof of concept for Bayesian inference of dynamic rating curve uncertainty&rdquo; (<a href="https://www.tandfonline.com/doi/full/10.1080/02626667.2024.2401094" target="_blank" rel="noopener">Cornelio et al. 2024, HSJ</a>). This version, in particular, contains two files in which the following changes were made from the earlier version (v2.0.1):</div> <div> <ul> <li><strong><em>250117_Lbn_RC_new.R</em></strong>&nbsp;is the updated R code. The argument for the random number generator (RNG) kind is defined for the set.seed() functions used in the script.&nbsp;</li> <li><strong><em>Lbn-DMs-csv0.csv</em></strong> is the updated input file containing the stage-discharge gaugings. The column for the stage values has been renamed to "H_rec" (instead of "H_m" as in the original CSV) to be consistent with the attribute name used throughout the R code.</li> </ul> <p>Except for the above files, all the input and output files in <a href="https://zenodo.org/records/12792513" target="_blank" rel="noopener">v2.0.1</a><span>&nbsp;remain unchanged.&nbsp;</span></p> </div> <p><u>&nbsp;</u></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Database of Cross-Linguistic Norms, Ratings, and Relations for Words and Concepts as CLDF dataset

<p>Cite the source of the dataset as:</p> <blockquote> <p>Tjuka, Annika, Robert Forkel, and Johann-Mattis List. 2022. Linking Norms, Ratings, and Relations of Words and Concepts Across Multiple Language Varieties. Behavior Research Methods 54. 864–884. DOI: 10.3758/s13428-021-01650-1</p> </blockquote>

opencc-by-4.0Nov 2022View details →
zenodo44/100

WikiMed and PubMedDS: Two large-scale datasets for medical concept extraction and normalization research

<p>Two large-scale, automatically-created datasets of medical concept mentions, linked to the <a href="https://uts.nlm.nih.gov/uts/umls/home">Unified Medical Language System (UMLS)</a>.</p> <p><strong>WikiMed</strong></p> <p>Derived from Wikipedia data. Mappings of Wikipedia page identifiers to UMLS Concept Unique Identifiers (CUIs) was extracted by crosswalking Wikipedia, Wikidata, Freebase, and the NCBI Taxonomy to reach existing mappings to UMLS CUIs. This created a 1:1 mapping of approximately 60,500 Wikipedia pages to UMLS CUIs. Links to these pages were then extracted as mentions of the corresponding UMLS CUIs.</p> <p>WikiMed contains:</p> <ul> <li>393,618 Wikipedia page texts</li> <li>1,067,083 mentions of medical concepts</li> <li>57,739 unique UMLS CUIs</li> </ul> <p>Manual evaluation of 100 random samples of WikiMed found 91% accuracy in the automatic annotations at the level of UMLS CUIs, and 95% accuracy in terms of semantic type.</p> <p><strong>PubMedDS</strong></p> <p>Derived from biomedical literature abstracts from <a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>. Mentions were automatically identified using distant supervision based on Medical Subject Heading (MeSH) headers assigned to the papers in PubMed, and recognition of medical concept mentions using the high-performance <a href="https://allenai.github.io/scispacy/">scispaCy</a> model. MeSH header codes are included as well as their mappings to UMLS CUIs.</p> <p>PubMedDS contains:</p> <ul> <li>13,197,430 abstract texts</li> <li>57,943,354 medical concept mentions</li> <li>44,881 unique UMLS CUIs</li> </ul> <p>Comparison with existing manually-annotated datasets (NCBI Disease Corpus, BioCDR, and MedMentions) found 75-90% precision in automatic annotations. Please note this dataset is&nbsp;<em>not&nbsp;</em>a comprehensive annotation of medical concept mentions in these abstracts (only mentions located through distant supervision from MeSH headers were included), but is intended as data for <em>concept n</em><em>ormalization</em>&nbsp;research.</p> <p>Due to its size, PubMedDS is distributed as 30 individual files of approximately 1.5 million mentions each.</p> <p><strong>Data format</strong></p> <p>Both datasets use JSON format with one document per line. Each document has the following structure:</p> <pre><code class="language-json">{ "_id": "A unique identifier of each document", "text": "Contains text over which mentions are ", "title": "Title of Wikipedia/PubMed Article", "split": "[Not in PubMedDS] Dataset split: &lt;train/test/valid&gt;", "mentions": [ { "mention": "Surface form of the mention", "start_offset": "Character offset indicating start of the mention", "end_offset": "Character offset indicating end of the mention", "link_id": "UMLS CUI. In case of multiple CUIs, they are concatenated using '|', i.e., CUI1|CUI2|..." }, {} ] }</code></pre> <p><strong>Version history</strong></p> <table align="left"> <thead> <tr> <th scope="col">Version</th> <th scope="col">Notes</th> </tr> </thead> <tbody> <tr> <td>1.0.0</td> <td>Initial release</td> </tr> <tr> <td>1.0.1</td> <td>Corrected duplication error in WikiMed.zip file</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record