Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
243
datasets available to search
ShareScore release 0.9.0
Dataset results
243 results for “categories”
Evaluation of a simple score-based Natural Language Processing (NLP) algorithm: Category Confusion Matrix
<p>Resulting category confusion matrix for the experiment "Evaluation of a simple score-based Natural Language Processing (NLP) algorithm".</p>
Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes
<p>The dataset was made for the purposes of the author's master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">"Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo’s Core Vocabulary"</a>. The dataset includes two worksheets. The first is titled "Main table", and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled "Extended table", includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively). <br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&lang=slv">at this link</a>. </p>
Data for: Bulk microphysics schemes may perform better with a unified cloud-rain category
Open the record for dataset details and reuse information.
Dataset, statistical analysis code, and supplementary material of juvenile ravens' responses towards acoustic cues of different social categories
Open the record for dataset details and reuse information.
Rule based and information integration category learning
Open the record for dataset details and reuse information.
Annotated genes harboring major effect markers (R2 ≥ 15%). Highlighted in green are genes annotated from Rhodes et al. 2014,2017, in orange genes annotated as similar to Peroxidase, in yellow new annotations from sorghum genome in Atlas. In the first three columns start and stop position on the sorghum genome and transcript name, followed by the nearest marker name and the distance of the gene from the nearest marker, then a column where are shown the GWAS methods and target traits for which the linked SNP was significant, the last column shows the category of the genes.
<p><strong>We conducted a comprehensive genomics study to map genomic loci determining the production of antioxidants in sorghum grains. Encouraging results were obtained and published in peer-reviewed article with impact factor (https://doi.org/10.1371/journal.pone.0225979). Annotated genes harboring major effect markers (R<sup>2</sup> ≥ 15%) were identified and will be of worldwide interest. </strong></p>
Activity classes from different categories
<p>A total of 102 activity classes (ACs) were assembled from ChEMBL version 20 and were classified as “easy” (i.e. yielded generally high compound recall using different fingerprints in benchmarking calculations), “preferred/intermediate” (moderate compound recall), and “difficult” (low compound recall). In addition, 10000 randomly selected ZINC compounds were provided. For each compound, both MACCS and ECFP4 fingerprints were given. Furthermore, the general information for these ACs was provided in the excel file.</p> <p> </p> <p> </p>
CMIP5 Quality Control Level 2 Exception Codes with Categories
<p>This document contains a list of exception codes for the interpretation of the quality control level 2 (QC L2) results of CMIP5 including exception categories. It was originally available on a web page under the url: http://cera-www.dkrz.de/CMIP5/QC/2/qc2list.html. The QC L2 results are available at: http://cera-www.dkrz.de/WDCC/CMIP5/QCResult.jsp .</p>
Prioritization of unknown features based on predicted toxicity categories
<p> The fish toxicity data set, the filtered CompTox data set, and the pesticide mixture datasets used for the development of the prioritization algorithms described elsewhere.</p>
Dataset for Interactive Profiling Narrative (IPN) with Toxicity Tolerance Score of each individual user to different categories of toxicity
<p>he dataset is the collection of the results of the Interactive Profiling Narrative that was developed as a part of the project Listener Aware Content Detoxification.<br>It contains the toxicity tolerance scores of each individual user to different categories.<br>High scores indicate that the user is extremely sensitive to the particular category where as low scores indicate that the user isn't that triggered by the category.Medium scores show moderate tolerance.</p> <p>Columns:<br>1.UserId: The unique id given to each user (helps preserve anonymity).<br>2.Race: The user's toxicity tolerance to the category race. High score means racist comment trigger them. <br>3.Sex: The user's toxicity tolerance to the category sexuality. High scores indicate sensitivity to comments on sexual identity.<br>4.Body_image: The user's sensitivity to comments regarding body image. High scores indicate that remarks on body appearance strongly affect the user.<br>5.Disability:The user's sensitivity to comments about disabilities. A high score means the user is highly sensitive to potentially ableist remarks.<br>6.Religion_culture: The user's sensitivity to content involving religion or culture. High scores show the user is easily triggered by insensitive comments about religious or cultural aspects.<br>7.Physical_abuse: User's sensitivity to comments regarding physical abuse. High scores indicate the user is strongly affected by remarks on physical abuse.<br>8.Mental_health: This measures the user's sensitivity to comments on mental health. High scores suggest that the user is particularly affected by comments stigmatizing mental health issues.<br>9.Politics:The user's sensitivity to political content. High scores indicate that political discussions or comments are likely to evoke a strong reaction in the user</p>
Toxic Sentence Classification Dataset with labels of categories such as religion, mental health, race, sex, body image, disability, physical abuse, and politics
<p>The dataset has a collection of various toxic sentences belonging to different categories. It was collected from various sources. It indicates which category each sentence belongs to. The values of the category columns are binary 1 or 0 indicating whether the sentence belongs to that particular category or not. Each sentence belongs to only 1 category. </p> <p> </p> <p>Columns:<br>1.comment_text: Contains toxic sentences that are insensitive and offensive, focusing on various categories.<br>2.mental_health: Binary value 1 indicates that the sentence focuses on mental health.<br>3.Race:Binary value 1 indicates that the sentence is racist.<br>4.sex:Binary value 1 indicates that the sentence focuses on sexuality.<br>5.body_image:Binary value 1 indicates that the sentence focuses on body image.<br>6.disability:Binary value 1 indicates that the sentence focuses on physical disability and related issues.<br>7.religion:Binary value 1 indicates that the sentence can be triggering to people who are extremely religious.<br>8.physical_abuse:Binary value 1 indicates that the sentence focuses on physical abuse issues.<br>9.politics:Binary value 1 indicates that the sentence focuses on political issues.</p>
Fig. 1 in Mercury bioaccumulation in fish of commercial importance from different trophic categories in an Amazon floodplain lake
Fig. 1. Map of lago Grande de Mancapuru.
Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Matching articles with categories)
<p>This document includes which primary study falls into which category with respect to the RQs in the following study: “Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review”</p>
Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Matching articles with categories)
<p>This document includes which primary study falls into which category with respect to the RQs in the following study: “Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review”</p>
Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review (Matching articles with categories)
<p>This document includes which primary study falls into which category with respect to the RQs in the following study: “Quality Evaluation Models or Frameworks for Open Source Software: A Systematic Literature Review”</p>
Prevalence of psychopathological symptoms and their determinants in 4 healthcare workers categories during the second year of COVID-19 pandemic
<p>Dataset including data presented in the paper: <strong>Prevalence of psychopathological symptoms and their determinants in 4 healthcare workers categories during the second year of COVID-19 pandemic</strong></p>
The Food and Food Categories (FFoCat) Dataset
<p><strong>The Food and Food Categories (FFoCat) Dataset</strong></p> <p>The Food and Food Categories (FFoCat) Dataset contains 58.962 images of food annotated with the food label and the food categories of the Mediterranean Diet. It is one of the most complete datasets regarding the Mediterranean Diet as it is aligned with the standard AGROVOC and HeLiS ontologies and allows to study multitask learning problems in Computer Vision for food recognition and diet recommendation.</p> <p>The dataset is already divided into the <code>train</code> and <code>test</code> folder. The file <code>label.tsv</code> contains the food labels, the file <code>food_food_category_map.tsv</code> contains the food labels with the corresponding food category labels. The following table compares the FFoCat dataset with previous datasets for food recognition.</p> <p>This dataset has been published at the International Conference on Image Analysis and Processing (ICIAP - 2019). The source code for reproducing the experiments together with other information about the dataset is available <a href="https://github.com/ivanDonadello/Food-Categories-Classification">here</a>.</p> <p><strong>AGROVOC Alignment of Food Categories</strong></p> <p>The <code>AGROVOC_alignment.tsv</code> file contains the alignment of the food categories in the FFoCat dataset with AGROVOC, the standard ontology of the Food and Agriculture Organization (FAO) of the United Nations. This allows interoperability and linked open data navigation. Such alignment can be derived by querying <a href="https://horus-ai.fbk.eu/helis/">HeLis</a>, here we propose a shortcut.</p> <p><strong>Citing FFoCat</strong></p> <p>If you use FFoCat in your research, please use the following BibTeX entry.</p> <pre><code>@inproceedings{DonadelloD19Ontology, author = {Ivan Donadello and Mauro Dragoni}, title = {Ontology-Driven Food Category Classification in Images}, booktitle = {{ICIAP} {(2)}}, series = {Lecture Notes in Computer Science}, volume = {11752}, pages = {607--617}, publisher = {Springer}, year = {2019} } </code></pre>
Tree dynamic response and survival in a category-5 tropical cyclone: The case of super typhoon Trami
In the future with climate change, we expect more forest and tree damage due to the increasing strength and changing trajectories of tropical cyclones (TCs). However, to date, we have limited information to estimate likely damage levels, and nobody has ever measured exactly how forest trees behave mechanically during a TC. In 2018, a category-5 TC destroyed trees in our ongoing research plots, in which we were measuring tree movement and wind speed in two different tree spacing plots. We found damaged trees in only the wider spaced plot. Here, we present how trees dynamically respond to strong winds during a TC. Sustained strong winds obviously trigger the damage to trees and forests but inter-tree spacing is also a key factor because the level of support from neighboring trees modifies the effective "stiffness" against the wind both at the single tree and whole forest stand level.
wiki-category-consistency-cache
<p>A collection of SQLite database files containing all the data retrieved from the Wikidata JSON dump of 2022-05-02 and the Wikipedia SQL dumps of 2022-05-01 in the context of analyzing the consistency between Wikipedia and Wikidata categories.<br> <br> Detailed information can be found on the <a href="https://github.com/fusion-jena/wiki-category-consistency">Github page</a>.</p>
Replication package for the article 'The residential patterns of Swiss urban elites. Continuity and change across elite categories (1890-2000),' to appear in the journal 'European Societies,' authored by Pierre Benz; Michael A. Strebel; Roberto Di Capua & André Mach.
<p>The documents in this replication package serve to reproduce the analysis for the article <br>'The residential patterns of Swiss urban elites. Continuity and change across elite categories (1890-2000),' to appear in<br>the journal 'European Societies,' authored by Pierre Benz; Michael A. Strebel; Roberto Di Capua & André Mach.<br>This replication package is created by Pierre Benz (pierre.benz@unil.ch).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.