Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Semantic annotations”
MOSA: Music mOtion and Semantic Annotation dataset
<p>MOSA dataset is a large-scale music dataset containing 742 professional piano and violin solo music performances with 23 musicians (> 30 hours, and > 570 K notes). This dataset features following types of data:</p> <ul> <li><strong>High-quality 3-D motion capture data</strong></li> <li><strong>Audio recordings</strong></li> <li><strong>Manual semantic annotations</strong></li> </ul> <p>This is the dataset of the paper: Huang et al. (2024) MOSA: Music Motion with Semantic Annotation Dataset for Multimedia Anaysis and Generation. IEEE/ACM Transactions on Audio, Speech and Language Processing. DOI: 10.1109/TASLP.2024.3407529<br>https://arxiv.org/abs/2406.06375</p> <p> </p> <p>The description of dataset is avaiable on Github: https://github.com/yufenhuang/MOSA-Music-mOtion-and-Semantic-Annotation-dataset/blob/main/MOSA-dataset/dataset.md</p> <p> </p> <p>To request the access of full dataset, please sign in with Zenodo and submit the request from.</p>
Virus multi-Semantic Annotation Dataset
<p><span>In-depth research into the characteristics of high-risk oncogenic viruses is of paramount scientific significance for the early prevention and control of related cancers and the development of effective vaccines. The mechanism of viral carcinogenesis involves numerous risk factors, including viral genomic variations, lifestyle, and environmental factors. Based on literature data on 8 oncogenic viruses, we have created a large-scale, semantically rich corpus of viral carcinogenic factors, including 551,715 abstracts and 5,821,308 entities using natural language processing technology combined with expert knowledge. The dataset includes annotation information for eight viruses: HPV, HIV, EBV, MCV, HCV, HTLV-1, HBV, and KSHV, covering 38 external factors and 5 internal factors.</span></p>
Figure 1 from: Deans A, Mullins P, Kawada R, Balhoff J (2013) Corrigenda: Mullins PL, Kawada R, Balhoff JP, Deans AR (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1–38, doi: 10.3897/zookeys.223.3572. ZooKeys 278: 105-113. https://doi.org/10.3897/zookeys.278.5108
Figure 1 - Evaniscus sulcigenis Roman, 1917. Lateral habitus (whole body) of holotype.
Figure 2 from: Deans A, Mullins P, Kawada R, Balhoff J (2013) Corrigenda: Mullins PL, Kawada R, Balhoff JP, Deans AR (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1–38, doi: 10.3897/zookeys.223.3572. ZooKeys 278: 105-113. https://doi.org/10.3897/zookeys.278.5108
Figure 2 - Evaniscus sulcigenis Roman, 1917. Lateral habitus (mesosma and head) of holotype.
TweetsCOV19 - A Semantically Annotated Corpus of Tweets About the COVID-19 Pandemic (Part 4, January 2021 - August 2022)
<p><strong><a href="https://data.gesis.org/tweetscov19/">TweetsCOV19</a></strong><strong> </strong>is a semantically annotated corpus of Tweets about the COVID-19 pandemic. It is a subset of <a href="https://data.gesis.org/tweetskb">TweetsKB</a> and aims at capturing online discourse about various aspects of the pandemic and its societal impact. <strong>Metadata</strong> information about the tweets as well as extracted <strong>entities</strong>, <strong>sentiments</strong>, <strong>hashtags</strong>, <strong>user mentions</strong>, and <strong>resolved URLs </strong>are exposed in RDF using established RDF/S vocabularies*.</p> <p>We also provide a <em><strong>tab-separated values (tsv)</strong></em> version of the dataset. Each line contains features of a tweet instance. Features are separated by tab character ("\t"). The following list indicate the feature indices:</p> <ol> <li>Tweet Id: Long.</li> <li>Username: String. Encrypted for privacy issues*.</li> <li>Timestamp: Format ( "EEE MMM dd HH:mm:ss Z yyyy" ).</li> <li>#Followers: Integer.</li> <li>#Friends: Integer.</li> <li>#Retweets: Integer.</li> <li>#Favorites: Integer.</li> <li>Entities: String. For each entity, we aggregated the original text, the annotated entity and the produced score from <a href="https://github.com/yahoo/FEL">FEL</a> library. Each entity is separated from another entity by char ";". Also, each entity is separated by char ":" in order to store "original_text:annotated_entity:score;". If FEL did not find any entities, we have stored "null;".</li> <li>Sentiment: String. <a href="http://sentistrength.wlv.ac.uk/">SentiStrength</a> produces a score for positive (1 to 5) and negative (-1 to -5) sentiment. We splitted these two numbers by whitespace char " ". Positive sentiment was stored first and then negative sentiment (i.e. "2 -1").</li> <li>Mentions: String. If the tweet contains mentions, we remove the char "@" and concatenate the mentions with whitespace char " ". If no mentions appear, we have stored "null;".</li> <li>Hashtags: String. If the tweet contains hashtags, we remove the char "#" and concatenate the hashtags with whitespace char " ". If no hashtags appear, we have stored "null;".</li> <li>URLs: String: If the tweet contains URLs, we concatenate the URLs using ":-: ". If no URLs appear, we have stored "null;"</li> </ol> <p>To extract the dataset from <a href="https://data.gesis.org/tweetskb">TweetsKB</a>, we compiled a seed list of 268 COVID-19-related <a href="https://data.gesis.org/tweetscov19/keywords_v1.1.txt">keywords</a>.</p> <p><em>* For the sake of privacy, we anonymize user IDs and we do not provide the text of the tweets.</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.