Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

483 results for “semantics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Biolinks, datasets and algorithms supporting semantic-based distribution and similarity for scientific publications

<p><strong>Background: </strong>Finding articles related to a publication of interest remains a challenge in the Life Sciences domain as the number of scientific publications grows day by day. Publication repositories such as PubMed and Elsevier provides a list of similar articles. There, similarity is commonly calculated based on title, abstract and some keywords assigned to articles. Here we present the datasets and algorithms used in Biolinks. Biolinks uses ontological concepts extracted from publication and makes it possible to calculate a distribution score according to semantic groups as well as a semantic similarity based on either all identified annotations or narrowed to one or more particular semantic groups. Biolinks supports both title and abstract only as well as full-text.</p> <p><strong>Materials: </strong>In a previous work [1], 4,240 articles from the TREC-05 collection [2] were selected. The title-and-abstract for those 4,240 articles were annotated with Unified Medical Language System (UMLS) concepts, such annotations are refer to as our TA-dataset and correspond to the JSON files under the pubmed folder in the JSON-LD.zip file. From those 4,240 articles, full-text was available for only 62. The title-and-abstract annotations for those 62 articles, TAFT-dataset, are located under the pubmed-pmc folder in the JSON-LD.zip file, which also contains the full-text annotations under the folder pmc, FT-dataset. The list corresponding to articles with title-and-abstract is found in the genomics.qrels.large.pubmed.onlyRelevants.titleAndAbstract.tsv file, while those with full-text are recorded in the genomics.qrels.large.pmc.onlyRelevants.fullContent.tsv file.</p> <p>Here we include the annotations on title and abstract as well as those for full-text for all our datasets (profiles.zip). We also provide the global similarity matrices (similarity.zip).</p> <p><strong>Methods:</strong> The TA-dataset was used to calculate the Information Gain (IG) according to the UMLS semantic groups, see IG_umls_groups.PMID.xlsx. A new grouping is proposed for Biolinks, see biolinks_groups.tsv. The IG was calculated for Biolinks groups as well, IG_biolinks_groups.PMID.xlsx, showing a improvement around 5%.</p> <p>In order to assess the similarity metric regarding the cohesion of TREC-05 groups, we used Silhouette Coefficient analyses. An additional dataset Stem-TAFT-dataset was used and compared to TAFT and FT datasets.</p> <p>Biolinks groups were used to calculate a semantic group distribution score for each article in all our datasets. A semantic similarity metric based on PubMed related articles [3] is also provided; the Biolinks groups can be used to narrow the similarity to one or more selected groups. All the corresponding algorithms are open-access and available on GitHub under the license Apache-2.0, a frozen version, biotea-io-parser-master.zip, is provided here. In order to facilitate the analysis of our datasets based on the annotations as well as the distribution and similarity scores, some web-based visualization components were created. All of them open-access and available in GitHub under the license Apache-2.0; frozen versions are provided here, see files biotea-vis-annotation-master.zip, biotea-vis-similarity-master.zip, biotea-vis-tooltip-master.zip and biotea-vis-topicDistribution-master.zip. These components are brought together by biotea-vis-biolinks-master.zip. A demo is provided at http://ljgarcia.github.io/biotea-biolinks/; this demo was built on top of GitHub pages, a frozen version of the gh-pages branch is provided here, see biotea-biolinks-gh-pages.zip.</p> <p><strong>Conclusions: </strong>Biolinks assigns a weight to each semantic group based on the annotations extracted from either title-and-abstract or full-text articles. It also measures similarity for a pair of documents using the semantic information. The distribution and similarity metrics can be narrowed to a subset of the semantic groups, enabling researchers to focus on what is more relevant to them.</p> <p> </p> <p>[1] Garcia Castro, L.J., R. Berlanga, and A. Garcia, <em>In the pursuit of a semantic similarity metric based on UMLS annotations for articles in PubMed Central Open Access.</em> Journal of Biomedical Informatics, 2015. <strong>57</strong>: p. 204-218</p> <p>[2] Text Retrieval Conference 2005 - Genomics Track. <em>TREC-05 Genomics Track ad hoc relevance judgement</em>. 2005  [cited 2016 23rd August]; Available from: http://trec.nist.gov/data/genomics/05/genomics.qrels.large.txt</p> <p>[3] Lin, J. and W.J. Wilbur, <em>PubMed related articles: a probabilistic topic-based model for content similarity.</em> BMC Bioinformatics, 2007. <strong>8</strong>(1): p. 423</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

WikiMuTe: A web-sourced dataset of semantic descriptions for music audio

<p>This upload contains the supplementary material for our <a href="https://arxiv.org/abs/2312.09207" target="_blank" rel="noopener">paper</a> presented at the <a href="https://mmm2024.org/" target="_blank" rel="noopener">MMM2024 conference</a>.</p> <h2>Dataset</h2> <p>The dataset contains rich text descriptions for music audio files collected from Wikipedia articles.</p> <p>The audio files are freely accessible and available for download through the URLs provided in the dataset.</p> <h3>Example</h3> <p>A few hand-picked, simplified examples of the dataset.&nbsp;</p> <table> <tbody> <tr> <td> <p><strong>file</strong></p> </td> <td> <p><strong>aspects</strong></p> </td> <td> <p><strong>sentences</strong></p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/7/7a/Bongo_sound.wav" target="_blank" rel="noopener"><strong>🔈 Bongo sound.wav</strong></a></p> </td> <td> <p>['bongoes', 'percussion instrument', 'cumbia', 'drums']</p> </td> <td> <p>['a loop of bongoes playing a cumbia beat at 99 bpm']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/4/46/Example_of_double_tracking_in_a_pop-rock_song_%283_guitar_tracks%29.ogg" target="_blank" rel="noopener"><strong>🔈 Example of double tracking in a pop-rock song (3 guitar tracks).ogg</strong></a></p> </td> <td> <p>['bass', 'rock', 'guitar music', 'guitar', 'pop', 'drums']</p> </td> <td> <p>['a pop-rock song']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/6/62/OriginalDixielandJassBand-JazzMeBlues.ogg" target="_blank" rel="noopener"><strong>🔈 OriginalDixielandJassBand-JazzMeBlues.ogg</strong></a></p> </td> <td> <p>['jazz standard', 'instrumental', 'jazz music', 'jazz']</p> </td> <td> <p>['Considered to be a jazz standard', 'is an jazz composition']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/5/58/Colin_Ross_-_Etherea.ogg" target="_blank" rel="noopener"><strong>🔈 Colin Ross - Etherea.ogg</strong></a></p> </td> <td> <p>['chirping birds', 'ambient percussion', 'new-age', 'flute', 'recorder', 'single instrument', 'woodwind']</p> </td> <td> <p>['features a single instrument with delayed echo, as well as ambient percussion and chirping birds', 'a new-age composition for recorder']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/8/8b/Belau_rekid_%28instrumental%29.oga" target="_blank" rel="noopener"><strong>🔈 Belau rekid (instrumental).oga</strong></a></p> </td> <td> <p>['instrumental', 'brass band']</p> </td> <td> <p>['an instrumental brass band performance']</p> </td> </tr> <tr> <td> <p><strong>...</strong></p> </td> <td> <p>...</p> </td> <td> <p>...</p> </td> </tr> </tbody> </table> <h3>Dataset structure</h3> <p>We provide three variants of the dataset in the&nbsp;<code>data</code>&nbsp;folder.</p> <p>All are described in the paper.</p> <ol> <li><code>all.csv</code> contains all the data we collected, without any filtering.</li> <li><code>filtered_sf.csv</code> contains the data obtained using the&nbsp;<em>self-filtering</em> method.</li> <li><code>filtered_mc.csv</code> contains the data obtained using the <em>MusicCaps</em>&nbsp;dataset method.</li> </ol> <h3>File structure</h3> <p>Each CSV file contains the following columns:</p> <ul> <li><code>file</code>: the name of the audio file</li> <li><code>pageid</code>: the ID of the Wikipedia article where the text was collected from</li> <li><code>aspects</code>: the short-form (tag) description texts collected from the Wikipedia articles</li> <li><code>sentences</code>: the long-form (caption) description texts collected from the Wikipedia articles</li> <li><code>audio_url</code>: the URL of the audio file</li> <li><code>url</code>: the URL of the Wikipedia article where the text was collected from</li> </ul> <h3>Citation</h3> <div> <p>If you use this dataset in your research, please cite the following paper:</p> <div> <pre><code>@inproceedings{wikimute,</code><br><code> title = {WikiMuTe: {A} Web-Sourced Dataset of Semantic Descriptions for Music Audio},</code><br><code> author = {Weck, Benno and Kirchhoff, Holger and Grosche, Peter and Serra, Xavier},</code><br><code> booktitle = "MultiMedia Modeling",</code><br><code> year = "2024",</code><br><code> publisher = "Springer Nature Switzerland",</code><br><code> address = "Cham",</code><br><code> pages = "42--56",</code><br><code> doi = {10.1007/978-3-031-56435-2_4},</code><br><code> url = {https://doi.org/10.1007/978-3-031-56435-2_4},</code><br><code>}</code></pre> </div> </div> <h3>License</h3> <p>The data is available under the&nbsp;<a href="https://creativecommons.org/licenses/by-sa/3.0/" target="_blank" rel="noopener">Creative Commons Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) license</a>.</p> <p>Each entry in the dataset contains a URL linking to the article, where the text data was collected from.</p>

opencc-by-sa-3.0Dec 2023View details →
zenodo44/100

Data for Paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning"

<p><strong>Example Data for DeepReefMap</strong></p> <p>This dataset contains input videos in MP4 format taken with GoPro Hero 10 Cameras in Reefs in the Red Sea to demonstrate the DeepReefMap tool, which is described in the paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning" by Sauder et al.</p> <p>It contains a directory for model checkpoints for semantic segmentation, and for the 3D SLAM component:</p> <p>```<br>checkpoints/<br>&nbsp; &nbsp; &nbsp; &nbsp; segmentation_net.pth<br>&nbsp; &nbsp; &nbsp; &nbsp; sfm_net.pth<br>```</p> <p>It also contains videos to run the reconstruction with. See the detailed instructions for running reconstructions in https://github.com/josauder/mee-deepreefmap</p> <p>```<br>input_videos/<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_SINGLE_VIDEO.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_1_OF_2.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_2_OF_2.MP4<br>```</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Semantic segmentation model of construction waste landfill based on high-resolution satellite images

<p>CWLD_model project shows scripts and instructions on how to use this dataset (<a href="../records/10686118">https://zenodo.org/records/10686118</a>) to train a segmentation model. requirements.txt files provide the libraries you need to run your project. The README.md document details the deployment process and features of each module.</p> <p>You can also visit the GitHub page for scripts and instructions on how to use this dataset for visualizing and plotting basic statistics. The models and the code to execute them are released on&nbsp;<a href="https://github.com/huangleinxidimejd/CWLD_Model">https://github.com/huangleinxidimejd/CWLD_Model</a>.</p> <h2>Training details</h2> <p>The model was trained with two GPUs, an Nvidia GeForce RTX 2080Ti, and the following parameters:</p> <ul> <li>'train_batch_size': 4,</li> <li>'val_batch_size': 4,</li> <li>'train_crop_size': 512,</li> <li>'val_crop_size': 512,</li> <li>'lr': 0.001, # the learning rate used during training. It determines how quickly the model learns from the data</li> <li>'Epoch Times': 200,</li> <li>'gpu': correct,</li> <li>'weight_decay': 5E-4,</li> <li>'Momentum': 0.9,</li> <li>'print_freq': 100,</li> <li>'predict_step': 5,</li> </ul> <h2>usage</h2> <ul> <li>After downloading the dataset from Zenodo, place the train and val files from the Deep Learning Datasets file into the data folder of the CWLD semantic segmentation model.</li> <li>Open: CWLD_ Open the root directory in CWLD_model/dataset/ and start training with the WasteSeg_Train.py file. The modelss module provides five convolutional networks, Improved_DeeplabV3_plus, PSPNet, ResNet, SegNet, and UNet, which can be selected and modified accordingly.</li> <li>The utils package provides a large number of data processing tools to use.</li> <li>The trained model can be predicted from a EvalSeg.py file.</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo44/100

CLDF dataset derived Zalizniak et al.'s "Database of Semantic Shifts" from 2024

<p>Cite the source of the dataset as:</p> <blockquote> <p>Zalizniak A. et al (2024). Database of Semantic Shifts. Moscow: Institute of Linguistics, Russian Academy of Sciences. [Accessed on 05.02.2024]</p> </blockquote>

opencc-by-4.0Apr 2024View details →
zenodo44/100

MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish

<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets).&nbsp;</p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the&nbsp;documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud,&nbsp;the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others.&nbsp;</p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then,&nbsp;for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L &ndash; Scientific Literature:&nbsp;</strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials&nbsp;</strong></strong>contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;</li> <li><strong>[Subtrack 3 corpus] MESINESP-P &ndash; Patents:&nbsp;</strong>This corpus&nbsp;includes patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants&#39; predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p>&nbsp;</p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized&nbsp;(folder <em>separated</em>).&nbsp;This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>.&nbsp;</p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train:&nbsp; <ul> <li>training_set_track1_all.json: Full training set for subtrack 1.&nbsp;</li> <li>training_set_track1_only_articles.json:&nbsp;Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json:&nbsp;</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1.&nbsp;</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json:&nbsp;Manually annotated&nbsp;development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json:&nbsp;Manually annotated&nbsp;development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip&nbsp;</strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020&nbsp;set)</li> <li>List of synonyms (the descriptors and synonyms from&nbsp; Latin Spanish DeCS 2020&nbsp;set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo&nbsp;</strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p>&nbsp;</p> <p><strong>Data format&nbsp;description</strong></p> <p>The&nbsp;<strong>input text files</strong>&nbsp;for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial &gt;140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial &gt;130/80 y &lt;140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP&nbsp;<strong>entity mention files</strong>&nbsp;contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p>&nbsp;</p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L &ndash; Scientific Literature&nbsp;</strong>: &nbsp; <ul> <li><em><strong>Training set:&nbsp;</strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.&nbsp;We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;&nbsp;We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database.&nbsp;We distribute two different datasets: <ul> <li><strong>Articles training set:&nbsp;</strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set:&nbsp;</strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts&nbsp;from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: &nbsp; <ul> <li><strong>Training set:&nbsp;</strong>The training dataset contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a&nbsp;<a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team.&nbsp;</li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set:&nbsp;</strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P &ndash; Patents:&nbsp;</strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set:&nbsp;</strong>We provide a&nbsp;<strong>test set</strong>&nbsp;containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set.&nbsp;We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li>&nbsp;We provide this information to the participants as additional data in the &ldquo;Additional Data&rdquo; folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains&nbsp;entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p>&nbsp; </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2&nbsp;Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p>&nbsp;</p> <p>For further information, please&nbsp;email us at luis.gasco@bsc.es</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

A list of Swedish words that have experienced historical semantic changes

<p>This list contains a set of Swedish words that have experienced semantic change during the past centuries. The list has been collected during the VR funded project <a href="https://languagechange.org/">Towards Computational Lexical Semantic Change Detection</a>, ( 2018-01184) and is a work in progress. Because the work is currently on pause, we have chosen to release the list as is in the hope of facilitating collaboration and use in other research.</p> <p>The list has four columns in the following format:</p> <pre><code>Ord*, Vilken betydelseförändring har skett*, När skedde förändringen, Källor (exv SAOL) word, what change occurred, when the change occur, reference</code></pre> <p><br> Not all fields are filled for every word. Where there are multiple change periods, there are multiple lines with empty values for word and what change occurred, see the example with <em>egendomlig </em>below.</p> <pre><code>egendomlig,som utgör (ngns) egendom &gt; karaktäristisk (positiv) &gt; speciell (negativt!!),"A. Sen 1600tal, ",SAOB ,,B. Sen 1850, ,,"C. Sen ??, efter 190",</code></pre> <p>The .xlsx file contains links to the references.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

SEMFIRE forest dataset for semantic segmentation and data augmentation

<p><strong>SEMFIRE Datasets (Forest environment dataset)</strong></p> <p>These datasets are used for semantic segmentation and data augmentation and contain various forestry scenes. They were collected as part of the research work conducted by the Institute of Systems and Robotics, University of Coimbra <a href="https://isr.uc.pt/index.php/people?task=showprojects.show&amp;idProject=203">team</a> within the scope of the Safety, Exploration and Maintenance of Forests with Ecological Robotics (SEMFIRE, ref. <a href="http://semfire.ingeniarius.pt/">CENTRO-01-0247-FEDER-032691</a>) research project coordinated by <a href="https://ingeniarius.pt/">Ingeniarius Ltd.</a></p> <p>The semantic segmentation algorithms attempt to identify various semantic classes (e.g. background, live flammable materials, trunks, canopies etc.) in the images of the datasets.</p> <p>The datasets include diverse&nbsp;image types, e.g. original camera images and their labeled images. In total the SEMFIRE&nbsp;datasets include&nbsp;about 1700 image pairs. Each dataset includes corresponding .bag files.</p> <p>To launch those .bag files on your ROS environment, use the instructions on the following Github <a href="https://github.com/Forestry-Robotics-UC/fruc_rosbags">repository</a></p> <p>Description of<strong> </strong>each <strong>dataset:</strong></p> <ol> <li><strong>2019_2020_quinta_do_bolao_coimbra:</strong> Robot moving on a path through a forest environment</li> <li><strong>2020_ctcv_parking_lot_coimbra:</strong> Robot moving in a circle in a parking lot for testings</li> <li><strong>2020_sete_fontes_forest: </strong>A set of forest images acquired by hand-held apparatus</li> </ol> <p>Each <strong>dataset</strong> consists of following <strong>directories:</strong></p> <ol> <li><strong>images directory: </strong>diverse&nbsp;image types, e.g. original camera images and their labeled images</li> <li><strong>rosbags directory: </strong>.bag files, which correspond to the image directory</li> </ol> <p>Each <strong>images directory </strong>consists of following <strong>directories:</strong></p> <ul> <li><strong>img:</strong> original camera images</li> <li><strong>lbl:</strong> single channel images (ground truth) with corresponding labels for each image in<strong> img</strong></li> <li><strong>lbl_colored: </strong>camera <strong> </strong>images in&nbsp;<strong>lbl</strong> colorized according to different semantic classes (for more details see the datasets descriptions)</li> <li><strong>lbl_overlaid: </strong>camera images in <strong>img </strong>overlaid with corresponding labels (colored)</li> </ul> <p>Each <strong>rosbags directory </strong>contains .bag files with the following <strong>topics:</strong></p> <ul> <li><strong>2019_2020_quinta_do_bolao_coimbra_rosbags: </strong> <ul> <li>/back_lslidar_packet</li> <li>/dalsa_camera_720p/compressed</li> <li>/flir_ax8/compressed</li> <li>/front_lslidar_packet</li> <li>/gps_fix</li> <li>/gps_time</li> <li>/gps_vel</li> <li>/imu/data</li> <li>/realsense/aligned_depth_to_color/image_raw</li> <li>/realsense/color/camera_info</li> <li>/realsense/color/image_raw/compressed</li> <li>/realsense/depth/camera_info</li> <li>/realsense/depth/image_rect_raw/compressed</li> <li>/realsense/extrinsics/depth_to_color</li> </ul> </li> <li><strong>2020_ctcv_parking_lot_coimbra_rosbags:</strong> <ul> <li>/dalsa_camera_720p/compressed</li> <li>/gps_fix</li> <li>/gps_ime</li> <li>/fused_point_cloud</li> <li>/imu/data</li> <li>/imu/mag</li> <li>/imu/rpy</li> </ul> </li> <li><strong>2020_sete_fontes_forest_rosbags: </strong> <ul> <li>/realsense/camera_info</li> <li>/realsense/depth_compressed/compressedDepth</li> <li>/realsense/nir/left/compressed</li> <li>/realsense/nir/right/compressed</li> <li>/realsense/rgb/compressed</li> </ul> </li> </ul> <p>All datasets include a detailed description as a text file. In addition, they include a rosbag_info.txt file with a description for each ROS inside&nbsp;the .bag files as well as a description for each ROS topic.</p> <p>&nbsp;</p> <p>The following table shows the statistical description of typical portuguese woodland configurations with structured plantations of <em>Pinus pinaster </em>(<em>Pp, </em>pine trees) and <em>Eucalyptus globulus </em>(<em>Eg, </em>eucalyptus).</p> <table> <tbody> <tr> <td>&nbsp;</td> <td><strong>&quot;Low density&quot; structured plantation</strong></td> <td><strong>&quot;High density&quot; structured plantation</strong></td> </tr> <tr> <td><strong>Tree density (assuming plantation in rows spaced 3m apart in all cases)</strong></td> <td> <p><em>Eg</em>: 900 trees/ha</p> <p><em>Pp</em>: 450 trees/ha</p> </td> <td> <p><em>Eg</em>: 1400 trees/ha</p> <p><em>Pp</em>: 1250 trees/ha</p> </td> </tr> <tr> <td> <p><strong>Average heights and corresponding ages of plantation trees</strong></p> </td> <td> <p><em>Eg</em>: 12m (6 years old)</p> <p><em>Pp</em>: 10m (15 years old)</p> </td> <td> <p><em>Eg</em>: 12m (6 years old)</p> <p><em>Pp</em>: 10m (15 years old)</p> </td> </tr> <tr> <td> <p><strong>Maximum heights and corresponding fully-matured ages of plantation trees</strong></p> </td> <td> <p><em>Eg</em>: 20m (11 years old)</p> <p><em>Pp</em>: 30m (40 years old)</p> </td> <td> <p><em>Eg</em>: 20m (11 years old)</p> <p><em>Pp</em>: 30m (40 years old)</p> </td> </tr> <tr> <td> <p><strong>Diameter at chest level (DCL &ndash; 1,3m) of plantation trees (average/maximum)</strong></p> </td> <td> <p><em>Eg</em>: 15cm/25cm</p> <p><em>Pp</em>: 20cm/50cm</p> </td> <td> <p><em>Eg</em>: 15cm/25cm</p> <p><em>Pp</em>: 20cm/50cm</p> </td> </tr> <tr> <td> <p><strong>Natural density of herbaceous plants</strong></p> </td> <td> <p>30% of woodland area</p> </td> <td> <p>30% of woodland area</p> </td> </tr> <tr> <td> <p><strong>Natural density of bush and shrubbery</strong></p> </td> <td> <p>30% of woodland area</p> </td> <td> <p>30% of woodland area</p> </td> </tr> <tr> <td> <p><strong>Natural density of arboreal plants (not part of plantation)</strong></p> </td> <td> <p>5% of woodland area</p> </td> <td> <p>5% of woodland area</p> </td> </tr> </tbody> </table> <ul> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

CafeteriaSA corpus: Scientific abstracts annotated across different food semantic resources

<p>In the last decades, a great amount of work has been done in predictive modeling of issues related to human and environmental health. Resolution of issues related to healthcare is made possible by the existence of several biomedical vocabularies and standards, which play a crucial role in understanding health information, together with a large amount of health-related data. However, despite the large number of available resources and work done in the health and environmental domains, there is a lack of semantic resources that can be utilized in the food and nutrition domain, as well as their interconnections. For this purpose, in an European Food Safety Authority-funded project CAFETERIA, we have developed the first annotated corpus of 500 scientific abstracts that consists of 6,407 annotated food entities with regard to Hansard taxonomy, 4,299 for FoodOn, and 3,623 for SNOMED-CT.&nbsp; The CafeteriaSA corpus will enable further development of natural language processing methods for food information extraction from textual data that will allow extracting of food information from scientific textual data.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Prioritization of semantic over visuo- perceptual aspects in multi-item working memory

<p>All data and code supporting&nbsp;Prioritization of semantic over visuo- perceptual aspects in multi-item working memory</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources

<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse

<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon.&nbsp;</p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts.&nbsp;</p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website:&nbsp;<a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a>&nbsp;on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint:&nbsp;<a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO:&nbsp;<a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2:&nbsp;<a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3:&nbsp;<a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH:&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a>&nbsp;should not be used as a relation. We have replaced it with&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use&nbsp;<a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of&nbsp;<a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future:&nbsp;<a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesj&ouml;, QLIT, G&ouml;teborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kristr&ouml;m, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p>&nbsp;</p> <p>Thank you very much for your interest in our project!</p>

opengpl-3.0-or-laterJul 2024View details →
zenodo44/100

Improving Hypernymy Extraction with Distributional Semantic Classes

<p>In this paper, we show for the first time how distributionally-induced semantic classes can be helpful for extraction of hypernyms. We &nbsp;present a method for (1) inducing sense-aware semantic classes using distributional semantics and (2) using these induced semantic classes for filtering noisy hypernymy relations. Denoising of hypernyms is performed by labeling each semantic class with its hypernyms. On one hand, this allows us to filter out wrong extractions using the global structure of the distributionally similar senses. On the other hand, we infer missing hypernyms via label propagation to cluster terms. We conduct a large-scale crowdsourcing study showing that processing of automatically extracted hypernyms using our approach improves the quality of the hypernymy extraction both in terms of precision and recall. Furthermore, we show the utility of our method in the domain taxonomy induction task, achieving the state-of-the-art results on a benchmarking dataset.</p> <p>This particular page contains datasets related to the paper. Namely the input induced word senses, a database of hypernyms, and the output clusters of senses labeled with hypernyms -- the distributional semantic classes. The semantic classes are of two granularities, as described in the paper (coarse and fine grained).&nbsp;</p>

opencc-by-sa-4.0Feb 2018View details →
zenodo44/100

SemUr - Semantic Databases for Uralic Languages

<p>These databases are translated from <a href="http://mikakalevi.com/semfi">SemFi</a> by using Giellatekno XML dictionaries. The included python script can be used to update these databases or to create new ones for other languages.</p> <p>Currently, SemUr has the following languages</p> <ul> <li>SemSms - Skolt Sami</li> <li>SemKpv - Komi Zyrian</li> <li>SemMyv - Erzya</li> <li>SemMdf - Moksha</li> </ul> <p>&nbsp;</p> <p><strong>Cite as</strong></p> <p>H&auml;m&auml;l&auml;inen, Mika. (2018).&nbsp;<a href="https://helda.helsinki.fi//bitstream/handle/10138/282733/paper9.pdf?sequence=1">Extracting a Semantic Database with Syntactic Relations for Finnish to Boost Resources for Endangered Uralic Languages</a>. In The Proceedings of Logic and Engineering of Natural Language Semantics 15 (LENLS15)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2018View details →
zenodo44/100

SemFi - Finnish Semantic Database with Syntactic Relations

<p>SemFi is a semantic database for Finnish in which the words are linked to each other by the syntactic relations and their frequency in a big corpus.</p> <p>SemFi is based on the syntactic bigrams of The Finnish Internet Parsebank provided by Turku University.</p> <p>The semfi.db file is an SQLite database and it is the one that should be used. The results_json.zip is mainly intended for those who are interested in working with SemUr which is a translated version of SemFi.</p> <p>The previous version of this dataset has successfully been used in the hard AI task of creating Finnish poetry automatically.&nbsp;That data still powers the computationally creative system,<a href="http://runokone.cs.helsinki.fi/"> Poem Machine</a>.</p> <p>More information and an online UI to browse the data&nbsp;is available on&nbsp;<a href="https://mikakalevi.com/semfi">https://mikakalevi.com/semfi/</a>.</p> <p><strong>Cite as</strong></p> <p>H&auml;m&auml;l&auml;inen, Mika. (2018).&nbsp;<a href="https://helda.helsinki.fi//bitstream/handle/10138/282733/paper9.pdf?sequence=1">Extracting a Semantic Database with Syntactic Relations for Finnish to Boost Resources for Endangered Uralic Languages</a>. In The Proceedings of Logic and Engineering of Natural Language Semantics 15 (LENLS15)</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

A decade of Semantic Web research through the lenses of a mixed methods approach (Resources)

<p>This work has been submitted to&nbsp;<a href="http://www.semantic-web-journal.net/content/decade-semantic-web-research-through-lenses-mixed-methods-approach">Semantic Web Journal</a>. We provide here resources to reproduce our approach.</p> <p>In this paper, we aim to provide a broader and more complete picture of Semantic Web topics and trends by adopting a mixed methods methodology, which allows a combined use of both qualitative and quantitative approaches. Concretely, we build on a qualitative analysis of the main seminal papers, which adopt a top-down approach, and on quantitative results derived with three bottom-up data-driven approaches (<a href="https://technologies.kmi.open.ac.uk/Rexplore/">Rexplore</a>, <a href="http://saffron.insight-centre.org/">Saffron</a>, <a href="https://www.poolparty.biz/">PoolParty</a>), on a corpus of Semantic Web papers published in the last decade. In this process, we both use the latter for &ldquo;fact-checking&rdquo; on the former and also to derive key findings in relation to the strengths and weaknesses of top-down and bottom-up approaches to research topic identification.</p> <p>Please access the full set of resources at:&nbsp;<a href="https://aic.ai.wu.ac.at/qadlod/SW/">https://aic.ai.wu.ac.at/qadlod/SW/</a></p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

JeSemE models for lexical semantic change

<p>Models for diachronic lexical semantics used by&nbsp;the&nbsp;<a href="http://jeseme.org">Jena Semantic Explorer (JeSemE)</a> web site described in our&nbsp;<a href="http://aclweb.org/anthology/C18-2003">COLING 2018 paper &quot;JeSemE: A Website for Exploring Diachronic Changes in Word Meaning and Emotion&quot;</a>.</p> <p>Also described and applied in Johannes Hellrich&#39;s Ph.D. thesis &quot;Word Embeddings: Reliability &amp; Semantic Change&quot; who was funded by the&nbsp;Deutsche Forschungsgemeinschaft (DFG) within the graduate school &quot;The Romantic Model&quot;&nbsp;(GRK 2041/1).</p> <p>One ZIP&nbsp;file per corpus, each containing several CSV files:</p> <ul> <li>CHI.csv with&nbsp;&chi;<sup>2&nbsp;</sup>word association values (structure: word-id, word-id, time, value)</li> <li>EMBEDDING.csv with SVD-PPMI word embeddings (aligned;&nbsp;structure: word-id, time, values)</li> <li>EMOTION.csv with VAD&nbsp;word emotion values (structure: word-id, time, values)</li> <li>FREQUENCY.csv with relative word&nbsp;frequency values (structure: word-id, time, value)</li> <li>PPMI.csv with PPMI<sup>&nbsp;</sup>word association values (structure: word-id, word-id, time, value)</li> <li>SIMILARITY.csv with word embedding derived word similarity&nbsp;values (structure: word-id, word-id, time, value)</li> <li>WORDIDS.csv mapping words to their corpus specific&nbsp;IDs</li> </ul> <p>Corpora are:</p> <ul> <li> <p>coha:&nbsp;Corpus of Historical American English</p> </li> <li> <p>dta:&nbsp;Deutsches Textarchiv &#39;German Text Archive&#39;</p> </li> <li> <p>google_fiction:&nbsp;Google Books N-Gram corpus, English fiction subcorpus</p> </li> <li> <p>google_german:&nbsp;Google Books N-Gram corpus, German subcorpus</p> </li> <li> <p>rsc: Royal Society Corpus&nbsp;</p> </li> </ul>

opencc-by-4.0Mar 2018View details →
zenodo44/100

Semantic Segmentation Vineyard Rows

<p>Test dataset for semantic segmentation.<br> The datasets includes 500 RGB - images with the relative single-channel binary masks.</p> <p>Images are taken from the vineyards in Grugliasco - Turin - Piedmont Region -Italy</p> <p>&nbsp;</p> <p><strong>For more info please check out our work <a href="https://arxiv.org/abs/2107.00700">here</a></strong></p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Breaking Bad? Semantic Versioning and Impact of Breaking Changes in Maven Central (Dataset)

<p>The content presented in this repository accompanies the paper &quot;Breaking Bad? Semantic Versioning and Impact of Breaking Changes in Maven Central&quot; authored by Lina Ochoa, Thomas Degueule, Jean-R&eacute;my Falleri, and Jurgen Vinju. The paper was submitted and accepted in the Journal of Empirical Software Engineering (EMSE&#39;21). This study is an external and differentiated replication study of the paper <a href="https://jstvssr.github.io/assets/pdf/semantic-versioning-maven.pdf">&quot;Semantic Versioning and Impact of Breaking Changes in the Maven Repository&quot;</a>&nbsp;presented by Steven Raemaekers, Arie van Deursen, and Joost Visser.</p> <p><strong>Content</strong></p> <ul> <li><strong>README.md: </strong>document with the main description to start&nbsp;exploring&nbsp;the bundle.</li> <li><strong>data.zip: </strong>contains the datasets used within the study. These datasets must be used to get the same results like the ones presented in the article.</li> <li><strong>maven-api-dataset.zip:</strong> contains the code used to generate the datasets and to analyse the obtained results. Check the README.md&nbsp;file within this bundle for more information.</li> </ul> <p><strong>Relevant Links</strong></p> <ul> <li><strong>maven-api-dataset repository:</strong>&nbsp;<a href="https://github.com/tdegueul/maven-api-dataset">https://github.com/tdegueul/maven-api-dataset</a></li> <li><strong>maracas repository:</strong>&nbsp;<a href="https://github.com/crossminer/maracas">https://github.com/crossminer/maracas</a></li> <li><strong>Companion webpage:</strong>&nbsp;<a href="https://crossminer.github.io/maracas/2021/08/16/emse21/">https://crossminer.github.io/maracas/2021/08/16/emse21/</a></li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Data for: Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification

<p>This data was used in analyses for &quot;Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification&quot;.</p>

opencc-by-4.0Aug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record