Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.7.1
Dataset results
86 results for “semantic data”
Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"
<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>
The POPREBEL semantic social network data
<p>The <a href="https://populism-europe.com/poprebel/">POPREBEL project</a> explores the phenomenon of populism in Europe. As part of it, a team of ethnographers generated and coded the corpus contained in this dataset. It consists of coded interviews, realized between spring 2021 and spring 2022, to Internet users in Czechia, Germany and Poland, who used social media to gather information about the COVID-19 pandemic. The dataset is pseudonymized. POPREBEL is supported by the European Union's Horizon 2020 programme, grant n. 822682.</p> <ul> <li><a href="https://zenodo.org/record/7494327">Final ethnographic report.</a> Section 1.2 contains a detailed description of how and why data were collected.</li> <li><a href="https://wellbeing.edgeryders.eu">Funnel website</a> of the project.</li> <li><a href="https://hal.archives-ouvertes.fr/hal-02478720/document">About semantic social networks</a>.</li> <li><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data)</li> </ul>
The TREASURE semantic social network data on the circular economy aspect of automotive manufacturing
<p>The <a href="https://www.treasureproject.eu/">TREASURE</a> project looks at industrial innovation to address the problem of making onboard electronics in the automotive industry easier to recycle, increasing the industry's contribution to the circular economy. As part of it, a team of ethnographers generated and coded the corpus contained in this dataset. interviews conducted between January 2022 and June 2023 with car owners and enthusiasts at car industry events. The interviews focus on experiences with car electronics and perspectives on sustainability and the circular economy. The dataset is pseudonymized. TREASURE is supported by the European Union's Horizon 2020 programme, grant n. 101003587.</p>
Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Språkbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. "Korp-the corpus infrastructure of Språkbanken." <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dannélls, and Nina Tahmasebi. "Exploring the Quality of the Digital Historical Newspaper Archive KubHist." <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length > 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p> </p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>
Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10
<p>Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10</p> <p>Github: https://github.com/phucty/mtab4dbpedia<br> ---------------------------------------------------------------------------------------------------------------------------------------</p> <p>CEA: </p> <ul> <li> <p>Keep only valid entities in DBpedia 2016-10</p> </li> <li> <p>Resolve percentage encoding</p> </li> <li> <p>Add missing redirect entities</p> </li> </ul> <p>CTA: </p> <ul> <li> <p>Keep only valid types</p> </li> <li> <p>Resolve transitive types (parents and equivalent types of the specific type) with DBpedia ontology 2016-10</p> </li> </ul> <p>CPA:</p> <ul> <li> <p>Add equivalent properties</p> </li> </ul> <p>Statistic of Adapted Tabular data SemTab 2019</p> <pre><code>| | CEA | | | CPA | | | CTA | | | |---------|:--------:|:-------:|:------:|:--------:|:-------:|:------:|:--------:|---------|--------| | | Orginal | Adapted | Change | Orginal | Adapted | Change | Orginal | Adapted | Change | | Round 1 | 8418 | 8406 | -0.14% | 116 | 116 | 0.00% | 120 | 120 | 0.00% | | Round 2 | 463796 | 457567 | -1.34% | 6762 | 6762 | 0.00% | 14780 | 14333 | -3.02% | | Round 3 | 406827 | 406820 | 0.00% | 7575 | 7575 | 0.00% | 5762 | 5673 | -1.54% | | Round 4 | 107352 | 107351 | 0.00% | 2747 | 2747 | 0.00% | 1732 | 1717 | -0.87% |</code></pre> <p> </p> <p>---------------------------------------------------------------------------------------------------------------------------------------<br> DBpedia 2016-10 extra resources: (Original dataset http://downloads.dbpedia.org/2016-10/)</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_classes_2016-10.csv</p> <p>Information: DBpedia classes and parents: (We remove the abstract types: Agent, Thing)</p> <p>Total: 759 classes</p> <p>Structure: [class, parents (separate with space)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "City","Location Place PopulatedPlace Settlement"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_properties_2016-10.csv</p> <p>Information: DBpedia properties and these equivalents</p> <p>Total: 2865 properties</p> <p>Structure: [property, it’s equivalent properties] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "restingDate","deathDate"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_domains_2016-10.csv</p> <p>Information: DBpedia properties and these domain types</p> <p>Total: 2421 properties (have types as their domain)</p> <p>Structure: [property, type (domain)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "deathDate","Person"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_entities_2016-10.jsonl.bz2 </p> <p>Information: DBpedia entity dump</p> <p>Format: json list bz2 (bz2 Compressed json list)</p> <p>Source: DBpedia dump 2016-10 core</p> <p>Total: 5,289,577 entities (No disambiguation entities)</p> <p>Structure:</p> <p>An entity: for example “Tokyo”: (datatype: dictionary),</p> <p>{</p> <p>'wd': 'Q1322032', (Wikidata ID, datatype: string)</p> <p>'wp': 'Tokyo', (Wikipedia ID, add prefix <a href="https://en.wikipedia.org/wiki/">https://en.wikipedia.org/wiki/</a> + wp to get the Wikipedia URL, datatype: string)</p> <p>'dp': 'Tokyo', (DBpedia ID, add prefix <a href="http://dbpedia.org/resource/">http://dbpedia.org/resource/</a> + dp to get the DBpedia URL, datatype: string)</p> <p>'label': 'Tokyo', (Entity label, datatype: string)</p> <p>'aliases': ['To-kyo', 'Tôkyô Prefecture', ..], (Other entity names, datatype: list) </p> <p>'aliases_multilingual': ['东京小子', 'طوكيو', ...], (Other entity names in multilingual, datatype: list)</p> <p>'types_specific': 'City', (Entity direct type, datatype: string) </p> <p>'types_transitive': ['Human settlement', 'City', 'PopulatedPlace', 'Location', 'Place', 'Settlement'], (Entity transitive types, datatype: list)</p> <p>'claims_entity': { (entity statements, datatype: dictionary. Keys: properties, Values: list of tail entities)</p> <p>'governingBody': ['Tokyo Metropolitan Government'], </p> <p> 'subdivision': ['Honshu', 'Kantō region'],</p> <p>...</p> <p>},</p> <p>'claims_literal': {</p> <p>'string': { (String literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>'postalCode': ['JP-13'], </p> <p>'utcOffset': ['+09:00', '+9'],</p> <p>…</p> <p>}</p> <p>'time': { (Time literal: datatype: dictionary. Keys: properties, Values: list of date time</p> <p>'populationAsOf': ['2016-07-31'], </p> <p>...</p> <p>}), </p> <p>'quantity': { (Numerical literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>populationDesity: [6224.66, 6349.0], </p> <p>'maximumElevation': [2017], </p> <p>...</p> <p>},</p> <p>'pagerank': 2.2167366040153352e-06 (Entity page rank score calculated on DBpedia Graph)</p> <p>}</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Data for Paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning"
<p><strong>Example Data for DeepReefMap</strong></p> <p>This dataset contains input videos in MP4 format taken with GoPro Hero 10 Cameras in Reefs in the Red Sea to demonstrate the DeepReefMap tool, which is described in the paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning" by Sauder et al.</p> <p>It contains a directory for model checkpoints for semantic segmentation, and for the 3D SLAM component:</p> <p>```<br>checkpoints/<br> segmentation_net.pth<br> sfm_net.pth<br>```</p> <p>It also contains videos to run the reconstruction with. See the detailed instructions for running reconstructions in https://github.com/josauder/mee-deepreefmap</p> <p>```<br>input_videos/<br> GX_SINGLE_VIDEO.MP4<br> GX_VIDEO_1_OF_2.MP4<br> GX_VIDEO_2_OF_2.MP4<br>```</p>
MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish
<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets). </p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p> </p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud, the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others. </p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then, for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L – Scientific Literature: </strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials </strong></strong>contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. </li> <li><strong>[Subtrack 3 corpus] MESINESP-P – Patents: </strong>This corpus includes patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants' predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p> </p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized (folder <em>separated</em>). This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>. </p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train: <ul> <li>training_set_track1_all.json: Full training set for subtrack 1. </li> <li>training_set_track1_only_articles.json: Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json: </li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1. </li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json: Manually annotated development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json: Manually annotated development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip </strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020 set)</li> <li>List of synonyms (the descriptors and synonyms from Latin Spanish DeCS 2020 set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo </strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p> </p> <p><strong>Data format description</strong></p> <p>The <strong>input text files</strong> for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial >140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial >130/80 y <140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP <strong>entity mention files</strong> contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p> </p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L – Scientific Literature </strong>: <ul> <li><em><strong>Training set: </strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish. We have filtered out empty abstracts and non-Spanish abstracts. We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database. We distribute two different datasets: <ul> <li><strong>Articles training set: </strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set: </strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: <ul> <li><strong>Training set: </strong>The training dataset contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a <a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team. </li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set: </strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P – Patents: </strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set: </strong>We provide a <strong>test set</strong> containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes “A61P” and “A61K31”.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li> We provide this information to the participants as additional data in the “Additional Data” folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p> </p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p> </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2 Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p> </p> <p>For further information, please email us at luis.gasco@bsc.es</p>
SEMFIRE forest dataset for semantic segmentation and data augmentation
<p><strong>SEMFIRE Datasets (Forest environment dataset)</strong></p> <p>These datasets are used for semantic segmentation and data augmentation and contain various forestry scenes. They were collected as part of the research work conducted by the Institute of Systems and Robotics, University of Coimbra <a href="https://isr.uc.pt/index.php/people?task=showprojects.show&idProject=203">team</a> within the scope of the Safety, Exploration and Maintenance of Forests with Ecological Robotics (SEMFIRE, ref. <a href="http://semfire.ingeniarius.pt/">CENTRO-01-0247-FEDER-032691</a>) research project coordinated by <a href="https://ingeniarius.pt/">Ingeniarius Ltd.</a></p> <p>The semantic segmentation algorithms attempt to identify various semantic classes (e.g. background, live flammable materials, trunks, canopies etc.) in the images of the datasets.</p> <p>The datasets include diverse image types, e.g. original camera images and their labeled images. In total the SEMFIRE datasets include about 1700 image pairs. Each dataset includes corresponding .bag files.</p> <p>To launch those .bag files on your ROS environment, use the instructions on the following Github <a href="https://github.com/Forestry-Robotics-UC/fruc_rosbags">repository</a></p> <p>Description of<strong> </strong>each <strong>dataset:</strong></p> <ol> <li><strong>2019_2020_quinta_do_bolao_coimbra:</strong> Robot moving on a path through a forest environment</li> <li><strong>2020_ctcv_parking_lot_coimbra:</strong> Robot moving in a circle in a parking lot for testings</li> <li><strong>2020_sete_fontes_forest: </strong>A set of forest images acquired by hand-held apparatus</li> </ol> <p>Each <strong>dataset</strong> consists of following <strong>directories:</strong></p> <ol> <li><strong>images directory: </strong>diverse image types, e.g. original camera images and their labeled images</li> <li><strong>rosbags directory: </strong>.bag files, which correspond to the image directory</li> </ol> <p>Each <strong>images directory </strong>consists of following <strong>directories:</strong></p> <ul> <li><strong>img:</strong> original camera images</li> <li><strong>lbl:</strong> single channel images (ground truth) with corresponding labels for each image in<strong> img</strong></li> <li><strong>lbl_colored: </strong>camera <strong> </strong>images in <strong>lbl</strong> colorized according to different semantic classes (for more details see the datasets descriptions)</li> <li><strong>lbl_overlaid: </strong>camera images in <strong>img </strong>overlaid with corresponding labels (colored)</li> </ul> <p>Each <strong>rosbags directory </strong>contains .bag files with the following <strong>topics:</strong></p> <ul> <li><strong>2019_2020_quinta_do_bolao_coimbra_rosbags: </strong> <ul> <li>/back_lslidar_packet</li> <li>/dalsa_camera_720p/compressed</li> <li>/flir_ax8/compressed</li> <li>/front_lslidar_packet</li> <li>/gps_fix</li> <li>/gps_time</li> <li>/gps_vel</li> <li>/imu/data</li> <li>/realsense/aligned_depth_to_color/image_raw</li> <li>/realsense/color/camera_info</li> <li>/realsense/color/image_raw/compressed</li> <li>/realsense/depth/camera_info</li> <li>/realsense/depth/image_rect_raw/compressed</li> <li>/realsense/extrinsics/depth_to_color</li> </ul> </li> <li><strong>2020_ctcv_parking_lot_coimbra_rosbags:</strong> <ul> <li>/dalsa_camera_720p/compressed</li> <li>/gps_fix</li> <li>/gps_ime</li> <li>/fused_point_cloud</li> <li>/imu/data</li> <li>/imu/mag</li> <li>/imu/rpy</li> </ul> </li> <li><strong>2020_sete_fontes_forest_rosbags: </strong> <ul> <li>/realsense/camera_info</li> <li>/realsense/depth_compressed/compressedDepth</li> <li>/realsense/nir/left/compressed</li> <li>/realsense/nir/right/compressed</li> <li>/realsense/rgb/compressed</li> </ul> </li> </ul> <p>All datasets include a detailed description as a text file. In addition, they include a rosbag_info.txt file with a description for each ROS inside the .bag files as well as a description for each ROS topic.</p> <p> </p> <p>The following table shows the statistical description of typical portuguese woodland configurations with structured plantations of <em>Pinus pinaster </em>(<em>Pp, </em>pine trees) and <em>Eucalyptus globulus </em>(<em>Eg, </em>eucalyptus).</p> <table> <tbody> <tr> <td> </td> <td><strong>"Low density" structured plantation</strong></td> <td><strong>"High density" structured plantation</strong></td> </tr> <tr> <td><strong>Tree density (assuming plantation in rows spaced 3m apart in all cases)</strong></td> <td> <p><em>Eg</em>: 900 trees/ha</p> <p><em>Pp</em>: 450 trees/ha</p> </td> <td> <p><em>Eg</em>: 1400 trees/ha</p> <p><em>Pp</em>: 1250 trees/ha</p> </td> </tr> <tr> <td> <p><strong>Average heights and corresponding ages of plantation trees</strong></p> </td> <td> <p><em>Eg</em>: 12m (6 years old)</p> <p><em>Pp</em>: 10m (15 years old)</p> </td> <td> <p><em>Eg</em>: 12m (6 years old)</p> <p><em>Pp</em>: 10m (15 years old)</p> </td> </tr> <tr> <td> <p><strong>Maximum heights and corresponding fully-matured ages of plantation trees</strong></p> </td> <td> <p><em>Eg</em>: 20m (11 years old)</p> <p><em>Pp</em>: 30m (40 years old)</p> </td> <td> <p><em>Eg</em>: 20m (11 years old)</p> <p><em>Pp</em>: 30m (40 years old)</p> </td> </tr> <tr> <td> <p><strong>Diameter at chest level (DCL – 1,3m) of plantation trees (average/maximum)</strong></p> </td> <td> <p><em>Eg</em>: 15cm/25cm</p> <p><em>Pp</em>: 20cm/50cm</p> </td> <td> <p><em>Eg</em>: 15cm/25cm</p> <p><em>Pp</em>: 20cm/50cm</p> </td> </tr> <tr> <td> <p><strong>Natural density of herbaceous plants</strong></p> </td> <td> <p>30% of woodland area</p> </td> <td> <p>30% of woodland area</p> </td> </tr> <tr> <td> <p><strong>Natural density of bush and shrubbery</strong></p> </td> <td> <p>30% of woodland area</p> </td> <td> <p>30% of woodland area</p> </td> </tr> <tr> <td> <p><strong>Natural density of arboreal plants (not part of plantation)</strong></p> </td> <td> <p>5% of woodland area</p> </td> <td> <p>5% of woodland area</p> </td> </tr> </tbody> </table> <ul> </ul>
CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources
<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>
Data for: Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification
<p>This data was used in analyses for "Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification".</p>
Dataset: Comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research.
<p>Supplementary material for a comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research. We conducted a relevance evaluation with 6 users over 19 search questions in two search interfaces.</p> <p>The users provided up to five search questions and relevant keywords from their research background. We setup a dataset search over a corpus of ~92,000 randomly selected metadata files from GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>). For each of their own search queries, the users got two result sets presented. The first one displayed results obtained from a keyword search. The second panel contained dataset results from a prototypical semantic search. Instead of results with exact mentions of the query terms, the semantic search also presented related results with synonyms and more specific terms or terms obtained from concept nodes of a higher hierarchy level.</p> <p>Each user rated the relevance of his/her own search queries on a 7-point Likert scale for both search results.<br> In addition, users also assessed the expanded keywords for each question.</p> <p>More information can be found in our publication:</p> <p>Löffler, F. and Klan, F. (2016): Does Term Expansion Matter for the Retrieval of Biodiversity Data? in Joint Proceedings of the Posters and Demos Track of the 12th International Conference on Semantic Systems - SEMANTiCS2016 and the 1st International Workshop on Semantic Change & Evolving Semantics (SuCCESS'16), co-located with the 12th International Conference on Semantic Systems (SEMANTiCS 2016),2016, <a href="http://ceur-ws.org/Vol-1695/paper2.pdf">http://ceur-ws.org/Vol-1695/paper2.pdf</a></p> <p> </p>
SEMAFORA Semantic Reference Data Models
<p>To support the aim of the Semafora project, a series of Semantic Reference Data Models were created to provide a target semantic structure for the integration of standard archaeological survey data. </p> <p> </p> <p>The following models constitute the Semafora SRDM package:</p> <p> </p> <p>Place: This model is used to document any places associated with the archaeological survey.</p> <p> </p> <p>Institution: This model is used to document any institution associated with the survey.</p> <p> </p> <p>Period: This model is used to document the generic historical period assigned to the production of artefacts, existence of sites or other observable archaeological and historical events.</p> <p> </p> <p>Feature: This model is used to document any physical features, such as walls and other human-made structures observable on the field.</p> <p> </p> <p>Project: This model is used to document the overarching project, a part of which is the archaeological survey. Some projects may involve surveys, excavations, and other archaeological activities.</p> <p> </p> <p>Site: This model is used to document a site declared as archaeological as a result of the survey process.</p> <p> </p> <p>Digital Object: This model is used to document any type of digital asset associated with the survey.</p> <p> </p> <p>Survey Unit: This model is used to document a defined survey unit where the survey activity happens. It has both the properties of a place with dimensions and coordinates and of a physical thing from which samples can be collected.</p> <p> </p> <p>Collection: This model is used to document a collection of physical things, usually artefacts, collected while surveying.</p> <p> </p> <p>Artefact: This model is used to document individual artifacts collected from while surveying as a part of a larger collection of material things or as a singular artefact collection or documentation.</p> <p> </p> <p>Image: This model is used to document any image representing components of the archaeological survey, such as artefacts, features, places, people, etc.</p> <p> </p> <p>Observation: This model is used to document the act of observation usually associated with archaeological sites or survey units and the properties assigned to those as a result of the observation.</p> <p> </p> <p>Bibliography: This model is used to document any textual object associated with the survey or any components of it.</p> <p> </p> <p>Sample: This model is used to document a material sample of the survey unit. It partially overlaps with collection but acts as a parent sample that may contain other physical things besides human-made objects.</p> <p> </p> <p>Person: This model is used to document an individual person (alive or dead) involved in some way in the survey process.</p> <p> </p> <p>These models are intended to be used in order to guide semantic data mapping processes as well as to provide instructions for the creation of a target data semantic data management system.</p> <p> </p> <p>Each model’s semantic reference data model description is stored here as a csv. The ongoing curation and updating of these SRDMs is undertaken using the Zellij system and can be accessed here:</p> <p> </p> <p><a href="https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c">https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c</a></p>
Accompanying Dataset migr_asyappctzm for Efficient Analytical Queries on Semantic Web Data Cubes
<p>This dataset shows how the Eurostat data cube in the orginal publicatin is modelled in QB4OLAP.</p> <p>This data is based on statistical data about asylum applications to the European Union, provided by Eurostat on</p> <p><a href="http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm">http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm</a></p> <p>Further data has been integrated from: https://github.com/lorenae/qb4olap/tree/master/examples</p>
Semantic Enrichment of the Laboratory Data Dictionary of the Study of Health in Pomerania (SHIP-START-4) with LOINC; Detailed Mapping Results
<p>Unlike West Germany, high morbidity and mortality have been observed in East Germany over the last century. The regional population-based Study of Health in Pomerania (SHIP) therefore investigates the long-term progression of sub-clinical findings, their determinants and prognostic values, to acquire knowledge that facilitates early diagnosis and thus helps prevent the progression of disease. The SHIP covers various areas of patient health. Each SHIP data set is accompanied by a data dictionary (DD) which provides descriptions of variables and definitions.</p> <p>This work shows the detailed mapping results of the semantic enrichment of the SHIP-START-4 medical laboratory data dictionary with LOINC codes. This work also provides detailed descriptions of the concepts applied in the semnatic enrichment. The results of this work serve as a critical step towards improving its interoperability and hence FAIRness for the SHIP laboratory-related measurements. </p>
Supplementary Data for "Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"
<p>This dataset contains quality assessment results for 26 vocabularies. The assessment was conducted using the <a href="https://skos-play.sparna.fr/skos-testing-tool/">qSKOS vocabulary quality assessment tool</a>.</p> <p>The 26 assessed vocabularies were converted from their original formats into the Simple Knowledge Organization System (SKOS) data model using the approach described in our paper titled <a href="https://doi.org/10.1007/978-3-031-62362-2_9">"Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"</a>, presented at the <a href="https://doi.org/10.1007/978-3-031-62362-2">24th International Conference on Web Engineering (ICWE 2024)</a>.</p> <p>The dataset contains a quality assessment for the following vocabularies:</p> <ol> <li>A Taxonomy of Evaluation Towards Standards</li> <li>Cross-Device Taxonomy</li> <li>What Makes a Data-driven Business Model? A Consolidated Taxonomy</li> <li>DDI Aggregation Method</li> <li>DDI Mode of Collection</li> <li>Building a New Taxonomy for Data Discretization Techniques</li> <li>Demopaedia</li> <li>Data Science Glossary</li> <li>A Taxonomy of Evaluation Approaches in Software Engineering</li> <li>Evaluation Thesaurus</li> <li>The Glossary of Human Computer Interaction</li> <li>Human-Factors Taxonomy</li> <li>A Taxonomy to Structure and Analyze Human–Robot Interaction</li> <li>A Taxonomy of Interaction for Instructional Multimedia</li> <li>A Taxonomy of Interrogation Methods</li> <li>Design Vocabulary for Human–IoT Systems Communication</li> <li>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors</li> <li>Thesaurus Mass Communication</li> <li>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey</li> <li>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction</li> <li>A Human-Centered Taxonomy of Interaction Modalities and Devices</li> <li>A Taxonomy of Spatial Interaction Patterns and Techniques</li> <li>A Taxonomy of Social Errors in Human-Robot Interaction</li> <li>Taxonomy of Digital Research Activities in the Humanities</li> <li>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions </li> <li>Cross-Device Interaction</li> </ol>
SemTab 2024: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets - WikidataTables2024R1 and WikidataTables2024R2
<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Round 1, Wikidata Tables. Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata https://www.wikidata.org/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Data for the Article: Cross-validation of a semantic segmentation network for natural history collection specimens
<p>This deposit contains six datasets which were used for testing and validating a semantic segmentation network. The purpose was to evaluate the suitability of the segmentation network for use in the processing of images from Natural History Collections.</p>
Data used to evaluate ORBITS: Optimal Repair-Based Inconsistency-Tolerant Semantics
<p>This dataset provides the input files that were used in the evaluation of the ORBITS system (Optimal Repair-Based Inconsistency-Tolerant Semantics, <a href="https://github.com/bourgaux/orbits">https://github.com/bourgaux/orbits</a>). A detailed description is available in a technical report on arXiv (<a href="https://arxiv.org/abs/2202.07980">https://arxiv.org/abs/2202.07980</a>).</p> <p><strong>Content:</strong></p> <p>Folders <em>cqapri_benchmark</em>, <em>food_inspection_benchmark</em>, and <em>physicians_benchmark</em> contain JSON files of conflict graphs and candidate queries and their causes.<br> These files are named using the following pattern: files of candidate answers and their causes are named <database>_<query>_answers_causes.json, and conflict graphs are named <database>_conflictGraph_<priority relation>.json where <priority relation> says whether the priority relation is score-structured (prio_score) or not (prio_non_score) and the probability (p<proba>) or number of scores (n<number>) used to build the priority relation.</p> <p>Folder <em>original_datasets_and_queries</em> contains the Food Inspection and Physicians datasets used to generate files from <em>food_inspection_benchmark</em> and <em>physicians_benchmark</em>.<br> Files from <em>cqapri_benchmark</em> have been generated from the CQAPri benchmark available at <a href="https://lahdak.lri.fr/CQAPri/CQAPri.php">https://lahdak.lri.fr/CQAPri/CQAPri.php</a>.<br> In all cases, we use ProvSQL (<a href="https://github.com/PierreSenellart/provsql">https://github.com/PierreSenellart/provsql</a>) to build conflict graphs and causes from the datasets.</p>
M4.4 (associated data) - FAIR-IMPACT Review of Semantic Artefact Catalogues and technologies
<p>This dataset (version 1) takes the form of a spreadsheet corresponding to the associated data described by <a href="../records/12799796" target="_blank" rel="noopener"><strong>M4.4 - Review of Semantic Artefact Catalogues and guidelines for serving FAIR semantic artefacts in EOSC</strong></a></p> <p>The spreadsheept (available here in ODS format) contains the listing of Semantic Artefact Catalogues (SACs) done within FAIR-IMPACT's WP4, their classifications (by status, type, discipline and technology) and the evaluation of their FAIR-enabling dimensions. </p> <p>A "live" version of this spreadsheet is available as an open Google Sheet open for comments and suggestions. We will take external contributions and comments into consideration when producing new verion of this dataset. Contributions could be of several types: </p> <ul> <li>New SAC or modification of the ones currently identified;</li> <li>New SAC technology or modification of the ones currently identified;</li> <li>New FAIR-enabling assessment or modification of the ones currently available.</li> </ul> <p>For more information, interested parties may contact the authors.</p> <p> </p>
Figure 4. After merging, overview is more transparent. Tens of persons were merged together into clusters in order to clarify the visualization. Firms and persons are recognized based on their icons.-Browsing Semantic Data in Slovakia
<p>The usefulness of such visualization has its key points regarding connections. Thanks to SBR browsing module, we were able to get 22 firm records for “Váhostav” query. Between any 2 companies, connections may be (and often are) not bidirectional, so, in order to navigate through connections, we have refined all 22 records. Although, even being filtered, graph is still complex. And it is possible to further navigate and search for outgoing connections, for example firm “MERLIN TRADE, a.s.” on Fig.4 contains item on “Ján Kato”, which is already included in our graph and connected to “VÁHOSTAV&SK&DEVELOPEMENT” on bottom left side and “VÁHOSTAV&SK, a.s.” in the center. Edge coloring and drawing is helpful with overlapped edges. For methods of visualization, including coloring, we refer to studies of H. Omote and K. Sugiyama (2006), and I. Herman, G. Melanon, and M. S. Marshall (2000) or our study on graph clutter filtering and connectivity distance (Mojzis & Laclavik, 2014).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.