Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo44/100

An Annotated Corpus of Tonal Piano Music from the Long 19th Century

<p>This corpus has been created within the&nbsp;<a href="https://github.com/DCMLab/dcml_corpora">DCML corpus initiative</a>&nbsp;and employs&nbsp;the&nbsp;<a href="https://github.com/DCMLab/standards">DCML harmony annotation standard</a>.</p> <p><strong>Version 1</strong>&nbsp;has been released for submitting it as part of the data&nbsp;report&nbsp;<code>Hentschel, J., Rammos, Y., Neuwirth, M., Rohrmeier, M. (forthcoming). An Annotated Corpus of Tonal Piano Music from the Long 19th Century</code>&nbsp;that accompanies nine corpora grouped under the DOI&nbsp;<a href="https://doi.org/10.5281/zenodo.7483349">10.5281/zenodo.7483349</a>.</p> <p><strong>Version 1.1</strong>&nbsp;comes with a complete set of metadata and score headers.&nbsp;Among more accurate composition dates,&nbsp;the&nbsp;metadata now include URIs that identify the compositions in terms of&nbsp;the&nbsp;<a href="https://viaf.org/">Virtual International Authority File (VIAF)</a>,&nbsp;<a href="https://www.wikidata.org/">Wikidata</a>,&nbsp;<a href="https://imslp.org/">IMSLP</a>&nbsp;and&nbsp;<a href="https://musicbrainz.org/">MusicBrainz</a>.&nbsp;The data has been re-extracted from the scores&nbsp;using&nbsp;<a href="https://pypi.org/project/ms3/">ms3 1.1.1</a>.</p> <p>The publication covers the following corpora (the DOI links always point at the latest version respectively):</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7473560">Ludwig van Beethoven - Piano Sonatas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473566">Fr&eacute;d&eacute;ric Chopin - Mazurkas</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473568">Claude Debussy - Suite Bergamasque</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473576">Anton&iacute;n Dvoř&aacute;k - Silhouettes</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473580">Franz Liszt - Ann&eacute;es de P&egrave;lerinage</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473528">Nikolai Medtner - Tales</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473582">Robert Schumann - Kinderszenen</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473586">Pyotr Tchaikovsky - The Seasons</a></li> <li><a href="https://doi.org/10.5281/zenodo.7473578">Edvard Grieg - Lyric Pieces</a></li> </ul> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2022View details →
zenodo44/100

The Tsez Annotated Corpus Project

<p>Cite the source of the dataset as:</p> <blockquote> <p>Abdulaev, A.K. &amp; I. K. Abdullaev. 2010. Cezyas folklor/Dido (Tsez) folklore/Didojskij (cezskij) fol´klor. Leipzig–Makhachkala: "Lotos".</p> </blockquote>

opencc-by-4.0Sep 2022View details →
zenodo44/100

The grammatically annotated corpus of the pericopes of the Old Lithuanian Postil of Jonas Bretkūnas

<p>This grammatically annoted corpus aims at facilitating linguistic research on Old Lithuanian based on the Postil of Jonas Bretkūnas from the year 1591. In addition to the two subcorpora "pericopes" and "homilies" of the version 1.0, the current version 2.0<br>includes also "passion harmony", "pericope (prophet)" and "prayer". In this new version, the tokens occurring in biblical citations can be specifically searched for. The search results also show the corresponding passage in the BrP pericopes.<br>For full description see the BrP_2.0._documentation.md included in the files.</p> <p>The pericopes were annotated using SIL Toolbox and converted to be used in the search-tool ANNIS using the conversion tool PEPPER.</p> <p>Three formats are provided in this release: 1. the Toolbox files, 2. the transitional Excel files and 3. a zipped folder to be imported into ANNIS.</p> <p>Created in the project B02, <em>Emergence and change of registers: The case of Lithuanian and Latvian</em> of the CRC 1412 "Register" (funded by the Deutsche Forschungsgemeinschaft: DFG, German Research Foundation: 416591334).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo44/100

PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities

<h2><a href="https://marianaossilva.github.io/pportal_ner/" target="_blank" rel="noopener">PPORTAL_ner</a></h2> <h3>An Annotated Dataset of Portuguese Literary Entities</h3> <p>The corpus is tailored to Brazilian and Portuguese literary texts, containing annotations for five entity categories, including PER, LOC, GPE, ORG, and DATE. Within a diverse collection of 25 literary works, it offers a total of 125,059 tokens and 5,266 annotated entities. This dataset contributes to the development of potentially more accurate and context-aware NER models, as well as to encourage further exploration within Portuguese literature.<br><br></p> <h2>Corpus Statistics</h2> <p>Our corpus is sourced from&nbsp;<a href="https://doi.org/10.5281/zenodo.5178063">PPORTAL</a>, an extensive repository of metadata containing over 80,000 public domain literary works in the Portuguese language, predominantly derived from Brazil and Portugal.&nbsp;<a href="https://doi.org/10.5281/zenodo.5178063">PPORTAL</a>&nbsp;aggregates data from three digital libraries:&nbsp;<a href="https://www.dominiopublico.gov.br/">Dom&iacute;nio P&uacute;blico</a>,&nbsp;<a href="https://projectoadamastor.org/">Projecto Adamastor</a>, and&nbsp;<a href="https://www.literaturabrasileira.ufsc.br/">Biblioteca Digital de Literatura dos Pa&iacute;ses Lus&oacute;fonos (BLPL)</a>.</p> <p>To simplify referencing, this new dataset is called PPORTAL_ner. PPORTAL_ner selection process contains a diverse range of 25 individual literary works, spanning different authors and literary styles. All of these texts were published prior to 1953, adhering to the current criteria for public domain status in Brazil, with the majority falling within the timeframe spanning from 1554 to 1938.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Genome annotation workflow for Effrenium voratum RCC1521

<p>Scripts of complete genome annotation workflow for Effrenium voratum RCC1521, associated with the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in&nbsp;<em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p>See <strong>README_EvRCC1521.txt</strong> for more detail.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Penicillium fuscoglaucum Pf_T2 Genome Assembly and Annotation

<p>During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum&nbsp;Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of 5 scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum, and contained a lineage specific, intact cylcopaizonic acid (CPA) gene cluster.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

A collection of annotated soundscape recordings from western Kenya

<p>This collection contains 35 soundscape recordings of 32 hours total duration, which have been annotated with 10,294 labels for 176 different bird species from western Kenya. The data were recorded in 2021 and 2022 west and southwest of Lake Baringo in Baringo County, Kenya. This collection has partially been featured as test data in the 2023 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>For this collection, AudioMoths and SWIFT recording units were deployed at multiple locations west and southwest of Lake Baringo, Baringo County, Kenya between Dezember 2021 and February 2022. Recording locations cover a variety of habitats from open grasslands to semi-arid scrubland and mountain forests. Recordings were originally sampled at 48 kHz and converted to MP3 for faster file transfer. For publication, all files were resampled to 32 kHz and converted to FLAC.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>A total of 32 hours of audio from various sites west and southwest of Lake Baringo were selected for annotation. Annotators were tasked with identifying and labeling each bird call they could discern, excluding any calls that were too weak or indiscernible. The annotation process was carried out using Audacity. Provided labels mark the center of each bird call. In this collection, we use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list). Parts of this dataset have previously been used in the 2023 BirdCLEF competition.&nbsp;</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the &ldquo;soundscape_data.zip&rdquo; file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in EAT (UTC+3). As an example, the file &ldquo;KEN_001_20211207_153852.flac&rdquo; has sequential ID 001 and was recorded on December 7th 2021 at 15:38:52 EAT. Ground truth annotations are listed in &ldquo;annotations.csv&rdquo; where each line specifies the corresponding filename, start and end time in seconds, and an eBird species code. These species codes can be assigned to scientific and common name of a species with the &ldquo;species.csv&rdquo; file. The approximate recording location with longitude and latitude can be found in the &ldquo;recording_location.txt&rdquo; file.</p> <p><strong>Acknowledgements</strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection. In particular, our thanks go to Francis Cherutich for setting up recording units, collecting and annotating data, and to Alain Jacot for assisting in programming the units and transporting the recorders to Kenya.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

ZibaldonED: Silver annotations for Entity Disambiguation from Digitalzibaldone

<p>The <em>Zibaldoned</em> dataset provides silver annotations for entity disambiguation, extracted from <em>DigitalZibaldone</em>, the HTML digital edition of Giacomo Leopardi&rsquo;s <em>Zibaldone di pensieri</em>, curated by Silvia Stoyanova and Ben Johnston [1].</p> <p>The dataset was collected through a simple web scraping algorithm and includes 2,941 references to people, places, and bibliographical sources across 957 paragraphs. These annotations offer a valuable resource for tasks such as Named Entity Recognition (NER) and Entity Linking (EL) in the context of literary and linguistic research. The dataset can assist in training AI models to automatically recognize and disambiguate entities in literary works.</p> <p>The GitHub repository containing the code used to extract the dataset and train NER and EL models is available here: <a href="https://github.com/sntcristian/zibaldoned" target="_blank" rel="noopener">GitHub - zibaldoned</a>.</p> <p><br>[1] Silvia Stoyanova and Ben Johnston (Eds.), <em>Giacomo Leopardi's Zibaldone di pensieri: a digital research platform</em>.</p>

opencc-by-nc-sa-4.0Oct 2024View details →
zenodo44/100

Annotated Dataset for Uncertainty Mining : Gold Standard

<p>&nbsp;</p> <h1>Description of the dataset</h1> <p>In order to study the expression of uncertainty in scientific articles, we have put together an interdisciplinary corpus of journals in the fields of Science, Technology and Medicine (STM) and the Humanities and Social Sciences (SHS).&nbsp;The selection of journals in our corpus is based on the Scimago Journal and Country Rank (SJR) classification, which is based on Scopus, the largest academic database available online.&nbsp;We have selected journals covering various disciplines, such as medicine, biochemistry, genetics and molecular biology, computer science, social sciences, environmental sciences, psychology, arts and humanities.&nbsp;For each discipline, we selected the five highest-ranked journals. In addition, we have included the journals PLoS ONE and Nature, both of which are interdisciplinary and highly ranked.</p> <p>Based on the corpus of articles from different disciplines described above, we created a set of annotated sentences as follows:</p> <ul> <li>593 were pre-selected automatically, by studying the occurrences of the lists of uncertainty indices proposed by Bongelli et al. (2019), Chen et al. (2018) and Hyland (1996).</li> <li>The remaining sentences were extracted from a subset of articles, consisting of two randomly selected articles per journal. These articles were examined by two human annotators to identify sentences containing uncertainty and to annotate them.</li> <li>600 sentences not expressing scientific uncertainty were manually identified and reviewed by two annotators<br><br></li> </ul> <p>The sentences were annotated by two independent annotators following the annotation guide proposed by Ningrum and Atanassova (2024). The annotators were trained on the basis of an annotation guide and previously annotated sentences in order to guarantee the consistency of the annotations.&nbsp;<br>Each sentence was annotated as expressing or not expressing uncertainty (<strong>Uncertainty</strong> and <strong>No Uncertainty)</strong>.<br>Sentences expressing uncertainty were then annotated along five dimensions: Reference , Nature, Context , Timeline and Expression.&nbsp;<br>The annotators reached an average agreement score of 0.414 according to Cohen's Kappa test, which shows the difficulty of the task of annotating scientific uncertainty.<br>Finally, conflicting annotations were resolved by a third independent annotator.</p> <p><br>Our final corpus thus consists of a total of 1 840 sentences from 496 articles in 21 English-language journals from 8 different disciplines.<br>The columns of the table are as follows:</p> <ol> <li><strong>journal</strong>: name of the journal from where the article originates</li> <li><strong>article_title</strong>: &nbsp;title of the article from where the sentence is extracted</li> <li><strong>publication_year</strong>: year of publication of the article</li> <li><strong>sentence_text</strong>: text of the sentence expressing or not expressing uncertainty</li> <li><strong>uncertainty</strong>: 1 if the sentence expresses uncertainty and 0 otherwise;</li> <li><strong>ref, nature, context, timeline, expression</strong>: annotations of the type of uncertainty according to the annotation framework proposed by Ningrum and Atanassova (2023). The annotation of each dimension in this dataset are in numeric format rather than textual. The mapping betwen textual and numeric labels is presented in the Table below.</li> </ol> <table> <tbody> <tr> <td>Dimension</td> <td>1</td> <td>2</td> <td>3</td> <td>4</td> <td>5</td> </tr> <tr> <td>Reference</td> <td>Author</td> <td>Former</td> <td>Both</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>Nature</td> <td>Epistemic</td> <td>Aleatory</td> <td>Both</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>Context</td> <td>Background</td> <td>Methods</td> <td>Res&amp;Disc</td> <td>Conclusion</td> <td>Others</td> </tr> <tr> <td>Timeline</td> <td>Past</td> <td>Present</td> <td>Future</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>Expression</td> <td>Quantified</td> <td>Unquantified</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <p><br>This gold standard has been produced as part of the <a href="https://project-inscim.github.io/">ANR InSciM (Modelling Uncertainty in Science) project.</a>&nbsp;</p> <h1>References</h1> <p><br>Bongelli, R., Riccioni, I., Burro, R., &amp; Zuczkowski, A. (2019). Writers&rsquo; uncertainty in scientific and popular&nbsp;biomedical articles. A comparative analysis of the British Medical Journal and Discover Magazine&nbsp;[Publisher: Public Library of Science]. PLoS ONE, 14 (9). <a href="https://doi.org/10.1371/journal.pone.0221933">https://doi.org/10.1371/journal.pone.0221933</a></p> <p>Chen, C., Song, M., &amp; Heo, G. E. (2018). A scalable and adaptive method for finding semantically equivalent cue words of uncertainty. Journal of Informetrics, 12 (1), 158&ndash;180. <a href="https://doi.org/10.1016/j.joi.2017.12.004">https://doi.org/10.1016/j.joi.2017.12.004</a></p> <p><br>Hyland, K. E. (1996). Talking to the academy forms of hedging in science research articles [Publisher: SAGE Publications Inc.]. Written Communication, 13 (2), 251&ndash;281. <a href="https://doi.org/10.1177/0741088396013002004">https://doi.org/10.1177/0741088396013002004</a></p> <p>Ningrum, P. K., &amp; Atanassova, I. (2023). Scientific Uncertainty: An Annotation Framework and Corpus Study in Different Disciplines. 19th International Conference of the International Society for Scientometrics and Informetrics (ISSI 2023). <a href="https://doi.org/10.5281/zenodo.8306035">https://doi.org/10.5281/zenodo.8306035</a></p> <p>Ningrum, P. K., &amp; Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. Scientometrics. <a href="https://doi.org/10.1007/s11192-024-05009-z">https://doi.org/10.1007/s11192-024-05009-z</a></p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Genome and annotations of cotton rat (Sigmodon hispidus)

<p>The chromosome level reference genome of <em>Sigmodon hispidus</em> based on third-generation high fidelity (HiFi) reads, high-throughput chromosome conformation capture (Hi-C), and second-generation sequencing techniques.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

LifeWatch observatory data: phytoplankton annotated image library by FlowCam imaging for the Belgian part of the North Sea.

<p>In the framework of the Lifewatch marine observatory a number of fixed stations in the Belgian Part of the North Sea (BPNS) are sampled for phytoplankton monitoring. Samples are processed using a VS-4 FlowCAM model at 4X magnification, size range imaged is 55-300&micro;m. The identification of the image data is done with the use of a classifier and followed by a manual validation step. These dataset comprises the full annotated image dataset which can be sampled for training of convolutional neural networks.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

AWOFRO : Annotated tweet corpus of mixed Wolof-French for detecting obnoxious messages

<p>These data are tweets of mixed Wolof-French codes annotated by three(3) annotators.&nbsp;<br>They were extracted during the period from 1 January 2021 to 31 May 2023.</p> <p>Content description :</p> <p><strong>Corpora.rar</strong> : The dataset contains 3510 annotated tweets</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments

<p>This data is supplementary to the paper:</p> <blockquote> <p><em>Manika Lamba, You Peng, Sophie Nikolov, and J. Stephen Downie. 2024.&nbsp;<strong>AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments</strong>. In The 2024 ACM/IEEE Joint Conference on Digital&nbsp;Libraries (JCDL &rsquo;24), December 2024, Hong Kong, China. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3677389.3702594</em></p> </blockquote>

opencc-by-nc-sa-4.0Dec 2023View details →
zenodo44/100

Tree Annotation Vocabulary (TAV) - Knowledge Graph and Annotated Dataset

<p>This dataset contains all the files used in developing the Tree-KG, the knowledge graph to capture the tree annotations in the works of Vladimir Nabokov.&nbsp;</p> <p>In the Annotated Dataset folder, 6 spreadsheets in excel (.xlsx) format are provided. They are numbered. Note that annotated data are all in English as the consulted works are the English translations of the literary works of Nabokov.</p> <p>(1) contains the tree annotations from the novels originally written in Russian by Vladimir Nabokov.</p> <p>(2) contains the tree annotations from the novels originally written in English by Vladimir Nabokov.</p> <p>(3) contains the tree annotations from the short stories originally written in Russian and English by Vladimir Nabokov.</p> <p>(4) is the knowledge base (KB) developed to link the annotated trees to Wikidata and DBPedia.</p> <p>(5) is the benchmarking results of some entity recognition tools. It includes the relevant passages from Nabokov's novels that were used in the experiments as well as the prompts used in getting the results.</p> <p>(6) represents the complete bibliographic details of the works of Vladimir Nabokov (https://thenabokovian.org/abbreviations).</p> <p>In the Ontology Versions folder, four ontology (TAV) files in turtle (.ttl) format are provided. They are all numbered and dated to represent their different versions. Some sample SPARQL queries are provided in a .txt file. The KG was developed on Prot&eacute;g&eacute;.&nbsp;</p> <p>(1) contains the essential schema for the TAV vocabulary.</p> <p>(2) contains the schema for TAV vocabulary with links to external vocabularies (Schema.Org; Open Annotation, etc.).&nbsp;</p> <p>(3) contains the Tree-KG in so far it reflects data from three novels (Mary; King, Queen, Knave; Glory).</p> <p>(4) contains the entire Tree-KG based on all the works mentioned in the excel sheets (20 books).</p> <p>(5) contains some sample SPARQL queries (.txt) file.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

scGPT: End-to-End Protocol for Fine-tuned Retina Cell Type Annotation

<h1>Abstract</h1> <p>Single-cell research faces challenges in accurately annotating cell types at high resolution, especially when dealing with large-scale datasets and rare cell populations. To address this, foundation models like scGPT offer flexible, scalable solutions by leveraging transformer-based architectures. This protocol provides a comprehensive guide to fine-tuning scGPT for cell-type classification in single-cell RNA sequencing (scRNA-seq) data. We demonstrate how to fine-tune scGPT on a custom retina dataset, highlighting the model&rsquo;s efficiency in handling complex data and improving annotation accuracy achieving 99.5% F1-score. This protocol automates key steps, including data preprocessing, model fine-tuning, and evaluation. This protocol enables researchers to efficiently deploy scGPT for their own datasets. The provided tools, including a command-line script and Jupyter Notebook, simplify the customization and exploration of the model, proposing an accessible workflow for users with minimal Python and Linux knowledge. The protocol offers an off-the-shell solution of high-precision cell-type annotation using scGPT for researchers with intermediate bioinformatics.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Breast Micro-Calcifications Dataset with Precisely Annotated Sequential Mammograms

<p><strong>Dataset Version 3 Update</strong></p> <p><strong>The ground truth images (.jpg) match the dimensions of the corresponding original images (.dcm), ensuring consistency across the dataset.</strong><br><br></p> <p><strong>Breast Micro-Calcifications Dataset with Precisely Annotated Sequential Mammograms</strong></p> <p><strong>Citing the Dataset</strong></p> <p>The dataset is released under a Creative Commons Attribution license, so please cite the dataset if it is used in your work in any form. Published academic papers should use the academic paper citation for our paper. &nbsp;Personal works, such as projects or blog posts, should provide a URL to this Zenodo page, though a reference to our paper would also be appreciated.</p> <p><em>Academic paper citation</em></p> <p>Loizidou, K., Skouroumouni, G., Pitris, C.&nbsp;<em>et al.</em>&nbsp;Digital subtraction of temporally sequential mammograms for improved detection and classification of microcalcifications.&nbsp;<em>Eur Radiol Exp</em>&nbsp;<strong>5,&nbsp;</strong>40 (2021). https://doi.org/10.1186/s41747-021-00238-w</p> <p><em>Personal use citation</em></p> <p>Include a link to this Zenodo page - 10.5281/zenodo.14859694</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>This research is funded by the European Union&rsquo;s Horizon 2020 research and innovation program under grant agreement No. 739551 (KIOS CoE) and from the Republic of Cyprus through the Directorate General for European Programs, Coordination and Development.</p> <p><strong>Contact Information</strong></p> <p>If you would like further information about the dataset, or if you experience any issues downloading files, please contact us at cloizi01@ucy.ac.cy.</p> <p><strong>General Information</strong></p> <p>This dataset consists of 100 pairs of mammograms, from two temporally sequential rounds. Specifically, this dataset includes the prior and recent mammograms of CC and MLO view of each patient. This is a complete dataset for the detection and BI-RADS classification of breast micro-calcifications, using digital mammograms. It contains normal (BI-RADS 1), benign (BI-RADS 2), and suspicious (BI-RADS 4-5) cases, and for each mammogram, an image with precise annotation of each individual micro-calcification, by two expert radiologists, is provided. In 32 suspicious cases, the biopsy results are also available.</p> <p><strong>More details are available in the README.txt</strong></p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Manually Annotated Drone Imagery (RGB) Dataset for automatic coastline delineation of Southern Baltic Sea, Poland with polyline annotations (0.1.1)

<p><strong>Overview:</strong></p> <p>The &nbsp;Manually Annotated Drone Imagery Dataset (MADRID) consists of hand annotated high resolution RGB images taken in two different types of coasts in Poland, Miedzyzdroje - cliff coast and in Mrzezyno - dune coast in 2022-2023. All images were converted into a uniform format of 1440x2560 pixels, polyline annotated and set into file structure format suited for semantic segmentation tasks (See "Usage" notes below for more details).</p> <p>The raw images of our dataset were captured Zenmuse L1 Sensor (RGB) mounted on a DJI Matrice 300 RTK Drone. Total of 4895 images were captured, however the dataset contains 3876 images with each image annotated with coastline. The dataset only include images with coastlines that are visually identifiable with the human eye. For the annotations of the images, CVAT v2.13 open-source software was utilized.</p> <p><strong>Usage:</strong></p> <p>The compressed RAR file contains two folders train and test. Each folder contains the file that represents the date at which the image was captured in the format of (year, month, day), number of the image and the name of the drone utilized to capture the image. For example, DJI_20220111140051_0051_Zenmuse-L1-mission and DJI_20220111140105_0053_Zenmuse-L1-mission. Additionally, the test folder contains annotations (one per image) which are extracted from the original XML annotation file provided in the CVAT 1.1 image format.</p> <p>Archives were compressed using RAR compression. They can be decompressed in a terminal by opening and extracting Madrid_v0.1_Data.zip.</p> <p>The subset of the data with the name Madrid_subset_data.zip has been added which contains a small portion of train and test images for purpose of inspecting the dataset without downloading the entire dataset.</p> <p>The training images for both training data and testing data are structured as follows.</p> <pre><code>Train/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...<br><br>Test/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...</code></pre> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

An extended and improved CCFv3 annotation and Nissl atlas of the entire mouse brain

<p>This archive contains the dataset produced by the Blue Brain Project (BBP) for improving the Common Coordinate Framework version 3 (CCFv3) mouse brain atlas from the Allen Institute for Brain Science (AIBS). The dataset nomenclature is aligned with AIBS standards, utilizing the Allen Reference Atlas (ARA) Nissl-stained volume sectioned in the coronal incidence (ARA NisslCOR) and the annotation file version 3 (ANNOTv3). Additional data were used, such as the AIBS Nissl-stained volume sectioned in the sagittal incidence (AIBS NisslSAG, Allen Mouse Brain Atlas ID 100042147) as well as a Waxholm (WAXH) Nissl-stained volume sectioned in the horizontal incidence (WAXH NisslHOR; https://www.nitrc.org/projects/incfwhsmouse).</p> <p>Here is a list of the files produced and shared below with their descriptions:</p> <p><strong>ara_bbp_nisslCOR_25 - 10</strong>: ARA Nissl-stained volume sectioned in the coronal incidence at 25 and 10 &mu;m isotropic resolution accurately aligned in the CCFv3.</p> <p><strong>arav3a_bbp_nisslCOR_25 - 10</strong>: ARA Nissl-stained volume sectioned in the coronal incidence at 25 and 10 &mu;m isotropic resolution accurately aligned in the CCFv3 and extended for covering the entire brain.</p> <p><strong>annotv3am_bbp_manual_25</strong>: Expert manual delineation of the extended tissue based on ARAv3aBBP NisslCOR at 25 isotropic resolution as well as of the granular and molecular layers in the cerebellum.</p> <p><strong>annotv3a_bbp_25 - 10</strong>: Extended CCFv3 annotation at 25 and 10 &mu;m isotropic resolution covering the entire mouse brain plus including new granular and molecular layers in all cerebellar lobules, assessed using ANNOTv3am.</p> <p><strong>annotv3c_bbp</strong>: Extended CCFv3aBBP annotation covering the mouse central nervous system including spinal cord as well as barrel columns in the isocortex.</p> <p><strong>aibs_bbp_nisslSAG_25</strong>: AIBS sagittal Nissl-stained volume aligned in the CCFv3aBBP at 25 &mu;m isotropic resolution.</p> <p><strong>waxh_bbp_nisslHOR_25</strong>: WAXH horizontal Nissl-stained volume aligned in the CCFv3aBBP at 25 &mu;m isotropic resolution.</p> <p><strong>annotation_bbp_atlas_pipeline_25</strong>: Annotation file from the Blue Brain cell atlas pipeline including the extended version annotv3a_bbp_25, plus some additional sublayers such as layer 2 and layer 3, as well as the barrel columns in the isocortex.</p> <p><strong>hierarchy_bbp_atlas_pipeline</strong>: Hierarchy file attached to the annotation_bbp_atlas_pipeline_25 version.</p> <p><strong>average_nissl_init_25_v3a_CBcorrected</strong>: Average Nissl-stained template composed of the average of arav3a_bbp_nisslCOR_25, aibs_bbp_nisslSAG_25, and waxh_bbp_nisslHOR_25 and including some automated corrections of the artifacts in the cerebellum. This was used as a reference and initialization for building the average Nissl-stained template.</p> <p><strong>average_nissl_template</strong>: Symmetric (symmetric_full) and non symmetric (nissl_average_full) &nbsp;averaged Nissl-stained template in the CCFv3aBBP, as well as the number of occurences per voxel (frequency) in the averaging process.</p> <p><strong>QuickNII-CCFv3a-extended</strong>: Extended CCFv3a atlas file compatible with QuickNII software.</p> <p><strong>VisuAlign-v0.91</strong>: Extended CCFv3a atlas file compatible with VisuAlign software.</p> <p>An additional video (<strong>FullBrainAtlas_bbp</strong>) is provided in that archive, presenting the different mouse brain annotations from AIBS to BBP ones, as well as the BBP computed neuron distribution among the entire mouse brain colored by regions given AIBS standards. The code for creating the data in the video is accessible at https://github.com/favreau/BioExplorer/tree/master/bioexplorer%2Fpythonsdk%2Fnotebooks%2Fccfv3.</p> <p>For accessing to the code related to that work, please go to the corresponding GitHub repository: https://github.com/BlueBrain/ccfv3a-extended-atlas.</p> <p>--</p> <p>Citation:</p> <p>Piluso, S., Veraszt&oacute;, C., Carey, H., Delattre, &Eacute;., L&rsquo;Yvonnet, T., Colnot, &Eacute;., Romani, A., Bjaalie, J. G., &amp; Keller, D. (2024). An extended and improved CCFv3 annotation and Nissl atlas of the entire mouse brain. Zenodo. <a href="https://doi.org/10.5281/zenodo.13640418" target="_blank" rel="noopener noreferrer">https://doi.org/10.5281/zenodo.13640418</a></p> <p>--</p> <p>Reference paper:</p> <p>S&eacute;bastien Piluso,&nbsp;Csaba Veraszt&oacute;,&nbsp;Harry Carey,&nbsp;&Eacute;milie Delattre,&nbsp;Thibaud L&rsquo;Yvonnet,&nbsp;&Eacute;lo&iuml;se Colnot,&nbsp;Armando Romani,&nbsp;Jan G. Bjaalie,&nbsp;Henry Markram,&nbsp;Daniel Keller; An extended and improved CCFv3 annotation and Nissl atlas of the entire mouse brain.&nbsp;<em>Imaging Neuroscience</em>&nbsp;2025; doi:&nbsp;<a href="https://doi.org/10.1162/imag_a_00565" target="_blank" rel="noopener">https://doi.org/10.1162/imag_a_00565</a></p> <p>--</p> <p>Funding:</p> <p><em>This study was supported by funding to the Blue Brain Project, a research center of the &Eacute;cole polytechnique f&eacute;d&eacute;rale de Lausanne (EPFL), from the Swiss government&rsquo;s ETH Board of the Swiss Federal Institutes of Technology. This project/research has received funding from the European Union&rsquo;s Research and Innovation Program Horizon Europe under Grant Agreement no. 101147319 (EBRAINS 2.0).</em></p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish

<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets).&nbsp;</p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the&nbsp;documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud,&nbsp;the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others.&nbsp;</p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then,&nbsp;for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L &ndash; Scientific Literature:&nbsp;</strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials&nbsp;</strong></strong>contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;</li> <li><strong>[Subtrack 3 corpus] MESINESP-P &ndash; Patents:&nbsp;</strong>This corpus&nbsp;includes patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants&#39; predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p>&nbsp;</p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized&nbsp;(folder <em>separated</em>).&nbsp;This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>.&nbsp;</p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train:&nbsp; <ul> <li>training_set_track1_all.json: Full training set for subtrack 1.&nbsp;</li> <li>training_set_track1_only_articles.json:&nbsp;Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json:&nbsp;</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1.&nbsp;</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json:&nbsp;Manually annotated&nbsp;development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json:&nbsp;Manually annotated&nbsp;development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip&nbsp;</strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020&nbsp;set)</li> <li>List of synonyms (the descriptors and synonyms from&nbsp; Latin Spanish DeCS 2020&nbsp;set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo&nbsp;</strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p>&nbsp;</p> <p><strong>Data format&nbsp;description</strong></p> <p>The&nbsp;<strong>input text files</strong>&nbsp;for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial &gt;140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial &gt;130/80 y &lt;140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP&nbsp;<strong>entity mention files</strong>&nbsp;contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p>&nbsp;</p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L &ndash; Scientific Literature&nbsp;</strong>: &nbsp; <ul> <li><em><strong>Training set:&nbsp;</strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.&nbsp;We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;&nbsp;We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database.&nbsp;We distribute two different datasets: <ul> <li><strong>Articles training set:&nbsp;</strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set:&nbsp;</strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts&nbsp;from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: &nbsp; <ul> <li><strong>Training set:&nbsp;</strong>The training dataset contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a&nbsp;<a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team.&nbsp;</li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set:&nbsp;</strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P &ndash; Patents:&nbsp;</strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set:&nbsp;</strong>We provide a&nbsp;<strong>test set</strong>&nbsp;containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set.&nbsp;We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li>&nbsp;We provide this information to the participants as additional data in the &ldquo;Additional Data&rdquo; folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains&nbsp;entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p>&nbsp; </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2&nbsp;Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p>&nbsp;</p> <p>For further information, please&nbsp;email us at luis.gasco@bsc.es</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

HOME-Alcar: Aligned and Annotated Cartularies

<p>The HOME-Alcar (Aligned and Annotated Cartularies) corpus was produced as part of the European research project HOME History of Medieval Europe (https://www.heritageresearch-hub.eu/project/home/), led under the coordination oflinebreakof Institut de Recherche et d&#39;Histoire des Textes (PI: D. Stutzmann), with the Universitat Politecnica de Valencia (PI: E. Vidal), the National Archives of the Czech Republic in Prague (PI: J. Kreckova), and Teklia SAS (PI: C. Kermorvant)<br> The HOME-Alcar (Aligned and Annotated Cartularies) corpus is a resource created to train Handwritten Text Recognition (HTR) and Named Entity Recognition (NER), and presents a collection of<br> (i) digital images of 17 medieval manuscripts;<br> (ii) scholarly editions thereof;<br> (iii) coordinates linking images and text at line level;<br> (iv) annotations of Named Entities (place and person names).<br> The 17 medieval manuscripts in this corpus are cartularies, i.e. books copying charters and legal acts, produced between the 12th and 14th centuries.</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record