Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Kobalt: Extension Corpus and Annotation Guidelines for Verb Classification and Dependency Adjustments

<p>Kobalt (Zinsmeister et al. 2012) is a task-based corpus of essays written by learners and native speakers of German. This repository contains data that was not included in the original corpus and new layers of annotation to the original and the extended corpus, specifically morphological and syntactic classification of verbs and corrections and changes to dependency parses. Please refer to the annotation guidelines included in this repository for further information.<br> &nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Actinidia chinensis Red5 genome assembly (version 2) and annotation files

<p>We present version 2 of the genome assembly for <em>Actinidia chinensis</em> var. <em>chinensis</em> genotype Red5. The Red5 genome was originally assembled using short read Illumina data (Pilkington et al, 2018; <a href="https://doi.org/10.1186/s12864-018-4656-3">https://doi.org/10.1186/s12864-018-4656-3</a>). In version 2 we employed Pacific BioSciences Sequel Single Molecule Real Time (SMRT) sequencing technology in place of Illumina paired end read sequencing for the main assembly but leveraged that short read data (Pilkington et al, 2018) for post assembly base correction of long read assembly contigs. Additionally the Illumina long insert libraries from Pilkington et al (2018) were used for post assembly scaffolding of contigs. Scaffold assignment to linkage groups leveraged the genetic map described in Pilkington et al (2018) as well as consensus evidence from DNA synteny comparisons to existing whole genome sequences from <em>Actinidia</em>.</p> <p>To meet the file size restrictions some dataset components have been split into multiple parts.</p> <p><strong>Assembly</strong></p> <p>The assembly work flow used the FALCON/FALCON-unzip assembly suite is described in Red5_version_2_genome_assembly.md. The assembly yielded both primary and haplotig contig data sets, the metrics for which are documented in this file. The CDS and predicted peptide fasta and GFF3 gene annotation for the primary and haplotig sets are provided in separate files.</p> <p><strong>File Descriptions</strong></p> <ul> <li>Files named chr1.fasta to chr29.fasta represent the primary assembly linkage group level assembly units</li> <li>Files named haplotig_part_1.fasta to haplotig_part_10.fasta represent the haplotig contig sets split into 10 parts to meet upload file size restrictions</li> <li>Files named primary_assembly.primary.gff3 and haplotig.gff3 contain the gene model annotations for the primary and haplotig assembly datasets respectively</li> <li>primary_assembly.cds.fasta and primary_assembly.pep.fasta contain the CDS and peptide sequences for the annotations on the primary contigs</li> <li>haplotig.cds.fasta and haplotig.pep.fasta contain the CDS and peptide sequences for the annotations on the haplotig contigs</li> <li>haplotigs.placements.tsv and haplotigs.reassignments.tsv describe the placement of haplotigs relative to the primary contigs as derived from purge_haplotigs</li> <li>The file Red5_version_2_genome_assembly.md describes the assembly work flow and code steps used as well as assembly metrics</li> <li>Files&nbsp;HYV3_1.v.R5V2_1.png to&nbsp;HYV3_29.v.R5V2_29.png depict Circos plots of DNA:DNA synteny based on 1coords alignment filter of nucmer alignments using dnadiff</li> </ul> <p>See Red5_version_2_genome_assembly.md for description of assembly methods and assembly metrics.</p> <p><strong>Funding</strong></p> <p>This work was funded by Kiwifruit Royalty Investment Program by The New Zealand Institute for Plant &amp; Food Research Ltd. with support from Zespri, and the CORE grant Endeavour Smart Idea Fund (UOOX1801) from the New Zealand Ministry of Business, Innovation and Employment (MBIE). The funding bodies had no role in the design of the study, the collection, analysis, or interpretation of data or writing this manuscript.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Metagenomes: gene function and family annotations

<p>Functional annotations of genes for all contigs in 1,782 metagenomes.</p> <p>Genes were annotated to three sources: (1) COGs, (2) Pfams, and (3) <em>de novo</em> families from reference sequences. These gene annotations are used to train and run PlasX.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

AUTH-OpenDR Mixed Image Annotated Dataset for Human-centric Perception Tasks

<p>The dataset was generated through a mixed (real and synthetic) image data generation method which utilizes real background images and DL-generated human models. It contains 50000 real images depicting urban scenes, populated by synthetic human models in various positions and poses&nbsp;and&nbsp;&nbsp; is suitable for training/evaluating (a) pose estimation, (b) person detection, (c) identity recognition methods. Annotations for 2D bounding boxes of the depicted humans, their&nbsp; IDs and&nbsp;2D keypoints etc are provided. The 133 3D human models, required by the method, were generated using the Pixel-aligned Implicit Function (PIFu) and full-body images of people from the Clothing Co-Parsing (CCP) dataset. As background images, a subset of the Cityscapes dataset was used. The Cityscapes license prohibits the distribution of any modified versions of itself. Thus, we provide code&nbsp;&nbsp;that can re-generate the exact same dataset, given that the Cityscapes dataset is downloaded by the website of its authors.</p> <p>Code and instructions for re-generating the dataset are provided <a href="https://github.com/opendr-eu/opendr/tree/master/projects/python/simulation/human_dataset_generation">here</a>.</p> <p>The dataset was developed by Aristotle University of Thessaloniki&nbsp; (AUTH) within the H2020 OpenDR Project.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Annotation of the the assembled genome of Fusarium oxysporum f. sp. albedinis strain 133, the causal agent of date palm dieback.

<p>Annotation of&nbsp;the the assembled genome of <em>Fusarium oxysporum f. sp. albedinis</em> strain 133 (Khayi et al., 2020). Gene prediction and annotation were carried out using funnotate pipeline v1.8.1 (Stajich, 2020), which&nbsp;includes masking, ab initio gene-prediction training, using Augustus and Genmark, with the EST dataset&nbsp;reported to the Ganoderma mycocosm repository, gene prediction, and the assignment of functional&nbsp;annotation to protein-coding gene models.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Reddit WSB Annotated Dataset 2021

<p>This is a WIP randomized sample of comments from the Wallstreetbets community on Reddit during the GameStop event during the 2021 rise. The sample contains 5000 observations of which 3000 were annotated and agreed upon by two annotators. The next 600 were annotated by two authors but due to time constraints, only 1 author corrected them. The remaining 1400 have only been annotated by one author and have not been compared.</p> <p>The second file is the annotation ruleset used to annotate the dataset, we briefly summarize the rules here:</p> <p>Annotations are broken into two main categories, support (the comment indicates some level of support for either the company GameStop, the stock price, or the narrative of &#39;us&#39; vs &#39;them&#39;.). Support can be either Y= Yes, N= No, U= Unsure, I= Informative.</p> <p>The second category, &#39;intent&#39; indicates the individual has expressed intentions or interest in the stock, or has already purchased the stock during the event period. Intent can be either Y= Yes, N= No, M= Maybe, U= Unsure, or I= Informative.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

BirdVox-25SD: a dataset of flight calls with species annotations

<pre>BirdVox 25 Species Dataset (BirdVox-25SD) ============= Version 1.0, Jan 2021. Created By ---------- Andrew Farnsworth (1), Benjamin Mark Van Doren (1), Steve Kelling (1), Vincent Lostanlen (2), Justin Salamon (3), Aurora Cramer (4), Juan Pablo Bello (4) (1): Cornell Lab of Ornithology (CLO) (2): Laboratoire des Sciences du Num&eacute;rique de Nantes (LS2N), CNRS (3): Adobe Research (4): New York University https://wp.nyu.edu/birdvox Description ----------- The BirdVox 25 Species Dataset (BirdVox-25SD) contains 26,124 audio clips of avian flight calls, each ranging from about 150 ms to 500 ms in duration. The clips are extracted from the <a href="http://https://doi.org/10.5281/zenodo.4603643">BirdVox-296h</a> dataset using the corresponding annotations. The recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the 2015 migration season (August - November). The dataset can be used, among other things, for the research, development and testing of bioacoustic classification models. For details on the hardware of ROBIN recording units, we refer the reader to [1]. [1] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016. Changes from BirdVox 14-SD ---------------------------- This dataset builds upon the <a href="http://https://doi.org/10.5281/zenodo.3667094">BirdVox 14 Species Dataset (BirdVox-14SD)</a>, adding ~12,000 audio clips and annotations. The annotation taxonomy has been expanded to add a new order, a new family, and 11 new species. Additionally, the audio clips are more accurately aligned to the annotation times. For backwards compatibility with the BirdVox-14SD taxonomy, we include the file `birdvox25sd-to-birdvox14sd-taxonomy-code-map.csv` which maps BirdVox-25SD taxonomy codes to BirdVox-14SD taxonomy codes. Taxonomic Annotations ----------------------- Classification annotations for each flight call are given at three taxonomic levels: order, family, and species. These annotations are condensed into a three-number-code which largely follow &quot;..&quot;. The specific numeric codes are: * Order * 1.\*.\* - Passeriformes * 2.\*.\* - Pelecaniformes * Family * 1.1.\* - American Sparrow * 1.2.\* - Cardinals * 1.3.\* - Thrushes * 1.4.\* - New World warblers * 2.1.\* - Herons * Species * 1.1.1 - American tree sparrow (ATSP) * 1.1.2 - Chipping sparrow (CHSP) * 1.1.3 - Savannah sparrow (SAVS) * 1.1.4 - White-throated sparrow (WTSP) * 1.1.5 - Song sparrow (SOSP) * 1.2.1 - Rose-breasted grosbeak (RBGR) * 1.3.1 - Gray-cheeked thrush (GCTH) * 1.3.2 - Swainson&#39;s thrush (SWTH) * 1.3.3 - Hermit thrush (HETH) * 1.3.4 - Veery (VEER) * 1.3.5 - Wood thrush (WOTH) * 1.4.1 - American redstart (AMRE) * 1.4.2 - Bay-breasted warbler (BBWA) * 1.4.3 - Black-throated blue warbler (BTBW) * 1.4.4 - Canada warbler (CAWA) * 1.4.5 - Common yellowthroat (COYE) * 1.4.6 - Mourning warbler (MOWA) * 1.4.7 - Ovenbird (OVEN) * 1.4.8 - Black-and-white warbler (BAWW) * 1.4.9 - Cape May warbler (CMWA) * 1.4.10 - Chestnut-sided warbler (CSWA) * 1.4.11 - Northern Parula (NOPA) * 1.4.12 - Wilson&#39;s warbler (WIWA) * 1.4.13 - Yellow-rumped warbler (YRWA) * 2.1.1 - Green heron (GRHE) Additionally, at any level of the taxonomy, the numeric code &quot;0&quot; is reserved for &quot;other&quot; and the code &quot;X&quot; refers to unknown. For example, 1.1.0 corresponds to an American Sparrow with a species outside of our scope of interest, and 1.1.X corresponds to an American Sparrow of unknown species. At the top level (family), the &quot;other&quot; codes (0.\*.\*) deviate from the family-order-species in order to capture a variety of other out-of-scope sounds, including anthropophony, non-avian biophony, and biophony of avians outside of the scope of interest. Please refer to `<a href="https://zenodo.org/record/5856260/files/BirdVox-296h_taxonomy.yaml">BirdVox-296h_taxonomy.yaml</a>` in <a href="http://https://doi.org/10.5281/zenodo.5856260">BirdVox-296h</a> for the details of this taxonomy structure. Data Files ------------ BirdVox-25SD contains the recordings as HDF5 files, sampled at 22,050 Hz, with a single channel (mono). Each HDF5 file contains flight call vocalizations of a particular species. The name of each HDF5 file follows the format: `BirdVox-25SD-v1pt0_{taxonomy_code}_original.h5`. The name of the HDF5 dataset in each file is &quot;waveforms&quot;, with the corresponding key for each audio recording following the format: `unit-{unit_num}`. Conditions of Use ---------------------- Dataset created by Andrew Farnsworth, Steve Kelling, Vincent Lostanlen, Justin Salamon, Aurora Cramer, and Juan Pablo Bello. The BirdVox-25SD dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International License. The dataset and its contents are made available on an &quot;as is&quot; basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, CLO is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-25SD dataset or any part of it. Feedback ----------- Please help us improve BirdVox-25SD by sending your feedback to: vincent.lostanlen@gmail.com and auroracramer@nyu.edu In case of a problem, please include as many details as possible. Acknowledgements ------------------------ Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes. We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy. The creation of this dataset was supported by NSF grants 1633259 (BIRDVOX).</pre>

opencc-by-4.0Jan 2022View details →
zenodo44/100

RookID: an annotated dataset of vocalisations produced by individually-identified rooks housed together in an outdoors aviary in France

<p>A dataset of annotated recordings of a captive colony of rooks, recorded in Strasbourg, France in&nbsp;2020 and 2021.&nbsp;Each rook was individually identifiable with leg rings.&nbsp;All recordings were taken in the morning&nbsp;a few hours after sunrise, when the birds were most vocally active.&nbsp;The colony was housed outdoors, so other noises are present, including both biotic (most notably various birds, human&nbsp;voices, and other animals)&nbsp;and abiotic (mostly car and train noises).</p> <p>Audio files (.wav): recorded at 48 kHz, 16-bit using 1 to 3 Song Meter 4 recorders (Wildlife Acoustics). Each recorder had two microphone with different gains to maximise dynamic range. The files were then manually synchronised and merged into multichannel (2 to 6) files.</p> <p>Label files (.tsv): Labels corresponding to each recording (each pair has the same name),&nbsp;noting the time stamps and individual emitter&nbsp;for each vocalisation. A single observer annotated all the recordings. Only rook vocalisations from the captive colony were annotated, not other bird vocalisations or the various noises in the data.&nbsp;The annotations consist of tables with 5 columns:&nbsp;</p> <ul> <li>Source: the individual producing the vocalisation. Note that only the bird&#39;s name is indicated. &quot;Inc&quot; and &quot;Pls&quot; are special cases: the first was&nbsp;for when identity could not be determined, the second when multiple individuals vocalised at once in such a manner that individuals could not be separated</li> <li>Start: starting time point for the vocalisation, in seconds (determined as the earliest point when the vocalisation was heard on any channel)</li> <li>End: ending time point for the vocalisation, in seconds (determined as the last point when the vocalisation was head on any channel)</li> <li>Event: gives information for the bird&#39;s activity at the time of the vocalisation, but largely in abbreviated form.&nbsp;One particular case is &quot;sing&quot;, which correspond to vocalisations part of a song bout (which are defined as sequences of different vocalisations separated by less than approximately 10 seconds).</li> <li>Comment: other observations regarding the vocalisation. These are usually not standardised compared to the Event column. One special case is for &quot;Pls&quot;: the Comment column then bears information regarding the identity of the individuals involved.</li> </ul> <p>&nbsp;</p> <p>This dataset was used in our article &quot;Acoustic detection and identification of individual rooks in field recordings using multi-task neural networks&quot;, to train neural networks to identify individual rooks. The dataset was therefore randomly&nbsp;split into train-validation-test datasets.&nbsp;For reproducibility, we provide the &quot;splitting.csv&quot; which contains the information pertaining to which files go in each dataset, and two scripts to do the split automatically.</p> <p>To do so: download and unpack the RookID folder somewhere on your computer, then download splitting.csv and either of the scripts to the same location. Both scripts will MOVE, not copy, the files to new folders corresponding to each dataset.</p> <ul> <li>with split_data.R: open the scrip in an RStudio environment, edit the out_path variable to the desired location, and run the script</li> <li>with split_data.py: run the following command line: python /path/to/split_data.py --out_path path/to/desired/location (note that the script will automatically create the necessary tree structure)</li> <li>Both scripts can be run without editing the out_path variables, in which case the new folders will be created at the same location</li> </ul> <p>&nbsp;</p> <p>For further information, see our code at&nbsp;<a href="https://gitlab.com/kimartin/rook-vocalisation-detection">https://gitlab.com/kimartin/rook-vocalisation-detection</a></p> <p>For any inquiries, please contact Killian Martin (<a href="mailto:killian.martin@ens-lyon.fr?subject=Inquiry%20about%20the%20RookID%20dataset">killian.martin@ens-lyon.fr</a>)</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Gila monster (Heloderma suspectum) genome assembly and annotation

<p><em>De novo</em>&nbsp;genome assembly and annotation of a male Gila monster (<em>Heloderma suspectum</em>). We annotated the genome&nbsp;using the Comparative Annotation Toolkit (CAT), and we have also included GFF3 files of the consensus gene set,&nbsp;output for each taxon included in this process.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

ENHG Annotation of the "Schwazer Berglehenbuch" (TLA Hs. 1587, approx. 1515)

<p>The dataset contains the TEI tags of the historical mining document&nbsp;&quot;Schwazer Berglehenbuch&quot; (Hs. 1587, approx. 1515) which is currently stored by the Tyrolean Regional Archive&nbsp;(Innsbruck, Austria).</p> <p>The following entities were annotated: person, place, mine, date.</p> <p>The citeable Transcript is online available (DOI: 10.5281/zenodo.6274928) as well as the related annotation guidelines (DOI: 10.5281/zenodo.6275197).</p> <p>The data was generated by the research team of the project &ldquo;Text Mining Medieval Mining Texts&rdquo; (T.M.M.M.T.). The research project (2019-2022) was carried out at the University of Innsbruck and funded by go!digital next generation programme of the Austrian Academy of Sciences.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

ENHG Annotation of "Verleihbuch der Rattenberger Bergrichter" (TLA Hs. 37, 1460-1463)

<p>The dataset contains the TEI tags of the historical mining document&nbsp;&quot;Verleihbuch der Rattenberger Bergrichter&quot; (TLA Hs. 37, 1460-1463) that is currently stored in the Tyrolean Regional archives (Innsbruck).</p> <p>The citeable Transcript is online available (DOI: 10.10.5281/zenodo.6274928) as well as the related annotation guidelines (DOI: 10.5281/zenodo.6275197)</p> <p>The data was generated by the research team of the project &ldquo;Text Mining Medieval Mining Texts&rdquo; (T.M.M.M.T.). The research project (2019-2022) was carried out at the University of Innsbruck and funded by go!digital next generation programme of the Austrian Academy of Sciences.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Structure Annotations of Assessment and Plan Sections from MIMIC-III

<p>Physicians record their detailed thought-processes about diagnoses and treatments as unstructured text in a section of a clinical note called the &quot;assessment and plan&quot;. This information is more clinically rich than structured billing codes assigned for an encounter but harder to reliably extract given the complexity of clinical language and documentation habits. To structure these sections we collected a dataset of annotations over assessment and plan sections from the publicly available and de-identified MIMIC-III dataset, and developed deep-learning based models to perform this task, described in the associated paper available as a pre-print at:&nbsp;<a href="https://www.medrxiv.org/content/10.1101/2022.04.13.22273438v1">https://www.medrxiv.org/content/10.1101/2022.04.13.22273438v1</a></p> <p>When using this data please cite our paper:</p> <pre><code>@article {Stupp2022.04.13.22273438, author = {Stupp, Doron and Barequet, Ronnie and Lee, I-Ching and Oren, Eyal and Feder, Amir and Benjamini, Ayelet and Hassidim, Avinatan and Matias, Yossi and Ofek, Eran and Rajkomar, Alvin}, title = {Structured Understanding of Assessment and Plans in Clinical Documentation}, year = {2022}, doi = {10.1101/2022.04.13.22273438}, publisher = {Cold Spring Harbor Laboratory Press}, URL = {https://www.medrxiv.org/content/early/2022/04/17/2022.04.13.22273438}, journal = {medRxiv} }</code></pre> <p>The dataset,&nbsp;presented&nbsp;here, contains annotations of assessment and plan sections of notes from the publicly available and de-identified MIMIC-III dataset, marking the active problems, their assessment description, and plan action items.&nbsp;Action items&nbsp;are additionally marked as one of 8 categories (listed below). The dataset contains over 30,000 annotations of 579 notes from distinct patients, annotated by 6 medical residents and students.&nbsp;</p> <p>The dataset is divided into 4 partitions&nbsp;-&nbsp; a training set (481 notes), validation set (50 notes), test set (48 notes) and an inter-rater set. The inter-rater set contains the annotations of each of the raters over the test set. Rater 1 in the inter-rater set should be regarded as an intra-rater comparison&nbsp;(details in the paper).&nbsp;The labels underwent automatic normalization to capture entire word boundaries and remove flanking non-alphanumeric characters.</p> <p>Code for transforming labels into TensorFlow examples and training models as described in the paper will be made available at GitHub:&nbsp;<a href="https://github.com/google-research/google-research/tree/master/assessment_plan_modeling">https://github.com/google-research/google-research/tree/master/assessment_plan_modeling</a></p> <p>In order to use these annotations, the user additionally needs to obtain the text of the notes which is&nbsp;found in the NOTE_EVENTS table from MIMIC-III, access to which is to be acquired independently (<a href="http://mimic.mit.edu">https://mimic.mit.edu/</a>)</p> <p>Annotations are given as character spans in a CSV file with the following schema:</p> <table> <tbody> <tr> <td>Field</td> <td>Type</td> <td>Semantics</td> </tr> <tr> <td>partition</td> <td>categorical (one of [train, val, test, interrater]</td> <td>The set of ratings the span belongs to.</td> </tr> <tr> <td>rater_id</td> <td>int</td> <td>Unique id for each the raters</td> </tr> <tr> <td>note_id</td> <td>int</td> <td>The note&rsquo;s unique note_id, links to the MIMIC-III notes table (as ROW-ID).</td> </tr> <tr> <td>span_type</td> <td>categorical (one of [PROBLEM_TITLE,<br> PROBLEM_DESCRIPTION, ACTION_ITEM]</td> <td>Type of the span as annotated by raters.</td> </tr> <tr> <td>char_start</td> <td>int</td> <td>Character offsets from note start</td> </tr> <tr> <td>char_end</td> <td>int</td> </tr> <tr> <td>action_item_type</td> <td>categorical (one of [MEDICATIONS, IMAGING, OBSERVATIONS_LABS, CONSULTS, NUTRITION, THERAPEUTIC_PROCEDURES, OTHER_DIAGNOSTIC_PROCEDURES, OTHER])</td> <td>Type of action item if the span is an action item (empty otherwise) as annotated by raters.</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2022View details →
zenodo44/100

MediCause Dataset of Causal Sentences with Annotated Entities

<p>The MediCause dataset contains 1202 causal sentences from medical publications where the entities involved in the causal relations have been annotated according to the MediCause ontological model for causal relations. The entities are annotated according to the Inside-Outside-Beginning (IOB) format. The labels used for the annotation are B-C (Cause), B-VC (Causal Variable), B-CS (Beginning Causal Specifier), I-CS (Inside Causal Specifier), B-CON (Beginning Connective), I-CON (Inside Connective), B-EF (Effect), B-VE (Effect Variable), B-ES (Beginning Effect Specifier), I-ES (nside Effect Specifier), O (Outside).</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Towards a systematic approach to manual annotation of code smells - C# Dataset of Long Method and Large Class code smells

<p>This dataset includes open-source projects written in C# programing language, annotated for the presence of Long Method and God Class code smells. Each instance was manually annotated by at least two annotators.&nbsp;We explain our motivation and methodology for creating this dataset in our <a href="https://www.techrxiv.org/articles/preprint/Towards_a_systematic_approach_to_manual_annotation_of_code_smells/14159183/1">preprint</a>:</p> <p>Luburić, N., Prokić, S., Grujić, K.G., Slivka, J., Kovačević, A., Sladić, G. and Vidaković, D., 2021. Towards a systematic approach to manual annotation of code smells.&nbsp;</p> <p>The dataset contains two excel datasheets:</p> <ul> <li><em>DataSet_Large Class.xlsx</em> &ndash; C# classes annotated for the Large Class code smell severity.</li> <li><em>DataSet_Long Method.xlsx</em> &ndash; C# methods annotated for the Long method code smell severity.</li> </ul> <p>&nbsp;The columns in the datasheet represent:</p> <ul> <li><em>Code Snippet ID</em> &ndash; the full name of the code snippet.&nbsp; <ul> <li>For classes, this is the package/namespace name followed by the class name. The full name of inner classes also contains the names of any outer classes (e.g., <em>namespace.subnamespace.outerclass.innerclass</em>).</li> <li>For methods, this is the full name of the class and the methods&rsquo;s signature (e.g., <em>namespace.class.method(param1Type, param2Type)</em> ).</li> </ul> </li> <li><em>Link </em>&ndash; The GitHub link to the code snippet, including the commit and the start and end LOC.</li> <li><em>Code Smell </em>&ndash; code smell for which the code snippet is examined (Large Class or Long Method).</li> <li><em>Project Link </em>&ndash; the link to the version of the code repository that was annotated.</li> <li><em>Metrics </em>&ndash; a list of metrics for the code snippet, calculated by our <a href="https://github.com/Clean-CaDET/platform#readme">platform</a>. Our dataset provides 25 class-level metrics for Large Class detection and 18 method-level metrics for Long Method detection The list of metrics and their definitions is available <a href="https://github.com/Clean-CaDET/platform/blob/c4acff95ec00ff6c25fa62dde4818c1f40e39d39/CodeModel/CaDETModel/CodeItems/CaDETMetrics.cs">here</a>.</li> <li><em>Final annotation </em>&ndash; a single severity score calculated by a majority vote.&nbsp;</li> <li><em>Annotators </em>&ndash; each annotator&#39;s (1, 2, or 3) assigned severity score.</li> </ul> <p>To help guide their reasoning for evaluating the presence and the severity of a code smell, three annotators independently annotated whether the considered heuristics apply to an evaluated code snippet. We provide these results in two separate excel datasheets:</p> <ul> <li><em>LargeClass_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> <li><em>LongMethod_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> </ul> <p>The columns of these two datasheets are:</p> <ul> <li><em>Code Snippet ID </em>- the full name of the code snippet (matching the IDs from <em>DataSet_Large Class.xlsx </em>and <em>DataSet_Long Method.xlsx</em>)</li> <li><em>Annotators</em> &ndash; heuristics labelled by each of the annotators (1, 2, or 3).</li> <li><em>Heuristics </em>&ndash; whether the heuristic is applicable to the examined code snippet or not (Section 1.2.4 lists heuristics relevant for the Large Class detection, and Section 1.2.5 lists the heuristics relevant for the Long Method detection).</li> </ul>

opencc-by-4.0May 2022View details →
zenodo44/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Annotation V2.0)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is a&nbsp;new version of&nbsp;the Annotation of MELA dataset, in which we add a missing annotation for &#39;mela_0732&#39;. This file&nbsp;includes&nbsp;the annotations&nbsp;of the whole training set and validation set.&nbsp;</p> <p>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `public_id: anonymous patient ID to match images and annotations.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

The first annotated genome assembly of Macrophomina tecta associated with charcoal rot of sorghum

<p>Raw reads of Macrophomina tecta were obtained from Nanopore, Illumina,&nbsp;and NextSeq (RNA). Files with information about the genome annotation, functional prediction, repeats, effectors and orthologous genes are included.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Metadata and annotation data for the XSample corpus on German academic language

<p>The XSample corpus war created in the project <em>XSample</em> (https://www.izus.uni-stuttgart.de/fokus/fdm-projekte/xsample/) at Universit&auml;t Stuttgart in 2021 by Melanie Andresen and Axel Pichler. It contains 135 German academic journal articles, 45 each from the disciplines linguistics, literary studies and philosophy. The texts themselves cannot be made public for copyright reasons. However, metadata and some annotation data are published here.</p> <p><strong>xsample-metadata.csv</strong><br> This file contains metadata on the texts in the corpus, like journal, title, authors, text length, and the URL to the original paper. It also contains two analytical metrics, &#39;past-ratio&#39; and &#39;temp-expr-ratio&#39;, that are based on the annotations in the other two files. The variable &#39;past-ratio&#39; expresses the proportion of verbs in past tense relative to all finite verbs in the text. The variable &#39;temp-expr-ratio&#39; gives the number of temporal expressions per 1000 token.</p> <p><strong>xsample-heidel.csv</strong><br> This file contains all temporal expressions found and classified by the annotation tool <em>HeidelTime</em> (https://github.com/HeidelTime/heideltime, V. 2.2.1, Str&ouml;tgen &amp; Gertz 2013 ). The variable &#39;position&#39; expresses the position of the first character of the temporal expression in the text in characters.</p> <p><strong>xsample-sticker2.csv</strong><br> This file contains all finite verbs found and classified by the annotation tool <em>sticker2</em> (https://github.com/stickeritis/sticker2). The variable &#39;position&#39; expresses the position of the first character of the finite verb in the text in characters.</p> <p><strong>References</strong><br> Str&ouml;tgen, Jannik &amp; Michael Gertz. 2013. Multilingual and cross-domain temporal tagging. <em>Language Resources and Evaluation</em>. Springer 47(2). 269&ndash;298. <a href="https://doi.org/10.1007/s10579-012-9179-y">https://doi.org/10.1007/s10579-012-9179-y</a>.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Drosophila willistoni genome annotation

<p>Genome annotation of Drosophila willistoni de novo assembly using long reads. Protein coding genes were predicted with Funannotate. This software aligned proteins and transcripts from the Drosophila willistoni Flybase annotation against the assembly using minimap2, diamond and exonerate. It processed these hints to be included by augustus when predicting protein coding genes.</p> <p>LncRNAs were identified by FEELnc using stranded ribodepleted RNA-seq libraries from ovaries, testes, male accessory glands and whole body samples. Additional lncRNAs were identified by mapping lncRNAs sequences downloaded from RNA central identified for Drosophila willistoni.</p> <p>Ribosomal RNAs were predicted with RNAmmer; tRNAs, with tRNAscan2; and other miscellaneous ncRNA, with cmsearch and Rfam models.</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

"Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan" - Genomes and annotation

<p>File-dump associated with the manuscript:</p> <p>&quot;<strong>Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan&quot; (Not yet published)</strong></p> <p>MGVs refer to the viral genome database in the publication:&nbsp;https://www.nature.com/articles/s41564-021-00928-6&nbsp;</p> <p>&nbsp;</p> <p>Following uploaded:</p> <p>File 1: VOG Markers in vOTUs/vMAGs and MGV genomes</p> <p>File 2: Viral Tree Newick&nbsp;file with vOTUs/vMAGs and MGV genomes</p> <p>File 3: All vOTUs/vMAGs genomes</p> <p>File 4: Master table annotation of vOTUs/vMAGs</p> <p>File 5: Centenarian bacterial isolate proviruses</p>

opencc-by-4.0May 2022View details →
zenodo44/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set&nbsp;and Annotation of MELA dataset, including 110&nbsp;CTs and the annotations&nbsp;of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `public_id: anonymous patient ID to match images and annotations.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record