Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
271
datasets available to search
ShareScore release 0.7.1
Dataset results
271 results for “annotated dataset”
Reddit WSB Annotated Dataset 2021
<p>This is a WIP randomized sample of comments from the Wallstreetbets community on Reddit during the GameStop event during the 2021 rise. The sample contains 5000 observations of which 3000 were annotated and agreed upon by two annotators. The next 600 were annotated by two authors but due to time constraints, only 1 author corrected them. The remaining 1400 have only been annotated by one author and have not been compared.</p> <p>The second file is the annotation ruleset used to annotate the dataset, we briefly summarize the rules here:</p> <p>Annotations are broken into two main categories, support (the comment indicates some level of support for either the company GameStop, the stock price, or the narrative of 'us' vs 'them'.). Support can be either Y= Yes, N= No, U= Unsure, I= Informative.</p> <p>The second category, 'intent' indicates the individual has expressed intentions or interest in the stock, or has already purchased the stock during the event period. Intent can be either Y= Yes, N= No, M= Maybe, U= Unsure, or I= Informative.</p>
BirdVox-25SD: a dataset of flight calls with species annotations
<pre>BirdVox 25 Species Dataset (BirdVox-25SD) ============= Version 1.0, Jan 2021. Created By ---------- Andrew Farnsworth (1), Benjamin Mark Van Doren (1), Steve Kelling (1), Vincent Lostanlen (2), Justin Salamon (3), Aurora Cramer (4), Juan Pablo Bello (4) (1): Cornell Lab of Ornithology (CLO) (2): Laboratoire des Sciences du Numérique de Nantes (LS2N), CNRS (3): Adobe Research (4): New York University https://wp.nyu.edu/birdvox Description ----------- The BirdVox 25 Species Dataset (BirdVox-25SD) contains 26,124 audio clips of avian flight calls, each ranging from about 150 ms to 500 ms in duration. The clips are extracted from the <a href="http://https://doi.org/10.5281/zenodo.4603643">BirdVox-296h</a> dataset using the corresponding annotations. The recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the 2015 migration season (August - November). The dataset can be used, among other things, for the research, development and testing of bioacoustic classification models. For details on the hardware of ROBIN recording units, we refer the reader to [1]. [1] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016. Changes from BirdVox 14-SD ---------------------------- This dataset builds upon the <a href="http://https://doi.org/10.5281/zenodo.3667094">BirdVox 14 Species Dataset (BirdVox-14SD)</a>, adding ~12,000 audio clips and annotations. The annotation taxonomy has been expanded to add a new order, a new family, and 11 new species. Additionally, the audio clips are more accurately aligned to the annotation times. For backwards compatibility with the BirdVox-14SD taxonomy, we include the file `birdvox25sd-to-birdvox14sd-taxonomy-code-map.csv` which maps BirdVox-25SD taxonomy codes to BirdVox-14SD taxonomy codes. Taxonomic Annotations ----------------------- Classification annotations for each flight call are given at three taxonomic levels: order, family, and species. These annotations are condensed into a three-number-code which largely follow "..". The specific numeric codes are: * Order * 1.\*.\* - Passeriformes * 2.\*.\* - Pelecaniformes * Family * 1.1.\* - American Sparrow * 1.2.\* - Cardinals * 1.3.\* - Thrushes * 1.4.\* - New World warblers * 2.1.\* - Herons * Species * 1.1.1 - American tree sparrow (ATSP) * 1.1.2 - Chipping sparrow (CHSP) * 1.1.3 - Savannah sparrow (SAVS) * 1.1.4 - White-throated sparrow (WTSP) * 1.1.5 - Song sparrow (SOSP) * 1.2.1 - Rose-breasted grosbeak (RBGR) * 1.3.1 - Gray-cheeked thrush (GCTH) * 1.3.2 - Swainson's thrush (SWTH) * 1.3.3 - Hermit thrush (HETH) * 1.3.4 - Veery (VEER) * 1.3.5 - Wood thrush (WOTH) * 1.4.1 - American redstart (AMRE) * 1.4.2 - Bay-breasted warbler (BBWA) * 1.4.3 - Black-throated blue warbler (BTBW) * 1.4.4 - Canada warbler (CAWA) * 1.4.5 - Common yellowthroat (COYE) * 1.4.6 - Mourning warbler (MOWA) * 1.4.7 - Ovenbird (OVEN) * 1.4.8 - Black-and-white warbler (BAWW) * 1.4.9 - Cape May warbler (CMWA) * 1.4.10 - Chestnut-sided warbler (CSWA) * 1.4.11 - Northern Parula (NOPA) * 1.4.12 - Wilson's warbler (WIWA) * 1.4.13 - Yellow-rumped warbler (YRWA) * 2.1.1 - Green heron (GRHE) Additionally, at any level of the taxonomy, the numeric code "0" is reserved for "other" and the code "X" refers to unknown. For example, 1.1.0 corresponds to an American Sparrow with a species outside of our scope of interest, and 1.1.X corresponds to an American Sparrow of unknown species. At the top level (family), the "other" codes (0.\*.\*) deviate from the family-order-species in order to capture a variety of other out-of-scope sounds, including anthropophony, non-avian biophony, and biophony of avians outside of the scope of interest. Please refer to `<a href="https://zenodo.org/record/5856260/files/BirdVox-296h_taxonomy.yaml">BirdVox-296h_taxonomy.yaml</a>` in <a href="http://https://doi.org/10.5281/zenodo.5856260">BirdVox-296h</a> for the details of this taxonomy structure. Data Files ------------ BirdVox-25SD contains the recordings as HDF5 files, sampled at 22,050 Hz, with a single channel (mono). Each HDF5 file contains flight call vocalizations of a particular species. The name of each HDF5 file follows the format: `BirdVox-25SD-v1pt0_{taxonomy_code}_original.h5`. The name of the HDF5 dataset in each file is "waveforms", with the corresponding key for each audio recording following the format: `unit-{unit_num}`. Conditions of Use ---------------------- Dataset created by Andrew Farnsworth, Steve Kelling, Vincent Lostanlen, Justin Salamon, Aurora Cramer, and Juan Pablo Bello. The BirdVox-25SD dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International License. The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, CLO is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-25SD dataset or any part of it. Feedback ----------- Please help us improve BirdVox-25SD by sending your feedback to: vincent.lostanlen@gmail.com and auroracramer@nyu.edu In case of a problem, please include as many details as possible. Acknowledgements ------------------------ Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes. We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy. The creation of this dataset was supported by NSF grants 1633259 (BIRDVOX).</pre>
RookID: an annotated dataset of vocalisations produced by individually-identified rooks housed together in an outdoors aviary in France
<p>A dataset of annotated recordings of a captive colony of rooks, recorded in Strasbourg, France in 2020 and 2021. Each rook was individually identifiable with leg rings. All recordings were taken in the morning a few hours after sunrise, when the birds were most vocally active. The colony was housed outdoors, so other noises are present, including both biotic (most notably various birds, human voices, and other animals) and abiotic (mostly car and train noises).</p> <p>Audio files (.wav): recorded at 48 kHz, 16-bit using 1 to 3 Song Meter 4 recorders (Wildlife Acoustics). Each recorder had two microphone with different gains to maximise dynamic range. The files were then manually synchronised and merged into multichannel (2 to 6) files.</p> <p>Label files (.tsv): Labels corresponding to each recording (each pair has the same name), noting the time stamps and individual emitter for each vocalisation. A single observer annotated all the recordings. Only rook vocalisations from the captive colony were annotated, not other bird vocalisations or the various noises in the data. The annotations consist of tables with 5 columns: </p> <ul> <li>Source: the individual producing the vocalisation. Note that only the bird's name is indicated. "Inc" and "Pls" are special cases: the first was for when identity could not be determined, the second when multiple individuals vocalised at once in such a manner that individuals could not be separated</li> <li>Start: starting time point for the vocalisation, in seconds (determined as the earliest point when the vocalisation was heard on any channel)</li> <li>End: ending time point for the vocalisation, in seconds (determined as the last point when the vocalisation was head on any channel)</li> <li>Event: gives information for the bird's activity at the time of the vocalisation, but largely in abbreviated form. One particular case is "sing", which correspond to vocalisations part of a song bout (which are defined as sequences of different vocalisations separated by less than approximately 10 seconds).</li> <li>Comment: other observations regarding the vocalisation. These are usually not standardised compared to the Event column. One special case is for "Pls": the Comment column then bears information regarding the identity of the individuals involved.</li> </ul> <p> </p> <p>This dataset was used in our article "Acoustic detection and identification of individual rooks in field recordings using multi-task neural networks", to train neural networks to identify individual rooks. The dataset was therefore randomly split into train-validation-test datasets. For reproducibility, we provide the "splitting.csv" which contains the information pertaining to which files go in each dataset, and two scripts to do the split automatically.</p> <p>To do so: download and unpack the RookID folder somewhere on your computer, then download splitting.csv and either of the scripts to the same location. Both scripts will MOVE, not copy, the files to new folders corresponding to each dataset.</p> <ul> <li>with split_data.R: open the scrip in an RStudio environment, edit the out_path variable to the desired location, and run the script</li> <li>with split_data.py: run the following command line: python /path/to/split_data.py --out_path path/to/desired/location (note that the script will automatically create the necessary tree structure)</li> <li>Both scripts can be run without editing the out_path variables, in which case the new folders will be created at the same location</li> </ul> <p> </p> <p>For further information, see our code at <a href="https://gitlab.com/kimartin/rook-vocalisation-detection">https://gitlab.com/kimartin/rook-vocalisation-detection</a></p> <p>For any inquiries, please contact Killian Martin (<a href="mailto:killian.martin@ens-lyon.fr?subject=Inquiry%20about%20the%20RookID%20dataset">killian.martin@ens-lyon.fr</a>)</p>
MediCause Dataset of Causal Sentences with Annotated Entities
<p>The MediCause dataset contains 1202 causal sentences from medical publications where the entities involved in the causal relations have been annotated according to the MediCause ontological model for causal relations. The entities are annotated according to the Inside-Outside-Beginning (IOB) format. The labels used for the annotation are B-C (Cause), B-VC (Causal Variable), B-CS (Beginning Causal Specifier), I-CS (Inside Causal Specifier), B-CON (Beginning Connective), I-CON (Inside Connective), B-EF (Effect), B-VE (Effect Variable), B-ES (Beginning Effect Specifier), I-ES (nside Effect Specifier), O (Outside).</p>
Towards a systematic approach to manual annotation of code smells - C# Dataset of Long Method and Large Class code smells
<p>This dataset includes open-source projects written in C# programing language, annotated for the presence of Long Method and God Class code smells. Each instance was manually annotated by at least two annotators. We explain our motivation and methodology for creating this dataset in our <a href="https://www.techrxiv.org/articles/preprint/Towards_a_systematic_approach_to_manual_annotation_of_code_smells/14159183/1">preprint</a>:</p> <p>Luburić, N., Prokić, S., Grujić, K.G., Slivka, J., Kovačević, A., Sladić, G. and Vidaković, D., 2021. Towards a systematic approach to manual annotation of code smells. </p> <p>The dataset contains two excel datasheets:</p> <ul> <li><em>DataSet_Large Class.xlsx</em> – C# classes annotated for the Large Class code smell severity.</li> <li><em>DataSet_Long Method.xlsx</em> – C# methods annotated for the Long method code smell severity.</li> </ul> <p> The columns in the datasheet represent:</p> <ul> <li><em>Code Snippet ID</em> – the full name of the code snippet. <ul> <li>For classes, this is the package/namespace name followed by the class name. The full name of inner classes also contains the names of any outer classes (e.g., <em>namespace.subnamespace.outerclass.innerclass</em>).</li> <li>For methods, this is the full name of the class and the methods’s signature (e.g., <em>namespace.class.method(param1Type, param2Type)</em> ).</li> </ul> </li> <li><em>Link </em>– The GitHub link to the code snippet, including the commit and the start and end LOC.</li> <li><em>Code Smell </em>– code smell for which the code snippet is examined (Large Class or Long Method).</li> <li><em>Project Link </em>– the link to the version of the code repository that was annotated.</li> <li><em>Metrics </em>– a list of metrics for the code snippet, calculated by our <a href="https://github.com/Clean-CaDET/platform#readme">platform</a>. Our dataset provides 25 class-level metrics for Large Class detection and 18 method-level metrics for Long Method detection The list of metrics and their definitions is available <a href="https://github.com/Clean-CaDET/platform/blob/c4acff95ec00ff6c25fa62dde4818c1f40e39d39/CodeModel/CaDETModel/CodeItems/CaDETMetrics.cs">here</a>.</li> <li><em>Final annotation </em>– a single severity score calculated by a majority vote. </li> <li><em>Annotators </em>– each annotator's (1, 2, or 3) assigned severity score.</li> </ul> <p>To help guide their reasoning for evaluating the presence and the severity of a code smell, three annotators independently annotated whether the considered heuristics apply to an evaluated code snippet. We provide these results in two separate excel datasheets:</p> <ul> <li><em>LargeClass_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> <li><em>LongMethod_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> </ul> <p>The columns of these two datasheets are:</p> <ul> <li><em>Code Snippet ID </em>- the full name of the code snippet (matching the IDs from <em>DataSet_Large Class.xlsx </em>and <em>DataSet_Long Method.xlsx</em>)</li> <li><em>Annotators</em> – heuristics labelled by each of the annotators (1, 2, or 3).</li> <li><em>Heuristics </em>– whether the heuristic is applicable to the examined code snippet or not (Section 1.2.4 lists heuristics relevant for the Large Class detection, and Section 1.2.5 lists the heuristics relevant for the Long Method detection).</li> </ul>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Annotation V2.0)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is a new version of the Annotation of MELA dataset, in which we add a missing annotation for 'mela_0732'. This file includes the annotations of the whole training set and validation set. </p> <p>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</p> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set and Annotation of MELA dataset, including 110 CTs and the annotations of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
An annotated high-content fluorescence microscopy dataset with Hoechst 33342-stained nuclei and manually labelled outlines
<p>Here we present a benchmarking dataset of fluorescence microscopy images with Hoechst 33342-stained nuclei together with annotations of nuclei, nuclear fragments and micronuclei. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification. Labelling was performed by a single annotator and reviewed by a biomedical expert.</p> <p>The dataset contains 50 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging1. A brief article describing the dataset is also available (Arvidsson M, Kazemi Rashed S, Aits S. <a href="https://doi.org/10.1016/j.dib.2022.108769">10.1016/j.dib.2022.108769</a> )</p> <p><strong>Dataset description:</strong></p> <p>Fluorescence microscopy images: original .C01 files and files converted to 8-bit .png format (Grayscale)</p> <p>Annotations: 24-bit .png format (RGB)</p> <p>Script used to convert C01 to png images: C01_to_png.py file with python code and readme.md file with instructions to run it</p>
HISTORIAN: a large-scale HISTORIcal film dataset with cinematographic ANnotation
<p>Developing automated tools for sustainable film preservation of extensive historical film collections assumes an understanding of fundamental cinematographic settings. In order to be able to investigate new approaches to detect and classify cinematographic settings, this paper proposes a novel large-scale historical film dataset with cinematographic annotations (HISTORIAN), i.e., shot boundaries, shot types, camera movements. The dataset consists of 98 digitized original analog film reels related to the Second World War and 10593 film shots manually annotated by human film experts. Moreover, annotations for overscan areas such as sprocket holes are included. A baseline film analysis pipeline is introduced and evaluated. To the best of our knowledge, HISTORIAN is the first dataset that covers the challenges and characteristics of historical film documentaries and provides novel possibilities for exploring automatic film analysis tools.</p> <p>This repository presents a tiny set including a few examples for demonstration.</p> <p>A link to the Github repository (including helper scripts and readme) can be found <a href="https://github.com/dahe-cvl/historian_dataset">here</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
Joseph Haydn - String Quartets Op.20 - Harmonic Analysis Annotations Dataset
<p>This dataset accompanies the Master Thesis from the same author. It is a manually-annotated corpus of harmonic analysis in **harm syntax.</p> <p>The dataset contains the following scores:<br> Haydn, Joseph<br> 1. E-flat major, op. 20 no. 1, Hob. III-31<br> I. Allegro moderato<br> II. Menuetto. Allegretto<br> III. Affettuoso e sostenuto<br> IV. Finale. Presto<br> 2. C major, op. 20 no. 2, Hob. III-32 <br> I. Moderato<br> II. Capriccio. Adagio<br> III. Menuetto. Allegretto<br> IV. Fuga a 4 soggetti<br> 3. G minor, op. 20 no. 3, Hob. III-33<br> I. Allegro con spirito<br> II. Menuetto. Allegretto<br> III. Poco adagio<br> IV. Finale. Allegro molto<br> 4. D major, op. 20 no. 4, Hob. III-34<br> I. Allegro di molto<br> II. Un poco adagio e affettuoso<br> III. Menuet alla Zingarese & Trio<br> IV. Presto e scherzando<br> 5. F minor, op. 20 no. 5, Hob. III-35<br> I. Allegro moderato<br> II. Menuetto<br> III. Adagio<br> IV. Finale. Fuga a due soggetti<br> 6. A major, op. 20 no. 6, Hob. III-36<br> I. Allegro di molto e scherzando<br> II. Adagio. Cantabile<br> III. Menuetto. Allegretto<br> IV. Fuga a 3 soggetti. Allegro</p>
Breast Masses Dataset with Precisely Annotated Sequential Mammograms
<p><strong>BREAST MASSES DATASET WITH PRECISELY ANNOTATED SEQUENTIAL MAMMOGRAMS</strong></p> <p><strong>Citing the Dataset</strong></p> <p>The dataset is released under a Creative Commons Attribution license, so it is mandatory to cite the dataset if you use it in your work in any form. Published academic papers should use the academic paper citation of our paper. Personal works, such as projects or blog posts, should provide a URL to this Zenodo page, though a reference to our paper would also be appreciated.</p> <p><em>Academic paper citation</em></p> <p>TBA</p> <p><em>Personal use citation</em></p> <p>Include a link to this Zenodo page - 10.5281/zenodo.11446259</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>The publication of this paper is supported by the European Union’s Horizon 2020 research and innovation programme under grant agreement No 739551 (KIOS CoE) and the Government of the Republic of Cyprus through the Cyprus Deputy Ministry of Research, Innovation and Digital Policy.</p> <p><strong>Contact Information</strong></p> <p>If you would like further information about the dataset, or if you experience any issues downloading files, please contact us at cloizi01@ucy.ac.cy.</p> <p><strong>General Information</strong></p> <p>This dataset consists of 100 pairs of mammograms, from two temporally sequential rounds. Specifically, this dataset includes the prior and recent mammograms with two mammographic views for each patient. This is a complete dataset for the detection and classification of breast masses, using sequential mammograms. It contains normal (BI-RADS 1), benign (BI-RADS 2), and biopsy-confirmed malignant cases (BI-RADS 6). For each mammogram, an image with precise annotation of each individual mass, by two expert radiologists, is provided.</p> <p><strong>More details are available in the README.txt</strong></p>
CLDF dataset accompanying Chacon's "Annotated Swadesh Wordlists for Northwest Arawakan Languages" from 2022
<p>Cite the source of the dataset as:</p> <blockquote> <p>Chacon, Thiago C. (2022): Annotated Swadesh wordlists for Northwest Arawakan languages. Leipzig: Max Planck Institute for Evolutionary Anthropology.</p> </blockquote>
CLDF dataset derived from Starostin's "Annotated Swadesh Wordlists for the Karen Group" from 2017
<p>Cite the source of the dataset as:</p> <blockquote> <p>Starostin, George S. (2017): Annotated Swadesh Wordlists for the Karen Group. Moscow: The Global Lexicostatistical Database.</p> </blockquote>
Test Phenology Annotations From Herbarium Specimen Dataset - Prunus serotina
<p>This is a test dataset of phenology annotations for the Black Cherry, P. serotina, used in the submitted manuscript:</p> <p>Brenskelle, L., B. Stucky, J. Deck, R. Walls, R. P. Guralnick [submitted]. Integrating herbarium specimen observations into global phenology data systems. Applications in Plant Sciences.</p>
Words or terms? Saṃjñā annotated dataset
<p>These data were used for the study published in:</p> <p>Lugli, Ligeia. 2019. Words or terms? Models of terminology and the translation of Buddhist Sanskrit vocabulary. In Alice Collett (ed.) Buddhism and Translation: Historical and Contextual Perspectives, New York: SUNY.</p> <p>data include:<br> 1. concordance lines for saṃjñā used for the study mentioned above. The concordance lines have been exported from the Sketch Engine and come from an automatically segmented corpus (segmenter = Lugli's version 1). They have not been proofread and contain segmentation errors. <br> 2. csv file with Lugli's semantic annotation of the concordance lines for saṃjñā. The data was annotated by Ligeia Lugli in 2017; part of the data constitutes a much revised version of a dataset originally prepared by Roberto Garcia for the Buddhist Translators Workbench in 2016.<br> 3. a pre-publication version of the study</p> <p> </p> <p> </p> <p>The creation of these data was funded by the British Academy through a Newton International Fellowship; the research was conducted at King's College London.<br> </p>
CLDF Dataset derived from Zhivlov's "Annotated Swadesh wordlists for the Ob-Ugrian group" from 2011
<p>Cite the source of the dataset as:</p> <blockquote> <p>Zhivlov, M. (2011): Annotated Swadesh wordlists for the Ob-Ugrian group (Uralic family). The Global Lexicostatistical Database. Moscow: RGGU.</p> </blockquote>
UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>
ForTrunkDet - Image dataset of visible and thermal annotated images for forest tree trunk detection
<p>Forest dataset composed by visible and thermal images with trunk annotations. The images were acquired in three different portuguese forests and were captured by four different cameras:</p> <ul> <li>GoPro Hero6</li> <li>Allied Mako G-125</li> <li>FLIR M232</li> <li>ZED Stereo</li> </ul> <p>The images and annotations are stored in two zip files:</p> <ul> <li>forest_dataset_original.zip - original dataset</li> <li>forest_dataset_augmented - augmented dataset</li> </ul> <p>Also, the subsets that were used to train, validate and test some deep learning models are available in three .TXT files (train.txt, val.txt and test.txt), where each file line corresponds to an image name.</p> <p> </p>
ArXiV-Entity/Relation annotated dataset
<p>This dataset is a collection of abstracts from the CS section of ArXiV, each annotated with <a href="https://github.com/dwadden/dygiepp">DyGIE++</a> (SciERC model)</p> <p>The dataset can be used to train triple extractors or to cluster triples (in the Computer Science and AI domains).</p> <p>Supersedes the ArXiV-AIKG dataset as these triples are unconstrained (so they don't forcibly appear in AIKG)</p>
Mycobacteroides abscessus subp. bolletii strain associated with a persistent infection (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a <strong><em>Mycobacteroides abscessus subp. bolletti </em></strong>strain associated with a persistente infection. (genome anotation was performed using Bakta v1.2.2 https://github.com/oschwengers/bakta)</p> <p>The raw sequence reads were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB57933; Run Accession: ERR10554471).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.