Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo40/100

VivesDebate: A New Annotated Multilingual Corpus of Argumentation in a Debate Tournament

<p>The application of the latest Natural Language Processing breakthroughs in computational argumentation has shown promising results which have raised the interest in this area of research. However, the available corpora with argumentative annotations are often limited to a very specific purpose or are not of adequate size to take advantage of state-of-the-art deep learning techniques (e.g., deep neural networks). In this paper, we present VivesDebate, a large, richly annotated, and versatile professional debate corpus for computational argumentation research. The corpus has been created from 29 transcripts of a debate tournament in Catalan and has been machine-translated into Spanish and English. The annotation contains argumentative propositions, argumentative relations, debate interactions, and professional evaluations of the arguments and argumentation. The presented corpus can be useful for research on a heterogeneous set of computational argumentation underlying tasks such as argument mining, argument analysis, argument evaluation, or argument generation among others. All this makes VivesDebate&nbsp;a valuable resource for computational argumentation research within the context of massive corpora aimed at Natural Language Processing tasks.</p>

opencc-by-nc-sa-4.0Jul 2021View details →
zenodo40/100

Pennsylvania German word list (lemmatized and POS-annotated)

<p>The file presents the words used in the Pennsylvania German part of the ENDE corpus (www.deitsch.eu). The list contains every lemma with its associated word forms documented in the corpus, comprised of&nbsp;1761 lemmata and 2704 word forms.</p> <p>The ENDE corpus (&ldquo;English-Deitsch&nbsp;translation corpus&rdquo;) is the first POS-annotated and searchable text corpus in Pennsylvania German (= Deitsch;&nbsp;ISO language code: pdc), aligned to the English source texts. Despite many digital texts in Deitsch are available on the internet, there are, so far, no digital corpora for this language. This is due mainly to the lack of a generally recognized standard variety which could serve as a reference point for the linguistic analysis needed for lemmatization and annotation.</p> <p>Lemmatization was done with the help of different lexicographic resources (https://www.deitsch.eu/news/view/9) most of which follow other spelling conventions. A fair number of word forms,&nbsp;especially English loanwords of some sort, cannot be found in the dictionaries. Moreover, the&nbsp;variety used here&nbsp;is characterized by a high variability regarding not only the spelling but also other aspects of the&nbsp;language.</p> <p>Part-of-speech tags were assigned manually (see tagsets A and B below). These tagsets for part-of-speech annotation of Deitsch texts are based on the 2017 version of the STTS system created and widely used for German (https://ids-pub.bsz-bw.de/frontdoor/deliver/index/docId/6063/file/Westpfahl_Schmidt_Jonietz_Borlinghaus_STTS_2_0_2017.pdf), which has been slightly modified and adapted to the corpus texts written in the Plain Deitsch variety. Tagset A gives a broader view and refers to the lemma level, tagset B is more fine-grained and suitable for&nbsp;&nbsp;the single word forms documented in the corpus. Only those tags are listed which are actually employed for the annotation of the corpus texts. Foreign items not integrated in the Deitsch text flow (e.g. English quotations) have been omitted.</p> <p>For more details about the corpus and the project please refer to the above mentioned website.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Genome assembly and annotation of Pisum sativum cultivar ZW6 (PeaZW6)

<p>This reposity stores the genome assembly and gene annotation of Pisum sativm cultivar ZW6 (PeaZW6)</p> <p>Current Version : Release Candidate Version 2 (RC2)</p> <p>Associated NCBI BioProject :&nbsp;<strong>PRJNA730094</strong></p> <p>Correspondance&nbsp;: gaoshh@im.ac.cn</p> <p>&nbsp;</p> <p>pea.assembly.ZW6.RC2.fasta.gz&nbsp;- Full&nbsp;genome sequences</p> <p>pea.assembly.ZW6.RC2.chr.fasta.gz - Genome sequences with only chromosome molecules&nbsp;</p> <p>pea.assembly.ZW6.RC2.annotated.gff3 / gtf / bed - Gene annotation in GFF3 / GTF / BED formats</p> <p>pea.assembly.ZW6.RC2.annotated.cds.fasta - Gene coding sequences</p> <p>pea.assembly.ZW6.RC2.annotated.proteins.fasta - Gene protein sequences</p> <p>pea.assembly.ZW6.RC2.annotated.annotations.txt - Additional annotation information in tabular text format (TSV)</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Paradinium-Oithona annotated images and count data

<p>Human annotated images, machine learning derived count data, and environmental data associated with our paper &quot; Discovery and dynamics of a cryptic marine copepod-parasite interaction&quot; published in Marine Ecology Progress Series. The work is based on imagery generated by the <a href="https://doi.org/10.1002/lom3.10394">Scripps Plankton Camera System</a> deployed at the Scripps Pier in La Jolla, California, USA.</p> <p>This repository contains:</p> <ul> <li><strong>human_annotated_SPC_data.zip </strong>- Regions of Interested collected by the SPC sorted by a human expert into relevant classes: Oithona, Oithona with parasite, and Oithona with eggs. Access to the &quot;other&quot; class is available upon request.</li> <li><strong>corrected_counts_091817.txt</strong> - Count data from the summer of 2015 generated a fine tuned deep neural network with associated human corrected counts.</li> <li><strong>oith_parasite_081420.txt </strong>-<strong> </strong>Count data from March 2015 - April 2016. Data collected after August 2015 is machine generated and should considered a maximum estimate of relative abundance.</li> <li><strong>autoss_a8bc_287c_07d6.csv </strong>- Environmental data measured by the automated shore station deployed at Scripps Pier and maintained the <a href="https://sccoos.org/">Southern California Coastal Ocean Observing System</a>. Complete shore station data can be accessed via <a href="https://erddap.sccoos.org/erddap/tabledap/HABs-ScrippsPier.html?Location_Code,latitude,longitude&amp;distinct()">NOAA ERDDAP</a>.<strong> </strong></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

EPIC: Annotated epileptic EEG independent components for artifact reduction

<p>Scalp electroencephalogram is a non-invasive multi-channel biosignal that records the brain&rsquo;s electrical activity. It is highly&nbsp;susceptible to noise that might overshadow important data. Independent component analysis is one of the most used artifact&nbsp;removal methods. Independent component analysis separates data into different components, although it can not automatically&nbsp;reject the noisy ones. Therefore, experts are needed to decide which components must be removed before reconstructing&nbsp;the data. To automate this method, researchers have developed classifiers to identify noisy components. However, to&nbsp;build these classifiers, they need annotated data. Manually classifying independent components is a time-consuming task.&nbsp;Furthermore, few labeled data are publicly available. This dataset&nbsp;is composed of a&nbsp;source of annotated electroencephalogram&nbsp;independent components acquired from patients with epilepsy (EPIC Dataset). This dataset&nbsp;contains 77,426 independent&nbsp;components obtained from approximately 613 hours of&nbsp;electroencephalogram, visually inspected by two experts, which was&nbsp;already successfully utilized to develop independent component classifiers.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Polifonia_Corpus_Wikipedia_Annotations_DE

<p>Polifonia_Corpus_Wikipedia_Annotations_DE</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Polifonia_Corpus_Pilots_Annotations_Meetups

<p>Polifonia_Corpus_Pilots_Annotations_Meetups</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Polifonia_Corpus_Pilots_Annotations_Child

<p>Polifonia_Corpus_Pilots_Annotations_Child</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Polifonia_Corpus_Wikipedia_Annotation_EN

<p>Polifonia_Corpus_Wikipedia_Annotation_EN</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

SocialDisNER corpus: gold standard annotations for detection of disease mentions in Spanish tweets

<p><strong>If you use any data from this repository, please cite our scientific paper instead of the Zenodo repo:&nbsp;</strong></p> <p>Luis Gasco S&aacute;nchez, Darryl Estrada Zavala, Eul&agrave;lia Farr&eacute;-Maduell, Salvador Lima-L&oacute;pez, Antonio Miranda-Escalada, and Martin Krallinger. 2022.&nbsp;<a href="https://aclanthology.org/2022.smm4h-1.48">The SocialDisNER shared task on detection of disease mentions in health-relevant content from social media: methods, evaluation, guidelines and corpora</a>. In&nbsp;<em>Proceedings of The Seventh Workshop on Social Media Mining for Health Applications, Workshop &amp; Shared Task</em>, pages 182&ndash;189, Gyeongju, Republic of Korea. Association for Computational Linguistics.</p> <pre><code class="language-json">@inproceedings{gasco2022socialdisner, title = "The {S}ocial{D}is{NER} shared task on detection of disease mentions in health-relevant content from social media: methods, evaluation, guidelines and corpora", author = "Gasco S{\'a}nchez, Luis and Estrada Zavala, Darryl and Farr{\'e}-Maduell, Eul{\`a}lia and Lima-L{\'o}pez, Salvador and Miranda-Escalada, Antonio and Krallinger, Martin", booktitle = "Proceedings of The Seventh Workshop on Social Media Mining for Health Applications, Workshop {\&amp;} Shared Task", month = oct, year = "2022", address = "Gyeongju, Republic of Korea", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2022.smm4h-1.48", pages = "182--189" }</code></pre> <p>&nbsp;</p> <p><strong>Introduction:</strong><br> The&nbsp;<strong>SocialDisNER corpus</strong>&nbsp;of the SMM4H 2022 &ndash; Task 10 task focus on the recognition of disease mentions in tweets written in Spanish after selecting primarily<strong><em>&nbsp;first-hand experience of diseases</em></strong>&nbsp;and other health-relevant content (from patient associations, professional healthcare institutions, and through&nbsp;followers of patient association accounts of a&nbsp;<em>diversity of pathologies</em>&nbsp;including rare diseases, mental health, cancer, etc..).</p> <p><strong>SocialDisNER Gold Standard</strong></p> <p>The Gold Standard corpus&nbsp;was manually annotated by medical experts following the <a href="https://doi.org/10.5281/zenodo.6983041">SMM4H-SocialDisNER guidelines</a>. These guidelines were adapted from previous efforts used to annotate patient clinical records and medical literature. It covers&nbsp;rules for annotating&nbsp;<strong>mentions&nbsp;of diseases</strong>&nbsp;in health-related tweets in Spanish,</p> <p>The training set consists of 5000 tweets written in Spanish and the validation set consists of 2500 tweets written in Spanish. Both sets have been manually annotated by healthcare professionals. The test dataset contains 23430 tweets, although only 2000 will be used to evaluate the systems participating in the task (the rest is background set).&nbsp;We don&#39;t plan to publish the test set, but if you want you can test your system from <a href="https://codalab.lisn.upsaclay.fr/competitions/3531">SocialDisNER Codalab</a>.</p> <p><strong>SocialDisNER Large Scale Corpus</strong></p> <p>The large-scale data contains mentions automatically extracted from a set of 85000 tweets. Separate datasets are shown for each entity including diseases, drugs, symptoms, professions, procedures, species, morphology neoplasm, and persons.</p> <p><strong>SocialDisNER co-mention networks</strong></p> <p>We have computed a co-occurrence matrix of the extracted diseases, as well as several co-mention matrices between the disease mentions and the rest of the entities in the large-scale corpora.</p> <p>&nbsp;</p> <p><strong>File structure:</strong></p> <p>The structure of the corpus is:&nbsp;</p> <ul> <li><strong>SocialDisNER_Data:</strong> <ul> <li>training-validation-data folder <ul> <li><strong><em>train-valid-txt-files</em></strong>:&nbsp;&nbsp;folder with training and validation text files. One text file per tweet, the file name corresponds to the tweet id.&nbsp;One sub-directory per corpus split (train and valid). The files named&nbsp;<em>ids_dev_set.txt</em>&nbsp;and<em>&nbsp;ids_train_set.txt&nbsp;</em>contain the list of file identifiers for each of the data splits (validation and train).</li> <li><strong><em>mentions.tsv</em></strong>:&nbsp;This file contains the manually annotated disease mentions. The file has the following fields: <ul> <li><em>tweets_id</em>: This is the id of the tweet, using Twitter API you can query the content of the tweet.</li> <li><em>Begin</em>: This is the position in the tweet where the annotation was found.</li> <li><em>End</em>: This is the position of the last character&nbsp;of the annotation in the tweet.</li> <li><em>Type:&nbsp;</em>This is the type of entity found, in&nbsp;our case &quot;ENFERMEDAD&quot;.</li> <li><em>Extraction</em>: This is the literal extraction, in other words, the fragment of text which refers to the annotation.&nbsp;</li> </ul> </li> </ul> </li> <li>test-data folder: <ul> <li><strong>test-data-txt-files</strong>: folder with test text files. One file per tweet, the file name corresponds to the tweet id. The folder contains 23430 tweets to be used as test set of the task. Of them, 2000 will be used to evaluate the participating systems.</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>SocialDisNER_LargeScale_additionaldata:</strong> <ul> <li>socialdisner_diseases: <ul> <li><strong>tweets_txt:</strong>&nbsp;Folder with large-scale tweet database. One text file per tweet, the file name corresponds to the tweet id.</li> <li><strong>diseases_mentions.tsv</strong>: This file contains the automatically annotated disease mentions from the large-scale SocialDisNER corpus (Silver Standard). The structure is the same than the Golden Standard annotations.</li> </ul> </li> <li>socialdisner_ENTITY: Each folder with this naming convention contains the following data structure. Corpora have been generated with mentions of diseases, drugs, symptoms, professions, procedures, species, morphology neoplasm and persons <ul> <li><strong>tweets_txt:</strong>&nbsp;Folder with large-scale tweet database. One text file per tweet, the file name corresponds to the tweet id.</li> <li><strong>ENTITY_mentions.tsv</strong>: This file contains the automatically annotated mentions of type &ldquo;ENTITY&rdquo; from the large-scale SocialDisNER corpus (Silver Standard). The structure is the same than the Golden Standard annotations.</li> </ul> </li> <li>socialdisner_networks: This folder contains tsv files containing the co-mention matrices between the diseases and the rest of the entities of the large-scale socialdisner data. Each file follows the following naming convention: <ul> <li><strong>socialdisner_disease-ENTITY_net.tsv</strong><em>:&nbsp; </em>The tsv file contains a series of columns and rows corresponding to the mentions used for building the matrix. Each column is separated by &ldquo;;&rdquo;. The type of each mention is identified by the label in parentheses of each title. The count represents the number of times that mention x and mention y were found in the same tweet of the large-scale dataset.</li> <li><strong>socialdiser_disease_net.tsv</strong>:&nbsp;This tsv file contains the array of socialdisner-disease large-scale corpus co-mentions separated by &quot;;&quot;. This file can be loaded into NetworkX to perform disease co-morbidity analysis on the socialdisner-disease large-scale data.</li> </ul> </li> </ul> </li> </ul> <p><em>Note:&nbsp;In previous versions of the dataset the order of the columns in the mentions.tsv file was not in the correct order. From this version onwards the order is correct and adequate to send the predictions of the task.</em></p> <p>&nbsp;</p> <p>For further information, please visit <a href="https://temu.bsc.es/socialdisner/">https://temu.bsc.es/socialdisner/</a></p> <p><strong>Summary statistics:</strong></p> <table> <caption>Manually annotated data</caption> <thead> <tr> <th scope="row">&nbsp;</th> <th scope="col">Training set</th> <th scope="col">Development set</th> </tr> </thead> <tbody> <tr> <th scope="row"># tweets</th> <td>5000</td> <td>2500</td> </tr> <tr> <th scope="row"># characters</th> <td>1253431</td> <td>516768</td> </tr> <tr> <th scope="row"># tokens</th> <td>211555</td> <td>84478</td> </tr> <tr> <th scope="row">Avg. char / tweet</th> <td>250.69</td> <td>206.71</td> </tr> <tr> <th scope="row">Avg. tok. / tweet</th> <td>42.31</td> <td>33.79</td> </tr> <tr> <th scope="row"># mentions</th> <td>15173</td> <td>4252</td> </tr> <tr> <th scope="row"># unique mentions</th> <td>4407</td> <td>1413</td> </tr> </tbody> </table> <p>&nbsp;</p> <table> <caption>Large-scale annotated data (Silver Standard)</caption> <tbody> <tr> <td>&nbsp;</td> <td><em>Socialdisner-diseases</em></td> <td><em>Socialdisner-pharma</em></td> <td><em>Socialdisner-morphology_neoplasms</em></td> <td><em>Socialdisner-symptoms</em></td> <td><em>Socialdisner-professions</em></td> <td><em>Socialdisner-Procedures</em></td> <td><em>Socialdisnerv-Person</em></td> <td><em>Socialdisner-Species</em></td> </tr> <tr> <td><strong># tweets</strong></td> <td>85077</td> <td>1759</td> <td>8518</td> <td>12624</td> <td>15831</td> <td>11462</td> <td>41033</td> <td>12118</td> </tr> <tr> <td><strong># characters</strong></td> <td>19920670</td> <td>435141</td> <td>2082574</td> <td>3023784</td> <td>4063114</td> <td>2873791</td> <td>10273278</td> <td>2933925</td> </tr> <tr> <td><strong># tokens</strong></td> <td>3236411</td> <td>68269</td> <td>332539</td> <td>521503</td> <td>660071</td> <td>467059</td> <td>1689479</td> <td>486249</td> </tr> <tr> <td><strong>Avg. char / tweet</strong></td> <td>234.15</td> <td>247.38</td> <td>244.49</td> <td>239.53</td> <td>256.66</td> <td>250.72</td> <td>250.37</td> <td>242.11</td> </tr> <tr> <td><strong>Avg. tok. / tweet</strong></td> <td>38.04</td> <td>38.81</td> <td>39.04</td> <td>41.31</td> <td>41.69</td> <td>40.75</td> <td>41.17</td> <td>40.13</td> </tr> <tr> <td><strong># mentions</strong></td> <td>116260</td> <td>1029</td> <td>8943</td> <td>12896</td> <td>18590</td> <td>10080</td> <td>58007</td> <td>14014</td> </tr> <tr> <td><strong># unique mentions</strong></td> <td>16034</td> <td>530</td> <td>541</td> <td>6991</td> <td>3667</td> <td>3841</td> <td>3446</td> <td>1676</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Do not share the data with other individuals/teams without permission from the task organizer. Tweets IDs are the primary source of information. Tweet texts are provided as support material. By downloading this resource, you agree to the Twitter <a href="https://twitter.com/en/tos">Terms of Service</a>, <a href="https://twitter.com/en/privacy">Privacy Policy</a>, <a href="https://developer.twitter.com/en/developer-terms/agreement">Developer Agreement</a>, and <a href="https://developer.twitter.com/en/developer-terms/policy">Developer Policy</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

<p>The European starling, <em>Sturnus vulgaris</em>, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (<em>S. vulgaris</em> vAU), and a second North American genome (<em>S. vulgaris</em> vNA), as complementary reference genomes for population genetic and evolutionary characterisation. <em>S. vulgaris</em> vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (<em>Taeniopygia guttata</em>) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using <em>S. vulgaris</em> vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, <em>S. vulgaris</em> vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.</p>

opencc-zeroJul 2022View details →
zenodo40/100

Training data for 'Functional annotation of protein sequences' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for functional annotation of protein sequences.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Training data for 'Refining Manual Genome Annotations with Apollo (eukaryotes)' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for manual curation of eukaryotic genome annotation using Apollo.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)

<p>Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Saccharomyces cerevisiae Bud-Annotation pipeline: Napari Example Data

<p>Example dataset for bud-annotation plugin Napari.</p> <p>Here we provide two sets of 3 images&nbsp;of single molecule mRNA FISH on Saccheromyces cerevisiae (BY4741)&nbsp;strains containing an&nbsp;mRNA bud localization reporter at the DOA1 locus.&nbsp;<br> The image set&nbsp;labeled EXPERIMENT contains images of the strain in which a localization element was present in the reporter mRNA. The reporter can be seen to localize to the bud.&nbsp;<br> The image set&nbsp;labeled CONTROL&nbsp;contains images of the control strain&nbsp;without localization element present in the reporter mRNA. The reporter can be seen to be&nbsp;randomly distributed throughout the cell.&nbsp;</p> <p>Images were acquired as 41 z-stacks per&nbsp;fluorescence channel&nbsp;CY5, CY3.5, CY3 and DAPI. For each fluorescence image a corresponding DIC image was acquired as well.&nbsp;</p> <p>The following smFISH probes were used:&nbsp;<br> - CY5:&nbsp;probes targeting the endogenous ASH1 and CLB2 mRNA to be&nbsp;used as bud marker&nbsp;(Quasar 670)<br> - CY3.5:&nbsp;probes targeting the DOA1&nbsp;mRNA reporter (CAL Fluor Red 610)<br> - CY3: probes targeting the MS2v6 sequence inserted at the 3&#39; end of the DOA1 mRNA reporter&nbsp;(Quasar 570)&nbsp;</p> <p>FISH-QUANT spot analysis results for the CY3.5 channel from these images are&nbsp;included. This data can be used to extract the spot information and assign spots to either mother cell or bud.&nbsp;Examples of nuclear, bud and cell masks for all images are provided as well.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

JOSSE: A Software Development Effort Dataset Annotated with Expert Estimates

<p>The JIRA Open-Source Software Effort (JOSSE) dataset consists of software development and maintenance tasks collected from the JIRA issue tracking system for Apache, JBoss, And Spring open-source projects. All the issues were annotated with actual effort and 19% of them were annotated with expert estimates. JOSSE is a task-based dataset with a textual attribute represented as a task description for each data point. This paper explains how the data were collected and details six data quality refinement procedures of the data points.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Aquitalea palustris nov. sp. strains MWU14-2217T and MWU14-2470 RASTtk annotations

<p>RASTtk annotation of the genomes of <em>Aquitalea palustris</em> nov. sp. strains MWU14-2217 (type isolate) and MWU14-2470 isolated from wild cranberry bog soil and berry surfaces, respectively, in the Cape Cod National Seashore during a 2014 culture-dependent survey of bacteria from wetlands bogs.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Common guillemots in the Baltic Sea studied with video surveillance and object detection: raw data, annotations, model, and model outputs

<p>The data comes from common guillemots studied at Stora Karlsö, Sweden between 2019 and 2021. The common guillemots breed at an artificial cliff, and has been filmed continusly from above over three breeding seasons. Using the video material, a YOLOv5 model has been trained to detect adult birds, chicks and eggs. The dataset contains annotations (bounding boxes) used for training the model, the model itself, and outputs from the model (object detections).</p> <p>The data can be used and shared freely.</p>

opencc-zeroSep 2022View details →
zenodo40/100

Annotated-VocalSet: A Singing Voice Dataset

<p>This dataset provides annotations for the <a href="https://doi.org/10.5281/zenodo.1442513">VocalSet dataset</a>, which is available online at</p> <pre><a href="https://doi.org/10.5281/zenodo.1442513">https://doi.org/10.5281/zenodo.1442513</a></pre> <p>.</p> <p>The annotations generated for the VocalSet audio files include fundamental frequency contour, note onset, note offset, the transition between notes, note F0, note duration, Midi pitch, and lyrics.</p> <p><a href="https://doi.org/10.5281/zenodo.1442513">VocalSet</a> consists of more than 10 hours of monophonic recorded audio of professional singers in a variety of vocal techniques (n = 17) and several singers (m = 20) with several WAV files (p = 3560). However, although several categories, including techniques, singers, tempo, and loudness, are considered in the dataset, the sung notes were not annotated. Therefore, this dataset aims to annotate VocalSet to make it a more powerful dataset for researchers.</p> <p>Details of the dataset are provided in the following academic journal paper.</p> <p><a href="https://www.mdpi.com/2076-3417/12/18/9257">Faghih, Behnam, and Joseph Timoney. 2022. &quot;Annotated-VocalSet: A Singing Voice Dataset&quot;&nbsp;<em>Applied Sciences</em>&nbsp;12, no. 18: 9257. https://doi.org/10.3390/app12189257</a></p> <p>Please use the above paper to cite this dataset.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Sei whole-genome sequence class annotations

<p>Sei sequence class whole-genome annotations are available in the following files:</p> <ul> <li> <p>sorted.hg38.tiling.bed.ipca_randomized_300.labels.merged.bed - The sorted, merged sequence class assignments from Louvain community clustering of the 30 million sequences, uniformly tiling the whole human genome. The fourth column is the sequence class number, with any sequence classes numbering 40-61 excluded from our analyses in the publication. Sequence classes 0-39 can be mapped to the following labels:&nbsp;<a href="https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names">https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names</a></p> </li> <li> <p>sorted.hg19.tiling.bed.ipca_randomized_300.labels.merged.bed - lifted over version of the hg38 BED file.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record