Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
Annotated Benchmark of Real-World Data for Approximate Functional Dependency Discovery
<p><strong>Annotated Benchmark of Real-World Data for Approximate Functional Dependency Discovery</strong></p> <p>This collection consists of ten open access relations commonly used by the data management community. In addition to the relations themselves (please take note of the references to the original sources below), we added three lists in this collection that describe approximate functional dependencies found in the relations. These lists are the result of a manual annotation process performed by two independent individuals by consulting the respective schemas of the relations and identifying column combinations where one column implies another based on its semantics. As an example, in the <em>claims.csv</em> file, the <em>AirportCode</em> implies <em>AirportName</em>, as each code should be unique for a given airport.</p> <p>The file <em>ground_truth.csv</em> is a comma separated file containing approximate functional dependencies. <em>table</em> describes the relation we refer to, <em>lhs</em> and <em>rhs</em> reference two columns of those relations where semantically we found that <em>lhs</em> implies <em>rhs</em>.</p> <p>The file <em>excluded_candidates.csv</em> and <em>included_candidates.csv</em> list all column combinations that were excluded or included in the manual annotation, respectively. We excluded a candidate if there was no tuple where both attributes had a value or if the <em>g3_prime</em> value was too small.</p> <p><strong>Dataset References</strong></p> <ul> <li><em>adult.csv</em>: Dua, D. and Graff, C. (2019). <a href="http://archive.ics.uci.edu/ml">UCI Machine Learning Repository</a>. Irvine, CA: University of California, School of Information and Computer Science.</li> <li><em>claims.csv</em>: TSA Claims Data 2002 to 2006, <a href="https://www.dhs.gov/tsa-claims-data">published by the U.S. Department of Homeland Security</a>.</li> <li><em>dblp10k.csv</em>: Frequency-aware Similarity Measures. Lange, Dustin; Naumann, Felix (2011). 243–248. <a href="https://hpi.de/naumann/projects/repeatability/datasets/dblp-dataset.html">Made available as DBLP Dataset 2</a>.</li> <li><em>hospital.csv</em>: Hospital dataset used in Johann Birnick, Thomas Bläsius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270–2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> <li><em>t_biocase_...</em> files: t_bioc_... files used in Johann Birnick, Thomas Bläsius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270–2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> <li><em>tax.csv</em>: Tax dataset used in Johann Birnick, Thomas Bläsius, Tobias Friedrich, Felix Naumann, Thorsten Papenbrock, and Martin Schirneck. 2020. Hitting set enumeration with partial information for unique column combination discovery. Proc. VLDB Endow. 13, 12 (August 2020), 2270–2283. https://doi.org/10.14778/3407790.3407824. <a href="https://owncloud.hpi.de/s/j6Z0yvXC0qhtGCk/download">Made available as part the dataset collection to that paper.</a></li> </ul>
An annotated compilation of chronometric dates for the Middle-Upper Palaeolithic Transition (45-30 ka BP) in northern Iberia (Spain). Source Data.
<p>This repository contains the files of the chronometric dates framed between 45-30 ka BP in Northern Iberia.</p>
Multifaceted quality assessment of gene repertoire annotation with OMArk
<p>Dataset associated to the OMArk paper.</p><p>Contain eight archives:</p><p>Supplementary_Tables</p><p>The Supplementary Table files referred to in the paper</p><p>OMAmerDB:</p><p>The OMAmer database constructed using the whole dataset of the OMA database (November 2022 Release) and used in the paper. An OMAmer database is necessary to run OMArk.</p><p>Simulation:<br>Proteomes with artificially introduced errors, contaminants or depleted completeness, used to assess OMArk's performance. The archive contains the generated proteomes (Simulated_Data) and their OMArk quality assessments (omark). They also contains the OMAmer results (OMAmerResults) that were used to run OMArk and BUSCO completeness assessments (BUSCO).</p><p>*Note that for storage efficiency, only the non-redundant part of the data (added errors, added contamination, random fraction of proteomes) are stored there. The full modified proteome can be regenerated from these data and the source proteomes.</p><p>Reference Proteomes:</p><p>The UniProt Reference Proteomes (Proteomes) (2021_04) and their proteome quality assesment results according to OMArk. The archive contains the source proteome FASTA (Source folder), OMAmer results for these proteomes (omamer folder) , OMArk results (omark folder), and BUSCO completeness assesments (BUSCO folder). It also contains a subfolder that contains part of the Contamination detection experiment (Contamination folder).</p><p>Ensembl_Metazoa_AssemblyChange.<br><br>Contains Ensembl Metazoa proteomes with version change between version 52 and 54 as well as their quality assesment resuls for both version. The archive contains the source proteomes FASTA (Source folder), a Splice file that group together all proteins coded by the same gene (Splice folder), omamer results for the proteomes (omamer folder) and the omark results (omark folder)</p><p>MissingGenesBLAST<br><br>Contains sequences of HOGs considered as missing in the Human proteome, that was used to look for sequences in the human genome.</p><p>Ensembl_NCBI_Results</p><p>Contains OMArk and BUSCO results for Ensembl and NCBI proteomes. These results were then used to evaluate OMArk biais due to source of proteomes in the OMA database.</p><p>Notebooks<br>Jupyter Notebooks that were used to perform the analysis described in the paper<br><br> </p>
Enzymes from the BRENDA and CAZy databases annotated with organism growth temperatures and predicted Topt
<p>This repo is an updated version of repo <strong>Gang Li, & Martin KM Engqvist. (2019). Enzymes from the BRENDA database annotated with organism growth temperatures and predicted <em>T</em><sub>opt</sub> (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.2539114. </strong></p> <p>Experimental as well as predicted organism growth temperatures were used to annotate enzymes from the BRENDA database (doi: 10.1093/nar/gky1048, https://www.brenda-enzymes.org) version 2018.2 (July 2018) and CAZy database (http://www.cazy.org/). </p> <p>An updated machine learning model was applied to predict the optimal functional temperature of enzymes from BRENDA and CAZy. </p> <p>There are four files in this repo:</p> <p>1. 'annotated_brenda.tsv' is a tab-seperated file that contains the annotated enzymes from BRENDA. There are 9 columns in the file: index column; "ec", EC number; "uniprot_id", protein id in Uniprot database; "domain", the domain of life (superkingdom), either Archaea, Bacteria, or Eukarya; "organism", species name; "ogt", optimal growth temperature of the organism; "ogt_note", whether the experimental or predicted ogt is used; "topt", the optimal functional temperature of the enzyme; "topt_note", whether the experimental or predicted topt is used.</p> <p>2. 'annotated_cazy.tsv' is a tab-seperated file that contains the annotated enzymes from CAZy. There are 12 columns in the file: index column; "family", CAZy family id; "genbank", genbank id; "Protein Name", the protein name from CAZy database; "ec", EC number; "organism", strain name; "uniprot_id", protein id in Uniprot database; "PDB/3D", structure id in PDB database; "ogt", optimal growth temperature of the organism; "ogt_note", whether the experimental or predicted ogt is used; "topt", the optimal functional temperature of the enzyme; "topt_note", whether the experimental or predicted topt is used.</p> <p>3. 'brenda.sql', which is a SQLite3 database version of 'annotated_brenda.tsv', with an additional column of enzyme sequences.</p> <p>4. 'cazy.sql', which is a SQLite3 database version of 'annotated_cazy.tsv'', with an additional column of enzyme sequences.</p> <p>The SQLite3 databases are for the Tome tool (<a href="https://github.com/EngqvistLab/Tome">https://github.com/EngqvistLab/Tome</a>), version 2.0.</p>
MiRoR11 - P2 - Annotated dataset for spin-related types of statements (statements of similarity and within-group comparisons)
<p>180 abstracts / 2401 sentences annotated for 2 types of spin-related statements: statements of similarity and within-group comparisons.</p>
Enhanced genome annotation strategy provides novel insights on the phylogeny of 'Flaviviridae': Supplementary material
<p>SUPPLEMENTARY MATERIAL</p> <p><strong>Index</strong></p> <ul> <li> <p>Table S1 (tableS1.csv): genomic data.</p> </li> <li> <p>Table S2 (tableS2.csv): character categorization for selected nodes.</p> </li> <li> <p>Table S3 (tableS3.csv): programs and parameters.</p> </li> <li> <p>Table S4 (tableS4.csv): annotation efficiency.</p> </li> <li> <p>File S1 (fileS1.gff): gene annotation.</p> </li> <li> <p>File S2 (fileS2.xml): configuration file for BEAST 2 (configuration.xml).</p> </li> <li> <p>Figure S1 (figureS1.pdf): dendrogram depicting the hierarchical clusters of trees based on match-split distances.</p> </li> <li> <p>Figure S2 (figureS2.pdf): full version of the working phylogenetic hypothesis (tree No. 0 in table 1).</p> </li> </ul> <p><strong>Figure captions</strong></p> <ul> <li>Figure S1: A dendrogram depicting the hierarchical clusters of trees based on match-split distances. Outgroup sequences (<em>Hepacivirus</em>, <em>Pegivirus</em>, and <em>Pestivirus</em>) were removed to guarantee the compared tree topologies would have the same terminals. Tree numbers correspond to those in table 1 of the manuscript. I. No outgroup sequences; some matrices were partitioned. II. Outgroup sequences and partitioned matrices. *This tree was produced without outgroup sequences.</li> <li>Figure S2: Full version of the working phylogenetic hypothesis (tree No. 0 in table 1). Branch lengths represent an estimation of the number of substitutions per site. Node labels indicate SH-aLRT support / ultrafast bootstrap (only shown if one of there is a value below 90%). Clade names correlate to the character categorization analysis (see table S2). Branch labels represent the four genera: I = <em>Pestivirus</em>; II = <em>Pegivirus</em>; III = <em>Hepacivirus</em>; IV = <em>Flavivirus</em>. * The Ecuador Paraiso Escondido virus (EPEV) was isolated from sand flies (<em>Psathyromyia abonnenci</em>). The EPEV was the first sand fly-borne <em>Flavivirus</em> identified in the New World.</li> </ul> <p><strong>Manuscript title</strong></p> <p>FLAVi: an enhanced annotator for viral genomes of <em>Flaviviridae</em>.</p> <p><strong>Authors</strong></p> <ul> <li> <p>de Bernardi Schneider, Adriano. University of California San Diego. ORCID: 0000-0001-7487-266X.</p> </li> <li> <p>Jacob Machado, Denis. University of North Carolina at Charlotte. ORCID: 0000-0001-9858-4515. Corresponding author.</p> </li> <li> <p>Guirales, Sayal. University of North Carolina at Charlotte.</p> </li> <li> <p>Janies, Daniel. University of North Carolina at Charlotte.</p> </li> </ul> <p><em>First author</em>: Adriano de Bernardi Schneider and Denis Jacob Machado have contributed equally to the manuscript.</p> <p><strong>Contact information</strong></p> <ul> <li> <p>Corresponding author: Denis Jacob Machado, Ph.D.</p> </li> </ul> <ul> <li> <p>OrcID: 0000-0001-9858-4515.</p> </li> </ul> <ul> <li> <p>Email: dmachado [at] uncc.edu.</p> </li> </ul> <p><strong>Other additional material</strong></p> <ul> <li>In addition to the material listed above, all 31 tree topologies and 15 alignment matrices discussed in this manuscript will are available in TreeBASE (<a href="http://purl.org/phylo/treebase/phylows/study/TB2:S24096">http://purl.org/phylo/treebase/phylows/study/TB2:S24096</a>) after the publication of the manuscript.</li> <li>The FLAVi pipeline and all the original scripts are available at GitLab (<a href="https://gitlab.com/MachadoDJ/FLAVi">https://gitlab.com/MachadoDJ/FLAVi</a>).</li> <li>The web application can be accessed at <a href="http://flavi-web.com">http://flavi-web.com</a>.</li> </ul>
Manually Annotated Instances of Ich ('I') from the German KoLas Corpus
<p>Dataset used in Andresen/Knorr (2020). The dataset comprises 360 instances of <em>ich</em> ('I') taken from the German learner corpus KoLaS (Andresen/Knorr 2017, see <a href="http://hdl.handle.net/11022/0000-0001-B732-8">http://hdl.handle.net/11022/0000-0001-B732-8</a> for full corpus access) and manually annotated with categories taken from Steinhoff (2007).</p> <p>Column descriptions:</p> <ul> <li>document: name of the document by which it can be found in the KoLaS corpus</li> <li>code_annotator1 - code_annotator4: Annotations by four annotators. Possible values: Verfasser-<em>Ich</em> (author <em>I</em>), Forscher-<em>Ich</em> (researcher <em>I</em>), Erzähler-<em>Ich</em> (narrator <em>I</em>)</li> <li>max_agreement_freq: Highest number of anntators that agreed on one label</li> <li>max_agreement_label: Label on which the highest number of annotators agreed</li> <li>context_before: 150 characters of context before the match</li> <li>match: the match itself (either <em>ich</em> or <em>Ich</em>)</li> <li>context_after: 150 characters of context after the match</li> </ul> <p><strong>References</strong></p> <p>Andresen M, Knorr D. KoLaS – Ein Lernendenkorpus in der Schreibberatungsausbildung einsetzen. <em>Zeitschrift Schreiben</em>. Published online July 5, 2017:10-17.</p> <p>Andresen M, Knorr D. Exploring the Use of the Pronoun I in German Academic Texts with Machine Learning. In: Burghardt M, Müller-Birn C, eds. <em>Methoden und Anwendungen der Computational Humanities</em>. Lecture Notes in Informatics (LNI). Gesellschaft für Informatik; 2020.</p> <p>Steinhoff T. Zum ich-Gebrauch in Wissenschaftstexten. <em>Zeitschrift für germanistische Linguistik</em>. 2007;35(1-2):1–26.</p>
Annotated checklist of the beetles of Abeti Soprani, a silver fir forest of Central Italy
<p>The checklist contains 179 species of beetles which belong to 48 families. The species were collected during a field study carried out in the years 2012 and 2013 and aimed at describing the beetles of the area. The collection methods consisted in window flight traps and emergence traps. The study area, named Abeti Soprani, is a silver fir (<em>Abies alba</em>) forest located in the Central Apennines.</p> <p>The checklist is annotated with information on the taxonomy of the species (order and family), number of individuals, geographic position, habitat type (following EUNIS habitat classification 2017), sampling protocol, collector name, specialist name, IUCN Red List categories of the saproxylic species (Carpaneto et al. 2015). </p> <p>The terms used for the dataset fields follows the Darwin Core Maintenance Group. 2020. List of Darwin Core terms. Biodiversity Information Standards (TDWG). <a href="https://dwc.tdwg.org/list/">https://dwc.tdwg.org/list/</a></p> <p>Investigations on spatial patterns and diversity have been based on this dataset and published (Parisi et al. 2016, 2020).</p> <p>The harmonization of the dataset to the point of view of taxa, authorship, LSID and the massive upgrading of the related identifiers in Zenodo record was performed by the use of R script using respectively dplyr, taxize (Chamberlain and Szöcs, 2013) and zen4r (Blondel and Barde, 2020) packages.</p>
Annotated checklist of the beetles of chestnut agroforestry systems in Aspromonte, Southern Italy
<p>The checklist contains 255 species of beetles which belong to 49 families. The species were collected during a field study carried out in the years 2017 and aimed at describing the community of beetles. The collection methods consisted of window flight traps. The study area included 3 sites, two coppice stands, young and mature (38.180221 N, 15.784308 E), and a traditional fruit orchard (38.06018 N, 15.781616 E), located in the Italian Southern Apennines on the borders of the Aspromonte National Park.</p> <p>The checklist is annotated with information on the taxonomy of the species (order and family), number of individuals, locality, habitat type (following EUNIS habitat classification 2017), sampling protocol, collector name, specialist name, IUCN Red List categories of the saproxylic species (Carpaneto et al. 2015). </p> <p>The terms used for the dataset fields follows the Darwin Core Maintenance Group. 2020. List of Darwin Core terms. Biodiversity Information Standards (TDWG). <a href="https://dwc.tdwg.org/list/">https://dwc.tdwg.org/list/</a></p> <p>The Diversity of saproxylic beetle communities have been analysed and published (Parisi et al. 2020).</p> <p>The harmonization of the dataset to the point of view of taxa, authorship, LSID and the massive upgrading of the related identifiers in Zenodo record was performed by the use of R script using respectively dplyr, taxize (Chamberlain and Szöcs, 2013) and zen4r (Blondel and Barde, 2020) packages.</p>
Annotated checklist of the beetles of beech forests in Matese National Park, Central Italy
<p>The checklist contains 165 species of beetles which belong to 37 families. The species were collected during a field study carried out in the year 2018 and aimed at describing the community of beetles. The collection methods consisted of window flight traps. The study activities were carried out in four distinct beech forest stands based on their altitude (High and Low) and exposure (South and North) and located in the Italian Central Apennines. The sites are included in the Natura 2000 site IT 7222287 “La Gallinola - Monte Miletto - Monti del Matese” and Matese National Park.</p> <p>The checklist is annotated with information on the taxonomy of the species (order and family), number of individuals, geographic position, habitat type (following EUNIS habitat classification 2017), sampling protocol, collector name, specialist name, IUCN Red List categories of the saproxylic species (Carpaneto et al. 2015). </p> <p>The terms used for the dataset fields follows the Darwin Core Maintenance Group. 2020. List of Darwin Core terms. Biodiversity Information Standards (TDWG). <a href="https://dwc.tdwg.org/list/">https://dwc.tdwg.org/list/</a></p> <p>The discovery of a new species of beetle (Elateridae) for the Italian fauna was based on this dataset (Parisi et al., 2020).</p> <p>The harmonization of the dataset to the point of view of taxa, authorship, LSID and the massive upgrading of the related identifiers in Zenodo record was performed by the use of R script using respectively dplyr, taxize (Chamberlain and Szöcs, 2013) and zen4r (Blondel and Barde, 2020) packages.</p>
Annotated checklist of the beetles of three beech forests in Gran Sasso National Park, Central Italy
<p>The checklist contains 163 species of beetles which belong to 36 families. The species were collected during a field study carried out in the years 2013 and 2016 and aimed at describing the community of beetles. The collection methods consisted of window flight traps and emergence traps. The study area included 3 beech forest sites, named Prati di Tivo (42.5096 N, 13.5679 E), Venacquaro (42.4988 N, 13.5139 E) and Incodara (42.5123 N, 13.4735 E) located in the Italian Central Apennines. The sites are included in the Natura 2000 site IT7110202 “Gran Sasso”.</p> <p>The checklist is annotated with information on the taxonomy of the species (order and family), number of individuals, locality, habitat type (following EUNIS habitat classification 2017), sampling protocol, collector name, specialist name, IUCN Red List categories of the saproxylic species (Carpaneto et al. 2015). </p> <p>The terms used for the dataset fields follows the Darwin Core Maintenance Group. 2020. List of Darwin Core terms. Biodiversity Information Standards (TDWG). <a href="https://dwc.tdwg.org/list/">https://dwc.tdwg.org/list/</a></p> <p>Investigations on stand structure and forest biodiversity (Sabatini et al. 2016) and faunistic analysis (Zanetti and Parisi 2019) have been based on this dataset.</p> <p>The harmonization of the dataset to the point of view of taxa, authorship, LSID and the massive upgrading of the related identifiers in Zenodo record was performed by the use of R script using respectively dplyr, taxize (Chamberlain and Szöcs, 2013) and zen4r (Blondel and Barde, 2020) packages.</p>
A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)
<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie "Forrest Gump" (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>
Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10
<p>Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10</p> <p>Github: https://github.com/phucty/mtab4dbpedia<br> ---------------------------------------------------------------------------------------------------------------------------------------</p> <p>CEA: </p> <ul> <li> <p>Keep only valid entities in DBpedia 2016-10</p> </li> <li> <p>Resolve percentage encoding</p> </li> <li> <p>Add missing redirect entities</p> </li> </ul> <p>CTA: </p> <ul> <li> <p>Keep only valid types</p> </li> <li> <p>Resolve transitive types (parents and equivalent types of the specific type) with DBpedia ontology 2016-10</p> </li> </ul> <p>CPA:</p> <ul> <li> <p>Add equivalent properties</p> </li> </ul> <p>Statistic of Adapted Tabular data SemTab 2019</p> <pre><code>| | CEA | | | CPA | | | CTA | | | |---------|:--------:|:-------:|:------:|:--------:|:-------:|:------:|:--------:|---------|--------| | | Orginal | Adapted | Change | Orginal | Adapted | Change | Orginal | Adapted | Change | | Round 1 | 8418 | 8406 | -0.14% | 116 | 116 | 0.00% | 120 | 120 | 0.00% | | Round 2 | 463796 | 457567 | -1.34% | 6762 | 6762 | 0.00% | 14780 | 14333 | -3.02% | | Round 3 | 406827 | 406820 | 0.00% | 7575 | 7575 | 0.00% | 5762 | 5673 | -1.54% | | Round 4 | 107352 | 107351 | 0.00% | 2747 | 2747 | 0.00% | 1732 | 1717 | -0.87% |</code></pre> <p> </p> <p>---------------------------------------------------------------------------------------------------------------------------------------<br> DBpedia 2016-10 extra resources: (Original dataset http://downloads.dbpedia.org/2016-10/)</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_classes_2016-10.csv</p> <p>Information: DBpedia classes and parents: (We remove the abstract types: Agent, Thing)</p> <p>Total: 759 classes</p> <p>Structure: [class, parents (separate with space)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "City","Location Place PopulatedPlace Settlement"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_properties_2016-10.csv</p> <p>Information: DBpedia properties and these equivalents</p> <p>Total: 2865 properties</p> <p>Structure: [property, it’s equivalent properties] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "restingDate","deathDate"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_domains_2016-10.csv</p> <p>Information: DBpedia properties and these domain types</p> <p>Total: 2421 properties (have types as their domain)</p> <p>Structure: [property, type (domain)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "deathDate","Person"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_entities_2016-10.jsonl.bz2 </p> <p>Information: DBpedia entity dump</p> <p>Format: json list bz2 (bz2 Compressed json list)</p> <p>Source: DBpedia dump 2016-10 core</p> <p>Total: 5,289,577 entities (No disambiguation entities)</p> <p>Structure:</p> <p>An entity: for example “Tokyo”: (datatype: dictionary),</p> <p>{</p> <p>'wd': 'Q1322032', (Wikidata ID, datatype: string)</p> <p>'wp': 'Tokyo', (Wikipedia ID, add prefix <a href="https://en.wikipedia.org/wiki/">https://en.wikipedia.org/wiki/</a> + wp to get the Wikipedia URL, datatype: string)</p> <p>'dp': 'Tokyo', (DBpedia ID, add prefix <a href="http://dbpedia.org/resource/">http://dbpedia.org/resource/</a> + dp to get the DBpedia URL, datatype: string)</p> <p>'label': 'Tokyo', (Entity label, datatype: string)</p> <p>'aliases': ['To-kyo', 'Tôkyô Prefecture', ..], (Other entity names, datatype: list) </p> <p>'aliases_multilingual': ['东京小子', 'طوكيو', ...], (Other entity names in multilingual, datatype: list)</p> <p>'types_specific': 'City', (Entity direct type, datatype: string) </p> <p>'types_transitive': ['Human settlement', 'City', 'PopulatedPlace', 'Location', 'Place', 'Settlement'], (Entity transitive types, datatype: list)</p> <p>'claims_entity': { (entity statements, datatype: dictionary. Keys: properties, Values: list of tail entities)</p> <p>'governingBody': ['Tokyo Metropolitan Government'], </p> <p> 'subdivision': ['Honshu', 'Kantō region'],</p> <p>...</p> <p>},</p> <p>'claims_literal': {</p> <p>'string': { (String literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>'postalCode': ['JP-13'], </p> <p>'utcOffset': ['+09:00', '+9'],</p> <p>…</p> <p>}</p> <p>'time': { (Time literal: datatype: dictionary. Keys: properties, Values: list of date time</p> <p>'populationAsOf': ['2016-07-31'], </p> <p>...</p> <p>}), </p> <p>'quantity': { (Numerical literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>populationDesity: [6224.66, 6349.0], </p> <p>'maximumElevation': [2017], </p> <p>...</p> <p>},</p> <p>'pagerank': 2.2167366040153352e-06 (Entity page rank score calculated on DBpedia Graph)</p> <p>}</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Reginsmál annotation
<p>The dataset contains annotations of Reginsmál's poem. The annotations are on the Codex Regius manuscript (facsimile, diplomatic and normalised annotations) and from a linguistic point of view (the text is lemmatised and partially grammatically, i.e. parts of speech, analysed).</p>
Updated Files from the revision 1 (986 cMAGs and prokka annotations, all supplemental files)
<p>Included are the following files<br><br>1) Assembly files (47) in the form of .fasta for each of the 47 different samples (8 participants x 5 or 6 time points). This is the results of the assembly from hybrid long read data (all Pacbio Revio + ONT Promethion > Q20). Assemblies performed using metaMDBG. This file includes all contigs from the pipeline described in the manuscript and will include both high-quality and complete MAGs. <br>"hybrid_assemblies.tar.gz"<br><br>2) Assembly files (40) in the form of .fasta for the short read metaspades and tell-seq assembly methods<br>"SR_tellseq_assemblies.tar.gz"<br>"SR_metaspades_assemblies.tar.gz"</p> <p>3) Assembly files from sub-sampling experiment: 3 samples (A6, D5, and H6) were deep sequenced with both PB Revio and ONT Promethion lsk114 R10.4.1 SUP400. Data was randomly subsampled to various depths (1, 5, 10, 20, 30, 40 Gbp, and all) and then assembled with both metaMDBG and metaFlye. <br>-A6: total of 26 files (PB was less than 40 Gbp thus no 40 Gbp subsampling depth)<br>-D5: total of 28 files<br>-H6: total of 28 files<br>-total: 82 files<br>"LR_ONTPB_sub_assemblies.tar.gz"</p> <p>4) Tar file of all circular contigs from assemblies. This will include 47 separate fasta files (1 for each participant and time point). This data was used for the viral and plasmid analyses. <br>"Hsap_circ_contigs_47_assemblies.tar.gz"<br><br>5) File containing all of the 985 cMAGs with annotations generated using prokka</p> <p>"985cMAGs_prokkaannotation.tar.gz</p> <p>6) Deep taxonomic profiling feature table: Pacbio samples (45) with features classified using the GTDB and 985 cMAGs custom database<br>"PB_985cMAG-sourmash_45s_2162f" #feature table<br>"mapping-file_PB45s_2162f_deID.txt" #mapping file, deidentified<br><br>7) tar file containing all genomes in the updated quality 'c986 cMAGs' <br>(>90% completeness, <5% contamination, contig=1<br><br>8) tar file containing all prokka annotation files associated with the 986 cMAGs<br><br>9) All supplemental tables or data files used in the revision</p>
AmphibiaWeb: AmphibiaWeb text w/traits based on Pensoft Annotator
from <p></p>http://amphibiaweb.org/. AmphibiaWeb is an online system enabling anyone with a Web browser to search and retrieve information relating to amphibian biology and conservation. This site was inspired by the global declines of amphibians, the study of which has been hindered by the lack of multidisplinary studies and a lack of coordination in monitoring, in field studies, and in lab studies. One of its major goals is to encourage a shared vision for the study of global amphibian declines and the conservation of remaining amphibians.<p></p><p></p>https://annotator.pensoft.net/about
LifeWatch observatory data: phytoplankton annotated trainingset by FlowCam imaging in the Belgian Part of the North Sea
<h1>Training dataset </h1> <p>The images were collected in the framework of the Belgian Lifewatch Research Infrastructure. During multidisciplinary campaigns, a number of fixed stations in the Belgian Part of the North Sea (BPNS) are visited on a monthly (onshore stations) or seasonal (offshore stations) basis. Samples are taken using a 55µm mesh size Apstein net and fixed in Lugol's iodine solution. In the lab, the samples are processed using a VS-4 FlowCAM model at 4X magnification targeting a particle size range of 55-300µm. The identification of the image data is done with the use of a CNN and followed by a manual validation step. Since May 2017, this dataset has provided micro- and phytoplankton observations, mainly covering diatoms, dinoflagellates and cilliates, for the Belgian Part of the North Sea (BPNS).</p> <p> </p> <p>This dataset comprises a trainings datasplit of 337,514 images distributed across 95 classes, with each class containing a minimum of 100 and a maximum of 10,000 images. The goal of this dataset is to be able to facilitate model training, here we have organized the data into a standard split, with 80% allocated for training, 10% for validation, and another 10% for testing purposes. This dataset structure ensures a balanced representation and supports scientific rigor in subsequent analyses.</p> <h1>Technical details </h1> <h2>Data preprocessing </h2> <p>Raw FlowCam output data is fully processed using in-house datapipelines, the VisualSpreadsheet software is only used for data acquisition during the lab run of the sample. Raw images and binary images are never saved during the FlowCam run, we only work on the image collages saved at the end of the run. Single images are cut from these collages using each image coordinates width and height pulled from the .lst file using in-house python code. The background of the images is not removed. These images are then predicted and annotated in-house at VLIZ.</p> <h2>Data splitting </h2> <p>The training dataset is 80% used for training, 10% for validation and 10% for prediction. </p> <h2>Classes, labels and annotations</h2> <p>The dataset comprises 337,514 images distributed across 95 classes, with each class containing a minimum of 100 and a maximum of 10,000 images. Taxonomic coverage of the dataset comprises mainly of diatoms, dinoflagellates and cilliates, but to a lesser extent also zooplankton and other protists.</p> <h2>Parameters </h2> <p>The images are read using cv2.imread and the values are used as parameters.</p> <p><strong>Metadata Parameter Descriptions:</strong></p> <ul> <li> <p><strong>image_path</strong>: The relative path to the image file showing the plankton or particle, usually structured by taxon or project folder.</p> </li> <li> <p><strong>sample_datetime</strong>: The exact date and time (in YYYY-MM-DD HH:MM:SS.sss format) when the sample was acquired using the imaging instrument.</p> </li> <li> <p><strong>flowcam_version</strong>: The software or hardware version of the FlowCAM instrument used to capture the image and process the sample.</p> </li> <li> <p><strong>station</strong>: The sampling location code where the image was taken. These station codes typically refer to predefined geographic or monitoring points in the field.</p> </li> <li> <p><strong>accepted_label</strong>: The final validated taxonomic label (e.g., genus or species) assigned to the organism or particle in the image, often based on expert review.</p> </li> <li> <p><strong>accepted_aphia_id</strong>: The unique identifier (AphiaID) corresponding to the accepted taxonomic label in the <strong>World Register of Marine Species (WoRMS)</strong> database, which ensures standardized taxonomic reference.</p> </li> <li> <p><strong>original_reference_id</strong>: A unique identifier (often a UUID) assigned to the image or sample in the original classification system (e.g., EcoTaxa), useful for traceability and linking back to the source record.</p> </li> </ul> <h2>Data sources </h2> <p>Images are collected during the monthly monitoring of phytoplankton communities in the Belgian Part of the North Sea during the LifeWatch multidisciplinary campaigns by FlowCam VS-4 benchmodel (Fluid Imaging Technologies, Yarmouth, Maine, U.S.A.).</p> <h2>Data quality </h2> <p>All images are predicted and subsequently manually validated to ensure the quality of the trainingset.</p> <h2>Image resolution </h2> <p>The size range imaged is 55-300µm. Images are acquired using a Sony XCD SC90 digital gray-scale camera. Images are during training of CNN resized to 100px by 100px.</p> <h2>Spatial coverage </h2> <p>The data comes from a number of fixed stations in the Belgian Part of the North Sea (BPNS). </p> <p>Nine stations onshore are visited monthly:</p> <table> <tbody> <tr> <td><strong>Station</strong></td> <td><strong>Longitude</strong></td> <td><strong>Latitude</strong></td> </tr> <tr> <td>130</td> <td>2.90535</td> <td>51.27055</td> </tr> <tr> <td>780</td> <td>3.057283</td> <td>51.471367</td> </tr> <tr> <td>330</td> <td>2.809083</td> <td>51.434117</td> </tr> <tr> <td>230</td> <td>2.85035</td> <td>51.308683</td> </tr> <tr> <td>710</td> <td>3.138283</td> <td>51.441217</td> </tr> <tr> <td>215</td> <td>2.61075</td> <td>51.274867</td> </tr> <tr> <td>ZG02</td> <td>2.500717</td> <td>51.33515</td> </tr> <tr> <td>120</td> <td>2.702483</td> <td>51.186083</td> </tr> <tr> <td>700</td> <td>3.221017</td> <td>51.377</td> </tr> </tbody> </table> <p>Eight additional offshore stations are visited seasonally:</p> <table> <tbody> <tr> <td><strong>Station</strong></td> <td><strong>Longitude</strong></td> <td><strong>Latitude</strong></td> </tr> <tr> <td>LW01</td> <td>2.256</td> <td>51.568667</td> </tr> <tr> <td>LW02</td> <td>2.556</td> <td>51.8</td> </tr> <tr> <td>435</td> <td>2.790333</td> <td>51.580667</td> </tr> <tr> <td>W07bis</td> <td>3.012517</td> <td>51.588033</td> </tr> <tr> <td>W08</td> <td>2.35</td> <td>51.458333</td> </tr> <tr> <td>W09</td> <td>2.7</td> <td>51.75</td> </tr> <tr> <td>W10</td> <td>2.416667</td> <td>51.683333</td> </tr> <tr> <td>421</td> <td>2.45</td> <td>51.4805</td> </tr> </tbody> </table> <h2> </h2> <h2>Temporal coverage </h2> <p>The monitoring was initiated in May 2017 and has been running continuously every month.</p> <h2>Contact information </h2> <p>For technical questions about training dataset, you can contact wout.decrop@vliz.be. </p> <p> </p> <p> </p> <p> </p> <p> </p>
RRID Gold Set of annotations for software tools in the biomedical literature
<p>This data is a subset of a larger human curated, machine assisted gold standard data set of RRID citations within the text of the scientific literature. The full set is accessible via Hypothes.is at https://hypothes.is/users/scibot?q=group%3A__world__ and via individual RRID records such as RRID:SCR_016250 https://scicrunch.org/resolver/SCR_016250/mentions?q=&i=rrid:scr_016250 </p><p>The data is based on authors who added RRIDs into their manuscripts. The data was then extracted by SciBot (RRID:SCR_016250), into Hypothes.is (RRID:SCR_000430) and then manually checked by a curator to determine if the author and the database agreed. The list of annotators is available in Hypothes.is user group: SciBotCurationGroup.</p><p>There are 78,140 rows and each row contains an annotation (annotation id, URI), linked to a paper (paper identifiers: PMID, DOI, PMC) and linked to the RRID (scr_id, exact, text_quote_selector). </p><p>A second spreadsheet contains a list of 8,322 software tools from the SciCrunch Registry (available here https://scicrunch.org/resources/data/source/nlx_144509-1/search ), enhanced by additions by thousands of individual authors, and curated over 10 years (Ozyurt et al., 2016). </p><p>The third spreadsheet contains a data dictionary and links to related ontologies, and tagging sets. </p><p> </p><p> </p>
1,237 Annotated Developer Apologies from GitHub
<p><em><strong>Software Developer Apologies</strong></em></p> <p>This dataset contains 1,237 GitHub comments with apology annotations (apology vs not apology), released as part of the following publication:</p> <ul> <li>Benjamin S. Meyers. <a href="https://scholarworks.rit.edu/theses/11609/">Human Error Assessment in Software Engineering.</a> Rochester Institute of Technology. 2023. </li> </ul> <p><em><strong>Included Files</strong></em></p> <p>The "github_apologies.csv" file contains the full dataset of 1,237 GitHub comments with apology annotations. In total, there are 365 comments containing an apology (872 non apologies). The comments themselves are a subset of those included in <a href="../records/5603093">88.6 Million Developer Comments from GitHub</a>.</p> <p><em><strong>Annotation Details</strong></em></p> <p>Full details are provided in the above publication. We implemented a naive classifier (Precision: 41.7%, Recall: 99.7%, F1: 86.9%, Accuracy: 91.1%) using counts of apology lemmas. 91% of developer comments containing at least one apology lemma matched our manual annotations. Agreement between raters was almost perfect (Cohen's Kappa = 0.94).</p> <p><em><strong>CSV Fields</strong></em></p> <ul> <li><strong>ID</strong>: Unique identifier for the comment.</li> <li><strong>SOURCE</strong>: Whether this comment originates from a commit, issue, or pull request.</li> <li><strong>COMMENT_URL</strong>: The URL linking to the comment.</li> <li><strong>COMMENT_TEXT</strong>: The raw comment text.</li> <li><strong>NUM_APOLOGY_LEMMAS</strong>: The count of apology lemmas present in the comment.</li> <li><strong>CLASSIFIER_LABEL</strong>: The automatically assigned label ("Apology" or "Not Apology").</li> <li><strong>RATER_1_LABEL</strong>: The manually assigned label ("Apology" or "Not Apology") from Rater 1.</li> <li><strong>RATER_2_LABEL</strong>: The manually assigned label ("Apology" or "Not Apology") from Rater 2.</li> <li><strong>AGREED_LABEL</strong>: The agreed upon label ("Apology" or "Not Apology") after Rater 1 and Rater 12 resolved disagreements.</li> </ul> <p><em><strong>Contact</strong></em></p> <p>Please contact Benjamin S. Meyers (<a href="mailto:bsm9339@rit.edu">email</a>) with questions about this data and its collection.</p> <p><em><strong>Acknowledgments</strong></em></p> <p>Collection of this data has been sponsored in part by the National Science Foundation (grant 1922169), by the NSA Science of Security Lablet program (grant H98230-17-D-0080/2018-0438-02), and by a Department of Defense DARPA SBIR program (grant 140D63-19-C-0018).</p>
200 Annotated Developer Human Errors from GitHub
<p><em><strong>Software Engineers' Human Errors</strong></em></p> <p>This dataset contains 200 GitHub comments with manual human error annotations, released as part of the following publication:</p> <ul> <li>Benjamin S. Meyers. <a href="https://scholarworks.rit.edu/theses/11609/">Human Error Assessment in Software Engineering.</a> Rochester Institute of Technology. 2023.</li> </ul> <p><em><strong>Included Files</strong></em></p> <p>The "developer_human_errors.csv" file contains the full dataset of 200 software defect descriptions annotated with human error types (slips, lapses, mistakes) and T.H.E.S.E. categories.</p> <p><em><strong>CSV Fields</strong></em></p> <ul> <li><strong>ID</strong>: Unique identifier for the comment.</li> <li><strong>SOURCE</strong>: Whether this comment originates from a commit, issue, or pull request.</li> <li><strong>COMMENT_URL</strong>: The URL linking to the comment.</li> <li><strong>COMMENT_TEXT</strong>: The raw comment text.</li> <li><strong>HUMAN_ERROR_TYPE</strong>: Whether the software defect described is a slip, lapse, or mistake.</li> <li><strong>THESE_V4_ID</strong>: Manually assigned T.H.E.S.E. category with labels corresponding to Version 4 of T.H.E.S.E.</li> <li><strong>THESE_NAME</strong>: Name corresponding to manually assigned T.H.E.S.E. category.</li> </ul> <p><em><strong>Annotation Details</strong></em></p> <p>Human error types span slips, lapses, and mistakes from James Reason's Generic Error Modelling System (GEMS):</p> <ul> <li><strong>Slips</strong>: Failures of attention.</li> <li><strong>Lapses</strong>: Failures of memory.</li> <li><strong>Mistakes</strong>: Failures of planning.</li> </ul> <p>T.H.E.S.E. categories are summarized below:</p> <ul> <li>S01: Typos & Misspellings</li> <li>S02: Syntax Errors</li> <li>S03: Overlooking documented Information</li> <li>S04: Multitasking Errors</li> <li>S05: Hardware Interaction Errors</li> <li>S06: Overlooking Proposed Code Changes</li> <li>S07: Overlooking Existing Functionality</li> <li>S08: General Attentional Failure</li> <li>L01: Forgetting to Finish a Development Task</li> <li>L02: Forgetting to Fix a Defect</li> <li>L03: Forgetting to Remove Development Artifacts</li> <li>L04: Working with Outdated Source Code</li> <li>L05: Forgetting an Import Statement</li> <li>L06: Forgetting to Save Work</li> <li>L07: Forgetting Previous Development Discussion</li> <li>L08: General Memory Failure</li> <li>M01: Code Logic Errors</li> <li>M02: Incomplete Domain Knowledge</li> <li>M03: Wrong Assumption Errors</li> <li>M04: Internal Communication Errors</li> <li>M05: External Communication Errors</li> <li>M06: Solution Choice Errors</li> <li>M07: Time Management Errors</li> <li>M08: Inadequate Testing</li> <li>M09: Incorrect/Insufficient Configuration</li> <li>M10: Code Complexity Errors</li> <li>M11: Internationalization/String Encoding Errors</li> <li>M12: Inadequate Experience Errors</li> <li>M13: Insufficient Tooling Access Errors</li> <li>M14: Workflow Order Errors</li> <li>M15: General Planning Failure</li> </ul> <p><em><strong>Contact</strong></em></p> <p>Please contact Benjamin S. Meyers (<a href="mailto:bsm9339@rit.edu">email</a>) with questions about this data and its collection.</p> <p><em><strong>Acknowledgments</strong></em></p> <p>Collection of this data has been sponsored in part by the National Science Foundation (grant 1922169), by the NSA Science of Security Lablet program (grant H98230-17-D-0080/2018-0438-02), and by a Department of Defense DARPA SBIR program (grant 140D63-19-C-0018).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.