Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “annotation”
Cross-phyla protein annotation by structural prediction and alignment
<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is typically limited to a few model organisms. In non-model species, the sequence-based prediction of gene orthology can be used to infer protein identity, however this approach loses predictive power at longer evolutionary distances. Here we propose a workflow for protein annotation using structural similarity, exploiting the fact that similar protein structures often reflect homology and are more conserved than protein sequences.</p> <p><strong>Results:</strong> We propose a workflow of openly available tools for the functional annotation of proteins via structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with known homology in >90% cases, and annotates an additional 50% of the proteome beyond standard sequence-based methods. We uncover new functions for sponge cell types, including extensive FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes, proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>
Database of Annotated Core Arguments: English, Lao and Russian
<p>This database contains corpus examples of transitive clauses with annotated core arguments realized as syntactic subjects and objects (A and P) in English, Lao and Russian. The coding scheme was developed together with Alena Witzlack-Makarevich</p>
List of TEI rolename annotations in the ISicily EpiDoc corpus
<p>This CSV file details every instance of a 'roleName' tag in the I.Sicily (sicily.classics.ox.ac.uk) EpiDoc TEI files, reporting the ID number of the file in which it appears, and the value of the @type and @subtype attributes in each case - as such it serves as an index of roleName attestations in the I.Sicily dataset (also recoverable directly from the EpiDoc files). The file will be updated in future.</p>
Occurrence Record Dataset from "Annotated checklist of the bees of Bonaire, with a focus on host plants"
<p>This is the occurrence dataset created for the publication "Annotated checklist of the bees of Bonaire, with a focus on host plants" (<a href="https://natuurtijdschriften.nl/pub/1026875" target="_blank" rel="noopener">https://natuurtijdschriften.nl/pub/1026875</a>).</p> <p>Observation and specimen data were assembled for this dataset, with the majority of records obtained during the Bonaire Estafette Expeditie (BEE). All citizen science records from Observation.org and iNaturalist.org up to December 2023 have been critically reviewed.<br>A project was created (<a href="https://www.inaturalist.org/projects/flower-visitors-and-pollinators-of-the-caribbean" target="_blank" rel="noopener">Flower visitors and pollinators of the Caribbean</a>) to improve standardized data collecting of plant-pollinator interactions and on <a href="https://observation.org/">observation.org</a> the standardized fields for interactions were used.<br>Records from passive trapping methods are not included. All bees were either observed or collected by hand or insect net. The majority of specimens will be accessible in the collection of Naturalis Biodiversity Center (RMNH), Leiden (the Netherlands). A synoptic collection is retained at the University of Tartu Zoological Collections in Tartu, Estonia (TUZ).</p> <p>The occurrence dataset (Version 1.4 and later) is:</p> <ul> <li>conform Darwin Core (DwC): <a href="https://dwc.tdwg.org/terms/">https://dwc.tdwg.org/terms</a></li> <li>in the data format CSV (tab delimited values) and UTF-8 encoded</li> </ul> <p> </p> <p><strong>DwC terms (Column labels) used in the dataset with their description:</strong></p> <table> <tbody> <tr> <td><strong>Column label</strong></td> <td><strong>Column description</strong></td> </tr> <tr> <td>occurrenceID</td> <td>Unique identifier or URI (GUID) for each record, mainly unique URLs generated by the web-based data holder.</td> </tr> <tr> <td>catalogNumber</td> <td>Unique code derived from URI in occurrenceID. Each specimen bears a label with this identifier and multimedia are tagged with this identifier.</td> </tr> <tr> <td>recordNumber</td> <td>Sample field ID used to manage data of preserved specimen occurrence records.</td> </tr> <tr> <td>otherCatalogNumbers</td> <td>Other unique identifiers used on specimen labels, but not derived from an URI.</td> </tr> <tr> <td>scientificName</td> <td>The scientific name of the lowest taxonomic rank to which the individual(s) was identified.</td> </tr> <tr> <td>scientificNameAuthorship</td> <td>The author name and year of publication in accordance with ICZN rules.</td> </tr> <tr> <td>verbatimIdentification</td> <td>The original identification, including qualifiers if needed.</td> </tr> <tr> <td>individualCount</td> <td>The number of individuals present at the time of the occurrence.</td> </tr> <tr> <td>sex</td> <td>The sex of the individual(s). The values female, male or unknown are used, if a mixed group is observed multiple values are listed.</td> </tr> <tr> <td>lifeStage</td> <td>The life stage of the individual(s).</td> </tr> <tr> <td>basisOfRecord</td> <td>The specific nature of the data record at the time of the identification (e.g. PreservedSpecimen).</td> </tr> <tr> <td>identifiedBy</td> <td>The name of the person who made the identification in the field or based on collected evidence (e.g. specimen or photo).</td> </tr> <tr> <td>identificationQualifier</td> <td>In case the identification could be given only to a species group 'cf.' is recorded.</td> </tr> <tr> <td>dateIdentified</td> <td>The year when the identification was made.</td> </tr> <tr> <td>previousIdentifications</td> <td>The scientific name originally given to the observed or collected individual(s).</td> </tr> <tr> <td>order</td> <td>The name of the order (e.g. Hymenoptera).</td> </tr> <tr> <td>family</td> <td>The name of the family (e.g. Apidae).</td> </tr> <tr> <td>genus</td> <td>The name of the genus (e.g. Apis).</td> </tr> <tr> <td>subgenus</td> <td>The name of the subgenus (e.g. Apis).</td> </tr> <tr> <td>specificEpithet</td> <td>The name of the species, epithet as given in dwc:scientificName.</td> </tr> <tr> <td>taxonRank</td> <td>The taxonomic rank of the most specific name in dwc:scientificName.</td> </tr> <tr> <td>eventDate</td> <td>The date-time when the event was observed and recorded. The event date uses the ISO 8601-1:2019 standard, with the following formatting being used: format YYYY-MM-DD, or YYYY if only the year is known. If time of capture is known, then format is YYYY-MM-DDTHH:MM, with HH:MM the local time.</td> </tr> <tr> <td>year</td> <td>The year in which the event was observed and recorded.</td> </tr> <tr> <td>month</td> <td>The month in which the event was observed and recorded.</td> </tr> <tr> <td>day</td> <td>The day in which the event was observed and recorded.</td> </tr> <tr> <td>eventTime</td> <td>The time or interval during which the event occurred.</td> </tr> <tr> <td>samplingProtocol</td> <td>The name or description of the collecting or recording method used.</td> </tr> <tr> <td>behavior</td> <td>A description of the behavior shown by the individual(s) recorded in this occurrence.</td> </tr> <tr> <td>decimalLatitude</td> <td>The geographic latitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>decimalLongitude</td> <td>The geographic longitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>geodeticDatum</td> <td>The ellipsoid, geodetic datum, or spatial reference system (SRS) upon which the geographic coordinates given in dwc:decimalLatitude and dwc:decimalLongitude is based.</td> </tr> <tr> <td>verbatimLocality</td> <td>The original textual description of the place.</td> </tr> <tr> <td>island</td> <td>The name of the island.</td> </tr> <tr> <td>countryCode</td> <td>The standard ISO 3166-1 alpha-2 country code for the country.</td> </tr> <tr> <td>coordinateUncertaintyInMeters</td> <td> <p>The horizontal distance (in meters) from the given dwc:decimalLatitude and dwc:decimalLongitude describing the smallest circle containing the actual location, usually the EPE (Estimated Position Error) from the GPS device. The EPE is here measured as the horizontal position error in meters.</p> </td> </tr> <tr> <td>recordedBy</td> <td>A person, group, or organization observing and recording the occurrence.</td> </tr> <tr> <td>associatedTaxa</td> <td>The type of association and the scientific name of the host taxon is recorded that is associated/has relationship with the taxon in dwc:scientificName. The association/relationship is recorded using the format as in the following example: "floral host":"Lantana sp."</td> </tr> <tr> <td>occurrenceRemarks</td> <td>Comments or notes about the dwc:Occurrence.</td> </tr> <tr> <td>associatedSequences</td> <td>A list (concatenated and separated) of identifiers (publication, global unique identifier, URI) of genetic sequence information.</td> </tr> <tr> <td>typeStatus</td> <td>A list (concatenated and separated) of nomenclatural types (type status, typified scientific name, publication) applied to the subject.</td> </tr> <tr> <td>collectionCode</td> <td>The name, acronym, coden, or initialism identifying the collection or data set from which the record was derived.</td> </tr> <tr> <td>identificationRemarks</td> <td>Comments or notes about the identification.</td> </tr> <tr> <td>identificationReferences</td> <td>A reference or list of references (publication, global unique identifier, URI) used for the identification.</td> </tr> <tr> <td>nameAccordingTo</td> <td>A reference to the checklist or publication that was followed to record the name in dwc:scientificName.</td> </tr> <tr> <td>samplingEffort</td> <td>The amount of effort, expressed in minutes or hours, to obtain and record the occurrences.</td> </tr> <tr> <td>occurrenceStatus</td> <td>A statement about the presence or absence of a taxon during the time of an event.</td> </tr> <tr> <td>disposition</td> <td>The current state of a specimen with respect to a collection.</td> </tr> <tr> <td>language</td> <td>The language of the record using ISO 639-1 codes, e.g. en</td> </tr> </tbody> </table>
Software and suspect database for: "A large scale multi-laboratory suspect screening of pesticide metabolites in human biomonitoring: From tentative annotations to verified occurrences"
<p>This upload contains the pesticide suspect list aggregated among the laboratories of work package 16 of the HBM4EU (https://www.hbm4eu.eu) project for a large-scale pesticide suspect screening and the resolving search templates for each pesticide. Additionally, we provide the used software version of MetAlign applied in this screening.</p>
Rice straw degradation analysis with Kraken2/Bracken annotation
<p>Metadata and annotation of the reads obtained from the rice straw degradation process using Kraken2/Bracken.</p>
Mappings for "Developing a Scalable Annotation Method for Large Datasets That Enhances Alarms With Actionability Data to Increase Informativeness: Mixed Methods Approach"
<p>Studies identified false and non-actionnable alarms as a factor for alarm fatigue in intensive care units.</p> <p>To annotate patient alarms, and analyse the alarm situation in intensive care units, we conceptualized and performed data mappings related to airway management and medication interventions. The mappings were based on information retrieved from the patient data management system (PDMS) and clinical expertise. For the airway management mappings, we used additional resources such as ISO 19223:2019 or ventilator instruction manuals. The mappings do not include patient data.</p> <p>As the mappings are generic, they could be used in other contexts than alarm annotation and research.</p> <p><strong>1. Respiratory Management Mappings:</strong></p> <ul> <li>General tables summarizing the 1) categories based on ISO 19223:2019 to describe respiratory support therapies (RSTs), 2) defining the invasiveness level of a RST and 3) listing the abbreviations used in the mappings</li> <li> <p>Tables including PDMS entries for airway devices (ADs), ventilation devices (VDs), and ventilation modes (VMs)</p> </li> <li> <p>Mapping of AD entries (from the PDMS) to defined categories</p> </li> <li> <p>Mapping of VDs, VMs, and ADs to defined RSTs, including information on invasiveness</p> </li> <li> <p>Table specifying suitable ventilation parameters in the context of each RST</p> </li> </ul> <p><strong>2. Medication Mappings:</strong></p> <ul> <li> <p>General tables providing information on physiological alarm conditions (PACs), interventions, routes, and techniques of administration of interest</p> </li> <li> <p>Mapping of routes of administration to techniques of administration including PDMS entries</p> </li> <li> <p>Mapping of active ingredients (including SNOMED CT Fully Specified Names and Identifiers), related PDMS information, and routes and techniques of administration to defined PAC and interventions</p> </li> </ul>
Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets
<p>Murreviikko (literally 'Dialect week') is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko ('dialect week') and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, Häme, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni 'my book-my') are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata > pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id's, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union’s Digital Single Market directive (2019/790).</p>
DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus
<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150 bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package 'dammit' that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>
Effects of crown gall disease on natural microbiota of Vitis vinifera - genome annotations
<p>Young grapevines (Vitis vinifera) frequently die due to the crown gall (CG) disease induced by the plant pathogen Allorhizobium vitis (Rhizobiaceae). Virulent members of A. vitis harbour a tumor-inducing (Ti) plasmid and cause formation of CGs due to genes encoded on the T-DNA. Expression of the oncogenes by transformed host cells induce cell proliferation, metabolic and physiological changes. The CG produces opines uncommon to plants, which provide an important nutrient source for A. vitis harbouring opine catabolism enzymes. CGs host a defined bacterial community and the mechanisms establishing a CG-specific bacterial community are currently unknown. Thus, we were interested in whether genes homologous to those of the Ti-plasmid coexist in the genomes of the microbial species coexisting in CGs. We isolated eight bacterial strains from grapevine CGs, sequenced their genomes and tested their virulence and opine utilization ability in bioassays. In addition, the eight genome sequences were aligned to the sequences of a Ti-plasmid and seven published bacterial genomes, including closely related plant associated bacteria but not from CGs. Homologous genes for virulence and opine anabolism were only present in the virulent Rhizobiaceae. By contrast, homologs of the opine catabolism genes were present in all strains including the non-virulent members of the Rhizobiaceae and non-Rhizobiaceae, indicating horizontal gene transfer of the opine degradation cluster from virulent to non-virulent strains. These results along with those of the opine utilization assay support the important role of opine utilization for co-colonization of virulent and non-virulent bacteria in CGs, thereby shaping the CG community.</p> <p>This dataset contains the prokka annotations of the genomes as used in "Opportunistic bacteria of grapevine crown galls are equipped with the genomic repertoire for opine utilization"</p>
CODE-test: An annotated 12-lead ECG dataset
<pre># Annotated 12 lead ECG dataset Contain 827 ECG tracings from different patients, annotated by several cardiologists, residents and medical students. It is used as test set on the paper: "Automatic diagnosis of the 12-lead ECG using a deep neural network". https://www.nature.com/articles/s41467-020-15432-4. It contain annotations about 6 different ECGs abnormalities: - 1st degree AV block (1dAVb); - right bundle branch block (RBBB); - left bundle branch block (LBBB); - sinus bradycardia (SB); - atrial fibrillation (AF); and, - sinus tachycardia (ST). Companion python scripts are available in: https://github.com/antonior92/automatic-ecg-diagnosis -------- Citation ``` Ribeiro, A.H., Ribeiro, M.H., Paixão, G.M.M. et al. Automatic diagnosis of the 12-lead ECG using a deep neural network. Nat Commun 11, 1760 (2020). https://doi.org/10.1038/s41467-020-15432-4 ``` Bibtex: ``` @article{ribeiro_automatic_2020, title = {Automatic Diagnosis of the 12-Lead {{ECG}} Using a Deep Neural Network}, author = {Ribeiro, Ant{\^o}nio H. and Ribeiro, Manoel Horta and Paix{\~a}o, Gabriela M. M. and Oliveira, Derick M. and Gomes, Paulo R. and Canazart, J{\'e}ssica A. and Ferreira, Milton P. S. and Andersson, Carl R. and Macfarlane, Peter W. and Meira Jr., Wagner and Sch{\"o}n, Thomas B. and Ribeiro, Antonio Luiz P.}, year = {2020}, volume = {11}, pages = {1760}, doi = {https://doi.org/10.1038/s41467-020-15432-4}, journal = {Nature Communications}, number = {1} } ``` ----- ## Folder content: - `ecg_tracings.hdf5`: The HDF5 file containing a single dataset named `tracings`. This dataset is a `(827, 4096, 12)` tensor. The first dimension correspond to the 827 different exams from different patients; the second dimension correspond to the 4096 signal samples; the third dimension to the 12 different leads of the ECG exams in the following order: `{DI, DII, DIII, AVR, AVL, AVF, V1, V2, V3, V4, V5, V6}`. The signals are sampled at 400 Hz. Some signals originally have a duration of 10 seconds (10 * 400 = 4000 samples) and others of 7 seconds (7 * 400 = 2800 samples). In order to make them all have the same size (4096 samples) we fill them with zeros on both sizes. For instance, for a 7 seconds ECG signal with 2800 samples we include 648 samples at the beginning and 648 samples at the end, yielding 4096 samples that are them saved in the hdf5 dataset. All signal are represented as floating point numbers at the scale 1e-4V: so it should be multiplied by 1000 in order to obtain the signals in V. In python, one can read this file using the following sequence: ```python import h5py with h5py.File(args.tracings, "r") as f: x = np.array(f['tracings']) ``` - The file `attributes.csv` contain basic patient attributes: sex (M or F) and age. It contain 827 lines (plus the header). The i-th tracing in `ecg_tracings.hdf5` correspond to the i-th line. - `annotations/`: folder containing annotations csv format. Each csv file contain 827 lines (plus the header). The i-th line correspond to the i-th tracing in `ecg_tracings.hdf5` correspond to the in all csv files. The csv files all have 6 columns `1dAVb, RBBB, LBBB, SB, AF, ST` corresponding to weather the annotator have detect the abnormality in the ECG (`=1`) or not (`=0`). 1. `cardiologist[1,2].csv` contain annotations from two different cardiologist. 2. `gold_standard.csv` gold standard annotation for this test dataset. When the cardiologist 1 and cardiologist 2 agree, the common diagnosis was considered as gold standard. In cases where there was any disagreement, a third senior specialist, aware of the annotations from the other two, decided the diagnosis. 3. `dnn.csv` prediction from the deep neural network described in the paper. THe threshold is set in such way it maximizes the F1 score. 4. `cardiology_residents.csv` annotations from two 4th year cardiology residents (each annotated half of the dataset). 5. `emergency_residents.csv` annotations from two 3rd year emergency residents (each annotated half of the dataset). 6. `medical_students.csv` annotations from two 5th year medical students (each annotated half of the dataset). </pre>
Public metagenome datasets annotated using SingleM
<p>These data underlie the community profiles shown at <a href="https://sandpiper.qut.edu.au">https://sandpiper.qut.edu.au</a></p> <p> </p> <h2>Changelog</h2> <p>version 1.0.0</p> <ul> <li>Public metagenomes published before Feb 20, 2025 were analysed using SingleM pipe v0.18.3 (the default R220 metapackage), and then renewed using an R226 metapackage (S5.4.0.GTDB_r226.metapackage_20250331).</li> </ul> <p>version 0.3.0</p> <ul> <li>Update profiles to use GTDB R220, generated using SingleM renew v0.17.0.</li> </ul> <p>version 0.2.0</p> <ul> <li>Initial version. Created using a GTDB R214-based reference SingleM metapackage S3.2.1.GTDB_r214.metapackage_20231006 based on public datasets available Dec 15, 2021.</li> </ul>
Gene Annotations of 49 Bacillariophyta Genome Assemblies
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div> </div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div> </div> <div>To extract the dataset, execute the following command:</div> <div> </div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div> </div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div> </div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div> </div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>
scRNA-seq atlases for 3 Caenorhabditis species - annotated cell datasets
<p>Annotated datasets (monocle3 objects) of scRNA-seq data for <em>C. elegans</em>, <em>C. briggsae</em> and <em>C. tropicalis</em> L2 nematodes. The datasets are published together with the manuscript "Divergence in neuronal signaling pathways despite conserved neuronal identity among <em>Caenorhabditis</em> species".</p> <p><a href="https://doi.org/10.1016/j.cub.2025.05.036" target="_blank" rel="noopener">https://doi.org/10.1016/j.cub.2025.05.036</a></p> <p>Files deposited include cell datasets for all sequenced cells ("all_cds") and datasets for all cells annotated as neurons ("neu_cds"). </p> <p><em>C. elegans</em> strain - N2.</p> <p><em>C. briggsae</em> strain - AF16.</p> <p><em>C. tropicalis</em> strain - NIC203.</p>
Curlie Enhanced with LLM Annotations: Two Datasets for Advancing Homepage2Vec's Multilingual Website Classification
<h3>Advancing Homepage2Vec with LLM-Generated Datasets for Multilingual Website Classification</h3> <p>This dataset contains two subsets of labeled website data, specifically created to enhance the performance of Homepage2Vec, a multi-label model for website classification. The datasets were generated using Large Language Models (LLMs) to provide more accurate and diverse topic annotations for websites, addressing a limitation of existing Homepage2Vec training data.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>LLM-generated annotations:</strong> Both datasets feature website topic labels generated using LLMs, a novel approach to creating high-quality training data for website classification models.</li> <li><strong>Improved multi-label classification:</strong> Fine-tuning Homepage2Vec with these datasets has been shown to improve its macro F1 score from 38% to 43% evaluated on a human-labeled dataset, demonstrating their effectiveness in capturing a broader range of website topics.</li> <li><strong>Multilingual applicability:</strong> The datasets facilitate classification of websites in multiple languages, reflecting the inherent multilingual nature of Homepage2Vec.</li> </ul> <p><strong>Dataset Composition:</strong></p> <ul> <li><strong>curlie-gpt3.5-10k:</strong> 10,000 websites labeled using GPT-3.5, context 2 and 1-shot</li> <li><strong>curlie-gpt4-10k:</strong> 10,000 websites labeled using GPT-4, context 2 and zero-shot</li> </ul> <p><strong>Intended Use:</strong></p> <ul> <li>Fine-tuning and advancing Homepage2Vec or similar website classification models</li> <li>Research on LLM-generated datasets for text classification tasks</li> <li>Exploration of multilingual website classification</li> </ul> <p><strong>Additional Information:</strong></p> <ul> <li><strong>Project and report repository:</strong> https://github.com/CS-433/ml-project-2-mlp</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>This dataset was created as part of a project at EPFL's Data Science Lab (DLab) in collaboration with <a href="https://people.epfl.ch/robert.west">Prof. Robert West</a> and <a href="https://tizianopiccardi.github.io/" rel="nofollow">Tiziano Piccardi.</a></p>
VoroCrack3d: An annotated data set of 3d CT concrete images with synthetic crack structures
<p>VoroCrack3d is an annotated data set of 3d CT images of concrete with synthetic crack structures. Its main purpose is the training and testing of machine learning models for 3d crack segmentation. The data set comprises 1344 images together with their corresponding ground truths. The concrete backgrounds are cropped out sections of size 400x400x400 voxels of CT images of concrete. To this end, several different concrete samples were scanned (normal concrete (NC), high-performance concrete (HPC), ultra-high-performance concrete (UHPC), air pore concrete; without and with reinforcements (straight steel fibers, crimped steel fibers, hooked-end steel fibers, polypropylene fibers, fibers made of glass fiber-reinforced polymer). The original concrete images have a resolution between 2.8 and 106 micrometers.</p> <p>The crack structures are modeled via minimum-weight surfaces in Voronoi diagrams according to the paper</p> <p>[1] C. Jung, C. Redenbach, Crack Modeling via Minimum-Weight Surfaces in 3d Voronoi Diagrams, Journal of Mathematics in Industry, 13, 10 (2023). https://doi.org/10.1186/s13362-023-00138-1.</p> <p>The surfaces are discretized, dilated and superimposed on the concrete backgrounds.</p> <p>The data set offers a high variety regarding concrete types, noise levels and crack widths, shapes, regularity and branching. This makes it suitable for studying the generalizability and robustness of 3d crack segmentation methods.</p> <p>______________________________________________________________________________________________</p> <p>The folder 'data' contains seven subfolders, each containing the data generated from a specific concrete type (NC, HPC, air pore concrete, polypropylene fiber-reinforced concrete, steel fiber-reinforced concrete (straight, crimped and hooked-end steel fibers)).</p> <p>Each subfolder again contains four subfolders according to the point process model that was used for generating the 3d Voronoi diagrams. The point processes and Voronoi diagrams are restricted to windows of size 400x150x400. </p> <p>- 'hc': Hard core point process with 60% volume density and intensity 0.000025 obtained from force-biased sphere packing.<br>- 'matclust': Matérn cluster process with parent intensity 0.0002/50, offspring intensity 50 and cluster radius 20.<br>- 'ppp': Poisson point process with intensity 0.0002.<br>- 'ppp-scaled': Poisson point process with intensity 0.0002 (but inside 200x150x200 window). The resulting Voronoi diagram is stretched in x- and z- direction by a factor of 2.</p> <p>Each of these contains five subfolders: one for the 3d input images, two for the corresponding labels (ground truths; one with and one without pores/fibers), one for the input and label previews (slice z=200 for each of the images) and a misc folder containing the concrete background without crack and, if applicable, the pore/fiber segmentation image.</p> <p>The data itself then contains 48 images:<br>1a-1d: crack with up to seven branches; fixed crack width (~1 voxel).<br>2a-2d: crack with up to four branches; fixed crack width (~1 voxel).<br>3a-3d: crack with up to one branch; fixed crack width (~1 voxel).<br>4a-4d: crack with no branches; fixed crack width (~1 voxel).<br>5a-5d: crack with no branches; fixed crack width (~3 voxels).<br>6a-6d: crack with no branches; fixed crack width (~5 voxels).<br>7a-7d: crack with no branches; fixed crack width (~7 voxels).<br>8a-8d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.01);<br>9a-9d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.02);<br>10a-10d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.05);<br>11a-11d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.1);<br>12a-12d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.2);</p> <p>The names 'a'-'d' indicate level of added noise added to the image:<br>a: None.<br>b: Uniformly on [-sigma,sigma] <br>c: Uniformly on [-2*sigma,2*sigma] <br>d: Uniformly on [-4*sigma,4*sigma] <br>Negative values are mapped to 0. <br>For inputs of type int, noise values are rounded to the nearest integer.<br>(sigma = standard deviation of voxel greyvalues in image)</p> <p>Note that the grey values in the ground truths correspond to the local crack width. They can be thresholded to obtain binary masks.</p> <p>For more details, we refer to [1].</p>
3D Data Derivatives of Grotta di Fumane: GigaMesh-processed, Annotations and Segmentations
<p><strong>Overview:</strong></p> <p>This repository contains derivatives of the Open Access publication by Falcucci & Peresani [FP22]. Our derived dataset (n = 62) is used to demonstrate our segmentation algorithm [BHM23], as shown in [BLM22], [BLM23], [LBM23], and will serve as a benchmark dataset for future analyses. To date, and to the best of our knowledge, our dataset is the first dataset of annotated lithic artifacts. In addition to the annotated dataset, we will also provide the segmented [BLM23] and GigaMesh preprocessed datasets [Mar+10; MK13] (n = 732) in separate folders. </p> <p><strong>Repository description: </strong></p> <p>A detailed description of the data can be found in 3D_Data_Derivatives_of_GdF_overview.pdf.</p> <p>For information on the archaeological interpretation of the artifacts, please refer to the original data publication by Falcucci and Peresani (2022). In our publications, we have expanded the CSV file from Falcucci and Peresani (2022) to document the use of the extended dataset:</p> <ul> <li> <p>Annotated: All artifacts that are annotated are marked with a 1.</p> </li> <li> <p>GT_PLY: All artifacts that are annotated and included in this publication are referenced by their respective file, such as 31_gt_labels.ply.</p> </li> <li> <p>Bullenkamp_et_al_2022: Artifacts utilized in [BLM22] are marked with a 1 .</p> </li> <li> <p>Bullenkamp_et_al_2023: Artifacts utilized in [BLM23] are marked with a 1.</p> </li> <li> <p>Linsel_et_al_2023: Artifacts utilized in [LBM23] are marked with a 1.</p> </li> </ul>
Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)
<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a “91 eutherian mammals” multiple whole genome alignment in Ensembl. </p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells. This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p> </p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>
Timema genome sequences and annotations. Version 8.
<p>Genome sequence (fasta) files and annotation (gff) files for ten <em>Timema </em>species: <em>T. bartmani, T. cristinae, T. poppensis, T. californicum, T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and T. genevievae.</em><br> <br> Species are abbreviated as follows: Tbi = <em>T. bartmani</em>, Tce = <em>T. cristinae</em>, Tps = <em>T. poppensis</em>, Tcm = <em>T. californicum</em>, Tpa = <em>T. podura</em>, Tte = <em>T. tahoe</em>, Tms = <em>T. monikensis</em>, Tdi = <em>T. douglasi</em>, Tsi = <em>T. shepardi</em>, and Tge = <em>T. genevievae</em><br> </p> <p>For details of assembly and annotation see: <br> <br> Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, Z., Figuet, E., François, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540</p> <p> </p> <p><strong>File list:</strong><br> <br> Tbi_b3v08.fasta = T. bartmani genome sequence file<br> Tbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation file<br> Tce_b3v08.fasta = T. cristinae genome sequence file<br> Tce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation file<br> Tcm_b3v08.fasta = T. bartmani genome sequence file<br> Tcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation file<br> Tdi_b3v08.fasta = T. douglasi genome sequence file<br> Tdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation file<br> Tge_b3v08.fasta = T. genevievae genome sequence file<br> Tge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation file<br> Tms_b3v08.fasta = T. monikensis genome sequence file<br> Tms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation file<br> Tpa_b3v08.fasta = T. podura genome sequence file<br> Tpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation file<br> Tps_b3v08.fasta = T. poppensis genome sequence file<br> Tps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation file<br> Tsi_b3v08.fasta = T. shepardi genome sequence file<br> Tsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation file<br> Tte_b3v08.fasta = T. tahoe genome sequence file<br> Tte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file</p>
Annotations to direct and indirect image rotation estimation methods of orthopedic X-ray images
<p>The annotation file contains labels for AP wrist images of the MURA dataset on the center line of the radius bone. The annotations are stored in json format. For each annotated image file of the MURA dataset an entry is provided with the coordinates of the start and end point of the radius' center line.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.