Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
Supplementary files for Machine learning for histological annotation and quantification of cortical layers
<div> <h2>Creators</h2> <ul> <li><a href="https://orcid.org/0009-0000-9093-9385">Meystre Julie</a></li> <li><a href="https://orcid.org/0000-0002-7100-3749">Olivier Burri</a></li> </ul> <h2>Contributors</h2> <ul> <li><a href="https://orcid.org/0009-0002-0029-7951">Jean Jacquemier</a></li> </ul> </div> <h2>Description</h2> <p>This dataset contains 7 <a href="https://qupath.github.io/">QuPath</a> projects. The raw data images linked to these projects and located in other Zenodo datasets need to be downloaded as well.</p> <p>The raw data contains images of 14 hemispheres from height animals.</p> <ul> <li>Nissl_1 : <ul> <li>animal 1413827 Right Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_2 : <ul> <li>animal 1413829 Right Hemisphere</li> <li>animal 1413828 Right Hemisphere</li> <li>animal 1413827 Left Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_3 : <ul> <li>animal 1413828 Left Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_4 : <ul> <li>animal 1443459 Right Hemisphere</li> <li>animal 1443460 Right Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_5 : <ul> <li>animal 1443459 Left Hemisphere</li> <li>animal 1443460 Left Hemisphere</li> </ul> </li> </ul> <ul> <li>Nissl_6 : <ul> <li>animal 1449920 Left Hemisphere</li> <li>animal 1449921 Left Hemisphere</li> <li>animal 1449921 Right Hemisphere</li> <li>animal 1449922 Left Hemisphere</li> <li>animal 1449922 Right Hemisphere</li> <li> </li> </ul> </li> <li>QuPath_LayerBoundaries_GroundTruth_20220927: <ul> <li>This is the QuPath project that contains S1HL layers annotations done by the experts and which have been used to trained the Random forest Machine Learning method for the S1HL brain classification. It contains some images from all the eight animals.</li> </ul> </li> </ul> <p> </p> <div> <h3>Animals</h3> <p>All animal procedures were approved by the Veterinary Authorities and the Cantonal Commission for Animal Experimentation of the Canton of Vaud, according to the Swiss animal protection laws, under license number VD3516.</p> <p>Outbred Wistar Han rats (Janvier Laboratories, France) were ordered with their litter aged eight postnatal days (P8). Dams were housed individually and allowed to raise their own litters until experimentation on male offspring aged fourteen days (P14; N=8 animals; N=3 litters). Animals were housed in standard plastic laboratory cages, with bedding, nesting material and paper tube and ad libitum access to food (SAFE 150 SP-25) and water, cleaned once per week, and kept on a twelve-hour light-dark schedule with lights turned on at 06:30 AM, in rooms under controlled humidity and temperature. The sample size here is greater than those reported in other open source atlases <a href="https://www.zotero.org/google-docs/?1dkN18">(“Allen Reference Atlas - Mouse,” n.d.; “The Rat Brain in Stereotaxic Coordinates - 7th Edition,” n.d.)</a>.</p> <h3>Sample preparation</h3> <p>On postnatal day fourteen, rats were transferred to the experimental room in the morning to acclimate. The described procedure was conducted within a consistent 3-hour window of the day (09:00-12:00). Initially, the rats were deeply anesthetized using pentobarbital (intraperitoneal dose of 150 mg/kg; concentration of 150 mg/ml). This was succeeded by transcardial perfusion with ice cold 0.1 M phosphate buffer (PB; pH 7.4), followed by cold 4% paraformaldehyde (PFA) in 0.1 M PB. Subsequently, the brain was carefully removed from the skull, postfixed at 4°C in 4% PFA overnight, and then rinsed in 0.1 M PB. The brains underwent a sequential storage process: first in a 15% sucrose solution (in 0.1 M PB) at 4°C for approximately 24 hours, followed by a 30% sucrose solution at 4°C for an additional 24 hours. The hemispheres were carefully divided along the midline, after which both right and left hemispheres were precisely sliced sagittally using a cryostat (Leica, VT-1200S) at 50 µm employing an approximate angle rotation of 4 ± 1 degrees along the anterior-posterior axis to optimize alignment with apical dendrites. These brain slices were stored in a cryoprotectant solution (30% v/v ethylene glycol; 30% m/v sucrose in 0.1 M PB) at -20°C, preserving them until immunohistochemistry assays were executed (within a maximum of two weeks from extraction to immunohistochemistry).</p> <p>In order to determine the cell densities in P14 rat, brain slices were immunostained using cresyl violet, a stain specifically targeting cell bodies, including the endoplasmic reticulum, also known as Nissl substance or Nissl bodies. Free-floating sections of 50 µm thickness were transferred from cryoprotectant into 0.1 M PB to thaw and eliminate any cryoprotectant remnants. Subsequently, they were transferred into 0.01 M PB to minimize salt residues before being meticulously mounted onto SuperFrost© glass slides (Thermo Fisher Scientific Inc., Gerhard Menzel B.V. & Co. KG, GE). This mounting was carried out while considering the brain’s orientation relative to the midline, from its external to internal regions. Slide-mounted sections were processed using an automated slide stainer Tissue-Tek® Prisma Plus (Sakura Finetek-Europe, NL). These sections were incubated for 6 minutes at room temperature (RT = 20°C) in a 0.5% cresyl violet solution in water (with pH adjusted to 2.85 using acetic acid), followed by a brief wash in tap water. The sections underwent dehydration through a series of ethanol concentrations (70%, 70%, 96%, 100%, 100%) with each step lasting one minute at RT. Subsequently cleared with two steps of xylene for one minute each at RT, and the sections were mounted using Pertex (Sakura Finetek-Europe, NL) before being cover-slipped using the automated glass coverslipper Tissue-Tek® Glas™ g2 (Sakura Finetek-Europe, NL). A meticulous assessment of the coloration was conducted and if the staining appeared faint, a repeat staining procedure was carried out.</p> <p><strong>Immunostained slides were scanned using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel. Each brain slice was entirely scanned. Subsequently, the digital images obtained were meticulously organized and subjected to analysis using the open-source software QuPath v0.3.2 <a href="https://www.zotero.org/google-docs/?jnVnIg">(Bankhead et al., 2017)</a>. </strong></p> <p> </p> <h2>Intructions</h2> <p>The projects contained in this dataset have been created with QuPath v0.3.2 but could be opened with new QuPath version.</p> <ol> <li>Download the dataset</li> <li>untar the tar balls included in this dataset</li> <li>install <a href="https://qupath.github.io">QuPath</a></li> <li>Open QuPath</li> <li>Open a project within QuPath (Files->Project...->Open Project...)</li> </ol> </div>
CLDF dataset accompanying Chacon's "Annotated Swadesh Wordlists for Northwest Arawakan Languages" from 2022
<p>Cite the source of the dataset as:</p> <blockquote> <p>Chacon, Thiago C. (2022): Annotated Swadesh wordlists for Northwest Arawakan languages. Leipzig: Max Planck Institute for Evolutionary Anthropology.</p> </blockquote>
CLDF dataset derived from Starostin's "Annotated Swadesh Wordlists for the Karen Group" from 2017
<p>Cite the source of the dataset as:</p> <blockquote> <p>Starostin, George S. (2017): Annotated Swadesh Wordlists for the Karen Group. Moscow: The Global Lexicostatistical Database.</p> </blockquote>
Hypertension - Florida Annotated Corpus for Translational Science (FACTS), Vital Sign Ontology Annotations
<p>Florida Annotated Corpus for Translational Science (FACTS), which currently consists of 20 case reports about hypertension annotated with Vital Sign Ontology (VSO) classes (version 2012-04-25). </p>
New annotation of the Lolium perenne genome described by Bryne et al, (2015)
<p>Annotation of the Lolium perenne genome described by Bryne <em>et al,</em> (2015, DOI: 10.1111/tpj.13037).</p> <p>To identify genic regions RNA sequencing data were aligned to the genome using Tophat (Tophat version: V2.0.11; Bowtie2 version: 2.1.0). Isoforms, of genes, were identified using Cufflinks (Version: 2.2.0). Open reading frames (ORF), were found using using program ORFpredictor (version: 3.0). Frame selection was assisted by BLASTX searching the proteomes of <em>Arabidopsis thaliana</em> (TAIR, version: 10), <em>Oryza sativa</em> (Ensembl)<em>, Gycine max </em>(Ensembl)<em>, Populus trichocarpa </em>(Ensembl)<em> and Manihot esculenta </em>(cassava, v4.1). The predicted CDS was back translated to annotate the GFF file created by Cufflinks for CDS using scripts kindly provided by Palmieri<em> et al.,</em> 2012 (doi: 10.1371/journal.pone.0046415). These results are included in the file LG_V2_full.gtf.</p> <p>Functional annoatation was using three sources. First, protein sequences were search against the <em>A. thaliana</em> proteome using BLASTP. Second, the proteins were search against the Swiss-Prot non-redundant protein database (<a href="http://www.uniprot.org/downloads">http://www.uniprot.org/downloads</a> downloaded 14/03/2016, UniProt Consortium, 2014), again using BLASTP. In the third step, the protein sequences were scanned against InterPro's signatures using InterProScan (Version: 5.16-55). These data are included in the file ALLXLOC.txt.</p>
SSIX BREXIT Twitter Annotated Data Set
<p><strong>SSIX BREXIT Gold Standard</strong></p> <p>This repository contains the BREXIT Twitter Gold Standard produced by the SSIX Project <a href="https://ssix-project.eu/">https://ssix-project.eu/</a>.</p> <p>Only a sample is available here, to rebuild the full dataset, follow the instructions on the SSIX Project code repository: </p> <p><a href="https://bitbucket.org/ssix-project/brexit-gold-standard">https://bitbucket.org/ssix-project/brexit-gold-standard</a></p>
Gold standard corpus, ontologies, and Entity-Quality ontology annotations for evolutionary phenotypes
<p>This data set includes a gold-standard corpus of evolutionary phenotype descriptions (in the form of character state descriptions pulled from a variety of phylogenetic systematics studies), and their corresponding expert-curated annotations with ontology terms in the form of Entity-Quality (EQ) statements. EQ annotatons allow machine-reasoning (through the semantics encoded in the requisite ontologies from which the ontology terms are drawn), and machine-reasoning in turn enables computing metrics for quantifying the semantic similarity between different phenotype descriptions as represented by their EQ annotations.</p> <p>Also included are the ontologies, and the human expert-generated and Semantic Charaparser (i.e., machine) generated EQ annotations used to assess Semantic Charaparser performance relative to inter-curator variation and to the effect of having access to external knowledge. The ontologies include those used as input, the "augmented" ontologies created by human curators in each experiment round, and the merged ontology used to maximize Semantic Charaparser's performance.</p> <p>The production of the gold standard corpus, annotation experiments, and evaluation of the results are described in detail in the following manuscript:</p> <blockquote> <p>Dahdul et al (2018) Annotation of phenotypes using ontologies: a Gold Standard for the training and evaluation of natural language processing systems. BioRxiv https://doi.org/10.1101/322156. Submitted to Database.</p> </blockquote> <p>The analysis code for evaluating the gold standard corpus (and the input data and ontologies for that) are available separately from the following:</p> <blockquote> <p>Manda et al (2018) Code and data for analysis of evolutionary phenotype ontology annotations and gold standard corpus. Zenodo. https://doi.org/10.5281/zenodo.1218010</p> </blockquote> <p>In comparison to the previous version (v1.0.0), this record includes a file of MD5 checksums of the Gold Standard data files. The data files themselves are unchanged.</p>
Test Phenology Annotations From Herbarium Specimen Dataset - Prunus serotina
<p>This is a test dataset of phenology annotations for the Black Cherry, P. serotina, used in the submitted manuscript:</p> <p>Brenskelle, L., B. Stucky, J. Deck, R. Walls, R. P. Guralnick [submitted]. Integrating herbarium specimen observations into global phenology data systems. Applications in Plant Sciences.</p>
RUSSE'2018: Human-Annotated Sense-Disambiguated Word Contexts for Russian
<p>This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the <a href="https://russe.nlpub.org/2018/wsi/">RUSSE'2018</a> shared task on Word Sense Induction and Disambiguation for the Russian language; part of the <em>bts-rnc</em> evaluation dataset. These sense identifiers are disambiguated as according to the sense inventory of the <a href="http://gramota.ru/slovari/info/bts/">Large Explanatory Dictionary of Russian</a>.</p> <p>The annotation is done on December 1, 2017, on the <a href="https://tolokanyandex.com/">Yandex.Toloka</a> crowdsourcing platform. In particular, 80 pre-annotated contexts are used for training the human annotators, 2562 contexts are annotated by humans such that each context was annotated by 9 different annotators. The annotation reliability is indicated by a high value of Krippendorff's α = 0.83. After the annotation, every context was additionally inspected (“curated”) by the organizers of the shared task.</p> <p>The following words are represented: <em>акция</em> (action / stock), <em>байка</em> (yarn / tale), <em>гвоздика</em> (carnation / nail), <em>гипербола</em> (hyperbole), <em>град</em> (avalanche), <em>гусеница</em> (grub), <em>домино</em> (domino), <em>кабачок</em> (marrow / pub), <em>капот</em> (hood), <em>карьер</em> (mine / career), <em>кок</em> (cook), <em>крона</em> (top / crown), <em>круп</em> (croup), <em>мандарин</em> (mandarine), <em>рок</em> (fate / rock), <em>слог</em> (syllable), <em>стопка</em> (glass, stack), <em>таз</em> (bowl), <em>такса</em> (rate / badger-dog), <em>шах</em> (shah / check).</p> <p>The following files are included in this dataset:</p> <ul> <li>Toloka assignments (training: <em>tasks-train.tsv</em>, annotation: <em>tasks-test.tsv</em>)</li> <li>Toloka output (non-aggregated: <em>assignments_01-12-2017.tsv.xz</em>, aggregated: <em>aggregated_results_pool_1036853__2017_12_01.tsv</em>)</li> <li>annotator agreement report (<em>agreement.txt</em>)</li> <li>curated report (<em>report-curated.tsv.xz</em> and a supplementary file <em>tasks-eval.tsv.xz</em>)</li> <li>the final aggregated dataset (<em>bts-rnc-crowd.tsv</em>)</li> </ul> <p>The <em>bts-rnc-crowd.tsv</em> file has the following format: <em>id</em>, <em>lemma</em>, <em>sense_id</em>, <em>left</em> hand side context, <em>word</em> form, <em>right</em> hand side context, list of <em>senses</em>. The encoding is UTF-8 and the line breaks are LF (UNIX).</p>
Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes
<p> </p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p> </p> <p> RELEASE MAG-2018/01<br> --------------------------------------</p> <p> </p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p> </p> <p>2. LOCATION</p> <p> </p> <p> MAGs sequences are available under ENA project PRJEB25176.</p> <p> </p> <p> ENA_accession Isolate<br> -------------------- --------------<br> ERZ494218 AM_0118<br> ERZ494219 AM_0219<br> ERZ494220 AM_0226<br> ERZ494221 AM_0228<br> ERZ494222 AM_0233<br> ERZ494223 AM_0240<br> ERZ494224 AM_0244<br> ERZ494225 AM_0256<br> ERZ494226 AM_0268<br> ERZ494227 AM_0275<br> ERZ494228 AM_0466<br> ERZ494229 AM_0507<br> ERZ494230 AM_0510<br> ERZ494231 AM_0519<br> ERZ494232 AM_0528<br> ERZ494233 AM_0546<br> ERZ494234 AM_0608<br> ERZ494235 AM_0615<br> ERZ494236 AM_0616<br> ERZ494237 AM_0619<br> ERZ494238 AM_0621<br> ERZ494239 AM_0630<br> ERZ494240 AM_0643<br> ERZ494241 AM_0729<br> ERZ494242 AM_0764<br> ERZ494243 AM_0832<br> ERZ494244 AM_0849<br> ERZ494245 AM_0854<br> ERZ494246 AM_0876<br> ERZ494247 AM_0902<br> ERZ494248 AM_0936<br> ERZ494249 AM_1003<br> ERZ494250 AM_1104<br> ERZ494251 AM_1111<br> ERZ494252 AM_1205<br> ERZ494253 AM_1312<br> ERZ494254 AM_1409<br> ERZ494255 AM_1503<br> ERZ494256 AM_1603<br> ERZ494257 AM_1606<br> ERZ494258 AM_1801<br> ERZ494259 AM_1811<br> ERZ494260 AM_2104<br> ERZ494261 AM_2116<br> ERZ494262 AM_2124<br> ERZ494263 AM_2202<br> ERZ494264 AM_2207<br> ERZ494265 AM_2208<br> ERZ494266 AM_2324<br> ERZ494267 AM_2502<br> ERZ494268 AM_2804<br> </p> <p>3. ACKNOWLEDGEMENTS<br> </p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of São Carlos, São Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), as well as, the spanish funding organ Consejo Superior de Investigaciones Científicas (CSIC).</p> <p>This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001.</p> <p> </p> <p>4. CONTACT INFORMATION</p> <p> Current curators:</p> <p> - Célio Dias Santos Júnior (celio.diasjunior@gmail.com)<br> - Flavio Henrique-Silva (dfhs@ufscar.br)<br> - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> </p> <p>5. COPYRIGHT NOTICE</p> <p> Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> Copyright (C) 2018 The AMnrGC consortium.</p> <p> This database is provided “as is” and without any warranty of any kind,<br> of openly available. You can redistribute and/or modify it<br> as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p> https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>
MiRoR11 - P2 - Annotated corpus for primary and reported outcomes extraction
<p>Annotated corpus for outcome extraction</p> <p>This folder contains 2 subfolders:<br> 1. Primary_outcomes - a corpus annotated for primary outcomes<br> The folder contains the following files:<br> po_sent_marked_p1_1000.txt - sentences 1 - 1000 of the annotated corpus, ConstruKT format; coordinated outcomes annotated as single entity<br> po_sent_marked_p2_1000.txt - sentences 1001 - 2000 of the annotated corpus, ConstruKT format; coordinated outcomes annotated as single entity</p> <p>po_sent_marked_col_p1.txt - sentences 1 - 1000 of the annotated corpus, tabulated format; coordinated outcomes annotated as single entity<br> po_sent_marked_col_p2.txt - sentences 1001 - 2000 of the annotated corpus, tabulated format; coordinated outcomes annotated as single entity</p> <p>po_sent_marked_col_p1_coord.txt - sentences 1 - 1000 of the annotated corpus, tabulated format; coordinated outcomes annotated as separate entities<br> po_sent_marked_col_p2_coord.txt - sentences 1001 - 2000 of the annotated corpus, tabulated format; coordinated outcomes annotated as separate entities</p> <p>Subfolders:<br> po - the corpus for 10-fold cross-validation (10 subfolders with train/dev/test sets); coordinated outcomes annotated as separate entities<br> po_coord - the corpus for 10-fold cross-validation (10 subfolders with train/dev/test sets); coordinated outcomes annotated as separate entities</p> <p>2. Reported_outcomes - a corpus annotated for reported outcomes<br> The corpus contains sentences from Results and Conclusions sections of the articles for which primary outcomes were annotated. The first part of reported outcomes corpus contains sentences from articles for the first half of the primary outcomes corpus (sentences 1 - 1000). The second part of reported outcomes corpus contains sentences from articles for the second half of the primary outcomes corpus (sentences 1001 - 2000).</p> <p>The folder contains the following files:<br> res_sent_marked_p1.txt - first part of the annotated corpus, ConstruKT format<br> res_sent_marked_p2.txt - second part of the annotated corpus, ConstruKT format</p> <p>res_sent_marked_p1_col.txt - first part of the annotated corpus, tabulated format<br> res_sent_marked_p2_col.txt - first part of the annotated corpus, tabulated format</p> <p>Subfolders:<br> rep - the corpus for 10-fold cross-validation (10 subfolders with train/dev/test sets)</p>
Annotation of Phytozome V12 protein plant sequences using the ragp pipeline for hydroxyproline-rich glycoprotein mining
<p>Hydroxyproline aware annotation of hydroxyproline-rich glycoprotein (HRGP) sequences was performed on sequence data from 62 plant proteomes obtained from Phytozome database (<a href="https://phytozome.jgi.doe.gov/pz/portal.html">https://phytozome.jgi.doe.gov/pz/portal.html</a>, version 12) using the ragp R package (<a href="https://github.com/missuse/ragp">https://github.com/missuse/ragp</a>, version 0.3.0.0001). </p> <p>In each archive a single comma separated value table (.csv) is present along with a README.txt file describing the contents of the corresponding .csv file. The archives are:</p> <p>- phytozome_V12.tar.gz - sequences from 62 plant proteomes (phytozome V12) with a total of 2797062 protein sequences.</p> <p>- phytozome_V12_phobius.tar.gz -<strong> </strong> Signal peptide prediction using Phobius (<a href="http://phobius.sbc.su.se/">http://phobius.sbc.su.se/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_signalp.tar.gz -<strong> </strong> Signal peptide prediction using SignalP 4.1 (<a href="http://www.cbs.dtu.dk/services/SignalP-4.1/">http://www.cbs.dtu.dk/services/SignalP-4.1/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_targetp.tar.gz - Signal peptide prediction using TargetP 1.1 (<a href="http://www.cbs.dtu.dk/services/TargetP/">http://www.cbs.dtu.dk/services/TargetP/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_predict_hyp.tar.gz - Probability of proline hydroxylation for each proline from 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1).</p> <p>- phytozome_V12_maab.tar.gz - Motif and amino acid bias (MAAB) classification of hydroxyproline-rich glycoproteins performed on 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1). The number of predicted hydroxyprolines in each sequence is also indicated (based on predictions provided in phytozome_V12_predict_hyp.tar.gz).</p> <p>- phytozome_V12_scan_ag.tar.gz. - Hydroxyproline aware arabinogalactan motif scan performed on 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1). Hydroxyproline predictions are provided in phytozome_V12_predict_hyp.tar.gz. </p> <p>- phytozome_V12_scan_ag_hmmscan.tar.gz - Detection of domains in a subset of protein sequences which were found to contain arabinogalactan motifs (a subset of phytozome_V12_scan_ag.tar.gz).</p> <p>The list of the 62 plant species is provided in phytozome_V12.tar.gz README.txt.</p> <p>For questions contact mdragicevic@ibiss.bg.ac.rs.</p> <p> </p> <p> </p>
ST131_4071_genome_assembly_annotation_files
<p>The genome assembly annotation files of 4,071 E. coli ST131 genomes (see Decano & Downing 2019).</p>
Words or terms? Saṃjñā annotated dataset
<p>These data were used for the study published in:</p> <p>Lugli, Ligeia. 2019. Words or terms? Models of terminology and the translation of Buddhist Sanskrit vocabulary. In Alice Collett (ed.) Buddhism and Translation: Historical and Contextual Perspectives, New York: SUNY.</p> <p>data include:<br> 1. concordance lines for saṃjñā used for the study mentioned above. The concordance lines have been exported from the Sketch Engine and come from an automatically segmented corpus (segmenter = Lugli's version 1). They have not been proofread and contain segmentation errors. <br> 2. csv file with Lugli's semantic annotation of the concordance lines for saṃjñā. The data was annotated by Ligeia Lugli in 2017; part of the data constitutes a much revised version of a dataset originally prepared by Roberto Garcia for the Buddhist Translators Workbench in 2016.<br> 3. a pre-publication version of the study</p> <p> </p> <p> </p> <p>The creation of these data was funded by the British Academy through a Newton International Fellowship; the research was conducted at King's College London.<br> </p>
SFB genomes and annotations
<p>This dataset contains sequence files for a Metagenome Assembled Genome (MAG) from human metagenomes, as well as 5 SFB reference genomes:</p> <ul> <li>GCF_000270205 Candidatus Arthromitus sp. SFB-mouse-Japan</li> <li>GCF_000283555 Candidatus Arthromitus sp. SFB-rat-Yit</li> <li>GCF_000284435 Candidatus Arthromitus sp. SFB-mouse-Yit</li> <li>GCF_000709435 Candidatus Arthromitus sp. SFB-mouse-NL</li> <li>GCF_001655775 Candidatus Arthromitus sp. SFB-turkey isolate UMNCA01</li> </ul> <p>The dataset consists of 8 gzipped tar archives. Here's brief summary of their contents:</p> <ul> <li><strong>sfb_abundance</strong>: Counts of mapped reads and normalized counts for each contig in 825 samples (see <strong>sfb_map</strong>) Files named 'raw_counts' are number of reads assigned to each contig while files named 'tpm' are counts normalized to Transcripts Per Million. The 'percontig' files show numbers per contig while raw_counts.tab and tpm.tab files have counts summed for each genome.</li> <li><strong>sfb_abundance.cds</strong>: Counts of mapped reads and normalized counts as above but only for reads mapping to protein-coding regions.</li> <li><strong>sfb_annotations</strong>: Annotation files, from running the prokka pipeline on the genomes and subsequently eggnog-mapper, pfam_scan and dbCAN.</li> <li><strong>sfb_checkm</strong>: Results from running 'checkm lineage_wf' on the genomes.</li> <li><strong>sfb_collated</strong>: Collated counts of annotations in each genome.</li> <li><strong>sfb_fastani</strong>: Results from running fastANI on the genomes, with subsequent clustering of genomes based on 75% overlap and 95% ANI.</li> <li><strong>sfb_gtdb</strong>: Results from the 'gtdbtk classify_wf' on the genomes. This shows how the genomes are classified against the <a href="https://gtdb.ecogenomic.org/">Genome Taxonomy Database</a> (release86).</li> <li><strong>sfb_gtdb_denovo</strong>: Phylogeny as created using the following command on the genomes.</li> </ul> <pre><code class="language-bash">gtdbtk de_novo_wf --bac120_ms --outgroup_taxon p__Patescibacteria -x .fna --cpus 20 --rnd_seed 123</code></pre> <ul> <li><strong>sfb_map</strong>: Results from mapping reads from 825 samples to the 6 genomes. Reads were aligned using bowtie2 with '--very-sensitive --no-unal' settings and '--score-min C,0,0' to only report reads aligning without mismatches.Output was sorted by position and duplicates removed using MarkDuplicates of the picard tools suite. The archive contains a single merged bam file ('sfb.bam') where each sample has been assigned a ReadGroup inferred from its file name. Note that this mapping step was performed to investigate the presence of the SFB MAG in other metagenomes and was not part of the actual binning step.</li> <li><strong>SFB.unoise.vsearch.tsv: </strong>Count table of amplified 16S sequence variants with one sample per column and one Amplicon Sequence Variant (ASV) per row. The sixth column shows the assigned taxonomy, and the seventh, the sequence. Total DNA was amplified with the universal bacterial 16S primer pair 341f-805r. Primer sequences and low quality bases were removed from the raw reads with Cutadapt. ASVs were picked using Unoise3 with standard parameters. Taxonomy was assigned by the SINA classifier, based on the SILVA database v132.</li> </ul>
The terrestrial carnivorous plant Utricularia reniformis sheds light on environmental and life-form genome plasticity: Annotation, Gene Ontology and raw data
<p><strong>Description:</strong> In this work, we deeply sequenced (genome and transcriptome of different organs), assembled, and analyzed the 311-Mbp genome of the terrestrial carnivorous plant <em>U. reniformis</em> (Lentibulariaceae). This project presents great importance to the understanding of genomic, evolutive and functional aspects of<em> U. reniformis</em>, which may, with the next-generation sequencing and computational biology approaches shed light to a better understanding not only for the biology and evolution of <em>Utricularia</em> genus, but also for other genera and lineages of the Lentibulariaceae family. Here we present all the raw data generated, including annotation and gene ontology files.</p> <p><strong>External Information</strong></p> <p><a href="https://genomevolution.org/coge/GenomeInfo.pl?gid=54799">Genome Browser</a> avaliable at CoGe Portal (https://genomevolution.org/coge/GenomeInfo.pl?gid=54799)</p> <p><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">GenBank </a><a href="http://https://www.ncbi.nlm.nih.gov/bioproject/290588">Bioproject</a> (https://www.ncbi.nlm.nih.gov/bioproject/290588) for raw genomic and transcriptomic reads</p> <p><a href="https://bv.fapesp.br/en/auxilios/84264/genomics-and-transcriptomics-of-utricularia-reniformis-lentibulariaceae-an-evolutive-and-function/">FAPESP grant website</a> contaning the project abstract and other information.</p> <p><strong>Papers published related to <em>Utricularia reniformis</em> genome</strong></p> <pre><strong>[1]</strong> Silva SR, Diaz YC, Penha HA, Pinheiro DG, Fernandes CC, Miranda VF, MichaelTP, Varani AM. <strong>The Chloroplast Genome of Utricularia reniformis Sheds Light on the Evolution of the ndh Gene Complex of Terrestrial Carnivorous Plants from the Lentibulariaceae Family</strong>. PLoS One. 2016 Oct 20;11(10):e0165176. doi:<strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/27764252">10.1371/journal.pone.0165176</a></strong>. </pre> <pre><strong>[2] </strong>Silva SR, Alvarenga DO, Aranguren Y, Penha HA, Fernandes CC, Pinheiro DG, Oliveira MT, Michael TP, Miranda VFO, Varani AM. <strong>The mitochondrial genome of the terrestrial carnivorous plant Utricularia reniformis (Lentibulariaceae): Structure, comparative analysis and evolutionary landmarks.</strong> PLoS One. 2017 Jul19;12(7):e0180484. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/28723946">10.1371/journal.pone.0180484</a></strong>.</pre> <pre><strong>[3] </strong>Silva SR, Moraes AP, Penha HA, Julião MHM, Domingues DS, Michael TP, Miranda VFO, Varani AM. <strong>The Terrestrial Carnivorous Plant Utricularia reniformis Sheds Light on Environmental and Life-Form Genome Plasticity.</strong> Int J Mol Sci. 2019 Dec 18;21(1). pii: E3. doi: <strong><a href="https://www.ncbi.nlm.nih.gov/pubmed/31861318">10.3390/ijms21010003</a></strong>.</pre> <p><strong>Acknowledgements</strong></p> <p>This work was supported by Sao Paulo Research Foundation FAPESP, Grant ID: [1325164-6]</p> <p> </p> <p><strong>---------------------------------------------------------</strong><br> <strong>FILES DESCRIPTION</strong><br> <strong>---------------------------------------------------------</strong><br> <br> ----------------<br> <strong>ANNOT-vFinal.sql: </strong>MySQL database containing all integrated annotation information of Urenif and Ugibba<br> ----------------<br> <strong>TABLE fields description</strong><br> gene_name gene name generated by EVidence Modeler + PASA<br> length gene lenght<br> status duplicate_gene_classifier status (0:singleton, 1:dispersed, 2:proximal, 3: tandem, 4:WGD)<br> product gene product <br> GOterms Blast2GO/OmicsBox GOterms<br> GO_mapping Blast2GO/OmicsBox GOterms derived from direct mapping (UniProt)<br> GO_annotation Blast2GO/OmicsBox annotated GOterms<br> GO_interpro Blast2GO/OmicsBox derived from InterProScan<br> EC Blast2GO/OmicsBox EC number<br> EC_name Blast2GO/OmicsBox enzyme name<br> NOG_annot EggNOG annotation description<br> NOG_EC EggNOG EC number<br> NOG_GO EggNOG GOterms<br> NOG_class EggNOG COG/KOG classfication<br> KEGG_Pathway EggNOG KEGG pathyways<br> KEGG_ko EggNOG KEGG ko<br> CAZy EggNOG CAZy enzymes<br> TAIR_gene Closest A. thaliana gene name (homologous) TAIR database lasted version<br> TAIR_annot Closest A. thaliana gene product (homologous) TAIR database lasted version <br> ortho MCL clustering among Vvinifera, Athaliana, and Slycopersicum (S:singleton, C: clustered, Y: shared)<br> ortho_two MCL clustering among Urenif and Ugibba (S:singleton, C: clustered, Y: shared)<br> -<br> -<br> ----------------<br> <strong>CEGs.zip </strong> 336 shared and concatenated CEGs from Urenif, U. gibba, Genlisea nigrocaulis, G. hispidula, G. aurea, G. pygmaea, and G. repens.<br> ----------------</p> <p><strong>ProcessRepeats_mod</strong> Modified version of RepeatMasker, ProcessRepeats script for detection of plant evolutionary lineages<br> ----------------</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia gibba</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Ugibba</strong><strong>-no-masked.fa </strong> Ugibba genome excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ugibba-softmasked.fa</strong> Ugibba genome RepeatMasker softmasked and excluding organellar genomes (provided by Lan et al., 2017)<br> <strong>Ug.collinearity </strong> MCScanX collinearity file<br> <strong>Ug-duplicates.txt</strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ug.gene_type </strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ug.tandem </strong> Ugibba tandem genes generated by MCScanX tool<br> <strong>Ugibba_annot.annot </strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim) <strong>Ugibba_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong><strong> </strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Ugibba</strong><strong>.cDNA</strong> Ugibba cDNAs fasta file<br> <strong>Ugibba</strong><strong>.CDS </strong> Ugibba CDSs fasta file<br> <strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Ugibba GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Ugibba</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Ugibba GFF3 file fully annotated (genes only)<br> <strong>Ugibba_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Ugibba_fasta.fasta</strong> Blast2GO/OmicsBox Ugibba fasta proteins containg annotation (product and GO terms)<br> <strong>ugibba_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>ugibba_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>ugibba_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)</p> <p><strong>Ugibba_GAF.txt</strong> GAF file<br> <strong>Ugibba</strong><strong>.gene</strong> Ugibba gene fasta file<br> <strong>Ugibba_GOstat.txt </strong> GOstat file<br> <strong>Ugibba</strong><strong>-PASA-assemblies.fasta </strong> Ugibba PASA assemblies<br> <strong>Ugibba</strong><strong>-PASA.stats </strong> Ugibba annotation STATS<br> <strong>Ugibba</strong><strong>.</strong><strong>prot</strong><strong> </strong> Ugibba protein fasta file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff </strong> Ugibba RepeatMasker gff file<br> <strong>Ugibba</strong><strong>-RepeatMasker.gff3 </strong> Ugibba RepeatMasker gff3 file<br> <strong>Ugibba</strong><strong>-RepeatMasker.tbl </strong> Ugibba RepeatMasker results<br> <strong>Ugibba</strong><strong>-RepeatMasker-v2.gff3</strong> Ugibba RepeatMasker gff3 second version file<br> <strong>Ugibba</strong><strong>-RNAseq-assembled.fasta </strong> Ugibba RNAseq assembled transcriptome (Trinity)<br> <strong>Ugibba_TEs_DANTE_2019.fa </strong> Ugibba TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Ugibba_WEGO.txt </strong> WEGO file</p> <p><strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> <em>Utricularia reniformis</em> files<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong><br> <strong>Urenif</strong><strong>-no-masked.fa </strong> Urenif genome excluding organellar genomes<br> <strong>Urenif</strong><strong>-</strong><strong>softmasked</strong><strong>.fa</strong> Urenif genome RepeatMasker softmasked and excluding organellar genomes<br> <strong>Ur.collinearity </strong> MCScanX collinearity file<br> <strong>Ur-duplicates.txt </strong> MCScanX duplicate_gene_classifier short report<br> <strong>Ur.gene_type</strong> MCScanX duplicate_gene_classifier full report<br> <strong>Ur.tandem</strong> Urenif tandem genes generated by MCScanX tool<br> <strong>Urenif_annot.annot</strong> Blast2GO/OmicsBox annotation file (eudicotyledons filtered and Viridiplantae GOSlim)<br> <strong>Urenif_annot-</strong><strong>noclean</strong><strong>.</strong><strong>annot</strong> Blast2GO/OmicsBox annotation file (not filtered)<br> <strong>Urenif</strong><strong>.cDNA</strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>.CDS </strong> Urenif cDNAs fasta file<br> <strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA-ANNOTATED.gff3</strong> Urenif GFF3 file fully annotated (including gene products and GO terms)</p> <p><strong>Urenif</strong><strong>-EVM.all-no-TEs-PASA.gff3</strong> Urenif GFF3 file fully annotated (genes only)<br> <strong>Urenif_export.txt</strong> Blast2GO/OmicsBox full exported table<br> <strong>Urenif_fasta.fasta</strong> Blast2GO/OmicsBox Urenif fasta proteins containg annotation (product and GO terms)<br> <strong>urenif_frozen_cleaned-validated.box</strong> Full Blast2GO/OmicsBox file</p> <p><strong>urenif_frozen.box</strong> Full Blast2GO/OmicsBox file (containing TEs genes annotation)</p> <p><strong>urenif_nogs_emapper_annotations.box</strong> Full Blast2GO/OmicsBox EggNOG file (containing TEs genes annotation)<br> <strong>Urenif_GAF.txt </strong> GAF file<br> <strong>Urenif</strong><strong>.gene</strong> Urenif gene fasta file<br> <strong>Urenif_GOStat.txt </strong> GOstat file<br> <strong>Urenif</strong><strong>-PASA-assemblies.fasta</strong> Urenif PASA assemblies<br> <strong>Urenif</strong><strong>-PASA.stats </strong> Urenif annotation STATS<br> <strong>Urenif</strong><strong>.</strong><strong>prot</strong><strong> </strong> Urenif protein fasta file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff </strong> Urenif RepeatMasker gff file<br> <strong>Urenif</strong><strong>-RepeatMasker.gff3 </strong> Urenif RepeatMasker gff3 file<br> <strong>Urenif</strong><strong>-RepeatMasker.tbl </strong> Urenif RepeatMasker results<br> <strong>Urenif</strong><strong>-RepeatMasker-v2.gff3 </strong> Urenif RepeatMasker gff3 second version file<br> <strong>Urenif</strong><strong>-RNAseq-assembled.fasta </strong> Urenif RNAseq assembled transcriptome (Trinity)<br> <strong>Urenif_TEs_DANTE_2019.fa </strong> Urenif TEs library, detected by REPET and annotated by PASTEC and DANTE<br> <strong>Urenif_WEGO.txt </strong> WEGO file<br> <strong>----------------------------------------------------------------------------------------------------------------------------------------------<br> ----------------------------------------------------------------------------------------------------------------------------------------------</strong></p>
CLDF Dataset derived from Zhivlov's "Annotated Swadesh wordlists for the Ob-Ugrian group" from 2011
<p>Cite the source of the dataset as:</p> <blockquote> <p>Zhivlov, M. (2011): Annotated Swadesh wordlists for the Ob-Ugrian group (Uralic family). The Global Lexicostatistical Database. Moscow: RGGU.</p> </blockquote>
Sardinops sagax genome assemblies and annotations
<p>Files included are the genome assemblies for each haplotype (hap 1 and hap 2) of Sardinops sagax and the corresponding annotation files for each haplotype.</p> <p> </p> <p> </p>
Annotation historische Semantik von mhd. ungehiure
<p>Der Datensatz enthält alle Daten, die als Grundlage für den Aufsatz Marion Darilek, Ungeheuerlich. Zur historischen Semantik des Monströsen am Beispiel der computergestützten Textannotation von mhd. <em>ungehiure</em>, in: <em>Euphorion</em> 118 (2024), erhoben wurden.</p> <p>Die Struktur des Datensatzes und der ZIP-Dateien ist in der txt-Datei "Dokumentation_Datensatz_Annotationen_ungehiure_DEU_ENG" auf Deutsch und Englisch erläuert.</p> <p> </p> <p>The dataset contains all data used as a basis for the article Marion Darilek, Ungeheuerlich. Zur historischen Semantik des Monströsen am Beispiel der computergestützten Textannotation von mhd. <em>ungehiure</em>, in: <em>Euphorion</em> 118 (2024).</p> <p>The structure of the data set and of the ZIP-files is explained in the txt file "Dokumentation_Datensatz_Annotationen_ungehiure_DEU_ENG" in German and English.</p>
Brightfield images of cells and spheroids in wells annotated with bounding boxes
<p>The images in this dataset show cells in different developmental stages upon forming spheroids. They can go through several developmental sages: Starting from cells, they turninto compacted objects and then into spheroids. Ultimately, they can die and disintegrate. The objects are annotated with bounding boxes that carry these respective labels. The dataset in its current form can be upload to an OMERO server using the omero-cli-transfer package - simply download the zip file, log into your omero server and use the `omero transfer unpack` command as shown on the <a href="https://github.com/ome/omero-cli-transfer">omero-cli-transfer documentation</a>.</p> <p>The dataset can be used to train object detection models such as a <a href="https://docs.ultralytics.com/models/yolov8/">yolo classifier </a>- the linked repository provides a tutorial on how to do so.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.