Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,250
datasets available to search
ShareScore release 0.7.1
Dataset results
6,250 results for “Classification”
Fig. 8 in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera
Fig. 8. FESEM images of the test structure in Tubothalamea from the Jurassic of Gnaszyn, Poland. A. Ophthalmidium carinatum Pazdro, 1958, MWGUW → ZI/67/08/5.03; front views of the abraded test surface, showing the extrados and porcelain (A1, A2). B.?Cornuspira radiata (Terquem, 1886), MWGUW ZI/67/55/11; front view of the test surface (B1, B3); oblique view of the test cross section, showing the test as being entirely composed of needle-shaped crystallites (B2, B4). C. Planiinvoluta sp., MWGUW ZI/67/57/13; view of the inner test surface (C1); side view of the test cross section, showing irregular meshwork of needle-shaped crystallites (C2). Abbreviations: e, extrados; p, porcelain.
Fig. 7 in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera
Fig. 7. FESEM images of the test structure in Recent calcareous cemented agglutinated textulariid (Globothalamea; A, B) and miliolid (Tubothalamea; C) → foraminifers from Ronsard Bay, Western Australia. A. Textularia sp., MWGUW ZI/67/55/24, front view of the test, showing agglutinated grains and the calcareous nanogranular matrix (A1); details of test wall (A2, A3). B. Gaudryina sp., MWGUW ZI/67/61/16, front view of the test, showing agglutinated grains and the matrix (B1); details of nanogranular matrix (B2, B3). C. Quinqueloculina arenata Said, 1949, MWGUW ZI/67/57/02, front view of the test (C1); oblique cross-sectional view of test showing foreign particle partially embedded in the irregular meshwork of needle-shaped crystallites (C2). Abbreviations: g, foreign particle; m, calcareous matrix. Arrows indicate pores.
Figure 7 in Updates on the classification and numbers of marine fungi
Figure 7: Rhytidhysteron bruguierae. (A–C) Appearance of hysterothecia on host. (D, E) Vertical section through hysteriothecium. (F, G) Cells of peridium. (H, I) Pseudoparaphyses, asci, and ascospores. (I) Stained with lactophenol cotton blue. (J–O) Ascospores. Scale bars: (A) = 1 mm, (B–E) = 200 μm, (F–I) = 50 μm, (J–O) = 5 μm.
Figure 6 in Updates on the classification and numbers of marine fungi
Figure 6: Phylogram generated from maximum likelihood (ML) analysis based on combined LSU, ITS, and act1 sequence data representing Sarocladiaceae (Hypocreales). Thirty-seven strains are included in the combined analyses which comprised 1753 characters (504 characters for LSU, 478 characters for ITS, 771 characters for act1) after alignment. Acremonium variecolor strains CBS 130360 and FMR 11141 in Bionectriaceae (Hypocreales) were used as the outgroup taxa. The best scoring RAxML tree with a final likelihood value of −8526.836 is presented. The matrix had 435 distinct alignment patterns, 356 parsimony-informative, 73 singleton sites, and 1332 constant sites. Estimated base frequencies were as follows: A = 0.216, C = 0.298, G = 0.265, T = 0.221; substitution rates: AC = 2.09479, AG = 3.08268, AT = 2.09479, CG = 1.00000, CT = 10.05849, GT = 1.000000; gamma distribution shape parameter α = 0.574. Bootstrap support values for ML equal to or greater than 75 % are given above the nodes (left side). Bayesian posterior probabilities (BYPP) equal to or greater than 0.95 are given above the nodes (right side). Ex-type strains are in bold and newly generated sequences are in blue.
Figure 5 in Updates on the classification and numbers of marine fungi
Figure 5: Colony of Remispora submersa (MUM 20.48) in malt extract agar (MEA) for 7 days at 25 °C: (A) obverse, (B) reverse.
Figure 8 in Updates on the classification and numbers of marine fungi
Figure 8: Savoryella sarushimana. (A–C) Holoblastic conidiogenous cells. (D–J) Variously shaped conidia. Note the proliferating conidia in (D, H). Scale bar = 10 μm.
Figure 3 in Updates on the classification and numbers of marine fungi
Figure 3: Randomized axelerated maximum likelihood (RAxML) phylogenetic tree based on a combined analysis of the 28S and 18S rRNA sequence data. Bootstrap support values for ML (>70 %) are given above each branch. The newly transferred isolate is shown in blue. The tree is rooted to Ceratocystis adiposa CCFC212707.
Figure 4 in Updates on the classification and numbers of marine fungi
Figure 4: Monosporascus cannonballus. (A) Oil globules realised from squashed ascospore. (B) Ascus with single globose ascospore. (C) Thick-walled ascospore. (D) Ascus with three globose ascospores. (E, F) Released ascospores around the projecting neck of the ascoma. Scale bars: (A–D) = 20 μm, (E–F) = 300 μm.
Figure 2 in Updates on the classification and numbers of marine fungi
Figure 2: Microascus cinereus. (A–I) Asci. (J–N, P) Ascospores. (O, Q, R) Anamorphic Scopulariopsis cinerea of Microascus cinereus. Scale bars: (A–D, K, M, P–Q) = 20 μm, (E–J, L, N–O, R) = 10 μm.
Figure 1 in Updates on the classification and numbers of marine fungi
Figure 1: Microascus trigonosporus. (A–E) Perithecia with spore cirrus. (F) Young ascus. (G, H, O) Mature asci. (I–L) Ascospores. (M, N) Anamorphic Scopulariopsis trigonosporus of Microascus trigonosporus. Scale bars: (A–E) = 100 μm, (G, H, O) = 20 μm, (F, I–N) = 10 μm.
MARMICRODB database for taxonomic classification of (marine) metagenomes
<p><strong>UPDATE (April 2024):<br>Please note that the combined fasta file (MARMICRODB.faa.bz2) associated with version 1.0.0 was mistakenly uploaded from an early prototype of the database. </strong>As a result, some taxonomy identifiers appended to the sequence records are mismatched with the names.dmp, nodes.dmp, and the MARMICRODB_catalog.tsv files. <br><br><strong>Importantly, the taxon IDs in the FM-compressed Kaiju database are correct and valid for use with the names.dmp and nodes.dmp files. </strong>Any analysis using the compressed Kaiju database (MARMICRODB.fmi) from version 1.0.0 will remain valid. However, if you have used the combined fasta file (MARMICRODB.faa.bz2) to generate your own Kaiju database and you used the taxdump files from version 1.0.0 I cannot guarantee the validity of those results. It is <em><strong>strongly recommended</strong></em> that you rebuild the combined fasta file from scratch using the FTP links to the NCBI assemblies provided in the MARMICRODB_catalog.tsv file.</p> <p>The combined fasta file has been removed from this version of the MARMICRODB.</p> <p><strong>Disclaimer:<br></strong>MARMICRODB is optimized for metagenomic samples from the marine environment, in particular, planktonic microbes from the pelagic euphotic zone. We expect this database may also be useful for classifying other types of marine metagenomic samples (for example, mesopelagic, bathypelagic, or even benthic or marine host-associated), but it has not been tested for this purpose. The original purpose of this database was to quantify clades/ecotypes of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 in metagenomes from Tara Oceans Expedition and the GEOTRACES project. We carefully annotated and quality controlled genomes from these five groups, but the processing of the other marine taxa was automated and unsupervised. Taxonomy for other groups was copied over from the Genome Taxonomy Database version RS83 (GTDB) [19,39] and NCBI Taxonomy [23] so any inconsistencies in those databases at the time of access will be propagated to MARMICRODB. If the user’s goal is to focus on a particular organism/clade that we did not curate in the database then the user should spend some time curating those genomes (ie checking for contamination, dereplicating, building a genome phylogeny for custom taxonomy node assignment). Currently the custom taxonomy is hardcoded in the MARMICRODB.fmi index, but if users wish to modify MARMICRODB by adding or removing genomes, or reconfiguring taxonomic ranks the names.dmp and nodes.dmp files can easily be modified. However, the Kaiju index will need to be rebuilt, which will require a high-performance compute cluster.</p> <p><strong>Introduction:</strong><br>This sequence database (MARMICRODB) was introduced in the publication JW Becker, SL Hogle, K Rosendo, and SW Chisholm. 2019. Co-culture and biogeography of <em>Prochlorococcus</em> and SAR11. ISME J. doi:10.1038/s41396-019-0365-4. Please see the original publication and its associated supplementary material for the original description of this resource. </p> <p><strong>Motivation:</strong><br>We needed a reference database to annotate shotgun metagenomes from the Tara Oceans project [1] the <a href="https://www.geotraces.org/">GEOTRACES</a> cruises GA02, GA03, GA10, and GP13 and the HOT and BATS time series [2]. Our interests are primarily in quantifying and annotating the free-living, oligotrophic bacterial groups <em>Prochlorococcus</em>, <em>Pelagibacterales</em>/SAR11, SAR116, and SAR86 from these samples using the protein classifier tool Kaiju [3]. Kaiju’s sensitivity and classification accuracy depends on the composition of the reference database, and the highest sensitivity is achieved when the reference database contains a comprehensive representation of expected taxa from an environment/sample of interest. However, the speed of the algorithm decreases as the database size increases. Therefore, we aimed to create a reference database that maximized the representation of sequences from marine bacteria, archaea, and microbial eukaryotes while minimizing (but not excluding) the sequences from clinical, industrial, and terrestrial host-associated samples.</p> <p><strong>Results/Description:</strong><br>MARMICRODB consists of 56 million sequence non-redundant protein sequences from 18769 bacterial/archaeal/eukaryote genome and transcriptome bins and 7492 viral genomes optimized for use with the protein homology classifier Kaiju [3]. To ensure maximum representation of marine bacteria, archaea, and microbial eukaryotes, we included translated genes/transcripts from 5397 representative “specI” species clusters from the proGenomes database [4]; 113 transcriptomes from the Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP) [5]; 10509 metagenome assembled genomes from the Tara Oceans expedition [6,7], the Red Sea [8], the Baltic Sea [9], and other aquatic and terrestrial sources [10]; 994 isolate genomes from the Genomic Encyclopedia of Bacteria and Archaea [11]; 7492 viral genomes from NCBI RefSeq [12]; 786 bacterial and archaeal genomes from MarRef [13]; and 677 marine single cell genomes [14]. In order to annotate metagenomic reads at the clade/ecotype level (subspecies) for the focal taxa <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116, we generated custom MARMICRODB taxonomies based on curated genome phylogenies for each group. The curated phylogenies, Kaiju formatted Burrows-Wheeler index, translated genes, the custom taxonomy hierarchy, an <a href="https://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">interactive kronaplot of the taxonomic composition</a>, and scripts and instructions for how to use or rebuild the resource is available from <a href="https://doi.org/10.5281/zenodo.3520509">10.5281/zenodo.3520509</a>. </p> <p><strong>Methods:</strong><br>The curation and quality control of MARMICRODB single-cell, metagenome-assembled, and isolate genomes was performed as described in [15]. Briefly, we downloaded all MARMICRODB genomes as raw nucleotide assemblies from NCBI. We determined an initial genome taxonomy for these assemblies using checkM with the default lineage workflow [16]. All genome bins met the completion/contamination thresholds outlined in prior studies [7,17]. For single cell and metagenome-assembled genomes, especially those from Tara Oceans Mediterranean sea samples [18], we use the GTDB-Tk classification workflow [19] to verify the taxonomic fidelity of each genome bin. We then selected genomes with a checkM taxonomic assignment of <em>Prochlorococcus</em>, <em>Synechococcus</em>, SAR11/<em>Pelagibacterales</em>, SAR86, and SAR116 for further analysis and confirmed taxonomic assignment using blast matches to known Prochlorococcus/Synechococcus ITS sequences and by matching 16S sequences to the SILVA database [20]. To refine our estimates of completeness/contamination of <em>Prochlorococcus</em> genome bins we created a custom set of 730 single copy protein families (available from <a href="https://doi.org/10.5281/zenodo.3719132">10.5281/zenodo.3719132</a>) from closed, isolate <em>Prochlorococcus</em> genomes [21] for quality assessments with checkM. For <em>Synechococcus</em> we used the CheckM taxonomic-specific workflow with the genus <em>Synechococcus</em>. After the custom CheckM quality control, we excluded any genome bins from downstream analysis that had an estimated quality < 30, defined as %completeness – 5x %contamination resulting in 18769 genome/transcriptome bins. We predicted genes in the resulting genome bins using prodigal [22] and excluded protein sequences with lengths less than 20 and greater than 20000 amino acids, removed non-standard amino acid residues, and condensed redundant protein sequences to a single representative sequence to which we assigned a lowest common ancestor (LCA) taxonomy identifier from the NCBI taxonomy database [23]. The resulting protein sequences were compiled and used to build a Kaiju [3] search database. </p> <p>The above filtering criteria resulted in 605 <em>Prochlorococcus</em>, 96 <em>Synechococcus</em>, 186 SAR11/<em>Pelagibacterales</em>, 60 SAR86, and 59 SAR116 high-quality genome bins. We constructed a high-quality fixed reference phylogenetic tree for each taxonomic group based on genomes manually selected for completeness and the phylogenetic diversity. For example the <em>Prochlorococcus</em> and <em>Synechococcus</em> genomes for the fixed reference phylogeny are estimated > 90% complete, and SAR11 genomes are estimated > 70% complete. We created multiple sequence alignments of phylogenetically conserved genes from these genomes using the GTDB-Tk pipeline [19] with default settings. The pipeline identifies conserved proteins (120 bacterial proteins) and generates concatenated multi-protein alignments [17] from the genome assemblies using hmmalign from the hmmer software suite. We further filtered the resulting alignment columns using the bacterial and archaeal alignment masks from [17] (http://gtdb.ecogenomic.org/downloads). We removed columns represented by fewer than 50% of all taxa and/or columns with no single amino acid residue occurring at a frequency greater than 25%. We trimmed the alignments using trimal [24] with the automated -gappyout option to trim columns based on their gap distribution. We inferred reference phylogenies using multithreaded RAxML [25] with the GAMMA model of rate heterogeneity, empirically determined base frequencies, and the LG substitution model [26](PROTGAMMALGF). Branch support is based on 250 resampled bootstrap trees. This tree was then pruned to only allow a maximum average distance to the closest leaf (ADCL) of 0.003 to reduce the phylogenetic redundancy in the tree [27]. We then “placed” genomes that either did not pass completeness threshold or were considered phylogenetically redundant by ADCL within the fixed reference phylogeny for each group using pplacer [28] representing each placed genome as a pendant edge in the final tree. We then examined the resulting tree and manually selected clade/ecotype cutoffs to be as consistent as possible with clade definitions previously outlined for these groups [29–32]. We then gave clades from each taxonomic group custom taxonomic identifiers and we added these identifiers to the MARMICRODB Kaiju taxonomic hierarchy.</p> <p><strong>Software/databases used:</strong><br>checkM v1.0.11[16]<br>HMMERv3.1b2 (http://hmmer.org/)<br>prodigal v2.6.3 [22]<br>trimAl v1.4.rev22 [24]<br>AliView v1.18.1 [33] [34]<br>Phyx v0.1 [35]<br>RAxML v8.2.12 [36]<br>Pplacer v1.1alpha [28]<br>GTDB-Tk v0.1.3 [19]<br>Kaiju v1.6.0 [34]<br>GTDB RS83 (https://data.ace.uq.edu.au/public/gtdb/data/releases/release83/83.0/)<br>NCBI Taxonomy (accessed 2018-07-02) [23]<br>TIGRFAM v14.0 [37]<br>PFAM v31.0 [38]</p> <p><strong>Use example:</strong><br>Because we used custom taxonomic MARMICRODB users will find many reads assigned to non-standard NCBI taxonomy identifiers. However, these reads are easily parsable using the custom names.dmp and nodes.dmp files included with the database. We include a brief description of how to do this below.</p> <p>I typically run Kaiju like:</p> <pre><code>kaiju -z 20 -a greedy -e 5 -m 11 -s 65 -E 0.05 -x \ -t nodes.dmp -f MARMICRODB.fmi \ -i inputfile_R1.fastq.gz \ -j inputfile_R2.fastq.gz \ -o MYOUTPUT.kaiju</code></pre> <p>To obtain a parseable report that lists the custom taxonomic ranks from the nodes.dmp and names.dmp files run kaiju2krona on the output.</p> <pre><code>kaiju2krona -t nodes.dmp -n names.dmp -i MYOUTPUT.kaiju -o MYOUTPUT.kaiju.krona </code></pre> <p>This report shows counts assigned to each node in the custom taxonomy and will also include the names for each rank. You can easily parse this programmatically using a scripting language like python or by using unix utilities.</p> <p><strong>File descriptions:</strong></p> <p><em>MARMICRODB_catalog.tsv</em><br>Tabular file of NCBI assembly accessions and associated taxonomic information for every genome in MARMICRODB. Also includes literature references for each genome where available. Header description:</p> <ol> <li>genome: Unique identifier for each genome</li> <li>full_name: full organism name where available</li> <li>source: literature reference where available</li> <li>taxid: NCBI taxonomy ID for the assembly accession</li> <li>MARMICRODBtaxid: taxonomy ID used in the custom Kaiju database</li> <li>lineage_assignment: taxonomic lineage assignment from NCBI</li> <li>domain: archaea, bacteria, or eukaryote</li> <li>taxgroup: short descriptive group</li> <li>taxclade: higher resolution clade assignment where available</li> <li>habitat_source: whether genome derives from marine or aquatic source</li> <li>sequence_type: isolate, single cell genome (sag), metagenome assembled genome (mag), or transcriptome in case of eukaryotes</li> <li>assembly_ftp: NCBI ftp for assembly</li> <li>gbk_acc: assembly genbank or refseq accession number</li> <li>gtdb_taxonomy: taxonomic lineage assignment from GTDB-Tk v0.1.3 against GTDB v83</li> </ol> <p><em>MARMICRODB_kronaplot.html</em><br><a href="http://htmlpreview.github.io/?https://github.com/slhogle/MARMICRODB/blob/master/MARMICRODB_kronaplot.html">Interactive Kronaplot</a> for the exploration of taxonomic composition of MARMICRODB</p> <p><em>MARMICRODB.faa.bz2</em><br>Fasta file of all protein sequences in MARMICRODB</p> <p><em>scripts.tar.gz</em><br>directory containing scripts for generating Kaiju formatted database</p> <p><em>phylogenies.tar.gz</em><br>directory containing detailed phylogenies for SAR11, Prochlorococcus, SAR86, and SAR116</p> <p><em>MARMICRODB.fmi</em><br>Kaiju index for MARMICRODB</p> <p><em>nodes.dmp</em><br>nodes file for taxonomic assignment with Kaiju</p> <p><em>names.dmp</em><br>names file for generating Kaiju reports</p> <p><strong>References:</strong><br>1. Karsenti E, Acinas SG, Bork P, Bowler C, De Vargas C, Raes J, et al. A holistic approach to marine eco-systems biology. PLoS Biol. 2011;9: e1001177.<br>2. Biller SJ, Berube PM, Dooley K, Williams M, Satinsky BM, Hackl T, et al. Marine microbial metagenomes sampled across space and time. Scientific Data. 2018;5: 180176.<br>3. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>4. Mende DR, Letunic I, Huerta-Cepas J, Li SS, Forslund K, Sunagawa S, et al. proGenomes: a resource for consistent functional and taxonomic annotations of prokaryotic genomes. Nucleic Acids Res. 2017;45: D529–D534.<br>5. Keeling PJ, Burki F, Wilcox HM, Allam B, Allen EE, Amaral-Zettler LA, et al. The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): illuminating the functional diversity of eukaryotic life in the oceans through transcriptome sequencing. PLoS Biol. 2014;12: e1001889.<br>6. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>7. Tully BJ, Graham ED, Heidelberg JF. The reconstruction of 2,631 draft metagenome-assembled genomes from the global oceans. Sci Data. 2018;5: 170203.<br>8. Haroon MF, Thompson LR, Parks DH, Hugenholtz P, Stingl U. A catalogue of 136 microbial draft genomes from Red Sea metagenomes. Sci Data. 2016;3: 160050.<br>9. Hugerth LW, Larsson J, Alneberg J, Lindh MV, Legrand C, Pinhassi J, et al. Metagenome-assembled genomes uncover a global brackish microbiome. Genome Biol. 2015;16: 279.<br>10. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017.<br>11. Mukherjee S, Seshadri R, Varghese NJ, Eloe-Fadrosh EA, Meier-Kolthoff JP, Göker M, et al. 1,003 reference genomes of bacterial and archaeal isolates expand coverage of the tree of life. Nat Biotechnol. 2017;35: 676–683.<br>12. Haft DH, DiCuccio M, Badretdin A, Brover V, Chetvernin V, O’Neill K, et al. RefSeq: an update on prokaryotic genome annotation and curation. Nucleic Acids Res. 2018;46: D851–D860.<br>13. Klemetsen T, Raknes IA, Fu J, Agafonov A, Balasundaram SV, Tartari G, et al. The MAR databases: development and implementation of databases specific for marine metagenomics. Nucleic Acids Res. 2018;46: D692–D699.<br>14. Berube PM, Biller SJ, Hackl T, Hogle SL, Satinsky BM, Becker JW, et al. Single cell genomes of Prochlorococcus, Synechococcus, and sympatric microbes from diverse marine environments. Scientific Data. 2018;5: 180154.<br>15. Becker JW, Hogle SL, Rosendo K, Chisholm SW. Co-culture and biogeography of Prochlorococcus and SAR11. ISME J. 2019. doi:10.1038/s41396-019-0365-4<br>16. Parks DH, Imelfort M, Skennerton CT, Hugenholtz P, Tyson GW. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015;25: 1043–1055.<br>17. Parks DH, Rinke C, Chuvochina M, Chaumeil P-A, Woodcroft BJ, Evans PN, et al. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol. 2017. doi:10.1038/s41564-017-0012-7<br>18. Tully BJ, Sachdeva R, Graham ED, Heidelberg JF. 290 metagenome-assembled genomes from the Mediterranean Sea: a resource for marine microbiology. PeerJ. 2017;5: e3558.<br>19. Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics. 2019. doi:10.1093/bioinformatics/btz848<br>20. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013;41: D590–6.<br>21. Biller SJ, Berube PM, Berta-Thompson JW, Kelly L, Roggensack SE, Awad L, et al. Genomes of diverse isolates of the marine cyanobacterium Prochlorococcus. Sci Data. 2014;1: 140034.<br>22. Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics. 2010;11: 119.<br>23. Federhen S. The NCBI Taxonomy database. Nucleic Acids Res. 2012;40: D136–43.<br>24. Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009;25: 1972–1973.<br>25. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>26. Le SQ, Gascuel O. An improved general amino acid replacement matrix. Mol Biol Evol. 2008;25: 1307–1320.<br>27. Matsen FA, Gallagher A, McCoy C. Minimizing the average distance to a closest leaf in a phylogenetic tree. arXiv [q-bio.PE]. 2012. Available: http://arxiv.org/abs/1205.6867<br>28. Matsen F a., Kodner RB, Armbrust EV. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree. BMC Bioinformatics. 2010;11: 538.<br>29. Biller SJ, Berube PM, Lindell D, Chisholm SW. Prochlorococcus: the structure and function of collective diversity. Nat Rev Microbiol. 2014;13: 13–27.<br>30. Giovannoni SJ. SAR11 Bacteria: The Most Abundant Plankton in the Oceans. Ann Rev Mar Sci. 2016.<br>31. Dupont CL, Rusch DB, Yooseph S, Lombardo M-J, Alexander Richter R, Valas R, et al. Genomic insights to SAR86, an abundant and uncultivated marine bacterial lineage. ISME J. 2012;6: 1186–1199.<br>32. Yang S-J, Kang I, Cho J-C. Expansion of Cultured Bacterial Diversity by Large-Scale Dilution-to-Extinction Culturing from a Single Seawater Sample. Microb Ecol. 2016;71: 29–43.<br>33. Larsson A. AliView: a fast and lightweight alignment viewer and editor for large datasets. Bioinformatics. 2014;30: 3276–3278.<br>34. Menzel P, Ng KL, Krogh A. Fast and sensitive taxonomic classification for metagenomics with Kaiju. Nat Commun. 2016;7: 11257.<br>35. Brown JW, Walker JF, Smith SA. Phyx: phylogenetic tools for unix. Bioinformatics. 2017;33: 1886–1888.<br>36. Stamatakis A. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics. 2006;22: 2688–2690.<br>37. Haft DH, Selengut JD, Richter RA, Harkins D, Basu MK, Beck E. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 2013;41: D387–95.<br>38. Finn RD, Coggill P, Eberhardt RY, Eddy SR, Mistry J, Mitchell AL, et al. The Pfam protein families database: towards a more sustainable future. Nucleic Acids Res. 2016;44: D279–85.<br>39. Parks DH, Chuvochina M, Waite DW, Rinke C, Skarshewski A, Chaumeil P-A, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol. 2018.<br> </p>
FIG. 10. Strophocaulon unitum, Fawcett 641 in A Generic Classification of the Thelypteridaceae
FIG. 10. Strophocaulon unitum, Fawcett 641 (VT), A. pinna lobes, abaxial view. B. rhizome. C. laminar apex.
FIG. 9. Steiropteris deltoidea, Fawcett 464 in A Generic Classification of the Thelypteridaceae
FIG. 9. Steiropteris deltoidea, Fawcett 464 (VT) A. habit. B. pinna lobes, adaxial view.C. Mesophlebion sp., Karger 1796 (VT), pinna lobes, abaxial view.
FIG. 6. A. Plesioneuron hopeanum, Fawcett 613 in A Generic Classification of the Thelypteridaceae
FIG. 6. A. Plesioneuron hopeanum, Fawcett 613 (VT), pinna lobes, adaxial view. Amauropelta (subg. Nibaa) noveboracensis, Fawcett 567 (MICH, VT), B. pinna-lobes, adaxial view.C. habit. D. Thelypteris palustris, Fawcett 569 (VT), pinna lobes abaxial view.
FIG. 3 in A Generic Classification of the Thelypteridaceae
FIG. 3. Grypothrix cuspidata, Lu 12803 (VT), A. habit.B. pinna, abaxial view. C. G. simplex, Lu 9270 (VT), habit.
FIG. 2 in A Generic Classification of the Thelypteridaceae
FIG. 2. Morphology of the Thelypteridaceae. A. Coryphopteris kolombangarae, SITW07554 (TAIF), glandular indusia. B. Goniopteris yaucoensis, Proctor 43584 (US), stellate hairs along adaxial costa.C.Pneumatopteris glandulifera, SITW11069 (TAIF),peg-like aerophores at pinna bases.D. Sphaerostephanos polycarpos,SITW11070 (TAIF), marginally glandular indusia.E. Sphaerostephanos doodioides, SITW05702 (TAIF), setose indusia,spherical glands on indusia, sporangia,and laminar tissue. F. Goniopteris reptans, Wright 813 (US), sori with shriveled, setose indusia, some hairs furcate. G. Macrothelypteris polypodioides, SITW11122 (TAIF), stipe scales. H. Hoiokula pendens, Hobdy 2664 (UC), setulose sporangia. I. Plesioneuron imbricatum, SITW05370 (TAIF), exindusiate sori, sporangia with spherical glands. J. Christella dentata, SITW10525 (TAIF), shriveled, setulose indusia. K. Sphaerostephanos heterocarpos SITW11079 (TAIF), abruptly reduced proximal pinnae, adaxial surface with acicular hairs. L. Coryphopteris subbipinnata, SITW11667 (TAIF), appressed stipe base scales,these marginally setulose. All photos by Cheng-Wei Chen except B, F, and H by Susan Fawcett.
FIG. 4. A. Hoiokula sandwicensis, Aborn s.n in A Generic Classification of the Thelypteridaceae
FIG. 4. A. Hoiokula sandwicensis, Aborn s.n. (VT),pinna lobes, abaxial view.B. Hoiokula pendens, Hobdy 2664 (UC), pinna lobes, adaxial view.C. Leptogramma pilosa, Diaz 6621 (UC), abaxial pinna lobes. D. Stegnogramma aspidioides, Palmer 1035 (UC), abaxial pinna. E. Stegnogramma wilfordii, Boufford 20159 (VT), proximal portion of lamina, abaxial view.
Coryphopteris subbipinnata, Mount Chaunapaho, Guadalcanal, Solomon Islands. Photo by Cheng-Wei Chen. in A Generic Classification of the Thelypteridaceae
Coryphopteris subbipinnata, Mount Chaunapaho, Guadalcanal, Solomon Islands. Photo by Cheng-Wei Chen.
FIG. 1. A in A Generic Classification of the Thelypteridaceae
FIG. 1. A synoptical tree illustrating relationships among genera of Thelypteridaceae based on Fawcett et al. (in press).
Real-World Signed Graphs Annotated for Whole Graph Classification
<p><strong>Description: </strong>this corpus was designed as an experimental benchmark for a task of signed graph classification. It is composed of three datasets derived from external sources and adapted to our needs:</p> <ul> <li><strong>SpaceOrigin Conversations [1]: </strong>set of conversational graphs, each one associated to a situation of verbal abuse vs. normal situation. These conversations model interactions happening in chatrooms hosted by an MMORPG/ The graphs were originally unsigned: we attributed signed to the edges based on the polarity of the exchanged messages. </li> <li><strong>Correlation Clustering Instances [2]: </strong>set of graph generated randomly as instances of the Correlation Clustering problem, which consists in partitioning signed graphs. These graphs are not associated in any class in the original paper. We proposed a class based on certain features of the space of optimal solutions explored in [2].</li> <li><strong>European Parliament Roll-Calls [3]: </strong>vote networks extracted from the activity of French Members of the European Parliament. The original data does not have any class associated to the networks: we proposed one based on the number of political factions identified in each network in [3]. </li> </ul> <p>These data were used in [4] in order to train and assess various representation learning methods. The authors proposed Signed Graph2vec, a signed variant of Graph2vec; WSGCN, a whole-graph variant of Signed Graph Convolutional Networks (SGCN), and use an aggregated version of Signed Network Embeddings (SiNE) as a baseline. The article provides more information regarding the properties of the datasets, and how they were constituted.</p> <p><strong>Software: </strong>the software used to train the representation learning methods and classifiers is publicly available online: <a href="https://github.com/CompNet/SWGE">SWGE</a>.</p> <p><strong>References:</strong></p> <ol> <li>Papegnies, É.; Labatut, V.; Dufour, R. & Linarès, G. Conversational Networks for Automatic Online Moderation. <em>IEEE Transactions on Computational Social Systems, </em>2019<em>, </em>6:38-55. DOI: <a href="http://doi.org/10.1109/TCSS.2018.2887240">10.1109/TCSS.2018.2887240</a> ⟨<a href="https://hal.science/hal-01999546">hal-01999546</a>⟩</li> <li>Arınık, N.; Figueiredo, R. & Labatut, V. Multiplicity and Diversity: Analyzing the Optimal Solution Space of the Correlation Clustering Problem on Complete Signed Graphs. <em>Journal of Complex Networks, </em>2020<em>, </em>8(6):cnaa025. DOI: <a href="http://doi.org/10.1093/comnet/cnaa025">10.1093/comnet/cnaa025</a> ⟨<a href="https://hal.science/hal-02994011">hal-02994011</a>⟩</li> <li>Arınık, N.; Figueiredo, R. & Labatut, V. Multiple partitioning of multiplex signed networks: Application to European parliament votes. <em>Social Networks, </em>2020<em>, </em>60:83-102. DOI: <a href="http://doi.org/10.1016/j.socnet.2019.02.001">10.1016/j.socnet.2019.02.001</a> ⟨<a href="https://hal.science/hal-02082574">hal-02082574</a>⟩</li> <li>Cécillon, N.; Labatut, V.; Dufour, R. & Arınık, N. Whole-Graph Representation Learning For the Classification of Signed Networks. <em>IEEE Access</em>, 2024, 12:151303-151316. DOI: <a href="https://dx.doi.org/10.1109/ACCESS.2024.3472474">10.1109/ACCESS.2024.3472474</a> <a href="https://hal.archives-ouvertes.fr/hal-04712854" rel="nofollow">⟨hal-04712854⟩</a></li> </ol> <p><strong>Funding: </strong>part of this work was funded by a grant from the <em>Provence-Alpes-Côte-d'Azur</em> region (PACA, France) and the <em>Nectar de Code</em> company.</p> <p><strong>Citation: </strong>If you use this data or the associated source code, please cite article [4]:</p> <p><code>@Article{Cecillon2024,</code><br><code> author = {Cécillon, Noé and Labatut, Vincent and Dufour, Richard and Arınık, Nejat},</code><br><code> title = {Whole-Graph Representation Learning For the Classification of Signed Networks},</code><br><code> journal = {IEEE Access},</code><br><code> year = {2024},</code><br><code> volume = {12},</code><br><code> pages = {151303-151316},</code><br><code> doi = {10.1109/ACCESS.2024.3472474},</code><br><code>}</code></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.