Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

410

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

410 results for “eukaryotic”

Learn how ShareScore rates datasets ↗
edi52/100

Abundance of eukaryote picophytoplankton and Synechococcus from a moored submersible flow cytometer at Martha's Vineyard Coastal Observatory, ongoing since 2003 (NES-LTER since 2017)

This is a decadal-scale time series of the abundance of eukaryote picophytoplankton and Synechococcus at 4 meters depth at the Martha's Vineyard Coastal Observatory, about 3 km south of Katama Beach, Edgartown, Massachusetts, USA. Picophytoplankton were sensed in situ by a submersible flow cytometer (FlowCytobot, or FCB). Sampling frequency was continuous at approximately 20-minute intervals binned to hourly resolution with some exceptions (e.g., winter in some years). This time series is ongoing for Northeast U.S. Shelf Long-Term Ecological Research (NES-LTER).

openCC (other)Jun 2025View details →
edi52/100

Abundance, biovolume, and biomass of Synechococcus and eukaryote pico- and nano- plankton from continuous underway flow cytometry during NES-LTER Transect cruises, ongoing since 2018

These data represent the abundance, biovolume, and biomass of prokaryotic and eukaryotic picoplankton and nanoplankton sampled continuously underway during Northeast U.S. Shelf Long-Term Ecological Research (NES-LTER) Transect cruises, ongoing since 2018. Samples were obtained with an Attune NxT Flow Cytometer sampling at approximately 2-min intervals from the underway science seawater. Cells were identified and enumerated from the flow cytometry data files based on their scattering, phycoerythrin (575 nm) and chlorophyll (680 nm) fluorescence signals.

openCC (other)Feb 2023View details →
edi52/100

Abundance, biovolume, and biomass of Synechococcus, eukaryote pico- and nano- phytoplankton, and heterotrophic bacteria from flow cytometry for water column bottle samples on NES-LTER Transect cruises, ongoing since 2018

These data represent the abundance, biovolume, and biomass of prokaryotic phytoplankton, eukaryotic pico- and nano- phytoplankton, and heterotrophic bacteria from discrete flow cytometry samples collected during the Northeast U.S. Shelf Long-Term Ecological Research (NES-LTER) Transect cruises, ongoing since 2018. Samples were collected and preserved from the water column at multiple depths using Niskin bottles on a CTD rosette system along the NES-LTER transect, and analyzed post cruise. Cells were identified and enumerated from the flow cytometry data files based on their scattering, SYBR (525 nm), phycoerythrin (575 nm) and chlorophyll (680 nm) fluorescence signals. Gating was completed manually in the Attune NXT software interface.

openCC (other)Jan 2024View details →
zenodo48/100

Darwin: an amino acid sequence collection of complete proteomes from eukaryotes with different phylogenetic affinities (v. 03_2020_137)

<p><strong>Background</strong></p> <p>Every time we find an interesting gene in an organism of interest, the first question is often &ldquo;how widely is this gene distributed in the eukaryotic kingdom?&rdquo;. Naturally, one could use NCBI BLAST search against the non-redundant sequence database provided by GenBank to answer this question. However, it can be cumbersome to parse the results and assign them to taxonomic units. It is also not straightforward to get an overview of which eukaryotic groups are represented in the results. Top BLAST hits can be crowded with sequences from closely-related organisms making it difficult gain an overview of the overall distribution across eukaryotes. To streamline this process, we developed an in-house database of complete eukaryotic proteomes. We tagged each sequence with a eukaryotic group handle (two-character symbol) and combined them into a single data set searchable by standalone BLAST on one&rsquo;s own computer. We named this data set &ldquo;Darwin&rdquo; to reflect the diverse nature of the sequences it contains.&nbsp;</p> <p><strong>Methods</strong></p> <p>We downloaded predicted proteomes in FASTA format from different sources such as GenBank, Joint Genome Institute (Depart of Energy, USA), Broad Institute (Massachusetts Institute of Technology, USA), Phytozome and a number of other specialized websites catering for a specific organism such as the Arabidopsis Information Resource (TAIR), or the Saccharomyces Genome Database (SGD). All the organisms we included in Darwin are listed in Table 1. To reduce redundancy, we took care not to include the same species more than once unless subspecies were known to show wide diversity. Each sequence header was tagged with a eukaryotic group handle composed of two-character symbols (based on Keeling&nbsp;<em>et al</em>., 2005). These handles clearly appear in BLAST output and can be parsed easily. We combined sequences from all proteomes into a single data set and named it &ldquo;Darwin&rdquo;.</p> <p><strong>Results</strong></p> <p>The current version of Darwin (v. 03_2020_137) contains 2,601,132 amino acid sequences from 137 eukaryotes (Table 1, Data file 1). The sizes of the proteomes were diverse, ranging from ~4000 sequences in some alveolates to 60,000-76,000 in plants. Darwin represents most of the supergroups of eukaryotic kingdom described in Keeling&nbsp;<em>et al.,</em>&nbsp;(2005) except those in Rhizaria whose genomes were not available at the time of data set construction. The data set contains larger numbers of proteomes from fungi and plants reflecting areas of interest in our group.&nbsp;</p> <p><strong>Conclusions</strong></p> <p>Darwin is provided as a text fasta file that can be formatted for BLAST searches on standalone computers. The results from the BLAST searches can be parsed to determine how widely a gene of interest is distributed among different eukaryotes. Simple counting of the eukaryotic group handles would also yield an overview of the distribution across taxa. Darwin is also useful for rapidly finding out whether a gene is missing in particular taxa.</p> <p><strong>Reference</strong></p> <p>Keeling PJ, Burger G, Durnford DG, Lang BF, Lee RW, Pearlman RE, Roger AJ, Gray MW (2005) The tree of eukaryotes.&nbsp;<em>Trends Ecol. Evol.</em>&nbsp;<strong>20:</strong>&nbsp;670-676</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes

<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here:&nbsp;<a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). &nbsp;&nbsp;</p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2&nbsp;Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name:&nbsp;</strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe:&nbsp;</strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' &nbsp;values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' &nbsp;values over 50%: FLAG_RP63.</li> <li><strong>flag_sum:&nbsp;</strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value).&nbsp;</p> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis.&nbsp;</p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0&nbsp; functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes.&nbsp;<br><br></p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

LukProt - an animal evolution-centric eukaryotic protein database

<p>LukProt is the EukProt database with additional species added, mostly the undersampled animal and some holozoan taxa.&nbsp;The database is composed of sequences translated from annotated genomes, transcriptomes or ESTs. <strong>The main purposes of the database are to consolidate sequences from undersampled animal taxa</strong> and provide usable search tools. The publication associated with LukProt can be found here: <a href="https://doi.org/10.1093/gbe/evae231">https://doi.org/10.1093/gbe/evae231</a>.</p> <p>The current version of the database (v1.5.1) is based on <a href="https://doi.org/10.24072/pcjournal.173">EukProt v3</a>. The home of all public versions of LukProt is this page (Zenodo).</p> <p>Proteomes that are novel in LukProt are denoted as LPXXXXX and those coming from AniProtDB are called APXXXXX. The sequence IDs from EukProt are conserved in LukProt. This means that each sequence is assigned an ID in the following format:</p> <pre><code>(A/E/L)PXXXXX_Species_epithet_(strain)_PYYYYYY</code></pre> <p>where XXXXX is a number from 00001 to 99999 and YYYYYY is a number from 000001 to 999999. Each sequence is assigned a unique number YYYYYY, and each taxon XXXXXX. All the IDs are compatible with BLAST v5 "-parse_seqids" option and the database can be readily deployed, for example on a server running <a href="https://doi.org/10.1093/molbev/msz185">SequenceServer</a>. Within each of the source fasta files, the source sequence identifier was kept after a blank space, so that it can still be retrieved if needed.</p> <p>A publicly available BLAST server providing LukProt search is available at: <a title="LukProt BLAST server" href="https://lukprot.hirszfeld.pl/" target="_blank" rel="noopener">https://lukprot.hirszfeld.pl/</a>.</p> <p>Comparison of EukProt v2/v3, LukProt 1.4.1 and LukProt v1.5.1 in their main areas of difference:</p> <table> <tbody> <tr> <th>Taxogroup</th> <th>EukProt v2</th> <th>EukProt v3</th> <th>LukProt v1.4.1</th> <th>LukProt v1.5.1</th> </tr> <tr> <th> <p>Holozoa</p> <p>(excluding Metazoa)</p> </th> <td>31</td> <td>40</td> <td>39</td> <td>43</td> </tr> <tr> <th>Ctenophora</th> <td>2</td> <td>2</td> <td>35</td> <td>38</td> </tr> <tr> <th>Porifera</th> <td>4</td> <td>5</td> <td>30</td> <td>47</td> </tr> <tr> <th>Placozoa</th> <td>2</td> <td>2</td> <td>3</td> <td>6</td> </tr> <tr> <th>Cnidaria</th> <td>3</td> <td>5</td> <td>65</td> <td>88</td> </tr> <tr> <th>Bilateria</th> <td>51</td> <td>51</td> <td>94</td> <td>142</td> </tr> </tbody> </table> <p>Included with the database are:</p> <ul> <li>ready to use main database files: <ul> <li><em>LukProt_v1.5.1_single_species_FASTA.7z</em> &ndash; a FASTA file with the sequences - <a href="https://en.wikipedia.org/wiki/7z">7-zipped</a>, <strong>uncompressed size: 17.6 GB</strong><br> <ul> <li>to concatenate all into one file, run this in the parent directory: <code>for file in $(find . -type f -name "*.fasta"); do awk 'FNR==1{print ""}1' $file &gt;&gt; LukProt_v1.5.1.fa; done</code>. This will create single FASTA file with all the sequences in the parent directory. <code>awk</code> is used to insert a new line after every file because&nbsp;<code>cat</code> would sometimes merge the last sequence with the header of the first sequence.</li> </ul> </li> <li><em>LukProt_v1.5.1_full_BLAST_db.7z</em> &ndash; a preformatted, full BLAST database (NCBI BLAST database format version: v5, masked with segmasker), <strong>uncompressed size: 28.3 GB</strong></li> <li><em>LukProt_v1.5.1_taxogroup_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one taxogroup and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.3 GB</strong></li> <li><em>LukProt_v1.5.1_single_species_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one BLAST database and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.4 GB</strong></li> </ul> </li> <li>auxiliary database files: <ul> <li><em>LukProt_v1.5.1.cdhit70.7z</em> &ndash; the full database clustered at 70% identity using CD-HIT with the following command: <code>cd-hit -g 1 -d 0 -T 20 -M 90000 -c 0.7 -uL 0.2 -uS 0.9 -s 0.2</code>,&nbsp;<strong>uncompressed sizes: fasta file - 11 GB, clstr file - 2.5 GB</strong></li> <li><em>LukProt_IDs_mapped.txt.gz</em> &ndash; a text file mapping the LukProt IDs to the AniProtDB IDs and EukProt IDs that are different</li> <li><em>BUSCO_tables.ods</em> &ndash; a spreadsheet with full result tables generated by BUSCO analysis</li> <li><em>OMAmer_output.zip</em> &ndash; a folder with full results of OMAmer analyses (includes per-sequence taxonomy classification)</li> <li><em>OMArk_output.zip</em> &ndash; a folder with the results of all OMArk analyses</li> </ul> </li> <li>metadata: <ul> <li><em>README.md</em> &ndash; a README file describing the metadata</li> <li><strong><em>LukProt_metadata_sheet.ods</em> &ndash; main metadata file. A spreadsheet with information about each proteome (in an open .ods format, most compatible with <a href="https://www.libreoffice.org/">LibreOffice</a>)</strong></li> <li><em>LukProt_metadata_other.zip</em> &ndash; an archive with other metadata files, documented in the README. Contents include:<br> <ul> <li>the LukProt taxonomy in various formats</li> <li>supporting scripts for data manipulation and visualization</li> </ul> </li> <li>a recoloring script (modified by LFS, originally by Dr. Celine Petitjean). The script is in&nbsp;<a title="formatFigtree2" href="https://doi.org/10.5281/zenodo.10654583">public domain</a> and reuploaded here only for convenience.&nbsp;</li> <li>other files - see README</li> </ul> </li> <li><em>changelog.md</em> &ndash; database changelog</li> </ul> <p>Words of caution:</p> <ul> <li>The database has been synchronized to EukProt v3 in version v1.5.1. This means that identifiers were modified in comparison to LukProt v1.4.1. The convention is not expected to change any more in future updates.</li> <li>Many proteomes, especially those transcriptome-based, may contain contamination from different species. In addition, the translation algorithms often introduce errors (e.g. the transcript may not represent a full length protein). For this reason, to get accurate sequences from each organism, users are directed to source data and to the included OMAmer, OMArk and BUSCO data for details.</li> <li>The taxonomy is different to UniEuk/EukMap, but UniEuk data were integrated where possible.</li> <li>A few NCBI taxids are missing and will be added in due course.</li> <li>Proteomes from NCBI and UniProt will be updated to current versions.</li> <li>A number of proteomes present in some metadata, are unpublished and were held back.</li> <li>While the database contains metadata that present a particular phylogeny of animals, holozoans and other eukaryotes, no particular claims or hypotheses are made by the author(s). However, in the future efforts will be made to name clades officially, once they are more firmly established.</li> </ul> <p><strong>Please report any problems or suggestions to Lukasz Sobala: lukasz.sobala (at) hirszfeld.pl.</strong></p> <p>&nbsp;</p> <p>Acknowledgements:</p> <ul> <li> <p>Andrew E. Allen Lab for creating the original <a href="https://allenlab.ucsd.edu/data/" target="_blank" rel="noopener">PhyloDB</a>.</p> </li> <li> <p>Daniel Richter <em>et al.</em> for creating <a href="https://doi.org/10.6084/m9.figshare.12417881">EukProt</a> and keeping it updated.</p> </li> <li> <p>Members of <a href="https://multicellgenome.com/">the Multicellgenome Lab</a>, especially Michelle Leger (for donating her database), for the bioinformatics support and for doing great science.</p> </li> <li> <p>All the authors of the original data.</p> </li> <li> <p>National Science Centre of Poland for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this database.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo44/100

The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3

<p>The North Pacific Eukaryotic Gene Catalog consolidates eukaryotic metatranscriptome data from three latitudinal transects of the North Pacific transition zone and one cruise in the subtropical gyre. Metatranscriptomes were gathered from latitudinally-resolved surface samples, and diel-resolved temporal studies, with samples taken in triplicate or duplicate and collected on 0.2-100 &mu;m, 0.2-3 &mu;m, and 3 &mu;m-100 or 200 &mu;m size fractions. These metatranscriptome data were <em>de novo</em> assembled into 175 independent assemblies, totalling 182 million clustered nucleotide contigs. Assemblies were annotated by taxonomy and function. This catalog provides assembled environmental contigs, their translated peptide sequences, and their taxonomic and functional annotations with the aim of facilitating continued discoveries about the molecular ecology of microbial eukaryotes in the North Pacific.<br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., &amp; Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.</p> <div> <p>This dataset repository is associated with a codebase and documentation repository:<br><a href="https://github.com/armbrustlab/NPac_euk_gene_catalog" target="_blank" rel="noopener">https://github.com/armbrustlab/NPac_euk_gene_catalog</a><br>Please see this code repository for additional data and project updates<br><br>Translated and processed protein sequences and their annotations are available in this repository: <br><a href="../doi/10.5281/zenodo.10472589">https://zenodo.org/doi/10.5281/zenodo.10472589</a><br><br>99% identity clustered nucleotide sequences and kallisto enumerations are available here:<br><a href="../doi/10.5281/zenodo.10570448">https://zenodo.org/doi/10.5281/zenodo.10570448</a></p> </div> <div> <p>File contents: this repository contains five .tar.gz compressed tarballs with raw de novo Trinity assemblies of poly-A selected metatranscriptomes from the Gradients 1 through 3 cruises, and a plain-text file with the custom spike-in mRNA standards (CustomStandardSequences.txt)</p> </div> <div> <p><strong><br>Gradients1.KOK1606.PA.assemblies.tar.gz</strong><br>- Link to&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G1PA" target="_blank" rel="noopener">G1PA project github page</a><br>- Simons CMAP cruise page and datasets:&nbsp;<a href="https://simonscmap.com/catalog/cruises/KOK1606" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KOK1606</a><br>- Short read processing code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.process_short_reads.sh" target="_blank" rel="noopener">G1PA.process_short_reads.sh</a><br>- Trinity assembly code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.trinity_assemblies.sh" target="_blank" rel="noopener">G1PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients2.MGL1704.PA.assemblies.tar.gz</strong><br>- Link to&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G2PA" target="_blank" rel="noopener">G2PA project github page</a><br>- Simons CMAP cruise page and datasets:&nbsp;<a href="https://simonscmap.com/catalog/cruises/MGL1704" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/MGL1704</a><br>- Short read processing code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.process_short_reads.sh" target="_blank" rel="noopener">G2PA.process_short_reads.sh</a><br>- Trinity assembly code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.trinity_assemblies.sh" target="_blank" rel="noopener">G2PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients3.KM1906.PA.assemblies.tar.gz</strong><br>- Link go&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets:&nbsp;<a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.process_short_reads.sh" target="_blank" rel="noopener">G3PA_UW.process_short_reads.sh</a><br>- Trinity assembly code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_UW.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>G3_diel.KM1906.PA.assemblies.tar.gz</strong><br>- Link go&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets:&nbsp;<a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.process_short_reads.sh" target="_blank" rel="noopener">G3PA_diel.process_short_reads.sh</a><br>- Trinity assembly code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_diel.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>CustomStandardSequences.txt<br></strong>- Plain-text FASTA file with the spike-in standards used during mRNA extraction and sequencing prep<br>- Link to publication of spike-in standards methods:&nbsp;<a href="https://www.nature.com/articles/s41564-019-0507-5" target="_blank" rel="noopener">https://www.nature.com/articles/s41564-019-0507-5</a></p> </div> <div> <p>The 2015 SCOPE Diel metatranscriptome raw assemblies have been released in a previous Zenodo repository, and are not included again in this deposition. We provide the links to the Diel1 resources here:<br>- Diel1 raw metatranscriptome assembly Zenodo repository:&nbsp;<a href="../records/5009803" target="_blank" rel="noopener">https://zenodo.org/records/5009803</a><br>- Dataset DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.5009803" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.5009803</a><br>- Associated publication:&nbsp;<a href="https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full" target="_blank" rel="noopener">https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full</a><br>- Codebase:&nbsp;<a href="https://github.com/armbrustlab/diel_eukaryotes" target="_blank" rel="noopener">https://github.com/armbrustlab/diel_eukaryotes</a><br>- Simons CMAP cruise page and datasets:&nbsp;<a href="https://simonscmap.com/catalog/cruises/KM1513" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1513</a><br>- Short read processing code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.process_short_reads.sh" target="_blank" rel="noopener">D1PA.process_short_reads.sh</a><br>- Trinity assembly code:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.trinity_assemblies.sh" target="_blank" rel="noopener">D1PA.trinity_assemblies.sh</a></p> </div> <p><br><br></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region

<p>This dataset has been prepared as part of the Interreg Alpine Space project Eco-AlpsWater (ASP569) -&nbsp;<em>Innovative Ecological Assessment and Water Management Strategy for the Protection of Ecosystem Services in Alpine Lakes and Rivers</em>,&nbsp;<a href="https://www.alpine-space.eu/projects/eco-alpswater/en/home">https://www.alpine-space.eu/projects/eco-alpswater/en/home</a></p> <p>Individual archives include 16S rRNA (cyanobacteria) and 18S rRNA (microalgae) FASTA sequences and associated blastn results obtained from the high throughput sequencing of plankton and biofilm bulk/eDNA samples collected in 2019 in 37 lakes and 22 rivers across the Alpine region. These are supporting files for the paper by Salmaso et al., 2022.&nbsp;DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region. Science of the Total Environment, in press.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

EukRibo: a manually curated eukaryotic 18S rDNA reference database

<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4&nbsp;&nbsp; &nbsp;</strong>We allowed up to 50 missing positions in the relatively conserved area at the 5&#39; end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3&#39; end of the V4 fragment is included.<br> <strong>V9&nbsp;&nbsp; &nbsp;</strong>We allowed up to 30 missing positions in the relatively conserved area at the 3&#39; end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5&#39; end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5&#39; end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3&#39; end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence (&#39;Y&#39;) or absence (&#39;N&#39;) of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> &nbsp;&nbsp; &nbsp;- same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete (&#39;yes - complete&#39;) or missing positions at the 5&#39; end (&#39;yes - partial&#39;)<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete (&#39;yes - complete&#39;), missing positions at the 3&#39; end (&#39;yes - partial&#39;), or was excluded, and the 6 possible reasons why (&#39;no - missing&#39;, &#39;no - too incomplete&#39;, &#39;no - chimera&#39;, &#39;no - bad quality&#39;, &#39;no - deletion in V9&#39;, &#39;no - Ns in V9&#39;)<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life. Data file underpinning Figure 2A and Figure 2B

<div>These datasheets accompany the article "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life" in Frontiers in Science</div> <div>This file contains data processed from Catalog of Life on 31 December 2023. The catalog was downloaded and post-processed to</div> <div>remove prokaryotic taxa</div> <div>remove extinct and fossil taxa</div> <div>remove taxon names that were listed as junior synonyms</div> <div>remove taxon names listed as "invalid"</div> <div>Total living, valid eukaryotic genera 167,085</div> <div>The taxa were sorted by the nomenclatorial Code under which they were declared (to avoid namespace clashes)</div> <div>International Code for Algae, Fungi and Plants https://www.iapt-taxon.org/nomen/main.php</div> <div>Algal, Fungal, Plant code genera 31,076</div> <div>International Code of Zoological Nomenclature https://www.iczn.org/the-code/the-code-online/</div> <div>Zoological code genera 136,009</div> <div>The Code-sorted taxa were aggregated by the generic portion of their names, and two plots were generated:</div> <div>a plot aggregating the cumulative number of species in genera sorted by species number (Figure 2A)</div> <div>a plot illustrating the distribution of the size of genera (Figure 2B)</div> <div>This data file gives access to these processed data for</div> <div>Figure 2 A Data</div> <div>Figure 2 B Data</div> <div>The original data including the intermediate calculations of values, and the plotted graphs, are available as a GoogleDoc at https://docs.google.com/spreadsheets/d/1V-bTtWjIRasC3AgID0jGlyToKqI-H9h1aPeSxUpNrjk/edit?usp=sharing</div>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data for "A promiscuous mechanism to phase separate eukaryotic carbon fixation in the green lineage"

<p>This repository contains all raw data associated with the manuscript:</p> <p>"<em><strong>A promiscuous mechanism to phase separate eukaryotic carbon fixation in the green lineage</strong></em>"</p> <p>&nbsp;</p> <p>The files are organised according to their appearance as figures in the manuscript.</p> <p>Within each of the zipped figure folders is a <strong>_readme.txt</strong> file that contains information about the raw data provided for each figure.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts

<p>This data continues with the development of the NPEGC Trinity&nbsp;<em>de novo</em> metatranscriptome assemblies from the protein data repository of <a href="../doi/10.5281/zenodo.10472589">The North Pacific Eukaryotic Gene Catalog</a>. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:<br><br><em>NPac.G1PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G2PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA.bf100.id99.nt.fasta.gz</em><br><em>NPac.G3PA_diel.bf100.id99.nt.fasta.gz</em><br><em>NPac.D1PA.bf100.id99.nt.fasta.gz</em><br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., &amp; Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.<br><br>These nucleotide sequences have been sourced from the&nbsp;Zenodo repository for raw assemblies: <a href="../records/7332796">The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3</a></p> <p>Key processing steps are sampled below with links to the detailed code on the main github code repository: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog">https://github.com/armbrustlab/NPac_euk_gene_catalog</a></p> <p><br>Code used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here:&nbsp;<a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/NPEGC.nt_kallisto_counts.sh">NPEGC.nt_kallisto_counts.sh</a><br><br>There are two main steps:<br>1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts<br>2. Map the short reads from environmental samples back to the assembly index</p> <p>As generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (&gt;50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.</p> <p>The code in this template script was used for each project: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/nt_data/aggregate_kallisto_counts.R">aggregate_kallisto_counts.R</a><br>The output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here:&nbsp;</p> <p><em>G1PA.raw.est_counts.csv.gz</em><br><em>G2PA.raw.est_counts.csv.gz</em><br><em>G3PA.raw.est_counts.csv.gz</em><br><em>G3PA_diel.raw.est_counts.csv.gz</em><br><em>D1PA.raw.est_counts.csv.gz</em></p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Diel-regulated transcriptional cascades of microbial eukaryotes in the North Pacific Subtropical Gyre

<p>Trinity <em>de novo </em>assemblies of 24 poly-A+ selected, combined-replicate metatranscriptomes from&nbsp;HOE-Legacy 2 cruise KM1513&nbsp;(Jul 24 - Aug 6, 2015).&nbsp;KM1513 cruise information, plots, and associated environmental data for the HOE Legacy II&nbsp;cruise can be found online at <a href="http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html">http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html</a>. Raw&nbsp;metatranscriptome short-read sequence data is available in the NCBI Sequence Read Archive&nbsp;under BioProject ID PRJNA492142. Code associated with this project is available on&nbsp;Github (<a href="https://github.com/armbrustlab/diel_eukaryotes">https://github.com/armbrustlab/diel_eukaryotes</a>).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset for 53 gene families from 16 eukaryotes

<p>NEXUS file representing Guigo et al.&#39;s (1996) dataset for 53 gene families from 16 eukaryotes, used in a number of gene tree reconciliation studies. Original data from&nbsp;Guigo et al., this NEXUS encoding by Roderic Page.</p>

opencc-by-4.0Jan 2023View details →
edi44/100

Effects of freshwater salinization on a salt-naïve planktonic eukaryote community

Salinization of freshwater ecosystems is a widespread issue, but evidence of ecological effects on aquatic eukaryote communities remains scarce. We experimentally exposed naive planktonic communities of a north-temperate, freshwater lake to a gradient of chloride (Cl-) concentration (0.27-1400 mg Cl-.L-1) with in-situ mesocosms. Following six weeks of exposure, we measured changes in the diversity, composition, and abundance of eukaryotic 18S rRNA gene. Total phytoplankton biomass remained unchanged, but we observed a shift in dominant phytoplankton groups with elevated salt concentration, from Cryptophyta and Chlorophyta that dominated in lower chloride concentrations (<185 mg Cl-.L-1) to Ochrophyta that dominated at higher conductivity (>185 mg Cl-.L-1). Most zooplankton and rotifer taxa were sensitive to the salinity and disappeared at low chloride concentrations (<40 mg Cl-.L-1). While ciliates thrived at low chloride concentrations (<185 mg Cl-.L-1), fungal groups dominated at intermediate chloride concentrations (185 mg Cl-.L-1 to 640 mg Cl-.L-1), and only phytoplankton remained at the highest chloride concentrations (> 640 mg Cl-.L-1).

openCC0Oct 2021View details →
edi44/100

Plumes and Blooms: Microbial eukaryote diversity and composition

These are amplicon sequencing data collected during Plumes and Blooms (PnB) cruises conducted from March, 2011, through September, 2014. The V9 hypervariable region of the 18S rRNA gene derived from microbial eukaryotic communities was amplified and sequenced from 345 discrete seawater samples. Sample collection and laboratory methods are described in Catlett et al. 2020 and Catlett et al. in review. Bioinformatic and data manipulation methods follow those employed in Catlett et al. in review. The data are provided in two tables: one includes amplicon sequence variant (ASV) sequences and relative sequence abundances for each sampling event, and the other includes ASV taxonomy predictions for each ASV sequence. References: Catlett, D., P. G. Matson, C. A. Carlson, E. G. Wilbanks, D. A. Siegel, and M. D. Iglesias‐Rodriguez. 2020. Evaluation of accuracy and precision in an amplicon sequencing workflow for marine protist communities. Limnol. Oceanogr.: Methods. 18(1): 20-40. https://doi.org/10.1002/lom3.10343. Catlett, D., D. A. Siegel, P. G. Matson, E. K. Wear, C. A. Carlson, T. S. Lankiewicz, and M. D. Iglesias‐Rodriguez. In review. Integrating phytoplankton pigment and DNA meta-barcoding observations to determine phytoplankton community composition in the coastal ocean. Limnol. Oceanogr.

openCC (other)May 2022View details →
zenodo40/100

3D MTHFR models in different eukaryotic species

<p>3D models for&nbsp;<em>Mus musculus, Gallus gallus, Danio rerio, Acanthaster planci, Arabidopsis thaliana, C.elegans MTHFR&nbsp;</em> protein&nbsp;</p>

opencc-by-4.0Apr 2020View details →
dryad40/100

SPIKEPIPE: A metagenomic pipeline for the accurate quantification of eukaryotic species occurrences and intraspecific abundance change using DNA barcodes or mitogenomes

<p>The accurate quantification of eukaryotic species abundances from bulk samples remains a key challenge for community ecology and environmental biomonitoring. We resolve this challenge by combining shotgun sequencing, mapping to reference DNA barcodes or to mitogenomes, and three correction factors: (a) a percent‐coverage threshold to filter out false positives, (b) an internal‐standard DNA spike‐in to correct for stochasticity during sequencing, and (c) technical replicates to correct for stochasticity across sequencing runs. The SPIKEPIPE pipeline achieves a strikingly high accuracy of intraspecific abundance estimates (in terms of DNA mass) from samples of known composition (mapping to barcodes R<sup>2</sup> = .93, mitogenomes R<sup>2</sup> = .95) and a high repeatability across environmental‐sample replicates (barcodes R<sup>2</sup> = .94, mitogenomes R<sup>2</sup> = .93). As proof of concept, we sequence arthropod samples from the High Arctic, systematically collected over 17 years, detecting changes in species richness, species‐specific abundances, and phenology. SPIKEPIPE provides cost‐efficient and reliable quantification of eukaryotic communities.</p>

opencc-zeroAug 2019View details →
zenodo40/100

Figure 12. from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013

Figure 12. - Map of Croatia showing the locality of Eupolybothrus cavernicolus Komerički &amp; Stoev sp. n.

opencc-by-4.0Mar 2017View details →
zenodo40/100

Figure 10a. from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013

Figure 10a. - Eupolybothrus cavernicolus Komerički &amp; Stoev sp. n., male paratype. Figure 10a. close up of the tip of prefemoral spine p Figure 10b. coxal pore pit, meso-ventral view <br> close up of the tip of prefemoral spine p

opencc-by-4.0Mar 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record