Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,835

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,835 results for “dominance”

Learn how ShareScore rates datasets ↗
edi60/100

Species diversity and plant dominance influence grassland stability in response to extreme climatic events and anthropogenic drivers across three LTER sites: Cedar Creek, Konza Prairie, and Kellogg Biological Station, 1982-2023.

The data in this package is associated with the analysis for a manuscript titled "Multiple community properties drive ecosystem resistance and resilience to extreme climate events across mesic grasslands". The files include compiled data on plant biomass production, species abundance, experimental treatments, extreme climate event values, and calculated diversity and stability measures from grassland plots in experiments at CDR, KBS, and KNZ LTER sites.

openCC (other)Sep 2025View details →
edi60/100

Annual monitoring of high marsh plots dominated by Juncus and Borrichia at the GCE LTER from 2013 - 2025

Annual monitoring of high marsh plots dominated by Juncus and Borrichia. Plots were established at GCE sites 6 and 10 in 2013, and are monitored annually. The goal is to determine how annual variation in climate and other abiotic factors affects the vegetation composition.

openCC (other)Oct 2025View details →
edi60/100

Percent cover measurements of four site-dominant species from the GCE-LTER Seawater Addition Long-Term Experiment (SALTEx) Project

SALTEx (Seawater Addition Long-Term Experiment) is a field experiment designed to simulate saltwater intrusion in a tidal freshwater wetland to predict how chronic (Press) and acute (Pulse) salinization will affect this and other tidal freshwater ecosystems. The SALTEx experiment was initiated in 2012 and consists of 31 field plots, each 2.5 m on a side. There are three treatments (Press, Pulse, and Fresh) and two types of controls (with and without sides), each consisting of six replicates. The Press treatment plots receive regular (4 times each week) additions of a mixture of seawater and fresh river water. Pulse plots receive the same mixture of seawater and river water during September and October, which is historically a time of low flow in the river when natural saltwater intrusion occurs. The Fresh treatment plots receive regular additions of fresh river water. Treatment water is added during low tide to facilitate its infiltration into the soil, and all plots are inundated by astronomical tides at high tide. Percent cover was measured for four site-dominant species (Zizaniopsis miliacea, Pontederia cordata, Persicaria hydropiperoides, and Ludwigia repens) each July from 2013 to 2022.

openCC (other)Jul 2024View details →
edi56/100

Seasonal Soil Sampling of Grass-dominated, Mesquite-dominated, and Ecotone Sites at the Jornada Basin LTER site for the Analysis of Microbial Community Variance, 2022-2023

Fungal and bacterial soil communities were analyzed to assess the influence of woody shrub encroachment on soil microbial communities. Three study sites in the Jornada Long Term Ecological Research Site were selected to represent a grass-dominated site, a woody shrub dominated site, and an ecotone of woody shrubs and grass. The field sampling began in October 2022 and concluded in July 2023 with five sampling periods that aimed to capture seasonal variation: October 2022, January 2023, March 2023, May 2023, and July 2023. This dataset includes data pertaining to the soil microbial composition, environmental characteristics, microbial sequence processing, and documentation of the code utilized for data processing and statistical analyses. Data on soil microbial composition was collected from Phospholipid Fatty-Acid composition data from soil samples. Data on environmental characteristics were collected from on-site temperature probes, laboratory assessments of soil properties, and Jornada meteorological stations. Information pertaining to microbial sequence processing is included in the documented code as well as in the record of the primers utilized.

openCC0Apr 2025View details →
edi56/100

Plant species cover and biomass for Sevilleta dominant species removal experiment.

The purpose of this research project was to connect the removal of dominant grass species in grasslands at the Sevilleta National Wildlife Refuge to changes in plant community composition and subsequent changes in aboveground biomass. We used species cover data for 23 years of a dominant species removal experiment (https://doi.org/10.6073/pasta/fd3c777524231ae245bf1916715c9140) and converted percent cover values to aboveground standing biomass using methods from Rudgers et al. 2019 (https://doi.org/10.1111/1365-2435.13463). For this project, only two sites from the original study were used blue grama (site 1) and black grama (site 3) as they are referred to in the original study.

openCC0Jul 2025View details →
zenodo52/100

Pawpaws prevent predictability: A locally-dominant tree alters understory beta-diversity and community assembly

<p>Data used in "Pawpaws Prevent Predictability: A locally-dominant tree alters understory beta-diversity and community assembly" (Wassel and Myers) accepted for publication in Ecosphere.<br><br><strong>Metadata for Zenodo.pdf&nbsp;</strong>contains more information on the following data files including descriptions of the columns.&nbsp;</p> <p>The file <strong>understory_abundance_data2021.csv</strong>&nbsp;contains all species abundances in 1x1m plots. This data was used for analyses in publication. Each row is a plot, each column is a speceis or plot descriptor, values for columns 5 and higher are species abundances. Data was collected July-August 2021 by Anna Wassel in Missouri, USA.&nbsp;</p> <p>The file <strong>understory_species_list2021.csv&nbsp;</strong>contains a list of the species codes used in the first file with their scientific names and their status as herbs or woody. This was used to filter out herbaceous species from the data set for herbaceous-only analyses.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Dominant contribution of Asgard archaea to eukaryogenesis (2024) Tobiasson, V., Koonin, E. PROCESSED DATA AND METADATA

<h1>Main data deposit for "Dominant contribution of Asgard archaea to eukaryogenesis".&nbsp;</h1> <p>Victor Tobiasson, Jacob Luo, Yuri I Wolf, Eugene V Koonin</p> <p>Computational Biology Branch, Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA</p> <p><strong>The Origin of eukaryotes is one of the key problems in evolutionary biology. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes inform and constrain evolutionary scenarios of eukaryogenesis. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, without consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis.&nbsp;</strong></p> <div> <div> <div>Version 0.3, updated 180325</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>Main data repository for:</div> <div>Dominant contribution of Asgard archaea to eukaryogenesis (2024)&nbsp;</div> <div>Tobiasson, V., Koonin, E.</div> <div>&nbsp;</div> <div>Contains all final parsed data from the main Eukaryogenesis project&nbsp;</div> <div>investigating the evolutionary ancetries of eukaryotic protein families.&nbsp;</div> <div>&nbsp;</div> <div>Currently (non-static) available at:&nbsp;</div> <div>https://www.biorxiv.org/content/10.1101/2024.10.14.618318v2</div> <div>https://assets-eu.researchsquare.com/files/rs-5352492/v1/2f9c68ae-cf3e-420a-8d29-867b6fb1a878.pdf</div> <div>&nbsp;</div> <div>All code used to generate the data present within this repository available at:&nbsp;</div> <div>https://github.com/VictorTobiasson/eukgen&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### General information</div> <div>&nbsp;</div> <div>To identify associations between prokaryotic and eukaryotic protein families, separate</div> <div>hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed&nbsp;</div> <div>using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using&nbsp;</div> <div>mmseqs2, followed by a multistep data-reduction and multiple sequence alignment (MSA)&nbsp;</div> <div>procedure to generate HMM profiles using hhsuite.&nbsp;</div> <div>&nbsp;</div> <div>A prokaryotic database of 37 million protein sequences was curated from prokaryotic&nbsp;</div> <div>genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins&nbsp;</div> <div>extracted from 146 Asgard genome assemblies. To avoid inclusion of genes present only&nbsp;</div> <div>within a narrow subset of species, possibly resulting from horizontal transfer from&nbsp;</div> <div>eukaryotes post LECA, we reconstructed the &ldquo;soft-core&rdquo; pangenome for each of the 26&nbsp;</div> <div>curated prokaryotic taxonomic classes. These pangenomes include only those genes that&nbsp;</div> <div>are present in at least 67% of the families within each class of Bacteria and Archaea.&nbsp;</div> <div>The initial eukaryotic database consisted of 30 million protein sequences from 993&nbsp;</div> <div>species taken from EukprotV3 and cleaned using mmseqs2 to remove likely prokaryotic&nbsp;</div> <div>contaminants.&nbsp;</div> <div>&nbsp;</div> <div>Both databases were clustered and MSAs constructed for all non, singleton clusters&nbsp;</div> <div>and HMM profiles created. The resulting eukaryotic HMM dataset was queried against&nbsp;</div> <div>the prokaryotic dataset using hhblits to identify sets of homologous protein sequences.&nbsp;</div> <div>Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual</div> <div>&nbsp;sequence set, hereinafter referred to as an Eukaryotic/Prokaryotic Orthologous Cluster&nbsp;</div> <div>(EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes&nbsp;</div> <div>(each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic&nbsp;</div> <div>proteins can be present in multiple EPOCs) that were used for phylogenetic tree&nbsp;</div> <div>construction, annotation, and evolutionary hypothesis testing.&nbsp;</div> <div>&nbsp;</div> <div>To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC,&nbsp;</div> <div>rather than relying on the tree topology directly, we employed a probabilistic approach&nbsp;</div> <div>for evolutionary hypothesis testing using constraint trees. We exhaustively sampled all&nbsp;</div> <div>arrangements of likely sister clades and obtained Expected Likelihood Weights (ELW) for&nbsp;</div> <div>the set of possible sister clade models. As the ELW metric is analogous to model selection&nbsp;</div> <div>confidence, here we take it to be proportional to the probability of a sampled prokaryotic&nbsp;</div> <div>clade to be the true sister group of the given eukaryotic clade among a set of competing&nbsp;</div> <div>sister clades. For each EPOC, our analysis dynamically accounts for long branch outliers&nbsp;</div> <div>and is robust to phylogenetically non-homogenous clades. This analysis is further capable&nbsp;</div> <div>of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a&nbsp;</div> <div>single datapoint for downstream analysis. Our resulting data contains EPOCs annotated&nbsp;</div> <div>using profiles generated from KEGG Orthology Groups (KOGs), each with an MSA generated&nbsp;</div> <div>using muscle5, a maximum likelihood tree inferred using IQtree2 and associated ELW values&nbsp;</div> <div>for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was&nbsp;</div> <div>performed only for those eukaryotic clades that included more than 5 distinct taxonomic&nbsp;</div> <div>labels, with at least one coming from Amorphea and one from Diaphoretickes, the two&nbsp;</div> <div>expansive eukaryotic clades considered to represent either the first or the second&nbsp;</div> <div>bifurcation in the evolution of eukaryotes. Thus, these clades likely represent genes&nbsp;</div> <div>mapping back to the LECA.</div> <div>&nbsp;</div> <div>For further details please see main publication or contact</div> <div>victor.tobiasson@nih.gov</div> <div>eugene.koonin@nih.gov</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>### Included files</div> <div>&nbsp;</div> <div>Unless otherwise stated all files contained are tab separated and utf-8 encoded&nbsp;</div> <div>with the first row containing header information.&nbsp;</div> <div>All data entries encoding lists are &ldquo;|&rdquo; (pipe) separated.&nbsp;</div> <div>Fields without data values are filled with string entries of &ldquo;none&rdquo;.</div> <div>&nbsp;</div> <div>--- Databases ---</div> <div>euk72_ep.tar.gz</div> <div>prok2311_as.tar.gz</div> <div>Prok2311As_final_clusters.tsv</div> <div>Euk72Ep_final_clusters.tsv</div> <div>prok2311_as.hmmDB.tar.gz</div> <div>euk72_ep.hmmDB.tar.gz</div> <div>&nbsp;</div> <div>--- Annotation and Curation ---</div> <div>NCBI_taxonomy_species_addendum.tsv</div> <div>NCBI_taxonomy_class_addendum.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>KEGG_category_mapping.tsv</div> <div>KEGG_metadata.tsv</div> <div>&nbsp;</div> <div>--- EPOC data ---</div> <div>EPOC_data.tar.gz</div> <div>EPOC_annotation_KEGG.tsv</div> <div>EPOC_data.tsv</div> <div>EPOC_data.pangenomes_s10.tsv</div> <div>EPOC_data.pangenomes_s25.tsv</div> <div>EPOC_data.pangenomes_s67.tsv</div> <div>EPOC_data.GTDB.tsv</div> <div>&nbsp;</div> <div># euk72_ep.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files&nbsp;</div> <div>constituting the initial eukaryotic mmseqs2 database with taxonomy annotation.&nbsp;</div> <div>Constructed from a pre-selected list of 72 eukaryotic proteomes downloaded from&nbsp;</div> <div>NCBI as well as a &ldquo;clean&rdquo; version of Eukprot, lacking highly prokaryotic-like&nbsp;</div> <div>contaminant sequences.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files constituting the&nbsp;</div> <div>initial prokaryotic mmseqs2 database with taxonomy annotation. Constructed from&nbsp;</div> <div>47545 complete genomes retrieved from NCBI in November 2023.&nbsp;</div> <div>&nbsp;</div> <div># prok2311_as.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from prok2311_as non--singleton clusters, contains 26286 profiles.</div> <div>&nbsp;</div> <div># euk72_ep.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted&nbsp;</div> <div>from euk72_ep non-singleton clusters, contains 1631704 profiles.</div> <div>&nbsp;</div> <div># NCBI_taxonomy_species_addendum.tsv</div> <div>Taxonomy mapping file with manually curated &lsquo;class&rsquo; level annotation for poorly&nbsp;</div> <div>annotated species.&nbsp;</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>proposed_class_id: Manually assigned NCBI taxid</div> <div>proposed_class_label: NCBI class name</div> <div>org_name: NCBI organism name</div> <div>&nbsp;</div> <div># NCBI_taxonomy_class_addendum.tsv</div> <div>Class revision file mapping poorly populated class level entries to higher order&nbsp;</div> <div>manually curated labels. Also includes information for small classes with shallow&nbsp;</div> <div>taxonomy which are deleted from the EPOC analysis at the level of tree construction.</div> <div>&nbsp;</div> <div>taxid: NCBI taxid</div> <div>ncbi_class: NCBI taxid of rank corresponding to &lsquo;class&rsquo; following manual&nbsp;</div> <div>amendment as per NCBI_taxonomy_species_addendum.tsv</div> <div>revised_class_id: Manually assigned NCBI taxid of rank corresponding to &lsquo;class&rsquo;</div> <div>revised_class_label: Proposed cleartext name of manually revised revised_class_id&nbsp;</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Final taxonomy at NCBI rank &lsquo;class&rsquo; following revisions for all sequences in Euk72Ep or&nbsp;</div> <div>Prok2311As. These taxonomic labels are used for EPOC tree annotation.&nbsp;</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of manually revised NCBI rank &lsquo;class&rsquo; identifier for annotation</div> <div>&nbsp;</div> <div># Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Final taxonomy at GTDB rank &lsquo;phylum&rsquo; transferred using marker genes from GTDB release 220</div> <div>&nbsp;</div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya,&nbsp;</div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of assigne GTDB phylum</div> <div>&nbsp;</div> <div># Prok2311As_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation&nbsp;&nbsp;</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># Euk72Ep_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the&nbsp;</div> <div>final clusters used for HMM creation</div> <div>&nbsp;</div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div>&nbsp;</div> <div># EPOC_data.tar.gz</div> <div>Gunzip-ed directory containing 16035 EPOC folders. Each folder is named corresponding&nbsp;</div> <div>to the eukaryotic cluster representative which generated its profile as an ID&nbsp;</div> <div>Matches the tree_name field in EPOC_data_prok2311As.tsv</div> <div>contains the following files:</div> <div>&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta: sequences for all members of the EPOC</div> <div>&lt;EPOC_ID&gt;.merged.fasta.leaf_mapping: tsv separated file containing taxonomy and tree reduction data</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle: main cropped MSA for tree generation&nbsp;</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.iqtree: IQtree2 output from tree generation</div> <div>&lt;EPOC_ID&gt;.merged.fasta.muscle.treefile.annot: annotated newick tree file with final tree</div> <div>&lt;EPOC_ID&gt;.merged.tree_data.tsv: final parsed tree data with columns matching&nbsp; EPOC_data_prok2311As.tsv</div> <div>&nbsp;</div> <div>EPOCs with more than one possible eukaryotic sister phyla also contains&nbsp;</div> <div>a folder "constraint_analysis" with constraint tree information used for&nbsp;</div> <div>ELW value calculation.&nbsp;</div> <div>&nbsp;</div> <div># EPOC_data.tsv</div> <div>Main resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs)&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class. This is the main</div> <div>data to be used for genereting the core dataset and for data visualistation</div> <div>Contains information regarding tree breakdown, LCA membership and phylogenetic&nbsp;</div> <div>distances between all detected LCAs. Equivalent to the stacked dataframes from all&nbsp;</div> <div>EPOC directories in EPOC_data&nbsp;</div> <div>&nbsp;</div> <div>tree_name: unique index for each EPOC&nbsp;</div> <div>euk_clade_rep: unique index for each annotated eukaryotic clade within each tree_name</div> <div>euk_clade_size: number of original sequences represented by euk_clade_rep</div> <div>euk_clade_weight: metric for taxonomic purity for each euk_clade_rep</div> <div>euk_leaf_clade: boolean indicating whether euk_clade_rep contains a single leaf</div> <div>euk_LCA: lowest taxa spanning all members in euk_clade_rep</div> <div>euk_scope: list of all taxonomic classes in euk_clade_rep</div> <div>euk_scope_len: length of euk_scope list</div> <div>prok_clade_rep: unique index for each annotated prokaryotic clade for each euk_clade_rep</div> <div>prok_clade_size: number of original sequences represented by prok_clade_rep</div> <div>prok_clade_weight: metric for taxonomic purity for each prok_clade_rep</div> <div>prok_leaf_clade: boolean indicating whether prok_clade_rep contains a single leaf</div> <div>prok_taxa: lowest taxa spanning all members in prok_clade_rep</div> <div>dist: tree-distance from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>top_dist: graph-distance (node-distance) from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>raw_stem_length: tree-distance from lowest tree node containing the union of all members of prok_clade_rep and euk_clade_rep to the tree node containing all members of euk_clade_rep</div> <div>median_euk_leaf_dist: median value for all tree distances from the tree node containing all members of euk_clade_rep to the individual leaves</div> <div>stem_length: raw_stem_length/median_euk_leaf_dist</div> <div>logL: log likelihood of best constraint tree constructed</div> <div>deltaL: log likelihood difference between constraint tree for prok_clade_rep and best constraint tree constructed</div> <div>bp-RELL: validation metric from IQtree -trees, see iqtree.org</div> <div>bp-RELL_accept: as above</div> <div>p-KH: as above</div> <div>p-KH_accept: as above</div> <div>p-SH: as above</div> <div>p-SH_accept: as above</div> <div>c-ELW: as above</div> <div>c-ELW_accept: as above</div> <div>p-AU: as above</div> <div>p-AU_accept: as above</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s10.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 10% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s25.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 25% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.pangenomes_s67.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>based on pangenomes defined as including 67% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.GTDB.tsv</div> <div>Resulting data&nbsp; from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>under revised taxonomy from GTDB based on data from Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Identical file structure to EPOC_data.tsv</div> <div>&nbsp;</div> <div># EPOC_data.alpha_replicates.tsv</div> <div>Resulting data from 20 repetitions of Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated&nbsp;</div> <div>from a subset of Alphaproteobacterial-derived EPOCs.&nbsp;</div> <div>Identical file structure to EPOC_data.tsv with the addition of:</div> <div>&nbsp;</div> <div>rep: indicating technical replicate number, 0-19</div> <div>&nbsp;</div> <div># EPOC_annotation_KEGG.tsv</div> <div>Parsed HHblits output of HMM profiles generated from KEGG KOGs (KEGG Orthologous Groups)&nbsp;</div> <div>against eukaryotic profiles constituting each EPOC</div> <div>&nbsp;</div> <div>Query: query name equal to tree_name from EPOC_data</div> <div>Target: target name equal to kogid in KEGG_category_mapping and KEGG_metadata</div> <div>Prob: data from HHblits, see https://github.com/soedinglab/hh-suite/wiki</div> <div>E-value : as above</div> <div>P-value : as above</div> <div>Score: as above</div> <div>SS: as above</div> <div>Cols: as above</div> <div>Identities: as above</div> <div>Similarity: as above</div> <div>Sum_probs: as above</div> <div>Query-HMM-start: as above</div> <div>Query-HMM-end: as above</div> <div>Template-HMM-start: as above</div> <div>Template-HMM-end: as above</div> <div>Template_columns: as above</div> <div>Template_Neff : as above</div> <div>Pairwise_cov: calculated pairwise coverage from Query and Target start and end</div> <div>Description: category_name from KEGG_category_mapping</div> <div>&nbsp;</div> <div># KEGG_category_mapping.tsv</div> <div>Mapping of relevant KOG identifiers to their higher order categories as&nbsp;</div> <div>"Maps" "Modules" or "Reactions" as per KEGG see https://www.kegg.jp/kegg/pathway.html</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>category_id: KEGG map, module, or reaction number</div> <div>category_name: cleartext name for KOG identifier</div> <div>&nbsp;</div> <div># KEGG_metadata.tsv</div> <div>File mapping KOGs to BRITE classification and to additional databases of chemical properties.</div> <div>&nbsp;</div> <div>kogid: unique KOG identifier</div> <div>name: cleartext name for KOG identifier</div> <div>brite_A: list of BRITE-A sets including KOG</div> <div>brite_B: list of BRITE-A sets including KOG</div> <div>brite_C: list of BRITE-A sets including KOG</div> <div>EC: list of Enzyme commission numbers associated with KOG, see https://enzyme.expasy.org/</div> <div>TC: list of transporter classification numbers associated with KOG, see https://www.tcdb.org/</div> <div>RN: list of KEGG reaction numbers associated with KOG</div> <div>CA: list of CAZY numbers associated with KOG, see http://www.cazy.org/</div> <div>GO: list of GO terms associated with KOG, see https://geneontology.org/</div> </div> <div>&nbsp;</div> </div>

opencc-by-4.0Oct 2024View details →
edi52/100

LTREB: Marsh elevation change in control and fertilized plots in a Spartina alterniflora-dominated salt marsh, North Inlet, Georgetown, SC: 1990-2025.

Marsh elevation was measured with a Surface Elevation Table (SET) as a component of a long-term project seeking to understand how salt marsh primary production and sediment chemistry respond to anthropogenic (e.g. eutrophication) and natural (e.g. sea-level rise) environmental change. Feedbacks between plants, sediments, nutrients and flooding were investigated with particular attention to mechanisms that keep marshes in equilibrium with sea level. Other data collected as part of the project include aboveground annual primary productivity, plant biomass, plant density and porewater nutrient concentrations. These data have been used to develop the Marsh Equilibrium Model, an important tool for coastal resource managers. Sampling occurred at 7 Spartina alterniflora-dominated salt marsh sites in North Inlet, a relatively pristine estuary near Georgetown, SC on the SE coast of the United States. North Inlet is a tidally-dominated, bar-built estuary, with a semi-diurnal mixed tide and a tidal range of 1.4m. The 25-km2 estuary is comprised of about 20.5 km2 of intertidal salt marsh and mudflats, and 4.5 km2 of open water. Marsh elevation sampling began in 1990, 1991, 1996 or 2000, depending on the site. Sampling occurred approximately monthly or approximately annually through 2025. The study is on-going. Additionally, some plots were fertilized with nitrogen and phosphorus.

openCC0Jan 2026View details →
edi52/100

LTREB: Aboveground biomass, plant density, annual aboveground productivity, plant heights and snail observations in control and fertilized plots in a Spartina alterniflora-dominated salt marsh, North Inlet, Georgetown, SC: 1984-2025

Aboveground biomass and plant density were measured non-destructively as a component of a long-term project seeking to understand how salt marsh primary production and sediment chemistry respond to anthropogenic (e.g. eutrophication) and natural (e.g. sea-level rise) environmental change. Feedbacks between plants, sediments, nutrients and flooding were investigated with particular attention to mechanisms that keep marshes in equilibrium with sea level. Biomass was calculated from plant height measurements using allometric equations. Annual productivity was calculated from approximately-monthly biomass estimates. In addition to plant height measurements, observations of snails in sample plots were recorded. Other data collected as part of the project include marsh surface elevation and porewater nutrient concentrations. These data have been used to develop the Marsh Equilibrium Model, an important tool for coastal resource managers. Sampling occurred at Spartina alterniflora-dominated salt marsh sites in North Inlet, a relatively pristine estuary near Georgetown, SC on the SE coast of the United States. North Inlet is a tidally-dominated, bar-built estuary, with a semi-diurnal mixed tide and a tidal range of 1.4m. The 25-km2 estuary is comprised of about 20.5 km2 of intertidal salt marsh and mudflats, and 4.5 km2 of open water. Sampling began at one location in 1984, and at three additional locations in 1986. Sampling occurred approximately monthly through 2025. The study is on-going. There are four sampling locations at two sites. Two locations are in the low marsh; two locations are in the high marsh. One high marsh location had control sampling plots in addition to plots fertilized with nitrogen and phosphorus.

openCC0Jan 2026View details →
edi52/100

Porewater nutrient concentrations in control and fertilized plots in a Spartina alterniflora-dominated salt marsh, North Inlet, Georgetown, SC : 1993-2025

Porewater nutrient concentrations were measured as a component of a long-term project seeking to understand how salt marsh primary production and sediment chemistry respond to anthropogenic (e.g. eutrophication) and natural (e.g. sea-level rise) environmental change. Feedbacks between plants, sediments, nutrients and flooding were investigated with particular attention to mechanisms that keep marshes in equilibrium with sea level. Other data collected as part of the project include aboveground macrophyte biomass, plant density, marsh surface elevation and annual above ground primary productivity. These data have been used to develop the Marsh Equilibrium Model, an important tool for coastal resource managers. Sampling occurred at Spartina alterniflora-dominated salt marsh sites in North Inlet, a relatively pristine estuary near Georgetown, SC on the SE coast of the United States. North Inlet is a tidally-dominated, bar-built estuary, with a semi-diurnal mixed tide and a tidal range of 1.4m. The 25-km2 estuary is comprised of about 20.5 km2 of intertidal salt marsh and mudflats, and 4.5 km2 of open water. Sampling began at two locations in December 1993, and at three additional locations in January 1994. Sampling occurred approximately monthly at these 5 locations through 2025. Sampling occurred at a sixth location from 2006 to 2010. The site was a dieback site that had recovered by 2010. At the other sites, the study is on-going. Porewater was collected at multiple depths from diffusion samplers and was analyzed for sulfide, salinity, ammonium, phosphate, and iron concentrations. There are five sampling locations at three sites. Two locations are in the low marsh; three locations are in the high marsh. One high marsh location had control sampling plots in addition to plots fertilized with nitrogen and phosphorus.

openCC0Jan 2026View details →
edi52/100

Species cover, community biomass, and richness in global grasslands from NutNet (2007–2023): Dominant species predict plant richness and biomass in global grasslands

The Nutrient Network (NutNet) is a globally coordinated research initiative designed to investigate the impacts of human-driven alterations in nutrient availability and consumer presence on grassland ecosystems. Data were collected from over 130 herbaceous-dominated sites worldwide, spanning diverse environmental conditions from desert grasslands to arctic tundra. Standardized methodologies were employed across all sites to enable direct comparisons of productivity, diversity, and ecosystem responses. Experimental treatments included nutrient additions to assess co-limitation of plant growth by multiple nutrients, as well as grazer manipulations to examine their role in regulating biomass, species diversity, and community composition. By compiling these cross-site data, NutNet aims to enhance our understanding of productivity-diversity relationships and provide new insights into the ecological consequences of anthropogenic changes to nutrient cycles and food webs at a global scale.

openCC (other)Apr 2025View details →
edi52/100

Survival, growth and biomass estimates of two dominant palmetto species of south-central Florida from 1981 - 2022, ongoing at 5-year intervals

This data package is comprised of three datasets all pertaining to two dominant palmetto species, Serenoa repens and Sabal etonia, at Archbold Biological Station in south-central Florida. The first dataset, palmetto_data, contains survival and growth data across multiple years, habitats and experimental treatments. The second dataset, seedlings_data, follows the fate of marked putative palmetto seedlings in the field to assess survivorship and growth. The final dataset, harvested_palmetto_data, contains size data and estimated dry mass (biomass in grams) of 33 destructively harvested palmetto plants (17 S. repens and 16 S. etonia) of varying sizes and across habitats. Thirty-two of these were used to calculate estimated biomass, using regression equations, for palmettos sampled in the palmetto_data. Below we summarize experimental setup and data collected for each dataset. Palmetto data Demographic data were collected as three separate components. The first component compared growth among habitats. Starting in 1981, equal numbers of both palmetto species were marked across scrubby flatwoods (oak scrub) and flatwoods habitats (3 sites per habitat) for a total of 240 marked plants. These habitats had not burned within the last decade, but historically had experienced a natural fire return interval of 5 - 20 years prior to this studies initiation. The second component added an additional 400 palmettos (200 of each species), which were marked in sand pine scrub (n = 200) in 1985 and sandhill habitat (n = 200) in 1989 on Archbold's Red Hill. At the time of this project's initiation, all Red Hill management units were last burned in 1927 and were considered long unburned. Part of Archbold's management plan included restoring fire into some management units while leaving others long unburned to serve as reference units. Therefore, for our second component, we were able to create a 2x2 factorial design using habitat types on Red Hill and fire management as factors, with 100

openCC0Sep 2023View details →
edi52/100

[DEPRECATED] Aboveground biomass at control and fertilized plots in a Spartina patens-dominated marsh, Rowley River, Plum Island Ecosystem LTER, MA. (Reformatted to ecocomDP Design Pattern)

This ecocomDP data package is about populations rather than communities, and has therefore been deprecated. This data package is formatted according to the "ecocomDP", a data package design pattern for ecological community surveys, and data from studies of composition and biodiversity. For more information on the ecocomDP project see https://github.com/EDIorg/ecocomDP/tree/master, or contact EDI https://environmentaldatainitiative.org. This Level 1 data package was derived from the Level 0 data package found here: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-pie&identifier=33&revision=16 The abstract below was extracted from the Level 0 data package and is included for context:

openCC (other)Aug 2021View details →
edi52/100

Co-dominant removal and N and C fertilization experiment for moist meadow tundra, 2002 - 2018.

In 2002 seven experimental sites were set up in areas of moist meadow alpine tundra on Niwot Ridge. At each site, ten 1 m^2 plots were established where there was roughly even cover by two dominant plant species, Geum rossii (forb) and Deschampsia cespitosa (grass). In a factorial design, plots were assigned treatments of 1) removal of G. rossii, D. cespitosa, or control (no plant removal), and 2) nutrient addition of nitrogen (N), carbon (C), or control (no addition). A tenth plot was assigned a treatment of random biomass removal and no nutrient addition. Plots were visited annually to implement removal and fertilization treatments, and to measure plant species composition. Plant productivity was measured every other year and nematode communities once via 18S rRNA metabarcoding, together with soils. Carbon addition plots were dropped from the experiment in 2016.

openCC (other)Jun 2022View details →
edi52/100

Annual primary productivity at a Spartina patens-dominated salt marsh at Law's Point, Rowley River, Plum Island Ecosystem LTER, MA, 2000-2025.

Annual productivity is estimated from aboveground biomass data collected destructively from control and fertilized plots during the growing season at a Spartina patens-dominated salt marsh on the Rowley River within the Plum Island Ecosystem (PIE) LTER site, MA.

openCC (other)Dec 2025View details →
edi52/100

Aboveground biomass at control and fertilized plots in a Spartina patens-dominated salt marsh, Rowley River, Plum Island Ecosystem LTER, MA (2000-2025).

Aboveground biomass is determined destructively at control and fertilized sites approximately monthly during the growing season at a Spartina patens salt marsh on the Rowley River within the Plum Island Ecosystems (PIE) LTER site.

openCC (other)Dec 2025View details →
edi52/100

Plant heights at control and fertilized plots in a Spartina alterniflora-dominated marsh, Law's Point, Rowley River, Plum Island Ecosystem LTER, MA (1999-2025).

Plant heights are measured during the growing season in permanent plots at a Spartina alterniflora-dominated salt marsh on the Rowley River within the Plum Island Ecosystems (PIE) LTER site. Plant heights are converted to plant weight using an algorithm to generate a non-destructive estimate of aboveground plant biomass.

openCC (other)Dec 2025View details →
edi52/100

Aboveground plant biomass and density in control and fertilized plots in a Spartina alterniflora-dominated marsh, Rowley River, Plum Island Ecosystem LTER, MA (1999-2025).

Aboveground plant biomass and density is determined non-destructively during the growing season in permanent control and fertilized plots in a Spartina alterniflora-dominated salt marsh at Laws Point on the Rowley River within the Plum Island Ecosystems (PIE) LTER site.

openCC (other)Dec 2025View details →
edi52/100

Plant heights from permanent plots in a Spartina alterniflora-dominated marsh, Nelson Island, Parker River National Wildlife Refuge, Plum Island Ecosystems LTER, MA (2019-2025).

Plant heights are measured during the growing season in permanent plots at a Spartina alterniflora-dominated salt marsh on Nelson Island, Parker River National Wildlife Refuge, within the Plum Island Ecosystems (PIE) LTER site. Plant heights are converted to plant weight using an algorithm to generate a non-destructive estimate of aboveground plant biomass.

openCC (other)Dec 2025View details →
zenodo48/100

Generalization of a density-dependent ecosystem function in dominant aquatic macroinvertebrates

<div> <div> <p>This Zenodo record contains the supporting data and code for the publication 'Generalization of a density-dependent ecosystem function in dominant aquatic macroinvertebrates', published in Oikos (<a title="DOI to publication" href="https://doi.org/10.1111/oik.10774">https://doi.org/10.1111/oik.10774</a>). The data are described in detail in the corresponding publication. The data archive contains a ReadMe file, two text files with the empirical data, and a corresponding R script for analysis. All required data to reproduce the full analysis from the original publication are provided.</p> <p>In order to reproduce the analysis and figures, run&nbsp;<code>DensityDependenceAnalysis20240319.R</code>. Make sure that your working directory is the actual folder containing the data files&nbsp;<code>Data_Field.txt</code>&nbsp;and&nbsp;<code>Data_Lab.txt</code>. If run in Rstudio, this should happen automatically. Else this is easily achieved by (re)starting R (or R Studio) by double-clicking the R script file from the folder. The script will produce all the figures from the paper, organized in a folder&nbsp;<code>AnalysisYYYYMMDD</code>&nbsp;and two subfolders&nbsp;<code>CheckFigs</code>&nbsp;and&nbsp;<code>SuppFigs</code>. Figures are prepared as pixel graphics (PNG).</p> </div> </div>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record