Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,998
datasets available to search
ShareScore release 0.7.1
Dataset results
1,998 results for “microbial”
North Temperate Lakes LTER: Summary of Microbial Activity 2000 - 2002
Summary of Microbial Observatory data from the bacterial production, planktonic respiration and alakline phosphatase activity databases, plus bacterial cell counts from epifluorescence microscopy using DAPI cell stain. Information on integrated sample depth and incubation temperature is also included Sampling Frequency: fortnightly during ice-free season - every 6 weeks during ice-covered season Number of sites: 4
Microbial Bacterial Production in Lakes at North Temperate Lakes LTER 2000 - 2002
Net production of bacteria passing a 70 micron mesh, or bacteria passing a 1 micron mesh, calculated incorporation of 3H labeled leucine into cell proteins. Method based on microcentrifuge method of Smith, David C., and Farooq Azam. 1992. A simple, economical method for measuring bacterial protein synthesis rates in seawater using 3H-leucine, Marine Microbial Food Webs 6(2):107-114. Nanomolar treatments refer to leucine concentration for calculation of isotopic dilution and necessary leucine concentration to saturate uptake into bacterial cells. For a good example on how to calculate isotopic dilution see Pace. Michael L., and Jonathan J. Cole. Primary and bacterial production in lakes: are they coupled over depth? Plankton Research 16(6):661-672. Sampling Frequency: fortnightly during ice-free season - every 6 weeks during ice-covered season Number of sites: 4
Microbial Bacterial Respiration in Lakes at North Temperate Lakes LTER 2000 - 2002
Lake Mendota, Madison, Wisconsin. Bacterial respiration of bacteria passing a 70 micron mesh. Based on bacterial growth efficiency determined empirically on bacteria production and oxygen depletion on microbes passing a 1 micron mesh. All samples from integrated sample of epilimnion to thermocline or 12 meters, which ever was more shallow. Method based on Roland, F. and J. J. Cole. 1999. Regulation of bacterial growth efficiency in a large turbid estuary. Aquatic Microbial Ecology 20:31-38 Sampling Frequency: fortnightly during ice-free season - every 6 weeks during ice-covered season Number of sites: 1
Microbial Planktonic Respiration in Lakes at North Temperate Lakes LTER 2001
Respiration of total plankton passing a 70 micron mesh, and bacteria passing a 1 micron mesh, calculated from loss of oxygen in lake water incubated at in situ temperatures. Oxygen concentration was determined using the Winkler reaction with azide modification. Final product concentration determined via spectrometry or titration with sodium thiosulfate. Titrations that may have overrun the endpoint were not included. The following equation for calculation of dissolved oxygen concentration from titration of Winkler end product with thiosulfate (from Wetzel, R. G. and G. E. Likens. 1991. Limnological Analyses, 2nd ed. Springer-Verlag, New York). Thiosulfate with a molarity of 0.20 N was used for all titrations. mg O2 L-1 = (ml titrant)*(molarity of thiosulfate)*(8000)/((ml of sample titrated)*((ml of bottle -3)/(ml of bottle))). The following equation is for calculation of dissolved oxygen calculation from spectophotometric analysis of Winkler reaction end product (from Roland, F., N. F. Caraco, J. J. Cole. 1999. Rapid and precise determination of dissolved organic oxygen by spectrophotometry: Evaluation of interference from color and turbidity. Limnol. Oceanogr. 44(4):1148-1154). mg O2 L-1 = absorbance at 430 nm (in units of cm-1) * 8.1-0.41 Sampling Frequency: fortnightly during ice-free season - every 6 weeks during ice-covered season Number of sites: 4
Microbial Community Composition in Lakes - Taxonomic/Ecological characteristics of the sample at North Temperate Lakes LTER 2000 - 2007
Microbial community composition is inferred by a combination of automated ribosomal intergenic spacer analysis (ARISA) and PCR-generated clone library analysis. Clone libraries include both the 16S rRNA gene and the 16S-23S ribosomal intergenic spacer fragment. Phylogenetic assignments for individual ARISA fragments are obtained by comparing the ARISA fragment length from each clone to all of the profiles stored in our database. We have analyzed over 3900 clones obtained from 41 lakes that represent the range of trophic types found in temperate landscapes. Querying by a combination of taxonomic and ecological characteristics of the sample allows the user to retrieve sample information [sample IDs, sample dates, lake information (region, type, size, depth) and physical/chemical data (water temperature, clarity, pH, DOC, SUVA, TN, TP, nitrates/nitrites)] and clone information [clone IDs, sequence data, and characteristics of the sequence (length, chimera status, accession number, taxonomic affiliation)]. The data can be filtered by ecological characteristics of the sample [lake name, sample date, lake information (region, type, size, depth)] and taxonomic characteristics of the community members [clone ID, ARISA fragment length (raw or binned), and/or taxonomic characteristics (Phylum and Phylum-Class)]. The output can include links to individual sample records, which contain links to the taxonomic composition of the sample inferred by dynamically matching clones to ARISA fragments in the individual sample. The output can also include links to clone records directly (though this creates a very large number of lines in the output and is not recommended). Project ID's 30 Lakes - Survey of 30 lakes in northern and southern Wisconsin. June, August and October, 2002. See http://microbes.limnology.wisc.edu/lakes30.html. Lake Characteristics. CB0000 - Time series monitoring microbial community composition in Crystal Bog. 2000-2002. CBX_02 - Food web manipulation experiment i
Microbial Community Composition in lakes - Ecological characteristics of the sample at North Temperate Lakes LTER 2002 - 2007
Microbial community composition is inferred by a combination of automated ribosomal intergenic spacer analysis (ARISA) and PCR-generated clone library analysis. Clone libraries include both the 16S rRNA gene and the 16S-23S ribosomal intergenic spacer fragment. Phylogenetic assignments for individual ARISA fragments are obtained by comparing the ARISA fragment length from each clone to all of the profiles stored in our database. We have analyzed over 3900 clones obtained from 41 lakes that represent the range of trophic types found in temperate landscapes. Querying by ecological characteristics of the sample allows the user to retrieve sample IDs, sample dates, lake information (region, type, size, depth) and physical/chemical data (water temperature, clarity, pH, DOC, SUVA, TN, TP, nitrates/nitrites). The data can be filtered by lake name, sample date, lake information (region, type, size, depth), and physical/chemical data (water temperature, clarity, pH, DOC, SUVA, TN, TP, nitrates/nitrites). The output includes links to individual sample records, which contain links to the taxonomic composition of the sample inferred by dynamically matching clones to ARISA fragments in the individual sample
Microbial Community Composition in Lakes - Taxonomic characteristics of the clones at North Temperate Lakes LTER 2000 - 2007
Microbial community composition is inferred by a combination of automated ribosomal intergenic spacer analysis (ARISA) and PCR-generated clone library analysis. Clone libraries include both the 16S rRNA gene and the 16S-23S ribosomal intergenic spacer fragment. Phylogenetic assignments for individual ARISA fragments are obtained by comparing the ARISA fragment length from each clone to all of the profiles stored in our database. We have analyzed over 3900 clones obtained from 41 lakes that represent the range of trophic types found in temperate landscapes. Querying by taxonomic characteristics of the clone allows the user to retrieve clone IDs, sequence data, and characteristics of the sequence (length, chimera status, accession number, taxonomic affiliation). The data can be filtered by clone ID, ARISA fragment length (raw or binned), and/or taxonomic characteristics (Phylum and Phylum-Class). The output includes links to individual clone records, which contain more detailed information about how the clone was generated (researcher, library ID, project ID, primer sets used, etc.).
Microbial narrow-escape is facilitated by wall interactions: Simulation Supplementary material
<p>Simulation codes and simulation results for the paper "Microbial narrow-escape is facilitated by wall interactions".</p>
Peat characteristics, microbial PLFA, and fungal and actinobacterial sequences from Lakkasuo peatland drainage experiment, year 2004
<p>We analysed the response of microbial communities, characterized by phospholipid fatty acids (PLFAs), and fungal and actinobacterial communities, characterized by PCR-DGGE fingerprinting and direct sequencing, to changing hydrological conditions at three different sites in the boreal peatland complex Lakkasuo in southern Finland. Additionally, several peat characteristics were measured. The experimental design involved undrained controls as well as short-term (3 years) and long-term (43 years) water-level drawdown. The sites were, in their undrained state, a herb-rich sedge fen, a sedge fen, and a bog with hummock-lawn-hollow microtopography.</p> <p>Codes are explained in the Notes sheets of the Excel files. The contents of the csv files are identical to the corresponding Excel file data sheets.</p> <p>Please check the decimal separator! Comma is used in Finland, and that may have been carried over. All commas in data columns are decimal separators.</p>
Microbial and soil moisture impacts of compost amendments and rainfall pulses in a degraded dryland soil, Arizona, 2021-2023
Compost, an organic soil amendment, has been proposed to increase soil carbon storage and water-holding capacity in drylands, and this management strategy may be particularly impactful in degraded drylands with low soil organic content. Compost additions and rainfall variability may interact to affect soil moisture, which is an important catalyst for soil microbial activity. This dataset is from a study that investigated how variable compost application amounts and simulated rainfall pulses affect soil moisture, microbial activity, and carbon content in a laboratory incubation study. Soils were amended with different amounts of compost (0, 0.35, and 0.70 g cm -2) and water pulses (5, 10, and 15 mm) in a full-factorial design. Each treatment received the same cumulative amount of water throughout the incubation, but pulses occurred at different frequencies (every 5, 10, and 15 days). Soil moisture content and microbial respiration were measured daily. Soil carbon content was measured at the end of the experiment.
Microbial Sampling from Shark River Slough and Taylor Slough, Everglades National Park, South Florida, USA (FCE LTER), January 2001 - December 2023
We collected monthly observations of water microbial characteristics from each of the FCE LTER sites beginning in February 2001. Bacteria counts using DAPI epifluorescence, heterotrophic bacteria production with tritiated Thymidine uptake, and algal energetics and chl a using the PAM fluorometer were determined from each of the water samples. Bacteria counts in Shark River Slough in the freshwater end of the transect revealed highest bacteria numbers upon marsh rewetting after the dry down. The highest overall count of bacteria was found at the most freshwater point adjacent to the perimeter of the water conservation area but only in March of 2001. Peak numbers of bacteria in the saline end of Shark River transect occurred in December and July. There is some evidence that these higher bacteria numbers in the SR transect are associated with higher salinities. Highest bacteria counts along the Taylor slough transect were observed at the mouth of the Taylor River with peak counts in the wet season in October. Bacteria densities also increased from the eastern to western end of the transect in Florida Bay. Heterotrophic bacterial production did not relate closely bacteria numbers and exhibited highest numbers in the dry season. There was no discernible relationship between salinity and bacteria production as seen in bacteria numbers. Monthly observations using PAM fluorometry revealed that brown algae was the greatest contribution to the algal chal a pool and this relationship was consistent across all FCE LTER sites. Converse to the pattern of bacteria abundance along the Shark River transect, there was evidence that algal biomass decreases from freshwater to estuary. In the Shark River, algal energetics increased from marsh to estuary as well as in the panhandle region of the Taylor River transect. From these observations we conclude the behavior of the microbial loop and interactions generated between functional guilds is highly variable along individual transects an
Physical and microbial processing of dissolved organic nitrogen (DON) (Salinity Experiment) along an oligotrophic marsh/mangrove/estuary ecotone (Taylor Slough and Florida Bay) for August 2003 in Everglades National Park (FCE), South Florida, USA
A better understanding of the biogeochemical cycling of nutrients entering Florida Bay is a key issue regarding the restoration of the Everglades. In addition to precipitation, the other major source of freshwater to Florida Bay is from Taylor Slough and the C-111 Basin in the northeast section of the Bay. While it is known that these areas deliver significant amounts of N to the Bay, a significant portion of this is in the form of dissolved organic N (DON). The sources, environmental fate and bioavailability to microorganisms of this DON are however, not known. Should this DON be readily available, any increased load as a function of restoration changes might have an impact on internal phytoplankton bloom dynamics. No significant flocculation or precipitation of DOM occurred with increase in salinity, meaning that terrestrial DOM does not get trapped in the sediments but stays in the water column where it subjected to photolysis and advective transport. Sunlight has a significant effect on the chemical characteristics of DOM. While the DOC levels did not change significantly during photo-exposure, the optical characteristics of the DOM were modified. The environmental implications of this are conflicting: photo-induced polymerization may stabilize the DOM by reducing its bioavailability while photolysis may make the DOM more labile. Overall, DON bioavailability was relatively low in this region. Even though the amount of DON loaded to the bay may be significant, the fraction of DON available for microbial cycling is much smaller. The amount of N supplied by recycling may be a significant portion of the total DIN pool. All this must be considered in context with the proposed CERP modifications to flows. As of the latest initial Comprehensive Everglades Restoration Project (CERP) update, the flows to Taylor Slough and C-111/Panhandle Basis are not predicted to change very much from base conditions. Therefore we do not expect any great increases in TN loading in this
Physical and microbial processing of dissolved organic nitrogen (DON) (Photodegradation Experiment) along an oligotrophic marsh/mangrove/estuary ecotone (Taylor Slough and Florida Bay) for August 2003 in Everglades National Park (FCE), South Florida, USA
A better understanding of the biogeochemical cycling of nutrients entering Florida Bay is a key issue regarding the restoration of the Everglades. In addition to precipitation, the other major source of freshwater to Florida Bay is from Taylor Slough and the C-111 Basin in the northeast section of the Bay. While it is known that these areas deliver significant amounts of N to the Bay, a significant portion of this is in the form of dissolved organic N (DON). The sources, environmental fate and bioavailability to microorganisms of this DON are however, not known. Should this DON be readily available, any increased load as a function of restoration changes might have an impact on internal phytoplankton bloom dynamics. No significant flocculation or precipitation of DOM occurred with increase in salinity, meaning that terrestrial DOM does not get trapped in the sediments but stays in the water column where it subjected to photolysis and advective transport. Sunlight has a significant effect on the chemical characteristics of DOM. While the DOC levels did not change significantly during photo-exposure, the optical characteristics of the DOM were modified. The environmental implications of this are conflicting: photo-induced polymerization may stabilize the DOM by reducing its bioavailability while photolysis may make the DOM more labile. Overall, DON bioavailability was relatively low in this region. Even though the amount of DON loaded to the bay may be significant, the fraction of DON available for microbial cycling is much smaller. The amount of N supplied by recycling may be a significant portion of the total DIN pool. All this must be considered in context with the proposed CERP modifications to flows. As of the latest initial Comprehensive Everglades Restoration Project (CERP) update, the flows to Taylor Slough and C-111/Panhandle Basis are not predicted to change very much from base conditions. Therefore we do not expect any great increases in TN loading in this
Microbially-Mediated Transformations of Estuarine Dissolved Organic Matter from Doboy Sound, Sapelo Sound and the Altamaha River near Sapelo Island, Georgia in 2012-2013
Dissolved organic matter (DOM) samples were collected from a marsh-dominated estuarine system (Altamaha River estuary, Doboy Sound and Sapelo Sound near Sapelo Island, Georgia), inoculated with native microbiological samples, and incubated in the dark for up to 70 days. Control and incubated DOM samples were then anlyzed using ultrahigh resolution mass spectrometry to investiage microbially-mediated transformations at the molecular level. Samples were collected at high tide in November 2012 and May 2013.
Hubbard Brook Nitrogen Oligotrophication (HBNO): Microbial Biomass and Activity, 2021-2023
The goal of this project is to test the overarching hypothesis that positive feedback mechanisms involving changes in seasonal cycles that diminish N availability to plants such that plant N demand is not met by soil N availability in northern forests. Specifically, we hypothesize that increasing N demand by plants (induced by increasing temperatures, longer growing seasons, and other environmental changes) leads to greater N resorption by trees in autumn, increased C:N in litter, and greater net immobilization of N by soil microbes in the following spring. However, the timing of snowmelt and soil freezing in spring may further affect net mineralization and N availability for plants. These hypotheses are being tested with a combination of observational, experimental, and modeling approaches at Hubbard Brook Experimental Forest in New Hampshire: 1) measurements at 14 previously established sites along an elevation/aspect climate gradient; 2) litter and snow manipulation experiments at six sites along the climate gradient to create variation in soil climate conditions and microbial N immobilization during spring. We leveraged 14 sites previously established along an elevation and aspect-driven climate gradient at Hubbard Brook as a “natural climate experiment" to test our hypothesis that a positive feedback between N cycling during fall senescence and spring contributes to declining N availability in northern forests. This elevation gradient encompasses variation in mean annual air temperature of ~2.5 °C that is similar to the change projected to occur with climate change over the next 50–100 years in the northeastern U.S. There is relatively little variation in soils along the gradient. We are utilizing three sites at higher elevation (~550-660 m, north facing) and three sites at lower elevation (~375-500 m, south facing) for the litter and snow manipulation experiments to maximize the differences in temperature among the 14 sites. Litterbox manipulation: The objecti
OMS01 Microbial and soil data in the Belowground Plot Experiment at konza prairie since 2017
Data describe soil chemistry collected during the 2017-2021 growing seasons at the Belowground Plot Experiment to assess for legacies. Whole plot-scale nitrogen fertilization at the Experiment ceased in 2017. Four subplots within each historically fertilized plot were set up to continue the annual fertilization treatment for soil chemistry.
MCR LTER: Coral Reef: Water Column microbial community data in lagoons of Moorea, French Polynesia
Microbes process a significant fraction of organic material in marine systems, and the composition and activity of their communities are strongly modulated by fluctuations in nutrient availability. To investigate the distribution and dynamics of microbial communities in tropical lagoon ecosystems, bacteria and archaea were quantified from water column samples collected in lagoons around Moorea, French Polynesia during May 2021, April 2022, and April 2023 using eDNA sequencing. Environmental DNA was extracted and sequenced to assess microbial community composition, with data processed using the QIIME 2 platform (version 2023.7). Microbial communities in fringing reef habitats differed from those in mid-lagoon and back reef sites. These differences in microbial communities were related to patterns of water column nutrients and fluorescent dissolved organic matter, providing a baseline for understanding how lagoon microbial communities respond to spatial and temporal variability in reef environments.
Using the Tea Bag Index to unravel how interactions between an antibiotic (Trimethoprim) and endocrine disruptor (17a-estradiol) affect aquatic microbial activity: Supporting Dataset 1
<p>The constant release of complex mixture of pharmaceuticals, including antimicrobials and endocrine disruptors, into the aquatic environment. These have the potential to affect aquatic microbial metabolism and alter biogeochemical cycling of carbon and nutrients. We used the Tea Bag Index (TBI) for decomposition within a series of contaminant exposure experiments to test how interactions between an antibiotic (trimethoprim) and endocrine disruptor (17a-estradiol) affects microbial activity in an aquatic system. The TBI is a citizen science tool used to test microbial activity by measuring the differential degradation of green and rooibos tea as proxies for labile and recalcitrant organic matter decomposition. Here we present the raw data on pharmaceutical exposures and the mass loss of the Rooibos and Green tea bags within the experiment. From Tea Bag mass loss we then calculated the Stabilisation Factor (S) and Initial Decomposition Rate of the labile organic matter fraction.</p>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
Inter-Chemical Correlation results for the study: HHEARx2017-1977 (Microbial and metabolic variations mediate the influence of childhood and adolescent EDC and trace element exposure on breast density.)
Title: Microbial and metabolic variations mediate the influence of childhood and adolescent EDC and trace element exposure on breast density. <br>Species: Homo sapiens <br>Number of samples: 1116 <br>Number of named analytes: 41 <br>Datasource url: https://hheardatacenter.mssm.edu/PublicFile/ViewPublicFile?projectid=46 <br>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.