Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,293

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,293 results for “gene sequencing”

Learn how ShareScore rates datasets ↗
zenodo48/100

FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions

<p>Fasta sequence correspond to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequence corresponds&nbsp;to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;<em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1&nbsp;(Colanero et al., 2020)], <em>S. chilense&nbsp;</em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), &nbsp;Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequences correspond to the MYB encoding genes&nbsp;<em>An2-like </em>and <em>Ant1</em>. The genomic&nbsp;sequences were combined correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em>&nbsp;accession LA1141 (this study),&nbsp;<em>S.&nbsp;lycopersicum</em>&nbsp;variety OH8245 (this study),&nbsp;<em>S. lycopersicum</em>&nbsp;variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)],&nbsp;and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018),&nbsp;<em>S. lycopersicum&nbsp;</em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)],&nbsp;<em>S. chilense</em>&nbsp;accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at&nbsp;<a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI) (available at NCBI:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

DADA2 formatted 16S rRNA gene sequences for both bacteria & archaea

<p><strong><em>This version is to stay up to date with the improvements and increase in 16S rRNA gene sequences (SSU) added to the GTDB release 220.&nbsp; Please read this post for the stats on the updates. </em></strong><strong><em>https://gtdb.ecogenomic.org/stats/r220 </em></strong><strong><em>.</em></strong><strong><em> </em></strong></p> <p><strong><em>There has been no change to the RDP-RefSeq reference database please use previous versions.</em></strong></p> <p><strong><em>If anyone has concerns&nbsp;with MAG extracted 16S rRNA gene contamination concerns, then I suggest that they contact the curators of GTDB themselves because it is outside of my role with these resources designed for DADA2 usage only. </em></strong></p> <p><strong><em>Another concern that was raised was the orientation of the DB sequences, to get past this problem please use the tryRC = TRUE argument in the assignTaxonomy command within DADA2, this will search your ASVs in the reverse complement as well.&nbsp;&nbsp;</em></strong></p> <p>The bacterial and archaeal 16S rRNA gene sequence databases were collated from various sources and formatted to use the "assignTaxonomy" command within the DADA2 pipeline. The data was converted to suite DADA2 format by Alishum Ali.</p> <ol> <li>Genome Taxonomy Database (GTDB): The new version of our dada2 formatted GTDB reference sequences now contains 58102 bacteria and 3672 archaea full 16S rRNA gene sequences. If you wonder why there are fewer species with 16S rRNA, that is because some metagenomics-assembled genomes (MAGs) lack the 16S gene and thus cannot be extracted.&nbsp; The database was downloaded from <a href="https://data.ace.uq.edu.au/public/gtdb/data/releases/release95/">https://data.ace.uq.edu.au/public/gtdb/data/releases/</a> on 24/10/2024. Please read the release notes and file descriptions.&nbsp;</li> </ol> <p>The formatting to DADA2 was done using simple awk bash scripts. The script takes as input a fasta file and a tab-delimited taxonomy file (slightly edited to remove special characters) and then it outputs a fasta file with all 7 taxonomy ranks separated by ";" as required for DADA2 compatibility. Additionally, we have concatenated the unique sequence GTDB ID to the species entry (but replaced the "." with an " _". We see this as an important QC step to highlight the issues/confidence associated with short-read taxonomy assignment at the finer rank levels.</p> <p>Also, this update includes two other files that you can use with the assignTaxonomy and addSpecies commands in DADA2.</p>

opencc-by-4.0Jan 2019View details →
zenodo48/100

DADA2 formatted eHOMD 16S rRNA gene sequences databse

<p>eHOMD Refseq database (V15.22) formated to be used with dada2 <em>i.e.</em>, dada2::assignTaxonomy(seqs, &quot;eHOMD_RefSeq_dada2_V15.22.fasta.gz&quot; ) and dada2::addSpecies(taxa, &quot;eHOMD_RefSeq_dada2_assign_species_V15.22.fasta.gz&quot;, verbose=TRUE)</p> <p>Alternatively, you could use the metabaRpipe R package to directly update the taxonomy of a phyloseq object see: https://github.com/fconstancias/metabaRpipe#2-addingreplacing-taxonomical-table-in-a-phyloseq-object</p> <p>Example below:<br> source(&quot;https://raw.githubusercontent.com/fconstancias/metabaRpipe-source/master/Rscripts/functions.R&quot;)</p> <p>readRDS(&quot;dada2/phyloseq.RDS&quot;) %&gt;%<br> &nbsp; phyloseq_dada2_tax(physeq = .,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; threshold = 60,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; db =&quot;~/metabaRpipe/databases/eHOMD_RefSeq_dada2_V15.22.fasta.gz&quot;,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; db_species =&quot;~/metabaRpipe/databases/eHOMD_RefSeq_dada2_assign_species_V15.22.fasta.gz&quot;,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nthreads = 2,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; full_return = FALSE) -&gt; physeq_eHOMD_tax</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of 16S and 18S rRNA genes from bacterial and protistan planktonic communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2011-2013

Microbial communities in the coastal Arctic Ocean experience extreme variability in organic matter and inorganic nutrients driven by seasonal shifts in sea ice extent and freshwater inputs. Lagoons border more than half of the Beaufort Sea coast and provide important habitats for migratory fish and seabirds; yet, little is known about the planktonic food webs supporting these higher trophic levels. To investigate seasonal changes in bacterial and protistan planktonic communities, amplicon sequences of 16S and 18S rRNA genes were generated from samples collected during periods of ice-cover (April), ice break-up (June), and open water (August) from shallow lagoons along the eastern Alaska Beaufort Sea coast from 2011 through 2013. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA530074 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA530074. This data package is associated with the following publication: Kellogg CTE, McClelland JW, Dunton KH and Crump BC (2019) Strong Seasonality in Arctic Estuarine Microbial Food Webs. Front. Microbiol. 10:2628. doi: 10.3389/fmicb.2019.02628 Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provided site codes (column "site_name" here) and collection dates (column "collection_date" here) in each dataset. Note that the site codes in this package are without hyphens (e.g. JAA) while site codes in the above environmental data package have hyphens (e.g. JA-A). Instead of citing this package which is jus

openCC0Jan 2020View details →
edi48/100

16S rRNA gene sequence accessions from discrete water column samples collected from lakes in the McMurdo Dry Valleys, Antarctica (2013-2023, ongoing)

An important component of the McMurdo Dry Valleys Long Term Ecological Research (MCM LTER) project is monitoring spatial and temporal patterns in the biological composition of perennially ice-covered lakes in Antarctica’s McMurdo Dry Valleys. This data package contributes to this core research area by providing a curated table linking 16S rRNA gene sequence accession numbers archived in NCBI to MCM LTER limnological sampling campaigns conducted at specific depths along the water column of Lakes Fryxell, Hoare, Bonney, and Miers. These data enable integration of microbial community data with co-collected biological, chemical, and physical measurements.

openCC (other)Dec 2025View details →
zenodo44/100

Transcriptome analysis of the effect of over-expressing H2A.J mutants in proliferating WI38 fibroblasts for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences

<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>H2A.J differs from canonical H2A only by a valine at position 11 instead of alanine, and the 7 C-terminal amino acids containing a potential minimal phosphorylation site SQ for DNA-damage response kinases. To test the functional importance of these H2A.J-specific sequences, we mutated Val-11 to Ala as is found in all canonical H2A sequences, and we mutated Ser-123 to either Glu to mimic a phospho-serine residue or to Ala to prevent phosphorylation. We also substituted the C-terminus of H2A.J with the C-terminus of H2A. These mutants, WT-H2A.J and canonical H2A-type1 were ectopically expressed in proliferating fibroblasts, and their microarray transcriptomes were compared to that of proliferating and senescent fibroblasts without ectopic histone expression. Genome-wide transcriptome analysis indicated that senescent fibroblasts clustered distinctly from proliferating fibroblasts, and proliferating fibroblasts expressing the H2A.J-V11A and H2A.J-S123E mutants clustered distinctly from fibroblasts expressing the other H2A.J mutants, WT-H2A.J, and H2A. Hallmark gene set enrichment analysis of the transcriptomes of fibroblasts expressing H2A.J-V11A or H2A.J-S123E versus control proliferating fibroblasts indicated that they showed the same highly significant enrichment for the Epithelial-Mesenchyme Transition, TNF-Alpha Signaling Via NF-kB, and Inflammatory Response gene sets. Notable inflammatory genes including IL1A, IL1B, IL6, CXCL8, and CCL2 are contained in these gene sets and are often induced in senescence as part of the senescence-associated secretory phenotype. Heat maps showed that the H2A.J-V11A and H2A.J-S123E mutants were particularly apt at activating the expression of these inflammatory genes in proliferating fibroblasts</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

FASTA file containing the MYB encoding gene An2-like and Ant1 coding sequences corresponding to wild and cultivated tomato accessions

<p>The coding sequence (CDS) of the MYB encoding genes&nbsp;<em>Ant1</em> and <em>An2-like</em>.&nbsp;Sequences were retrieved from regions corresponding to the<em> Aft</em> locus from <em>Solanum galapagense </em>accession&nbsp;LA1141, <em>S. lycopersicum</em> variety OH8245, and&nbsp;&nbsp;84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences were compared to available&nbsp;CDS available from the Sol genomics network (SGN) and&nbsp;the National Center for Biotechnology Information. The CDS was&nbsp;retrieved from <em>S. lycopersicum</em>&nbsp;variety&nbsp;Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum</em> accession LA1996 [MN242011.1, EF433417.1( Sapir et al., 2008; Colanero et al., 2020)], and&nbsp;<em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], The orthologous CDS&nbsp;corresponding&nbsp;to the <em>Aft </em>MYB encoding genes from&nbsp;<em>Solanum tuberosum</em> L. Group Phureja clone DM1-3 genome (PGSC DM v4.03 Pseudomolecules) was retrieved from the Potato Genome Sequence Consortium (PGSC: Potato Genome Sequencing Consortium et al., 2011), and the Capsicum annum cv. CM334 genome was retrieved from&nbsp;<em>Capsicum annuum </em>cv CM334 genome chromosome release 1.55 (Hulse-Kemp et al. 2018). These CDS&nbsp;were obtained using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at https://solgenomics.net/tools/blast/). Comparison of syntenic chromosomal regions using known positions of tomato, potato, and pepper markers with comparative map viewer from&nbsp; SGN: (available at https://solgenomics.net/cview) on chromosome 10,&nbsp;was used as a quality check for S.<em> tuberosom</em> and <em>C. annuum.</em> Orthologous&nbsp;CDS corresponding to&nbsp;<em>Salvia miltiorrhiza,&nbsp;Arabidopsis thaliana</em>, [NM_105308.2, NM_105310.4 (Teng et al., 2005, Cominelli et al., 2008; Beradini et al., 2015)] were chosen based on tomato <em>Aft</em> sequence homology and gene annotations of&nbsp;positive R2R3 MYB regulation of anthocyanin. The CDS&nbsp;corresponding&nbsp;to the <em>Aft</em> genes were retrieved from the CDS reference genomes available from the Sol Genomics Network SGN: Tomato Genome CDS (ITAG release 4.0), Potato PGSC DM v3.4 CDS sequences, <em>Capsicum annuum </em>cv CM334 Genome CDS (release 1.55), or from the National Center for Biotechnology Information (NCBI: https://www.ncbi.nlm.nih.gov) reference sequences (RefSeq) section of the Genbank records. When accessed from Genank records, the CDS sequence was extracted from the &ldquo;features&rdquo; section and exported as a FASTA file.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Targeted Re-sequencing Identifies Candidate Fusiform Rust Resistance Genes in Loblolly Pine

<p>A fasta file containing the subset of the v2.01 Pita genome in addition to the novel NLR genes that were targeted by hybridization probes.&nbsp;</p> <p>A bed file describing the intervals targeted by the hybridization&nbsp;probes.</p> <p>Trinity assemblies of the 30 RNAseq libraries along with predictions by transdecoder of CDS and peptide sequences from those trinity assemblies.&nbsp;&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Human intestinal Bacteria Collection (HiBC): 16S rRNA gene sequences

<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the sequences of the 16S rRNA gene sequences of the isolates in the FASTA nucleotide format. Sequences ending in Sanger were obtained using the Sanger dideoxy sequencing technology. Sequences ending in Genome were obtained from the genome sequence using barrnap.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

16S rRNA sequencing gene datasets for CRC data

<p>Used datasets:&nbsp;</p> <table> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>Dataset</th> <th>16S rRNA Region</th> <th>Control (n)</th> <th>Adenoma (n)</th> <th>CRC (n)</th> <th>Available metadata</th> </tr> </thead> <tbody> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4823848/">Baxter</a></td> <td>V4</td> <td>171</td> <td>198</td> <td>120</td> <td>Gender, age, weight, height, BMI, country, race</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4221363/">Zackular</a></td> <td>V4</td> <td>30</td> <td>30</td> <td>30</td> <td>Gender, age, weight, height, BMI, country, race, FOBT, medication</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4299606/">Zeller</a></td> <td>V4</td> <td>50</td> <td>38</td> <td>41</td> <td>Gender, age, BMI, country, FOBT</td> </tr> <tr> <td><strong>TOTAL</strong></td> <td>V4</td> <td>251</td> <td>266</td> <td>191</td> <td><em>All of the above</em></td> </tr> </tbody> </table> </th> </tr> </thead> <tbody> <tr> <td>&nbsp;</td> </tr> </tbody> </table> <p>Data processing &amp; sharing</p> <p>All datasets were processed using&nbsp;<a href="https://docs.qiime2.org/2021.11/">qiime2</a>&nbsp;pipeline with&nbsp;<a href="https://benjjneb.github.io/dada2/">DADA2</a>&nbsp;for Sequence quality control and feature table construction and&nbsp;<a href="https://www.arb-silva.de/">SILVA</a>&nbsp;database for taxonomic assignment, and then a <em>phyloseq </em>object was constructed.</p> <ul> <li>Abundance table at genus level is in file <em>genus.csv</em> (Sample counts with NO filtering).</li> <li>Clean metadata is in <em>metadata.csv</em> file (Countries: CA - Canada. USA - United States of America. FRA - France.)</li> <li>Phyloseq object is in file <em>physeq.RDS</em> (Saved as an RDS object in R)</li> </ul> <p>More information is&nbsp;<a href="https://hackmd.io/nbsLqCLlSNSRFc5RBX9c5Q?view">here</a>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Vitamin_B12_related_genes_and_FASTA_sequences

<p>These two files are related to a&nbsp;submitted research paper on the study of a 7-year metagenomic time series carried monthly&nbsp;in a northwestern Mediterranean coastal site (SOLA, Banyuls-sur-Mer, FRANCE) :</p> <p><em>Seasonal succession of different vitamin B12 biosynthesis pathways and producers in coastal marine microbial communities.&nbsp;</em></p> <p>This study focuses on prokaryotes involved in vitamin B12 metabolism (biosynthesis, transport, remodeling) and on the metabolic pathways involved in these processes.</p> <p>The dataset in <strong>.xlsx</strong> format represents the abundance table of all the genes associated with 82 KEGGs involved in vitamin B12 metabolism, and the file in <strong>.fasta</strong> format corresponds to the FASTA sequences associated with these genes.</p>

opencc-by-4.0Dec 2022View details →
edi44/100

Inventory of soil prokaryotic and fungal microbiome (via 16S rRNA gene amplicons and ITS sequencing) from Shark River Slough and Taylor Slough, Everglades National Park (FCE LTER), Florida, USA, February 2019 - October 2020

Global sea-level rise is transforming coastal ecosystems, especially freshwater wetlands, in part due to increased saltwater exposure, leading to change in soil microbial communities and many important biogeochemical processes. Given the high spatial and temporal heterogeneity in coastal wetlands, especially in tropical or subtropical climates characterized by seasonal temperature, precipitation, and tidal fluctuations, it remains unclear which environmental factors influence the compositions of soil microbial communities in wetlands affected by varying degrees of sea-water intrusion. To address this, a two-year survey was conducted on microbial community structure in submerged surface soils from 14 wetland sites across the Florida Everglades, representing three major ecosystem types, i.e. freshwater marshes, mangrove forests, and seagrass meadows. Bulk surface soil samples of each site were collected from February 2019 to October 2020 to cover dry and wet seasons. In addition to bulk soil samples, soil cores were collected from each site in August 2020 to assess vertical gradients of microbial communities. The dataset contains amplicon sequencing data of 16S rRNA gene (both bulk soil and soil cores) and ITS gene (only the bulk soil). The 2019 to 2020 data are published in Zhao et al. 2023. A detailed list of sequence data and their accession numbers in GenBank is provided, and data collection is complete. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA804243 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804243), PRJNA804246 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804246), and PRJNA804228 (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA804228). This data package is associated with the following publication: Zhao, J., Chakrabarti, S., Chambers, R., Weisenhorn, P., Travieso, R., Stumpf, S., Standen, E., Briceno, H., Troxler, T., Gaiser, E., Kominoski, J., Dhillon, B., & Martens-H

openCC (other)Feb 2024View details →
zenodo40/100

GenBank accession numbers of the four marker genes and associated voucher specimens/tissues that were used in this study. For more details see Guo et al. (2014). Sequences of species in bold are unpublished and were provided by P. Guo as personal communication in Rediscovery of Andrea's keelback, Hebius andreae (Ziegler & Le, 2006): First country record for Laos and phylogenetic placement

GenBank accession numbers of the four marker genes and associated voucher specimens/tissues that were used in this study. For more details see Guo et al. (2014). Sequences of species in bold are unpublished and were provided by P. Guo as personal communication

opencc-by-4.0Mar 2019View details →
zenodo40/100

Xpresso: Predicting gene expression levels from genomic sequences

<p>Xpresso: Predicting gene expression levels from genomic sequences<br> <br> More info at:<br> Publication:&nbsp;https://doi.org/10.1016/j.celrep.2020.107663<br> Website: https://xpresso.gs.washington.edu/<br> Github:&nbsp;https://github.com/vagarwal87/Xpresso</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

16S gene and ASV sequences of bacteria isolated from soil and the phyllosphere of Arabidopsis thaliana

<p>Data from &quot;Induction of antibiotic specialized metabolism by co-culturing in a collection of phyllosphere bacteria&quot; by Qi et al.&nbsp;</p> <p>- 16S gene sequences of bacteria isolated from soil and the phyllosphere of Arabidopsis thaliana in FASTA format.&nbsp;</p> <p>- Filtered OTU (ASV) table across all samples&nbsp;</p> <p>- ASV sequences</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Data for: Faster rates of molecular sequence evolution in reproduction-related genes and in species with hypodermic sperm morphologies

<p>This repository contains a record of analysis scripts and sequence alignments used for the analyses presented in the manuscript.</p> <p>Some of the R scripts depend on supplementary tables associated with the manuscript.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Fig. 2 in Phylogenetic Relationships Of Malayan And Malagasy Pygmy Shrews Of The Genus Suncus (Soricomorpha: Soricidae) Inferred From Mitochondrial Cytochrome B Gene Sequences

Fig. 2. The neighbour-joining (A) and Bayesian (B) trees for Suncus inferred from 1140 base-pairs of cytochrome b gene sequence. Bootstrap and posterior probability values are given above branches.

opencc-by-4.0Aug 2011View details →
zenodo40/100

Fig. 1 in Phylogenetic Relationships Of Malayan And Malagasy Pygmy Shrews Of The Genus Suncus (Soricomorpha: Soricidae) Inferred From Mitochondrial Cytochrome B Gene Sequences

Fig. 1. Male Malayan pygmy shrew (Suncus malayanus) captured in the Cameron Highlands, Pahang, Peninsular Malaysia, in a pitfall trap set on the forest floor. Notice the characteristic large ears and dark fine pelage.

opencc-by-4.0Aug 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record