Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

52

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

52 results for “GenBank”

Learn how ShareScore rates datasets ↗
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of 16S and 18S rRNA genes from bacterial and protistan planktonic communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2011-2013

Microbial communities in the coastal Arctic Ocean experience extreme variability in organic matter and inorganic nutrients driven by seasonal shifts in sea ice extent and freshwater inputs. Lagoons border more than half of the Beaufort Sea coast and provide important habitats for migratory fish and seabirds; yet, little is known about the planktonic food webs supporting these higher trophic levels. To investigate seasonal changes in bacterial and protistan planktonic communities, amplicon sequences of 16S and 18S rRNA genes were generated from samples collected during periods of ice-cover (April), ice break-up (June), and open water (August) from shallow lagoons along the eastern Alaska Beaufort Sea coast from 2011 through 2013. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA530074 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA530074. This data package is associated with the following publication: Kellogg CTE, McClelland JW, Dunton KH and Crump BC (2019) Strong Seasonality in Arctic Estuarine Microbial Food Webs. Front. Microbiol. 10:2628. doi: 10.3389/fmicb.2019.02628 Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provided site codes (column "site_name" here) and collection dates (column "collection_date" here) in each dataset. Note that the site codes in this package are without hyphens (e.g. JAA) while site codes in the above environmental data package have hyphens (e.g. JA-A). Instead of citing this package which is jus

openCC0Jan 2020View details →
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of metagenomic DNA sequence analyses of bacterial and archaeal water column communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2012

In contrast to temperate systems, Arctic lagoons that span the Alaska Beaufort Sea coast face extreme seasonality. Nine months of ice cover up to ∼1.7 m thick is followed by a spring thaw that introduces an enormous pulse of freshwater, nutrients, and organic matter into these lagoons over a relatively brief 2–3 week period. Prokaryotic communities link these subsidies to lagoon food webs through nutrient uptake, heterotrophic production, and other biogeochemical processes, but little is known about how the genomic capabilities of these communities respond to seasonal variability. This study characterizes the metabolic capabilities of microbial communities across three seasons in two lagoons and one open coastal site along the eastern Alaska Beaufort Sea coast. We used metagenomic DNA sequence data of bacterial and archaeal water column communities to identify genes of relevant biogeochemical pathways. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA642637 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA642637. This data package is associated with the following publication: Baker, Kristina D., Colleen T. E. Kellogg, James W. McClelland, Kenneth H. Dunton, and Byron C. Crump. “The Genomic Capabilities of Microbial Communities Track Seasonal Variation in Environmental Conditions of Arctic Lagoons.” Frontiers in Microbiology 12 (2021). https://doi.org/10.3389/fmicb.2021.601901. Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provi

openCC0Apr 2021View details →
zenodo40/100

GenBank accession numbers of the four marker genes and associated voucher specimens/tissues that were used in this study. For more details see Guo et al. (2014). Sequences of species in bold are unpublished and were provided by P. Guo as personal communication in Rediscovery of Andrea's keelback, Hebius andreae (Ziegler & Le, 2006): First country record for Laos and phylogenetic placement

GenBank accession numbers of the four marker genes and associated voucher specimens/tissues that were used in this study. For more details see Guo et al. (2014). Sequences of species in bold are unpublished and were provided by P. Guo as personal communication

opencc-by-4.0Mar 2019View details →
zenodo40/100

Appendix. List of the 28S and 16S rRNA sequences recovered from GenBank. 28S = 28S rRNA GenBank accession number; 16S = 16S rRNA GenBank accession number. in Genetic and morphological evidence for cryptic species in Macrobrachium australe and resurrection of M. ustulatum (Crustacea, Palaemonidae)

Appendix. List of the 28S and 16S rRNA sequences recovered from GenBank. 28S = 28S rRNA GenBank accession number; 16S = 16S rRNA GenBank accession number.

opencc-by-3.0Feb 2017View details →
zenodo40/100

Genetic data and underlying taxa and GenBank sources of diatoms used in phylogenetic analysis for the diatom genus Nupela

<p>Supplementary material for the&nbsp;manuscript: Kulikovskiy M., Maltsev Y., Glushchenko A., Gusev E., Kapustin D., Kuznetsova I., Kociolek J.P. Preliminary molecular phylogeny of the diatom genus <em>Nupela</em> with the description of a new species and consideration of the interrelationships of taxa in the suborder Neidiineae D.G. Mann sensu E.J. Cox. Fottea</p> <p>Molecular investigation of diatom genera <em>Nupela</em> and <em>Brachysira</em> is conducted using strains from Indonesia and Vietnam. New species from the genus <em>Nupela indonesica</em> sp. nov. is described using combined approach. <em>Nupela lesothensis</em> (Schoeman) Lange-Bertalot is investigated using molecular data too. Phylogenetic analysis shows that <em>Nupela</em> and <em>Brachysira</em> are not closest genera. Morphology of <em>Nupela</em> and it differences from <em>Brachysira</em> is discussed. The genus <em>Nupela</em> is differs from all other diatom taxa by having coalescent hymenes ouside of areolae but not inside. Facultative development of raphe between different <em>Nupela</em> species is discussed.<br> SUPPLEMENT S1. Taxa and DNA sequence data used in phylogenetic analysis.<br> SUPPLEMENT S2. Final alignment of 2-gene DNA sequence data used for phylogenetic analysis in FASTA format.<br> SUPPLEMENT S3. Maximum Likelihood tree of <em>Nupela</em> species (indicated in bold) constructed from a concatenated alignment of 163 partial rbcL and partial 18S rDNA sequences of 1806 characters. Values near the horizontal lines (slash) are bootstrap support from RAxML analyses (&lt;50 are not shown). Species from the centric diatoms were used as an outgroup. Families indicated according COX (2015).</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Genbank accession numbers of sequences used in the phylogenetic analysis

<p><em>Phylogenetic analysis: </em>Sequences were assembled using Lasergene v15 (DNASTAR, Inc. Madison, USA), and combined with sequences obtained from Genebank</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Data from: Assembly ASM291031v2 (Genbank: GCA_002910315.2) identified as assembly of the Northern Dolly Varden (Salvelinus malma malma) genome, and not the Arctic char (S. alpinus) genome

<p>Here is the data that is a supplementary to the preprint: Shedko S.V. 2019. Assembly ASM291031v2 (Genbank: GCA_002910315.2) identified as assembly of the Northern Dolly Varden (Salvelinus malma malma) genome, and not the Arctic char (S. alpinus) genome // arXiv:1912.02474 <a href="https://arxiv.org/abs/1912.02474">https://arxiv.org/abs/1912.02474</a></p>

openother-openDec 2019View details →
zenodo40/100

GenBank accession numbers

Explanation note: GenBank accession numbers of the sequences used in the multi-gene phylogenetic analysis.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Genbank annotation and sequence of the genome of a Rickettsiales symbiont of Reticulomyxa filosa

<p>The genome sequence of a Rickettsiales symbiont was selectively from the genome sequencing reads of its host, Reticulomyxa filosa. Those sequences were obtained from a previous third-party study (doi: 10.1016/j.cub.2013.11.027). The selective assembly procedure was based on GC content and coverage of contigs obtained from the total read sets with SPAdes. Full details are provided in the manuscript file.<br>The obtained symbiont sequence was then annotated with Prokka, and the gbk output file selected.</p><p>This genome was obtained and analysed in the context of a larger genome comparative studies on the Rickettsiales, aimed to investigate the evolutionary patterns within the whole lineage.</p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Custom script for feature extraction from Genbank files

<p><span>Microbes are thought to be distributed and circulated around the world, but the connection between marine and terrestrial microbiomes is largely unknown. We use <em>Plantibacter</em>, a representative plant-associated genus, as our research model to show the global distribution and adaptation of plant-related bacteria in plant-free environments, especially in the remote Southern Ocean and the deep Atlantic Ocean. The marine isolates and their plant-associated relatives shared over 98% whole-genome average nucleotide identity (ANI), indicating recent divergence and ongoing speciation from plant-related niches to marine environments. Comparative genomics revealed that the marine strains acquired new genes via horizontal gene transfer from non-<em>Plantibacter</em> species and refined existing genes through positive selection to improve adaptation to new habitats. Meanwhile, marine strains retained the ability to interact with plants, such as modifying root system architecture and promoting germination. <em>Plantibacter </em>species were further found to be widely distributed in marine environments, revealing an unrecognized phenomenon that plant-associated microbiomes have colonized the ocean, which could serve as a reservoir for plant growth-promoting microbes. This study demonstrates the presence of an active reservoir of terrestrial plant growth-promoting bacteria in remote marine systems and advances our understanding of the microbial connections between plant-associated and plant-free environments at the genome level.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

FI GU R E 3 Maximum likelihood phylogenetic tree of the Hyalospheniformes with a focus on Apodera, Alocodera, and Padaungiella based on COI gene sequences. Bootstrap values (bs) and Bayesian posterior probabilities (p.p.) are indicated respectively between branches. COI sequences from genera other than Apodera were retrieved from GenBank in Superficially described and ignored for 92 years, rediscovered and emended: Apodera angatakere (Amoebozoa: Arcellinida: Hyalospheniformes) is a new flagship testate amoeba taxon from Aotearoa (New Zealand)

FI GU R E 3 Maximum likelihood phylogenetic tree of the Hyalospheniformes with a focus on Apodera, Alocodera, and Padaungiella based on COI gene sequences. Bootstrap values (bs) and Bayesian posterior probabilities (p.p.) are indicated respectively between branches. COI sequences from genera other than Apodera were retrieved from GenBank

opencc-by-4.0Aug 2021View details →
zenodo40/100

Linked collectors and determiners for: A new species of Polietina (Diptera: Muscidae) from South America, with an updated phylogeny of the genus and a review of species' identity in GenBank.

Natural history specimen data linked to collectors and determiners held within, "A new species of Polietina (Diptera: Muscidae) from South America, with an updated phylogeny of the genus and a review of species' identity in GenBank". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="http://bionomia.net/dataset/a30b0cfb-72ef-49cf-b982-d1a7f7be5027">https://bionomia.net/dataset/a30b0cfb-72ef-49cf-b982-d1a7f7be5027</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/a30b0cfb-72ef-49cf-b982-d1a7f7be5027">https://gbif.org/dataset/a30b0cfb-72ef-49cf-b982-d1a7f7be5027</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
zenodo40/100

Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group. in Distribution of extracellular enzyme-producing bacteria in the digestive tracts of 4 brackish water fish species

Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group.

opencc-by-4.0Dec 2013View details →
zenodo40/100

Figure. Interferon alpha-A based phylogenetic tree (neighbor joining method) constructed by MEGA 6.1 for Punjab urial in comparison with other mammalian species sequences available from GenBank (NCBI). in Characterization of interferon alpha of major histocompatibility complex class I in Punjab urial (Ovis vignei punjabiensis)

Figure. Interferon alpha-A based phylogenetic tree (neighbor joining method) constructed by MEGA 6.1 for Punjab urial in comparison with other mammalian species sequences available from GenBank (NCBI).

opencc-by-4.0Dec 2017View details →
zenodo40/100

Diamond NCBI Genbank Viral database for SOVAP

<p><strong>Diamond NCBI Genbank Viral database</strong></p> <p>Database type:&nbsp;Diamond database</p> <p>Database format version:&nbsp;3</p> <p>Label:&nbsp;2023-03-18_18-40-17</p> <p>Sequences:&nbsp;3,191,190</p> <p>Sum length: 824,564,244</p> <p>Assembly summary entries:&nbsp;58,201</p> <p>--------------------------------------------------------</p> <p><strong>SOVAP v.1.3:&nbsp;</strong><a href="https://github.com/poursalavati/SOVAP">GitHub</a></p> <p><strong><em>Soil Virome Analysis Pipeline</em></strong></p> <p>Description</p> <p>The study of viral communities in complex environmental samples, such as&nbsp;<strong>soil</strong>, can provide valuable insights into the diversity and functions of viral communities in the ecosystem. However, processing and analyzing of virome data can be a challenging task that requires the integration of various computational tools and techniques.</p> <p>To address these challenges, we have developed&nbsp;<strong>SOVAP</strong>&nbsp;pipeline that utilizes a suite of state-of-the-art tools for processing, analysis, and annotation viromics and metagenomics data.</p> <p>It utilizes various tools such as&nbsp;<strong>Fastp</strong>&nbsp;and&nbsp;<strong>Centrifuge</strong>&nbsp;for preprocessing and contamination removal,&nbsp;<strong>geNomad</strong>,&nbsp;<strong>Diamond</strong>&nbsp;and&nbsp;<strong>Megan</strong>&nbsp;for identification and annotation of viral contigs which are assembled and clustered using&nbsp;<strong>Megahit</strong>&nbsp;and&nbsp;<strong>CD-HIT</strong>. Additionally, this pipeline provides an&nbsp;<strong>estimate of the abundance</strong>&nbsp;of viral contigs, allowing for a more comprehensive understanding of the virome within the sample. The integration of these tools offers a reliable and effective means of taxonomy classification and annotation of viral contigs, aiding researchers in gaining insight into the composition and function of the virome within the analyzed sample.</p> <p>By integrating the SOVAP pipeline with&nbsp;<strong>IMG/VR</strong>&nbsp;and&nbsp;<strong>geNomad</strong>, it is possible to identify a wider range of viruses, including those that were previously unknown.</p> <p>The&nbsp;<strong>batch-mode</strong>&nbsp;script allows for the processing of multiple datasets using the SOVAP pipeline. This feature is particularly useful for&nbsp;<strong>large-scale</strong>&nbsp;analyses, such as those involving multiple environmental samples or large sequencing datasets.</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

Custom script for feature extraction from Genbank files

Open the record for dataset details and reuse information.

publicJan 2024View details →
zenodo36/100

GenBank + BOLD CO1 Eukaryotic representative sequence set

<p>This is a representative sequence set for cytochrome oxidase subunit 1 (CO1 or COI) combining all available eukaryotic CO1 sequences from GenBank and BOLD, clustered at 99% similarity.</p> <p>&nbsp;</p> <p>TODO:</p> <p>generate and add 7-level taxonomies for each sequence in this rep set.</p>

opencopyleft-next-0.3.1Feb 2020View details →
zenodo36/100

List of specimens, collection numbers, localities, and GenBank accessions of sequences. The neotype of Scinax x‑signatus is underlined. New sequences produced for this study are in bold. Abbreviations are as follow. Countries: ARG = Argentina, BOL = Bolivia, BRA = Brazil, GUF = French Guiana, GUY = Guyana, MTQ = Martinique, PER = Peru, SUR = Suriname; Brazilian states: AP = Amapá, BA = Bahia, CE = Ceará, ES = Espírito Santo, MA = Maranhão, MG = Minas Gerais, PE = Pernambuco, RJ = Rio de Janeiro, RS = Rio Grande do Sul, SP = São Paulo. An asterisk (*) indicates approximate coordinates taken from Google Earth. in A neotype for Hyla x-signata Spix, 1824 (Amphibia, Anura, Hylidae)

List of specimens, collection numbers, localities, and GenBank accessions of sequences. The neotype of Scinax x‑signatus is underlined. New sequences produced for this study are in bold. Abbreviations are as follow. Countries: ARG = Argentina, BOL = Bolivia, BRA = Brazil, GUF = French Guiana, GUY = Guyana, MTQ = Martinique, PER = Peru, SUR = Suriname; Brazilian states: AP = Amapá, BA = Bahia, CE = Ceará, ES = Espírito Santo, MA = Maranhão, MG = Minas Gerais, PE = Pernambuco, RJ = Rio de Janeiro, RS = Rio Grande do Sul, SP = São Paulo. An asterisk (*) indicates approximate coordinates taken from Google Earth.

opencc-by-nc-4.0Nov 2020View details →
zenodo36/100

Table of hsp65 OTUs (cutoff 99%), their inferred taxonomic allocations according to the hsp65 database and, for selected OTUs, closest species obtained from GenBank (BLAST) with percent identity.

<p>This table is part of the paper intitled &quot;Comparison of Actinobacteria communities from human-impacted and pristine karst caves&quot;</p>

opencc-by-4.0Feb 2022View details →
dryad36/100

Updating splits, lumps, and shuffles: Reconciling GenBank names with standardized avian taxonomies

Abstract Biodiversity research has advanced by testing expectations of ecological and evolutionary hypotheses through the linking of large-scale genetic, distributional, and trait datasets. The rise of molecular systematics over the past 30 years has resulted in a wealth of DNA sequences from around the globe. Yet, advances in molecular systematics also have created taxonomic instability, as new estimates of evolutionary relationships and interpretations of species limits have required widespread scientific name changes. Taxonomic instability, colloquially "splits, lumps, and shuffles," presents logistical challenges to large-scale biodiversity research because (1) the same species or sets of populations may be listed under different names in different data sources, or (2) the same name may apply to different sets of populations representing different taxonomic concepts. Consequently, distributional and trait data are often difficult to link directly to primary DNA sequence data without extensive and time-consuming curation. Here, we present RANT: Reconciliation of Avian NCBI Taxonomy. RANT applies taxonomic reconciliation to standardize avian taxon names in use in NCBI GenBank, a primary source of genetic data, to a widely used and regularly updated avian taxonomy: eBird/Clements. Of 14,341 avian species/subspecies names in GenBank, 11,031 directly matched an eBird/Clements; these link to more than 6 million nucleotide sequences. For the remaining unmatched avian names in GenBank, we used Avibase's system of taxonomic concepts, taxonomic descriptions in Cornell's Birds of the World, and DNA sequence metadata to identify corresponding eBird/Clements names. Reconciled names linked to more than 600,000 nucleotide sequences, ~9% of all avian sequences on GenBank. Nearly 10% of eBird/Clements names had nucleotide sequences listed under 2 or more GenBank names. Our taxonomic reconciliation is a first step towards rigorous and open-source curation of avian GenBank sequences and is available at GitHub, where it can be updated to correspond to future annual eBird/Clements taxonomic updates.

opencc-zeroAug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record