Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.7.1
Dataset results
47 results for “fasta”
FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions
<p>Fasta sequence correspond to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>
FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequence corresponds to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), <em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1 (Colanero et al., 2020)], <em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>
FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequences correspond to the MYB encoding genes <em>An2-like </em>and <em>Ant1</em>. The genomic sequences were combined correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI) (available at NCBI: <a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>
FASTA file containing the MYB encoding gene An2-like and Ant1 coding sequences corresponding to wild and cultivated tomato accessions
<p>The coding sequence (CDS) of the MYB encoding genes <em>Ant1</em> and <em>An2-like</em>. Sequences were retrieved from regions corresponding to the<em> Aft</em> locus from <em>Solanum galapagense </em>accession LA1141, <em>S. lycopersicum</em> variety OH8245, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences were compared to available CDS available from the Sol genomics network (SGN) and the National Center for Biotechnology Information. The CDS was retrieved from <em>S. lycopersicum</em> variety Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum</em> accession LA1996 [MN242011.1, EF433417.1( Sapir et al., 2008; Colanero et al., 2020)], and <em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], The orthologous CDS corresponding to the <em>Aft </em>MYB encoding genes from <em>Solanum tuberosum</em> L. Group Phureja clone DM1-3 genome (PGSC DM v4.03 Pseudomolecules) was retrieved from the Potato Genome Sequence Consortium (PGSC: Potato Genome Sequencing Consortium et al., 2011), and the Capsicum annum cv. CM334 genome was retrieved from <em>Capsicum annuum </em>cv CM334 genome chromosome release 1.55 (Hulse-Kemp et al. 2018). These CDS were obtained using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at https://solgenomics.net/tools/blast/). Comparison of syntenic chromosomal regions using known positions of tomato, potato, and pepper markers with comparative map viewer from SGN: (available at https://solgenomics.net/cview) on chromosome 10, was used as a quality check for S.<em> tuberosom</em> and <em>C. annuum.</em> Orthologous CDS corresponding to <em>Salvia miltiorrhiza, Arabidopsis thaliana</em>, [NM_105308.2, NM_105310.4 (Teng et al., 2005, Cominelli et al., 2008; Beradini et al., 2015)] were chosen based on tomato <em>Aft</em> sequence homology and gene annotations of positive R2R3 MYB regulation of anthocyanin. The CDS corresponding to the <em>Aft</em> genes were retrieved from the CDS reference genomes available from the Sol Genomics Network SGN: Tomato Genome CDS (ITAG release 4.0), Potato PGSC DM v3.4 CDS sequences, <em>Capsicum annuum </em>cv CM334 Genome CDS (release 1.55), or from the National Center for Biotechnology Information (NCBI: https://www.ncbi.nlm.nih.gov) reference sequences (RefSeq) section of the Genbank records. When accessed from Genank records, the CDS sequence was extracted from the “features” section and exported as a FASTA file.</p>
Vitamin_B12_related_genes_and_FASTA_sequences
<p>These two files are related to a submitted research paper on the study of a 7-year metagenomic time series carried monthly in a northwestern Mediterranean coastal site (SOLA, Banyuls-sur-Mer, FRANCE) :</p> <p><em>Seasonal succession of different vitamin B12 biosynthesis pathways and producers in coastal marine microbial communities. </em></p> <p>This study focuses on prokaryotes involved in vitamin B12 metabolism (biosynthesis, transport, remodeling) and on the metabolic pathways involved in these processes.</p> <p>The dataset in <strong>.xlsx</strong> format represents the abundance table of all the genes associated with 82 KEGGs involved in vitamin B12 metabolism, and the file in <strong>.fasta</strong> format corresponds to the FASTA sequences associated with these genes.</p>
HEE Burn-Control Site OTU FASTA
<p>Paper: Diversity and composition of fungal soil communities across prescribed burn areas in temperate hardwood forests</p> <p>Authors: S.D. Russell & M.C. Aime</p> <p>FASTA file containing the sequences of each OTU that was recovered from the burn and control survey areas using the methodology described in the paper above.</p>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
bacteria_masking:v21.1.1 complementary file (fasta)
<p>This archive contains the 42216 fasta files used to build the kraken database here : https://zenodo.org/records/11518607</p> <p>This is 24Gb large and could not be added in the original record.</p>
Micractinium rhizosphaerae NFX-FRZ genome annotation, fasta file, amino acid
<p>Micractinium rhizosphaerae NFX-FRZ annotation. FASTA file, amino acid.</p>
FASTA consensus sequences obtained using amplicon-based genome sequencing of SARS-CoV-2
<p>Set of 22 FASTA consensus sequences that were produced during routine SARS-CoV-2 sequencing obtained using amplicon-based sequencing (ARTIC protocol). Those sequences were compared to those generated in NASCarD applications.</p>
Rare variant replaced Korea reference genome fasta
<p>The rare variants of Korea reference genome fasta were replaced with using 396 Korean vcf information</p>
Alignment of mitogenome sequences (FASTA file) for a paleogenomic investigation of overharvest implications in an endemic wild reindeer subspecies
<p>Overharvest can severely reduce the abundance and distribution of a species and thereby impact its genetic diversity and threaten its future viability. Overharvest remains an ongoing issue for Arctic mammals, which due to climate change now also confront one of the fastest changing environments on Earth. The high-Arctic Svalbard reindeer (<em>Rangifer tarandus platyrhynchus</em>), endemic to Svalbard, experienced a harvest-induced demographic bottleneck that occurred during the 17–20th centuries. Here we investigate changes in genetic diversity, population structure, and gene-specific differentiation during and after this overharvesting event. Using whole-genome shotgun sequencing, we generated the first ancient and historical nuclear (n = 11) and mitochondrial (n = 18) genomes from Svalbard reindeer (up to 4000 BP) and integrated these data with a large collection of modern genome sequences (n = 90), to infer temporal changes. We show that hunting resulted in major genetic changes and restructuring in reindeer populations. Near-extirpation followed by pronounced genetic drift have altered the allele frequencies of important genes contributing to diverse biological functions. Median heterozygosity was reduced by 23%, while the mitochondrial genetic diversity was reduced only to a limited extent, likely due to already low pre-harvest diversity and a complex post-harvest recolonization process. Such genomic erosion and genetic isolation of populations due to past anthropogenic disturbance will likely play a major role in metapopulation dynamics (i.e., extirpation, recolonization) under further climate change. Our results from a high-arctic case study therefore emphasize the need to understand the long-term interplay of past, current, and future stressors in wildlife conservation.</p>
Curated Phage Database (CPD) fasta file
<p>This is the fasta file including phage genomes used to generate the Curated Phage Database (CPD) utilized in our in review manuscript "The circulating phageome reflects bacterial infections". The corresponding phage characteristic data will be present in the manuscript as a supplemental file, and can be used to connect a Genbank ID to identified bacteriophage host and phage taxonomic information if known.</p> <p>Please note that this database is built from phage sequences in the NCBI nucleotide repository. Due to field bias towards sequencing human disease-related bacteria and their phage, this database is reflective of this bias and is most representative of bacteriophage associated with human pathogens and as such underrepresents environmental phages in comparison - a limitation to keep in mind when utilizing to interpret potential phage sequences.</p>
GFA file and fasta files with read length 100 for Minnow
<p>de-Bruijn GFA file with read length 100 along with the reference transcript (produced by TwoPaCo) for both mouse and human. </p>
Sei framework resources (hg19 and hg38 FASTA files)
<p>This is the `resources` directory that should be downloaded in order to run the Sei framework code. It contains the hg19 and hg38 genome FASTA files downloaded from UCSC, as well as the index files generated by `pyfaidx` Python package for fast indexing and querying of genome coordinate sequences. </p>
Gene annotations and associated FASTA file of Antechinus flavipes (yellow-footed antechinus)
<p>Gene annotations of <em>Antechinus flavipes</em> (yellow-footed antechinus) associated with the manuscript "A chromosome-level genome of <em>Antechinus flavipes</em> provides a reference for an Australian marsupial genus with male death after mating" (<em>Molecular Ecology Resources</em>, accepted).</p> <p><strong>Annotation file and associated FASTA files for <em>A. flavipes</em> assembly AdamAnt. </strong><br> • AdamAnt.gff3.tar.gz: EVidenceModeler (EVM) gene model annotation file for <em>A. flavipes</em> assembly AdamAnt.<br> • AdamAnt.CDS.fa.tar.gz: EVM gene models coding sequences.<br> • AdamAnt.AA.fa.tar.gz: EVM gene models coding sequences translated into amino acid sequences.</p> <p><strong>Annotation file and associated FASTA files for <em>A. flavipes</em> transcriptome assembly generated using ROPUS (Reference-guided Ortholog Pipeline for Unannotated species; see doi:10.5281/zenodo.3722900). </strong><br> • antechinus_flavipes_ROPUS_CDS_annotated.fa.tar.gz: ROPUS CDS set. Annotated using human as the reference species.<br> • antechinus_flavipes_ROPUS_AA_annotated.fa.tar.gz: ROPUS AA set. Translated using the CDS sequence. <br> • antechinus_flavipes_ROPUS_annotated.gff3.tar.gz: annotation file for gene expression analysis using the <em>A. flavipes </em>ROPUS CDS FASTA.</p>
DNA sequences of transgenes detected via environmental DNA (raw ABI files, processed FASTA files, and reference alignments)
We demonstrate that simple, non-invasive environmental DNA (eDNA) methods can detect transgenes of genetically modified (GM) animals from terrestrial and aquatic sources in invertebrate and vertebrate systems. We detected transgenic fragments between 82-234 bp through targeted PCR amplification of environmental DNA extracted from food media of GM fruit flies (<i>Drosophila melanogaster</i>), feces, urine, and saliva of GM laboratory mice (<i>Mus musculus</i>), and aquarium water of GM tetra fish (<i>Gymnocorymbus ternetzi</i>). With rapidly growing accessibility of genome-editing technologies such as CRISPR, the prevalence and diversity of GM animals will increase dramatically. GM animals have already been released into the wild with more releases planned in the future. eDNA methods have the potential to address the critical need for sensitive, accurate, and cost-effective detection and monitoring of GM animals and their transgenes in nature.
GRCh38.p13 Reference FASTA (bgzip'd with faidx)
<p>This is derived from https://ftp.ncbi.nlm.nih.gov/genomes/all/GCF/000/001/405/GCF_000001405.39_GRCh38.p13/GRCh38_major_release_seqs_for_alignment_pipelines/GCA_000001405.15_GRCh38_full_plus_hs38d1_analysis_set.fna.gz.</p> <p>All non-primary sequences have been removed.</p> <p>It has then been recompressed with bgzip and indexed with samtools:</p> <pre><code class="language-bash">curl -#fSL https://ftp.ncbi.nlm.nih.gov/genomes/all/GCF/000/001/405/GCF_000001405.39_GRCh38.p13/GRCh38_major_release_seqs_for_alignment_pipelines/GCA_000001405.15_GRCh38_full_plus_hs38d1_analysis_set.fna.gz -o genomic.fna.gz gunzip genomic.fna.gz awk '{ if ((NR>1)&&($0~/^>/)) { printf("\n%s", $0); } else if (NR==1) { printf("%s", $0); } else { printf("\t%s", $0); } }' genomic.fna | grep -v "^>chr\S*_" - | tr "\t" "\n" > genomic.short.fna bgzip -c genomic.short.fna > reference.fna.bgz samtools faidx reference.fna.bgz tar -czvf GRCh38_reference_fasta.tar reference.fna.bgz reference.fna.bgz.fai reference.fna.bgz.gzi </code></pre> <p> </p> <p>This tar file contains:</p> <ul> <li>reference.fna.bgz</li> <li>reference.fna.bgz.fai</li> <li>reference.fna.bgz.gzi</li> </ul> <p> </p>
Fasta format protein sequences from assembled kyphosid fish gut metagenomes
<p>Predicted proteins sequences from kyposid fish gut metagenomic samples F5, F6, F7, and F8, obtained as described in the following study:</p> <p>Podell S, Oliver A, Kelly LW, Sparagon W, Plominsky, A, Nelson RS, Laurens LML, Augyte, S, Sims NA, Nelson CE, Allen EE. Herbivorous fish microbiome adaptations to sulfated dietary polysaccharides (2023)<br> manuscript submitted.</p>
Serratia fonticola EBS19 Whole genome sequence data fasta file annotated
<p>Whole genome sequence data of <em>Serratia fonticola</em> <strong>EBS19 </strong>strain.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.