Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “gff3”
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
BSGatlas GFF3 file
<p>The Bacillus Subtilis Genome atlas (BSGatlas) was obtained by integrating and curating existing annotation sources. The annotation is available as a genome browser hub and as a standardized GFF3 download.</p> <p>See: <a href="https://rth.dk/resources/bsgatlas/">https://rth.dk/resources/bsgatlas/ </a><br> The browser enables a user to directly compare their experimental data with the annotation which was not directly possible beforehand.</p> <p>This doi contains the annotation GFF3 file and the genome sequence. For increased technical portability (the various genome browser tools), the genome sequence is referred to as "basu168", yet it is identical to the NCBI reference assembly "AL009126.3".</p> <p>The scripts that generated this annotation are under https://doi.org10.5281/zenodo.4305872</p>
RxLR EFFECTROR SEQUENCES ALONG WITH GFF3 FILES
<p>Fasta sequences are the RxLR sequences of all 128 Phytophthora genomes. The GFF3 files contain the coordinates of the RxLRs. </p>
Our predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates
<p>These files contain predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates.</p>
(gff3) MMETSP re-assemblies
<p>Corresponding gff3 annotation files for the Trinity assembly files (https://doi.org/10.5281/zenodo.251828) generated from the dammit pipeline (https://github.com/camillescott/dammit).</p> <p>The Marine Microbial Eukaryotic Transcriptome Sequencing Project (MMETSP) data set contains cultured samples of pelagic and endosymbiotic marine eukaryotic species representing more than 40 phyla (Keeling et al. 2014).</p> <p>Methods for the de novo transcriptome assembly are described in the Eel pond khmer protocols (Brown et al. 2015).</p> <p>Scripts available on github: </p> <p>https://github.com/dib-lab/dib-MMETSP</p> <p>References:</p> <p>C. Titus Brown, Camille Scott, and Leigh Sheneman. 2015. The Eel Pond mRNAseq Protocol. https://khmer-protocols.readthedocs.io/en/ctb/mrnaseq/</p> <p>Keeling et al. 2014. The Marine Microbial Eukaryote Transcriptome Sequencing Project (MMETSP): Illuminating the Functional Diversity of Eukaryotic Life in the Oceans through Transcriptome Sequencing. http://dx.doi.org/10.1371/journal.pbio.100188</p>
Annotation file (gff3) from the killifish, Fundulus rathbuni (gill epithelium)
<p>Annotation, gff3 file of the reference de novo transcriptome assembly from the killifish, Fundulus rathbuni created with the dammit pipeline using the Pfam, Rfam, OrthoDB databases in addition to Fundulus heteroclitus amino acids (RefSeq).</p> <p>Fish were acclimated to either brackish or fresh water then exposed to an acute brackish water challenge. Transcriptome data from gill epithelium tissue were collected. A reference transcriptome assembly was generated from all individuals then used to analyze transcriptional responses to salinity.</p> <p> </p>
O. amberensis GFF3
<p>Gene annotation file to go with the O. amberensis genome assembly</p>
GFF3 Annotation files for 120 Phytophthora genomes
<p>The GFF3 files contain annotation data for 120 Phytophthora genomes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.