Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48,977
datasets available to search
ShareScore release 0.7.1
Dataset results
48,977 results for “gene”
Gene Expression and Tree Growth in the CTFS-ForestGEO Plot at Harvard Forest 2017-2019
Major goals in ecosystem ecology have been to scale from leaves to canopies and to determine whether individual-level, intra-species and inter-specific variation is critical for models projecting ecosystem processes now and in the future. The project is important in that it examines these issues in detail considering genotypes and levels of gene expression all the way up to canopy level CO2 flux. Ecological genomics and transcriptomics are nascent fields that have been primarily restricted to model species in natural and (mostly) controlled environments. To date, we have very few studies of non-model organisms in nature and/or studies of functional genomics through space and time. The research is producing extraordinarily rich datasets regarding the gene expression of trees across populations, through space in each population, across the growing season and across years and linking this information to growth and gas exchange. It will, therefore, provide tremendous insights into how much variation exists in nature thereby guiding sampling designs in future ecological 'omics projects. More importantly, it will provide unusually detailed phenotypic information for important non-model species that have large impacts on the CO2 flux of eastern US forests.
Global eutrophication and antibiotic resistance genes dataset for "Coupling mechanisms between cyanobacteria and antibiotic resistance genes in freshwater ecosystems"
This dataset compiles global records of cyanobacteria, antibiotic resistance genes (ARGs), and associated water quality parameters to support research on freshwater ecosystem dynamics. It includes 990 metagenomes, 16,648 chlorophyll-a (Chl-a) records, and over 90 documented cases of ARGs–cyanobacteria co-occurrence under comparable spatiotemporal conditions. The dataset covers the years 2000–2024 and provides both raw measurements and harmonized tables for cross-study comparisons. Data were extracted from previously published literature and public repositories, with references to source publications included. This archive is intended to facilitate reproducible analyses, enable large-scale meta-studies, and support further exploration of microbial interactions in freshwater systems.
Expression-based polygenic score from the amygdala 5HTT gene network
<p>This pipeline intends to facilitate the calculation of biologically informed polygenic scores from collected genomic data. This template can be adapted to create other expression-based polygenic risk scores. Data is 1) step by step description and 2) a list of genes that compose the gene network.</p>
Differential gene expression data of commercial compounds used to assess the performance of human TeraTox assay
<p>The dataset supplements the publication `Optimization of the <em>TeraTox</em> assay for preclinical teratogenicity assessment`. </p> <ul> <li>2022-02-18-TeraTox-commercial-logFC.gct: log2FC matrix of genes by compounds (in concentration ranges)</li> <li>2022-02-18-TeraTox-commercial-pScore.gct: p-scores (log 10 transformed p-values with the sign of logFC) of genes by compounds</li> <li>2022-02-18-TeraTox-commercial-featureData.txt: feature annotation in TSV format</li> <li>2022-02-18-TeraTox-commercial-phenoData.txt: sample annotation in TSV format</li> <li>2021-06-10-gcGeneFactorAnno-withPositiveCoefs.tsv: gene membership of germ-layer factors, with germ-layer annotation and average expression in copies per million (cpm).</li> </ul> <p>Citation: Jaklin, Manuela, Jitao David Zhang, Nicole Schäfer, Nicole Clemann, Paul Barrow, Erich Küng, Lisa Sach-Peltason, Claudia McGinnis, Marcel Leist, and Stefan Kustermann. “Optimization of the TeraTox Assay for Preclinical Teratogenicity Assessment.” <em>Toxicological Sciences</em> 188, no. 1 (July 1, 2022): 17–33. <a href="https://doi.org/10.1093/toxsci/kfac046">https://doi.org/10.1093/toxsci/kfac046</a>.</p>
Lab513/Landscape of Gene Expression Dataset
<p>Datasets from the article:</p> <p><strong>A microfluidic device for inferring metabolic landscapes in yeast monolayer colonies.</strong></p> <p>by Zoran S Marinkovic<sup>1,2,3</sup>, Clément Vulin<sup>1,4,5</sup>, Mislav Acman<sup>1,3</sup>, Xiaohu Song<sup>2</sup>, Jean Marc Di Meglio<sup>1</sup>, Ariel B. Lindner<sup>*,2,3</sup>, Pascal Hersen<sup>*,1</sup></p> <p>eLife 2019;8:e47951 DOI: <a href="https://doi.org/10.7554/eLife.47951">10.7554/eLife.47951</a></p> <p>first versions on BioRxiv : <a href="https://www.biorxiv.org/content/10.1101/527846v2">https://www.biorxiv.org/content/10.1101/527846v2</a></p> <p> </p>
Light-regulated gene expression and alternative splicing data from rice seedlings.
<p>This data contains analyzed data from the experiment conducted on rice seedlings under dark and light conditions. Seeds of rice (Oryza sativa spp. japonica cv. Nipponbare) were sown in the dark and germinated on day 2 and continued to grow in the dark for another 6 days. 3 biological replicates of the dark-grown etiolated shoots were harvested on day 8 after sowing. The remaining dark-grown seedlings were exposed to continuous white light at 120 mol/m2/sec for 48 hours or another 2 days (Days 9 and 10 after sowing). Three replicates of the light-treated green-colored seedling samples were harvested at the end of day 10. Harvested samples were frozen in liquid nitrogen and stored at -80C until further processing.</p>
Undinarchaeota illuminate DPANN phylogeny and the impact of gene transfer on archaeal evolution
<p><strong>General Description </strong></p> <p>Repository with all analyses described our paper: <a href="https://www.nature.com/articles/s41467-020-17408-w">Undinarchaeota illuminate DPANN phylogeny and the impact of gene transfer on archaeal evolution</a>.</p> <p>If you find this work useful for your own analyses, please cite this work.</p> <p> </p> <p><strong>Abstract</strong></p> <p>The evolution and diversification of Archaea is central to the history of life on Earth. Cultivation-independent approaches have revealed the existence of the DPANN archaea: a radiation of organisms with small cell and genome sizes. Currently, the placement of the various DPANN lineages and in turn the early evolution of metabolism and symbiosis are debated. Here, we reconstructed genomes of a thus far uncharacterized archaeal phylum-level lineage UAP2 (<em>Candidatus</em> Undinarchaeota). Comparative genomics revealed that members of the Undinarchaeota have small estimated genome sizes and, while potentially being able to conserve energy through fermentation, likely depend on partner organisms for the acquisition of vitamins, amino acids and other metabolites. In contrast to previous indications, our phylogenomic analyses robustly placed the Undinarchaeota as independent lineage between two major and highly supported clans of ‘DPANN’. Furthermore, our work suggests that DPANN archaea have exchanged core genes with their hosts by horizontal gene transfer, adding to the difficulty of placing DPANN in the tree of life (ToL). In several cases, this pattern is sufficiently dominant that known symbiont-host clades can be identified by inferring routes of HGT across the ToL. Together, our findings provide crucial insights into the origins and evolution of DPANN archaea and their hosts.</p> <p><strong>The annotation workflow for archaeal/bacterial genomes that was used for this paper is also available on github (<a href="https://github.com/ndombrowski/Genome_annotations">here</a>) and an updated version that includes the COG search is available on: </strong><a href="https://github.com/ndombrowski/Annotation_worfklow">https://github.com/ndombrowski/Annotation_workflow</a></p> <p> </p> <p><strong>Repository Contents</strong></p> <p><strong>1_Genome_files.tar.gz</strong> includes all Undinarchaeota (original name UAP2) metagenome-assembled genomes (MAGs). This includes: </p> <ol> <li>The original contigs for each UAP2 MAG (fna files)</li> <li>The prokka output for each UAP2 MAG (faa files)</li> <li>A concatenated file of all proteins from each UAP2 MAG and all archaeal reference genomes (364 genomes in total). This folder also includes a list of archaeal genomes investigated.</li> </ol> <p><strong>2_Phylogenies.tar.gz</strong> includes all files for the phylogenetic analyses. This includes the following folders:</p> <p>1. Files for the concatenated species trees for different taxa sets. These files are related to the following parts of the manuscript: Supplementary Table 6; Figure 1 and Supplementary Figures S8-S58. The folder includes the following:</p> <ul> <li>Folder '1_unaligned_sequences' includes individual protein sequences extract from the different taxa sets.</li> <li>Folder '2_alignments' includes the alignment files generated by MAFFT.</li> <li>Folder '3_alignments_trimmed' includes the alignments trimmed with BMGE.</li> <li>Folder '4_phylogenies' includes the IQ-TREE output for all phylogenies as well as color-annotation file for figtree. Additionally files rooted with minimal ancestor deviation (MAD) rooting (*.rooted) are provided. Note, that for the final figures the *treefile_renamed (i.e. the iqtree file with the full taxa string) were artificially rooted using the DPANN archaea. The numbering corresponds to Supplementary Table S6 of the main manuscript.</li> <li>Folder ' 5_pdfs' includes the PDFs for each tree</li> </ul> <p>2. Files for single gene trees that includes:</p> <ul> <li>The folder '1_arcogs' includes the unaligned proteins, alignments, trimmed alignments, trees and pdfs for the single gene trees based on the arCOG identifiers. The arCOGs were extract from 12 UAP2 MAGs + 352 archaeal + 3020 bacterial + 100 eukaryotic genomes. ArCOGs were only considered if they occurred in at least 3 UAP2 genomes. Notice, these files were used to investigate UAP2 for HGT events and correspond to the following parts of the manuscript: Figure 4 and Supplementary Tables 4, 5, 20-22. Additionally, the folder 0_parsing includes some information on how to generate count tables for each marker gene.</li> <li>The folder '151_markers' including the proteins, alignments, trimmed alignments, trees and pdfs for evaluating the 151 marker set used for the concatenated species tree. Files were provided for the 127 and 364 taxa set. These files were used as a basis for the concatenated species trees that were used to generated Supplementary Figures S8-S58. Additionally, the trees were used for ranking marker proteins and generating Supplementary Tables 4-5. For the 364 taxa set, the folder also included a subfolder 0_parsing that provides scripts to investigate some statistics for each marker protein, including the average protein length, average alignment length and average bootstrap support.</li> <li>The folder '3_other_individual_trees' includes the proteins, alignments and phylogenies for the 16S_23S, RubisCO and primase analyses. The data was used to generate the following parts of the manuscript: Supplementary Table 11, Supplementary Figures 3-5, 57 and 59.</li> </ul> <p><strong>3_Scripts.tar.gz</strong> includes all files for the phylogenetic analyses. This includes the following folders:</p> <p>1. The files for the main workflow for the annotations and phylogenies.</p> <ul> <li>This folder includes the workflow to generate annotations for archaeal genomes as well as an example script that was used to generate phylogenies. These analyses were typically run on a in-house bioinformatics cluster with 4x Xeon Gold 6140 2.3 GHz processors using bash, python and perl. The used system runs a Linux operating system, Red Hat Enterprise 7.5.</li> </ul> <p>2. A folder providing any required dependencies that include:</p> <ul> <li>any python or perl scripts that were used during this study and/or that are mentioned in the methods section</li> <li>Databases used for the annotations, esp. if these were slightly modified. Notice, changes typically include parsing of the mapping files or modifications of the sequence headers for easier parsing.</li> <li>mapping files needed to link the genome accession ids to the taxonomy string as well as lists of protein IDs used for different phylogenies (i.e. 14 + 48 arCOGs used for protein phylogenies)</li> </ul> <p>3. R scripts (including all needed input files) used to: </p> <ul> <li>generate tables and figures for the annotations, i.e. Figure 2 and 3 and Supplementary Tables 7, 8, 9, 12, 13-15 and Supplementary Figures 60, 62-64 . The input folder includes the raw output from the annotation workflow and includes annotations for the 12 UAP2 MAGs as well as 352 archaeal reference genomes.</li> <li>generate tables and figures for the HGT analyses, i.e. Figure 4 and Supplementary Tables S20-22 Here, proteins based on arCOGs were extracted from 364 archaeal, 3020 bacterial and 98 eukaryotic genomes and used to generate single protein phylogenies. The resulting trees were used to investigate horizontal gene transfer events and the necessary scripts are provided in this folder.</li> <li>generate tables and figures for the amino acid identify (AAI) comparisons, i.e. Supplementary Table S3 and Supplementary Figure S2. </li> <li>rank the marker genes for concatenated species trees for the 127 and 364 taxa set. These were used to generate Supplementary Tables S4 and S5.</li> </ul> <p><strong>General comment:</strong></p> <p>In contrast to the previous version, this datasets includes some small additional scripts generated during the revision process of the corresponding manuscript.</p> <p> </p>
Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention
<p>Single cell RNA seq datasets used for analysis in the Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention</p>
Pseudo-nitzschia multistriata gene models
<p>The resource contains fasta files with the <em>Pseudo-nitzschia multistriata</em> gene models and proteins, Pm-1.4_mRNA_v3.fa and Pm-1.4_peptide_v3.fa, a file with the annotation, psmu_mRNA_uniref_2015_06_filt_ann_out.txt, and a file with mapping information, genes_v3_WA.gff3.</p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div> </div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div> </div> <div>To extract the dataset, execute the following command:</div> <div> </div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div> </div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div> </div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div> </div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
CAbiNet: Joint clustering and visualization of cells and genes for single-cell transcriptomics
<p>We here provide the data sets to reproduce the results in our manuscript "CAbiNet: Joint clustering and visualization of cells and genes for single-cell transcriptomics". Our package "CAbiNet" can be downloaded from https://github.com/VingronLab/CAbiNet. The scripts to reproduce the results in our manuscript can be found from https://github.com/VingronLab/CAbiNet_paper.</p><p>You can find the description of folders in 'Data.zip' in the README.md file.</p>
FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions
<p>Fasta sequence correspond to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>
FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequence corresponds to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), <em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1 (Colanero et al., 2020)], <em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>
FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequences correspond to the MYB encoding genes <em>An2-like </em>and <em>Ant1</em>. The genomic sequences were combined correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI) (available at NCBI: <a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>
Supporting data for: Type 1 diabetes risk genes mediate pancreatic beta cell survival in response to proinflammatory cytokines
<p><strong>SUMMARY OF THE STUDY</strong></p> <p>We combined functional genomics and human genetics to investigate processes that affect type 1 diabetes (T1D) risk by mediating beta-cell survival in response to proinflammatory cytokines. We mapped 38,931 cytokine-responsive candidate <em>cis-</em>regulatory elements (cCREs) in beta-cells using ATAC-seq and snATAC-seq and linked them to target genes using co-accessibility and HiChIP. Using a genome-wide CRISPR screen in EndoC-βH1 cells we identified 867 genes affecting cytokine-induced survival, and genes promoting survival and up-regulated in cytokines were enriched at T1D risk loci. Using SNP-SELEX, we identified 2,229 variants in cytokine-responsive cCREs altering transcription factor (TF) binding, and variants altering binding of TFs regulating stress, inflammation and apoptosis were enriched for T1D risk. At the 16p13 locus, a fine-mapped T1D variant altering TF binding in a cytokine-induced cCRE interacted with <em>SOCS1</em>, which promoted survival in cytokine exposure. Our findings reveal processes and genes acting in beta-cells during inflammation that modulate T1D risk.</p> <p><strong>DESCRIPTION OF FILES:</strong></p> <ul> <li>Supplementary Data 1. List of islet cCREs annotated with cell type and cytokine response - also in GSE205853</li> <li>Supplementary Data 2. Coaccessible sites in untreated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 3. Coaccessible sites in cytokine-treated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 4. Coaccessible sites in cytokine treated and untreated beta cells and promoter annotations - also in GSE205853</li> <li>Supplementary Data 5. Chromatin interactions in EndoC-BH1 cells - also in GSE205853</li> <li>Supplementary Data 6. Variants selected for SNP-SELEX assay </li> <li>Supplementary Data 7. Variants with TF binding and allelic binding results from SNP-SELEX</li> <li>Supplementary Data 8. snATAC-seq barcodes and metadata - also in GSE205853</li> <li>Supplementary Data 9. CRISPR-KO screen results - also in GSE205853</li> <li>Supplementary Data 10. Bulk ATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 11. Bulk RNA-seq count matrix - also in GSE205853</li> <li>Supplementary Data 12. Alpha cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 13. Acinar cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 14. Beta cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 15. Stellate cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 16. Endothelial cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 17. Delta cells snATAC-seq count matrix - also in GSE205853</li> <li>Supplementary Data 18. Luciferase assay rs10483809</li> <li>Supplementary Data 19. SOCS1 knockdown qPCR results</li> <li>Supplementary Data 20. SOCS1 knockdown Apotracker (flow-cytometry)results</li> </ul> <p><strong>Raw data deposited at GEO, accessions GSE205853 and GSE118725.</strong></p> <p><em>Please refer to publication and GEO for details on methods.</em></p>
Gene expression ATLAS of Arabidopsis thaliana (accession Columbia) across its lifecycle
<p><strong>Abstract: </strong>Arabidopsis thaliana (accession- Columbia) is an important model plant. RNA-Seq based study of 36 gene expression libraries was carried out to explore transcriptional programs operating in different plant parts (seedling, rosette, root, inflorescence, flower, fruit silique, and seed) and developmental stages (2-leaf stage, 6-leaf stage, 12-leaf stage, senescence stage, dry mature and imbibed seed stage). For each tissue type and developmental stage, three individual plants were used as biological replicates.</p> <div><strong><span>Organism part: </span></strong><span>inflorescence, whole plant, seed, root, silique fruit, flower, rosette</span></div> <div> </div> <div><span><strong>Developmental stage:</strong> </span><span>LP.02 two leaves visible stage, IL.00 inflorescence just visible stage, fruit size 30 to 50% stage, LP.12 twelve leaves visible stage, root development stage, fruit size 70% to final stage, LP.06 six leaves visible stage, dry seed stage, flowering stage, seed imbibition stage, sporophyte senescent stage, inflorescence development stage</span></div> <div> </div> <div> <div><strong><span>Organism: </span></strong><span>Arabidopsis thaliana</span></div> <div> </div> <div><span><strong>Ecotype:</strong> </span><span>Col-0</span></div> <div> </div> <div><strong><span>Genotype: </span></strong><span>wild type genotype</span></div> <div> </div> <div><span><strong>Age:</strong> Samples are from </span><span>20-day, 49-day, 39-day, 15-day, 21-day, 9-day, 22-day, 55-day, 26-day, 45-day</span></div> <div> </div> <div><span><strong><span>Experimental Designs: </span></strong><span>growth chamber study<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://purl.obolibrary.org/obo/EO_0007269" target="_blank" rel="noopener"> EFO</a></span>, <span>development or differentiation design<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0001746" target="_blank" rel="noopener"> EFO</a></span>, <span>organism part comparison design<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0001750" target="_blank" rel="noopener"> EFO</a></span></span></div> <div> </div> <div><span>For more description of the data and sample types see the file <a href="../api/records/11133989/draft/files/PRJEB24664_Sample_descriptors.xlsx/content" target="_blank" rel="noopener noreferrer">PRJEB24664_Sample_descriptors.xlsx or visit </a> or visit <a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-6422/sdrf">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-6422/sdrf</a></span></div> <div> </div> <div><span>Original data was submitted from </span></div> <div> <ul> <li><span>EMBL-EBI ArraExpress: <a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-6422">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-6422</a></span></li> <li><span>NCBI SRA: <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB24664">https://www.ncbi.nlm.nih.gov/bioproject/PRJEB24664</a></span></li> </ul> <p><strong><span>Protocol description:</span></strong></p> <table> <tbody><tr> <th>Name</th> <th>Type</th> <th>Description</th> <th>Hardware</th> </tr> </tbody><tbody> <tr> <td>P-MTAB-71349</td> <td><span>growth protocol<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0003789" target="_blank" rel="noopener"> EFO</a></span></td> <td>Seeds were planted in pots containing commercial potting mix with fertilizers. Pots were covered with clear perforated plastic wrap and kept at 4 degrees celsius for 3 days to break the dormancy. After 3 days plants were transferred to the Intellus Ultra growth chamber (Percival Scientific, IA, USA) which was set to temperature 22-23 degrees celsius, light intensity 120-150 micromol/m2sec under the cycle of 16h light and 8h dark. Soil was kept moist by gently spraying with water every 72 hours to maintain humidity to 50-60%. Sampling time point is given in days after germination.</td> <td> </td> </tr> <tr> <td>P-MTAB-71350</td> <td><span>nucleic acid extraction protocol<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0002944" target="_blank" rel="noopener"> EFO</a></span></td> <td>Total RNA from frozen samples was extracted as a method described in Filichkin et al., 2010. Total RNA was used to isolate large RNA as per manufacturer's protocol for miRNeasy Mini kits (Qiagen Inc., USA), and RNase-free DNase (Life Technologies Inc., USA).</td> <td> </td> </tr> <tr> <td>P-MTAB-71351</td> <td><span>nucleic acid library construction protocol<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0004184" target="_blank" rel="noopener"> EFO</a></span></td> <td>True-Seq kit (Illumina Inc.) was used to prepare RNA-seq libraries, according to the manufacturer’s protocol.</td> <td> </td> </tr> <tr> <td>P-MTAB-71352</td> <td><span>nucleic acid sequencing protocol<a title="" href="https://www.ebi.ac.uk/ols4/ontologies/efo/terms?iri=http://www.ebi.ac.uk/efo/EFO_0004170" target="_blank" rel="noopener"> EFO</a></span></td> <td>101bp paired-end sequencing of mRNA was performed by using the standard protocols on Illumina HiSeq 3000.</td> <td>Illumina HiSeq 3000</td> </tr> </tbody> </table> </div> </div>
AMnrGC - Amazon river non-reduntant microbial genes catalogue
<p> </p> <p><strong>AMnrGC : Amazon river basin non-redundant microbial gene catalogue</strong></p> <p> </p> <p> RELEASE 2018/01<br> --------------------------------------</p> <p> </p> <p>1. INTRODUCTION</p> <p> AMnrGC is a collection of genes and proteins which were constructed<br> by use of Amazon river basin openly available metagenomes from<br> sequencing projects (SRP044326, PRJEB25171 and SRP039390). Briefly,<br> metagenomes were coassembled by groups made up their geographical<br> location with Megahit v.1.0 and the contigs were used to gene predictions<br> by Prodigal v.2.6.3. Genes sequences were length filtered (> 150 bp) and<br> clustered by CD-HIT-EST (version 4.6) at 95% of nucleotide identity and<br> 90% of overlap of the shorter gene. Theorical protein products were annotated<br> by the most completes databases up to date and their complete information<br> is available here.</p> <p> </p> <p>2. LOCATION</p> <p> AMnrGC versions will be available on the web only under the current ZENODO<br> repository: 10.5281/zenodo.1484504</p> <p> </p> <p>3. FORMAT</p> <p> Gene entries were named as ">AM_AGSSY_XXX" where XXX represents an unique numerical<br> identifier. Genes were deposited in their coding phase, because of this, all of them<br> can be used to generate the protein sequences by transeq function at ORF+1.<br> The protein entries correspond to genes artifical translation used in the annotations,<br> and also available, codified in the same way, but containing the indication "_1" in<br> the end of the header. Example:</p> <p> Gene:<br> >AM_AGSSY_151515</p> <p> Protein:<br> >AM_AGSSY_151515_1</p> <p> Annotations were provided as separate tables for each database used to annotate the<br> sequences. The header of these tables indicates the meaning of each value.</p> <p> </p> <p>2. FUTURE FORMAT CHANGES</p> <p> No major changes are expected for the main general format of the database.<br> New versions should include updated versions of annotations or even additional sequences,<br> numbered as subsequent entries.</p> <p> </p> <p>3. ACKNOWLEDGEMENTS<br> <br> This work is a joint effort of Laboratory of molecular biology from Federal<br> University of São Carlos, São Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), as well as, the spanish funding organ Consejo Superior de Investigaciones Científicas (CSIC).</p> <p> </p> <p> This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001.</p> <p> </p> <p>4. THE AMnrGC TEAM</p> <p> AMnrGC is maintained by a group of researchers. You can contact<br> the AMnrGC consortium.<br> <br> Current curators:</p> <p> - Célio Dias Santos Júnior (celio.diasjunior@gmail.com)<br> - Flavio Henrique-Silva (dfhs@ufscar.br)<br> - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> </p> <p>5. COPYRIGHT NOTICE</p> <p> AMnrGC - Amazon river basin non-redundant microbial gene catalogue<br> Copyright (C) 2018 The AMnrGC consortium.</p> <p> This database is provided “as is” and without any warranty of any kind,<br> of openly available for non-commerical purposes. You can redistribute and/or modify it<br> as you wish, under the terms of the ODBL 1.0 license:</p> <p> https://opendatacommons.org/licenses/odbl/1.0/</p> <p> For commercial purposes, please contact us. </p> <p>___________________</p> <p>The AMnrGC Consortium<br> 2018</p>
Bedrock radioactivity influences the rate and spectrum of mutation - Orthologous genes
<p>Alignments of the 2490 orthologous genes used in the article "Natural Bedrock radioactivity influences the rate and spectrum of mutation" to estimate the mutational spectrum and synonymous substitution rate.</p> <p>To compute accurate synonymous substitution rate, we removed genes with short sequences (<half of the alignment) and genes strongly supporting another phylogeny using ProfileNJ <a href="https://paperpile.com/c/Klqlpb/W5sS">(Noutahi et al. 2016)</a> with a bootstrap threshold of 90%, resulting in a subset of 769 genes listed in the file "List_769_1-to-1_orthologs_EvolutionRate.txt".</p> <p>Transcriptome paired-end reads used to define these orthologous genes have been deposited to the European Nucleotide Archive and are available under the study ID PRJEB14193.</p> <p>Sequences were aligned with Prank<a href="https://paperpile.com/c/Klqlpb/pilh"> (Löytynoja & Goldman 2008)</a> using a codon model and sites ambiguously aligned were removed with Gblocks <a href="https://paperpile.com/c/Klqlpb/c5kb">(Castresana 2000)</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.