Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

59

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

59 results for “Sequence databases”

Learn how ShareScore rates datasets ↗
zenodo52/100

Database of fitted spectra for: Changing-Look AGNs - I. Tracking the transition on the main sequence of quasars

<h3>Results from the spectral fitting for a sample of changing-look active galactic nuclei (AGNs) with SDSS spectroscopy using PyQSOFit.</h3>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Harmonised LUCAS database classified by crop sequence type

<p>Assessing the benefits of crop diversification &ndash; a pillar of the agroecological transition &ndash; on a large scale requires a description of current crop sequences as a baseline, which is lacking at the scale of the European Union (EU). This work is based on the Harmonised LUCAS in-situ land cover and use database for field surveys from 2006 to 2018 in the European Union (doi: <a href="http://doi.org/10.2905/f85907ae-d123-471f-a44a-8cca993485a2">10.2905/f85907ae-d123-471f-a44a-8cca993485a2)</a> to fill this gap, We completed this dataset with a crop sequence type information for each point under non-perennial agricultural land cover in 2012, 2015 and 2018.</p> <p>The dataset lucas_classified.csv includes 31 159 points. Variables &quot;point_id&quot;, &quot;nuts0&quot;, &quot;nuts2&quot;, &quot;th_lat&quot;, &quot;th_long&quot;, &quot;LC1_2012&quot;, &quot;LC1_2015&quot;, &quot;LC1_2018&quot; are inherited from the Harmonised LUCAS databse. Variables &quot;cereals&quot;, &quot;corn&quot;, &quot;rapeseed&quot;, &quot;sunflower&quot;, &quot;pulses&quot;, &quot;rootCrops&quot;, &quot;forageLeg&quot;, &quot;grassland&quot; correspond to the temporal frequencies of respectively cereals, corn, rapeseed, sunflower, pulses, root crops, forage legumes and grassland within the 2012, 2015 and 2018 crop sequence for each point. Variable &quot;crop_sequence_type&quot; is the crop sequence type assigned to each point, among eight options: cereals, corn and cereals, forage legumes and cereals, pulses and cereals, rapeseed and cereals, root crops and cereals, sunflower and cereals, temporary grasslands.</p> <p>This dataset could be used to map current dominant crop sequences in the European Union, as illustrated in the map attached, and to assess the benefits of future crop diversification.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.

<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Universit&eacute; Claude Bernard Lyon 1, VetAgro Sup, Research Team &ldquo;Bacterial Opportunistic Pathogens and Environment&rdquo; (BPOE), 69280 Marcy L&rsquo;Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bont&eacute; et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer&rsquo;s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O&#39;LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project &ldquo;Iouqmer&rdquo; EST 2016/1/120, l&#39;Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rh&ocirc;ne, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: C&eacute;line COLINON, Romain MARTI, Emilie BOURGEOIS, S&eacute;bastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257&ndash;4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161&ndash;168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bont&eacute;, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153&ndash;164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

IBP-database and environmental IBP sequences for functional analysis of microalgae

<p>Database for analysis of ice binding protein (IBP) sequences (Uhlig et al. (2015)):</p> <p>(1) DUF3494_seqs_Uniprot.fasta:  full length sequences with DUF3494 domain used for the calculation of the backbone tree in the phylogenetic placement</p> <p>(2) env_IBPs.fasta: potential IBP sequences from one Arctic and five Antarctic sea ice metatranscriptomes (Sanger or 454)</p> <p>(3) DUF3494_substree_fig2a_UniprotIDs.txt: UniProtIDs for subtree in Fig 2a</p> <p>(4) DUF3494_confirmed_IBPactivity_UniprotIDs.txt: UniProtIDs for sequences with confirmed IBP function of the protein</p> <p>If using this dataset please cite the following publication: Uhlig, C., Kilpert, F., Frickenhaus, S., Kegel, J.U., Krell, A., Mock, T., Valentin, K., Beszteri, B., (2015) The significance of antifreeze proteins for eukaryotic microbial communities of Arctic and Antarctic sea ice, The ISME Journal, 9, 2537–2540, doi:10.1038/ismej.2015.43</p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region

<p>This dataset has been prepared as part of the Interreg Alpine Space project Eco-AlpsWater (ASP569) -&nbsp;<em>Innovative Ecological Assessment and Water Management Strategy for the Protection of Ecosystem Services in Alpine Lakes and Rivers</em>,&nbsp;<a href="https://www.alpine-space.eu/projects/eco-alpswater/en/home">https://www.alpine-space.eu/projects/eco-alpswater/en/home</a></p> <p>Individual archives include 16S rRNA (cyanobacteria) and 18S rRNA (microalgae) FASTA sequences and associated blastn results obtained from the high throughput sequencing of plankton and biofilm bulk/eDNA samples collected in 2019 in 37 lakes and 22 rivers across the Alpine region. These are supporting files for the paper by Salmaso et al., 2022.&nbsp;DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region. Science of the Total Environment, in press.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

PSSH2 - database of protein sequence-to-structure homologies (including Sars-CoV-2 structures)

<p><strong>Protein sequence and structure data</strong></p> <p>This data set contains data from Uniprot (in the files called protein_sequence, protein_synonyms, protein_names, organism_synonyms) and PDB (in the files called PDB and PDB_chain) as used by the <a href="https://github.com/ODonoghueLab/Aquaria">Aquaria web resource</a> at the time of download (2022-02-08).</p> <p>&nbsp;</p> <p><strong>The&nbsp;PSSH2 data set</strong><br> <br> PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences.</p> <p>&nbsp;</p> <p><strong>Calculating PSSH2</strong></p> <p>The&nbsp;Swissprot and PDB data was downloaded in November 2021.<br> Generating PSSH2: We used <a href="https://gwdu111.gwdg.de/~compbiol/uniclust/2021_03/UniRef30_2021_03.tar.gz">UniRef30_2021_03</a> (originally called UniRef30_2021_06)&nbsp;from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30%. The HHblits code and the code for running the calculations&nbsp;was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively)&nbsp;at the respective&nbsp;time of calculation in the timeframe until December&nbsp;2021.&nbsp;<br> &nbsp;</p> <p><strong>PDB based sequence-to-structure alignments</strong></p> <p>In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These are contained in the PDBchain data.</p> <p>This data covers sequences and PDB structures in the timeframe until February 2022.&nbsp;</p> <p>&nbsp;</p> <p><strong>Evaluating PSSH2</strong></p> <p>The resulting alignment data was analysed using CATH domain assignments downloaded from&nbsp;/cath/releases/all-releases/v4_2_0/cath-classification-data/ to define correct hits and false hits:&nbsp;</p> <ul> <li>The set of query sequences is defined by the CATH non-redundant S40_overlap_60 dataset (ftp://orengoftp.biochem.ucl.ac.uk/cath/releases/all-releases/v4_2_0/non-redundant-data-sets/)</li> <li>The set of all expected hits are all pdb structures containing a domain with the same CATH code if contained in the set of processed sequences (-&gt; all) or&nbsp;only if also contained in the set of non redundant sequences (-&gt; nr40).</li> <li>The set of true positives is defined by sharing the same CATH code up to the level of homology (&quot;CATH&quot;) or up to the level of topology (&quot;CAT&quot;).</li> </ul> <p>The data was evaluated with respect to false discovery rate (FDR) and recall (true positive rate TPR) by cumulatively considering all hits with an E-value below the threshold (&quot;C&quot;) or in bins with an E-value between the threshold and one tenth of the threshold (&quot;B&quot;). This evaluation was carried out for the data obtained in November 2021 (202111)&nbsp;as well as previous data from October 2020 (202010), February 2020 (202002) and&nbsp;September 2017 (201709). The results are&nbsp;&nbsp;collected in&nbsp;<a href="https://zenodo.org/api/files/445add84-fcf1-4dfe-b8a1-63dc55f378ee/PSSH%20CATH%20validation.csv?versionId=a5df7473-6efd-442b-b422-3944e9452003">PSSH CATH validation.csv</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Known errors</strong></p> <p>Due to processing error, the profile of pdb structure 5fia A / B (sequence md5 052667679fc644184f40063c7602c9e1) is incomplete in the pdb_full hhblits database which led to further errors in generating sequence based alignments for sequences for 1vtm P (sequence md5 c844aff103449363cb8489c78c58ebf1) and 434t A / B (sequence md5 d67aa1c3a36492c719cb48b5e7ecc624).<br> <br> &nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

gapseq reference sequence databases for Bacteria and Archaea

<p>The repository contains the protein sequences used by <a href="https://github.com/jotech/gapseq">gapseq</a> to predict the presence of metabolic reactions and to construct metabolic models.</p> <p>The workflow using gapseq to generate this set of reference protein sequences:</p> <p>&nbsp;</p> <p>```sh</p> <p># delete all "old" data<br>rm dat/seq/Bacteria/rev/*.fasta<br>rm dat/seq/Bacteria/unrev/*.fasta<br>rm dat/seq/Bacteria/rxn/*.fasta<br>rm dat/seq/Archaea/rev/*.fasta<br>rm dat/seq/Archaea/unrev/*.fasta<br>rm dat/seq/Archaea/rxn/*.fasta</p> <p># run gapseq find to re-download everything#<br># the genome is irrelevant as no blasting is performed ('-x')<br>gapseq find -p all -t Bacteria -n -x -U toy/ecoli.faa.gz &gt; bac_update.log 2&gt;&amp;1<br>gapseq find -p all -t Archaea -n -x -U toy/ecoli.faa.gz &gt; ar_update.log 2&gt;&amp;1</p> <p># create all sequence .tar.gz archives (rev/unrev/rxn)<br>cd dat/seq/Bacteria/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../</p> <p># create md5sum table for all tar.gz archives<br>cd dat/seq/<br>find -mindepth 2 -type f -name "*.tar.gz" -exec md5sum {} \; &gt; md5sums.txt</p> <p># create taxon-specific final archive for Zenodo upload<br>tar -czvf Bacteria.tar.gz Bacteria/*/*.tar.gz<br>tar -czvf Archaea.tar.gz Archaea/*/*.tar.gz</p> <p># Upload Bacteria.tar.gz, Archaea.tar.gz, and md5sums.txt &nbsp;to Zenodo via the web-interface</p> <p>```</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

WorldCOM Deliverable 1: Prevalence of ESBL subtypes in bacterial pathogens and a sequence database of selected alleles

<p><strong>OHEJP Project: WorldCOM, Deliverable 1, Work Package 1.</strong></p> <p>This dataset is connected to Work Package 1, Task1 of the WorldCOM consortium grant within the One Health EJP group. The aim was to analyse publicly available sequences for antimicrobial resistance genes associated with <em>Salmonella</em>, <em>Campylobacter</em> and <em>E. coli</em>. For the initial phase of this work package, we have focused on ESBL-related AMR genes. As these genes are absent from <em>Campylobacter</em>, we have not included this bacterium in these analyses, and have used the important pathogens <em>Klebsiella</em> and <em>Acinetobacter</em>. All types and subtypes of Extended Spectrum &beta;-Lactamases (ESBLs) and plasmid-mediated colistin resistance genes have been analysed for frequency among reported and extracted sequences. High frequency resistant genes subtypes have been highlighted for further sequence analysis to illustrate geographic distribution and geographic-specific single nucleotide polymorphisms (SNPs). The data shown are work in progress.&nbsp;&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Inbred Strain Variant Database (ISVdb): A repository for probabilistically informed sequence differences among the Collaborative Cross strains and their founders

<p>Data files for the development of a database for storing (and a GUI for retrieving) the imputed variants for 72 Collaborative Cross strains of mice. Files include the inputs for the imputation, as well as the final results. See File_S1_Readme for more details on the included files.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Structures extracted from the AlphaFold database for the set of sequences used in Rosenberg et al. Nat. Comm. 2022.

<p>Structures extracted from the AlphaFold database for the set of sequences used in Rosenberg, A.A., Marx, A. &amp; Bronstein, A.M. Codon-specific Ramachandran plots show amino acid backbone conformation depends on identity of the translated codon. Nat Commum 13, 2815 (2022) <a href="https://doi.org/10.1038/s41467-022-30390-9">https://doi.org/10.1038/s41467-022-30390-9</a>.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

K-mer databases of plant virus sequences for use with the Kodoja workflow

<p><strong>Details</strong></p> <p>This is a gzipped tar file that includes the plant virus database files required to run the Kodoja workflow (https://github.com/abaizan/kodoja)[1]. Kodoja is a workflow for the detection of plant virus sequences in RNA-seq data files that uses two previoulsy published tools Kraken[2] and Kaiju[3].</p> <p>This file contains databases for Kraken [2] and Kaiju [3]. The file includes the kraken database files: database.idx, database.kdb, nodes.dmp, names.dmp and the kaiju database file kaij_library.fmi.</p> <p>These k-mer databases are based on virus sequences in RefSeq [4] (ttps://www.ncbi.nlm.nih.gov/refseq/) with plant hosts as defined in the Virus-Host Database [5] (https://www.genome.jp/virushostdb/).</p> <p><strong>Version</strong><strong> 1.0</strong></p> <p>kodojaDB_v1.0 is based on RefSeq v89 and the Virus-Host Database (accessed 03/09/2018 which is based on RefSeq 89 and Genbank 226.0). The viral partition of RefSeq v89 genome comprises 7946 viruses (ftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/viral/assembly_summary.txt).</p> <p>kodojaDB_v1.0 was created using kodoja_retrieve.py which is part of the kodoja workflow (v0.05) (https://github.com/abaizan/kodoja).</p> <p><strong>References</strong></p> <p>[1] Baizan-Edge, A, Cock, P, MacFarlane, S, McGavin, W, Torrance, T, Jones, S. Kodoja: A workflow for virus detection in plants using k-mer analysis of RNA-sequencing data (under review Nucleic Acids Research).&nbsp;</p> <p>[2] Wood,D.E. and Salzberg,S.L. (2014) Kraken: ultrafast metagenomic sequence classification using exact alignments. <em>Genome Biol.</em>, <strong>15</strong>, R46</p> <p>[3] Menzel,P., Ng,K.L. and Krogh,A. (2016) Fast and sensitive taxonomic classification for metagenomics with Kaiju. <em>Nat. Commun.</em>, <strong>7</strong>, 1&ndash;9.</p> <p>[4] O&rsquo;Leary,N.A., Wright,M.W., Brister,J.R., Ciufo,S., Haddad,D., McVeigh,R., Rajput,B., Robbertse,B., Smith-White,B., Ako-Adjei,D., <em>et al.</em> (2016) Reference sequence (RefSeq) database at NCBI: Current status, taxonomic expansion, and functional annotation. <em>Nucleic Acids Res.</em>, <strong>44</strong>, D733&ndash;D745.</p> <p>[5] Mihara,T., Nishimura,Y., Shimizu,Y., Nishiyama,H., Yoshikawa,G., Uehara,H., Hingamp,P., Goto,S. and Ogata,H. (2016) Linking virus genomes with host taxonomy. <em>Viruses</em>, <strong>8</strong>, 10&ndash;15</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

The hsp65 metabarcoding DNA sequence database for taxonomic allocations using the Mothur (Version 1.0.0)

<ul> <li>The <em>hsp65</em> gene codes for an Heat Shock Protein (Telenti et al., 1993) and is widespread in the Actinobacteria phylum. It is well suited for the species allocation of the Nocardia genus (Rodriguez-Nava et al., 2006).</li> <li>The <em>hsp65</em> database, named ACTIhsp65, was designed to apply the <em>hsp65</em>-metabarcoding analytical scheme published in Vautrin et al. (2021). It includes the full <em>hsp65</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 401 nucleotide-long <em>hsp65</em> sequences of 1066 unique taxa belonging to 198 genera.</li> <li>Nucleotide sequences of <em>hsp65</em> (range: 165-565 nucleotides) were either retrieved from public repositories (GenBank) or made available by Veronica Rodriguez-Nava.Vautrin et al. (2021) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>hsp65</em> sequences.</li> <li>ACTIhsp65 V1.0.0 (June 2018 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>hsp65 </em>sequences down to the species.</li> </ul>

opencc-by-4.0Oct 2021View details →
dryad40/100

Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases

Open the record for dataset details and reuse information.

publicMar 2024View details →
zenodo36/100

CusVarDB: A tool for building customized sample-specific variant protein database from Next-generation sequencing datasets

<p>CusVarDB is a windows based tool for creating a variant protein database from Next-generation sequencing datasets. The program supports variant calling for Genome, RNA-Seq and exome datasets.</p> <p>This repository will provide the resultant variant peptides identified in our study and its corresponding information. The detailed information of the table is given below.</p> <p>Supplementary Table 1. This table contains the resultant variant peptides along with its wild-type peptides from BT474, MDMAB157, MFM223, and HCC38 datasets. Along with mutant peptides, this section also provides additional information such as peptide-spectrum match (PSM), Protein accession, cross-correlation value from the search (Xcorr), and retention time (RT).</p> <p>Supplementary Table 2. This table provides the complete details of the resultant peptides. Here the mutant and corresponding wild-type peptides are mentioned in different sheets. For a given mutant peptide its wild-type peptide and corresponding information can be mapped using the VLOOKUP function in Excel by keeping column A (Sl.No) as lookup parameter.</p> <p>Supplementary Table 3. This table briefs about the variants which are already reported in other cancers.</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Uncovering hundreds of RNA viral RdRps amongst uncharacterised sequences in public protein databases.

<p>These data are associated with the following manuscript:</p> <p>Brown, K., Firth, A. E. (2025)<br>Uncovering hundreds of RNA viral RdRps amongst uncharacterised sequences in public protein databases.<br><br></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

gapseq sequence database

<p>The repository contains the protein sequences used by gapseq [https://github.com/jotech/gapseq] to predict the presence of metabolic reactions and to construct metabolic models.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

ProbioSML: Probiotic Sequences database based on Machine Learning

<p>ProbioSML is a database comprising 1,071 genes associated with microbial genera that have been demonstrated to possess probiotic properties.</p>

opencc-by-4.0Jun 2025View details →
zenodo36/100

Database of liquefaction phenomena triggered by the March 2021 Thessaly, Greece, seismic sequence

<p>This archive contains data related to the paper &ldquo;Floodplain evolution and its influence on liquefaction clustering: the case study of March 2021 Thessaly, Greece, seismic sequence&rdquo; by George Papathanassiou, Sotiris Valkaniotis, Athanassios Ganas, Alexandros Stampolidis, Dimitra Rapti, Riccardo Caputo.</p> <p><br> Submitted to <em>Engineering Geology, Elsevier</em></p> <p>&nbsp;</p> <p><strong>Liquefaction_DB_20210303</strong> contains shapefiles and spreadsheet for the mapped liquefaction phenomena<br> <strong>UAS_Liquefaction_Surveys_Larisa_EQ</strong> contains detailed geospatial data for mapped liquefaction phenomena from UAS surveys</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Database of 16S sequences from SILVA (r114), filtered, curated and annotated to be used easily by programs of taxonomic assignments

<p>The database used for the taxonomic assignment of reads generally comes from the SILVA database (http://www.arb-silva.de/). The logic behind this&nbsp;database is to use the&nbsp;information from the best one to the worst one. This is why the curated database was splitted in two parts : the [C] sequences for Complete sequences in&nbsp;terms of taxonomy, and the [I] and [E] sequences, for Incomplete and Environmental sequences.</p> <p>Each sequence included into the database must have a specific format summarizing&nbsp;all needed information (example below):<br> &gt;[I]AACY020336309;Archaea(superkingdom);Euryarchaeota(phylum);Thermoplasmata(class);Thermoplasmatales(order);Marine_Group_II(no_rank);;marine_metagenome</p> <p>This sequence is an incomplete one ([I]), with a specific accession number from NCBI or SILVA, or another database (AACY020336309). Then, all taxonomic data is&nbsp;separated using &#39;;&#39; characters, for each considered level (superkingdom, phylum,&nbsp;<br> class, order, family, and genus). The species name is the last one and separated by two &#39;;&#39; characters from the rest of the descriptive line. Finally, the descriptive line must not contain specific characters like spaces. If one or several levels are unknown, this is indicated by &#39;no_rank&#39;.</p> <p>Another example here for [C] sequences:<br> &gt;[C]AAAK03000010;Bacteria(superkingdom);Firmicutes(phylum);Bacilli(class);Lactobacillales(order);Enterococcaceae(family);Enterococcus(genus);;Enterococcus_faecium_DO<br> This sequence is a complete one ([C]), with a specific accession number from NCBI or SILVA, or another database (AACY020187844). Then, all taxonomic data is&nbsp;separated using &#39;;&#39; characters, for each considered level (superkingdom, phylum,&nbsp;<br> class, order, family, and genus). The species is the last one and separated by two &#39;;&#39; characters from the rest of the descriptive line. Complete sequences&nbsp;must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>Another example here for [E] sequences:<br> &gt;[E]U59968;Archaea(superkingdom);Thaumarchaeota(phylum);Soil_Crenarchaeotic_Group(SCG)(no_rank);;uncultured_crenarchaeote<br> This sequence is a environmental one ([E]), with a specific accession number from NCBI or SILVA, or another database (U59968). Then, all taxonomic data is&nbsp;separated using &#39;;&#39; characters, for each considered level (superkingdom, phylum,&nbsp;class, order, family, and genus). The species is the last one and separated by two &#39;;&#39; characters from the rest of the descriptive line. Complete sequences&nbsp;<br> must have six levels of information (superkingdom, phylum, class, order, family, and genus). If it is not the case, the sequence will be considered as Incomplete ([I]) (between three and five levels), or Environmental ([E]) (with only the superkingdom and the phylum levels).</p> <p>More details on the steps defined to clean and define this new database can be available on demand (sebastien.terrat@inra.fr).</p>

opencc-by-4.0Nov 2017View details →
zenodo36/100

Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing - RASE Database for EuSCAPE

<p>RASE databases used for the prediction of antibiotic phenotype in the paper titled "Rapid Inference of Antibiotic Susceptibility Phenotype of Uropathogens using Metagenomic Sequencing with Neighbour Typing". A database for&nbsp;<em>Klebsiella pneuomiae </em>constructed from EuSCAPE isolates<em>.</em></p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record