Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

103

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

103 results for “reference database”

Learn how ShareScore rates datasets ↗
zenodo48/100

Reference Windfarm database CNk4 90

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a conventionally neutral boundary layer flow with 90 degree inflow wind direction angle (Casename CNk4 90)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo48/100

18S V9 metabarcoding reference databases and naive-bayes classifier

<p>18S metabarcoding databases and naive-bayes classifiers specific to the V9 region. Built&nbsp;from&nbsp;the <a href="https://pr2-database.org/">PR2 database</a> using Qiime2 (version 2023.2)<a href="https://github.com/BenKaehler/q2-clawback">.</a> Includes&nbsp;a naive-bayes classifier for use with Qiime2. Sequences were dereplicated with Rescript --p-mode 'uniq' ,&nbsp;retaining identical sequence records that have differing taxonomies.</p><p>Primers used:</p><p>EMP 18S 1391f:&nbsp;GTACACACCGCCCGTC</p><p>EMP 18S EukBr:&nbsp;TGATCCTTCTGCAGGTTCACCTAC</p><p><strong>Stats</strong></p><p>19,470 unique sequences</p><p>39,170 total sequences</p><p>11,748 unique taxa&nbsp;</p><p>Note: there were 221,085 sequences in the original PR2 database. Many were filtered out due to the in-silico extraction with our V9 primers.</p><h3>File Descriptions</h3><p><strong>Files in bold are recommended for taxonomic classification.</strong></p><p>Create naive-bayes classifier for 18S PR2 database.md: &nbsp;Markdown with code used to generate databases |</p><p><strong>pr2_v5.0.0_SSU_18S-V9_uniq-classifier.qza</strong>: Unweighted naive-bayes classifier for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, dereplicated, generated by qiime2-2023.2 |</p><p><strong>pr2_version_5.0.0_SSU_18S-V9_uniq_seqs.qza</strong>: Sequences for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, dereplicated, generated by qiime2-2023.2 |</p><p><strong>pr2_version_5.0.0_SSU_18S-V9_uniq_tax.qza</strong>: Taxa for pr2_version_5.0.0_SSU_18S-V9_uniq_seqs.qza (dereplicated) |</p><p>pr2_version_5.0.0_SSU_18S-V9_seqs.qza: Sequences for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, NOT dereplicated, generated by qiime2-2023.2 |</p><p>pr2_version_5.0.0_SSU_18S-V9_tax.qza: Taxa for pr2_version_5.0.0_SSU_18S-V9_seqs.qza (NOT dereplicated)&nbsp;</p><p>pr2_version_5.0.0_SSU_mothur.fasta: SSU sequences downloaded from PR2 v 5.0.1 &nbsp;|</p><p>pr2_version_5.0.0_SSU_mothur.tax: SSU taxa downloaded from PR2 v5.0.1 |</p><p>pr2_version_5.0.0_taxonomy.xlsx: Detailed taxonomy downloaded from PR2 v5.0.1 |</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Database on Certified Reference Materials measured with PAT tools for validation and verification purposes

<p>The H2020 PAT4Nano project aims to develop and demonstrate Process Analytical Technologies (PAT) tools for nanosuspension characterization which have sufficiently high resolution, accuracy, and speed, for real-time industrial process monitoring and control. Real time monitoring is desired for example to obtain: small, high precision, specialty batch of materials, processing monitoring of nucleation/growth/milling of materials at different scales (lab, pilot, production), and for producing feedback loops (adapt T, pH, etc.,) needed for process control.<br> Laser diffraction (LD), Spatially Resolved Dynamic Light Scattering (SR-DLS), Cross-Correlation Dynamic Light Scattering (CC-DLS), Ultrasound Nanoparticle Sizer (UNPS), Raman, and Transmission Electron Microscopy (TEM) are the main PAT tools used in this project. For validation and verification purposes of these measurement techniques, polystyrene and silica samples (200 and 1000 nm particle size) were selected as (Certified) Reference Materials ((C))RMs) by the consortium partners. The results described in this database are particle size measurements using PAT methods in an offline mode. The particle size and particle size distribution data are presented as the D10, D50 and D90 and PDI/span measured with each PAT tool.<br> Raman spectra of the CRMs are presented as well. Here, particle size data was extracted by using chemometric software. Lastly, TEM images of the CRMs are included in the database to cross-correlate and cross-validate the results of the spectroscopic and scattering PAT tools.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Reference Windfarm database CNk2 30

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a conventionally neutral boundary layer flow with 30 degree inflow wind direction angle (Casename CNk2 30)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Reference Windfarm database CNk4 30

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a conventionally neutral boundary layer flow with 30 degree inflow wind direction angle (Casename CNk4&nbsp;30)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Reference Windfarm database PDk 30

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a pressure-driven high Reynolds number&nbsp;boundary layer flow with 30 degree inflow wind direction angle (Casename PDk&nbsp;30)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Reference Windfarm database PDk 0

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a pressure-driven high Reynolds number&nbsp;boundary layer flow with 0 degree inflow wind direction angle (Casename PDk&nbsp;0)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Reference Windfarm database CNk8 0

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a conventionally neutral boundary layer flow with 0 degree inflow wind direction angle (Casename CNk8&nbsp;0)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Reference Windfarm database CNk2 60

<p>Dataset for TotalControl reference windfarm&nbsp;database simulation of a conventionally neutral boundary layer flow with 60 degree inflow wind direction angle (Casename CNk2 60)</p> <p>Included Python files for loading and visualizing the data.&nbsp;Use the plot_*.py files.</p> <p>Further information, including description of the case and&nbsp;dataset can be found in the deliverable report at:&nbsp;</p> <p><a href="https://cordis.europa.eu/project/id/727680/results">https://cordis.europa.eu/project/id/727680/results</a></p> <p>&quot;Database for reference wind farms part 2: windfarm&nbsp;simulations&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

INDIGO Graffiti Reference Database

<p>There is an extensive body of popular and scholarly literature&nbsp;on ancient and recent&nbsp;graffiti. The&nbsp;<strong>INDIGO Graffiti Reference Database</strong> was created to collect much of this graffiti-related literature in one place. It also contains literature about street art because contemporary graffiti and street art are hard to disentangle,&nbsp;and their relationship and hierarchy depend&nbsp;upon the source one consults. The database is a product of the academic graffiti project INDIGO.</p> <p>This collection of references is made available in two forms:</p> <ul> <li>a closed-source <a href="https://www.citavi.com/en">Citavi</a> database&nbsp;(the *.ctv6archive file);</li> <li>an open-source <a href="https://www.zotero.org">Zotero</a> database <ul> <li>either <a href="https://www.zotero.org/groups/5192206/indigo_graffiti_reference_database/library">consultable online</a>&nbsp;(reading only),</li> <li>or downloadable here as a Zotero RDF or&nbsp;BibTex file (*.rdf and *.bib, respectively)&nbsp;for import&nbsp;into Zotero or other reference managers.</li> </ul> </li> </ul>

opencc-by-sa-4.0Sep 2023View details →
zenodo44/100

DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region

<p>This dataset has been prepared as part of the Interreg Alpine Space project Eco-AlpsWater (ASP569) -&nbsp;<em>Innovative Ecological Assessment and Water Management Strategy for the Protection of Ecosystem Services in Alpine Lakes and Rivers</em>,&nbsp;<a href="https://www.alpine-space.eu/projects/eco-alpswater/en/home">https://www.alpine-space.eu/projects/eco-alpswater/en/home</a></p> <p>Individual archives include 16S rRNA (cyanobacteria) and 18S rRNA (microalgae) FASTA sequences and associated blastn results obtained from the high throughput sequencing of plankton and biofilm bulk/eDNA samples collected in 2019 in 37 lakes and 22 rivers across the Alpine region. These are supporting files for the paper by Salmaso et al., 2022.&nbsp;DNA sequence and taxonomic gap analyses to quantify the coverage of aquatic cyanobacteria and eukaryotic microalgae in reference databases: Results of a survey in the Alpine region. Science of the Total Environment, in press.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Standard Reference Database : ITV-CORE

<p>Decisions demand data, and poor quality data can lead to wrong, inaccurate, or late decisions. The way in which data is collected, stored and shared will reverberate in its quality and accuracy, consequently, reflecting on the ability to understand the aspects they represent. The private sector acting in the environmental area demands objectivity and assertiveness, and that is why it is essential to treat the data that subsidize conservation and restoration actions with exceptional care. To assess the state of biodiversity and environmental impacts, extensive field surveys are often required; for this, independent service providers are hired, who are specialized in obtaining a variety of types of information. Consequently, different collection methods are applied, and almost always methodological and formatting inconsistencies can be found in the resulting data. For the subsequent integration of this data into databases, it will be necessary to extract, adjust and standardize them, generating an entirely new demand, consuming time, human effort and financial resources. In addition, this demand also increases the risk of misinterpretation, typing and digitization errors, which can compromise quality, or even lead to loss of information. The standardization of data used in the survey, inventories, storage and sharing processes is a strategic solution to increase efficiency, reduce costs and risks of information degradation and loss. Furthermore, it brings a number of other benefits, such as the transformation of the analogic field recording system (field notebooks) to an entirely digital format, with the integration of cameras, tablets, dataloggers, and other widely available technologies. When it comes to preparing a recommendation for the standardization of data in a comprehensive and inclusive way, we mapped the biodiversity data frequently used by researchers from the Biodiversity and Ecosystem Services group at The Instituto Tecnol&oacute;gico Vale. Through this mapping, we seek to understand the types of data that already exist, how they have been used, stored and shared in databases, but also their convergence and peculiarities. With the participation of researchers, we seek to develop and validate a preliminary system of terms and metadata, including recommendations for best practices, aiming to improve the use of environmental and biodiversity data. The mapping showed a series of correspondences regarding the types of data used by the BES-ITV group, especially in the data applied in studies of Conservation and Restoration, Landscape Ecology, Genomics and Radio Frequency Identification. But also a great diversity of research topics (Total=29), focusing on six large biological groups, aspects that demonstrate the high multidisciplinary and wide coverage of environmental, ecological, genetic and biodiversity data used by the group. Based on these results, a system of terms and metadata is being developed, as well as the idealization of a modular system for the automatic generation of field digital spreadsheets, in order to simplify data collection through exclusively digital means.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

EukRibo: a manually curated eukaryotic 18S rDNA reference database

<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4&nbsp;&nbsp; &nbsp;</strong>We allowed up to 50 missing positions in the relatively conserved area at the 5&#39; end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3&#39; end of the V4 fragment is included.<br> <strong>V9&nbsp;&nbsp; &nbsp;</strong>We allowed up to 30 missing positions in the relatively conserved area at the 3&#39; end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5&#39; end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5&#39; end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3&#39; end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence (&#39;Y&#39;) or absence (&#39;N&#39;) of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> &nbsp;&nbsp; &nbsp;- same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete (&#39;yes - complete&#39;) or missing positions at the 5&#39; end (&#39;yes - partial&#39;)<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete (&#39;yes - complete&#39;), missing positions at the 3&#39; end (&#39;yes - partial&#39;), or was excluded, and the 6 possible reasons why (&#39;no - missing&#39;, &#39;no - too incomplete&#39;, &#39;no - chimera&#39;, &#39;no - bad quality&#39;, &#39;no - deletion in V9&#39;, &#39;no - Ns in V9&#39;)<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

CVD2014 - A database for evaluating no-reference video quality assessment algorithms

<p>The CVD video database is developed to provide an useful tool for researchers in the validation and developing processes of no-reference (NR) objective video quality assessment (VQA) algorithms. It consists of 234 videos from five different scenes captured by 78 different cameras (mobile phones, compact camera, video camera, SLR). The subjective experiments are conducted following the Single-Stimulus (SS) procedure to collect ratings of video quality.</p> <p><strong>Setup</strong></p> <p>We implement our experiments according to the Single Stimulus methodology using VQone MATLAB toolboxon high quality monitors (Eizo ColorEdge CG241W) with 1920x1200 pixel resolution in a dark room (ambient light &lt; 20 lux). Video stimuli were displayed at their original size of VGA (640 x 480) or HD (1280 x 720). The subjects viewing distance (80 cm) was controlled by a string hanging from a ceiling and they were instructed to keep their head steady next to it. The monitors were calibrated to according to sRGB (target values were: 6500 K, 80 lux, and gamma 2.2) using EyeOne Pro calibrator (X-rite co.). The laboratory setup is showed in the figure below.</p> <p><strong>Subjects</strong></p> <p>Subjects (n = 30, 30, 28, 33, 30, 32 and 27 for Tests 1 - 7 respectively) were na&iuml;ve in a sense that they did not study or work with image quality or related fields. They were recruited through student mailing lists consisting mainly humanities and behavioral science students. Subjects&rsquo; vision was controlled for the near visual acuity, near contrast vision (near F.A.C.T.) and color vision (Farnsworth D15) before the participation. They received movie tickets as a reward.</p> <p><strong>Procedure</strong></p> <p>Subjects evaluated one video sample at a time and all video samples of one scene were presented in a row. The order of video samples and scenes was randomized. Subjects had the option to view video samples again as many times as they wanted.</p> <p><strong>Data</strong></p> <p>The results are processed and reported in the form of Mean Opinion Score (MOS) for the tested video samples. In addition, we provide the whole raw data from the subjective experiments instead of just pre-calculated mean opinion scores from each video sample. This allows further analyses to be made by those who wish to use this database and gives them better opportunity to utilize the data to its full potential.</p> <p>Realignment study (test 7) contains the data from the additional study in which the mappings from the test and scene specific quality scales (test 1-6) to the global quality scale were formed. The global scale is valuable when studying and developing VQA algorithms. With the global scale, all of the samples (234 video samples in the case of the CVD2014) are in the same scale, and the performance analysis for algorithms can be conducted with a high number of samples.</p> <p><strong>If you use this database in your research, we kindly ask that you follow The Copyright notice below and cite the following paper:</strong></p> <p>&nbsp;</p> <p>M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen and J. H&auml;kkinen, &quot;CVD2014&mdash;A Database for Evaluating No-Reference Video Quality Assessment Algorithms,&quot; in <em>IEEE Transactions on Image Processing</em>, vol. 25, no. 7, pp. 3073-3086, July 2016. doi: 10.1109/TIP.2016.2562513</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>-----------COPYRIGHT NOTICE STARTS WITH THIS LINE------------</p> <p>Copyright (c) 2014 The University of Helsinki<br> All rights reserved.</p> <p>Permission is hereby granted, without written agreement and without license or royalty fees, to use, copy, modify, and distribute this database (the videos, the images, the results and the source files) and its documentation for any purpose, provided that the copyright notice in its entirely appear in all copies of this database, and the original source of this database,Visual Cognition research group (www.helsinki.fi/psychology/groups/visualcognition/index.htm) and the Institute of Behavioral Science (www.helsinki.fi/ibs/index.html) at the University of Helsinki (www.helsinki.fi/university/), is acknowledged in any publication that reports research using this database. Individual videos and images may not be used outside the scope of this database (e.g. in marketing purposes) without prior permission.</p> <p>The database and our paper are to be cited in the bibliography as: M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen and J. H&auml;kkinen, &quot;CVD2014&mdash;A Database for Evaluating No-Reference Video Quality Assessment Algorithms,&quot; in <em>IEEE Transactions on Image Processing</em>, vol. 25, no. 7, pp. 3073-3086, July 2016.<br> doi: 10.1109/TIP.2016.2562513</p> <p>-----------------------------------------------------------------------------</p> <p>LIMITATION OF LIABILITY</p> <p>UNIVERSITY OF HELSINKI SHALL IN NO CASE BE LIABLE IN CONTRACT, TORT OR OTHERWISE FOR ANY LOSS OF REVENUE, PROFIT, BUSINESS OR GOODWILL OR ANY DIRECT, INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL OR PUNITIVE COST, DAMAGES OR EXPENSE OF ANY KIND HOWEVER CAUSED OR HOWEVER ARISING UNDER OR IN CONNECTION WITH THE USE OF THIS DATABASE.</p> <p>THE UNIVERSITY OF HELSINKI SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE DATABASE PROVIDED HEREUNDER IS ON AN &quot;AS IS&quot; BASIS, AND THE UNIVERSITY OF HELSINKI HAS NO OBLIGATION TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS.</p> <p>THIS AGREEMENT SHALL BE CONSTRUED AND INTERPRETED IN ACCORDANCE WITH THE LAWS OF FINLAND, EXCLUDING ITS RULES FOR CHOICE OF LAW.</p> <p>-----------COPYRIGHT NOTICE ENDS WITH THIS LINE------------</p>

opencc-by-4.0Jun 2016View details →
zenodo44/100

gapseq reference sequence databases for Bacteria and Archaea

<p>The repository contains the protein sequences used by <a href="https://github.com/jotech/gapseq">gapseq</a> to predict the presence of metabolic reactions and to construct metabolic models.</p> <p>The workflow using gapseq to generate this set of reference protein sequences:</p> <p>&nbsp;</p> <p>```sh</p> <p># delete all "old" data<br>rm dat/seq/Bacteria/rev/*.fasta<br>rm dat/seq/Bacteria/unrev/*.fasta<br>rm dat/seq/Bacteria/rxn/*.fasta<br>rm dat/seq/Archaea/rev/*.fasta<br>rm dat/seq/Archaea/unrev/*.fasta<br>rm dat/seq/Archaea/rxn/*.fasta</p> <p># run gapseq find to re-download everything#<br># the genome is irrelevant as no blasting is performed ('-x')<br>gapseq find -p all -t Bacteria -n -x -U toy/ecoli.faa.gz &gt; bac_update.log 2&gt;&amp;1<br>gapseq find -p all -t Archaea -n -x -U toy/ecoli.faa.gz &gt; ar_update.log 2&gt;&amp;1</p> <p># create all sequence .tar.gz archives (rev/unrev/rxn)<br>cd dat/seq/Bacteria/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Bacteria/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/unrev/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../<br>cd dat/seq/Archaea/rxn/ &amp;&amp; tar -czvf sequences.tar.gz ./*.fasta &amp;&amp; cd ../../../../</p> <p># create md5sum table for all tar.gz archives<br>cd dat/seq/<br>find -mindepth 2 -type f -name "*.tar.gz" -exec md5sum {} \; &gt; md5sums.txt</p> <p># create taxon-specific final archive for Zenodo upload<br>tar -czvf Bacteria.tar.gz Bacteria/*/*.tar.gz<br>tar -czvf Archaea.tar.gz Archaea/*/*.tar.gz</p> <p># Upload Bacteria.tar.gz, Archaea.tar.gz, and md5sums.txt &nbsp;to Zenodo via the web-interface</p> <p>```</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

16S V4-V5 metabarcoding reference databases and weighted naive-bayes classifiers, dereplicated

<p>16S metabarcoding databases and naive-bayes classifiers specific to the V4-V5 region. Built&nbsp;from&nbsp;the <a href="https://www.arb-silva.de/documentation/release-138/">Silva 138.1 SSU Ref NR 99</a> database using Qiime2 (version 2023.2) and the <a href="https://github.com/BenKaehler/q2-clawback">q2-clawback plugin.</a> Includes&nbsp;weighted classifiers for two Earth Microbiome Project Ontology (EMPO) 3 habitat types: &quot;sediment (saline)&quot;&nbsp;and &quot;water (saline)&quot;&nbsp;, with data&nbsp;downloaded from <a href="https://qiita.ucsd.edu/">Qiita</a>. Sequences were dereplicated with Rescript --p-mode &#39;uniq&#39; ,&nbsp;retaining identical sequence records that have differing taxonomies.</p> <p>Primers used:</p> <p>EMP 16S 515f:&nbsp;GTGYCAGCMGCCGCGGTAA</p> <p>EMP 16S 926r:&nbsp;CCGYCAATTYMTTTRAGTTT</p> <p><strong>Stats</strong></p> <p>286,948 unique sequences</p> <p>309,567 total sequences</p> <p>46,254 unique taxa (Level 7)</p> <table> <caption>File description</caption> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>File</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td>make new 16S silva V4-V5 database.md</td> <td>Markdown with code used to generate databases</td> </tr> <tr> <td>silva-138-99-seqs.qza</td> <td>Full length Silva 138.1 SSU 99 sequences</td> </tr> <tr> <td>silva-138-99-tax.qza</td> <td>Taxa for full length Silva 138.1 SSU 99 database</td> </tr> <tr> <td>silva-138_1-99-515f_926r-uniq-seqs.qza</td> <td>Sequences for 16S V4-V5 (primers 515f, 926r), extracted from Silva 138.1 SSU 99, generated by qiime2-2023.2 (forward compatible), dereplicated</td> </tr> <tr> <td>silva-138_1-99-515f_926r-uniq-taxa.qza</td> <td>Taxa for silva-138_1-99-515f_926r-seqs.qza database, dereplicated</td> </tr> <tr> <td>uniform-silva-138_1-99-515f_926r-uniq-classifier.qza</td> <td>Unweighted (uniform) naive-bayes classifier for 16S V4-V5 (primers 515f, 926r) extracted from Silva 138.1 SSU 99, generated by qiime2-2023.2 (forward compatible)</td> </tr> <tr> <td>silva-138_1-99-515f_926r-uniq-sediment-saline-classifier.qza</td> <td>Weighted naive-bayes classifier for 16S V4-V5 (primers 515f, 926r) extracted from Silva 138.1 SSU 99, weighted for sediment-saline, generated by qiime2-2023.2 (forward compatible)</td> </tr> <tr> <td>silva-138_1-99-515f_926r-q2_2023_2-uniq-sediment-saline-weights.qza</td> <td>Weights used to generate silva-138_1-99-515f_926r-q2_2023_2-sediment-saline-classifier.qza</td> </tr> <tr> <td>silva-138_1-99-515f_926r-uniq-water-saline-classifier.qza</td> <td>Weighted naive-bayes classifier for 16S V4-V5 (primers 515f, 926r) extracted from Silva 138.1 SSU 99, weighted for water-saline, generated by qiime2-2023.2 (forward compatible)</td> </tr> <tr> <td>silva-138_1-99-515f_926r-uniq-water-saline-weights.qza</td> <td>Weights used to generate silva-138_1-99-515f_926r-water-saline-classifier.qza</td> </tr> </tbody> </table> </th> <th scope="col">&nbsp;</th> </tr> </thead> <tbody> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Reference data to the low-wavenumber Raman spectral database of pharamceutical excipients

<p>Supplementary LF-785 and FT-Raman data of all the excipient samples included in the database. More information available in the following paper:&nbsp;<a href="https://doi.org/10.1016/j.vibspec.2020.103021">https://doi.org/10.1016/j.vibspec.2020.103021</a></p> <p>This version also contains all spectra in the&nbsp;.spc file format.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

PR2_V9, a SSU V9 rDNA reference database with functional annotations.

<p>The present data set provides a tab separated text file compressed in a gzip archive. The file includes 63,401 18S V9 rDNA reference sequences for 44,084 unique eukaryotic taxa and 9,759 16S V9 rDNA references sequences for 9,661 unique prokaryotic taxa. It includes the following fields : sequence = nucleic acid sequence of reference; lineage = taxonomic path of the reference sequence; refs = original accession numbers corresponding to the reference sequence; name = reference sequence identifier; taxogroup = high-taxonomic level assignation of the reference sequence. The file also includes six categories of functional annotations: (1) chloroplast: yes, presence of permanent chloroplast; no, absence of permanent chloroplast ; NA, undetermined. (2) symbiont (small partner): parasite, the species is a parasite; commensal, the species is a commensal; mutualist, the species is a mutualist symbiont, most often a microalgal taxon involved in photosymbiosis; no the species is not involved in a symbiosis as small partner; NA, undetermined. (3) symbiont (host): photo, the host species relies on a mutualistic microalgal photosymbiont to survive (obligatory photosymbiosis); photo_falc, same as photo, but facultative relationship; photo_klep, the host species maintains chloroplasts from microalgal prey(s) to survive; photo_klep_falc, same as photo_klep, but facultative; Nfix, the host species must interact with a mutualistic symbiont providing N2 fixation to survive; Nfix_falc, same as Nfix, but facultative; no, the species is not involved in any mutualistic symbioses; NA, undetermined. (4) silicification; yes, the species has a silicified skeleton; no, it does not; NA, undetermined. (5) calcification; yes, the species has a calcified skeleton; no, it does not; NA, undetermined. (6) strontification; yes, the species has a skeleton made of strontium; no, it does not; NA, undetermined.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

16S Reference Database for Fish

<p>Database of 16S reference genes from various fish familes, curated for use in R package DADA2. DBT refers to use for function; <em>assign.taxonomy</em>, DBS refers to use for function; <em>assign.species.&nbsp;</em></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Database of concrete heat of hydration parameters for concrete mixes used in piling (Grant reference EP/R511547/1)

<p>The heat of hydration&nbsp;parameters correspond to the model proposed by Liu et al. (2022). Parameters for five different concrete mixes are provided. They were obtained through finite element back analysis of the response&nbsp;of piles in the field assuming 1D axisymmetric conditions. Details of each of the concrete mixes and the thermal properties employed for soil and concrete in the FE analyses&nbsp;are provided. &nbsp;</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record