Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.9.0
Dataset results
103 results for “reference database”
rCRUX Generated Kelly Metazoan 16S Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: Kelly Metazoan 16S<br> Gene: 16S<br> Length of Target: 114-160<br> get_seeds_local() minimum length: 80<br> get_seeds_local() maximum length: 195<br> blast_seeds() minimum length: 50<br> blast_seeds() maximum length: 168<br> max_to_blast: 100<br> Forward Sequence (5'-3'): AGTTACYYTAGGGATAACAGCG<br> Reverse Sequence (5'-3'): CCGGTCTGAACTCAGATCAYGT<br> Reference: Kelly, R. P., O’Donnell, J. L., Lowell, N. C., Shelton, A. O., Samhouri, J. F., Hennessey, S. M., ... & Williams, G. D. (2016). Genetic signatures of ecological diversity along an urbanization gradient. PeerJ, 4, e2444. https://doi.org/10.7717/peerj.2444</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated MiDeca Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: MiDeca<br> Gene: 16S<br> Length of Target: 154-184<br> get_seeds_local() minimum length: 105<br> get_seeds_local() maximum length: 230<br> blast_seeds() minimum length: 66<br> blast_seeds() maximum length: 191<br> max_to_blast: 50<br> Forward Sequence (5'-3'): GGACGATAAGACCCTATAAA<br> Reverse Sequence (5'-3'): ACGCTGTTATCCCTAAAGT<br> Reference: Komai, T., Gotoh, R.O., Sado, T. and Miya, M., 2019. Development of a new set of PCR primers for eDNA metabarcoding decapod crustaceans. Metabarcoding and Metagenomics, 3, p.e33835. https://doi.org/10.3897/mbmg.6.76534</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated MarVer3 Marine Mammal Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: MarVer3 Marine Mammal<br> Gene: 16S<br> Length of Target: 232-274<br> get_seeds_local() minimum length: 160<br> get_seeds_local() maximum length: 345<br> blast_seeds() minimum length: 124<br> blast_seeds() maximum length: 309<br> max_to_blast: 100<br> Forward Sequence (5'-3'): AGACGAGAAGACCCTRTG<br> Reverse Sequence (5'-3'): GGATTGCGCTGTTATCCC<br> Reference: Valsecchi, E., Bylemans, J., Goodman, S. J., Lombardi, R., Carr, I., Castellano, L., ... & Galli, P. (2020). Novel universal primers for metabarcoding environmental DNA surveys of marine mammals and other marine vertebrates. Environmental DNA, 2(4), 460-476. https://doi.org/10.1002/edn3.72</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated 18s SSU3/SSU4 Metazoan Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: 18s SSU3/SSU4 Metazoan<br> Gene: 18S v7<br> Length of Target: 170<br> get_seeds_local() minimum length: 120<br> get_seeds_local() maximum length: 220<br> blast_seeds() minimum length: 98<br> blast_seeds() maximum length: 200<br> max_to_blast: 100<br> Forward Sequence (5'-3'): GGTCTGTGATGCCCTTAGATG<br> Reverse Sequence (5'-3'): GGTGTGTACAAAGGGCAGGG<br> Reference: McInnes, J. C., Alderman, R., Deagle, B. E., Lea, M. A., Raymond, B., & Jarman, S. N. (2017). Optimised scat collection protocols for dietary DNA metabarcoding in vertebrates. Methods in Ecology and Evolution, 8(2), 192-202. http://dx.doi.org/10.1111/2041-210X.12677</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated MiFish Universal 12S Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: MiFish Universal<br> Gene: 12S<br> Length of Target: 163–185<br> get_seeds_local() minimum length: 170<br> get_seeds_local() maximum length: 250<br> blast_seeds() minimum length: 140<br> blast_seeds() maximum length: 250<br> max_to_blast: 1000<br> Forward Sequence (5'-3'): GTGTCGGTAAAACTCGTGCCAGC<br> Reverse Sequence (5'-3'): CATAGTGGGGTATCTAATCCCAGTTTG<br> Reference: Miya, M., Sato, Y., Fukunaga, T., Sado, T., Poulsen, J. Y., Sato, K., ... & Kondoh, M. (2015). MiFish, a set of universal PCR primers for metabarcoding environmental DNA from fishes: detection of more than 230 subtropical marine species. Royal Society open science, 2(7), 150088. <a href="https://doi.org/10.1098/rsos.150088">https://doi.org/10.1098/rsos.150088</a></p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated Baker Dlp1.5-H/Oordlp4 Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: Scott Baker Dlp1.5-H/Oordlp4<br> Gene: D loop<br> Length of Target: 350-390<br> get_seeds_local() minimum length: 245<br> get_seeds_local() maximum length: 495<br> blast_seeds() minimum length: 205<br> blast_seeds() maximum length: 455<br> max_to_blast: 100<br> Forward Sequence (5'-3'): TCACCCAAAGCTGRARTTCTA<br> Reverse Sequence (5'-3'): GCGGGTTGCTGGTTTCACG<br> Reference: Baker, C. S., Steel, D., Nieukirk, S., & Klinck, H. (2018). Environmental DNA (eDNA) from the wake of the whales: Droplet digital PCR for detection and species identification. Frontiers in Marine Science, 5, 133. https://doi.org/10.3389/fmars.2018.00133</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated MiSebastes Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: MiSebastes<br> Gene: CytB<br> Length of Target: 153<br> get_seeds_local() minimum length: 107<br> get_seeds_local() maximum length: 199<br> blast_seeds() minimum length: 71<br> blast_seeds() maximum length: 163<br> max_to_blast: 100<br> Forward Sequence (5'-3'): AAGCTCATTCAAGTGCTT<br> Reverse Sequence (5'-3'): GACCACTTACACAATTCT<br> Reference: Min, M. A., Barber, P. H., & Gold, Z. (2021). MiSebastes: An eDNA metabarcoding primer set for rockfishes (genus Sebastes). Conservation Genetics Resources, 13(4), 447-456. https://doi.org/10.1007/s12686-021-01219-2</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated Taberlet c/h trnl Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: Taberlet c/h trnl<br> Gene: trnl<br> Length of Target: 150<br> get_seeds_local() minimum length: 105<br> get_seeds_local() maximum length: 150<br> blast_seeds() minimum length: 85<br> blast_seeds() maximum length: 128<br> max_to_blast: 100<br> Forward Sequence (5'-3'): CGAAATCGGTAGACGCTACG<br> Reverse Sequence (5'-3'): CCATTGAGTCTCTGCACCTATC<br> Reference: Taberlet, P., Gielly, L., Pautou, G., & Bouvet, J. (1991). Universal primers for amplification of three non-coding regions of. Plant molecular biology, 17, 1105-1109. & Taberlet, Pierre, Eric Coissac, François Pompanon, Ludovic Gielly, Christian Miquel, Alice Valentini, Thierry Vermat, Gerard Corthier, Christian Brochmann, and Eske Willerslev. "Power and limitations of the chloroplast trn L (UAA) intron for plant DNA barcoding." Nucleic acids research 35, no. 3 (2007): e14-e14.</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated Teleo Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: teleo L1848/H1913<br> Gene: 12S<br> Length of Target: 100<br> get_seeds_local() minimum length: 70<br> get_seeds_local() maximum length: 130<br> blast_seeds() minimum length: 40<br> blast_seeds() maximum length: 100<br> max_to_blast: 100<br> Forward Sequence (5'-3'): ACACCGCCCGTCACTCT<br> Reverse Sequence (5'-3'): CTTCCGGTACACTTACCATG<br> Reference: Valentini, A., Taberlet, P., Miaud, C., Civade, R., Herder, J., Thomsen, P. F., ... & Gaboriaud, C. (2016). Next‐generation monitoring of aquatic biodiversity using environmental DNA metabarcoding. Molecular Ecology, 25(4), 929-942. https://doi.org/10.1111/mec.13428</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
rCRUX Generated trnl plants Reference Database
<p>rCRUX generated reference database using NCBI nt blast database downloaded in December 2022.</p> <p>Primer Name: trnl plants<br> Gene: trnl<br> Length of Target: ~85<br> get_seeds_local() minimum length: 60<br> get_seeds_local() maximum length: 110<br> blast_seeds() minimum length: 23<br> blast_seeds() maximum length: 73<br> max_to_blast: 250<br> Forward Sequence (5'-3'): GGGCAATCCTGAGCCAA<br> Reverse Sequence (5'-3'): TTTGAGTCTCTGCACCTATC<br> Reference: Coissac, E., Pompanon, F., Gielly, L., Miquel, C., Valentini, A., Vermat, T., ... & Willerslev, E. (2007). Power and limitations of the chloroplast trnL (UAA) intron for plant DNA barcoding. Nucleic Acids Research 3 (35),.(2007). https://doi.org/10.1093%2Fnar%2Fgkl938</p> <p>We chose default rCRUX parameters for <em>get_blast_seeds</em>() of percent coverage of 70, percent identity of 70, evalue 3e+7, and max number of blast alignments = '100000000' and for <em>blast_seeds</em>() of coverage of 70, percent identity of 70, evalue 3e+7, rank of genus, and max number of blast alignments = '10000000'. </p>
Unlocking natural history collections to improve eDNA reference databases and biodiversity monitoring
Open the record for dataset details and reuse information.
Data from: A global FAOSTAT reference database of cropland nutrient budgets and nutrient use efficiency (1961–2020): nitrogen, phosphorus and potassium
Open the record for dataset details and reuse information.
Collaborative Reference Database as of 2017-02-17
<p><strong>Abstract</strong> (our paper)</p> <p><strong>Code</strong></p> <p>crd-reference-v2.pl:<br> mew.</p> <p><strong>Data</strong></p> <p>crd-reference-v2.txt.gz:<br> mew.</p> <p>crd-reference-v2-urls.txt.gz:<br> mew.</p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br> Sho Sato, Mitsuo Yoshida. Current Status of Reference Services in Japanese Libraries Analyzed with the Collaborative Reference Database. <em>Current Awareness (in Japanese)</em>. no.332, pp.xx-xx, 2017.<br> http://current.ndl.go.jp/ca_list</p> <p><strong>Note</strong></p> <p>The raw data is available in the following page.<br> http://crd.ndl.go.jp/jp/help/general/api.html</p>
AVATAR-Soils Database: A Database of 137Cs and 239+240Pu in Equatorial and Southern Hemisphere Reference Soils
<p>A total of 1122 reference soil profiles with <sup>137</sup>Cs and <sup>239+240</sup>Pu data from 135 publications were included to build a database under the AVATAR Project (“A reVised dATing framework for quantifying geomorphological processes during the AnthRopocene”), with a focus on compiling available data from the Equatorial and Southern Hemisphere regions. The AVATAR-Soils Database covers parts of the continents of Asia and the Sub-Saharan Africa, and the whole of Oceania and South America. The <sup>137</sup>Cs (decay-corrected to 2024) and <sup>239+240</sup>Pu data extracted from the literature include the inventory (in Bq/m<sup>2</sup>) and average activities (in Bq/kg) of the soil profile collected, and the <sup>137</sup>Cs/<sup>239+240</sup>Pu activity and <sup>239</sup>Pu/<sup>240</sup>Pu atomic ratios. In addition to the <sup>137</sup>Cs and <sup>239+240</sup>Pu data, the associated spatial, climatic, and topographic parameters and sampling details were also added in the database. The database contains the metadata describing the column names, separate tabs for <sup>137</sup>Cs and <sup>239+240</sup>Pu, and the list of publications.</p> <p> To cite the AVATAR-Soils Database, please use:</p> <p>Dicen, G., Guillevic, F., Gupta, S., Chaboche, P.-A., Meusburger, K., Sabatier, P., Evrard, O., and Alewell, C.: <strong>Distribution and sources of fallout <sup>137</sup>Cs and <sup>239+240</sup>Pu in equatorial and Southern Hemisphere reference soils</strong>, Earth Syst. Sci. Data, 17, 1529–1549, https://doi.org/10.5194/essd-17-1529-2025, 2025.</p> <p>***NOTES TO RESEARCHERS***</p> <p>Researchers who wish to add data to the AVATAR-Soils Database may do so through the following link: https://docs.google.com/spreadsheets/d/15R-rDMH6zW65B3ok3YRb0cdtyY67X_mdlC0H1t1atoY/edit?usp=drive_link</p>
RNA viral reference database
<p><span>The RNA viral reference database was composed of (1) 6,621 complete RNA viral genomes collected from the NCBI virus database (accessed on August 12<sup>th</sup>, 2022, [1]), (2) 378,253 RNA viral contigs from the RVMT database (v3)[2], and (3) 858 RNA viral genomes published in a terrestrial RNA viral study [3].</span></p> <p><span>[1] NCBI virus database, accessible at https://www.ncbi.nlm.nih.gov/labs/virus/vssi/#/</span></p> <p><span>[2] <span>Neri U, Wolf YI, Roux S, Camargo AP, Lee B, Kazlauskas D, Chen IM, Ivanova N, Allen LZ, Paez-Espino D.<strong><span> </span></strong>2022. Expansion of the global RNA virome reveals diverse clades of bacteriophages. Cell 185:4023-4037. e18.</span></span></p> <p><span>[3] <span>Chen Y-M, Sadiq S, Tian J-H, Chen X, Lin X-D, Shen J-J, Chen H, Hao Z-Y, Wille M, Zhou Z-C.<strong><span> </span></strong>2022. RNA viromes from terrestrial sites across China expand environmental viral diversity. Nature Microbiology 7:1312-1323.</span></span></p>
CM2 Reference Database Uniref100/KO
<p>CM2 DIAMOND reference database - Version 3. Compatible with DIAMOND version 2.1.11 specifically. </p>
SingleM reference database
<p>DEPRECATED: See <a href="https://aus01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fzenodo.org%2Frecord%2F5739612&data=05%7C01%7Crossen.zhao%40hdr.qut.edu.au%7Cc3c9e94152b54b06e14008dbc868dcfa%7Cdc0b52a368c544f7881d9383d8850b96%7C0%7C0%7C638324124934642739%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=wycEGojHU4c%2B6B3b7Ds9BwJ9IPUGfy%2BDmlXLOHn%2F3dY%3D&reserved=0">https://zenodo.org/record/5739612</a></p> <p>SingleM reference database. Metapackage containing packages for 59 broad-spectrum single-copy marker genes. Associated sequences and taxonomy based on GTDB r202 release.</p>
MDMcleaner reference database
<p>MDMcleaner reference database used at time of publication.</p> <p>based on:</p> <p>GTDB release r95</p> <p>RefSeq release 203</p> <p>Silva version 138.1</p>
Transitioning from environmental genetics to genomics using mitogenome reference databases
<p><span>Species detection using eDNA is revolutionizing the global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of complete mitogenomic sequence information to evaluate the effectiveness of such data to differentiate, identify and detect taxa. We created the Oregon Biodiversity Genome Project working group to utilize recent advances in sequencing technology to create a database of complete, near error-free mitogenomic sequences for all of Oregon's resident freshwater fishes. So far, we have successfully assembled the complete mitogenomes of 313 specimens of freshwater fish representing 7 families, 55 genera, and 129 (88%) of the 146 resident species and lineages. Our comparative analyses of these sequences illustrate that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays are not consistently diagnostic for species-level identification and that no single region is best for metabarcoding Oregon's fishes. However, often-overlooked intergenic regions of the mitogenome such as the D-loop have the potential to reliably diagnose and differentiate species. This project provides a blueprint for other researchers to follow as they build regional databases. It also illustrates the taxonomic value and limits of complete mitogenomic sequences, and how current eDNA assays and the "PCR-free" environmental genomics methods of the future can best leverage this information.</span></p>
DNA reference database for macroinvertebrates
<p>This repository hosts a ready-to-use database for the study of macroinvertebrate diversity using DNA metabarcoding with the fwhF2/EPTDr2n primer set (Leese et al. 2022).</p> <p>The database was developed from the MIDORI database and completed with some data from the BOLD database.</p> <p>The process of preparation and curation of the database is described below:</p> <p>1. The MIDORI CO1 RAW v247 database was downloaded from the official servers.</p> <p>2. Reference sequences were matched against the fwhF2 forward primer sequences. Non-matching sequences were removed. Bases preceding the forward primer were removed.</p> <p>3. Reference sequences with a length of less than 142 bp were removed. Bases following the position 142 were removed.</p> <p>4. Taxonomy was simplified to keep only 8 ranks: superkingdom, kingdom, phylum, class, order, family, genus and species.</p> <p>5. Taxonomic nomenclature was harmonized using `refdb::refdb_clean_tax_harmonize_nomenclature`</p> <p>6. Extra words were removed from taxonomic names using `refdb::refdb_clean_tax_remove_extra`</p> <p>7. Subspecific information were removed from taxonomic names using `refdb::refdb_clean_tax_remove_subsp`</p> <p>8. Missing taxonomic names were identified and normalized using `refdb::refdb_clean_tax_NA`.</p> <p>9. Hybrids and taxonomic names with qualifiers of uncertainty were converted to NA.</p> <p>10. Duplicates (identical sequences and taxonomy) were removed.</p> <p>11. Sequences with more than ambiguous nucleotides (N) were removed).</p> <p>12. Sequences with low taxonomic precision (above phylum) were removed.</p> <p>13. A random subset of 10 sequences was retained for each taxa.</p> <p>14. For a given taxonomic rank, records with NA values were removed if they were not the only representative of the upper clade.</p> <p>15. For Switzerland: Missing EPT species and IBCH families were searched in BOLD and merged with the reference database. The same filters and cleaning steps (2-14) were then applied.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.