Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

76

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

76 results for “reference libraries”

Learn how ShareScore rates datasets ↗
zenodo52/100

Liquid Chromatography - Tandem Mass Spectrometry (LC-MS/MS) and Gas Chromatography - Mass Spectrometry (GC-MS) Reference Libraries from Global Natural Products Social Molecular Networking (GNPS) and National Institute of Standards and Technology (NIST) WebBook Processed for Spectral Library Matching

<div>In order to obtain a high-quality LC-MS/MS reference database for spectral library matching, we selected 22 high-quality GNPS tandem mass spectrometry databases generated under the positive ion mode. Further preprocessing similar to Huber et al involving mass-to-charge (m/z) and intensity filtering yields the database found in the file LCMS_GNPS_reference_library.csv which contains 14,705 electrospray ionization (ESI) mass spectra, each of which corresponds to a unique compound. The NIST WebBook database was used to construct GC-MS database contained in the file GCMS_NIST_WebBook.csv. This database contains 23,721 electron ionization (EI) mass spectra, each of which corresponds to a unique non-hyphenated Chemical Abstract Service (CAS) Registry Number.</div> <div>&nbsp;</div> <div>Both LC-MS/MS and GC-MS databases are organized into three columns: one for the identifier, one for the m/z values, and one for the intensity values. For example, if spectrum A has 20 ion fragments, then there will be 20 rows corresponding to spectrum A in the corresponding database with the identifier A repeated 20 times with the corresponding m/z and intensity values.</div>

opencc-by-4.0Jul 2024View details →
zenodo48/100

MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes

<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here:&nbsp;<a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). &nbsp;&nbsp;</p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2&nbsp;Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name:&nbsp;</strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe:&nbsp;</strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' &nbsp;values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' &nbsp;values over 50%: FLAG_RP63.</li> <li><strong>flag_sum:&nbsp;</strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value).&nbsp;</p> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis.&nbsp;</p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0&nbsp; functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes.&nbsp;<br><br></p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples

<p>[This repository contains the source data for the manuscript &quot;A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples&quot; by deWaard et al., 2019. BioRxiv]</p>

opencc-zeroJun 2019View details →
dryad40/100

BAGS: an automated barcode, audit & grade system for DNA barcode reference libraries

<p>Biodiversity studies greatly benefit from molecular tools, such as DNA metabarcoding, which provides an effective identification tool in biomonitoring and conservation programmes. The accuracy of species-level assignment, and consequent taxonomic coverage, relies on comprehensive DNA barcode reference libraries. The role of these libraries is to support species identification, but accidental errors in the generation of the barcodes may compromise their accuracy. Here we present an R-based application, BAGS (Barcode, Audit &amp; Grade System; https://github.com/tadeu95/BAGS), that performs automated auditing and annotation of cytochrome c oxidase subunit I (COI) sequences libraries, for a given taxonomic group of animals, available in the Barcode of Life Data System (BOLD). This is followed by implementing a qualitative ranking system that assigns one of five grades (A to E) to each species in the reference library, according to the attributes of the data and congruency of species names with sequences clustered in Barcode Index Numbers (BINs). Our goal is to allow researchers to obtain the most useful and reliable data, highlighting and segregating records according to their congruency. Different tests were performed to perceive its usefulness and limitations. BAGS fulfils a significant gap in the current landscape of DNA barcoding research tools by quickly screening reference libraries to gauge the congruence status of data and facilitate the triage of ambiguous data for posterior review. Thereby, BAGS has the potential to become a valuable addition in forthcoming DNA metabarcoding studies, in the long term contributing to globally improve the quality and reliability of the public reference libraries.</p>

opencc-zeroAug 2020View details →
zenodo40/100

Fig 5 in A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library

Fig 5. Neighbour-joining tree showing all available DNA barcodes for species in family Rhinolophidae reported from Peninsular Malaysia. The percentage of pseudoreplicate trees (±70%) in which the DNA barcodes clustered together in the bootstrap test (500 pseudoreplicates) are shown above the branches. Abbreviation as follows: PM = Peninsular Malaysia, VN = Vietnam, BN = Borneo (including Sabah &amp; Sarawak of East Malaysia, Brunei and Kalimantan Indonesia), TH = Thailand, LA = Laos, SM = Sumatera Indonesia, JV = Java Indonesia, IND = India, CH = China, CM = Cambodia, MN = Myanmar. https://doi.org/10.1371/journal.pone.0179555.g005

opencc-by-4.0Jul 2017View details →
zenodo40/100

Fig 4 in A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library

Fig 4. Neighbour-joining tree showing all available DNA barcodes for species in family Hipposideridae reported from Peninsular Malaysia. The percentage of pseudoreplicate trees (±70%) in which the DNA barcodes clustered together in the bootstrap test (500 pseudoreplicates) are shown above the branches. Abbreviation as follows: PM = Peninsular Malaysia, VN = Vietnam, BN = Borneo (including Sabah &amp; Sarawak of East Malaysia, Brunei and Kalimantan Indonesia), TH = Thailand, LA = Laos, SM = Sumatera Indonesia, CH = China, CM = Cambodia. https://doi.org/10.1371/journal.pone.0179555.g004

opencc-by-4.0Jul 2017View details →
zenodo40/100

Fig 2 in A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library

Fig 2. Neighbour-joining tree showing all available DNA barcodes for species in family Pteropodidae reported from Peninsular Malaysia. The percentage of pseudoreplicate trees (±70%) in which the DNA barcodes clustered together in the bootstrap test (500 pseudoreplicates) are shown above the branches. Abbreviation as follows: PM = Peninsular Malaysia, VN = Vietnam, JV = Java, Indonesia, BN = Borneo (including Sabah, Sarawak, Brunei and Kalimantan), TH = Thailand, LA = Laos. https://doi.org/10.1371/journal.pone.0179555.g002

opencc-by-4.0Jul 2017View details →
zenodo40/100

Fig 3 in A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library

Fig 3. Neighbour-joining tree showing all available DNA barcodes for species in families Emballonuridae, Megadermatidae, Molossidae and Nycteridae reported from Peninsular Malaysia. The percentage of pseudoreplicate trees (±70%) in which the DNA barcodes clustered together in the bootstrap test (500 pseudoreplicates) are shown above the branches. Abbreviation as follows: PM = Peninsular Malaysia, VN = Vietnam, BN = Borneo (including Sabah &amp; Sarawak of East Malaysia, Brunei and Kalimantan Indonesia), TH = Thailand, LA = Laos, SM = Sumatera Indonesia, CH = China. https://doi.org/10.1371/journal.pone.0179555.g003

opencc-by-4.0Jul 2017View details →
zenodo40/100

Reference Sequence Library Resources - Maine-eDNA

<p>The following files and resources are associated with the Maine-eDNA Reference Library Research Group - aiming to create reference sequence library resources for researchers part of Maine-eDNA or otherwise interested in leveraging eDNA tools for research in the Gulf of Maine.</p> <p>These include:</p> <ul> <li> <p>RoughWorkflow.zip</p> </li> <ul> <li> <p>contains a NCBI scraping script to build reference databases based on an input species list, a configuration file for the script, and genbankr version - most useful for shorter species lists (time-intensive)</p> </li> </ul> <li> <p>12S_REFDB.fasta</p> </li> <ul> <li> <p>A DADA2-compliant reference library for 12S sequences, built with the RoughWorkflow based on the GNRMaineSpecies_May2024 species list</p> </li> </ul> <li> <p>COI_REFDB.fasta</p> </li> <ul> <li> <p>A DADA2-compliant reference library for COI sequences, built with the RoughWorkflow based on the GNRMaineSpecies_May2024 species list</p> </li> </ul> <li> <p>GitHub Repo - referee - <a href="https://github.com/BigelowLab/referee">https://github.com/BigelowLab/referee</a></p> </li> <ul> <li> <p>Scripts for building reference databases for the Maine-eDNA project through downloading GenBank - this workflow is suggested especially for large species lists</p> </li> </ul> <li> <p>GitHub Repo - refdbtools - <a href="https://github.com/BigelowLab/refdbtools">https://github.com/BigelowLab/refdbtools</a>&nbsp;</p> </li> <ul> <li> <p>R language package to assist in making eDNA reference databases</p> </li> </ul> <li> <p>SpeciesListMeta_Shareable.xlsx</p> </li> <ul> <li> <p>Describes the sources from which the GNRMaineSpecies_May2024.csv and GNRMaineTaxonomiedSpecies_May2024.csv species lists were compiled - sources not associated with a link were found as separate files and are hosted elsewhere. Species list sources are courtesy of public datasets, Maine-eDNA researchers, and collaborators</p> </li> </ul> <li> <p>GNRMaineSpecies_May2024.csv</p> </li> <ul> <li> <p>A Maine (and surrounding area) species list ran through taxize&rsquo;s gnr_resolve to resolve species names and fill out taxonomy (full results)</p> </li> </ul> <li> <p>GNRMaineTaxonomiedSpecies_May2024.csv</p> </li> <ul> <li> <p>A Maine (and surrounding area) species list ran through taxize&rsquo;s gnr_resolve to resolve species names and fill out taxonomy (only species results that could be resolved with taxonomy)</p> </li> </ul> </ul> <p>&nbsp;</p> <p>Contact Beth Y. Davis - bethy.davis4@gmail.com for questions</p> <p>&nbsp;</p> <p>###</p> <p>Changelog:</p> <p>May 16, 2023 (Version 1.0) - Initial upload</p> <p>July 11, 2023 (Version 1.1) - Did additional cleaning to the MaineSpeciesList_Clean file and uploaded the new version - July2023_SpeciesList</p> <p>May 20, 2024 (Version v3) -&nbsp; Additional cleaning to correct deduplication errors and ran taxize's gnr_resolve to resolve names and fill out taxonomy. The version update adds two files, GNRMaineSpecies_May2024.csv containing the full result of gnr_resolve, and GNRMaineTaxonomiedSpecies_May2024.csv contains only those species from the original list that could be resolved with taxonomy. The SpeciesListMeta_Shareable.csv has not been updated but is still an accurate tracker of the sources from which the species names were gained from.</p> <p>November 27, 2024 (Version 4.0) - Updated Zenodo description and added the RoughWorkflow R files, 12S and COI files</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

InpactorDB: A Plant classified lineage-level LTR retrotransposon reference library for free-alignment methods based on Machine Learning

<p>LTR retrotransposons are mobile elements that make up the major part of most plant genomes. Their identification and annotation via bioinformatics approaches represent a major challenge in the era of massive plant genome sequencing. In addition to their involvement in the variation in genome size, these elements are also associated in the function and structure of different chromosomal regions and in the alteration of the function of coding regions, among others. Several plant retrotransposon sequence databases of LTR retrotransposons are available with public access such as PGSB, RepetDB or restricted access such as Repbase. Although they are useful for approaches to identify LTR-RTs in new genomes by similarity, the elements of these databases are not classified down to the lineage/family level. with great depth.&nbsp;</p> <p>Here, we present InpactorDB a semi-curated dataset composed of 130,511 elements from 195 plant genomes (belonging to 108 plant species), classified down to the lineage level. This data set has been used to train two deep neural networks (one fully connected and one convolutional) for fast classification of elements. Used in lineage-level classification approaches, we obtain a score above 98% of F1-score, precision and recall.&nbsp;</p> <p>In order to classify elements of the &lsquo;LTR_STRUC&rsquo; and &lsquo;EDTA&rsquo; datasets, we used the methodology proposed by Inpactor, which uses homology-based strategy with known coding domains belonging to LTR-RTs. We utilized the RexDB &nbsp;domain library as reference. LTR-RTs were classified into superfamilies, Gypsy (RLG) or Copia (RLC) and sub-classified into lineages according to the similarities of five different amino acid reference domains (GAG, AP, RT, RNAseH, and INT domains). In addition, we applied filters to remove keep only intact elements:</p> <p>1) to remove predicted elements with domains from two different superfamilies (i.e. Gypsy and Copia),</p> <p>2) or elements with domains belonging to two or more different lineages,</p> <p>3) to remove elements with lengths different than those reported by Gypsy Database (GyDB) with a tolerance of 20%,</p> <p>4) to delete incomplete elements which has less than three identified domains, and</p> <p>5) to remove elements with insertions of TE class II (reported in Repbase).&nbsp;</p> <p>The final non-redundant version of InpactorDB consists of 67,305 LTR retrotransposons. Both redundant and non-redundant versions of InpactorDB are available in Fasta&nbsp;format in which sequences have identifiers with the following general&nbsp;Identification code:</p> <p>&gt;Superfamily-Lineage-plant_family-specie-source-length-ID,</p> <p>Where Superfamily&nbsp;can&nbsp;is either RLC (for Copia) or RLG (for&nbsp;Gypsy), Lineage/family&nbsp;follows&nbsp;following&nbsp;the RexDB nomenclature, source&nbsp;(can be&nbsp;Repbase, RepetDB, PGSB, LTR_STRUC or EDTA&nbsp;datasets), length, and ID,&nbsp;is&nbsp;a unique number which identify each element inside&nbsp;the&nbsp;InpactorDB.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Figure 3 in Butterfly-parasitoid-hostplant interactions in Western Palaearctic Hesperiidae: a DNA barcoding reference library

Figure 3. Mounted specimens illustrating the species of Microgastrinae recovered in this study. A, Cotesia glabrata Telenga ex Carcharodus alceae, Italy. Adult plus cocoons. Gregarious parasitoid; brood sizes vary considerably, host usually well grown or prepupal when killed. The other Cotesia species (near glabrata) look similar and behave in the same way. B, Dolichogenidea sp. near sicaria Marshall, ex Carcharodus alceae, Spain. Adult plus cocoon. Solitary parasitoid, killing the host while still quite young. C, Microgaster australis Thomson, ex Muschampia stauderi, Greece. Adult plus cocoon. Solitary parasitoid, usually killing the host as a prepupa. D, Microgaster nobilis Reinhard, ex Carcharodus alceae, Spain. Adult plus cocoon. Solitary parasitoid, usually killing the host as a prepupa. All specimens are in the collection of the National Museums of Scotland.

opencc-by-4.0Sep 2022View details →
zenodo40/100

Figure 4 in Butterfly-parasitoid-hostplant interactions in Western Palaearctic Hesperiidae: a DNA barcoding reference library

Figure 4. Interaction matrices showing the recorded interactions of Hesperiidae and their hostplants (A), Hesperiidae and their parasitoids (B) and parasitoids and hostplants of Hesperiidae (C). White squares indicate recorded interactions between the taxa in the corresponding row and column, while blue squares indicate lack of interaction.

opencc-by-4.0Sep 2022View details →
zenodo40/100

Figure 2 in Butterfly-parasitoid-hostplant interactions in Western Palaearctic Hesperiidae: a DNA barcoding reference library

Figure 2. Circular cladogram showing ecological interactions among European and North African Hesperiidae, their hostplants, and their microgastrine parasitoids, recovered through DNA barcoding for Hesperiidae and/or parasitoids. Hesperiid, parasitoid and plant cladograms are coloured in orange, blue and green, respectively. Lines representing interactions with parasitoids are coloured in blue, while lines involving hostplant interactions are coloured in green.

opencc-by-4.0Sep 2022View details →
zenodo40/100

Figure 1 in Butterfly-parasitoid-hostplant interactions in Western Palaearctic Hesperiidae: a DNA barcoding reference library

Figure 1. Representation of the study system. Hesperiid larvae feeding on their hostplants can be attacked by a number of parasitoids, which can in turn be attacked by various hyperparasitoids. A, Spialia rosae on its hostplant Rosa sicula. B, third instar larva of Sp. rosae on a silk shelter. C, Microgaster australis parasitizing an L3 Sp. rosae larva. D, Gelis sp. parasitizing M. australis on its cocoon after emerging from the Sp. rosae larva. Drawings by Martí Franch.

opencc-by-4.0Sep 2022View details →
zenodo40/100

Public reference library for edible dormouse calls

<p>The acoustics of small mammals, particularly dormice species, are generally understudied. We explored the vocalisations and acoustic behaviour of the edible dormouse (<em>Glis glis</em>) in various environments in Catalonia (northern Iberian Peninsula) using ultrasound recorders between 2022 and 2023. Up to five different types of calls were identified in various environmental conditions (captivity and free-ranging animals) and developmental stages, from pups to adults. Additionally, one new call type was described, highlighting the plasticity of their vocalisations, and ultrasonic sound production was discovered in pups, suggesting ontogenetic changes in the vocal repertoire. With this project, we emphasise the potential of the acoustic method as a non-invasive tool for studying ecological behaviours and interactions, or early detection of the species in the natural environment.</p> <p>We, hereby, provide an open reference library for edible dormouse calls, laying the groundwork for a better understanding of its acoustics, behaviour and conservation. The compilation provides five clean sequences of high quality for each of the five call types recorded under different environmental conditions (totalling 35 recordings). This includes the chirp (captivity and wild), blow (captivity), aggressive call (captivity and wild), pups during manipulation (wild) and free-ranging pups (wild). The collection of sounds was prepared under the Dormouse Project (<a href="http://www.dormice.org/" target="_blank" rel="noopener">www.dormice.org</a>)</p> <p>This work was supported by the Barcelona Zoo Fundation under the Research and Conservation Program Grant; Generalitat de Catalunya under Grant number ARD264/23/000001; and Diputaci&oacute; de Barcelona under Grant number 2023/0005732.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Linked collectors and determiners for: A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library.

Natural history specimen data linked to collectors and determiners held within, "A checklist of the bats of Peninsular Malaysia and progress towards a DNA barcode reference library". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/6ea2cc5c-857b-4b47-8135-8bff7efbd1fc">https://bionomia.net/dataset/6ea2cc5c-857b-4b47-8135-8bff7efbd1fc</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/6ea2cc5c-857b-4b47-8135-8bff7efbd1fc">https://gbif.org/dataset/6ea2cc5c-857b-4b47-8135-8bff7efbd1fc</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
dryad40/100

Data from: Memory-bound k-mer selection for large and evolutionary diverse reference libraries

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad40/100

BAGS: an automated barcode, audit & grade system for DNA barcode reference libraries

Open the record for dataset details and reuse information.

publicSep 2020View details →
zenodo36/100

eDNAExplorer.org reference libraries

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
dryad36/100

A new taxonomist-curated reference library of DNA barcodes for Neotropical electric fishes (Teleostei: Gymnotiformes)

<p>DNA barcoding is a useful tool for identifying species; however, successful barcode-based identification requires a reference library of barcode sequences from accurately identified specimens. Here we present a reference library of <em>co1</em> barcode sequences for the Neotropical electric knifefish order Gymnotiformes (Teleostei: Ostariophysi), a model taxon for studies of tropical diversification and biogeography, genomics, behaviour, and neurobiology. Our library contains barcodes for 167 of the ca. 270 valid species of gymnotiforms derived from geo-referenced museum voucher specimens, and includes sequences from 26 type specimens and 21 specimens from type localities, most of which we collected. To assess the state of gymnotiform barcodes in two main public barcode repositories, GenBank and BOLD, we compared the barcodes in these databases to our reference library. Our analysis shows that a considerable proportion of gymnotiform barcodes in GenBank and BOLD are mis- or unidentified. We encourage taxonomists to develop and publish barcode reference libraries composed of carefully curated barcode sequences.</p>

opencc-zeroJun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record