Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
103
datasets available to search
ShareScore release 0.9.0
Dataset results
103 results for “reference database”
mSWEEP/mGEMS E. faecalis reference database
<p>E. faecalis reference database.</p>
mSWEEP/mGEMS E. coli reference database
<p>E. coli reference for mSWEEP/mGEMS/Themisto</p>
CAMITAX reference databases
<p>Reference databases for CAMITAX: Taxon labels for microbial genomes.</p> <ul> <li>Mash sketches for all RefSeq genomes (for sketch size 1000),</li> <li>NCBI Taxonomy (names.dmp, nodes.dmp, merged.dmp), and</li> <li>Centrifuge and Kaiju indices for the proGenomes database.</li> </ul>
A Database of References to Different Strands of Characters' Identity in Elye Bokher's 'Bovo D'Antona'
Open the record for dataset details and reuse information.
The Brill Knowledge Graph: A Database of Bibliographic References and Index Terms extracted from Books in Humanities and Social Sciences
<p>We present a complete dataset of linked bibliography and index data, partially disambiguated and augmented with references to external resources, extracted from the Brill’s archive in the field of Classics. Processed book identifiers are listed in a separate text file. Text fragments extracted from different books via this process are then parsed and compared using a string-based similarity metric to form clusters of bibliographic references to the same published work or (variants of) the same subjects discussed in these books. The entire set of references was then disambiguated using Google Books and Crossref APIs.</p> <p><a href="https://jdmdh.episciences.org/11062">Paper about extraction pipeline</a></p> <p><a href="https://www.nkokash.com/documents/KIEM-RDJ.pdf">Paper about extracted KG</a></p> <p> </p>
Reference sequence database for eDNA metabarcoding of San Francisco estuary fishes and invertebrates
<p>Environmental DNA (eDNA) methods complement traditional monitoring and can be configured to detect multiple species simultaneously. One such approach, eDNA metabarcoding, uses high-throughput DNA sequencing to indirectly detect many different organisms, spanning broad taxonomic boundaries, from water samples. We are optimizing a non-invasive, low cost eDNA metabarcoding protocol to be used in conjunction with existing monitoring programs. One resource that is currently lacking for metabarcoding studies in general, including those in the San Francisco Estuary (SFE), is a comprehensive database of DNA barcode reference sequences. Without this foundational data, many species go undetected or misidentified in metabarcoding studies. To meet this need, we generated a custom barcode sequence database for the SFE by DNA sequencing and mining of public DNA seqeunce data for estuarine and freshwater species of interest to monitoring programs and ecological studies. Here we present custom reference sequence databases for three barcodes: Cytochrome C Oxidase I (COI), 12S MiFish and 16S.</p>
Kraken2 Human Pangenome Reference Consortium database
<p>A kraken2 database built from the genome assemblies used by the Human Pangenome Reference Consortium (https://projects.ensembl.org/hprc/). This archive contains the three files required by kraken2, hash.k2d, opts.k2d, and taxo.k2d, along with inspect.txt, which is obtained by running kraken2-inspect on the database, ktaxonomy.tsv, which contains the taxonomy information of the database (obtained by running https://github.com/jenniferlu717/KrakenTools#make_ktaxonomypy).</p> <p>The genomes for this database were downloaded using the assembly summary text file included in this dataset and genome_updater.sh (v0.6.3; https://github.com/pirovc/genome_updater)</p> <pre><code>genome_updater.sh -m -a -f "genomic.fna.gz" -t 8 -e "hprc_assembly_summary.txt" -o HPRC_genomes/</code></pre> <p>The python script prepare_kraken_fasta.py was then used to prepare the assemblies for use in kraken with the following command</p> <pre><code>python prepare_kraken_fasta.py -r -T 9606 -o HPRC.fna HPRC_genomes/</code></pre> <p>The database was then built with kraken2 using the following commands</p> <pre><code>kraken2-build --download-taxonomy --db db/ kraken2-build --add-to-library HPRC.fna --db db/ kraken2-build --build --db db/ --threads 16</code></pre>
Reference sequence database for eDNA metabarcoding of San Francisco estuary fishes and invertebrates
Open the record for dataset details and reuse information.
TK6 toxicogenomics reference database
GEO Series GSE58431. Homo sapiens. 64 samples. Type: Expression profiling by array.
Mare-MAGE database A curated reference database of fish mitochondrial genes
<div> <p>Biodiversity assessment approaches based on molecular biology techniques such as NGS, metabarcoding, RAD-seq, or SnaPshot sequencing, are increasingly used in assessing marine and aquatic ecosystems. In this study, we present a new reference database for fish meta-barcoding based on mitochondrial genes. The <strong><em>Mare-MAGE</em></strong> database contains quality-checked sequences of the mitochondrial genes for 12S ribosomal RNA and Cytochrome c Oxidase I. All sequences were obtained from the National Center for Biotechnology Information- GenBank (NBCI-GenBank) and the European Nucleotide Archive (ENA) and have undergone intensive processing. They were checked for false annotations and non-target anomalies, according to the Integrated Taxonomic Information System (ITIS) and FishBase. The dataset is compiled in ARB-Home, FASTA and Qiime2 formats, and is publicly available from the <strong><em>Mare-MAGE</em></strong> database website (<a href="http://mare-mage.weebly.com/">http://mare-mage.weebly.com/</a> and <a href="https://figshare.com/projects/MARE-MAGE_database_Fish/90917">https://figshare.com/projects/MARE-MAGE_database_Fish/90917</a>). It includes altogether 231,333 COI and 12S rRNA gene sequences of fish covering 19,506 species of 4,058 genera and 586 families.</p> </div> <div> <div> </div> </div>
Used in VirID, self-constructed rRNA reference database
Open the record for dataset details and reuse information.
Creation of a Pediatric Reference Database for the Kerpape-Rennes-EMG-Based Gait Index
ClinicalTrials.gov study NCT07212257. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Assessment of an Automated Optical Coherence Tomography and Camera : Reference Database
ClinicalTrials.gov study NCT06511440. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Prospective Database of Factors Associated With Faecal vs. Double Incontinence in Patients Referred for High Resolution Anorectal Manometry.
ClinicalTrials.gov study NCT05550675. IPD Sharing: NO. Countries: 1. Publications: 0.
Development of a Reference Interval Database With the NeuroCatch™ Platform
ClinicalTrials.gov study NCT03835962. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
IMOvifa Perimeter Reference Database
ClinicalTrials.gov study NCT05792046. IPD Sharing: Not stated. Countries: 1. Publications: 0.
P200TE US Reference Database Study
ClinicalTrials.gov study NCT05844852. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Reference Database & Longitudinal Registry of the Normal and Pathological Aging Brain
ClinicalTrials.gov study NCT02875496. IPD Sharing: NO. Countries: 1. Publications: 0.
OCT Reference Database
ClinicalTrials.gov study NCT01986478. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Topcon 3D OCT-1 Maestro Reference Database Study II
ClinicalTrials.gov study NCT02447120. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.