Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Genus level DNA sequence data for three genes (matK, rbcL, trnH-psbA) for the paper: A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability

<p>This data file contains the consensus DNA sequences in fasta format, of 64 tree genera found in Mediterranean Europe, following the checklist of M&eacute;dail et al. (2019).&nbsp;</p> <p>The data are used in a manuscript submitted for publication to Botany Letters and currently under revision. The manuscript is entitled: &quot;<em>A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability</em>&quot;. Its authors are: Marwan Cheikh Albassatneh, Marcial Escudero, Loic Ponge<sup>*</sup>, Anne-Christine Monnet, Juan Arroyo, Toni Nikolic, Gianluigi Bacchetta, Francesca Bagnoli, Panayotis Dimopoulos, Agathe Leriche, Fr&eacute;d&eacute;ric M&eacute;dail, Anne Roig, Ilaria Spanu, Giovanni Giuseppe Vendramin, Arndt Hampe, Bruno Fady.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Human sequence alignment data set used for analysis of SPDI algorithm and tools

<p>Collection of alignment segments produced on October 30, 2019.&nbsp;&nbsp;The ADS currently consists of over 2,680,000 pairwise alignment segments generated from over 350,000 distinct input sequences.&nbsp;&nbsp;</p> <ul> <li> <p>Old assembly to current Genome Reference Consortium (GRC) <a href="http://f1000.com/work/citation?ids=111899&amp;pre=&amp;suf=&amp;sa=0">(Church et al., 2011)</a> primary assemblies (e.g. GRCh36(hg18) or GRCh37(hg19) with GRCh38(hg38))</p> </li> </ul> <ul> <li> <p>Patches, alternative loci, or pseudoautosomal regions (PAR) to GRC primary assembly</p> </li> <li> <p>RefSeq <a href="http://f1000.com/work/citation?ids=2599029&amp;pre=&amp;suf=&amp;sa=0">(O&rsquo;Leary et al., 2016)</a> and select GenBank <a href="http://f1000.com/work/citation?ids=6183037&amp;pre=&amp;suf=&amp;sa=0">(Benson et al., 2018)</a> transcripts to selected RefSeq genomic regions, also known as RefSeqGene (NG), a member of the Locus Reference Genome (LRG) collaboration <a href="http://f1000.com/work/citation?ids=3225699&amp;pre=&amp;suf=&amp;sa=0">(Dalgleish et al., 2010)</a>.</p> </li> <li> <p>Current RefSeq transcripts (NM/NR/XM/XR) and RefSeq genomic regions (NG) to the latest Assembly</p> </li> <li> <p>Previous versions of NG and RefSeq transcripts (NM/NR) to GRC primary assembly</p> </li> </ul>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Sequencing data from: Divergent lineages in a young species: the case of Datilillo (Yucca valida), a broadly distributed plant from the Baja California Peninsula

<div> <div> <p><strong>Premise:&nbsp;</strong>Globally, barriers triggered by climatic changes have caused habitat fragmentation and population allopatric divergence. Across North America, oscillations during the Quaternary have played important roles in the distribution of wildlife. Notably, diverse plant species from the Baja California Peninsula in western North America, isolated during the Pleistocene glacial&ndash;interglacial cycles, exhibit strong genetic structure and highly concordant divergent lineages across their ranges. A representative plant genus of the peninsula is&nbsp;<em>Yucca</em>, with&nbsp;<em>Y. valida</em>&nbsp;having the widest range. Although a dominant species, it has an extensive distribution discontinuity between 26&deg; N and 27&deg; N, suggesting restricted gene flow. Moreover, historical distribution models indicate the absence of an area with suitable conditions for the species during the Last Interglacial, making it an interesting model for studying genetic divergence.<br>Methods: We assembled 4411 SNPs from 147 plants of&nbsp;<em>Y. valida</em>&nbsp;throughout its range to examine its phylogeography to identify the number of genetic lineages, quantify their genetic differentiation, reconstruct their demographic history and estimate the age of the species.<br>Results: Three allopatric lineages were identified based on the SNPs. Our analyses support that genetic drift is the driver of genetic differentiation among these lineages. We estimated an age of less than 1 million years for the common ancestor of&nbsp;<em>Y. valida</em>&nbsp;and its sister species.<br>Conclusions: Habitat fragmentation caused by climatic changes, low dispersal, and an extensive geographical range gap acted as cumulative mechanisms leading to allopatric divergence in&nbsp;<em>Y. valida</em>.</p> </div> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Fig. 3 in New data on Thelohanellus nikolskii Achmerov, 1955 (Myxosporea, Myxobolidae) a parasite of the common carp (Cyprinus carpio, L.): The actinospore stage, intrapiscine tissue preference and molecular sequence

Fig. 3. Thelohanellus cysts on the scales of an aged common carp specimen.

opencc-by-4.0Aug 2021View details →
zenodo36/100

Fig. 1 in The neglected diversity: Description and molecular characterisation of Trypanosoma haploblephari Yeld and Smit, 2006 from endemic catsharks (Scyliorhinidae) in South Africa, the first trypanosome sequence data from sharks globally

Fig. 1. Map of sampling sites on the south coast of South Africa.

opencc-by-4.0Aug 2021View details →
zenodo36/100

Fig. 1 in New data on Thelohanellus nikolskii Achmerov, 1955 (Myxosporea, Myxobolidae) a parasite of the common carp (Cyprinus carpio, L.): The actinospore stage, intrapiscine tissue preference and molecular sequence

Fig. 1. Thelohanellus nikolskii cysts on the fins of carp fingerlings.

opencc-by-4.0Aug 2021View details →
zenodo36/100

Flow Cytometry Data from "Bacterial cell surface characterization by phage display coupled to high-throughput sequencing"

<p>This record contains the flow cytometry data from the manuscript "Bacterial cell surface characterization by phage display coupled to high-throughput sequencing."</p> <p>Files are in <a href="https://docs.flowjo.com/flowjo/advanced-features/fj-acs/">Archive Cytometry Standard (ACS) format</a> . Each <code>.acs</code> file is a zip container which holds both the raw <code>.fcs</code> files and a FlowJo workspace (<code>.wsp</code>) file.</p> <p>Keywords in the workspace file identify which primary antibody (<code>primary</code>) was used and which cell genotype (<code>strain</code>) was used for each sample. The workspace also encodes the&nbsp;gating scheme and compensation matrix applied to each sample. Plots in the manuscript are exported from Layout views in the workspace.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Chronic social defeat stress induces meningeal neutrophilia via type I interferon signaling: single cell RNA sequencing data

<p>Meningeal single cell RNA sequencing data</p> <p>Meningeal samples were collected from both dorsal and ventral skull, avoiding inclusion of choroid plexus. Samples were digested in 2.5 mg/mL Collagenase D (Cat. #11088858001; Roche) and 12.5 &mu;L of 0.5 mg/mL DNAseI (Cat. #L5002139; Worthington), put on a shaker at 370C for 30 m, diluted with cold HBSS + 0.1% BSA, and mashed through a 70 &mu;m cell strainer prior to sorting.</p> <p>Data represent live, nucleated, singlet cells (DAPI-DRAQ5+) sorted on a BD FACS Aria Fusion into HBSS + 10% FBS prior to droplet encapsulation using 10x Genomics&rsquo; Drop-seq platform (Chromium v2).</p> <p>10X chip lane is indicate by 'group' column</p> <p>Group 1 = 4 pooled homecage control (unstressed) mice</p> <p>Group 2 = 4 pooled homecage control (unstressed) mice</p> <p>Group 3 = 4 pooled mice exposed to chronic social defeat for 14 days; tissue was collected 2 hours following final defeat</p> <p>See the following repositories for data processing:</p> <p><a href="https://github.com/maryellenlynall/2019_bcell_stress/blob/master/bcellstress20.Rmd">https://github.com/maryellenlynall/2019_bcell_stress/</a> (processing from raw files starts at bcellstress020.Rmd)</p> <p><a href="https://github.com/staceykigar/meningeal_neut/">https://github.com/staceykigar/meningeal_neut/</a></p> <p>We also provide a processed dataset (processed.RData) with assays 'counts' and 'logcounts' which is the processed single cell object saved at line "# Save object for upload to Zenodo" in script&nbsp;<a href="https://github.com/staceykigar/meningeal_neut/">https://github.com/staceykigar/meningeal_neut/</a>neutrophilstress01.Rmd&nbsp;</p> <p>Cluster annotations are in sce$Annotation</p> <p>Neutrophil subcluster annotations are in sce$Subcluster</p> <p>Sample condition is in sce$cond, where "HC" indicates homecage control and "SD" indicates chronic social defeat</p> <p>10X chip lane is in sce$group</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data for paper: Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences

<p>This deposit contains data for the paper entitled: "<strong>Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences</strong>"</p> <p>Contents include:</p> <ul> <li><strong>00_sequencing-metadata.xlsx</strong> - Metadata for each sequencing sample.</li> <li><strong>01_raw-fastqs.zip</strong> - Raw basecalled FASTQ data.</li> <li><strong>02_clean-fastas.zip</strong> - Cleaned FASTA files (removal of adapters and barcode sequences).</li> <li><strong>03_read-statistics.zip</strong> - Read statistics for all samples.</li> <li><strong>04_sequence-analysis.zip</strong> - Sequence analysis output for all samples.</li> <li><strong>05_qpcr-data.xlsx</strong> - qPCR data for all the dNTP mixes studied.</li> <li><strong>06_analysis-scripts.zip</strong> - Analysis scripts.</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data mining antibody sequences for database searching in bottom-up proteomics

<p>Mass spectrometry (MS)-based proteomics is a powerful method for identifying and quantifying antibodies. Among the various MS approaches, bottom-up proteomics is especially effective for analyzing thousands of antibodies in complex mixtures. In this method, proteins are enzymatically digested into smaller peptides, typically using the protease trypsin, which are then analyzed via mass spectrometry. These peptides are matched to sequences in standard databases like UniProt or NCBI-RefSeq for identification.</p> <p>However, a major limitation of this approach is the absence of comprehensive disease-specific antibody databases. Current databases, such as UniProt, include only a fraction of the antibody sequences present in the human body. For instance, as of January 2024, UniProt contains just 38,800 immunoglobulin sequences, far short of the billions of antibodies the human immune system can produce. As a result, relying on such limited databases can lead to under-detection of antibodies, particularly those associated with specific diseases. Expanding antibody databases with disease-specific sequences is crucial for improving the accuracy of MS-based proteomics in identifying antibodies relevant to human health.</p> <p>Recently, through next-generation sequencing of antibody gene repertoires, it has become possible to obtain billions of antibody sequences (in amino acid format) by annotating, translating, and numbering antibody gene sequences. These large numbers of sequences are now available in public databases such as the&nbsp;<a href="https://opig.stats.ox.ac.uk/webapps/oas/" rel="nofollow">Observed Antibody Space</a>. We hypothesize that using these theoretical antibody sequences as new databases for bottom-up proteomics could address the current lack of antibody coverage in standard databases.</p> <p>We developed a workflow to create disease-specific antibody peptide databases for bottom-up proteomics. The workflow details are available on <a href="https://github.com/trinhxt/SDU_Immunoinformatics">GitHub</a>. The database and metadata files generated by this workflow are stored in this Zenodo dataset, and they are used in DAT-DB &mdash; a web application that allows researchers to obtain FASTA files of disease-specific antibody peptides for direct use in bottom-up proteomics (see <a href="https://trinhxt.shinyapps.io/DAT-DB/">Demo version</a>).</p> <p>Each database file in this dataset is in <em>.duckdb</em> format and contains tables with 10 columns: Sequence, Filename, Patient, BSource, BType, Isotype, N_patient, N_antibody, Length_aa, and CDR3. The "<strong>Sequence</strong>" column contains tryptic peptides. "<strong>Filename</strong>" is the file where the data was collected. "<strong>Patient</strong>" refers to the patient number as listed in <em>metadata2.csv</em>. "<strong>BSource</strong>" refers to the B-cells' source, and "<strong>BType</strong>" refers to the type of B-cells. "<strong>Isotype</strong>" specifies the antibody isotype (IgA, IgD, IgE, IgG, IgM, or Bulk). "<strong>N_patient</strong>" indicates the number of patients having this peptide, and "<strong>N_antibody</strong>" specifies the number of antibodies containing this peptide. "<strong>Length_aa</strong>" indicates the number of amino acids in the peptide, while "<strong>CDR3</strong>" shows whether the peptide is found in the CDR3 region.</p> <p>The file <em>metadata1.csv</em> contains information about each database file, while <em>metadata2.csv</em> provides details about the sources of the collected antibodies.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Data accompanying "In silico analysis of the profilaggrin sequence indicates alterations in the stability, degradation route, and intracellular protein fate in filaggrin null mutation carriers" article.

<p>This research was supported by the National Science Centre, Poland, grant PRELUDIUM number 2021/41/N/NZ1/03473 to NS, National Science Centre, Poland, grant SONATA BIS number 2019/34/E/NZ6/00354 to DG-O, as well as POIR.04.04.00-00-21FA/16&ndash;00 grant, carried out within the First TEAM programme of the Foundation for Polish Science co-financed by the European Union under the European Regional Development Fund (awarded to DG-O). WP was supported by the National Science Centre, Poland, grant SONATA-BIS number 2021/42/E/NZ1/00190. SB is supported by a Wellcome Trust Senior Research Fellowship (220875/Z/20/Z).</p>

opencc-by-4.0May 2023View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Human Genome Diversity Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the HGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>2b388c1fa446ecec70e33ea0471e06f8</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>50b80ed32b1ae542c8967cc31986dd19</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>995f30b74c4bb094a674b1a994853246</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>e79e61efab4c491fa2825b7d1853df58</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Simons Genome Diversity Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the SGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <table> <tbody> <tr> <th>Public Dataset</th> <th>EGP Result File Type</th> <th>MD5</th> </tr> </tbody> <tbody> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>86b09553f80926c1c29c57000ec1a88f</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>010026d77bee81e7b8daf5836bd12da3</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>f1ea3edf4a82b42f2028467fb3544dc4</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>4872eeb792c214ad49662e98e4b14620</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 2504 Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 2504 Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 2504 Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/1000G_2504_high_coverage.sequence.index</code>. Please note that the index files are there as well. You just have to append a <code>.crai</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>dbf39d6ff0e4389b900f9d985f2e6c64</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>4d53ef60ec16f3e4b566c45fdf0fb977</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Variant Tables</td> <td>16925b546051d37cce27df8ec57ccc5e</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Copy Number</td> <td>365c1b360ea327795d981356064658a6</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 698 Related Whole-Genome Sequencing Data

<div> <p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 698 Related Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 698 Related Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp-trace.ncbi.nlm.nih.gov/1000genomes/ftp/1000G_2504_high_coverage/additional_698_related/1000G_698_related_high_coverage.sequence.index</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>322038d61b4da2e937b32410613c3532</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>36c782c12245100478903f7fa191a402</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Variant Tables</td> <td>68b2a51361ffae4e7ad9d420b8becd38</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Copy Number</td> <td>1e83c8ae132b0a7ef33b090757b29063</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div> <p>&nbsp;</p> </div> <h2>&nbsp;</h2>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Gambian Genome Variation Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GGVP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Gambian Genome Variation Project</strong>:&nbsp;Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>d21e1e91e8b4c00627171fae79a1f54d</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>b359d1068d4f84f7746d1ebde82df29a</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>ee2b93aa93d2177d92ec0f8f308b43ed</td> </tr> <tr> <td>Gambian Genome Variation Project</td> <td>Mitochondrial Genome Copy Number</td> <td>fda509ba1d2bf33fd2d6b77b92e76c03</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on GIAB Whole-Genome Sequencing Data

<div> <p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GIAB. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on GIAB</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://raw.githubusercontent.com/genome-in-a-bottle/giab_data_indexes/refs/heads/master/AshkenazimTrio/alignment.index.AJtrio_Illumina300X_wgs_novoalign_GRCh37_GRCh38_NHGRI_07282015</code></p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>5eac6ec7d36307aa401fd5441b38a506</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>1151ae74c8e515f4f39bef816bb55d6a</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Variant Tables</td> <td>b3e342fe9827df2e399f5685f84cd4dc</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Copy Number</td> <td>3c45c19f76f71b3ccaad155565ead5e4</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div> <p>&nbsp;</p> <p>&nbsp;</p> </div> <h2>&nbsp;</h2>

opencc-by-4.0Sep 2024View details →
zenodo36/100

DAS Data for the figure in the paper entitled "A Hybrid Earthquake Detection Method for Distributed Acoustic Sensing Array Data and Its Application to the 2022 Menyuan Earthquake Sequence"

<p>The DAS data can be loaded using numpy. The sampling rate is 100 Hz, and each row is a time series for that channel.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Feces - deep sequencing data

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo36/100

Feces - shallow sequencing data

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record