Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,868

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,868 results for “variant”

Learn how ShareScore rates datasets ↗
zenodo48/100

PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets

<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>

opencc-by-4.0Aug 2020View details →
zenodo48/100

Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)

<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). &nbsp;Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>&nbsp;Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>

opencc-by-4.0Oct 2024View details →
zenodo48/100

TAILVAR (Terminal codon Analysis and Improved prediction of Lengthened VARiants)

<p>This dataset includes relevant files for developing the TAILVAR score designed to assess the functional impact of <strong>stop-loss variants</strong> occurring at stop codons (TAA, TGA, TAG). <strong>TAILVAR</strong>&nbsp;is built using a Random Forest model that predicts the pathogenicity of&nbsp;<strong>stop-loss variants</strong>. By integrating a combination of in-silico prediction scores, transcript features, and protein context information,&nbsp;<strong>TAILVAR</strong> provides a score ranging from 0 to 1, indicating the probability of a variant being pathogenic.</p> <p>For more information, please visit&nbsp;<a href="https://github.com/dr-yoon/TAILVAR">https://github.com/dr-yoon/TAILVAR</a></p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Performance Criteria and Example Parameter Sets Comparing Different Variants of the Ensemble Kalman Filter as Applied to Volcanology

<p>This dataset contains the results of various Ensemble Kalman Filter (EnKF) inversions in which synthetic GNSS and InSAR observations from an inflating magma system are assimilated into numerical models of rock deformation around a pressurized ellipsoidal magma reservoir. Each inversion uses a different variant of the EnKF, with changes to workflow meta-parameters such as the number of ensemble members or the particular update algorithm used. In particular, each filter variant is evaluated by comparing the final output model to the original synthetic model. The specific performance criteria used include (1) the root mean square error (RMSE) between the model predictions and the assimilated observations, as well as normalized misfit terms measuring the filter&#39;s ability to resolve (2) reservoir wall tensile stress, (3) easily-observable unique parameters such as reservoir position and aspect ratio, and (4) difficult-to-derive non-unique parameters such as the specific size and internal pressure of the reservoir. The assimilated data include two different scenarios, one in which inflation is caused by pressurization and another in which it is driven by a lateral reservoir expansion. Both datasets are tested with each EnKF variant. Finally, we include example matrices from within an EnKF update step to demonstrate inter-parameter correlations that develop during the assimilation and how they can be mitigated through randomization.</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants

<p>This dataset&nbsp;consists of the reference data files, metadata and processed results files for the paper &quot;Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants,&quot; which&nbsp;investigates clonality in normal human dermal fibroblast cell populations in 32 cell lines from distinct donors, using bulk whole-exome sequencing and single-cell RNA-sequencing data.</p> <p>This dataset contains everything required to reproduce the results presented in the paper from&nbsp;processed data and results of our data processing workflows. Our analyses can be reproduced using the <a href="https://github.com/davismcc/fibroblast-clonality">source code</a>&nbsp;and instructions available at our <a href="https://davismcc.github.io/fibroblast-clonality/">project website</a>.</p> <p>The <em>entire</em> analysis workflow from raw data to final results is also reproducible but&nbsp;is substantially more complicated and computationally intensive.&nbsp;It also requires large datasets to be obtained from other repositories. Specifically, single-cell RNA-seq data have been deposited in the ArrayExpress database at EMBL-EBI under accession number E-MTAB-7167. Whole-exome sequencing data is available through the HipSci portal (www.hipsci.org). Combined with the dataset in this repository and following the instructions on the project website, it is possible to run our entire analysis pipeline.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Variant Data from Pooled Sequencing of Hybrid Kiwifruit

<p>Variant data from pooled sequencing of hybrid <em>Actinidia</em> families segregating for fruit size and Vitamin C Content.</p>

opencc-by-4.0Sep 2018View details →
zenodo48/100

Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'

<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

GTEx analysis for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial gland cells

<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

cGTEx_dataset:A multi-tissue atlas of regulatory variants in cattle

<p>The files are raw data of the cGTEX dataset used in the publication&nbsp;<strong>https://doi.org/10.1038/s41588-022-01153-5</strong>. For details, please read the Methods section.&nbsp;</p> <p>1. cGTEx_meta_data_8646sample.xlsx</p> <p>Metadata consists of sample names with their sample accession, including&nbsp;information such as data size, cleaned reads, mapping rate, and age. The data is extracted from&nbsp;SRA (<a href="https://www.ncbi.nlm.nih.gov/sra">https://www.ncbi.nlm.nih.gov/sra/</a>) and BIGD (<a href="https://bigd.big.ac.cn/bioproject/">https://bigd.big.ac.cn/bioproject/</a>) ( samples starting with CRS)</p> <p>2.&nbsp;cGTEx_count_8646sample_27607gene.txt.gz</p> <p>Data consist of raw RNA-seq read count of 27607 genes (column names as Ensembl gene id )of 8646&nbsp;samples (as row&nbsp;names)&nbsp;</p> <p>3.&nbsp;cGTEx_TPM_8646sample_27607gene.txt.gz</p> <p>Data consist of TPM values of 27607 genes (column names as Ensembl gene id) in&nbsp; samples (8646 samples as row&nbsp;names)</p> <p>4.&nbsp;cGTEx_imputed_vcf.tar.gz</p> <p>Imputed genotypes&nbsp;(SNP) of 7297 RNA-seq samples in 29 autosomes.</p> <p>5.&nbsp;cGTEx_exon_junction_8646sample.tar.gz</p> <p>Exon junction files of 8646 files&nbsp;</p> <p>Note: Small discrepancies in some sample&nbsp;names or the absence of headers in some data sets compared to https://cgtex.roslin.ed.ac.uk/&nbsp;are sorted out in this upload.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Transcriptome analysis of the effect of over-expressing H2A.J mutants in proliferating WI38 fibroblasts for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences

<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>H2A.J differs from canonical H2A only by a valine at position 11 instead of alanine, and the 7 C-terminal amino acids containing a potential minimal phosphorylation site SQ for DNA-damage response kinases. To test the functional importance of these H2A.J-specific sequences, we mutated Val-11 to Ala as is found in all canonical H2A sequences, and we mutated Ser-123 to either Glu to mimic a phospho-serine residue or to Ala to prevent phosphorylation. We also substituted the C-terminus of H2A.J with the C-terminus of H2A. These mutants, WT-H2A.J and canonical H2A-type1 were ectopically expressed in proliferating fibroblasts, and their microarray transcriptomes were compared to that of proliferating and senescent fibroblasts without ectopic histone expression. Genome-wide transcriptome analysis indicated that senescent fibroblasts clustered distinctly from proliferating fibroblasts, and proliferating fibroblasts expressing the H2A.J-V11A and H2A.J-S123E mutants clustered distinctly from fibroblasts expressing the other H2A.J mutants, WT-H2A.J, and H2A. Hallmark gene set enrichment analysis of the transcriptomes of fibroblasts expressing H2A.J-V11A or H2A.J-S123E versus control proliferating fibroblasts indicated that they showed the same highly significant enrichment for the Epithelial-Mesenchyme Transition, TNF-Alpha Signaling Via NF-kB, and Inflammatory Response gene sets. Notable inflammatory genes including IL1A, IL1B, IL6, CXCL8, and CCL2 are contained in these gene sets and are often induced in senescence as part of the senescence-associated secretory phenotype. Heat maps showed that the H2A.J-V11A and H2A.J-S123E mutants were particularly apt at activating the expression of these inflammatory genes in proliferating fibroblasts</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

DATA ANALYSIS - SARS-COV-2 ( Del69-70 VARIANT ) – NEW UK MUTANTS

<p>The data for S - genome sequence analysis known as Del69-70 is under variant of concern ( VOC ) . It is also termed as variant of investigation ( VUI ) . The data for VUI is statistically analysed by datewise and regionwise . The software used for data analysis is CURVE FINDER V.1.4 . The reproducibility of correlation and standard error is reported here for analysis of scattered data an attempt to study the Rational Fit and Harris Fit .</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

Structural variant discovery and genotyping in next-generation sequencing data

<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels &gt;50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>

opencc-by-4.0Oct 2016View details →
zenodo44/100

List of human genes and their probabilities of being intolerant to heterozygous protein truncating variants.

<p>Recalculation of the Supplementary Table 2 (doi:10.1371/journal.pcbi.1004647.s002) of the journal article "The Characteristics of Heterozygous Protein Truncating Variants in the Human Genome" by Bartha and Rausell published in PLoS Computational Biology (http://dx.doi.org/10.1371/journal.pcbi.1004647). Probabilities in this dataset were computed using human variation data from the Exome Aggregation Consortium (http://exac.broadinstitute.org/).</p> <p>Methods described in that article is relevant for this dataset. All author and affiliation information in that article is relevant for this dataset.</p> <p>Credit for the original human variation data is for the Exome Aggregation Consortium (http://exac.broadinstitute.org/, doi:10.1038/nature19057).</p>

opencc-by-4.0Jan 2017View details →
zenodo44/100

Characterization of cefiderocol resistant spontaneous mutant variants of Klebsiella pneumoniae producing NDM-5 with single mutation in cirA

<p>Cefiderocol (CFDC) is a siderophore-cephalosporin antibiotic designed to combat highly resistant Gram-negative bacterial infections. Its mechanism involves a strong affinity for iron and active transport into bacterial cells, providing an alternative against strains resistant to common antibiotics. However, the emergence of CFDC resistance in Klebsiella is a growing concern. Recent reports highlight increasing CFDC resistance in K. pneumoniae, particularly associated with mutations in the cirA gene, responsible for encoding a siderophore receptor. Co-localization of blaNDM-like gene and cirA mutations correlates with higher CFDC resistance. The study focuses on a carbapenem-resistant K. pneumoniae strain (Kp-1) with carbapenemases blaNDM-5 and blaOXA-181, recovered from a post-surgery patient. The strain exhibited resistance to all tested antibiotics but susceptibility to CFDC. Heteroresistant populations with the halo on inhibition of CFDC were observed. Genomic analysis identified a novel mutation (W123*) in the cirA gene associated with CFDC resistance. Additionally, increased blaNDM-5 expression in Kp-1 IHC (intra-halo colony) compared to Kp-1 was noted. The coexistence of blaNDM-like and cirA variants, along with high blaNDM-5 expression, explains the observed 21-fold increase in Minimum Inhibition Concentration (MIC) in Kp-1 IHC. The study contributes to understanding the molecular mechanisms driving the emergence of cefiderocol resistance, emphasizing the significance of coexisting mutations in cirA and blaNDM-like genes.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Atrophy Pattern Maps of Frontotemporal Dementia variants (bvFTD, svPPA, pnfaPPA)

<p>The files contain voxel-wise t-statistics maps contrasting deformation based morphometry (DBM) measurements of frontotemporal dementia (FTD) patients, divided according to subtype diagnosis, against matched normal controls.</p> <p>BV = behavioral-variant Frontotemporal Dementia</p> <p>SV = semantic-variant Primary Progressive Aphasia</p> <p>PNFA = non-fluent-variant Primary Progressive Aphasia</p> <p>&nbsp;</p> <p>FTD map is based on NIFD data, available at: https://memory.ucsf.edu/research/studies/nifd</p> <p>For more information regarding the participants and method details, see:</p> <p>Dadar, Mahsa, et al. "White matter hyperintensities are associated with grey matter atrophy and cognitive decline in Alzheimer's disease and frontotemporal dementia."&nbsp;<em>Neurobiology of aging</em> 111 (2022): 54-63.</p> <p>Metz, Amelie, et al. "Brain Atrophy and White Matter Hyperintensities in Frontotemporal Dementia Variants and their Impact on Cognition" [Conference presentation] CCNA 2024 Partners Forum and Science Days (2024, March 19-21).</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Growth of wild cucumber (Echinocystis lobata) in three variants of support (5, 20, 50 cm) and shoot geometry

<p><span>The experiments included three variants of support: 5 cm, 20 cm, and 50 cm. </span><span>The data was collected in the years 2020-2023. After developing their first pair of mature leaves, when the plants were about 20-35 cm in length, the plants were photographed. The shutter was released every 15 minutes. The resulting images were combined into a 10 fps video. The following were used for the analysis: 13 recordings of the 5 cm variant and 12 recordings of the 20 and 50 cm variants.</span></p> <p><span>Next, an analysis was performed based on the resulting video using Tracker</span><span> for kinetic analysis of video objects. The recordings were used to measure the plant growth parameters. The tape measure and point mass tools were used to determine the length of the shoots and to change the position of the apex relative to the areas of the X- and Y-axis photographed over time. To compensate for this, a trend line was drawn (polynomial of the 2nd degree, due to the very good fit, R<sup>2</sup> = 0,975&ndash;0,995) and new plant length parameters were calculated.&nbsp;</span></p> <p><span>To enhance the research, cross-sections of the shoots of ten plants were scanned. Samples were taken every 5 cm from the base of the plant. Using ImageJ</span><span> the following parameters were measured: shoot cross-sectional area, tissue area, perimeter, and circularity.&nbsp;</span></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Analysis of variant-dependent m6A modifications within the Human genome

<p>Interactive and machine-readable results produced by the&nbsp;<a href="https://github.com/cumbof/m6Ad-SNVs" target="_blank" rel="noopener">m6Ad-SNVs</a> tool to asses if m6A-distal SNVs affect DRACH site accessibility, specifically by evaluating the alteration of base-pairing of nucleotides within segments of the DRACH motif.</p> <p>These results contain the predicted m6Ad-SNV candidates with the length of the reference and m6Ad-SNV-containing alternate sequences limited to 250 base pairs. This constraint has been applied to maintain the reliability of the results predicted by RNAFold (<a href="https://www.tbi.univie.ac.at/RNA/">ViennaRNA</a> package). The sequence composition contains up to 100 base pairs from 3'UTRs, with the remaining base pairs limited to the last two exons.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.

openmit-licenseApr 2024View details →
zenodo44/100

Hepatocystis alignments and variant calls

<p>This contains BAM files of <em>Hepatocystis</em> reads identified in <em>Papio</em> and <em>Chlorocebus</em> samples mapped to the <em>Hepatocystis</em> reference genome as well as major and minor allele calls for cHEP and pHEP called with ANGSD, both individually and jointly. Nucleotide alignments have been added in the latest version.</p>

opencc-by-sa-4.0Jun 2024View details →
zenodo44/100

Dataset variants used in "Task-Driven Knowledge Graph Filtering Improves Prioritizing Drugs for Repurposing"

<p>This file contains all datasets and variants thereof used in the linked paper. We do not take credit for constructing the datasets, which has been done by the respective original authors (<a href="https://github.com/hetio/hetionet">https://github.com/hetio/hetionet</a>,&nbsp;<a href="https://github.com/gnn4dr/DRKG">https://github.com/gnn4dr/DRKG</a>). For our work we produced modified versions (called &quot;subset&quot; in the file) by applying our metapath based filtering approach. For validation purposed we also constructed ablation versions where one specific type of entities is missing (i.e. &quot;nogene&quot;, &quot;noside&quot;, etc).</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record