Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
152
datasets available to search
ShareScore release 0.9.0
Dataset results
152 results for “Rare Diseases”
Rare Disease analysis in Mondo
<p>To answer the question of 'How many rare diseases are there?' we analyzed terms in Mondo to get a total count of Rare Diseases as defined in Mondo Disease Ontology (Mondo).</p> <p> </p> <p>Methods</p> <p>This analysis was performed on the <a href="http://purl.obolibrary.org/obo/mondo/releases/2019-09-30/mondo.json">Mondo 2019-09-30 release</a>.</p> <p><strong>1. Get all 'Disease' terms from Mondo</strong></p> <p>First we get all the terms in Mondo that are a descendants of <code>MONDO:0000001 'Disease'</code>.</p> <p>There are <code>21633</code> Mondo disease terms.</p> <p><strong>2. Filter terms that are descendants of 'disease susceptibility'</strong></p> <p>We then filter out terms that are descendants of <code>MONDO:0042489 'disease susceptibility'</code>, to avoid counting ambiguous terms that are related to disease susceptibility and not the actual disease itself.</p> <p>This gives us a list of <code>21563</code> Mondo rare disease terms.</p> <p><strong>3. Identify terms that are 'rare'</strong></p> <p>Any disease term in Mondo is considered rare if the term, or its ancestor, has modifier <code>MONDO:0021136 'Rare'</code> in the ontology.</p> <p>There are <code>12914</code> Mondo rare disease terms.</p> <p><strong>4. Consider terms in 'gard_rare' subset</strong></p> <p>There are <code>3176</code> Mondo disease terms that are in <code>gard_rare</code> subset which contains Mondo terms that are yet to be treated as 'rare'.</p> <p>We add these terms to our set of Mondo rare disease terms.</p> <p>This increases the Mondo rare disease term count to <code>13866</code>.</p> <p>But for this analysis, we are interested in terms that are both rare and are leaf nodes in the ontology.</p> <p>After considering only leaf nodes, we get <code>10394</code> as the final count of Mondo rare disease terms.</p> <p> </p> <p>Results</p> <p>all-mondo-disease-terms.tsv: As part of our analysis, we generated a TSV containing 21633 Mondo disease terms, each with annotations that signifies whether the term is a rare disease term and whether that term is a leaf node in the ontology.</p>
Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"
<p>These are the burden test results (summary statistics) for the publication:</p> <p>"Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer’s Disease",</p> <p>Nature Genetics, 2022.</p> <p> </p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on 6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL>=75, LOF+REVEL>=50, LOF+REVEL>=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p> </p>
Interviews for New Business Models for Pharmaceutical Innovation and Access to Medicines - Rare Diseases
<p>These supplementary materials represent the partial dataset in the form of semi-structured interviews, collected and analyzed in the research article "Alternative innovation models of pharmaceutical development for rare disease drugs: how (and) do they work?: A qualitative study". This article is one of the outcomes of the "New Business Models for Pharmaceutical Innovation and Global Access to Medicines" research project, conducted at the Global Health Center, within the Geneva Graduate Institute. The dataset contains 10/11 interviews collected and used in this article, which are published with the informed consent of the interviewees.</p> <p>Details about the research project can be found at: <a href="https://www.graduateinstitute.ch/NBM">https://www.graduateinstitute.ch/NBM</a></p>
Rare Diseases hand-annotated news articles and research articles
<p>This dataset was produced in 2023 from the data collected throughout 2022 from MEDLINE (scientific articles) and from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p><p>The data is distributed across 16 diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p><p>The available data comes in 2 kinds and file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br>JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p><p>The CSV files with name starting in "f1_", "pr_", "re_" are the results of the F1/Precision/Recall evaluation for each of the cases.</p><p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdolšek Draksler (domain expert) </p>
Systematic analysis of disease-linked rare germline variants reveals new classes of cancer predisposing genes
<ul> <li>GEMs_Liver-HCC: 312 cancer patient-specific genome-scale metabolic models (GEMs) for Liver-HCC reconstructed using the RNA-seq data from PCAWG-TCGA Liver-HCC samples and generic human GEM 'Recon 2M.2'</li> <li>GEMs_Lung-SCC: 493 cancer patient-specific GEMs for Lung-SCC reconstructed using the RNA-Seq data from PCAWG-TCGA Lung-SCC samples and generic human GEM 'Recon 2M.2'</li> </ul> <p>All the patient-specific GEMs were generated using a previously developed method (i.e., tINIT algorithm with a rank-based weight function), which is available at <a href="https://bitbucket.org/kaistmbel/recon-manager">https://bitbucket.org/kaistmbel/recon-manager</a>.</p>
Tissue-aware interpretation of genetic variants advances the etiology of rare diseases
<p>Pathogenic variants underlying Mendelian diseases often disrupt the normal physiology of a<br>few tissues and organs. However, variant effect prediction tools that aim to identify<br>pathogenic variants are typically oblivious to tissue contexts. Here we report a machine-<br>learning framework, denoted ‘Tissue Risk Assessment of Causality by Expression for<br>variants’ (TRACEvar, https://netbio.bgu.ac.il/TRACEvar/), that offers two advancements.<br>First, TRACEvar predicts pathogenic variants that disrupt the normal physiology of specific<br>tissues. This was achieved by creating 14 tissue-specific models that were trained on over<br>14,000 variants and combined 84 attributes of genetic variants with 495 attributes derived<br>from tissue omics. TRACEvar outperformed 10 well-established and tissue-oblivious variant<br>effect prediction tools. Second, the resulting models are interpretable, thereby illuminating<br>variants' mode-of-action. Application of TRACEvar to variants of 52 rare-disease patients<br>highlighted pathogenicity mechanisms and relevant disease processes. Lastly, interpretation<br>of large-scale models revealed that top-ranking determinants of pathogenicity included<br>attributes of disease-affected tissues, particularly cellular process activities. Hence, tissue<br>contexts and interpretable machine-learning models can greatly enhance the etiology of rare<br>diseases.</p> <p>Article link: https://www.embopress.org/doi/full/10.1038/s44320-024-00061-6</p> <p> </p>
Classification of Text Data on Rare Diseases
<div>A dataset with text and labels for 3 categories: </div> <div> <div> </div> <div>- Rare Diseases</div> <div>- Non-Rare Diseases</div> <div>- Other</div> </div> <div> </div> <div>It is a subset of abstracts obtained from PubMed and sorted into the 3 classes on the basis of their MeSH terms.</div> <div> </div> <div>The dataset is provided for demonstration and methodology validation purposes. The original PubMed data was randomly under-sampled. </div> <p>The dataset consists of 3 files in the Tab-Separated Values (TSV) format, corresponding to the 3 splits used in the article:</p> <p>Rei L, Pita Costa J, Zdolšek Draksler T. Automatic Classification and Visualization of Text Data on Rare Diseases. _Journal of Personalized Medicine_. 2024; 14(5):545. https://doi.org/10.3390/jpm14050545</p>
Rare disease resources
<p>A repository listing rare-disease resources.</p>
Wikidata Dump Rare diseases
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/1005">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Rare Diseases hand-annotated news articles: Angelman, De Lange, Fragile X, Kleefstra
<p>This dataset was produced in 2023 from the data collected throughout 2023 from Event Registry (news) for the development of the Rare Diseases Mining project (https://idefine-europe.org/medline)</p> <p>The data is distributed across 4 specific diseases supporting the research paper "Automatic text classification and interactive data visualization of published scientific and news articles on Rare Diseases"</p> <p>The available data comes in the file formats:<br>CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings</p> <p>This work was prepared by Joao Pita Costa (researcher) and curated by Tanja Zdolšek Draksler (domain expert) </p>
Formalized information from the study protocols of three decentralized studies on rare diseases within the CORD project
<p>The file contains formalized information from the study protocols of decentralized studies on rare diseases within the CORD project. This includes the diagnoses that are coded with ICD-10-GM. The following diseases or use cases are described:</p><ul><li>Cystic Fibrosis and Pregnancy or Delivery</li><li>Phenylketonuria, Comorbidities and Pregnancy or Delivery</li><li>Kawasaki and Pediatric Inflammatory Multisystem Syndrome (PIMS)</li></ul>
Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.
<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>
Interactions of pharmaceutical companies with world countries, cancers and rare diseases from Wikipedia network analysis
<p>Using the English Wikipedia network of more than 5 million articles we analyze interactions and interlinks between the 34 largest pharmaceutical companies, 195 world countries, 47 rare renal diseases and 37 types of cancer. The recently developed algorithm using a reduced Google matrix (REGOMAX) allows us to take account both of direct Markov transitions between these articles and also of indirect transitions generated by the pathways between them<br> via the global Wikipedia network. This approach therefore provides a compact description of interactions between these articles that allows us to determine the friendship networks between them, as well as the PageRank sensitivity<br> of countries to pharmaceutical companies and rare renal diseases. We also show that the top pharmaceutical companies in terms of their Wikipedia PageRankvare not those with the highest market capitalization.</p>
Network analysis reveals rare disease signatures across multiple levels of biological organization - Co-expression dataset
<p>The GTEx-derived co-expression data in 38 tissues generated in Buphamalai et.al., Network analysis reveals rare disease signatures across multiple levels of biological organization, Nature Communications 2021. Please see the publication's Methods section for details.</p>
DOCUMENTED HUMAN OSTEOLOGICAL COLLECTIONS AS BIOBANKS: RELEVANCE FOR RARE DISEASES IDENTIFICATION IN THE PAST
<p><em><strong>Presented at: 23rd Paleopathology Association European Meeting, Vilnius, Lituânia, 25-29 Agosto. Paleopathology Association European (Vilnius, Lituânia)</strong></em></p> <p>Disease identification in paleopathology relies on the exercise of differential diagnosis, and interpretation. Only a few diseases leave macroscopic pathognomonic traits in bone, and even in cases where microscopic, biochemical and biomolecular analyses are used, diagnosis is invariably inconclusive. Additionally, bone response to a variety of etiologies tends to be homogenous, with mosaic pattern(s) of bone formation and destruction. Therefore, access to pathological cases from human remains of Documented Human Osteological Collections (DHOC) is an exceptional approach. The access to biographical data of the individuals incorporated into the DHOC includes the cause of death, ancestry, sex, age, clinical data and other information akin to clinical data allowing for the possibility of hypothesis-driven research in which bones changes correlate with causes of death - hence providing tested and informed differential diagnosis. In this sense, DHOC may be viewed as a biobank equivalent, i.e. biorepository that stores biological samples for research in the identification of bone changes related to diseases associated with clinical and personal data. This paper will explore known cases of diseases’ diagnoses, such as lepra, neoplasias, tuberculosis, syphilis, and diffuse idiopathic skeletal hyperostosis that have used DHOC as diagnostic testing grounds, to explore bone changes and methodological advancements. The paper also introduces the idea of DHOC as biobanks dedicated to the study of rare diseases, as rarely reposted diseases, in paleopathology. </p> <p><strong>Keywords: </strong>Health, biorepository, biobanks, DHOC, differential diagnosis</p>
FACE for Children With Rare Diseases
ClinicalTrials.gov study NCT04855734. IPD Sharing: NO. Countries: 1. Publications: 6.
Ready to Sail 2: A Pilot Study of Sail-Assisted Telerehabilitation in Rare Skeletal Diseases
ClinicalTrials.gov study NCT07102875. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Institutional Registry of Rare Diseases
ClinicalTrials.gov study NCT06573723. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Setmelanotide in Pediatric Participants With Rare Genetic Diseases of Obesity
ClinicalTrials.gov study NCT04966741. IPD Sharing: NO. Countries: 4. Publications: 1.
Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) - code and data sets
<p>This repository contains the scripts for the paper in revision to the AAPS J: Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) </p> <p>Authors: Niels Hendrickx, MSc, France Mentré, MD, PhD, Andreas Traschütz, MD, PhD, Cynthia Gagnon, PhD, Rebecca Schüle, MD, ARCA Study Group, EVIDENCE-RND consortium, Matthis Synofzik, MD, Emmanuelle Comets, PhD</p> <p>A simulated dataset (<strong>simulated_arsacs.csv</strong>) has been included in the repository to make the code executable as a standalone. Four main scripts have been provided in addition with the present Readme describing the files. The repository also includes 3 R objects and 2 folders which will be overwritten when the scripts are run, and are included as examples of the expected outputs. The main scripts are:</p> <p>- <strong>Script_imputation_selection.R</strong>: runs the covariate selection method. It uses a simulated dataset provided in the depot. The multiple imputation model is hardcoded as an input to the mice package to generate 10 imputed datasets, saved in current_directory/imputed_data_sets/df_arsacs_mi_i.csv. The script then runs the covariate selection method. The script prints out the list of selected covariates and returns a saemixObject containing the fit of the selected covariate model.<br> After the script executes, a list will be saved with the name of the selected covariates in the current directory (an example is included under the name "cov_matrix_model.RData" in the repository), the output of the selection, containing the whole history of runs will be saved under "final_covariate_model.RData", the list of selected covariate names will be saved under "list_covariates.RData".</p> <p>- <strong>source_mi.R</strong>: contains the functions used by Script_imputation_selection.R</p> <p>- <strong>script_bootstrap_indfit.R</strong>: This script loads "cov_matrix_model.RData" containing the matrix of covariate effects (used by saemix) and "list_covariates.RData", the list of covariates included, fits the model on the imputed data sets and computes its bootstrap distribution for each imputed data set (in the script, using only 20 samples for computation time, saved in current_directory/bootstrap/boot.arsacs.case.mi.i). It then computes the mean parameter and relative standard error of each parameter. It then computes the conditional distribution of each patient in each bootstrap samples and returns a data frame of individual predictions. The script will then plot 4 indivudal predictions. </p> <p>-<strong> source_bootstrap.R</strong>: contains the functions used by script_bootstrap_indfit.R</p> <p>Both scripts need the saemix package to run, which we haven’t included in the repository as it is freely available on the CRAN (https://cran.r-project.org/web/packages/saemix/index.html). Additional libraries we make use of in the code (MICE, tidyverse, ggplot2) also need to be installed prior to execution. <br>The R code provided can be further customised to be adapted to different scenarios.</p> <p>For the code to run, it is preferable to unzip the whole folder and set the working directory to the source file location as the script uses the "bootstrap" and "imputed_data_sets" sub-folders</p> <p>To execute this code, assuming the required libraries are available in the local R installation, please open an R session and run:<br>source("Script_imputation_selection.R") # for the covariate selection method (runtime: 3h on a i7-8565U laptop)<br>source("script_bootstrap_indfit.R") # to obtain individual trajectories (runtime: 1h on a i7-8565U laptop)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.