Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,919

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,919 results for “Cohort”

Learn how ShareScore rates datasets ↗
zenodo44/100

CINECA synthetic cohort NA Canada CHILD [CC-BY-NC-SA]

<p>The &quot;CINECA synthetic cohort NA Canada CHILD&quot; dataset is a synthetic dataset developed to provide insight into how data is structured for select common attributes in the <a href="https://childstudy.ca/">CHILD Cohort Study</a>, but not reveal any personal or identifiable information associated with cohort participants. Such synthetic datasets are valuable for software developers to be able to see specific examples of data for common attributes (i.e. a minimal metadata model of a selection of common variables usually present in cohorts). This dataset comprises 100 variables for 150 synthetic participants which have faked phenotypic data that reflects CHILD cohort data. In addition, there is genetic data based on the <a href="https://www.nature.com/articles/nature15393">1000 Genomes</a> project. This dataset was created within the context of the <a href="https://www.cineca-project.eu/">CINECA</a> project. More information about the creation of this dataset can be found in the included documentation.&nbsp;</p> <p><br> <em>Please note this preamble must be included with any distribution of this dataset:&nbsp;</em>This synthetic dataset (with cohort &ldquo;participants&rdquo; / &rdquo;subjects&rdquo; marked with FAKE) has no identifiable data and cannot be used to make any inference about CHILD cohort data or results. The purpose of this dataset is to aid development of technical implementations for cohort data discovery, harmonization, access, and federated analysis. In support of FAIRness in data sharing, this dataset is made freely available under the Creative Commons Licence (CC-BY; <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>). Please ensure this preamble is included with this dataset and that the CHILD project and the CINECA project (funding: EC H2020 grant 825775 and CIHR grant 404896) are acknowledged. If you have any questions about this dataset contact Fiona Brinkman at brinkman@sfu.ca or Erin Gill at egill@sfu.ca.</p> <p>&nbsp;</p> <p><strong>CINECA synthetic cohorts</strong></p> <ul> <li><a href="https://zenodo.org/record/4955933">CINECA synthetic cohort Africa H3ABioNet</a></li> <li><a href="https://zenodo.org/record/5082689">CINECA synthetic cohort Europe CH SIB</a></li> <li><a href="https://ega-archive.org/datasets/EGAD00001006673">CINECA synthetic cohort Europe&nbsp;UK1</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Dataset used for: Effectiveness of Acute Malnutrition Treatment at Health Center and Community Level with a Simplified, Combined Protocol in Mali: An Observational Cohort Study

<p>This dataset contains the variables used in the analysis of the body composition and&nbsp;outcomes of the Acute Malnutrition Treatment at Health Center and Community Level with a Simplified, Combined Protocol in Mali pilot study, from December 2018 to December 2021</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Multidimensional pain profiling in people living with obesity and attending weight management services: a protocol for a longitudinal cohort study

<p><strong>Please note: The final dataset will not be available until data collection has been completed in the Autumn of 2024.<br> This dataset currently contains the following:</strong></p> <p>1. Details of the study, including authorship, ethical approval, funding, registration details, an abstract for the protocol of the study and details of data being collected (both in Microsoft Word and open access .txt formats)</p> <p>2. Outline of the data in the process of collection (Microsoft Excel)</p> <p>3. Ethical approval letters from relevant Research Ethics Committees (PDF)<br> <br> &nbsp;</p> <p>&nbsp;</p> <p><strong>Project Title: </strong>Multidimensional pain profiling in people living with obesity and attending weight management services: a longitudinal cohort study</p> <p>&nbsp;</p> <p><strong>Authors:</strong> Keith M. Smart<sup>1,2</sup>, Natasha Hinwood<sup>1</sup>, Colin G. Dunlevy<sup>3</sup>, Catherine Doody<sup>1</sup>, Catherine Blake<sup>1</sup>, Brona Fullen<sup>1</sup>, Jean O&rsquo;Connell<sup>3</sup>, Carel W. Le Roux<sup>4</sup>, Clare Gilsenan<sup>5</sup>, Francis M. Finucane<sup>6,7</sup>, Gr&aacute;inne O&rsquo;Donoghue<sup>1</sup>.</p> <p>&nbsp;</p> <p><strong>Corresponding author</strong>: Natasha Hinwood</p> <p><strong>Address:</strong> UCD School of Public Health, Physiotherapy and Sport Science, University College Dublin, Dublin, Ireland</p> <p><strong>Email:</strong> <a href="mailto:natasha.hinwood@ucdconnect.ie">natasha.hinwood@ucdconnect.ie</a></p> <p><strong>Phone:</strong> +353 1 716 6511</p> <p>&nbsp;</p> <p>Full name, department, institution, city, and country of all co-authors.</p> <p><sup>1</sup>UCD School of Public Health, Physiotherapy and Sport Science, University College Dublin, Dublin, Ireland</p> <p><sup>2</sup>Physiotherapy Department, St. Vincent&rsquo;s University Hospital, Dublin, Ireland</p> <p><sup>3</sup>Weight Management Service, St Columcille&rsquo;s Hospital, Dublin, Ireland</p> <p><sup>4</sup>Diabetes Complications Research Centre, University College Dublin, Dublin, Ireland</p> <p><sup>5</sup>Physiotherapy Department, Beaumont Hospital, Dublin, Ireland</p> <p><sup>6</sup> School of Medicine, College of Nursing and Health Sciences, University of Galway</p> <p><sup>7</sup>Bariatric Medicine Service, Centre for Diabetes, Endocrinology and Metabolism, Galway University Hospitals</p> <p>&nbsp;</p> <p><strong>ORCID</strong></p> <p>1. Keith M. Smart: 0000-0002-1598-5215</p> <p>2. Natasha Hinwood: 0000-0001-9382-716X</p> <p>3. Colin G. Dunlevy:</p> <p>4. Catherine Doody:</p> <p>5. Catherine Blake: 0000-0002-0600-629X</p> <p>6. Brona Fullen: 0000-0003-4408--2063</p> <p>7. Carel W. Le Roux: 0000-0001-5521-5445</p> <p>8.&nbsp;Jean O&rsquo;Connell: 0000-0001-7241-8025</p> <p>9. Clare Gilsenan:</p> <p>10. Francis M. Finucane: 0000-0002-5374-7090</p> <p>11. Gr&aacute;inne O&rsquo;Donoghue: 0000-0002-9126-2094</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Project abstract (Protocol): </strong></p> <p><em>Introduction</em>:</p> <p>Pain is prevalent in people living with overweight and obesity. Obesity is associated with increased self-reported pain intensity and pain-related disability, reductions in physical functioning and poorer psychological well-being. People living with obesity tend to respond less well to pain treatments or management compared to people living without obesity. Mechanisms linking obesity and pain are complex and may variously include contributions from and interactions between physiological, behavioural, psychological, socio-cultural, biomechanical, and genetic factors. Our aim is to study the multidimensional pain profiles of people living with obesity, over time, in an attempt to better understand the relationship between obesity and pain.<br> &nbsp;</p> <p><em>Methods and analysis: </em></p> <p>This longitudinal observational cohort study will recruit (n=216) people living with obesity and who are newly attending three&nbsp;weight management services in Ireland. Participants will complete questionnaires that assess their multidimensional biopsychosocial pain experience at baseline and at 3, 6, 12 and 18-months post-recruitment. Quantitative analyses will characterise the multidimensional pain experiences and trajectories of the cohort as a whole and in defined sub-groups.<br> &nbsp;</p> <p><em>Ethics and dissemination: </em></p> <p>The study protocol has been approved by the Ethics and Medical Research Committee of St Vincent&rsquo;s Healthcare Group, Dublin, Ireland (Reference No.: RS21-059) and the University College Dublin Human Research Ethics Committee (Reference No.: LS-E-22-41-Hinwood-Smart). Findings will be disseminated through peer-reviewed journals, conference presentations, public and patient advocacy groups, and social media.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Epidemiology and Disease Burden of Neurocritical Disorders: A Cohort Study - Data Sharing

<p>Dataset of a Neurocritical Brazil cohort study whose summary is described below.</p> <p><strong>Abstract</strong></p> <p><strong>Objective:</strong> To describe a cohort of neurocritical patients and their differences based on primary neurological diagnoses and identify predictors of mortality and unfavorable outcome along with the disease burden of each neurological condition on intensive care unit (ICU) admission. <strong>Methods:</strong> Prospective cohort study including patients admitted to 36 ICUs in Brazil and followed up for 30 days. <strong>Results:</strong> Of 4245 patients admitted to the participating ICUs during the study period, 1194 (28.1%) were neurocritical patients and were included in the study. Neurocritical patients had a mean mortality rate 1.7 times higher than non-neurocritical patients admitted to the same ICUs (17.21% versus 10.1%, respectively). The most frequent primary neurological diagnoses on ICU admission were postoperative care of elective neurosurgery, traumatic brain injury, ischemic stroke, and encephalopathy. The estimated total disability-adjusted life-years (DALYs) were 4482.94 in the overall cohort, and the diagnosis with the highest DALYs was traumatic brain injury (1634.42). DALYs were significantly impacted by the patients&rsquo; primary neurological diagnosis, sex, age group, and number of secondary neurological injuries.&nbsp;<strong>Conclusion:</strong> We accurately described the epidemiology of neurocritical patients and estimated their overall and relative disease burden. The findings of this study are important to direct policies regarding education, prevention, and treatment of severe neurocritical diseases.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Complement activation induces excessive T cell cytotoxicity in severe COVID-19: Analysis of single cell data cohort 1 (Berlin).

<p>This repository contains the R Markdown files with the analysis of CyTOF and scRNA-seq data corresponding to cohort 1 (Berlin) analysed in Georg et al. 2021 &quot;Complement activation induces excessive T cell cytotoxicity in severe COVID-19&quot;. Additionally, here we&nbsp;include&nbsp;the necessary CyTOF data to reproduce this&nbsp;analysis.</p> <p>CyTOF data:</p> <ul> <li>The debarcoded fcs files (before batch-correction) can be found in&nbsp;<a href="https://flowrepository.org/id/FR-FCM-Z4P5">https://flowrepository.org/id/FR-FCM-Z4P5</a>. \</li> <li>Here you can find the necessary data to reproduce the analysis (cytof_analysis.Rmd, cytof_analysis.html): <ul> <li>data_norm_all.csv: single-cell protein expression data (after batch-normalization and in linear scale).</li> <li>data_Tcells_annotated.csv: single-cell protein expression of gated T cells with cluster annotation.</li> <li>phenograph_CD4_k30.csv, phenograph_CD8_k30.csv, phenograph_TCRgd_k30.csv: output from Louvain Clustering computed with PhenoGraph (<a href="https://github.com/jacoblevine/PhenoGraph">https://github.com/jacoblevine/PhenoGraph</a>) per T cell compartment.</li> <li>clusterannotation.csv: annotation for each cluster and metacluster</li> </ul> </li> </ul> <p>scRNA-seq data:</p> <ul> <li>The raw data can be found in&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE175450">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE175450</a></li> <li>Other files&nbsp;to reproduce the analysis (scRNAseq_analysis_1preprocessing.Rmd, scRNAseq_analysis_2clustering.Rmd, scRNAseq_analysis_3convalescent.Rmd): <ul> <li><a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_Sawitzki_RECAST_09_2021.xlsx">scRNAseq_Sawitzki_RECAST_09_2021.xlsx</a>: Single-cell metadata.</li> <li>scRNAseq_samples.tsv: Samples metadata.</li> <li><a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_genelist_annotation.xlsx">scRNAseq_genelist_annotation.xlsx</a>:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE175450">G</a>ene list for the annotation of T cells (Also in Mendeley, see&nbsp;Data and Code Availability).</li> <li><a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_GO_RESPONSE_TO_TYPE_I_INTERFERON.txt">scRNAseq_GO_RESPONSE_TO_TYPE_I_INTERFERON.txt</a>,&nbsp;<a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_GO_DEFENSE_RESPONSE_TO_VIRUS.txt">scRNAseq_GO_DEFENSE_RESPONSE_TO_VIRUS.txt</a>,&nbsp;,&nbsp;<a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_GO_T_CELL_MEDIATED_CYTOTOXICITY.txt">scRNAseq_GO_T_CELL_MEDIATED_CYTOTOXICITY.txt</a>: Gene lists for the signatures &ldquo;Response to Type I Interferon&rdquo; , &ldquo;Defense Response to virus&rdquo; and &ldquo;Cytotoxicity&rdquo; used for GSEA. (Also in&nbsp;Table S2).</li> <li><a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_traj18_trav10.txt">scRNAseq_traj18_trav10.txt</a>,<a href="https://zenodo.org/api/files/76286c93-628d-4251-9118-52130d4a75c6/scRNAseq_trbv25.txt">scRNAseq_trbv25.txt</a>: sequences to determine&nbsp;the proportion of TRAV10-TRAJ18-TRBV25 pairing T cell clones across all T cell clusters.</li> </ul> </li> </ul>

opencc-by-4.0Dec 2021View details →
edi44/100

White Spruce NPP: average NPP per tree by age cohort for 4 time periods between 1993 and 2008

This data set is derived from BNZ LTER inventory plots, where DBH of all trees is measured every 3 to 4 years, and increment changes in AG biomass are derived from allometric equations. Data included in this file were obtained from 1993, 1997, 2000, 2004 and 2008 inventories, generating 4 growth increments between 1997 and 2008. Data are reported by age cohort, and include NPP increments (averaged across all trees within each landscape and successional stage) for all trees of known age (determined from coring).

openOpenNov 2009View details →
zenodo40/100

Lung ultrasonography features and risk stratification in 80 patients with COVID-19: a prospective observational cohort study

<p><strong>Background</strong></p> <p>Point-of-care lung ultrasound (LUS) is a promising and pragmatic risk stratification tool in COVID-19. This study describes and compares early LUS characteristics across of range of clinical outcomes.</p> <p><strong>Method</strong></p> <p>Prospective observational study of PCR-confirmed COVID-19 patients in the emergency department (ED) of Lausanne University Hospital. A trained physician recorded LUS images using a standardized protocol. Two experts retrospectively reviewed images blinded to patient outcome. We describe and compare early LUS findings (acquired within 24hours of presentation at the ED) between patient groups based on their outcome at 7-days after inclusion:&nbsp; 1) self-resolving outpatients, 2) hospitalised and 3) intubated/death. The LUS score was used to discriminate between groups.</p> <p><strong>Findings</strong></p> <p>Between March 6 and April 3 2020, we included 80 patients (18 outpatients, 41 hospitalized and 21 intubated/dead). 73 patients (91%) had abnormal LUS (72% outpatients, 95% hospitalised and 100% intubated/death; p=0.004). The proportion of involved zones was lower in outpatients compared with other groups (median&nbsp; 30% [IQR 0-40%], 44% [33-70%] and 70% [50-88%], p&lt;0.001). Predominant abnormal patterns were bilateral and multifocal spread thickening of the pleura with pleural line irregularities (77%), confluent B lines (66%) and pathologic B lines (55%). Posterior inferior zones were more often affected. Median LUS score had a good level of discrimination between outpatients and others with area under the ROC of 0.80 (95% CI 0.66-0.95).</p> <p><strong>Interpretation</strong></p> <p>Systematic LUS is a reliable, cheap and easy-to-use triage tool for the early stratification of risk in COVID-19 patients presenting at emergency departments.</p> <p><strong>Funding</strong></p> <p>Leenaards Foundation</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Volatile organic compound analysis, a new tool in the quest for preterm birth prediction – an observational cohort study

<p>Vaginal swabs were taken in pregnancy in high risk asymptomatic women attending a preterm prevention clinic. Women in the study attended the clinical due to a history of preterm birth or midtrimester pregnancy loss, or due to a history of cervical surgery. Individualised management plans were made depending upon individual patient risk factors. During their attendance to the clinic vaginal swabs were taken for VOC analysis. Swabs were taken between 15 and 28 weeks gestation. Women consented to vaginal swabs at each of their visits to the clinic. The dataset contains GC-IMS VOC data from a G.A.S. GC-IMS and includes a Spreadsheet of demographics.</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Data for "Incidence, clinical course and risk factor for recurrent PCR positivity in discharged COVID-19 patients in Guangzhou, China: a prospective cohort study"

<p>Data for &quot;Incidence, clinical course and risk factor for recurrent PCR positivity in discharged COVID-19 patients in Guangzhou, China: a prospective cohort study&quot;</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Data for Are Changes in Alcohol Use and Personality Traits associated? A Cohort Study among Young Swiss Men

<p>These are the data and metadata for the article&nbsp;</p> <p><strong>Are Changes in Alcohol Use and Personality Traits associated? A Cohort Study among Young Swiss Men</strong></p> <p>by&nbsp;</p> <p><strong>Gerhard Gmel, Simon Marmet, Joseph Studer, and Matthias Wicki</strong></p> <p><strong>to be published in Frontiers of Psychiatry</strong></p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Identification of co-infections in a cohort of patients diagnosed with Lyme Disease

<p>Serlogy test data used in the study:&nbsp;Identification of co-infections in a cohort of patients diagnosed with Lyme Disease</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

16S rRNA Sequencing Data of Fecal Microbiota in an Italian Cohort of Patients with CDKL5 Deficiency Disorder

<h3>Summary of the study&nbsp;</h3> <p>CDKL5 deficiency disorder (CDD) is a neurodevelopmental condition characterized by global developmental delay, early-onset seizures, intellectual disability, visual and motor impairments, distinct from Rett Syndrome (RTT) due to the absence of a clear regression period. Gastrointestinal (GI) disturbances and signs of subclinical immune dysregulation are common in CDD patients, yet the underlying causes are unknown. Recent studies hint at a possible link between neurological disorders and gut microbiota, an unexplored area in CDD. In this groundbreaking study, we examined fecal microbiota in CDD patients and their healthy relatives, revealing differences in bacterial diversity and composition. We further investigated microbiota changes based on various factors, including the severity of GI issues, seizure frequency, sleep disorders, food intake type, neuro-behavioral features (assessed through the RTT Behaviour Questionnaire &ndash; RSBQ), and ambulation capacity.&nbsp;</p> <p>Our findings suggest a potential connection between CDD, microbiota, and symptom severity. This study represents the first exploration of the gut-microbiota-brain axis in CDD patients, contributing to the growing body of research on the role of gut microbiota in neurodevelopmental disorders. It opens doors to potential interventions targeting intestinal microbes to enhance the well-being of individuals with CDD.</p> <h3>Mehods</h3> <p>The Dataset represent the raw data (.fastq) obtained from the sequencing of the fecal samples from 17 Italian Patients with CDD, and 17 Healthy Relatives (i.e. siblings or mother), collected at a single time-point.</p> <p>Samples from Patients affected by CDD are called CDD, samples from Healthy Relatives are called HC-CDD (i.e. healthy controls of patients affected by CDD). For details about the sample names see the &ldquo;Explanation Table&rdquo;.</p> <p>Bacterial DNA was extracted from fecal samples using the QIAmp Powerfexal DNA Kit (Qiagen, Germany) following the manufacturer's protocol. The 16S rRNA sequencing and analysis was performed by a service offered by Zymo Research (Germany).</p> <p><em>Targeted Library Preparation</em>: The DNA samples were prepared for targeted sequencing with the Quick-16S&trade; NGS Library Prep Kit (Zymo Research). The primer sets used were Quick-16S&trade; Primer Set V3-V4 (Zymo Research). The sequencing library was prepared using an innovative library preparation process in which PCR reactions were performed in real-time PCR machines to control cycles and therefore limit PCR chimera formation. The final PCR products were quantified with qPCR fluorescence readings and pooled together based on equal molarity. The final pooled library was cleaned up with the Select-a-Size DNA Clean &amp; Concentrator&trade;, then quantified with TapeStation&reg; (Agilent Technologies, Santa Clara, CA) and Qubit&reg; (Thermo Fisher Scientific, Waltham, WA).&nbsp;&nbsp;</p> <p><em>Sequencing:</em> The final library was sequenced on Illumina&reg; MiSeq&trade; with a v3 reagent kit (600 cycles).&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Prospective with historical control, case-matched cohort study of the tertiary survey beneficial in critically severe trauma patients.

<p>Raw data, Tertiary survey record form, and STROBE checklist of "Prospective with historical control, case-matched cohort study of the tertiary survey beneficial in critically severe trauma patients" study.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Mental health, physical health, training load and subjective performance during the COVID-19 pandemic – a Swiss elite athletes' cohort study

<p>Dataset of&nbsp;Swiss elite athletes (n=203) participating in a repeated online survey evaluating mental and physical health factors, as well as training and performance related metrics. After the first survey during the first lockdown between April and May 2020, there were monthly follow-up surveys over a 6-month period.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Significant sQTLs and splicing TWAS reference panels (AMP-AD brain and EADB Belgian LCL cohorts)

<p>This dataset is part of the manuscript &quot;<em><strong>New insights into the genetic etiology of Alzheimer&rsquo;s disease and related dementias</strong></em>&quot; by Bellenguez, K&uuml;&ccedil;&uuml;kali, et al. Nature Genetics 2022.</p> <p>Publication link:&nbsp;<a href="https://www.nature.com/articles/s41588-022-01024-z">https://www.nature.com/articles/s41588-022-01024-z</a></p> <p>GitHub repository of all QTL/TWAS data shared&nbsp;for this study:&nbsp;<a href="https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS">https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS</a></p> <p>For details, please see the publication.&nbsp;For any questions, please contact Fahri K&uuml;&ccedil;&uuml;kali (<a href="mailto:fahri.kucukali@uantwerpen.vib.be">fahri.kucukali@uantwerpen.vib.be</a>) and Kristel Sleegers (<a href="mailto:Kristel.Sleegers@uantwerpen.vib.be">Kristel.Sleegers@uantwerpen.vib.be</a>).</p> <p>Significant sQTL catalogues are compressed with <em>gzip </em>and tar achieve of splicing TWAS reference panels are compressed with <em>bzip2</em>.</p> <p><strong>sQTL catalogues</strong></p> <p>The files show significant sQTL - splice junction pairs mapped in AMP-AD brain and EADB Belgian LCL cohorts. The catalogues are in hg38/GRCh38 human genome build.&nbsp;Most of the columns in the files&nbsp;are based on FastQTL output (<a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a>).</p> <p><em>sQTL file columns:</em></p> <ol> <li>variant_id - ID of the significant sQTL variant based on the dbSNPv151 rsID annotation or hg38/GRCh38 CHR_POS_REF_ALT ID if rsID not available.</li> <li>junc_id - sJunction ID assigned by Regtools/Leafcutter pipeline. For stranded datasets (ROSMAP and EADB Belgian) strand info is provided with &quot;+&quot; or &quot;-&quot; symbols, and if not stranded, &quot;?&quot; symbol is used.</li> <li>junc_distance - Genomic distance between sQTL variant and splice junction start</li> <li>ma_samples - Number of samples carrying the minor allele</li> <li>ma_count - Total count of minor alleles</li> <li>maf - Minor allele frequency</li> <li>pval_nominal - Nominal P-value&nbsp;of the association</li> <li>slope - Slope of the association&nbsp;with respect to alternative (ALT)&nbsp;allele indicated on column 14</li> <li>slope_se - Standard error of the slope</li> <li>pval_nominal_threshold - Nominal&nbsp;<em>P</em>-value significant threshold for thissJunction</li> <li>min_pval_nominal - Most significant&nbsp;<em>P</em>-value observed for this sJunction</li> <li>pval_beta -&nbsp;permutation&nbsp;<em>P</em>-value&nbsp;obtained via beta approximation and later used to calculate Storey q-values</li> <li>junction - sJunction in chr:start-end splice junction format</li> <li>cluster - The splice cluster of this sJunction</li> <li>genes - Genes overlapping with this sJunction (if any), based on GENCODEv24 (AMP-AD) and GENCODEv32 (EADB Belgian)</li> <li>GRCh38_chr_pos - Genomic position of the variant, separated by underscore</li> <li>ref_alt - Reference (REF) and alternative (ALT) allele of the variant, separated by &quot;&gt;&quot; sign. ALT is the tested (A1) allele</li> </ol> <p>Of note, we also mapped the significant eQTLs in the same datasets (please see the data availability section of the manuscript or the GitHub repository).</p> <p><strong>Splicing TWAS reference panels</strong></p> <p>Custom splicing TWAS reference panels prepared using FUSION pipeline (<a href="http://gusevlab.org/projects/fusion/">http://gusevlab.org/projects/fusion/</a>) in AMP-AD brain and EADB Belgian LCL cohorts. All data in hg38/GRCh38 genome build.&nbsp;In each directory, you will find &quot;.pos&quot;, .&quot;profile&quot;, and &quot;.profile.err&quot; files. These are explained in the FUSION website as well, but briefly these are:</p> <ol> <li><strong>.pos:</strong> This is a position file that describes the 1Mb extended splice junction start and end coordinates for each calculated weight file for splice junction phenotype. Used for scanning the variants in those coordinates for TWAS.</li> <li><strong>.profile: </strong>This informs about all prediction weights calculated, in terms of number of variants in the model, heritability information, and R2 info for each prediction model used (top1, blup, enet, bslmm, lasso; bslmm was not used therefore has NA values).</li> <li><strong>.profile.err: </strong>This summarizes the reference panel in terms of average hsq (with SD), and which model is the best performing.</li> </ol> <p>Each TWAS weight&nbsp;is provided in a .RDat file under <strong>All_Splicing_Weights</strong>, and in this data we included all calculated functional weights independent of the fact that they are heritable features or not. In our TWAS analyses, we included the heritable functional weights at a hsq&nbsp;<em>P</em>-value &le; 0.05 level.</p> <p>Please also see the expression TWAS reference panels we prepared in the same datasets (see the data availability section of the manuscript or the GitHub repository). If you need an LD reference data in hg38/GRCh38 genome build based on 1000 Genomes NFE samples (whose variant ID annotation are matching to these functional weights), suitable for running the TWAS/FUSION pipeline, please contact us.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities: discovery cohort meta data and parsed TCR repertoire data

<p>Meta data corresponding the the discovery cohort for the paper, &quot;Combining genotypes and T cell receptor distributions to infer genetic loci determining V(D)J recombination probabilities&quot;&nbsp;by Magdalena L Russell, Aisha Souquette, David M Levine, Stefan A Schattgen, E Kaitlynn Allen, Guillermina Kuan, Noah Simon, Angel Balmaseda, Aubree Gordon, Paul G Thomas, Frederick A Matsen IV, and Philip Bradley. These meta data include:&nbsp;</p> <p>(1) a file mapping the SNP data subject IDs&nbsp;to the TCR repertoire data&nbsp;subject IDs (gwas_id_mapping.tsv)<br> (2) a file including the PCAir PCs, self-reported ancestry, and genomic ancestry for each subject (all_pc_air.txt)<br> (3) a file including the PCAir variance explained by each PC (all_pc_air_variance.txt)<br> (3)&nbsp;a file including the SNP ID, chromosome, hg19 position, allele, rsid, and quality control metrics&nbsp;for each SNP in the SNP array (emerson_snp_rs_data.tsv)<br> (4) a file including IMGT genes and sequences used for parsing TCRB repertoire data (human_vj_allele_cdr3_nucseqs.tsv)<br> (5) a file including predicted TRBD2 allele genotypes for each subject (emerson_trbd2_alleles.tsv)<br> (6)&nbsp;Parsed TCRB repertoire data.&nbsp;These raw data were&nbsp;first published in Emerson et. al,&nbsp;<em>Nature Genetics&nbsp;</em>2017. (emerson_parsed_tcrb.tgz)</p> <p><strong>Corresponding discovery&nbsp;cohort raw TCR repertoire data is available here:&nbsp;</strong>https: //doi.org/10.21417/B7001Z (ImmuneACCESS database)<br> <strong>Corresponding discovery cohort SNP data is available here:</strong>&nbsp;https: //www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs001918.v1.p1 (The database of Genotypes and Phenotypes,&nbsp;accession number: phs001918)<br> <br> <strong>Software tools designed to work with these data are available here:</strong>&nbsp;https://github.com/phbradley/tcr-gwas</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Audio recordings of COVID-19 positive individuals from the prospective Predi-COVID cohort study with their fatigue status

<p>We uploaded <strong>3544 </strong>audio recordings originating from <strong>296 </strong>distinct participants with COVID-19 in the prospective <strong>Predi-COVID cohort study</strong> recruited between May 2020 and May 2021. The audios have been converted from their original format into WAV files and normalized. The audio name structure integrates the participant ID, the recording date and time of the audio recording, the type of audio (Type 1: text reading, Type2: holding the [a] vowel without breathing), the original audio format, the gender (W: women, M: Men), and the <strong>fatigue status</strong> of the participant&nbsp;(1: Fatigue, 0: No fatigue) as such:</p> <p>Predi-COVID_{participant ID}{recording date and time}{type of audio}{original format}{gender}{fatigue status}.wav</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).

<p>*******************************************************************<br> MetaDrugs workflow<br> *******************************************************************</p> <p>Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).</p> <p>For questions and requests, please contact:<br> Sofia K. Forslund (sofia.forslund@mdc-berlin.de)<br> and Till Birkner&nbsp; (till.birkner@mdc-berlin.de)</p> <p>*******************************************************************<br> Contents:<br> -------------------------------------------------------------------<br> Data files:<br> metadata.tar.gz - archived cohort metadata files*<br> input_features.tar.gz - archived preprocessed serum and urine metabolome and gut microbiome features<br> output_complete.tar.gz - archived example analysis output files for each of the input feature file<br> output_rerun.tar.gz - archived empty directory for generating test output files as described in this document<br> <br> *Please note: Due to conflicts with Danish Data Protection laws, metadata from the Danish subset of the cohort were removed in this repository. Please reach out for a potential case-by-case access request for access to the complete set of metadata.<br> -------------------------------------------------------------------<br> Text files:<br> archived in feature_names.tar.gz:<br> atcs_names - full names for atcs drug compounds<br> contrast_names - full names for disease comparison groups<br> file_names - brief description of the files in input_features folder<br> gmm_names - full names of GMM modules<br> kegg_names - full names of KEGG modules<br> ko_names - full names of KO modules<br> metadata_names - full names of metadata features<br> mOTU_names - species names for metagenomics data<br> taxon_names - taxon names for metagenomics data<br> -------------------------------------------------------------------<br> Scripts:<br> -------------------------------------------------------------------<br> runFrame.r - main wrapper script envoking the analysis pipeline<br> -------------------------------------------------------------------<br> runFrame_rel_comb.r - script calculating drug combination effects<br> runFrame_rel.r - script calculating dosage effects<br> testCombPresenceSeparate.r - testing of significant drug combination effects beyond single drug effects<br> testDosagePresenceSeparate.pl - testing of significant drug dosage effects beyond single drug effects<br> testDosagePresenceSeparateNegative.pl - testing of unique drug dosage effects beyond single drug effects<br> -------------------------------------------------------------------<br> prettifyResults_uncollapsed.pl - wrapper scripts to create and format a single analysis output file<br> makeTables.r - wrapper script to make excel tables with analysis results<br> -------------------------------------------------------------------<br> Example output file:<br> -------------------------------------------------------------------<br> output_all_formatted_noc_uncollapsed_complete.tsv - contains all disease-drug-host-microbiome feature analysis results in one place.<br> *******************************************************************</p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

Molecular Signatures of Tumour and its Microenvironment for Precise Quantitative Diagnosis of Oral Squamous Cell Carcinoma: An Interna-tional Multi-cohort Diagnostic Validation Study

<p><strong>Supplementary Materials: </strong>The following supporting information can be downloaded at: www.mdpi.com/xxx/s1, <strong>Table ST1</strong> &ndash; qMIDS<sup>V2 </sup>Gene panel primer sequences; <strong>Figure S1</strong> &ndash; qMIDS<sup>V1</sup> vs qMIDS<sup>V2</sup> 384-well assay format and protocols; <strong>Figure S2.</strong> Individual target gene expression pattern in 1761 samples; <strong>Figure S3.</strong> Various statistical methods used for gene selection analysis on 1761 clinical samples; <strong>Figure S4. </strong>Diagnostic performance comparison between qMIDS<sup>V2</sup> vs qMIDS<sup>V2* </sup>(with 4 less effective genes removed from the panel of 14 target genes of qMIDS<sup>V2</sup>); <strong>Figure S5</strong>. Effect of removing individual genes from the 14-target gene panel qMIDS<sup>V2</sup> (qV2) on diagnostic test performance based on the UK patient cohort data.</p>

opencc-by-4.0Feb 2022View details →
dryad40/100

Identification of infectious agents in early marine Chinook and Coho salmon associated with cohort survival

<p>Recent decades have seen an increased appreciation for the role infectious diseases can play in mass mortality events across a diversity of marine taxa. At the same time many Pacific salmon populations have declined in abundance as a result of reduced marine survival. However, few studies have explicitly considered the potential role pathogens could play in these declines. Using a multi-year dataset spanning 59 pathogen taxa in Chinook and Coho salmon sampled along the British Columbia coast, we carried out an exploratory analysis to quantify evidence for associations between pathogen prevalence and cohort survival, and between pathogen load and body condition. While a variety of pathogens had moderate to strong negative correlations with body condition or survival for one host species in one season, we found that <em>Tenacibaculum maritimum</em> and Piscine orthoreovirus had consistently negative associations with body condition in both host species and seasons, and were negatively associated with survival for Chinook salmon collected in the fall and winter. Our analyses, which offer the most comprehensive examination of associations between pathogen prevalence and Pacific salmon survival to date, suggest that pathogens in Pacific salmon warrant further attention, especially those whose distribution and abundance may be influenced by anthropogenic stressors.</p>

opencc-zeroDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record