Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,868

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,868 results for “variant”

Learn how ShareScore rates datasets ↗
zenodo44/100

Embeddings from protein language models predict conservation and variant effects

<p>For this work, we used protein language model representations (embeddings) to predict sequence conservation without multiple sequence alignments (MSAs). Embeddings alone predicted residue conservation almost as accurately from single sequences as ConSeq using MSAs (two-state Matthew Correlation Coefficient &ndash; MCC - for ProtT5 embeddings of 0.596&plusmn;0.006 vs. 0.608&plusmn;0.006 for ConSeq).</p> <p><strong><em>ConSurf10k</em>- Dataset for the development of ProtT5cons:</strong> The method (ProtT5cons) predicting residue conservation used <em>ConSurf-DB </em>(Ben Chorin et al. 2020). This resource provided sequences and conservation for 89,673 proteins. For all, experimental high-resolution three-dimensional (3D) structures were available in the Protein Data Bank (PDB) (Berman et al. 2000). As standard-of-truth for the conservation prediction, we used the values from ConSurf-DB generated using HMMER (Mistry et al. 2013), CD-HIT (Fu et al. 2012), and MAFFT-LINSi (Katoh and Standley 2013) to align proteins in the PDB (Burley et al. 2019). For proteins from families with over 50 proteins in the resulting MSA, an evolutionary rate at each residue position is computed and used along with the MSA to reconstruct a phylogenetic tree. The ConSurf-DB conservation scores ranged from 1 (most variable) to 9 (most conserved). The PISCES server (Wang and Dunbrack 2003) was used to redundancy reduce the data set such that no pair of proteins had more than 25% pairwise sequence identity. We removed proteins with resolutions &gt;2.5&Aring;, those shorter than 40 residues, and those longer than 10,000 residues. The resulting data set (ConSurf10k) with 10,507 proteins (or domains) was randomly partitioned into training (9,392 sequences), cross-training/validation (555) and test (519) sets.</p> <p>Uploaded data:</p> <ul> <li>ConSuf10k_PDBid_seq_cons.fasta: fasta file with PDBid, sequence and conservation annotation</li> <li>consurf10k_test_ids.txt: txt file with id&#39;s of test set</li> <li>consurf10k_train_ids.txt: txt file with id&#39;s of train set</li> <li>consurf10k_val_ids.txt: txt file with id&#39;s of cross-validation set</li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Dataset related to article "Molecular Studies and ex vivo Complement assay on Endothelium Highlight the Genetic Complexity of Atypical Hemolytic Uremic Syndrome: The Case of a Pedigree With a Null CD46 Variant".

<p><em>The files contain&nbsp;raw data related to the article&nbsp;&quot;Molecular Studies and ex vivo Complement assay on Endothelium Highlight the Genetic Complexity of Atypical Hemolytic Uremic Syndrome: The Case of a Pedigree With a Null CD46 Variant&quot;, available from&nbsp;<a href="https://www.frontiersin.org/articles/10.3389/fmed.2020.579418/full">https://www.frontiersin.org/articles/10.3389/fmed.2020.579418/ful</a>l.</em></p> <p>File <strong>&quot;Genetic and clinical data&quot;</strong>:</p> <ul> <li>In the sheet &quot;485 aHUS patients&quot; are reported data obtained from the screening of 485 unrelated patients with aHUS including rare variants (RVs) in complement disease-associated genes (<em>CFH, CD46, CFI, C3, CFB </em>and <em>THBD</em>), the presence of <em>CFH-CFHR</em> genomic rearrangements and/or anti-FH antibodies.</li> <li>In the sheet &quot;Pedigrees with c.286+2T&gt;G&quot; are listed all pedigrees carrying the c.286+2T&gt;G variant, the diseases status of all subjects and the age of disease onset of patients. In bold are indicated pedigrees (n=7) used to study the penetrance of aHUS in c.286+2T&gt;G carriers.</li> <li>In the sheet &quot;Haplotypes&quot; are reported genotypes used to evaluate the association between the presence of <em>CFH-H3</em> and <em>CD46<sub>GGAAC</sub></em> risk haplotypes and aHUS. Results of this analysis are reported in Table 3 of the published paper.</li> <li>In the sheet &quot;Raw data Fig.2&quot; are reported data of &quot;platelet count&quot; and &quot;serum creatinine&quot; of the proband used to elaborate Figure 2.</li> </ul> <p>In the file <strong>&quot;C3 and C5b-9 deposition&quot;</strong> is reported the quantification of serum-induced C3 and C5b-9 deposition on human microvascular endothelial cell line (HMEC-1). The fluorescent staining was evaluated with Image J and expressed as pixel<sup>2 </sup>per field analyzed. The fields with the lowest and highest values were excluded from calculation. These values were used to elaborate data included in Table 2 and in Figure 5.</p> <p>In the file <strong>&quot;CD46 protein expression&quot;</strong> are reported data of CD46 expression on peripheral blood mononuclear cells (PBMCs) isolated from the proband, his relatives and healthy volunteers. Data of specific expression of CD46 (evaluated for SCR1 or for SCR4 as reported in the materials and methods section) are indicated as median fluorescence intensity (MFI) percentage compared with the control.</p> <p>In the ppt file <strong>&quot;cDNA amplification and sequencing results&quot;</strong> is reported:</p> <ul> <li>the agarose gel image of the amplified cDNA from the control (ctr), the proband (IV-8) and his healthy father (III-7).</li> <li>Electropherograms obtained from the cDNA sequencing of the control (ctr), the proband (IV-8) and his healthy father (III-7).</li> </ul> <p>Additional data will be made available by the authors, without undue reservation, to any qualified researcher.&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Mammary single-cell RNA-seq analysis and prostate cancer survival as a function of H2AFJ expression for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial cells

<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

A common NFKB1 variant detected through antibody analysis in UK Biobank predicts risk of infection and allergy: Summary statistics - Health records

<p>Infectious agents contribute significantly to the global burden of diseases, through both acute infection and their chronic sequelae. We leveraged the UK Biobank to identify genetic loci that influence humoral immune response to multiple infections. From 45 genome-wide association studies in 9,611 participants from UK Biobank, we identified NFKB1 as a locus associated with quantitative antibody responses to multiple pathogens including those from the herpes, retro- and polyoma-virus families. An insertion-deletion variant thought to affect NFKB1 expression (rs28362491), was mapped as the likely causal variant. This variant has persisted throughout hominid evolution and could play a key role in regulation of the immune response. Using 121 infection and inflammation related traits in 487,297 UK Biobank participants, we show that the deletion allele was associated with an increased risk of infection from diverse pathogens but had a protective effect against allergic disease. We propose that altered expression of NFKB1, as a result of the deletion, modulates haematopoietic pathways, and likely impacts cell survival, antibody production, and inflammation. Taken together, we show that disruptions to the tightly regulated immune processes may tip the balance between exacerbated immune responses and allergy, or increased risk of infection and impaired resolution of inflammation.&nbsp;</p> <p>-------------------------------------------------------------------------------------</p> <p>This dataset contains GWAS summary statistics for infection, inflammation, and allergy related traits in&nbsp;487,297 individuals</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

A common NFKB1 variant detected through antibody analysis in UK Biobank predicts risk of infection and allergy: Summary statistics - Serology

<p>Infectious agents contribute significantly to the global burden of diseases, through both acute infection and their chronic sequelae. We leveraged the UK Biobank to identify genetic loci that influence humoral immune response to multiple infections. From 45 genome-wide association studies in 9,611 participants from UK Biobank, we identified NFKB1 as a locus associated with quantitative antibody responses to multiple pathogens including those from the herpes, retro- and polyoma-virus families. An insertion-deletion variant thought to affect NFKB1 expression (rs28362491), was mapped as the likely causal variant. This variant has persisted throughout hominid evolution and could play a key role in regulation of the immune response. Using 121 infection and inflammation related traits in 487,297 UK Biobank participants, we show that the deletion allele was associated with an increased risk of infection from diverse pathogens but had a protective effect against allergic disease. We propose that altered expression of NFKB1, as a result of the deletion, modulates haematopoietic pathways, and likely impacts cell survival, antibody production, and inflammation. Taken together, we show that disruptions to the tightly regulated immune processes may tip the balance between exacerbated immune responses and allergy, or increased risk of infection and impaired resolution of inflammation.&nbsp;</p> <p>-------------------------------------------------------------------------------------</p> <p>This dataset contains GWAS summary statistics for quantitative antibody responses&nbsp;in 9611 individuals and results for a&nbsp;meta-analysis of UK Biobank and CoLaus/PsyCoLaus antibody responses.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Dataset and structure database for an ML model to predict diffusivity in ZIF variants

<p>This dataset accompanies the publication titled &quot;Data Mining for Predicting Gas Diffusivity in Zeolitic-imidazolate Frameworks (ZIFs)&quot; (DOI:&nbsp;<a href="https://doi.org/10.1039/D2TA02624D">https://doi.org/10.1039/D2TA02624D</a>)</p> <p><a href="https://zenodo.org/api/files/b80f6d07-3bf4-484c-97ac-5d579fb0cc27/ESI_2_dataset.xlsx?versionId=dc4525d0-1c5c-478c-9bef-a156587ad69b">ESI_2_dataset.xlsx</a>: Descriptors for all ZIFs of the publication and simulations output, in the form of diffusivities of gas molecules (He up to iso-butane), in all ZIFs.</p> <p>ZIF_database.zip: ZIP file containing all ZIFs prepared by the authors (as discussed in the publication), through various units replacements, in the SOD topology, in .pdb&nbsp;format.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Supporting data for "CoVEffect: Interactive System for Mining the Effects of SARS-CoV-2 Mutations and Variants Based on Deep Learning"

<p>This repository contains the datasets created and extracted for the paper:</p> <p>Giuseppe Serna Garc&iacute;a, Ruba Al Khalaf, Francesco Invernici, Stefano Ceri, and Anna Bernasconi. 2022.<br> &quot;<strong>CoVEffect</strong>: Interactive System for Mining the <strong>Effects of SARS-CoV-2 Mutations and Variants</strong> Based on Deep Learning&quot;. (Available online at http://gmql.eu/coveffect)</p> <p>--------------------------------------------------------------------------------<br> LIST OF FILES WITH DESCRIPTION:<br> --------------------------------------------------------------------------------</p> <p>AdditionalFile1-effects-taxonomy:<br> Descriptions of legal values for the &#39;Effect&#39; field, based on a categorized taxonomy.</p> <p>AdditionalFile2-levels-taxonomy:<br> Descriptions of legal values for the &#39;Level&#39; field.</p> <p>AdditionalFile3-training_dataset_target:<br> List of target tuples (manually annotated) of 221 abstracts considered for training the model. For each abstract, target tuples&nbsp; follow the schema ID, DOI, title, entity, effect, level, type (mutation or variant), tuples_count (&gt;1 when an effect/level is shared by multiple entities, #abstracts containing the same effect described in the tuple).</p> <p>AdditionalFile4-validation_dataset_target:<br> List of target tuples (manually annotated) of 50 abstracts considered for validating the prepared prediction model.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile5-validation_dataset_highlighted:<br> Textual abstracts of the 50 manuscripts considered for validation; the text used to support the manual target annotations has been highlighted in yellow.</p> <p>AdditionalFile6-validation_dataset_prediction:<br> List of predicted annotations of 50 abstracts considered for validating the prepared prediction model. The file is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p> <p>AdditionalFile7-keywords_query_list:<br> Keyword-based search run on the CORD-19 dataset to extract a relevant subset of abstracts regarding the scope of interest of CoVEffect. The Boolean logic used to combine keywords is explained in the section &#39;Annotations of the biology-related CORD-19 cluster&#39;.</p> <p>AdditionalFile8-CORD-19_batch_dataset_metadata:<br> Metadata of the 7,230 papers extracted by the keyword-based query in AdditionalFile7.<br> These abstracts have been annotated by the prediction framework.</p> <p>AdditionalFile9-CORD-19_batch_dataset_prediction:<br> List of predicted annotations of 7,230 abstracts extracted from the biology-related cluster of CORD-19.</p> <p>AdditionalFile10-test_dataset_target:<br> List of target tuples (manually annotated) of 100 abstracts randomly selected from the 7,230 extracted as in AdditionalFile8.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile11-test_dataset_prediction:<br> List of predicted annotations of 100 abstracts considered for testing the prediction model on a subset of the CORD-19 biology-related cluster. As AdditionalFile6, it is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p>

opencc-zeroDec 2022View details →
zenodo44/100

Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data

<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11):&nbsp;2429&ndash;2442.</p> <p>Code available at:&nbsp;https://github.com/kelsj/EVICORD</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Variant calls for 'Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration'

<p>A VCF of SNP calls used for input to GWAS in https://elifesciences.org/articles/26255</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Modeling islet enhancers using deep learning identifies candidate causal variants at loci associated with T2D and glycemic traits

<p>Genetic association studies have identified hundreds of independent genetic signals associated with type 2 diabetes (T2D) and related traits. Despite these successes, the identification of specific causal variants underlying a genetic association signal remains challenging. In this study, we describe a deep learning method to analyze the impact of sequence variants on enhancers. Focusing on pancreatic islets, a relevant T2D tissue, we show that our model learns islet-specific transcription factor (TF) regulatory patterns and can be used to prioritize candidate causal variants. At 101 genetic signals associated with T2D and related glycemic traits where multiple variants occur in linkage disequilibrium, our method nominates a single causal variant for each association signal, including three variants previously shown to alter reporter activity in islet-relevant cell types. For another signal associated with blood glucose levels, we biochemically test all candidate causal variants from statistical fine-mapping using a pancreatic islet beta cell line and show biochemical evidence of allelic effects on TF binding for the model-prioritized variant. To aid in future research, we publicly distribute our model and islet enhancer perturbation scores across ~67 million variants. We anticipate that deep learning methods like the one presented in this study will enhance the prioritization of candidate causal variants for functional studies.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Keccak variant data

<p>This dataset contains all data referring to the Keccak implementation used in the optimized variant of Elephant.</p><p>Specifically, it contains:</p><ul><li>The source code after being processed by CIL</li><li>The CIL code after being manually processed further</li><li>The while+ code after being transformed from CIL</li><li>The while+ code, after being manually processed further</li><li>The annotated program resulting from the first pass derivation rules being applied on the while+ code</li><li>The annotated programs resulting from the second pass derivation rules being applied on the results of the first pass</li></ul><p>Note that the second pass derivation trees also contain a small amount of natural language, to explain what is being done, apart from simply applying the derivation rules.</p><p>&nbsp;</p><p>This new version contains files that were previously missing.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Genetic variants (chr. 6) from Old World Schistosoma mansoni exomes

<p>Variant calling file (VCF) produced from exome libraries of <em>Schistosoma mansoni</em> (bloodfluke) samples from the Old Wold (West Africa (Senegal, Niger), East Africa (Tanzania), and Middle East (Oman)). One sample form the New World (Caribbean (HR9)) was added for comparison. The variants were called on the 3 Mb of chromosome 6 centered on the <em>SmSULT-OR</em> gene. This gene is involved in resistance to the drug oxamniquine&nbsp; (OXA). The aim of the related article was to investigate the origin of OXA resistant mutations in the New Wolrd by identifying sequence variation in <em>SmSULT-OR</em> in <em>S. mansoni</em> from the Old World, where OXA has seen minimal usage.</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Paralog variant classification and scoring

<p><em>Para_zscore </em>data</p> <p>Input data, annotation of all hg19 missense variants, score for every gene having a paralog in the human genes. This dataset is a supplement for the publication Lal. et al.</p> <p>Information on the files, scripts to generate and use the <em>para_zscore</em> are available under</p> <p>https://git-r3lab.uni.lu/genomeanalysis/paralogs.</p> <p>Version 3582386 updates:</p> <p>- Annovar annotation file for hg38 added</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Assembly and variant calling results of strain mixtures of HCMV

<p>The tgz files contain&nbsp;the assembly contigs of 10&nbsp;different assemblers and variant calling results&nbsp;of 6&nbsp;callers on the strain mixture dataset of the HCMV virus.&nbsp;</p>

opencc-byOct 2019View details →
zenodo40/100

Data for Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery

<p><strong>Description of the datasets</strong></p> <p>Data are organized as folders and compressed with tar.gz.</p> <p>There are two compressed data folder: <strong>data </strong>which used for cattle genome graphs experiment and&nbsp;<strong>data_human</strong> which we used for human genome graphs experiment.&nbsp;</p> <p><strong>Cattle genome graphs experiments</strong></p> <p>First you need to unzip the file using command <em>tar -xvzf data.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>Utilities: contain bovine ARS-UCD 1.2 fasta reference with the accompanying index.</li> <li>Bin: contain the softwares used in the paper (vg, liftover, vcf2diploid)</li> <li>Part1: data for analysis in variant prioritization section, further subdivided into: <ul> <li>vcf_sim: variant files from four animal in each breed used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul> </li> <li>Part2: data used for analysis in the section of graph mapping with breeds-filtered variants, further subdivided into: <ul> <li>vcf_breed: variant files used to graphs construction.</li> </ul> </li> <li>Part3: data used for analysis in the section of consensus genome, further subdivided into: <ul> <li>read_sims: simulated reads as in the part1, but the coordinates are liftovered to the new consensus genomes.</li> <li>reference: contain the original reference and consensus references.</li> <li>vcf_consensus: contain major allele variants to construct consensus genomes.</li> </ul> </li> <li>Part4: data analysis in the section of whole genome graph construction and variant genotyping. <ul> <li>vcf_construct: variants from chromosome 1-29 from 82 Brown Swiss used to construct BSW whole genome graph.</li> <li>BSW_graph: whole genome Brown Swiss graph with the three accompanying indexes (xg,gcsa, and gbwt).</li> </ul> </li> </ul> <p><strong>Human genome graphs experiments</strong></p> <p>First you need to unzip the <em>data_human</em> file using command <em>tar -xvzf data</em><em>_hum.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>reference: the g1k_v37 reference used as a graph backbone</li> <li>vcf_sim: variant files from four individuals in each population used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul>

opencc-by-4.0Dec 2019View details →
zenodo40/100

i2QTL HipSci Structural Variant and Short Tandem Repeat Genotypes

<p>Here we provide structural variant and short tandem repeat variant calls from 204 HipSci donors as described in the manuscript Jakubosky et al.&nbsp;&quot;Discovery and quality analysis of a comprehensive set of structural variants and short tandem repeats&quot;. &nbsp;&nbsp;</p> <p><a href="https://www.biorxiv.org/content/10.1101/713198v2">https://www.biorxiv.org/content/10.1101/713198v2</a></p> <p>Jakubosky D, Smith EN, D&rsquo;Antonio M, Bonder MJ, Young Greenwald WW, Matsui H, D&rsquo;Antonio-Chronowska A, (Hipsci), Stegle O, Montgomery SB, DeBoever C, Frazer KA. Discovery and Quality Analysis of a Comprehensive Set of Structural Variants and Short Tandem Repeats. bioRxiv. January 2019:713198. doi:10.1101/713198.</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Intersections between clinical variant databases

<p>The upset plots visualize the distribution of 990 unique validated variants from patients with MDS/AML detected in seven public datasets (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA388411) among clincal variant databases as well as the overlap between the different databases.</p> <p><strong>A</strong> Distribution of the variants with therapeutic implications (tier I)</p> <p><strong>B </strong>Distribution of the variants with diagnostic and/or prognostic implications or disease association (tier II)</p> <p><strong>C</strong> Distribution of the variants of unclear significance (tier III)<br> &nbsp;</p> <p><strong>List of abbreviations:</strong></p> <p>Cancer Genome Interpreter&rsquo;s variants database (CGI)</p> <p>Clinical Interpretations of Variants in Cancer database (CIVIC)</p> <p>Clinical Variants database of the National Center for Biotechnology Information (ClinVar)</p> <p>Catalogue Of Somatic Mutations In Cancer (COSMIC)</p> <p>Database of Curated Mutation (DOCM)</p> <p>Human Gene Mutation Database (HGMD)</p> <p>Jackson Laboratory Clinical Knowledgebase (JAX)</p> <p>MolecularMatch database (MM)</p> <p>Oncology Knowledge Base (OncoKB)</p> <p>Pharmacogenomics Knowledgebase (PharmGKB)</p> <p>Phenotype for ENCODE (PhenCode)</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

VCF file of Variants in a wide onion cross segregating for bolting

<p>VCF (variant call format) file of&nbsp;variants from bulked segregant RNA PoolSeq (BSR-seq) of F2 progeny pools of a wide onion cross segregating for bolting (precocious flowering). The reference assembly is&nbsp;GBGJ00000000.1&nbsp;http://www.ncbi.nlm.nih.gov/nuccore/656904698</p> <p>The pools were taken from materials sampled for validation of the <em>AcBlt1</em> locus as described in&nbsp;http://www.ncbi.nlm.nih.gov/pubmed/24247236</p> <p>Pools were from bolting or non-bolting plants homozygous at the most closely linked marker to&nbsp;<em>AcBlt1 &nbsp;</em>either for the bolt-associated genotype (AA) or the non-bolt genotype (BB).&nbsp;</p>

opencc-zeroAug 2014View details →
zenodo40/100

Common Genetic Variants in FOXP2 are Not Associated with Individual Differences in Language Development

<p>Three data sets used in&rdquo; Common Genetic Variants <em>in FOXP2</em> Are Not Associated with Individual Differences in Language Development&rdquo; are provided.&nbsp; The discovery data set was comprised of 834 children who were members of a Longitudinal sample and children who were members of a School sample.&nbsp; Both samples are contained in the Iowa data set.&nbsp; The Iowa data set contains a quantitative variable LCOMP that represents a composite z-score representing oral language ability.&nbsp; The data set also identifies which sample the children belonged to and the allele calls for 13 tag SNPs located across <em>FOXP2</em>.&nbsp; A second data file, ELVS, contains data from a separate sample of children who were used to test for replication of inconsistent evidence of an association between language and the SNP rs1916988.&nbsp; The ELVS file provides a composite oral language score scaled in standard score units (mean=100, SD=15) and the genotype calls for the &nbsp;SNP rs1916988.</p>

opencc-zeroJan 2016View details →
zenodo40/100

Raw BRCA1/2 variants in breast cancer patients and healthy relatives produced with GATK.

<p>Aligned sequencing data is available in the NCBI Sequence Read Archive (SRA, https://www.ncbi.nlm.nih.gov/sra/) under accession SRP095082. Variants were called using GATK HaplotypeCaller (version 3.6). After joint performing joint genotyping multi-sample vcf file was generated. Next, SNPs and indels were extracted into two different vcf files and specific set of filters were applied for each case.</p> <p> </p> <p><strong>File descriptions</strong></p> <p><strong><em>Datasets</em></strong></p> <p><strong>BRCA_SNVs.vcf</strong> - this file contains SNPs called with GATK and hard filters applied. Following filtering options were applied: "QD &lt; 2.0", "FS &gt; 60.0", "MQ &lt; 40.0",  "MQRankSum &lt; -12.5", "ReadPosRankSum &lt; -8.0", "SB &lt; -0.10" , "DP &lt; 10" , "GQ &lt; 30" , and "SOR &gt; 3.0"</p> <p><strong>BRCA_indels.vcf</strong> - This file contains indels called with GATK and hard filters applied. Following filtering options were applied: "QD &lt; 2.0", "FS &gt; 200.0", "ReadPosRankSum &lt; -20.0", "InbreedingCoeff &lt; -0.8", "SOR &gt; 10.0".</p> <p> </p> <p><strong><em>Scripts package (scritps.zip)</em></strong></p> <p>Scripts.zip file contains scripts and supporting files for genotype calling and filtering. </p> <p><strong>raw.variant.caling.sh </strong>– bam files preprocessing, alignment refining and raw genotype calling with HaplotypeCaller.</p> <p><strong>genotyping_and_filtering.sh </strong>– joint genotyping, variant hard filtering and callset refinement.</p> <p><strong>LIST.txt</strong> – supporting file that contains bam filenames containing aligned reads.</p> <p><strong>sample_order.txt</strong> – supporting file for sample renaming.</p> <p> </p> <p><strong><em>Reference files (hg19) used in variant calling scripts</em></strong></p> <p>Reference files can be downloaded from GATK bundle web-site at https://software.broadinstitute.org/gatk/download/bundle.  </p> <p><strong>ucsc.hg19.fasta</strong> - human genome assembly;</p> <p><strong>Mills_and_1000G_gold_standard.indels.hg19.sites.vcf.gz</strong> – set of known indels to be used for local realignment;</p> <p><strong>1000G_phase1.indels.hg19.sites.vcf.gz</strong> – set of known indels to be used for local realignment;</p> <p><strong>dbsnp_138.hg19.vcf.gz</strong> – a recent dbSNP release (build 138); </p> <p><strong>1000G_phase3_v4_20130502.hg19.lifted.sites.vcf</strong> – the latest set from 1000G phase 3 (v4) for genotype refinement.</p> <p> </p>

opencc-by-4.0Dec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record