Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,868

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,868 results for “variant”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supplementary data for publication Global distribution of mcr gene variants in 214K metagenomic samples

<p># Supplementary data for the manuscript &quot;Global distribution of mcr gene variants in 214,095 metagenomic samples&quot;</p> <p>SD1_mapped_runids.csv : tab-separated file with columns of run_accessions downloaded from ENA and whether the metagenome were positive for at least one of the mcr genes.</p> <p>SD2_mcr_df.csv : compositional table of mcr-positive metagenomes with associated metadata (collection_year, country, and host) for each run_accession, as well as mapping results.</p> <p>SD3_mcr_contigs.fa : FASTA file with contigs carrying mcr genes. The header contains the run_accession ID.</p> <p>SD4_aldex2_results.csv: CSV file containing ALDEx2 results. The columns are as follows:<br> * group: metadata category (year, country or host). If the column contains more than one label, e.g., &quot;Denmark - 2020 - Pigs&quot;, significance is tested within Danish pig samples from 2020.<br> * rab.all:&nbsp; median clr value for all samples in the feature<br> * rab.win.conditionA:&nbsp; median clr value for the condition A of samples<br> * rab.win.conditionB: median clr value for the condition B of samples<br> * diff.btw: median difference in clr values between A and B conditions<br> * diff.win: median of the largest difference in clr values within A and B conditions<br> * effect : median effect size: diff.btw / max(diff.win) for all instances<br> * overlap : proportion of effect size that overlaps 0 (i.e. no effect)<br> * we.ep: Expected P value of Welch&rsquo;s t test<br> * we.eBH: Expected Benjamini-Hochberg corrected P value of Welch&rsquo;s t test<br> * wi.ep: Expected P value of Wilcoxon rank test<br> * wi.eBH: Expected Benjamini-Hochberg corrected P value of Wilcoxon test<br> * parts: gene name<br> * conditionA: label of condition A that is compared against condition B<br> * conditionB: label of condition B that is compared against condition A<br> * conditions.A.vs.B: label to explain condition A compared against condition B<br> NOTE: see for more explanation of the output of ALDEx2 https://www.bioconductor.org/packages/release/bioc/vignettes/ALDEx2/inst/doc/ALDEx2_vignette.html#5_ALDEx2_outputs</p> <p>SD5: Multi-VCF file containing SNP information on mcr alleles. Can be used to construct consensus sequences.</p> <p>SD6: FASTA file containing all unique consensus sequences reported in the manuscript.</p> <p>SD7: CSV file with an overview of which metagenome contains which unique consensus sequence.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

CADD-SV - A framework to score the effects of structural variants in health and disease

<p>Required annotation data-set to run the CADD-SV framework; a method to retrieve and integrate a wide set of annotations to predict the effects of SVs. Pre-scored variants as well as additional information on used features.<br> A webserver for online scoring as well as data downloads is available at: https://cadd-sv.bihealth.org/<br> Source code for CADD-SV is available at GitHub: https://github.com/kircherlab/CADD-SV</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

A Large-scale Dataset of (Open Source) License Text Variants

<p>We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive&mdash;the largest publicly available archive of FOSS source code with accompanying development history&mdash;all versions of files whose names are commonly used to convey licensing terms to software users and developers.<br> The dataset consists of 6.5 million unique license files that can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing.<br> Additional metadata about shipped license files are also provided, making the dataset ready to use in various contexts; they include: file length measures, detected MIME type, detected SPDX license (using ScanCode), example origin (e.g., GitHub repository), oldest public commit in which the license appeared.<br> The dataset is released as open data as an archive file containing all deduplicated license blobs, plus several portable CSV files for metadata, referencing blobs via cryptographic checksums.</p> <p>For more details see the included&nbsp;README file and companion paper:</p> <ul> <li>Stefano Zacchiroli.&nbsp;<a href="https://doi.org/10.1145/3524842.3528491"><em>A Large-scale Dataset of (Open Source) License Text Variants</em></a>. In proceedings of the&nbsp;<a href="https://conf.researchr.org/home/msr-2022">2022 Mining Software Repositories Conference (MSR 2022)</a>. 23-24 May 2022 Pittsburgh, Pennsylvania, United States. ACM 2022.</li> </ul> <p>If you use this dataset for research purposes, please acknowledge its use by citing the above paper.</p> <ul> </ul>

opencc-by-4.0Mar 2022View details →
zenodo44/100

data set to bioRxiv preprint 'Persistent cross-species SARS-CoV-2 variant infectivity predicted via comparative molecular dynamics simulation

<p>This is supporting data and software code for the following preprint in bioRxiv</p> <p><strong>Persistent cross-species SARS-CoV-2 variant infectivity predicted via comparative molecular dynamics simulation</strong></p> <p>https://www.biorxiv.org/content/10.1101/2022.04.18.488629v1</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Phase I trial of CX-5461, a first-in-class G-quadruplex stabilizer in patients with advanced solid tumors enriched for DNA-repair deficiencies (CCTG IND.231) - Variant Calls

<p>Variant Calls from Phase I trial of CX-5461, a first-in-class G-quadruplex stabilizer in patients with&nbsp; advanced solid tumors enriched for DNA-repair deficiencies (CCTG IND.231)</p> <p>See publication for methodology.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

PRJEB16533 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to project accession:&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB16533">PRJEB16533</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PRJNA516678 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to project accession:&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJNA516678">PRJNA516678</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PRJEB39369 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to project accession: <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB39369">PRJEB39369</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PRJNA596071 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to project accession:&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJNA596071">PRJNA596071</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PRJNA311274 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to project accession:&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJNA311274">PRJNA311274</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

PRJNA363032 raw variants

<p>Sequencing reads were aligned to the Amel_HAv3.1 reference genome using BWA-MEM v0.7.17. Reads were sorted with SAMtools v1.9 and duplicates marked (MarkDuplicates) with GATK v4.0.11.0. Variants for each sample were called using GATK&rsquo;s HaplotypeCaller with the following non-default parameters --ERC GVCF, --sample-ploidy 1 and -A AlleleFraction. Joint variant calling was performed across all samples collated for AmelHap using GATK&rsquo;s GenomicDBImport and GenotypeGVCFs with --sample-ploidy 1 and a window size of 10 Mb. This dataset comprises the raw variant calls only for samples belonging to&nbsp;project accession:&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJNA363032">PRJNA363032</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

SETDB1 knock-out and reexpression of a WT, a catalytic-dead variant (CA) or a NLS-tagged SETDB1 protein

<p>In order to generate the Setdb1 cKO allele, exons 15 and 16, which encode the core amino acids of the catalytic domain, have been flanked with two lox-P sites recognized by Cre recombinase enzyme. Cre-oestrogen receptor fusion gene Mer-Cre-Mer has been introduced to induce acute Setdb1 KO after Tamoxifen treatment. Setdb1 cKO mESCs where Setdb1 expression is stably rescued by wild type 3xFlag-Setdb1 (WT) or by the catalytic dead mutant 3xFlag-Setdb1 (CA) were generously given by Pr Yoichi Shinkai. The catalytic dead mutant was obtained with a single lysine mutation. In the lab, Setdb1 cKO mESCs, where Setdb1 expression is stably rescued by NLS-3xFlag-Setdb1 which localizes only in the nucleus, have been established (3 NLS sequences have been added in order to retain Setdb1 in the nucleus).</p> <p>The quality of the RNA was determined on the Agilent 2100 Bioanalyzer (Agilent Technologies, Palo Alto, CA, USA), the RNA integrity number was above 8 for all the samples. To construct libraries, 1 mg of high-quality total RNA sample was processed using Truseq&reg; stranded total RNA kit (Illumina&reg;). After the removal of ribosomal RNAs (using Ribo-zero&reg; rRNA), confirmed by QC control on pico chipTM on the Agilent 2100 Bioanalyzer (Agilent Technologies, Palo Alto, CA, USA), total RNA molecules are fragmented and reverse-transcribed using random primers. Replacement of dTTP by dUTP during the second strand synthesis will permit to achieve the strand specificity. Addition of a single A base to the cDNA is followed by ligation of adapters. Libraries were quantified by qPCR using the KAPA Library Quantification Kit for Illumina Libraries (KapaBiosystems) and library profiles were assessed using the DNA High SensitivityTMHS kit on an Agilent Bioanalyzer 2100. Libraries were sequenced on an Illumina&reg; Nextseq 500 instrument using 75 base-lengths read V2 chemistry in a paired-end mode.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Functional genomics analysis to disentangle the role of genetic variants in major depression - Supplementary Tables

<p>This entry contains the data generated by the study &quot;Functional genomics analysis to disentangle the role of genetic variants in major depression&quot; that are part of the Supplementary information of&nbsp;the article describing the study.</p> <p>The entry contains the following data:</p> <p><strong>Supplementary Tables S1-S7</strong></p> <p>Supplementary Table S1. Summary of resources.</p> <p>Supplementary Table S2. Causal GVs for MD.</p> <p>Supplementary Table S3. pGenes functional and disease enrichment analysis.</p> <p>Supplementary Table S4. Fine-mapped MD causal GVs disease enrichment analysis.</p> <p>Supplementary Table S5. Colocalizing GWAS-eQTLs association to disease.</p> <p>Supplementary Table S6. TFBS analysis.</p> <p>Supplementary Table S7. GVs state annotation.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"

<p>These are the burden test results (summary statistics) for the publication:</p> <p>&quot;Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer&rsquo;s Disease&quot;,</p> <p>Nature Genetics, 2022.</p> <p>&nbsp;</p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on&nbsp;6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL&gt;=75, LOF+REVEL&gt;=50, LOF+REVEL&gt;=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Variant dataset and code for "Population-level whole genome sequencing of Ascochyta rabiei identifies genomic loci associated with isolate aggressiveness"

<p>This dataset contains genetic variants (SNPs) of <em>Ascochyta rabiei</em> isolates and the R code used in their analysis to generate the results and figures described in the manuscript "<strong>Population-level whole genome sequencing of <em>Ascochyta rabiei</em> identifies genomic loci associated with isolate aggressiveness</strong>".</p> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Dataset for genes and gene variants from familial gastroschisis

<p>The dataset consists of records from whole exome sequecing and bioinformatic analysis which includes genes and gene variants from a Mexican family with recurrence for gastroschisis (two affected half-sisters with gastroschisis, mother, and father of the proband).</p> <p>Release of this dataset was based on the Human Genome annotation, GRCh37/hg19.</p> <p>The full list of tables is described in the file READ ME and remain available in csv files.</p> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Genetic Variants Representation Learning (GV-Rep)

<p>This dataset is used for Genetic Variants (GV) representation learning.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Assembly of enterohemorrhagic Escherichia coli type IV pilin PpdD and its variants

<p>Assembly of the EHEC major type IV pilin PpdD was analyzed in a reconstituted TP assembly system described in LunaRico et al. Mol Microbiol. 2019 Mar;111(3):732-749. doi: 10.1111/mmi.14188. Bacteria of strain BW25113 F&#39;tet harboring plasmids pMS41 and pCHAP8565 (or its variants) were grown for 2 days at 30&deg;C on M9 plates containing 0.5% glycerol, amplicillin (100 ug/ml) chloramphenicol&nbsp; (25 ug/ml) and 1 mM IPTG.</p> <p>Bacteria were collected and fractionated as described in Luna Rico et al&nbsp;Methods Mol Biol. 2018;1764:291-305. doi: 10.1007/978-1-4939-7759-8_18. Cell and sheared fractions were analysed by electrophoresis on 10 % Tris-Tricin gels, transferred on nitrocellulose and probed with anti-MalE-PpdD polyclonal antibodies. The fluorescence signal was developed with ECL2 (Thermo) and recorded with Typhoon FLA9000 imager (GE).</p> <p>The signal was quantified using ImageJ. The fractions of PpdD assembled into pili were quantified and analysed using Prism9.</p> <p>The images uploaded here are the raw data used to produce the Fig. 4B of the article Karami et al., Structure, 2021.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Phenopackets for case reports of structural variants

<p>A&nbsp;collection of 188 published deleterious structural variants based on 182 cases published in 146 clinical case reports that describe individuals with Mendelian diseases.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

A vaccine-induced public antibody protects against SARS-CoV-2 and emerging variants

<p>These are the<strong> processed</strong> BCR repertoire bulk&nbsp;sequencing data described in <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz,&nbsp;Turner &amp;&nbsp;Liu et al., Immunity, 2021</a>.&nbsp;The <strong>raw</strong> sequence data are available on SRA under BioProjects <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">PRJNA731610</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA741267">PRJNA741267</a>.&nbsp;</p> <p><strong>Summary</strong>:&nbsp;Bulk-sorted total plasmablasts and IgDlo enriched B cells&nbsp;from PBMCs&nbsp;and germinal centre&nbsp;B cells from lymph nodes from various timepoints&nbsp;after primary immunization from 22&nbsp;BNT162b2&nbsp;vaccinees who had no prior history of infection with SARS-CoV-2.&nbsp;</p> <p><strong>Metadata file</strong>:&nbsp;WU368_schmitz_et_al_immunity_2021_meta.tsv</p> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>PB = plasmablast</li> <li>GC = germinal center</li> <li>mAb = monoclonal antibody</li> </ul> <p><strong>BCR data file</strong>:&nbsp;WU368_schmitz_et_al_immunity_2021_bcr.tsv.gz</p> <p>In addition to the processed bulk sequences, also included are the&nbsp;heavy chains of 37 mAbs (including 2C08)&nbsp;first reported in <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner &amp; O&#39;Halloran et al., Nature, 2021</a>&nbsp;that had been validated to be spike-binding. The mAbs are annotated as &quot;mab&quot; in the &quot;seq_type&quot; column.</p> <p><strong>Sequence data column description</strong></p> <p>The columns largely follow the&nbsp;<a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s are used, as opposed to IMGT-defined &quot;junctions&quot;. Non-standard columns are noted below.</p> <ul> <li>v_call_genotyped:&nbsp;V gene annotation reassigned after individualized genotyping&nbsp;by&nbsp;<a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a></li> <li>isotype: IGH[ADEGM]</li> <li>cdr3: CDR3 nucleotide sequence</li> <li>cdr3_length: CDR3 nucleotide sequence length</li> <li>cdr3_aa: CDR3 amino acid sequence</li> <li>donor: vaccinee ID</li> <li>sample: sample ID (arbitrary)</li> <li>timepoint: time point at which sample was collected</li> <li>tissue: tissue from which sample was collected</li> <li>sorting: FACS sorting</li> <li>seq_type: sequence type (mAb or bulk)</li> </ul>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record