Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “number of alleles”

Learn how ShareScore rates datasets ↗
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S1 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S0 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
dryad36/100

Revisiting the number of self‐incompatibility alleles in finite populations: From old models to new results

<p>Under gametophytic self-incompatibility (GSI), plants are heterozygous at the self-incompatibility locus (S-locus) and can only be fertilized by pollen with a different allele at that locus. The last century has seen a heated debate about the correct way of modeling the allele diversity in a GSI population that was never formally resolved. Starting from an individual-based model, we derive the deterministic dynamics as proposed by Fisher (1958), and compute the stationary S-allele frequency distribution. We find that the stationary distribution proposed by Wright (1964) is close to our theoretical prediction, in line with earlier numerical confirmation. Additionally, we approximate the invasion probability of a new S-allele, which scales inversely with the number of resident S-alleles. Lastly, we use the stationary allele frequency distribution to estimate the population size of a plant population from an empirically obtained allele frequency spectrum, which complements the existing estimator of the number of S-alleles. Our expression of the stationary distribution resolves the long-standing debate about the correct approximation of the number of S-alleles and paves the way to new statistical developments for the estimation of the plant population size based on S-allele frequencies.</p>

opencc-zeroAug 2022View details →
dryad36/100

Revisiting the number of self‐incompatibility alleles in finite populations: From old models to new results

Open the record for dataset details and reuse information.

publicJun 2024View details →
zenodo32/100

Inferring allele-specific copy number aberrations and tumor phylogeography from spatially resolved transcriptomics (output data)

<p>This contains the output results of CalicoST (inferred CNAs and cancer clones), results of comparison methods, and CNAs inferred from WES data of 13 samples across four cancer types.</p> <p>In this updated version, we also included the simulated data and the results from CalicoST and other methods in CalicoST_simulation_deposit.zip. README contains the details of deposited files.</p>

opencc-by-4.0Sep 2024View details →
dryad32/100

Data from: Number of alleles as a predictor of the relative assignment accuracy of STR and SNP baselines for chum salmon

Short tandem repeat (STR) markers, which exhibit many alleles per locus, are commonly used to assign fish to their populations of origin. Single nucleotide polymorphisms (SNPs), which have many technical advantages over STRs, typically exhibit only two alleles per locus. Simulation studies have indicated that number of independent alleles is a good predictor of accuracy of genetic markers for fishery applications. Extant STR baselines for salmon contain hundreds of alleles, and it has been extrapolated that hundreds of SNP markers need to be developed before SNP baselines will compare to these STR baselines. We compared 15 STRs exhibiting 349 independent alleles to 61 SNP assays exhibiting 66 independent alleles for accuracy in assigning to closely related populations of chum salmon. The SNP baseline yielded slightly higher mean accuracies for proportional assignment and comparable accuracies for individual assignment. Overall the SNP baseline performed considerably better, relative to the microsatellite baseline, than predicted based on the number of independent alleles in each baseline. We suggest that this discrepancy is due to the fact that the simulation studies do not capture the impacts of the different strategies commonly employed for discovering and selecting STR and SNP markers.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Number of alleles as a predictor of the relative assignment accuracy of STR and SNP baselines for chum salmon

Open the record for dataset details and reuse information.

publicApr 2011View details →
geo24/100

A SNF2 protein targets variable copy number repeats and thereby influences allele-specific expression

GEO Series GSE22162. Homo sapiens; Mus musculus. 10 samples. Type: Genome binding/occupancy profiling by array; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2010View details →
geo24/100

Genome-wide copy number and allele-specific copy number analysis of choroid plexus tumors (I)

GEO Series GSE60899. Homo sapiens. 55 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenNov 2014View details →
geo24/100

Genome-wide copy number and allele-specific copy number analysis of choroid plexus tumors (II)

GEO Series GSE61363. Homo sapiens. 20 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenNov 2014View details →
geo24/100

Landscape of somatic allelic imbalances and copy number alterations in HER2-amplified breast cancer

GEO Series GSE31645. Homo sapiens. 26 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.

openGEO-OpenDec 2011View details →
geo24/100

Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios

GEO Series GSE11977. Homo sapiens. 6 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenOct 2008View details →
geo24/100

Allele-specific copy number analysis of tumor samples with aneuploidy and tumor heterogeneity

GEO Series GSE26302. Homo sapiens. 14 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenSep 2011View details →
geo20/100

Characterizing genetic transitions of copy number alterations and allelic imbalances in metastatic process of the oral tongue carcinoma

GEO Series GSE76014. Homo sapiens. 40 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenSep 2016View details →
geo20/100

Landscape of somatic allelic imbalances and copy number alterations in human lung carcinoma

GEO Series GSE29065. Homo sapiens. 78 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenOct 2012View details →
geo20/100

Single-Stranded Annealing Induced by Re-Initiation of Replication Origins Provides a Novel and Efficient Mechanism for Generating Copy Number Expansion via Non-Allelic Homologous Recombination.

GEO Series GSE41259. Saccharomyces cerevisiae. 333 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenJan 2013View details →
geo16/100

Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios [CytoSNP]

GEO Series GSE190390. Homo sapiens. 19 samples. Type: Genome variation profiling by array; SNP genotyping by SNP array.

openGEO-OpenDec 2021View details →
geo16/100

Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios [Omni2.5]

GEO Series GSE190392. Homo sapiens. 80 samples. Type: Genome variation profiling by array; SNP genotyping by SNP array.

openGEO-OpenDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record