Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “number of alleles”
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S1 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S0 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Revisiting the number of self‐incompatibility alleles in finite populations: From old models to new results
<p>Under gametophytic self-incompatibility (GSI), plants are heterozygous at the self-incompatibility locus (S-locus) and can only be fertilized by pollen with a different allele at that locus. The last century has seen a heated debate about the correct way of modeling the allele diversity in a GSI population that was never formally resolved. Starting from an individual-based model, we derive the deterministic dynamics as proposed by Fisher (1958), and compute the stationary S-allele frequency distribution. We find that the stationary distribution proposed by Wright (1964) is close to our theoretical prediction, in line with earlier numerical confirmation. Additionally, we approximate the invasion probability of a new S-allele, which scales inversely with the number of resident S-alleles. Lastly, we use the stationary allele frequency distribution to estimate the population size of a plant population from an empirically obtained allele frequency spectrum, which complements the existing estimator of the number of S-alleles. Our expression of the stationary distribution resolves the long-standing debate about the correct approximation of the number of S-alleles and paves the way to new statistical developments for the estimation of the plant population size based on S-allele frequencies.</p>
Revisiting the number of self‐incompatibility alleles in finite populations: From old models to new results
Open the record for dataset details and reuse information.
Inferring allele-specific copy number aberrations and tumor phylogeography from spatially resolved transcriptomics (output data)
<p>This contains the output results of CalicoST (inferred CNAs and cancer clones), results of comparison methods, and CNAs inferred from WES data of 13 samples across four cancer types.</p> <p>In this updated version, we also included the simulated data and the results from CalicoST and other methods in CalicoST_simulation_deposit.zip. README contains the details of deposited files.</p>
Data from: Number of alleles as a predictor of the relative assignment accuracy of STR and SNP baselines for chum salmon
Short tandem repeat (STR) markers, which exhibit many alleles per locus, are commonly used to assign fish to their populations of origin. Single nucleotide polymorphisms (SNPs), which have many technical advantages over STRs, typically exhibit only two alleles per locus. Simulation studies have indicated that number of independent alleles is a good predictor of accuracy of genetic markers for fishery applications. Extant STR baselines for salmon contain hundreds of alleles, and it has been extrapolated that hundreds of SNP markers need to be developed before SNP baselines will compare to these STR baselines. We compared 15 STRs exhibiting 349 independent alleles to 61 SNP assays exhibiting 66 independent alleles for accuracy in assigning to closely related populations of chum salmon. The SNP baseline yielded slightly higher mean accuracies for proportional assignment and comparable accuracies for individual assignment. Overall the SNP baseline performed considerably better, relative to the microsatellite baseline, than predicted based on the number of independent alleles in each baseline. We suggest that this discrepancy is due to the fact that the simulation studies do not capture the impacts of the different strategies commonly employed for discovering and selecting STR and SNP markers.
Data from: Number of alleles as a predictor of the relative assignment accuracy of STR and SNP baselines for chum salmon
Open the record for dataset details and reuse information.
A SNF2 protein targets variable copy number repeats and thereby influences allele-specific expression
GEO Series GSE22162. Homo sapiens; Mus musculus. 10 samples. Type: Genome binding/occupancy profiling by array; Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide copy number and allele-specific copy number analysis of choroid plexus tumors (I)
GEO Series GSE60899. Homo sapiens. 55 samples. Type: Genome variation profiling by SNP array.
Genome-wide copy number and allele-specific copy number analysis of choroid plexus tumors (II)
GEO Series GSE61363. Homo sapiens. 20 samples. Type: Genome variation profiling by SNP array.
Landscape of somatic allelic imbalances and copy number alterations in HER2-amplified breast cancer
GEO Series GSE31645. Homo sapiens. 26 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.
Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios
GEO Series GSE11977. Homo sapiens. 6 samples. Type: Genome variation profiling by SNP array.
Allele-specific copy number analysis of tumor samples with aneuploidy and tumor heterogeneity
GEO Series GSE26302. Homo sapiens. 14 samples. Type: Genome variation profiling by SNP array.
Characterizing genetic transitions of copy number alterations and allelic imbalances in metastatic process of the oral tongue carcinoma
GEO Series GSE76014. Homo sapiens. 40 samples. Type: Genome variation profiling by SNP array.
Landscape of somatic allelic imbalances and copy number alterations in human lung carcinoma
GEO Series GSE29065. Homo sapiens. 78 samples. Type: Genome variation profiling by genome tiling array.
Single-Stranded Annealing Induced by Re-Initiation of Replication Origins Provides a Novel and Efficient Mechanism for Generating Copy Number Expansion via Non-Allelic Homologous Recombination.
GEO Series GSE41259. Saccharomyces cerevisiae. 333 samples. Type: Genome variation profiling by genome tiling array.
Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios [CytoSNP]
GEO Series GSE190390. Homo sapiens. 19 samples. Type: Genome variation profiling by array; SNP genotyping by SNP array.
Normalization of Illumina Infinium whole-genome SNP data improves copy number estimates and allelic intensity ratios [Omni2.5]
GEO Series GSE190392. Homo sapiens. 80 samples. Type: Genome variation profiling by array; SNP genotyping by SNP array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.