Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,283
datasets available to search
ShareScore release 0.7.1
Dataset results
1,283 results for “Copying”
Computational validation of clonal and subclonal copy number alterations from bulk tumour sequencing
<p>Somatic variant identification from WGS data is a crucial step in the analysis of cancer genomes. Several tools are available to perform mutations calling, however, the noisiness of data requires appropriate quality control assessment. In the preprint work available at https://doi.org/10.1101/2021.02.13.429885 we present CNAqc, an R package devised to assess the quality of allele-specific Copy Number Alterations (CNA), somatic mutations, and tumor purity estimates. In order to test the model, we ran CNAqc on 2778 single-sample PCAWG whole-genomes and 48 TCGA whole-exomes. We uploaded the results of our analysis using the release of the tool available at https://github.com/caravagnalab/CNAqc/releases/tag/rr_22_0.1 in the form of .rds files. All the necessary details for the reproduction of our results are reported in the Supplementary Materials of the above-mentioned manuscript.</p>
The genomic architecture of the passerine MHC region: high repeat content and contrasting evolutionary histories of single copy and tandemly duplicated MHC genes
<p><span>The Major Histocompatibility Complex (MHC) is of central importance to the immune system, and an optimal MHC diversity is believed to maximize pathogen elimination. Birds show substantial variation in MHC diversity, ranging from few genes in most bird orders to very many genes in passerines. Our understanding of the evolutionary trajectories of the MHC in passerines is hampered by lack of data on genomic organization. Therefore, we assemble and annotate the MHC genomic region of the great reed warbler (<em>Acrocephalus arundinaceus</em>), using long-read sequencing and optical mapping. The MHC region is large (>5.5Mb), characterized by structural changes compared to hitherto investigated bird orders and shows higher repeat content</span><span> than the genome average. These features were supported by analyses in three additional passerines. MHC genes in passerines are found in two different chromosomal arrangements, either as single copy MHC genes located among non-MHC genes, or as tandemly duplicated tightly linked MHC genes. Some single copy MHC genes are old and putative orthologs among species. In contrast tandemly duplicated MHC genes are monophyletic within species and have evolved by simultaneous gene duplication of several MHC genes. Structural differences in the MHC genomic region among bird orders seem substantial compared to mammals and have possibly been fuelled by clade-specific immune system adaptations. Our study provides methodological guidance in characterizing complex genomic regions, constitutes a resource for MHC research in birds, and calls for a revision of the general belief that avian MHC has a conserved gene order and small size compared to mammals.</span></p>
Loghub-version6(Copy)
<p>This is a copy of LogHub version 6, which is an open benchmark for log testing.</p> <p>Original link: <a href="https://zenodo.org/record/1596245#.Yv3zYOxBxqs">https://zenodo.org/record/1596245#.Yv3zYOxBxqs</a></p> <p>Original Paper Reference: He, Shilin, et al. "Loghub: a large collection of system log datasets towards automated log analytics." <em>arXiv preprint arXiv:2008.06448</em> (2020).</p> <p>Original Paper Link: https://arxiv.org/pdf/2008.06448v1.pdf</p>
Supplementary material 1 from: Tedersoo L, Liiv I, Kivistik PA, Anslan S, Kõljalg U, Bahram M (2016) Genomics and metagenomics technologies to recover ribosomal DNA and single-copy genes from old fruit-body and ectomycorrhiza specimens. MycoKeys 13: 1-20. https://doi.org/10.3897/mycokeys.13.8140
Full information and metadata about the genomic and metagenomic samples : Explanation note: Detailed information about metadata, DNA quality and genomic/metagenomic results of fruit-body and EcM root tip samples.
Plasma metabolomic signatures for copy number variants and COVID-19 risk loci in Northern Finland Populations
<p>A MySQL database for metabolomic signatures for The Northern Finland Birth Cohorts, <strong>NFBC1966 </strong>and <strong>NFBC1986</strong>, a longitudinal research program based at the Medical Faculty, University of Oulu, Finland. This dataset is a supplement to the manuscript titlted "Plasma metabolomic signatures for copy number variants and COVID-19 risk loci in Northern Finland Populations"</p>
Fig. 4 in ELAV Intron 8: a single-copy sequence marker for shallow to deep phylogeny in Eupulmonata Hasprunar & Huber, 1990 and Hygrophila Férussac, 1822 (Gastropoda: Mollusca)
Fig. 4 Comparison of ML phylogenetic reconstructions based on ELAVI8 and concatentated ITS1 +2 sequence from 21 California, USA, Haplotrema specimens representing 9 described and 3 undescribed species. The ELAVI8 Haplotrema tree was rooted on a Discidae, a chondrinid, two Orthurethra and two Limacoidea species. This was not possible in
Fig. 5 in ELAV Intron 8: a single-copy sequence marker for shallow to deep phylogeny in Eupulmonata Hasprunar & Huber, 1990 and Hygrophila Férussac, 1822 (Gastropoda: Mollusca)
Fig. 5 Comparison of ML phylogenetic reconstructions based on ELAVI8 and 28S sequence from 50 specimens/37 genera representing all of the major panpulmonate land snail clades identified by
Fig. 3 Base pair coverage across the 1296 in ELAV Intron 8: a single-copy sequence marker for shallow to deep phylogeny in Eupulmonata Hasprunar & Huber, 1990 and Hygrophila Férussac, 1822 (Gastropoda: Mollusca)
Fig. 3 Base pair coverage across the 1296 aligned sites in the ELAVI8 MSA (see ELAVI8_panpul.fas in the Supporting Information). The y-axis represents the percentage of sites that are represented at that location across all specimens
Fig. 2 ELAVI8 PCR yield and stringency across using a 1.1 in ELAV Intron 8: a single-copy sequence marker for shallow to deep phylogeny in Eupulmonata Hasprunar & Huber, 1990 and Hygrophila Férussac, 1822 (Gastropoda: Mollusca)
Fig. 2 ELAVI8 PCR yield and stringency across using a 1.1% Agarose gel in 1 × TBE buffer with GoldView Dye and 3 µl of PCR product from each reaction. Panel A represents 50° C annealing temperature for 40 cycles while B represents the modified touchdown procedure of 54° C anneal for 10 cycles followed by 50° C for 30 cycles. The
GToTree prepackaged single-copy gene HMMs
<p>SCG HMM datasets from the initial publication of GToTree, based on NCBI, updated with modified "Universal" set from the initial GToTree release thanks to notes from Molly Chen</p> <p> - PF00181 ("Ribosomal_L2") was changed to PF03947 ("Ribosomal_L2_C")<br> - the C-terminal (which PF03947 covers) is better conserved<br> - PF00827 ("Ribosomal_L15") was changed to PF00828 ("Ribosomal_L27A")<br> - PF00827 was archaea/euk only, PF00828 holds the bac/arc L15 also<br> - PF17135 ("Ribosomal_L18") was changed to PF00861 ("Ribosomal_L18p")<br> - PF00861 is better conserved</p> <p>For general process, see: https://github.com/AstrobioMike/GToTree/wiki/SCG-sets</p>
Dataset for "Luminal breast epithelial cells from wildtype and BRCA mutation carriers harbor copy number alterations commonly associated with breast cancer"
<p>Processed single cell whole genome sequencing data from:</p> <p>Luminal breast epithelial cells from wildtype and BRCA mutation carriers harbor copy number alterations commonly associated with breast cancer Williams, Vinci Oliphant et al 2024</p> <p>Included are processed copy number profiles for all cells included in the study.</p>
Molecular Dynamics Trajectories of Membrane-bound Influenza Hemagglutinin (A/duck/Alberta/35/76) in Complex with 0-3 Copies of FISW84 Fab Fragments
<p>MD simulation trajectories of influenza hemagglutinin (A/duck/Alberta/35/76) in a bilayer mimicking the viral membrane, with 0-3 copies of FISW84 antibody's Fab domains bound.</p> <p>Files are named as (copy number of Fab).(simulation replica ID).(file type extension). The systems were constructed in CHARMM format PSF files and the trajectories were recorded in DCD format.</p> <p>Simulations were performed with NAMD3. The trajectories were re-centered and re-wrapped about the periodic boundary from the raw trajectories, in order to keep the HA-Fab complex and the lipid bilayer appearing as a single continuous entity instead of isolated molecules at the opposite side of the periodic boundary. Water molecules in the simulation were removed in these trajectories due to file size limitations.</p>
Figure 5 in Exploring phylogenetic informativeness and nuclear copies of mitochondrial DNA (numts) in three commonly used mitochondrial genes: mitochondrial phylogeny of peppermint, cleaner, and semi-terrestrial shrimps (Caridea: Lysmata, Exhippolysmata, and Merguia)
Figure 5. Phylogenetic informativeness of three mtDNA gene fragments (16S, 12S, and COI) in peppermint, cleaner, and semi-terrestrial shrimps. (A) Phylogenetic informativeness (PI) profiles of the three different mtDNA gene fragments studied through relative time in shrimps from the genera Lysmata, Exhippolysmata, and Merguia. The sum of the instantaneous asymptotic informativeness of all sites in each gene is plotted. The arrows and numbers above or below them indicate the relative time (arrow) and magnitude (numbers) at which PI reaches its maximum value. (B) Tree topology resulting from the maximum-likelihood analysis of the sequences studied with a relative time-enforced branch length. This phylogeny was used to calculate the PI profiles in panel (A). Species pertaining to the different monophyletic clades previously revealed by the combined analyses of the three mtDNA gene fragments are highlighted with different colours, as in Figure 3.
Figure 7. Neighbour-nets generated using SplitsTree4 in Exploring phylogenetic informativeness and nuclear copies of mitochondrial DNA (numts) in three commonly used mitochondrial genes: mitochondrial phylogeny of peppermint, cleaner, and semi-terrestrial shrimps (Caridea: Lysmata, Exhippolysmata, and Merguia)
Figure 7. Neighbour-nets generated using SplitsTree4 from the three mtDNA gene fragments studied (16S, 12S, and COI) in shrimps from the genera Lysmata, Exhippolysmata, and Merguia. Species pertaining to the different monophyletic clades previously revealed by the combined analyses of the three mtDNA gene fragments are highlighted with different colours, as in Figure 3. Abbreviations: LA, Lysmata ankeri; LABP, Lysmata cf. vittata; LAM, Lysmata amboinensis; LARG, Lysmata argentopuctata; LBA, Lysmata bahia; LBO, Lysmata boggessi; LCA, Lysmata californica; LD, Lysmata debelius; LGA, Lysmata galapagensis; LGB, Lysmata grabhami; LGR, Lysmata gracilirostris; LH, Lysmata hochi; LHO, Lysmata holthuisi; LI, Lysmata intermedia; LIM2, Lysmata cf. intermedia; LK, Lysmata kuekenthali; LM, Lysmata moorei; LN, Lysmata nayaritensis; LNI, Lysmata nilita; LO, Lysmata olavoi; LP, Lysmata pederseni; LRA, Lysmata rafa; LSET, Lysmata seticaudata; LT, Lysmata cf. ternatensis; LV, Lysmata vittata; LU, Lysmata udoi; LWEF, Lysmata wurdemanni EFL; LWG, Lysmata wurdemanni TX; LWWF, Lysmata wurdemanni WFL; EXO, Exhippolysmata oplophoroides; EXE, Exhippolysmata ensirostris; MO, Merguia oligodon; MR, Merguia rhizophorae; and NSP, Nikoides sp.
Figure 3 in Exploring phylogenetic informativeness and nuclear copies of mitochondrial DNA (numts) in three commonly used mitochondrial genes: mitochondrial phylogeny of peppermint, cleaner, and semi-terrestrial shrimps (Caridea: Lysmata, Exhippolysmata, and Merguia)
Figure 3. Tree topology resulting from the combined analysis of the three mtDNA gene fragments studied (16S, 12S, and COI) for shrimps from the genus Lysmata (29 taxa), Exhippolysmata (two taxa), Merguia (two taxa), and one out-group (Nikoides sp.), under maximum likelihood (ML). Numbers above or below the branches represent the bootstrap values obtained from the maximum likelihood (ML) analysis in TREEFINDER and posterior probabilities from the Bayesian inference (BI) analysis in MrBayes (ML/BI). The general topology of the trees obtained from ML and BI analyses was the same.
Figure 1 in Exploring phylogenetic informativeness and nuclear copies of mitochondrial DNA (numts) in three commonly used mitochondrial genes: mitochondrial phylogeny of peppermint, cleaner, and semi-terrestrial shrimps (Caridea: Lysmata, Exhippolysmata, and Merguia)
Figure 1. Amino acid usage analysis (mean amino acid count per sequence) for COI reference sequences (from selected species of crustaceans: Macrobrachium rosenbergii, Exopalaemon caricaudinata, Halocaridina rubra, and Cherax destructor), for COI orthologous sequences obtained from shrimps from the genus Lysmata, and for COI-like cloned sequences from Lysmata seticaudata. The error bars in each graph represent the highest and lowest amino acid counts per sequence in the three data sets. Amino acid determination and naming follows the invertebrate mitochondrial translation code, and was performed in MEGA 5.
Figure 2 in Exploring phylogenetic informativeness and nuclear copies of mitochondrial DNA (numts) in three commonly used mitochondrial genes: mitochondrial phylogeny of peppermint, cleaner, and semi-terrestrial shrimps (Caridea: Lysmata, Exhippolysmata, and Merguia)
Figure 2. Tree topologies resulting from the analysis of COI-like cloned sequences from Lysmata seticaudata and mtDNA COI gene fragments for shrimps from the genus Lysmata (29 taxa), Exhippolysmata (two taxa), Merguia (two taxa), and one out-group (Nikoides sp.), under maximum likelihood (ML) and Bayesian inference (BI). Numbers above or below the branches represent the bootstrap values obtained from the ML analysis in TREEFINDER, and posterior probabilities from the BI analysis in MrBayes.
A copy of website at https://www.naenamarray2020.info
<p>This is the full copy of the website 'https://www.naenamarray2020.info'</p>
Whole Proteome Copy Number Dataset in Primary Mouse Cortical Neurons
<p>Raw data-Supplementary table- "Whole Proteome Copy Number Dataset in Primary Mouse Cortical Neurons"</p>
Single-cell somatic copy number variants in brain using different amplification methods and reference genomes
<p>Variable and constant sized bins for GRCh38 and T2T-Chm13 were generated using the buildGenome scripts provided with Ginkgo (<a href="https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts">https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts</a>).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.