Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

63

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

63 results for “linkage mapping”

Learn how ShareScore rates datasets ↗
zenodo44/100

An updated map of GRCh38 linkage disequilibrium blocks based on European ancestry data

<p>A map of approximately independent linkage disequilibrium (LD) blocks has many uses in statistical genetics. Current publicly available LD block maps are based on sparse recombination maps and are only available for GRCh37 (hg19) and prior genome assemblies. We generated LD blocks in GRCh38 for European (EUR) ancestry populations using a recent recombination map based on more than 115,000 individuals. This new map consists of 1,361 independent LD blocks across the 22 autosomal chromosomes and can be accessed at https://github.com/jmacdon/LDblocks_GRCh38</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Data for Linkage mapping of root shape traits associated with market class in two biparental carrot populations

<p>&nbsp;</p> <p>This repository contains essential data to support the findings presented in the forthcoming publication titled &quot;Linkage Mapping of Root Shape Traits Associated with Market Class in Two Biparental Carrot Populations.&quot; It includes VCF files for two distinct carrot biparental populations, as well as R code for filtering, constructing linkage maps, and conducting QTL analysis. Furthermore, the repository hosts phenotypic data gathered from these two biparental populations during the years 2020 and 2021.</p> <p>Two carrot genetic maps, one for each population, have been made available alongside their respective phenotypic data.</p> <p>The provided R code contains absolute working directory paths that may not function as intended on your system. The primary purpose of sharing this code is to offer readers insight into the techniques employed in this study. You may need to adapt the directory paths to suit your specific setup.&nbsp;</p> <p>To assist readers in understanding the logical sequence of steps involved in our linkage mapping project, the R code scripts have been sequentially numbered from 0 to 10.</p> <p>For more info contact: vegaalfaro@wisc.edu.</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

High-density genetic linkage mapping in Sitka spruce advances the integration of genomic resources in conifers

<p><span>In species with large and complex genomes such as conifers, dense linkage maps are a useful for supporting genome assembly and laying the genomic groundwork at the structural, populational and functional levels. However, most of the 600+ extant conifer species still lack extensive genotyping resources, which hampers the development of high-density linkage maps. In this study, </span><span><span>we developed a linkage map relying on 21,570 SNP makers in </span></span><span>Sitka spruce (<em>Picea sitchensis</em> [Bong.] Carr.)</span><span><em><span>, </span></em></span><span><span>a long-lived conifer from western North America that is widely planted for productive forestry in the British Isles. </span></span><span>We used a single-step mapping approach to efficiently combine RAD-Seq and genotyping array SNP data for 528 individuals from two full-sib families. As expected for spruce taxa, the saturated map contained 12 linkages groups with a total length of 2,142 cM. The positioning of 5,414 unique gene coding sequences allowed us to compare our map with that of other Pinaceae species, which provided evidence for high levels of synteny and gene order conservation in this family. We then developed an integrated map for <em>P. sitchensis</em> and <em>P. glauca</em> based on 27,052 makers and 11,609 gene sequences. Altogether, these two linkage maps, the accompanying catalog of 286,159 SNPs and the genotyping chip developed herein opens new perspectives for a variety of fundamental and more applied research objectives, such as for the improvement of spruce genome assemblies, or for marker-assisted sustainable management of genetic resources in Sitka spruce and related species.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

Data in support of "An exploration of linkage fine-mapping on sequences from case-control studies"

<p>These data were simulated for an exploration of linkage fine-mapping on sequences from case-control studies. The code to generate and analyze the data is available on GitHub in the scripts at&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.&nbsp;Queries may be directed to Payman Nickchi at&nbsp;<a href="mailto:pnickchi@sfu.ca">pnickchi@sfu.ca</a>&nbsp;or Charith (Bhagya) Karunarathna at&nbsp;<a href="mailto:ch757276@dal.ca">ch757276@dal.ca</a>.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data in support of "An exploration of linkage fine-mapping on sequences from case-control studies"

<p>These data were simulated for an exploration of linkage fine-mapping on sequences from case-control studies. The scripts&nbsp;to generate and analyze the data are&nbsp;available&nbsp;at&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.&nbsp;Queries may be directed to Payman Nickchi at&nbsp;<a href="mailto:pnickchi@sfu.ca">pnickchi@sfu.ca</a>&nbsp;or Charith (Bhagya) Karunarathna at&nbsp;<a href="mailto:ch757276@dal.ca">ch757276@dal.ca</a>.</p> <p><strong>README file for All_data directory</strong></p> <p><strong>Directory structure</strong></p> <p>The&nbsp;All_data&nbsp;directory consists of this README file and 500 sub-directories named&nbsp;DatasetX, for&nbsp;X=1 to 500. Within each&nbsp;DatasetX&nbsp;sub-directory are further sub-directories named&nbsp;alt&nbsp;and&nbsp;null&nbsp;containing files named&nbsp;pop_data.RData&nbsp;and&nbsp;sample_data.RData.</p> <p><strong>alt&nbsp;<em>versus</em>&nbsp;null&nbsp;directories</strong></p> <p>The files in the&nbsp;alt&nbsp;and&nbsp;null&nbsp;directories contain the same variant data but different phenotype data. In particular, under the null hypothesis, disease status is simulated at random according to a 5% prevalence in the population, whereas under the alternative hypothesis disease status is simulated according to a penetrance model that depends on causal SNVs. The R script to simulate data<br> under the alternative hypothesis is in the file&nbsp;1_SimulateData.R&nbsp;in the Github repository&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.</p> <p><strong>pop_data.RData&nbsp;and&nbsp;sample_data.RData&nbsp;files</strong></p> <p>The data structures contained in the&nbsp;pop_data.RData&nbsp;and&nbsp;sample_data.RData&nbsp;files are described below. The structure is the same under both the null and alternative hypothesis.</p> <p><strong>pop_data.RData</strong></p> <p>From R,&nbsp;load(&quot;pop_data.RData&quot;)&nbsp;loads a list named&nbsp;pop_data&nbsp;whose elements describe the population&rsquo;s haplotype and phenotype data. The list elements are as follows.</p> <ul> <li>Variants: a matrix of variants for the population of 6200 haplotypes <ul> <li>rows are SNVs,</li> <li>columns are sequences</li> </ul> </li> <li>Positions: a data frame of SNV positions <ul> <li>rows are SNVs,</li> <li>column 1 is the SNV name and column 2 is the SNV position in base pairs</li> </ul> </li> <li>Population.Mapping: a data frame telling us how the sequences are paired into individuals <ul> <li>rows are individuals</li> <li>First column 1 is an individual ID from 1,&hellip;,3100; columns 2 and 3 are the sequence IDs of the first and second sequence for that individual where the sequence IDs are the column names of the&nbsp;Variants&nbsp;matrix.</li> </ul> </li> <li>Genotype.Matrix: a matrix of genotypes (i.e.&nbsp;variant counts) for the 3100 individuals <ul> <li>rows are SNVs</li> <li>columns are the individuals</li> </ul> </li> <li>causal_region: a vector containing the lower- and upper-limit of the causal region in base pairs.</li> <li>cSNV: a vector containing the IDs of the causal SNVs, where the SNV IDs are the row names of the&nbsp;Variants&nbsp;matrix.</li> <li>DISCRETE: a list with the following elements. <ul> <li>CaseIndividuals: vector of IDs of the affected individuals in the population.</li> <li>ControlIndividuals: vector of IDs of the unaffected in the population.</li> <li>BinaryTrait: a vector of trait status (0=unaffected, 1=affected) for each individual.</li> </ul> </li> </ul> <p><strong>Note:</strong>&nbsp;Within the same&nbsp;DatasetX&nbsp;directory, the only difference between the&nbsp;pop_data&nbsp;data structures under the null and alternative hypothesis is the phenotype information contained in their respective&nbsp;DISCRETE&nbsp;list elements. Both the null and alternative pop_data data structure share&nbsp;list elements: Variants,&nbsp;Positions,&nbsp;Population.Mapping,&nbsp;Genotype.Matrix,&nbsp;causal_region&nbsp;and&nbsp;cSNV.</p> <p><strong>sample_data.RData</strong></p> <p>From R,&nbsp;load(&quot;sample_data.RData&quot;)&nbsp;loads a list whose elements describe the sequences and phenotypes of the sample of 50 affected individuals (cases) and 50 unaffected individuals (controls) from the population.</p> <ul> <li>Haps: a list with two elements. <ul> <li>sample_haps: a matrix of 200 sequences for the 50 cases and 50 controls. Rows are SNVs and columns are sequences, with the sequences of sampled cases appearing first (i.e.&nbsp;first 100 columns), followed by the sequences of sampled controls (i.e.&nbsp;last 100 columns). Sequences include only those SNVs that are polymorphic in the sample.</li> <li>ccStatus: a vector indicating the case/control status of the individual to which the sequence belongs, with case=1 and control=0.</li> </ul> </li> <li>Genos: a list with two elements. <ul> <li>sample_genos: a matrix of 100 genotypes for the 50 cases and 50 controls. Rows are SNVs and columns are genotypes, with genotypes of cases appearing first, followed by genotypes of controls.</li> <li>ccStatus: a vector indicating the case/control status of each individual, with case=1 and control=0.</li> </ul> </li> <li>Posn: a data frame of SNV positions for each SNV that is polymorphic in the sample. The first column is the SNV name and the second is the SNV position in base pairs.&nbsp;Posn&nbsp;is a subset of&nbsp;pop_data$Positions.</li> <li>poly_cSNV: a vector of IDs for causal SNVs that are polymorphic in the sample.</li> <li>CaseIND: a vector of individual IDs for the case individuals (see&nbsp;pop_data$Population.Mapping).</li> <li>ControlIND: a vector of individual IDs for the control individuals (see&nbsp;pop_data$Population.Mapping).</li> <li>CaseHapID: a vector of IDs for the sequences that belong to cases (see the sequence IDs in the column names of the matrix&nbsp;pop_data$Variants).</li> <li>ControlHapID: a vector of IDs for the sequences that belong to controls (see the sequence IDs in the column names of the matrix&nbsp;pop_data$Variants).</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

RAD-SEQ LINKAGE MAPPING AND PATTERNS OF SEGREGATION DISTORTION IN SEDGES: MEIOSIS AS A DRIVER OF KARYOTYPIC EVOLUTION IN ORGANISMS WITH HOLOCENTRIC CHROMOSOMES" in Journal of Evolutionary Biology

<p>This a data set from the paper RAD-SEQ LINKAGE MAPPING AND PATTERNS OF SEGREGATION DISTORTION IN SEDGES: MEIOSIS AS A DRIVER OF KARYOTYPIC EVOLUTION IN ORGANISMS WITH HOLOCENTRIC CHROMOSOMES&quot; to be published in Journal of Evolutionary Biology</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Constructing a high-density linkage map to infer the genomic landscape of recombination rate variation in European Aspen (Populus tremula)

<p>Data sets and files for linkage map construction and for inferring recombination rate variation in <em>Populus tremula</em>. Associated scripts for analyses can be found at <a href="https://github.com/parkingvarsson/Recombination_rate_variation">https://github.com/parkingvarsson/Recombination_rate_variation</a>&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Linkage maps and genotype data of strawberry produced with skim-sequencing data

<p>The following set of files contain the results and scripts to produce those results, described in Chapter 5 of the PhD thesis of Alejandro Th&eacute;r&egrave;se Navarro, entitled &quot;How to map a million markers: linkage mapping of skim-sequencing data in strawberry&quot;. In this study, a large dataset of markers produced by whole genome resequecning of a strawberry (<em>Fragaria&nbsp;</em>x&nbsp;<em>ananassa</em>) biparental population are used to generate linkage maps. To that end the software <a href="https://github.com/Alethere/SmoothDescent">Smooth Descent</a>&nbsp;is used, since it is oriented to obtaining linkage maps in usin low quality (error-prone) genotype data. With this methodology we were able to produce a linkage map of 27 out of 28 chromosomes of strawberry which containing 1.85M markers in ~2400 unique genetic mpositions. We also compare this map with a linkage map produced using SNP array data and with the genome sequence assembly&nbsp;&quot;Camarosa&quot;.</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

High-density genetic linkage mapping in Sitka spruce advances the integration of genomic resources in conifers

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad36/100

Updated Pinus lambertiana high-density linkage maps

<p>Sugar pine (<em>Pinus</em> <em>lambertiana</em>) is a five-needle pine of great economical and ecological importance in western North America. Genome-wide analyses in the species have been limited by the absence of a chromosome-scale reference genome. The genome of the species is one of the largest in any plant diploid species, with a size of 31 Gb, making assembly and annotation very challenging.</p> <p>Here we present high-density linkage maps that were used to help scaffolding a new chromosome-scale genome of sugar pine (assembly version 2.0, unpublished). </p>

opencc-zeroSep 2023View details →
dryad36/100

Data from: Genome of an iconic Australian bird: High-quality assembly and linkage map of the superb fairy-wren (Malurus cyaneus)

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad36/100

Updated Pinus lambertiana high-density linkage maps

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad32/100

Construction of genetic linkage map based on SNP markers, QTL mapping and detection of candidate genes of growth-related traits in Pacific abalone using genotyping-by-sequencing

<p><a name="_Hlk72585736"><span>Pacific abalone (<i>Haliotis discus hannai</i>) is a commercially important high valued molluscan species. Its wild population has decreased in recent years. Pacific abalone is widely cultured in Korea. Traditional breeding programs have been implemented for hatchery production of abalone seeds. To obtain more genetic information for the molecular breeding program, a high-density linkage map and quantitative trait locus (QTL) for three growth-related traits was constructed for Pacific abalone. F1 cross population with two parents were sampled to construct the linkage map using genotyping by sequencing (GBS). A total of 664,630,534 clean reads and 56,686 SNPs were generated. In sum, 3,345 segregating SNPs were used to construct a consensus linkage map. The map spanned 1,747.023 cM with 18 linkage groups and an average interval of 0.55 cM. QTL analysis revealed two significant QTL in LG10 on the consensus linkage map in each growth-related trait. Both the QTLs are located in the telomere region of the chromosome. Moreover, four potential candidate genes for growth-related traits were identified in the QTL region. Expression analysis revealed that identified genes are involved in growth regulation of abalone. The newly constructed genetic linkage map, growth-related QTLs and potential candidate genes identified in the present study can be used as valuable genetic resources and will be useful for marker-assisted selection (MAS) of Pacific abalone in molecular breeding program.</span></a></p>

opencc-zeroJun 2021View details →
dryad32/100

Data from: Linkage mapping reveals strong chiasma interference in sockeye salmon: implications for interpreting genomic data

Meiotic recombination is fundamental for generating new genetic variation and for securing proper disjunction. Further, recombination plays an essential role during the rediploidization process of polyploid-origin genomes because crossovers between pairs of homeologous chromosomes retain duplicated regions. A better understanding of how recombination affects genome evolution is crucial for interpreting genomic data; unfortunately, current knowledge mainly originates from a few model species. Salmonid fishes provide a valuable system for studying the effects of recombination in nonmodel species. Salmonid females generally produce thousands of embryos, providing large families for conducting inheritance studies. Further, salmonid genomes are currently rediploidizing after a whole genome duplication and can serve as models for studying the role of homeologous crossovers on genome evolution. Here, we present a detailed interrogation of recombination patterns in sockeye salmon (Oncorhynchus nerka). First, we use RAD sequencing of haploid and diploid gynogenetic families to construct a dense linkage map that includes paralogous loci and location of centromeres. We find a nonrandom distribution of paralogs that mainly cluster in extended regions distally located on 11 different chromosomes, consistent with ongoing homeologous recombination in these regions. We also estimate the strength of interference across each chromosome; results reveal strong interference and crossovers are mostly limited to one per arm. Interference was further shown to continue across centromeres, but metacentric chromosomes generally had at least one crossover on each arm. We discuss the relevance of these findings for both mapping and population genomic studies.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Landscape connectivity for wildlife: development and validation of multi-species linkage maps

The ability to identify regions of high functional connectivity for multiple wildlife species is of conservation interest with respect to forest management and corridor planning. We present a method that does not require independent, field-collected data, is insensitive to the placement of source and destination sites (nodes) for modeling connectivity, and does not require the selection of a focal species. In the first step of our approach, we created a cost surface that represented permeability of the landscape to movement for a suite of species. We randomly selected nodes around the perimeter of the buffered study area and used circuit theory to connect pairs of nodes. When the buffer was removed, the resulting current density map represented, for each grid cell, the probability of use by moving animals. We found that using nodes that were randomly located around the perimeter of the buffered study area was less biased by node placement than randomly selecting nodes within the study area. We also found that a buffer of ≥ 20% of the study area width was sufficient to remove the effects of node placement on current density. We tested our method by creating a map of connectivity in the Algonquin to Adirondack region in eastern North America, and we validated the map with independently collected data. We found that amphibians and reptiles were more likely to cross roads in areas of high current density, and fishers (Pekania [Martes] pennanti) used areas with high current density within their home ranges. Our approach provides an efficient and cost-effective method of predicting areas with relatively high functional connectivity.

opencc-zeroDec 2013View details →
dryad32/100

Data from: A high-density linkage map for Astyanax mexicanus using genotyping-by-sequencing technology

The Mexican tetra, Astyanax mexicanus, is a unique model system consisting of cave-adapted and surface-dwelling morphotypes which diverged &gt;1My ago. This remarkable natural experiment has enabled powerful genetic analyses of cave adaptation. Here, we describe the application of next-generation sequencing technology to the creation of a high-density linkage map. Our map comprises over 2200 markers populating 25 linkage groups constructed from genotypic data generated from a single genotyping-by-sequencing project. We leveraged emergent genomic and transcriptomic resources to anchor hundreds of anonymous Astyanax markers to the genome of the zebrafish (Danio rerio), the most closely related model organism to our study species. This facilitated the identification of 784 distinct connections between our linkage map and the Danio rerio genome, highlighting several regions of conserved genomic architecture between the two species despite ~150My of divergence. Using a Mendelian cave-associated trait as a proof-of-principle, we successfully recovered the genomic position of the albinism locus near the gene Oca2. Further, our map successfully informed the positions of unplaced Astyanax genomic scaffolds within particular linkage groups. This ability to identify the relative location, orientation and linear order of unaligned genomic scaffolds will facilitate ongoing efforts to improve upon the current early draft and assemble future versions of the Astyanax physical genome. Moreover, this improved linkage map will enable higher resolution genetic analyses and catalyze the discovery of the genetic basis for cave-associated phenotypes.

opencc-zeroDec 2014View details →
dryad32/100

Data from: A microsatellite-based linkage map for song sparrows (Melospiza melodia)

Although linkage maps are important tools in evolutionary biology, their availability for wild populations is limited. The population of song sparrows (Melospiza melodia) on Mandarte Island, Canada, is among the more intensively studied wild animal populations. Its long-term pedigree data, together with extensive genetic sampling, have allowed the study of a range of questions in evolutionary biology and ecology. However, the availability of genetic markers has been limited. We here describe 191 new microsatellite loci, including 160 high-quality polymorphic autosomal, 7 Z-linked and 1 W-linked markers. We used these markers to construct a linkage map for song sparrows with a total sex-averaged map length of 1731 cM and covering 35 linkage groups, and hence, these markers cover most of the 38–40 chromosomes. Female and male map lengths did not differ significantly. We then bioinformatically mapped these loci to the zebra finch (Taeniopygia guttata) genome and found that linkage groups were conserved between song sparrows and zebra finches. Compared to the zebra finch, marker order within small linkage groups was well conserved, whereas the larger linkage groups showed some intrachromosomal rearrangements. Finally, we show that as expected, recombination frequency between linked loci explained the majority of variation in gametic phase disequilibrium. Yet, there was substantial overlap in gametic phase disequilibrium between pairs of linked and unlinked loci. Given that the microsatellites described here lie on 35 of the 38–40 chromosomes, these markers will be useful for studies in this species, as well as for comparative genomics studies with other species.

opencc-zeroDec 2014View details →
dryad32/100

Data from: First-generation linkage map for the European tree frog (Hyla arborea) with utility in congeneric species

Background: Western Palearctic tree frogs (Hyla arborea group) represent a strong potential for evolutionary and conservation genetic research, so far underexploited due to limited molecular resources. New microsatellite markers have recently been developed for Hyla arborea, with high cross-species utility across the entire circum-Mediterranean radiation. Here we conduct sibship analyses to map available markers for use in future population genetic applications. Findings: We characterized eight linkage groups, including one sex-linked, all showing drastically reduced recombination in males compared to females, as previously documented in this species. Mapping of the new 15 markers to the ~200 My diverged Xenopus tropicalis genome suggests a generally conserved synteny with only one confirmed major chromosome rearrangement. Conclusions: The new microsatellites are representative of several chromosomes of H. arborea that are likely to be conserved across closely-related species. Our linkage map provides an important resource for genetic research in European Hylids, notably for studies of speciation, genome evolution and conservation.

opencc-zeroDec 2013View details →
ClinicalTrials.gov32/100

Mapping of End Stage Renal Disease Genetic Susceptibility in African Americans by Admixture Linkage Disequilibrium

ClinicalTrials.gov study NCT00559767. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Marasmius oreades genome assembly v2: Linkage map & repeat library

Open the record for dataset details and reuse information.

publicMar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record