Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

183

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

183 results for “association mapping”

Learn how ShareScore rates datasets ↗
zenodo52/100

CryoEM Maps and Associated Data Submitted to the 2015/2016 EMDataBank Map Challenge

<p>Files and metadata associated with the EMDataBank/Unified Data Resource for 3DEM 2015/2016 Map Challenge hosted at challenges.emdatabank.org are deposited.</p> <p>All members of the Scientific Community--at all levels of experience--were invited to participate as Challengers, and/or as Assessors.</p> <p>Seven benchmark raw image datasets were selected for the challenge. Six are selected from recently described single particle structure determinations with image data collected as multi-frame movies; one is based on simulated (in silico) images. All of the raw image datasets are archived at pdbe.org/empiar.</p> <p>27 Challengers created 66 single particle reconstructions from the targets, and then uploaded their results with associated details.&nbsp; 15 of the reconstructions were calculated using the SDSC Gordon supercomputer.</p> <p>This map challenge was one of two community-wide challenges sponsored by EMDataBank in 2015/2016 to critically evaluate 3DEM methods that are coming into use, with the ultimate goal of developing validation criteria associated with every 3DEM map and map-derived model.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2017View details →
zenodo48/100

COVID19 Flow-Maps Mobility-Associated-Risk

<p><strong>The Mobility Associated Risk</strong></p> <p>The Mobility Associated Risk is a risk score combines mobility and COVID-19 incidence to estimate how many cases could theoretically be exported/imported between different origin-destination pairs of regions.</p> <p>For more information about how the MAR is calculated visit: <a href="https://flowmaps.life.bsc.es/flowboard/board_what_is_risk#what_is_risk">https://flowmaps.life.bsc.es/flowboard/board_what_is_risk#what_is_risk</a></p> <p>Dashboard The Mobility Associated Risk combines mobility and COVID-19 incidence to estimate how many cases could theoretically be exported/imported between different origin-destination pairs of regions.</p> <p>For more information about how the MAR is calculated visit: <a href="https://flowmaps.life.bsc.es/flowboard/board_what_is_risk#what_is_risk">https://flowmaps.life.bsc.es/flowboard/board_what_is_risk#what_is_risk</a></p> <p>Dashboard <a href="https://flowmaps.life.bsc.es/flowboard/">https://flowmaps.life.bsc.es/flowboard/</a></p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Flood Hazard Maps and Associated Data for Case Study: Funding rules that promote equity in climate adaptation outcomes

<p>Inundation grids for multiple return periods and multiple scenarios. Please see the underlying study for more details about the methods. The data here can be reproduced following the code and instructions at this repository: https://github.com/CoRE-Lab-UCF/Pollack_et_al_2024/tree/main. Also available here: https://doi.org/10.5281/zenodo.14515896.&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Second release of the data associated with the paper entitled 'Cluster-enhanced ensemble learning for mapping global monthly surface ozone from 2003 to 2019'

<p>This is the second release of the data associated with the paper entitled &#39;Cluster-enhanced ensemble learning for mapping global monthly surface ozone from 2003 to 2019&#39;.</p> <p>The paper was published&nbsp;in&nbsp;Geophysical Research Letters. We provide the data that has been smoothed by moving filter&nbsp;and not. The data can be loaded by the <em>raster </em>package in <em>R.</em>&nbsp;Note that the unit is ppmv.</p> <p>Please note that both of these files must be in the same directory to open in <em>R</em> properly<em>.</em></p> <p>Please get in touch with the authors if you have any issues, email: xliu21@smail.nju.edu.cn or wanghk@nju.edu.cn</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Genotypes for Neurospora nested association mapping population

<p>Genotype file that contains genotypes all the Neurospora association mapping population developed in Kronholm lab. See https://github.com/ikron/Neurospora_NAM_population</p> <p>The file is a tab-delimited text file in hapmap format and contains first contains the columns: rs &nbsp; &nbsp;alleles &nbsp; &nbsp;chrom &nbsp; &nbsp;pos &nbsp; &nbsp;strand &nbsp; &nbsp;assembly &nbsp; &nbsp;center &nbsp; &nbsp;protLSID &nbsp; &nbsp;assayLSID &nbsp; &nbsp;panel &nbsp; &nbsp;GCcode &nbsp;&nbsp; and then the subsequent columns are strains identifiers. The column 'rs' is the name of the SNP, column 'alleles' shows which two alternative bases occur, the column 'chrom' is the chromosome, column 'pos' is the coordinate, columns 'strand', 'center', 'protLSID', 'assayLSID', 'panel' and 'GCcode' contain all missing data. They are included for hapmap format compatibility. The column 'assembly' indicates the version of the Neurospora crassa reference genome that the coordinates are based on. And in this case it is NC12 for all SNPs.</p> <p>Missing data is indicated by 'N'</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Data for Linkage mapping of root shape traits associated with market class in two biparental carrot populations

<p>&nbsp;</p> <p>This repository contains essential data to support the findings presented in the forthcoming publication titled &quot;Linkage Mapping of Root Shape Traits Associated with Market Class in Two Biparental Carrot Populations.&quot; It includes VCF files for two distinct carrot biparental populations, as well as R code for filtering, constructing linkage maps, and conducting QTL analysis. Furthermore, the repository hosts phenotypic data gathered from these two biparental populations during the years 2020 and 2021.</p> <p>Two carrot genetic maps, one for each population, have been made available alongside their respective phenotypic data.</p> <p>The provided R code contains absolute working directory paths that may not function as intended on your system. The primary purpose of sharing this code is to offer readers insight into the techniques employed in this study. You may need to adapt the directory paths to suit your specific setup.&nbsp;</p> <p>To assist readers in understanding the logical sequence of steps involved in our linkage mapping project, the R code scripts have been sequentially numbered from 0 to 10.</p> <p>For more info contact: vegaalfaro@wisc.edu.</p>

opencc-by-4.0Oct 2023View details →
edi44/100

Hubbard Brook Experimental Forest: Hyperspectral Foliar N map and associated field data, 2012

A canopy nitrogen map was created for the Hubbard Brook Experimental Forest and watersheds using airborne imaging spectrometer data collected by SpecTIR LLC (Reno, NV) on August 7, 2012, and associated field data. Leaf samples collected in the field were analyzed for nitrogen concentration, scaled to plot (whole canopy) level, and related to airborne imaging spectrometer reflectance data using partial least squares regression modeling to derive spatially explicit estimates of canopy nitrogen concentration (mass-based) for the spatial extent of the airborne imagery.

openCC (other)Feb 2021View details →
zenodo40/100

Fine-scale map reveals highly variable recombination rates associated with genomic features in the Eurasian blackcap

<p>This project estimates historical recombination rates of the Eurasian blackcap&nbsp;(<em>Sylvia atricapilla</em>) and the garden warbler (<em>Sylvia borin)</em>&nbsp;using a linkage disequilibrium-based approach. The population-scaled recombination rates were inferred using 'pyrho' considering the population's demographic information and mutation rates. The mutation rate was taken from a closely related species (flycatcher,&nbsp;<em>Ficedula albicollis</em>).</p> <p>Here we report the recombination rate per site per generation (r, output of pyrho) for all the chromosomes along the genome for both species.&nbsp;</p> <p>Blackcap dataset file: WholeGenome_chromosomes_blackap_n38_Pen20_W50.gz</p> <p>Garden warbler dataset file: WholeGenome_chromosomes_SylBor_n10_Pen20W50.gz</p> <p>&nbsp;</p> <p>Additionally, we annotated the blackcap genome to evaluate the association between recombination rates and specific annotation features (e.g., genes, transcription start sites, promoters, 5'UTR, and 3'UTR regions). The annotation of the genome was based on RNA-seq and Iso-Seq data. It can be found with the name bSylAtr1.1.gff.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Data from: Genome-wide association mapping within a local Arabidopsis thaliana population more fully reveals the genetic architecture for defensive metabolite diversity

<p>A paradoxical finding from genome-wide association studies (GWAS) in plants is that variation in metabolite profiles typically maps to a small number of loci, despite the complexity of underlying biosynthetic pathways. This discrepancy may partially arise from limitations presented by geographically diverse mapping panels. Properties of metabolic pathways that impede GWAS by diluting the additive effect of a causal variant, such as allelic and genic heterogeneity and epistasis, would be expected to increase in severity with the geographic range of the mapping panel. We hypothesized that a population from a single locality would reveal an expanded set of associated loci. We tested this in a French <em>Arabidopsis thaliana</em> population (&lt; 1 km transect) by profiling and conducting GWAS for glucosinolates, a suite of defensive metabolites that have been studied in depth through functional and genetic mapping approaches. For two distinct classes of glucosinolates, we discovered more associations at biosynthetic loci than previous GWAS with continental-scale mapping panels. Candidate genes underlying novel associations were supported by concordance between their observed effects in the TOU-A population and previous functional genetic and biochemical characterization. Local populations complement geographically diverse mapping panels to reveal a more complete genetic architecture for metabolic traits.</p>

opencc-zeroMay 2024View details →
zenodo40/100

Association mapping identified novel candidate loci affecting wood formation in Norway spruce

<p>Data sets associated with the study for the Association mapping and identification of novel candidate loci affecting wood formation in Norway spruce</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Identification of novel genes involved in phosphate accumulation in Lotus japonicus through Genome Wide Association mapping of root system architecture and anion content

<p>130 Lotus japonicus accessions were used. The names and accession numbers are<br> listed in S6 Table. Seeds were scarified with sandpaper and then sterilized 14 minutes in 0.05%<br> sodium hypochlorite. Subsequently, seeds were rinsed and washed 5 times in sterile distilled<br> water. For the germination, seeds were positioned in imbibed filter paper, in sterile Petri dishes,<br> and wrapped in aluminium foil. After 3 days at 21&deg;C, young seedling were transferred to square<br> plates (12 x 12 cm) containing growth medium. Both media used in this<br> study were based on Long-Ashton solution (with two levels of phosphate concentration -20 or<br> 750 &mu;M, LP or HP, respectively) with 0.8% MES buffer (Duchefa Biochemie,<br> Haarlem, The Netherlands), 0.8% agarose (to minimize phosphate contamination), and adjusted<br> to pH 5.7 with 1M KOH. After adding the medium, plates were dried, closed, overnight in a<br> sterile laminar flow hood. Two accessions, with four replicates per each accession, were placed<br> on each plate. Each plate was replicated, with mirrored position of each accession to minimize<br> any positional growth effects. Plates were placed vertically, and plants grown under long-day<br> conditions (21&deg;C, 16 h light/8 h dark cycle) with white light bulbs emitting 50 &mu;mol/m 2 /s and<br> roots were exposed to light. Every day at the same time, the racks were transported to the image<br> acquisition room where images of each plate were acquired with eight Epson V600 CCD flatbed<br> color image scanners (Seiko Epson) and then immediately returned to the growth chamber.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Mapping of solar panels and Fukushima Daiichi Nuclear Power Plant Accident-associated radioactive waste storage in 2022 and 2023, Fukushima, Japan

<p>The policy of reconstruction after the Fukushima Daiichi Nuclear Plant accident has led to a radical transformation of the landscapes of Fukushima Prefecture especially in two aspects (Asanuma-Brice et al., 2023). The first is related to an extensive decontamination policy which resulted in the removal of more than 13 million m<sup>3</sup> of contaminated soil (MOEJ, 2021). The second is the widespread installation of solar panels, which demonstrate the transition decided by the Prefecture and the inhabitants in terms of energy policy.</p> <p>A systematic mapping of these features was carried out from the satellite imagery of Google Map (2023) within the boundaries of Fukushima Prefecture. The objective was to highlight the evolution of specific land use features that are captured imprecisely by automatic detection mapping. We focused on the main visible change in the landscape in terms of land use since Fukushima Daiichi nuclear accident:&nbsp; contaminated waste disposal areas and solar panel fields. These zones were delineated allowing a calculation of the corresponding surface areas (m<sup>2</sup>).<strong> The dataset is composed of 4 shapefile layers: contaminated waste deposits in 2022 and 2023, solar panels in 2022 and 2023. For the year 2022 the last update was conducted in July 2022 and for the year 2023 the last update took place in March 2023.</strong></p> <p>As the land use is in constant and rapid transition (Asanuma-Brice, 2021), we considered as contaminated waste deposits, the permanent storage centers as well as the sites where there are still bags of contaminated waste in varying numbers, knowing that they will be removed and stored on other dedicated sites (Evrard et al., 2019). This choice was made to potentially identify, when the map was updated, the future uses of the land where this waste was stored temporarily.</p> <p>This dataset is part of a larger project that aims to provide the community with an interactive tool (https://mitatelab.cnrs.fr/mitate-labs-map-of-solar-panel-and-contaminated-wasted-land/) that makes available various types of information essential to the analysis of the reconstruction, such as: the delineation of the evacuated zone (which evolved throughout time), the delineation of the municipality boundaries affected by the reconstruction policy, the main services found in these localities, the location of the memorials of the disaster in the zone, as well as the geo-localization of the soil/sediment samples collected by other Mitate lab researchers in order to investigate the redistribution of radionuclides in the environment (Evrard et al., 2021).</p>

opencc-by-4.0Apr 2023View details →
dryad40/100

Data from: Genome-wide association mapping within a local Arabidopsis thaliana population more fully reveals the genetic architecture for defensive metabolite diversity

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Genome-wide association mapping to identify genetic loci for cold tolerance and cold recovery during germination in rice

<p>To investigate the genetic architecture underlying cold tolerance during germination in rice (<i>Oryza sativa</i>), we conducted a genome-wide association study (GWAS) using a novel diversity panel of 257 rice accessions from around the world and 5,185 SNP markers from a 7K SNP marker array. Genotyping was performed using a 7K Illumina iSelect custom-designed array by following the Infinium HD Array Ultra Protocol. The 7K array, called the C7AIR, was designed by Dr. Susan McCouch's Lab at Cornell University and consists of 7,098 SNPs (Morales et al. 2020, under review). After genotyping 257 rice accessions with the 7K array (C7AIR), poor-performing SNP markers (SNPs of call rate &lt;90%; minor allele frequency &lt;5%; or heterozygosity &gt;20%) were removed from the dataset. For our study, a subset of 5,185 high-quality SNP markers obtained after filtering was used to perform the genome-wide association analysis.  The dataset representing the genotype data of 5,185 SNP markers by 257 rice accessions is presented here.</p>

opencc-zeroDec 2018View details →
zenodo36/100

Voxel-wise maps for the paper: Longitudinal associations of magnetic susceptibility with clinical severity in Parkinson's disease

<p>This upload contains voxel-wise group level QSM data, and statistical maps for group level results associated with the paper: Longitudinal associations of magnetic susceptibility with clinical severity in Parkinson's disease.</p> <p>The assocaited code for reproducing these statistical maps can be found here: <a href="https://github.com/gecthomas/QSM_PD_longitudinal">gecthomas/QSM_PD_longitudinal (github.com)</a></p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Data repository associated with 'A Functional Map of the Human Intrinsically Disordered Proteome'

<p><strong>ES_MAP.zip</strong></p> <ul> <li>a hierarchically clustered map of the human IDR-ome</li> <li>.cdt and .gtr files -&nbsp;outputs of Cluster3.0 software</li> <li>can be visualized using JavaTreeView (see Tutorial_ES.pdf)</li> </ul> <p><strong>TUTORIAL.zip</strong>, information on:</p> <ul> <li>visualization and analysis of the human IDR-ome map</li> <li>search for proteins of interest and exploratory analyses of clusters</li> <li>automatic export and analysis of exported clusters (code available at https://github.com/IPritisanac/ES_PW)</li> </ul> <p><strong>IDROME_SEQUENCES.zip</strong></p> <ul> <li>human proteome fasta file</li> <li>IDRome fasta file</li> <li>SPOT-Disorder v1.0 disorder boundaries <ul> <li>13 044 unique protein sequences with at least one IDR (&gt;=30 amino acids)</li> <li>21 252 total unique human IDRs</li> </ul> </li> </ul> <p><strong>IDR_ALN.zip</strong></p> <ul> <li>alignments of IDR sequences across ENSEMBL orthologs</li> <li>19 459 IDR alignments</li> <li>UniProt ID and IDR boundaries for the human sequence are indicated in the name of the file</li> </ul> <p><strong>FAIDR_TSTATS.zip</strong></p> <ul> <li>hierarchical clustering of FAIDR t-statistics for 148 GO terms<br> <ul> <li>.cdt, .gtr files from Cluster3.0</li> <li>can be visualized using JavaTreeView</li> <li>reveals the most predictive molecular features for the top performing 148 models</li> </ul> </li> </ul> <p><strong>CLUSTERS_EXPLORE.zip</strong></p> <ul> <li>clusters obtained through exploratory analysis of the map provided in ES_MAP.zip</li> <li>93 exported clusters in .cdt file format</li> </ul> <p><strong>CLUSTERS_AUTO.zip</strong></p> <ul> <li>clusters extracted from the hierarchically clustered IDR-ome map at a range of distance thresholds (0.4 - 0.8) in .cdt file format</li> <li>distance refers to the uncentered correlation distance between vectors of Z-scores representing human IDRs</li> <li>clusters extracted at different distance thresholds are split into separate archives</li> <li>AUTO_GO_FEATS.xlsx - summary of GO-term overrepresentation and feature enrichment analyses; each distance threshold is in a separate sheet</li> </ul> <p><strong>FAIDR_HIGH_AUC_PPV_GO.zip</strong></p> <ul> <li>target files with annotations of 148 GO terms for which good quality FAIDR models could be obtained (AUC &gt;= 0.7, PPV &gt;= 0.4)</li> <li>file format: three columns; 1st: IDR ID (includes IDR boundaries); 2nd: protein UniProt ID; 3rd: annotation of the protein to a GO term (1 if known to be associated with the GO term, 0 if not)</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Fine-mapping gene-based associations via knockoff analysis of biobank-scale data with applications to UK Biobank

<p>The results of BIGKnock analyses of manuscript&nbsp;&#39;&#39;Fine-mapping gene-based associations via knockoff analysis of biobank-scale data with applications to UK Biobank&#39;&#39;</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Association mapping identified novel candidate loci affecting wood formation in Norway spruce

<p>Genotypic data set for the association mapping in Norway spruce for wood formation and tracheid traits.</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

Bivariate GWA mapping reveals associations between aliphatic glucosinolates and plant responses to thrips and heat stress

<p>Supplemental data on bivariate GWA mapping of stress phenotypes and metabolomes of Arabidopsis.</p>

opencc-by-4.0Aug 2024View details →
dryad36/100

Data from: Integrating Bayesian genomic cline analyses and association mapping of morphological and ecological traits to dissect reproductive isolation and introgression in a Louisiana Iris hybrid zone

Hybrid zones provide unique opportunities to examine reproductive isolation and introgression in nature. We utilized 45,384 Single Nucleotide Polymorphism (SNP) loci to perform association mapping of 14 floral, vegetative, and ecological traits that differ between Iris hexagona and Iris fulva, and to investigate, using a Bayesian Genomic Cline (BGC) framework, patterns of genomic introgression in a large and phenotypically diverse hybrid zone in southern Louisiana. Many loci of small effect-size were consistently found to be associated with phenotypic variation across all traits, and several individual loci were revealed to influence phenotypic variation across multiple traits. Patterns of genomic introgression were quite heterogeneous throughout the Louisiana Iris genome, with I. hexagona alleles tending to be favored over those of I. fulva. Loci that were found to have exceptional patterns of introgression were also found to be significantly associated with phenotypic variation in a small number of morphological traits. However, this was the exception rather than the rule, as most loci that were associated with morphological trait variation were not significantly associated with excess ancestry. These findings provide insights into the complexity of the genomic architecture of phenotypic differences and are a first step towards identifying loci that are associated with both trait variation and reproductive isolation in nature.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record