Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

37

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

37 results for “Genetic assignment”

Learn how ShareScore rates datasets ↗
edi52/100

Genetic assignments for Spring Evolutionary Significant Unit reanalysis, Central Valley Chinook Salmon populations, CA, 2011-2024

Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units (ESU). As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central V

openCC (other)Oct 2025View details →
edi48/100

Chinook Salmon genetic assignments for the Central Valley Project (CVP) and State Water Projects (SWP), Sacramento and San Joaquin Delta Waters, CA, 2024-25

Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley

openCC (other)Oct 2025View details →
zenodo40/100

F I G U R E 4 Assigned ancestry plot using AIC criteria determining optimal K in Riverscape genetics of the orangethroat darter complex

F I G U R E 4 Assigned ancestry plot using AIC criteria determining optimal K of 5. Each color corresponds to a distinct genetic cluster, and each bar represents an individual's assigned ancestry to each cluster. Names above chart correspond to river basin (HUC6). EB—E. burri, EP— E. pulchellum, ES—E. spectabile, and EU—E. uniporum. Glaciated and unglaciated correspond to areas that were/were not covered in ice during the last glacial maximum.

opencc-by-4.0Dec 2023View details →
dryad40/100

Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information

<p class="western">Dispersal has a crucial role determining eco-evolutionary dynamics through both gene flow and population size regulation. However, to study dispersal and its consequences, one must distinguish immigrants from residents. Dispersers can be identified using telemetry, capture-mark-recapture (CMR) methods, or genetic assignment methods. All of these methods have disadvantages, such as, high costs and substantial field efforts needed for telemetry and CMR surveys, and adequate genetic distance required in genetic assignment. In this study, we used genome-wide 200K Single Nucleotide Polymorphism data and two different genetic assignment approaches (GSI_SIM, Bayesian framework; BONE, network-based estimation) to identify the dispersers in a house sparrow (<i>Passer domesticus</i>) metapopulation sampled over 16 years. Our results showed higher assignment accuracy with BONE. Hence, we proceeded to diagnose potential sources of errors in the assignment results from the BONE method due to variation in levels of inter-population genetic differentiation, intra-population genetic variation and sample size. We show that assignment accuracy is high even at low levels of genetic differentiation and that it increases with the proportion of a population that has been sampled. Finally, we highlight that dispersal studies integrating both ecological and genetic data provide robust assessments of the dispersal patterns in natural populations.</p>

opencc-zeroJul 2021View details →
dryad40/100

Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad36/100

Data from: Validating dispersal distances inferred from autoregressive occupancy models with genetic parentage assignments

1.Dispersal distances are commonly inferred from occupancy data but have rarely been validated. Estimating dispersal from occupancy data is further complicated by imperfect detection and the presence of unsurveyed patches. 2.We compared dispersal distances inferred from seven years of occupancy data for 212 wetlands in a metapopulation of the secretive and threatened California black rail (Laterallus jamaicensis coturniculus) to distances between parent-offspring dyads identified with 16 microsatellites. 3.We used a novel autoregressive multi-season occupancy model that accounted for both unsurveyed patches and imperfect detection to quantify patch isolation using buffer radius (BRM) and incidence function (IFM) connectivity measures at 15 scales (1–10, 15, 20, 25, and 30 km). Connectivity measures were then fit as colonization covariates in occupancy models to estimate a model-averaged dispersal distance. 4.As predicted, colonization was more strongly related to connectivity at small spatial scales (&lt; 10 km). AIC weights were greatest at 7 km for BRM and at 4 km for IFM. 5.Model-averaged dispersal distances (BRM = 7.46 km; IFM = 5.48 km) showed good agreement with the mean (± SE) dispersal distance from 23 parent-offspring dyads (5.58 ± 1.92 km), indicating reasonably accurate mean dispersal distances can be inferred from occupancy data when isolation strongly affects colonization.

opencc-zeroDec 2017View details →
dryad36/100

Data from: Inferring the timing of long-distance dispersal between Rail metapopulations using genetic and isotopic assignments

Open the record for dataset details and reuse information.

publicAug 2016View details →
dryad36/100

Data from: Validating dispersal distances inferred from autoregressive occupancy models with genetic parentage assignments

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad36/100

Genetic assignment of individuals to source populations using network estimation tools

Open the record for dataset details and reuse information.

publicNov 2019View details →
dryad32/100

Data from: Genetic assignment of large seizures of elephant ivory reveals Africa's major poaching hotspots

Poaching of elephants is now occurring at rates that threaten African populations with extinction. Identifying the number and location of Africa's major poaching hotspots may assist efforts to end poaching and facilitate recovery of elephant populations. We genetically assign origin to 28 large ivory seizures (≥0.5 tons) made between 1996-2014, also testing assignment accuracy. Results suggest that the major poaching hotspots in Africa may be currently concentrated in as few as two areas. Increasing law enforcement in these two hotspots could help curtail future elephant losses across Africa and disrupt this organized transnational crime.

opencc-zeroDec 2014View details →
dryad32/100

Data from: SNPs reveal a genetic cline across the northeast Atlantic and enable powerful population assignment in the European lobster

Resolving stock structure is crucial for fisheries conservation to ensure that the spatial implementation of management is commensurate with that of biological population units. To address this in the economically important European lobster (Homarus gammarus), genetic structure was explored across the species' range using a small panel of single nucleotide polymorphisms (SNPs) previously isolated from restriction-site associated DNA sequencing; these SNPs were selected to maximise differentiation at a range of both broad- and fine-scales. After quality control and filtering, 1,278 lobsters from 38 sampling sites were genotyped at 79 SNPs. The results revealed a pronounced phylogeographic break between the Atlantic and Mediterranean basins, while structure within the Mediterranean was also apparent, partitioned between lobsters from the central Mediterranean and the Aegean Sea. In addition, a genetic cline across the northeast Atlantic was revealed using both putatively neutral and outlier SNPs, but the precise driver(s) of this clinal pattern –isolation-by-distance, secondary contact, selection across an environmental gradient, or a combination of these factors– remains undetermined. Putatively neutral markers differentiated lobsters from Oosterschelde, an estuary on the Dutch coast, a finding likely explained by past bottlenecks and limited gene flow with adjacent North Sea populations. Building on the findings of our spatial genetic analysis, we were able to test the accuracy of assigning lobsters at various spatial scales, including to basin of origin (Atlantic or Mediterranean), region of origin and sampling location. The predictive model assembled using 79 SNPs correctly assigned 99.7 % of lobsters not used to build the model to their basin of origin, but accuracy decreased to region of origin and again to sampling location. These results are of direct relevance to managers of lobster fisheries and hatcheries, and provide the basis for a genetic tool for tracing the origin of European lobsters in the food supply chain.

opencc-zeroJul 2019View details →
dryad32/100

Data from: Bayesian clustering analyses for genetic assignment and study of hybridization in oaks: effects of asymmetric phylogenies and asymmetric sampling schemes

Bayesian clustering methods have been widely used for studying species delimitation and genetic introgression. In order to test the effect of phylogenetic relationships and sampling scheme on the inferred clustering solution and on the performance of Bayesian clustering analysis, I simulated genotypes of the interfertile oak species Quercus robur, Quercus petraea, and Quercus pubescens and I run analyses using two popular software programs, STRUCTURE and BAPS. First, based on purebred simulations, I compared clustering solutions resulting from different sample size configurations. While clustering solution generally reflected the taxonomic relationships when equal samples of each species were included, spurious partition was inferred by STRUCTURE when some species were represented by larger and others by smaller samples. In very unbalanced configurations, STRUCTURE failed to identify the three species, even if three subpopulations were assumed. By contrast, BAPS could properly identify the three species under any sampling scheme. Second, based on simulations of purebreds and hybrids, I tested the performance of individual assignments with variable number of loci. This analysis showed that STRUCTURE can detect introgressed individuals more efficiently than BAPS. However, BAPS could assign purebreds more efficiently with a lower number of loci. Method performance also depended on phylogenetic relationships. In the case of Q. petraea, Q. pubescens, and their hybrids, method performance was lower due to their phylogenetic affinity. Inclusion of three instead of two species into the analysis led to reduction of performance, and to misclassification of hybrids, which often reflected the phylogenetic affinity between Q. petraea and Q. pubescens.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Genetic assignment with isotopes and habitat suitability (GAIAH), a migratory bird case study

1. Identifying migratory connections across the annual cycle is important for studies of migrant ecology, evolution, and conservation. While recent studies have demonstrated the utility of high-resolution SNP-based genetic markers for identifying population-specific migratory patterns, the accuracy of this approach relative to other intrinsic tagging techniques has not yet been assessed. 2. Here, using a straightforward application of Bayes' Rule, we develop a method for combining inferences from high-resolution genetic markers, stable isotopes, and habitat suitability models, to spatially infer the breeding origin of migrants captured anywhere along their migratory pathway. Using leave-one-out cross validation, we compare the accuracy of this combined approach with the accuracy attained using each source of data independently. 3. Our results indicate that when each method is considered in isolation, the accuracy of genetic assignments far exceeded that of assignments based on stable isotopes or habitat suitability models. However, our joint assignment method consistently resulted in small, but informative increases in accuracy and did help to correct misassignments based on genetic data alone. We demonstrate the utility of the combined method by identifying previously undetectable patterns in the timing of migration in a North American migratory songbird, the Wilson's warbler. 4. Overall, our results support the idea that while genetic data provides the most accurate method for tracking animals using intrinsic markers when each method is considered independently, there is value in combining all three methods. The resulting methods are provided as part of a new computationally-efficient R-package, GAIAH, allowing broad application of our statistical framework to other migratory animal systems.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Applications of random forest feature selection for fine-scale genetic population assignment

Genetic population assignment used to inform wildlife management and conservation efforts requires panels of highly informative genetic markers and sensitive assignment tests. We explored the utility of machine-learning algorithms (random forest, regularized random forest, and guided regularized random forest) compared with FST ranking for selection of single nucleotide polymorphisms (SNP) for fine-scale population assignment. We applied these methods to an unpublished SNP dataset for Atlantic salmon (Salmo salar) and a published SNP data set for Alaskan Chinook salmon (Oncorhynchus tshawytscha). In each species, we identified the minimum panel size required to obtain a self-assignment accuracy of at least 90% using each method to create panels of 50-700 markers Panels of SNPs identified using random forest-based methods performed up to 7.8 and 11.2 percentage points better than FST-selected panels of similar size for the Atlantic salmon and Chinook salmon data, respectively. Self-assignment accuracy ≥90% was obtained with panels of 670 and 384 SNPs for each dataset, respectively, a level of accuracy never reached for these species using FST-selected panels. Our results demonstrate a role for machine-learning approaches in marker selection across large genomic datasets to improve assignment for management and conservation of exploited populations.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Seeds in motion: genetic assignment and hydrodynamic models demonstrate concordant patterns of seagrass dispersal

Movement is fundamental to the ecology and evolutionary dynamics within species. Understanding movement through seed dispersal in the marine environment can be difficult due to the high spatial and temporal variability of ocean currents. We employed a mutually enriching approach of population genetic assignment procedures and dispersal predictions from a hydrodynamic model to overcome this difficulty and quantify the movement of dispersing floating fruit of the temperate seagrass Posidonia australis Hook.f. across coastal waters in southwestern Australia. Dispersing fruit cohorts were collected from the water surface over two consecutive years and seeds were genotyped using microsatellite DNA markers. Likelihood-based genetic assignment tests were used to infer the meadow of origin for seed cohorts and individuals. A three-dimensional hydrodynamic model was coupled with a particle transport model to simulate the movement of fruit at the water surface. Floating fruit cohorts were mainly assigned genetically to the nearest meadow, but significant genetic differentiation between cohort and most-likely meadow of origin suggested a mixed origin. This was confirmed by genetic assignment of individual seeds from the same cohort to multiple meadows. The hydrodynamic model predicted 60% of fruit dispersed within 20 km, but that fruit were physically capable of dispersing beyond the study region. Concordance between these two independent measures of dispersal provide insight into the role of physical transport for long distance dispersal (LDD) of fruit and the consequences for spatial genetic structuring of seagrass meadows.

opencc-zeroDec 2017View details →
zenodo32/100

FIGURE 3 in Genetic and morphological variability among the populations assigned to the genus Tropiocolotes Peters, 1880 (Squamata: Gekkonidae) in south Iran

FIGURE 3. Bayesian inference phylogenetic tree of Tropiocolotes populations in southern Iran using two mtDNA genes (COI and 16S). Tropiocolotes steudneri sensu stricto from Egypt was used as the outgroup. Numbers next to the nodes are the MP and ML bootstrap values and BI posterior probabilities (MP/ML/BI).

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURE 1 in Genetic and morphological variability among the populations assigned to the genus Tropiocolotes Peters, 1880 (Squamata: Gekkonidae) in south Iran

FIGURE 1. Map of southern Iran showing sampling localities for the populations of Tropicolates. Blue circles denote T. naybandensis and red circles denote Tropiocolotes cf. steudneri. The type locality of T. naybandensis is marked with a star.

opennotspecifiedDec 2017View details →
zenodo32/100

FIGURE 2 in Genetic and morphological variability among the populations assigned to the genus Tropiocolotes Peters, 1880 (Squamata: Gekkonidae) in south Iran

FIGURE 2. Ordination of principal component 1 (PC1) against principal component 2 (PC2) for differentiated characters of the genus Tropiocolotes in southern Iran.

opennotspecifiedDec 2017View details →
dryad32/100

Data from: The twenty amino acids are identified by unique numbers assigned to the uracil, cytosine, adenine, and guanine found in the three base positions of the sixty-four messenger RNA genetic codons

<p>A codon's three bases consist of any combination of uracil, cytosine, adenine, or guanine and these encode the twenty amino acids. When the codon's first two bases are given specific values, and those values are multiplied, then the third base of the codon is used during translation only when the product is greater than three. Here we show that those values plus more variables within the ribosomal decoding site results in specific flow values for each of the twenty amino acid groups. These results are demonstrated in a flow chart showing the unidirectional flow which is expected during the translation process. All twenty amino acids can be represented by numbers that describe their relationship to each other and to the decoding site. We anticipate our findings will increase discussion about using a number system to better understand the translation process.</p>

opencc-zeroJan 2024View details →
dryad32/100

Data from: Genetic sex assignment in wild populations using GBS data: a statistical threshold approach

Establishing the sex of individuals in wild systems can be challenging and often requires genetic testing. Genotyping-by-sequencing (GBS) and other reduced representation DNA sequencing (RRS) protocols (e.g., RADseq, ddRAD) have enabled the analysis of genetic data on an unprecedented scale. Here, we present a novel approach for the discovery and statistical validation of sex-specific loci in GBS datasets. We used GBS to genotype 166 New Zealand fur seals (NZFS, Arctocephalus forsteri) of known sex. We retained monomorphic loci as potential sex-specific markers in the locus discovery phase. We then used (i) a sex-specific locus threshold (SSLT) to identify significantly male-specific loci within our dataset and (ii) a significant sex-assignment threshold (SSAT) to confidently assign sex in silico the presence or absence of significantly male-specific loci to individuals in our dataset treated as unknowns (98.9% accuracy for females; 95.8% for males, estimated via cross-validation). Furthermore, we assigned sex to 86 individuals of true unknown sex using our SSAT, and assessed the effect of SSLT adjustments on these assignments. From 90 verified sex-specific loci, we developed a panel of three sex-specific PCR primers that we used to ascertain sex independently of our GBS data, which we show amplify reliably in at least three other pinniped species. Using monomorphic loci normally discarded from large SNP datasets is an effective way to identify robust sex-linked markers for non-model species. Our novel pipeline can be used to identify and statistically validate monomorphic and polymorphic sex-specific markers across a range of species and RRS datasets.

opencc-zeroDec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record