Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15,277
datasets available to search
ShareScore release 0.9.0
Dataset results
15,277 results for “All Genetics”
Using machine learning to integrate genetic and environmental data to model genotype-by-environment interactions
<p>Files generated from the study described in <a href="https://doi.org/10.1101/2024.02.08.579534">Fernandes et. al (2024)</a> .</p> <p>The file "cvs_h2s.csv" comprises the coefficient of variation and the Cullis heritability for each environment.</p> <p>The file "all_predictions.csv" contains the predictions from all the models evaluated, in different cross-validation (CV) scenarios.</p> <p>The file "coincidence_index.csv" has the Coincidence Index (CI) for each CV and models evaluated in our study.</p> <p>Our study used the multi-environment maize yield trials data from the Genomes to Fields 2022 initiative (<a href="https://doi.org/10.1186/s13104-023-06421-z">Lima et. al 2024</a>).</p>
Data from Neutral genetic structuring of pathogen populations during rapid adaptation
<p><strong>Datasets and temporary dataframes relating to the article "Neutral genetic structuring of pathogen populations during rapid adaptation".</strong></p> <p>These datasets and temporary dataframes are necessary to run the scripts from the public GitLab repository: <a href="https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation">https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation</a>. Please refer to this public GitLab repository for the latest version of the codes and to perform all analyses presented in the article.</p> <p>Original datasets from the demogenetic model:</p> <ul> <li>Output_RandomDesign.txt</li> <li>Output_RegularDesign_With_host_alternation.txt</li> <li>Output_RegularDesign_Without_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_With_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_Without_host_alternation.txt</li> </ul> <p>All remaining files correspond to temporary dataframes generated by the scripts in the GitLab repository, provided here for reproducibility of the results and to save time at certain time-consuming scripts.</p>
Frictionless Tabular Data Package for GC-MS data from the 'Rose Genome' article published in Nature genetics, June, 2018
<p>This dataset, in the form of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable <a href="https://github.com/ISA-tools/stato">STATO</a> terms. </p> <p>The data was extracted from a supplementary material table, available from <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a> and published alongside the Nature Genetics manuscript identified by the following doi: <a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018. This supplementary material table was deposited to Zenodo and is identified by the following doi: <a href="https://doi.org/10.5281/zenodo.2598799">https://doi.org/10.5281/zenodo.2598799</a></p> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project: <a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a> with all the necessary information, executable code and tutorials in the form of Jupyter notebooks.</p>
Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015
<p>This dataset, in the form of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable <a href="https://github.com/ISA-tools/stato">STATO</a> terms. </p> <p>The data were extracted from:</p> <ul> <li>a supplementary material table, available from <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a> and published alongside the Nature Genetics manuscript identified by the following doi: <a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018</li> <li>a supplementary material table available as a pdf from "Biosynthesis of monoterpene scent compounds in roses" by Magnard et al, Science 03 Jul 2015 identified by the following doi: <a href="https://doi.org/10.1126/science.aab0696">https://doi.org/10.1126/science.aab0696</a></li> </ul> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project: <a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a> with all the necessary information, executable code and tutorials in the form of Jupyter notebooks.</p> <p> </p>
Frictionless Tabular data package for GC-MS data from Rose Genome article published in Nature genetics, June, 2018
<p>This dataset, in the form of a Frictionless Tabular Data Package (https://frictionlessdata.io/specs/tabular-data-package/), holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxId) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The data was extracted from a supplementary material table, available from https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip and published alongside the Nature Genetics manuscript identified by the following doi: https://doi.org/10.1038/s41588-018-0110-3, published in June 2018. This dataset is used to demonstrate how to make data Findeable, Accessible, Discoverable and Interoperable(FAIR) and how Tabular Data Package representations can be easily mobilized for re-analysis and data science. It is associated to the following project available from github at: https://github.com/proccaserra/rose2018ng-notebook with all necessary information and Jupyter notebooks.</p>
Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024
<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (“full” matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (“reduced” matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the “new technology” search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping. </p>
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
Role of information in consumers' preferences for eco-sustainable genetic improvements in plant breeding - DATASET
<p>Data-set and variables description related to the paper titled “Role of information in consumers’ preferences for eco-sustainable genetic improvements in plant breeding“, by Massimiliano Borrello, Luigi Cembalo, Riccardo Vecchio. PLOS-ONE, 2021. DOI: 10.1371/journal.pone.0255130</p>
Solutions and Genetic algorithm dataset of the Scenarios used for the Validation of the Conflict Detection and Resolution Use Case (ARTIMATION )
<p>This dataset contains the <strong>solution </strong>of the scenarios used for one of the validation of the ARTIMATION project: Conflict Detection and Resolution (CD&R) use case (link).</p> <p>The solution are computed by a Genetic Algorithm developped by Nicolas Durand.<br> <br> Inside, one can find:</p> <p>-One archive, "GA_Scenario_Solution_Dataset.zip", containing 10 couple of files (so 20 files). Each couple of file "sol_X_1.csv" and "sols_X_1.csv" are reciprocally the solutino given by the Genetic Algorithm to scenario X, and all the candidate solution explroed by the GA while solving scenario X. This archive also contain other versions of the solutions made by the GA with other parameters.<br> <br> -One archive, "GA_Toy_Dataset.zip" , containing solution to random scenarios, used to develop the first interfaces.</p> <p>Those solutions are used to developp the heatmatrix and heatmaps of the project (link), and visualisations for the validation (link).</p>
Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat
<p>SNPs obtained by UNEAK pipeline for <em>Habromys schmidlyi </em>and <em>Reithrodontomys microdon</em>. </p> <p>Pleae cite as: </p> <p>Colunga-Salas P., T Marines-Macías, G Hernández-Canchola, S Barbosa, C Ramírez, JB Searle, L León-Paniagua. 2022. <strong>Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat</strong>. Mammalian Reasearch. Doi: 10.1007/s13364-022-00667-x</p>
Data to reproduce analysis in "Systematic analysis of transcriptional and epigenetic effects of genetic variation in Kupffer cells enables discrimination of cell intrinsic and environment-dependent mechanisms"
<p>Here you can find the datasets necessary to reproduce all analyses described in the Glass lab paper by <a href="https://www.biorxiv.org/content/10.1101/2022.09.22.509046v1">Bennett et al</a>. The python and R code for reproducing analysis and figures can be found on our linked <a href="https://github.com/HunterBennett/KupfferCell_NaturalGeneticVariation">github repository.</a></p> <p>Briefly, this paper explores the effect of natural genetic variation <em>in vivo</em>, using Kupffer cells as a model cell type. We collect and analyze transcriptional and epigenetic data (ATAC-seq, H3K27Ac ChIP-seq) to identify putative <em>trans</em> regulators driving differential gene expression across inbred strains of mice. Additionally, we provide evidence that <em>trans</em> effects control a majority of strain differential genes at homeostasis while <em>cis</em> effects dominate the transcriptional response to an external signal (lipopolysaccharide).</p> <p>References:</p> <p>Hunter Bennett, Ty D. Troutman, Enchen Zhou, Nathanael J. Spann, Verena M. Link, Jason S. Seidman, Christian K. Nickl, Yohei Abe, Mashito Sakai, Martina P. Pasillas, Justin M. Marlman, Carlos Guzman, Mojgan Hosseini, Bernd Schnabl, Christopher K. Glass bioRxiv 2022.09.22.509046; doi: <a href="https://doi.org/10.1101/2022.09.22.509046">https://doi.org/10.1101/2022.09.22.509046</a></p> <p> </p>
Shared genetic factors between stress-related disorders and cardiovascular disease
<p>The ultimate goal of this study is to advance our understanding of the biological mechanisms of stress-related disorders and CVD, through demonstrating pleiotropic genes and pathways underlying their comorbidity that are potentially testable as targets of future interventions in experimental investigations.</p>
Supplementary data: Agro-morphological and molecular characterization reveal deep insights in promising genetic diversity and marker-trait associations in Fagopyrum esculentum and F. tataricum
<p>Our study focuses on the global/European buckwheat germplasm collected as part of the ECOBREDD project. The potential of this highly diverse collection for organic buckwheat breeding was evaluated at two complementary levels: phenotypic and genetic. Here, we characterized the phenotypic and genetic diversity of a global collection of the two cultivated buckwheat species <em>Fagopyrum esculentum</em> and <em>F. tataricum</em> (190 and 51 accessions, respectively) using 37 agro-morphological traits and 24 SSR markers (Simple Sequence Repeats) (see publication and info sheet of the data).</p>
Central Valley Project, Genetic Determination of Population of Origin 2011-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Knights Landing, California Department of Fish and Wildlife, Genetic Determination of Population of Origin 2017 through 2019
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Sacramento trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Chipps Island trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Delta Smelt (Hypomesus transpacificus) biomarker and genetic data from supplemental release into the San Francisco Estuary, 2022
Delta Smelt (Hypomesus transpacificus) is an endangered fish that is endemic to the San Francisco Estuary. As a conservation strategy, hatchery-reared Delta Smelt have been released into the San Francisco Estuary to supplement the wild population. State and federal agencies surveyed the abundance of Delta Smelt, collected fish specimens, and recorded associated environmental data from sampling sites. Delta Smelt specimens were preserved and transported to the University of California, Davis where a variety of biomarkers were assessed on individual fish. This project, including the supplementation of hatchery-reared Delta Smelt and data collection, is ongoing.
Chinook Salmon genetic assignments for the Central Valley Project (CVP) and State Water Projects (SWP), Sacramento and San Joaquin Delta Waters, CA, 2024-25
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley
Genetic characterization of 24 Angus × Hereford cows from the Jornada Experimental Range, Las Cruces, NM, USA
The southwestern US is increasingly facing dry and variable climate conditions, requiring beef operations to adopt novel strategies to meet these emerging challenges. One potential approach is the use of locally adapted cattle breeds or biotypes. A distinctive Angus x Hereford (AH) research herd at the USDA Agricultural Research Service Jornada Experimental Range provides an opportunity to explore the genetic makeup of a desert-adapted cattle herd bred for over four decades under the extreme and harsh conditions of New Mexico’s Chihuahuan Desert. The objective of this study was to analyze the population structure, genetic diversity and signatures of selection of the AH research herd (n = 24). All cows were genotyped using a 64K SNP chip. Principal component and admixture analyses confirmed the mixed genetic background of the AH cows, predominantly of Angus ancestry. The heterozygosity level, effective population size, and inbreeding coefficient indicated that the AH cows maintain moderate genetic diversity and inbreeding levels. Genomic regions under positive selection revealed genes and Quantitative Trait Loci associated with beneficial carcass traits, milk composition, fertility, body homeostasis, antioxidant activity, immune response, and terrain utilization. This research herd could potentially serve as a valuable genetic resource for improving the adaptability and productivity of commercial beef cattle in harsh semi-arid and arid environments, balancing hardiness and performance.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.