Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

41

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

41 results for “soybean genomics”

Learn how ShareScore rates datasets ↗
zenodo40/100

Supplementary Data - Using landscape genomics to infer genomic regions involved in environmental adaptation of soybean genebank accessions

<p><strong>File: 50K_GenotypesEU_raw_UHOH_SoySNP50K.csv.tgz </strong></p> <p>Genotyping data of SoySNP50k SNP array of 170 European soybean varieties.</p> <p>The array includes 51.955 SNP markers.</p> <p>Genotypes of each variety are in columns and each row is a SNP marker. Naming of markers follows the annotation of the soybean genome.</p> <p><strong>File: EUvarieties_infos.csv </strong></p> <p>Description of European varieties</p> <p>Contains variety name, country of origin, EU region and maturity group assignment.</p> <p>&nbsp;</p> <p><strong>File: Supplementary_Data_Haupt_Schmid.xlsx</strong></p> <p>Additional data derived from data analysis. Description of data contained within file (Worksheet &quot;Summary&quot;)</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Improved genome assembly and annotation of the soybean aphid (Aphis glycines Matsumura)

<p>Updated genome assembly and annotation of <em>Aphis&nbsp;glycines</em> biotype 4.</p> <p><strong>Overview of files included in this release:</strong></p> <p><strong>Frozen release:</strong></p> <p>Updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gz&nbsp;</p> <p>BRAKER2 gene models for updated <em>A. glycines </em>biotype 4 genome assembly: Aphis_glycines_4.v2.1.scaffolds.fa.gff</p> <p>BRAKER2 protein sequences:&nbsp;Aphis_glycines_4.v2.1.scaffolds.fa.gff.aa.fa</p> <p>BRAKER2 nucleotide coding sequences:&nbsp;&nbsp;Aphis_glycines_4.v2.1.scaffolds.fa.gff.CDS.fa</p> <p><strong>Unfiltered raw intermediate genome assemblies:</strong></p> <p>Canu assembly of biotype 4 PacBio data from Wenger et. al. (2017):&nbsp;canu.fa.gz</p> <p>DBG2OLC hybrid assembly of selected biotype 4 MiSeq data and biotype 4 PacBio data from&nbsp;Wenger et. al. (2017):&nbsp;DBG2OLC.fa.gz</p> <p>Merged Canu and DBG2OLC assembly created with quickmerge:&nbsp;quickmerge.fa.gz</p> <p>Pilon polished (2 rounds) quickmerge assembly:&nbsp;quickmerge.pilon_r2.fa.gz</p> <p><strong>Mitochondrial and endosymbiont contigs extracted from the pilon polished quickmerge assembly:&nbsp;</strong></p> <p><em>A. glycines </em>biotype 4 mitochondrial genome:&nbsp;Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4&nbsp;<em>Buchnera aphidicola</em>&nbsp;contigs:&nbsp;Aphis_glycines_4_Buchnera_v1.fa</p> <p><em>A. glycines </em>biotype 4&nbsp;<em>Wolbachia</em> contigs:&nbsp;Aphis_glycines_4_Buchnera_v1.fa</p> <p><strong>Other files:</strong></p> <p>MUSCLE alignment of <em>A. glycines </em>v1, <em>A. glycines </em>biotype 4 v2.1 and <em>Drosophila&nbsp;melanogaster</em> R6.22 Osiris proteins in fasta format:&nbsp;D_mel_v1_v2_osiris.prots.muscle.fasta</p> <p>FastTree Maximum Likelihood phylogeny based on the MUSCLE alignment of Osiris genes in newick format:&nbsp;D_mel_v1_v2_osiris.prots.muscle.FastTree.nwk</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
dryad36/100

Genomic relationships of Glycine remota, a recently discovered perennial relative of soybean, within the legume genus Glycine

<p><span>The legume genus, <em>Glycine</em>, which includes the Asian annual cultivated soybean, also includes a group of Australian perennial species comprising the subgenus <em>Glycine</em>. Because the subgenus <em>Glycine</em> represents the tertiary gene pool for one of the world's most important crops, the group has been the target of collection and study for decades, resulting in a steady growth in the number of formally recognized species, from six in the 1970s to over 20 at present, as well as a number of additional informal taxa. These studies have also produced a system of nuclear diploid "genome groups" corresponding to clades in molecular phylogenies. The aptly named <em>G</em>. <em>remota</em> is known only from a single isolated population in the Kimberley region of northwestern Australia and was named only in 2015. The species is unique within <em>Glycine</em> in having unifoliolate leaves; its discoverers hypothesized that <em>G</em>. <em>remota</em>, if diploid, is related to species of the I-genome that are also native to the Kimberley region. We produced low-coverage short-read genome sequencing data from an herbarium specimen of <em>G</em>. <em>remota</em>. Genome size estimates from the sequencing data suggest that <em>G</em>. <em>remota</em> is a diploid, while ploidy estimation is inconclusive likely due to the history of whole genome duplication in <em>Glycine</em>. Phylogenomic analyses of genome-wide SNPs, as well as phylogenetic analyses of the low copy nuclear gene (histone H3D), the entire ribosomal RNA cistron, and the internal transcribed spacer all placed the species unequivocally in the diploid I-genome clade. A complete plastome sequence was also generated and its placement with a plastome phylogeny is also consistent with membership in the I-genome.</span></p>

opencc-zeroMar 2023View details →
dryad36/100

Genomic relationships of Glycine remota, a recently discovered perennial relative of soybean, within the legume genus Glycine

Open the record for dataset details and reuse information.

publicMar 2023View details →
zenodo32/100

Genome-wide identification and characterization of the soybean SOD Family during alkaline stress

<p>Supplementary materials</p>

opencc-by-4.0Jul 2019View details →
dryad32/100

Incorporation of soil-derived covariates in progeny testing and line selection to enhance genomic prediction accuracy in soybean breeding

<p>The availability of high-dimensional molecular markers has allowed plant breeding programs to maximize their efficiency through the genomic prediction of a phenotype of interest. Yield is a highly complex and quantitative trait whose expression is sensitive to environmental stimuli. In this research, we investigated the potential of incorporating soil texture and its interaction with molecular markers through covariance structures to enhance predictive ability. A total of 797 advanced soybean breeding lines derived from 367 unique bi-parental populations were genotyped using the Illumina Infinium BARCSoySNP6K BeadChip and tested for yield for five years in Tiptonville silt loam, Sharkey clay, and Malden fine sand environments. Four statistical models were considered, including a default GBLUP model (M1), a reaction norm model (M2) accounting for the interaction between molecular markers and the environment (GE), an expansion of M2 including soil type (S), and the interaction between soil type and molecular markers (GS) (M3), and an alternative version of M3 without the GE term. Four cross-validation scenarios simulating progeny testing and line selection were implemented (CV2, CV1, CV0, and CV00). Across environments, the addition of GS in M3 decreased the amount of variability captured by both the environment (-30.4%) and residual (-39.2%) terms as compared to M1. Within environments, the GS term in M3 reduced the variability captured by the residual term by roughly 60% and 30% when compared to M1 and M2, respectively. M3 outperformed all models in CV2 (0.577), CV1 (0.480), and CV0 (0.488). The addition of soil texture seems to structure the environment term revealing its components that could enhance or hinder the predictability of a model. The availability of soil texture before the growing season may maximize the functionality of covariance structures, particularly in scenarios with untested genotypes in untested environments. Genomic selection can optimize the efficiency of a soybean breeding program by allowing the reconsideration of field experimental design, allocation of resources, reduction of preliminary trials, and shortening of the breeding cycle.</p>

opencc-zeroOct 2022View details →
dryad32/100

Data from: Chromosome-level reference genome of X12, a highly virulent race of the soybean cyst nematode Heterodera glycines

Open the record for dataset details and reuse information.

publicJul 2019View details →
dryad32/100

Incorporation of soil-derived covariates in progeny testing and line selection to enhance genomic prediction accuracy in soybean breeding

Open the record for dataset details and reuse information.

publicOct 2022View details →
zenodo28/100

The Genome sequences of Calonectria ilicicola (anamorph Cylindrocladium parasiticum) causing Cylindrocladium black rot of peanut and red crown rot of soybean

<p>The fungus <em>Calonectria ilicicola</em> (anamorph <em>Cylindrocladium parasiticum</em>) is an important plant pathogen causing Cylindrocladium black rot (CBR) on peanut and Red crown rot (RCR) on soybean. CBR infection of peanuts cause symptoms such as chlorosis of leaves, blackening of taproots, and wilting, while RCR infection of soybeans lead to root and interveinal necrosis of soybean. In the present study, we sequenced the genome of four CBR-related<em> Ca. ilicicola</em> strains and three RCR-related<em> Ca. ilicicola</em> strains. The draft genome of <em>Ca. ilicicola</em>, ranged from 68.63 Mb to 70.35 Mb in genome size, containing 18536 to 19361 protein-coding genes. The described genome sequences will provide insights into factors that contribute to pathogenicity toward peanut and soybean and will be useful for future research in population genomics and molecular diagnostic marker development to quickly detect this pathogen.</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Structural equation models to interpret genome-wide association studies for morphological and productive traits in soybean [Glycine max (L.)]

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
dryad28/100

Data from: Genome-wide analysis and endo-β-mannanase gene families expression profiling of tomato and soybean

Open the record for dataset details and reuse information.

publicMar 2018View details →
geo24/100

Comprehensive analyses of microRNA gene evolution in paleopolyploid soybean genome

GEO Series GSE72902. Glycine max. 23 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenSep 2015View details →
geo24/100

Whole genome-wide transcript profiling to identify differentially expressed genes associated with seed field emergence in two soybean low phytate mutants

GEO Series GSE83152. Glycine max. 18 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2016View details →
geo24/100

A genome-wide view of transcriptional responses during Aphis glycines Matsumura infestation in soybean

GEO Series GSE141720. Glycine max. 54 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2020View details →
geo24/100

Dynamic changes in genome-wide histone methylation and gene expression of soybean roots in response to salt stress (RNA-seq dataset)

GEO Series GSE133574. Glycine max. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2019View details →
geo24/100

Genome-wide MNase hypersensitive sites in soybean

GEO Series GSE167578. Glycine max. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenOct 2023View details →
geo24/100

Genome-Wide Transcript Profiling During Soybean Seed Development and Throughout the Soybean Life Cycle

GEO Series GSE29163. Glycine max. 12 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2011View details →
geo24/100

Genome-wide transcriptome analyses of developing seeds from low and normal phytic acid soybean lines

GEO Series GSE75575. Glycine max. 30 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2016View details →
geo24/100

Genome-wide identification of binding sites for NAC and YABBY transcription factors and co-regulated genes during soybean seedling development by ChIP-Seq and RNA-Seq.

GEO Series GSE42422. Glycine max. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2013View details →
geo24/100

The composition and origins of intravarietal genomic heterogeneity in soybean

GEO Series GSE25294. Glycine max. 6 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenDec 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record