Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
199
datasets available to search
ShareScore release 0.7.1
Dataset results
199 results for “Reference genome”
bksnake_reference_genomes_2022-07-15_hg38
<p>Human hg38 reference genomes and annotations files from RefSeq/NCBI and Ensembl. Archived in a tar.gz file.</p> <p>Files are processed and formatted so that they can be used together with the bulk RNASeq pipeline "bksnake".</p> <p>STAR index files for the genome file are also included.</p> <p>The tar.gz file contains the following data folders</p> <ul> <li>genomes_2022-07-15_hg38/fasta</li> <li>genomes_2022-07-15_hg38/star_2.7.10b</li> <li>genomes_2022-07-15_hg38/gtf/ensembl</li> <li>genomes_2022-07-15_hg38/gtf/refseq</li> </ul> <p>"fasta" contains fasta genomic sequences and index files, as well as ribosomal transcripts "intervals" and chromosome "dictionary" files. "star_2.7.10b" contains STAR aligner index files. "gtf/ensembl" and "gtf/refseq" contain Ensembl and RefSeq gtf files as well as "genePred" and "refFlat" files.</p> <p>The recipe how to generate these files is going to be published in github.</p> <p>The bulk RNASeq pipeline "bksnake" is going to be published in github.</p>
Genomic datasets of Laminaria digitata: Paired-end reads from dd-RADseq, reference genome assembly and filtered VCF
<p>The long-term persistence of species in the face of climate change can be evaluated by examining the interplay between selection and genetic drift in the contemporary evolution of populations. In this study, we focused on spatial and temporal genetic variation in four populations of the cold-water kelp Laminaria digitata using thousands of SNPs (ddRAD-seq). These populations were sampled from the center to the south margin in the North Atlantic at two different time points, spanning at least two generations. By conducting genome scans for local adaptation from a single time point, we successfully identified candidate loci that exhibited clinal variation, closely aligned with the latitudinal changes in temperature. This finding suggests that temperature may drive the adaptive response of kelp populations, although other factors, such as the species' demographic history should be considered. Furthermore, we provided compelling evidence of selection through the examination of allele frequency changes over time, by taking into the impact of genetic drift. Specifically, we detected candidate loci exhibiting temporal differentiation that surpassed the levels typically attributed to genetic drift at the south margin, confirmed through simulations. This finding was in sharp contrast with the lack of detection of outlier loci based on temporal differentiation in a population from the North Sea, exhibiting low and decreasing levels of genetic diversity. These contrasting evolutionary scenarios among populations can be primarily attributed to the differential prevalence of selection relative to genetic drift. In conclusion, our study highlights the potential of temporal genomics to gain deeper insights into the contemporary evolution of marine foundation species in response to rapid environmental changes.</p>
Mitochondrial reference genome for Teleopsis dalmanni
Open the record for dataset details and reuse information.
Automated improvement of stickleback reference genome assemblies with Lep-Anchor software
Open the record for dataset details and reuse information.
Reference genome choice and filtering thresholds jointly influence phylogenomic analyses
Open the record for dataset details and reuse information.
Variant calling in the Goldilocks Zone: how reference genome choice and read mapping stringency impact heterozygosity estimates and phylogenetic analyses
Open the record for dataset details and reuse information.
Data from: Refinement of the Antarctic fur seal (Arctocephalus gazella) reference genome increases continuity and completeness
Open the record for dataset details and reuse information.
Genomic datasets of Laminaria digitata: Paired-end reads from dd-RADseq, reference genome assembly and filtered VCF
Open the record for dataset details and reuse information.
Data from: De novo reference genome of a Geomyid rodent, Botta’s pocket gopher (Thomomys bottae)
Open the record for dataset details and reuse information.
Updated functional annotation of the Mycobacterium bovis AF2122/97 reference genome - datasets
<p>This is an archive of the repository held under https://github.com/dmnfarrell/gordon-group/tree/master/mbovis_annotation</p> <p>It contains a notebook and required input files for updating the MTB/Mbovis AF2122/97 genomes with new protein product annotations from literature.</p> <p>D Farrell UCD February 2020</p> <p>Updates to M.bovis genome (Aug 2019):<br> added 611 product annotations<br> added 689 gene names<br> added 5 manually edited entries from pdb hits<br> removed locus_tags for repeat_region features<br> added prefix to tRNA and mobile_element features for consistency</p> <p><br> References<br> Updated functional annotation of the Mycobacterium bovis AF2122/97 reference genome (https://www.biorxiv.org/content/10.1101/757823v1)</p>
Data From: TERRA-REF, An open reference data set from high resolution genomics, phenomics, and imaging sensors
<p>The ARPA-E funded TERRA-REF project is generating open-access reference datasets for the study of plant sensing, genomics, and phenomics. Sensor data were generated by a field scanner sensing platform that captures color, thermal, hyperspectral, and active flourescence imagery as well as three dimensional structure and associated environmental measurements. This dataset is provided alongside data collected using traditional field methods in order to support calibration and validation of algorithms used to extract plot level phenotypes from these datasets.</p> <p>Data were collected at the University of Arizona Maricopa Agricultural Center in Maricopa, Arizona. <br> This site hosts a large field scanner with fifteen sensors, many of which are capable of capturing mm-scale images and point clouds at daily to weekly intervals.</p> <p>These data are intended to be re-used, and are accessible as a combination of files and databases linked by spatial, temporal, and genomic information. In addition to providing open access data, the entire computational pipeline is open source, and we enable users to access high-performance computing environments.</p> <p>The study has evaluated a sorghum diversity panel, biparental cross populations, and elite lines and hybrids from structured sorghum breeding populations. <br> In addition, a durum wheat diversity panel was grown and evaluated over three winter seasons.<br> The initial release includes derived data from from two seasons in which the sorghum diversity panel was evaluated.<br> Future releases will include data from additional seasons and locations.</p> <p>The TERRA-REF reference dataset can be used to characterize phenotype-to-genotype associations, on a genomic scale, that will enable knowledge-driven breeding and the development of higher-yielding cultivars of sorghum and wheat. <br> The data is also being used to develop new algorithms for machine learning, image analysis, genomics, and optical sensor engineering.</p>
Rare variant replaced Korea reference genome fasta
<p>The rare variants of Korea reference genome fasta were replaced with using 396 Korean vcf information</p>
A chromosome-scale reference genome and genome-wide genetic variations elucidate adaptation in yak
<p>Yak is an important livestock for the people who lived in harsh and oxygen-deprived Qinghai-Tibetan Plateau and Hindu-Kush Himalayan Mountains. Although there is a yak genome be sequenced in 2012, the assembly is quite fragmented due to the limitation of Illumina sequencing technology. An accurate and complete reference genome is critical for studying genetic variation of a specie. Long-read sequences are more complete than short-read ones, and they have been successfully used for high-quality genome assembly in several species. Here, we present a high-quality assembly of the yak genome (PB_v1.0) at chromosome scale, which was constructed using long-read sequencing technology assisted by chromatin interaction technology. Compared to the previous yak genome assembly (BosGru_v2.0), the PB_v1.0 assembly has substantially improved chromosome sequence continuity, minimized repetitive structure ambiguity, and achieved gene model completeness. To intensively characterize genetic variation of yak, we generated de novo genome assemblies based on Illumina short reads of seven recognized domestic yak breeds from Tibet and Sichuan as well as one wild yak from Hoh Xil. By comparing these eight assemblies to the PB_v1.0 genome, we obtained a comprehensive map of yak genetic diversity at whole genome level and identified a few protein-coding genes that were absent from the PB_v1.0 assembly. Although wild yak suffered bottleneck effect, the genetic diversity of wild yak is still higher than that of domestic yak. By whole genome alignment, we identified breed-specific sequences and genes, this will help the breeds identification of yak.</p>
Data from: Hare pseudo-reference genome from: the genomic impact of historical hybridization with massive mitochondrial DNA introgression
<p><b>Background:</b> The extent to which selection determines interspecific patterns of genetic exchanges enlightens the role of adaptation in evolution and speciation. Often reported extensive interspecific introgression could be selection-driven, but also result from demographic processes, especially in cases of invasive species replacements, which can promote introgression at their front. Because invasion and selective sweeps similarly mold variation, population genetics evidence for selection can only be gathered in an explicit demographic framework. The Iberian hare, <i>Lepus granatensis</i>, displays in its northern range extensive mitochondrial DNA introgression from <i>L. timidus</i>, an arctic/boreal species that it replaced locally after the last glacial maximum. We use whole-genome sequencing to infer geographic and genomic patterns of nuclear introgression and fit a neutral model of species replacement with hybridization, allowing us to evaluate how selection influenced introgression genome-wide, including for mtDNA.</p> <p><b>Results:</b> Although the average nuclear and mtDNA introgression patterns are strongly contrasted, they fit a single neutral model of post-glacial invasive replacement of <i>timidus</i> by <i>granatensis</i>. Outliers of elevated introgression include several genes related to immunity, spermatogenesis, and mitochondrial metabolism. Introgression is reduced on the X-chromosome and in low recombining regions.</p> <p><b>Conclusion:</b> General nuclear and mtDNA patterns of introgression can be explained by purely demographic processes. Hybrid incompatibilities and interplay between selection and recombination locally modulate levels of nuclear introgression. Selection promoted introgression of some genes involved in conflicts, either interspecific (parasites) or possibly cytonuclear. In the latter case, nuclear introgression could mitigate the potential negative effects of alien mtDNA on mitochondrial metabolism and male-specific traits.</p>
Mash Sketch of RefSeq Bacterial Reference Genomes
<p>The mash reference that can be downloaded from <a href="https://mash.readthedocs.io/en/latest/data.html">the mash documentaion</a> is for RefSeq version 70.</p> <p>I do not inherently have a problem with RefSeq version 70, but RefSeq is well past version 200 now. </p> <p>RefSeq updates four times year, and I needed an easy way to create and distribute a mash sketch file of the representative bacterial/prokaryotic genomes.<br><br>This is intended to be a place to hold the mash sketches from <a href="https://github.com/erinyoung/update_mash_dist">https://github.com/erinyoung/update_mash_dist</a>.<br><br>The mash sketch file from erinyoung/update_mash_dist requires git lfs to be installed when cloning the repository, which is cumbersome for some users.<br><br>The update requency is intended to mirror that of RefSeq (i.e. 4 time a year), but... is likely to be less frequent than that.<br><br>Don't hesitate to <a href="https://github.com/erinyoung/update_mash_dist/issues">submit an issue</a> if this needs to get updated.<br><br>I do have some prior zenodo repositories (https://zenodo.org/records/10519852 , https://zenodo.org/records/7887021 , and https://zenodo.org/records/7348463 ) which hold the same mash sketch reference, but the refseq version is in the title. I'd rather have one repository that gets updated rather than create new repositories each time.<br><br>This is how the mash reference file was created:<br><br></p> <pre><code># Step 1. Download Datasets and Dataformat </code></pre> <pre><code>wget https://ftp.ncbi.nlm.nih.gov/pub/datasets/command-line/v2/linux-amd64/datasets wget https://ftp.ncbi.nlm.nih.gov/pub/datasets/command-line/v2/linux-amd64/dataformat chmod +x datasets dataformat</code></pre> <pre><code> # Step 2. Download Mash <br></code></pre> <pre><code> wget https://github.com/marbl/Mash/releases/download/v2.3/mash-Linux64-v2.3.tar tar -xvf mash-Linux64-v2.3.tar </code></pre> <pre><code><br> # Step 3. Get a list of all the genomes # Note: this also changes how some of the names are represented datasets summary genome taxon bacteria --reference --as-json-lines | \ dataformat tsv genome --fields accession,organism-name --elide-header | \ sed 's/\[//g' | \ sed 's/\]//g' | \ sed 's/["'\'']//g' | \ sed 's/endosymbiont of /endosymbiont_of_/g' > \ ids.txt # Step 4. Download the reference files and sketch them # Note: Since this is done in Github Actions (GA), I need to keep everything below 30G. # The best way to do this is to download the process each reference file individually, and then combine it to the whole. # This obviously does not need to be followed if not under those same limitations. while read line do id=$(echo $line | awk '{print $1}') ge=$(echo $line | awk '{print $2}') if [ ! -n "$ge" ] ; then ge="unknown" ; fi sp=$(echo $line | awk '{print $3}') if [ ! -n "$sp" ] ; then sp="unknown" ; fi datasets download genome accession $id unzip ncbi_dataset.zip cp ncbi_dataset/data/*/*_genomic.fna ${ge}_${sp}_${id}.fasta if [ ! -f RefSeqSketches_${version}.msh ] then mash sketch ${ge}_${sp}_${id}.fasta -o RefSeqSketches_${version} else mash sketch ${ge}_${sp}_${id}.fasta -o ${ge}_${sp}_${id} mv RefSeqSketches_${version}.msh tmp.msh mash paste RefSeqSketches_${version} tmp.msh ${ge}_${sp}_${id}.msh rm tmp.msh ${ge}_${sp}_${id}.msh fi rm ${ge}_${sp}_${id}.fasta rm -rf ncbi_dataset/ rm ncbi_dataset.zip rm README.md rm md5sum.txt done < ids.txt</code></pre> <pre><br><br>To use</pre> <pre><code># download file wget <insert url for file> mask sketch sample.fasta RefSeqSketches_<version>.msh > mash_results.txt # These results are unsorted, so many find it useful to sort them. sort -gk3 mash_results.txt > sorted_mash_results.txt</code></pre> <p> <br>The should look like the following:</p> <pre><code>2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_pyogenes_GCF_900475035.1.fasta 0.0116661 0 643/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_dysgalactiae_GCF_016128095.1.fasta 0.0782587 0 107/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_canis_GCF_900636575.1.fasta 0.132399 2.34894e-153 32/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_agalactiae_GCF_001552035.1.fasta 0.164662 1.32611e-72 16/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_castoreus_GCF_000425025.1.fasta 0.174408 2.34302e-58 13/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_didelphis_GCF_000380005.1.fasta 0.182269 8.30736e-49 11/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_uberis_GCF_900475595.1.fasta 0.186761 5.62934e-44 10/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_iniae_GCF_000831485.1.fasta 0.191731 3.33152e-39 9/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_ictaluri_GCF_000188015.2.fasta 0.197292 1.75608e-34 8/1000 2024CK-00429-UT-M03999-240412_contigs.fa Streptococcus_phocae_GCF_001302265.1.fasta 0.203604 2.46548e-30 7/1000</code></pre> <pre> </pre>
Hybridization dynamics and extensive introgression in the Daphnia longispina species complex: new insights from a high-quality Daphnia galeata reference genome
<p>Supplementary data for the Genome Biology and Evolution paper <a href="http://dx.doi.org/10.1093/gbe/evab267">10.1093/gbe/evab267</a></p>
Transitioning from environmental genetics to genomics using mitogenome reference databases
<p><span>Species detection using eDNA is revolutionizing the global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of complete mitogenomic sequence information to evaluate the effectiveness of such data to differentiate, identify and detect taxa. We created the Oregon Biodiversity Genome Project working group to utilize recent advances in sequencing technology to create a database of complete, near error-free mitogenomic sequences for all of Oregon's resident freshwater fishes. So far, we have successfully assembled the complete mitogenomes of 313 specimens of freshwater fish representing 7 families, 55 genera, and 129 (88%) of the 146 resident species and lineages. Our comparative analyses of these sequences illustrate that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays are not consistently diagnostic for species-level identification and that no single region is best for metabarcoding Oregon's fishes. However, often-overlooked intergenic regions of the mitogenome such as the D-loop have the potential to reliably diagnose and differentiate species. This project provides a blueprint for other researchers to follow as they build regional databases. It also illustrates the taxonomic value and limits of complete mitogenomic sequences, and how current eDNA assays and the "PCR-free" environmental genomics methods of the future can best leverage this information.</span></p>
Landscape connectivity and genetic structure in a mainstem and a tributary stonefly (Plecoptera) species using a novel reference genome
<p>Abstract Understanding how environmental variation influences population genetic structure can help predict how environmental change influences population connectivity, genetic diversity, and evolutionary potential. We used riverscape genomics modelling to investigate how climatic and habitat variables relate to patterns of genetic variation in two stonefly species, one from mainstem river habitats (Sweltsa coloradensis) and one from tributaries (Sweltsa fidelis) in 40 sites in northwest Montana, USA. We produced a draft genome assembly for S. coloradensis (N50 = 0.251 Mbp, BUSCO &gt; 95% using "insecta_ob9" reference genes). We genotyped 1930 SNPs in 372 individuals for S. coloradensis and 520 SNPs in 153 individuals for S. fidelis. We found higher genetic diversity for S. coloradensis compared to S. fidelis, but nearly identical genetic differentiation among sites within each species (both had global loci median FST = 0.000), despite differences in stream network location. For landscape genomics and testing for selection, we produced a less stringently filtered data set (3454 and 1070 SNPs for S. coloradensis and S. fidelis, respectively). Environmental variables (mean summer precipitation, slope, aspect, mean June stream temperature, land cover type) were correlated with 19 putative adaptive loci for S. coloradensis. but there was only one putative adaptive locus for S. fidelis (correlated with aspect). Interestingly, we also detected potential hybridization between multiple Sweltsa species which has never been previously detected. Studies like ours, that test for adaptive variation in multiple related species are needed to help assess landscape connectivity and the vulnerability of populations and communities to environmental change.</p>
RirC3: Rhizophagus irregularis reference genome and annotation
<p>Reference genome assembly of <em>Rhizophagus irregularis</em> isolate C3, made with PacBio Sequel II SMRT sequencing, and polished with Illumina reads. Annotation was produced with FunAnnotate.</p>
A chromosomal-scale reference genome of the New World Screwworm, Cochliomyia hominivorax
<p>The New World Screwworm, <em>Cochliomyia hominivorax</em> (Calliphoridae), is the most important myiasis-causing species in America. Screwworm myiasis is a zoonosis that can cause severe lesions in livestock, domesticated and wild animals, and occasionally in people. Beyond the sanitary problems associated with this species, these infestations negatively impact economic sectors, such as the cattle industry.</p> <p>Here, we present a chromosome-scale assembly of <em>C</em>. <em>hominivorax</em>'s genome, organized in 6 chromosome-length and 515 unplaced scaffolds spanning 534 Mb. There was a clear correspondence between the <em>D</em>. <em>melanogaster</em> linkage groups A-E and the chromosomal-scale scaffolds. Chromosome Quotient (CQ) analysis identified a single scaffold from the X chromosome that contains most of the orthologs of genes that are on the <em>D</em>. <em>melanogaster</em> fourth chromosome (linkage group F or dot chromosome). CQ analysis also identified potential X and Y unplaced scaffolds and genes. Y-linkage for selected regions was confirmed by PCR with male and female DNA. Some of the long chromosome-scale scaffolds include Y-linked sequences, suggesting misassembly of these regions. These resources will provide a basis for future studies aiming at understanding the biology and evolution of this devastating obligate parasite.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.