Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

167

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

167 results for “coding sequences”

Learn how ShareScore rates datasets ↗
zenodo48/100

Strong sequence dependence in RNA/DNA hybrid strand displacement kinetics supplementary data and code

<p>Supplementary data and code needed to replicate figures and results for the paper: Strong sequence-dependence in RNA/DNA hybrid strand displacement kinetics - Francesca G. Smith, John P. Goertz, Molly M. Stevens and Thomas E. Ouldridge. README is included to explain each folder and file in the repository.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

FASTA file containing the MYB encoding gene An2-like and Ant1 coding sequences corresponding to wild and cultivated tomato accessions

<p>The coding sequence (CDS) of the MYB encoding genes&nbsp;<em>Ant1</em> and <em>An2-like</em>.&nbsp;Sequences were retrieved from regions corresponding to the<em> Aft</em> locus from <em>Solanum galapagense </em>accession&nbsp;LA1141, <em>S. lycopersicum</em> variety OH8245, and&nbsp;&nbsp;84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences were compared to available&nbsp;CDS available from the Sol genomics network (SGN) and&nbsp;the National Center for Biotechnology Information. The CDS was&nbsp;retrieved from <em>S. lycopersicum</em>&nbsp;variety&nbsp;Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum</em> accession LA1996 [MN242011.1, EF433417.1( Sapir et al., 2008; Colanero et al., 2020)], and&nbsp;<em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], The orthologous CDS&nbsp;corresponding&nbsp;to the <em>Aft </em>MYB encoding genes from&nbsp;<em>Solanum tuberosum</em> L. Group Phureja clone DM1-3 genome (PGSC DM v4.03 Pseudomolecules) was retrieved from the Potato Genome Sequence Consortium (PGSC: Potato Genome Sequencing Consortium et al., 2011), and the Capsicum annum cv. CM334 genome was retrieved from&nbsp;<em>Capsicum annuum </em>cv CM334 genome chromosome release 1.55 (Hulse-Kemp et al. 2018). These CDS&nbsp;were obtained using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at https://solgenomics.net/tools/blast/). Comparison of syntenic chromosomal regions using known positions of tomato, potato, and pepper markers with comparative map viewer from&nbsp; SGN: (available at https://solgenomics.net/cview) on chromosome 10,&nbsp;was used as a quality check for S.<em> tuberosom</em> and <em>C. annuum.</em> Orthologous&nbsp;CDS corresponding to&nbsp;<em>Salvia miltiorrhiza,&nbsp;Arabidopsis thaliana</em>, [NM_105308.2, NM_105310.4 (Teng et al., 2005, Cominelli et al., 2008; Beradini et al., 2015)] were chosen based on tomato <em>Aft</em> sequence homology and gene annotations of&nbsp;positive R2R3 MYB regulation of anthocyanin. The CDS&nbsp;corresponding&nbsp;to the <em>Aft</em> genes were retrieved from the CDS reference genomes available from the Sol Genomics Network SGN: Tomato Genome CDS (ITAG release 4.0), Potato PGSC DM v3.4 CDS sequences, <em>Capsicum annuum </em>cv CM334 Genome CDS (release 1.55), or from the National Center for Biotechnology Information (NCBI: https://www.ncbi.nlm.nih.gov) reference sequences (RefSeq) section of the Genbank records. When accessed from Genank records, the CDS sequence was extracted from the &ldquo;features&rdquo; section and exported as a FASTA file.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Simulation Data for "Community-Driven Code Comparisons for Three-Dimensional Dynamic Modeling of Sequences of Earthquakes and Aseismic Slip"

<p>Simulation data from Jiang et al. (2022), "Community-Driven Code Comparisons for Three-Dimensional Dynamic Modeling of Sequences of Earthquakes and Aseismic Slip," <em>Journal of Geophysical Research:&nbsp;Solid Earth</em><em>.</em></p> <p>The archive includes simulation data for 3D SEAS benchmarks BP4-QD and BP5-QD that are analyzed in our paper (descriptions in NOTES.txt)&nbsp;</p> <p><strong>BP4-QD Benchmark Simulations:</strong><br>1000 m: &nbsp;jiang.5, lambert.8, barbot.3, barbot.2, dliu.2, li.4<br>500 m:&nbsp; jiang.3, lambert.3, barbot.5, barbot.7, ozawa</p> <p><strong>BP5-QD Benchmark Simulations:</strong><br>2000 m: &nbsp;jiang.6, lambert.8, &nbsp;liu.4, cattania.5, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;dli.7, barbot.3, dliu.10, li.3<br>1000 m:&nbsp; jiang.2, lambert.7, &nbsp;liu.5, cattania.3, ozawa, &nbsp; dli.5, barbot, &nbsp; dliu.6, &nbsp;li.2<br>500 m:&nbsp; jiang.4, lambert.9, &nbsp;liu.6, cattania.4, ozawa.2, dli.6, barbot.2, dliu.8<br>250 m:&nbsp; lambert.10, liu.7</p> <p><strong>BP5-QD with Off-Fault Data:</strong><br>1000 m: &nbsp;lambert.7, dli.5, barbot, &nbsp; dliu.6, li.2<br>500 m:&nbsp; lambert.9, dli.6, barbot.2, dliu.8</p> <p>Tables 2&ndash;4 in our paper summarizes details of numerical codes and selected simulations.</p> <p>The benchmark descriptions and the full suite of simulation data are available at SEAS online platform https://strike.scec.org/cvws/seas/.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Variant dataset and code for "Population-level whole genome sequencing of Ascochyta rabiei identifies genomic loci associated with isolate aggressiveness"

<p>This dataset contains genetic variants (SNPs) of <em>Ascochyta rabiei</em> isolates and the R code used in their analysis to generate the results and figures described in the manuscript "<strong>Population-level whole genome sequencing of <em>Ascochyta rabiei</em> identifies genomic loci associated with isolate aggressiveness</strong>".</p> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Ferry et al. 2024 - Prey that is attractive but not repelled by predators suggests an asymmetric investment in the encounter-avoid-escape sequence. - R Code and Datasets

<p>R code for formating data and running PAMMs for all different combinations of predator-prey.</p> <p>Data of camera trap observation.</p> <p>Data of environmental variable associated to camera trap sites.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Sequencing data and code - Prince et al 2023

<p>Raw sequencing data, code and intermediate analysis files from &quot;Antiviral activity of molnupiravir precursor NHC against SARS-CoV-2 Variants of Concern (VOCs) and implications for the therapeutic window and resistance&quot; (Prince et al, 2023).</p> <p>Please see sample_metadata.xlsx for all metadata relating to the files contained in this repository.</p> <p>Code for data visualisation can be found in: mut_sub_all_pts_serial-pass.R. Specific paths to data will have to be changes to refer to where you have downloaded the data in this repository. Metadata for use with the R script is&nbsp;serial-pass-nimagen-metadata-forR.csv.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Protein codes for tetrameric and dimeric 6-phosphogluconate dehydrogenases and their signature sequence relevant to the cofactor specificity in the 6PGDH family.

<p>Here we present the Uniprot or Genbank code for tetrameric and dimeric 6-phosphogluconate dehydrogenases collected from different organisms. Besides, we include the sequence of the&nbsp;<span class="math-tex">\(\beta2-\alpha2\)</span> motif regarding the cofactor specificity in the 6PGDH family.&nbsp;</p> <p>The sequence data allows&nbsp;generating a phylogenetic tree of the 6PGDH family.&nbsp;The&nbsp;enzymes cluster first by their oligomerization state and then by their sequence in the&nbsp;<span class="math-tex">\(\beta2-\alpha2\)</span>&nbsp;motif, as is shown in the annexed figure.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
dryad40/100

Code and sequence data pertaining to: A phylogenomic perspective on interspecific competition

<p>Evolutionary processes may have substantial impacts on community assembly, but evidence for phylogenetic relatedness as a determinant of interspecific interaction strength remains mixed. In this perspective, we consider a possible role for discordance between gene trees and species trees in the interpretation of phylogenetic signal in studies of community ecology. Modern genomic data show that the evolutionary histories of many taxa are better described by a patchwork of histories that vary along the genome rather than a single species tree. If a subset of genomic loci harbor trait-related genetic variation, then the phylogeny at these loci may be more informative of interspecific trait differences than the genome background. We develop a simple method to detect loci harboring phylogenetic signal and demonstrate its application through a proof of principle analysis of Penicillium genomes and pairwise interaction strength. Our results show that phylogenetic signal that may be masked genome-wide could be detectable using phylogenomic techniques and may provide a window into the genetic basis for interspecific interactions.</p>

opencc-zeroDec 2023View details →
zenodo40/100

Code for generating figures and analyzing amplicon sequencing of human mRNA and reporter mRNA targeted with type III-A CRISPR complex from Streptococcus thermophiles

<p>This dataset contains code for analyzing amplicon sequencing data and generating figures in the manuscript by Anna Nemudraia, Artem Nemudryi, and Blake Wiedenheft (2024), "Repair of CRISPR-guided RNA breaks enables site-specific RNA excision in human cells."&nbsp;</p> <p>Amplicon sequencing data has been deposited to NCBI Sequence Read Archive (SRA) under BioProject PRJNA1099688. The description of read files deposited to SRA can be found in the spreadsheet ./code_for_sequencing_data_analysis/SRA_read_files_description.xlsx</p> <p>The code for analyzing amplicon sequencing data can be found in the archive "code_for_sequencing_data_analysis.tar.gz." Output files from this analysis were used to generate figures. Figures were generated using the ggplot2 package in R and finalized in CorelDRAW.</p> <p>Code for generating figures can be found in the archive "code_for_generating_figures.tar.gz".&nbsp;</p> <p>Any questions or requests regarding the data or the code should be addressed to Dr. Artem Nemudryi at artem.nemudryi@gmail.com.</p>

opencc-by-4.0Apr 2024View details →
dryad40/100

Dorsal striatum coding for the timely execution of action sequences

<p>The automatic initiation of actions can be highly functional. But occasionally these actions cannot be withheld and are released at inappropriate times, impulsively. Striatal activity has been shown to participate in the timing of action sequence initiation and it has been linked to impulsivity. Using a self-initiated task, we trained adult male rats to withhold a rewarded action sequence until a waiting time interval has elapsed. By analyzing neuronal activity we show that the striatal response preceding the initiation of the learned sequence is strongly modulated by the time subjects wait before eliciting the sequence. Interestingly, the modulation is steeper in adolescent rats, which show a strong prevalence of impulsive responses compared to adults. We hypothesize this anticipatory striatal activity reflects the animals' subjective reward expectation, based on the elapsed waiting time, while the steeper waiting modulation in adolescence reflects age-related differences in temporal discounting, internal urgency states, or explore-exploit balance. </p>

opencc-zeroJan 2022View details →
zenodo40/100

GWAS summary statistics and code for "Sequence variants affecting voice pitch in humans"

<p>Contents:&nbsp;GWAS summary statistics for voice pitch (median F0 in reading) and code for acoustic analysis</p> <p>Please refer to the corresponding publication:</p> <p>Gisladottir et al. Sequence variants affecting voice pitch in humans.&nbsp;<em>Science Advances</em></p> <p>The GWAS summary statistics is also available at:&nbsp;https://www.decode.com/summarydata/</p> <p>The code for acoustic analysis is also available at:&nbsp;https://github.com/cadia-lvl/deCODE</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Data and code for "Differential methylation analysis of reduced representation bisulfite sequencing experiments using edgeR"

<p>This data set provides data files and R code to accompany the article <em>Differential methylation analysis of reduced representation bisulfite sequencing experiments using edgeR</em> published by F1000Research.</p> <p>The data consists of Reduced Representation BS-seq methylation profiles of epithelial populations from the mouse mammary gland, with n=2 biological replicates for each of three cell populations.</p> <p>RNA-seq expression profiles of luminal and basal mammary epithelial populations are also provided.</p> <p>The R code undertakes an differential methylation analysis of the BS-seq profiles and demonstrates a strong negative correlation between the differential methylation and differential expression results.</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Code and data associated with Christiansen et al. 2021 "Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing"

<p>All code and data input and output files (except reference genome and raw sequencing data) needed to reproduce the results of Christiansen et al. 2021&nbsp;as released on&nbsp;<a href="https://github.com/notothen/radpilot">https://github.com/notothen/radpilot</a> alongside journal publication. See published paper:</p> <p>Christiansen, H., Heindler, F.M., Hellemans, B.&nbsp;<em>et al.</em>&nbsp;Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing.&nbsp;<em>BMC Genomics</em>&nbsp;<strong>22,&nbsp;</strong>625 (2021). <a href="https://doi.org/10.1186/s12864-021-07917-3">https://doi.org/10.1186/s12864-021-07917-3</a></p>

openother-openJun 2021View details →
zenodo40/100

Accurate annotation of protein coding sequences with IDTAXA - Training Data

<p>Training data used to test IDTAXA, HMMER, and BLAST performance of classification of amino acid and nucleotide sequences.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Accurate annotation of protein coding sequences with IDTAXA - classification results

<p>Raw classification results generated by IDTAXA, BLAST, and HMMER on a training set scraped from KEGG of &gt; 1.5M sequences. HMMER and BLAST results are also converted into IDTAXA like objects for ease of comparison.</p>

opencc-by-4.0Jul 2021View details →
dryad40/100

Data from: Coding-sequence evolution does not explain divergence in petal anthocyanin pigmentation between Mimulus luteus var. luteus and M. l. variegatus

<p><span>Biologists have long been interested in understanding genetic constraints on the evolution of development. For example, noncoding changes in a gene might be favored relative to coding changes due to being less constrained by pleiotropic effects. Here we evaluate the importance of coding-sequence changes to the recent evolution of a novel anthocyanin pigmentation trait in the monkeyflower genus <em>Mimulus</em>. The magenta-flowered <em>Mimulus</em> <em>luteus</em> var. <em>variegatus</em> recently gained petal lobe anthocyanin pigmentation via a single-locus Mendelian difference from its sister taxon, the yellow-flowered <em>M. l. luteus</em>. Previous work showed that the differentially expressed transcription factor gene <em>MYB5a</em>/<em>NEGAN</em> is the single causal gene. However, it was not clear whether <em>MYB5a</em> coding-sequence evolution (in addition to the observed patterns of differential expression) might also have contributed to increased anthocyanin production in <em>M. l. variegatus</em>. Quantitative image analysis of tobacco leaves, transfected with <em>MYB5a</em> coding sequence from each taxon, revealed robust anthocyanin production driven by both alleles. Counter to expectations, significantly higher anthocyanin production was driven by the allele from the low-anthocyanin <em>M. l. luteus.</em> Together with previously-published expression studies, this supports the hypothesis that petal pigment in <em>M. l. variegatus</em> was not gained by protein-coding changes, but instead solely via non-coding cis-regulatory evolution. Finally, while constructing the transgenes needed for this experiment, we unexpectedly discovered two sites in <em>MYB5a</em> that appear to be post-transcriptionally edited – a phenomenon that has been rarely reported, and even less often explored, for nuclear-encoded plant mRNAs.</span></p>

opencc-zeroJun 2023View details →
zenodo40/100

RefSeq bacterial protein coding (nucleotide) sequences

<p><strong>Bacteria_Nucleotide.fas.gz</strong></p><p>151,835,459 protein coding (nucleotide) sequences extracted from 44,831 randomly selected bacterial genomes from NCBI's RefSeq (release 220). Sequences are named by their accession number, followed by "|" and their PGAP predicted function ("protein" tag). For example, the first sequence is named:</p><blockquote><p>WP_125174066.1|iron ABC transporter permease</p></blockquote><p>The process of creating the file involved the following steps.<br><strong>Step 1.</strong> Download 318,613 faa and fna files associated with a bacterial assembly in RefSeq. The following query was used:<br><i>esearch -db assembly -query '"Bacteria"[Organism] AND "latest refseq"[properties] AND "refseq has annotation"[properties]' | esummary | xtract -pattern DocumentSummary -element FtpPath_RefSeq</i><br><strong>Step 2.</strong> Verify all protein coding sequences match the expected protein sequence lengths within three codons, otherwise skip the assembly.<br><strong>Step 3.</strong> Remove all redundant protein coding or protein sequences in a genome. Only exact duplicates were removed, but they were removed from both nucleotides and proteins. Hence, a duplicated amino acid sequence would be discarded along with its coding sequence even if the coding sequence was unique. This was done to keep the two sets of sequences consistent.<br><strong>Step 4.</strong> Name sequences by their accession and PGAP predicted function, separated by a "|" character. The PGAP predicted function is generally uniform, although there are subtle difference between some taxon specific predictions. The predicted function is reasonably dependable but certainly not perfect.<br><strong>Step 5.</strong> Discard any sequences without a predicted function ("hypothetical protein"). These were discarded under the assumption that the protein's function would be required for downstream uses of the sequences.<br><strong>Step 6.</strong> Append protein and protein coding (nucleotide) sequences from randomly ordered assemblies to separate gzipped FASTA formatted files until the Zenodo file size limit was met for either file. Hence, there are many exact duplicate sequences in the set, but none for the sequences from each genome.</p><p>The final sets of sequences are intended to provide large sets of matched protein coding (nucleotide) and protein (amino acid) sequences with consistent labels. The FASTA descriptions in both files are identical. Note, the protein coding sequences do not exactly translate into the protein sequences because of slight differences in length (typically inclusion/exclusion of the first or last codon), as well as use of different translation tables depending on the organism.</p><p>See <i>Related works</i> for the companion file of protein (amino acid) sequences (DOI: 10.5281/zenodo.10030000).</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Data from: Coding-sequence evolution does not explain divergence in petal anthocyanin pigmentation between Mimulus luteus var. luteus and M. l. variegatus

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad40/100

Dorsal striatum coding for the timely execution of action sequences

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad40/100

Code and sequence data pertaining to: A phylogenomic perspective on interspecific competition

Open the record for dataset details and reuse information.

publicDec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record