Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data supporting "Slowest-first translation scheme: Structural asymmetry along protein sequences and co-translational folding"

<p>Contains data for a set of 16,200 non-redundant protein structures taken from the Protein Data Bank. Associated code can be found at https://github.com/jomimc/FoldAsymCode.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Accession passport and sequence data

<p>Accession passport data corresponding to <em>An2-like</em> and <em>Ant1&nbsp;</em>sequence data.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

MOREST Experiment Data with test sequence

<p>MOREST experiment data that contains the test results of MOREST vs. 3 other competing methods. The json files are the raw data generated by the tool, which contains the 500 responses corresponding to bugs information.&nbsp;</p> <p>&nbsp;</p> <p>The detailed test sequence outputs are in the &quot;Test sequence details.zip&quot;</p>

opencc-byNov 2021View details →
dryad40/100

Data from: Pitfalls and pointers: an accessible guide to marker gene amplicon sequencing in ecological applications

<p>Next Generation Sequencing (NGS) is a powerful tool that has been rapidly adopted by many ecologists studying microbial communities. Despite the exciting demonstration of NGS technology as a tool for ecological research, cryptic pitfalls inherent to its use can obscure correct interpretation of NGS data. Here, we provide an accessible overview of a NGS process that uses marker gene amplicon sequences (MGAS) that will allow scientists, particularly community ecologists, to make appropriate methodological choices and understand limits on inference about community composition and diversity that can be drawn from MGAS data.</p> <p>We describe the MGAS pipeline, focusing specifically on cryptic sources of variation that have received less emphasis in the ecological literature, but which may substantially impact inference about microbial community diversity and composition. By simulating communities from published microbiome data, we demonstrate how these sources of variation can generate inaccurate or misleading patterns.</p> <p>We specifically highlight sample dilution without researcher awareness and lane-to-lane variability, two cryptic sources of variation arising during the MGAS pipeline. These sources of variation affect estimates of species presence and relative abundance, particularly for species with moderate to low abundances. Each of these sources of bias can lead to errors in the estimation of both absolute and relative abundance within, and turnover among, microbial communities.</p> <p>Awareness and understanding of what happens and, specifically, why it happens during MGAS generation is key to generating a strong data set and building a robust community matrix. Requesting sample dilution information from the sequencing center, including technical replicates across sequencing lanes, and understanding how sampling intensity and community taxa distribution patterns shape the measurement of community richness, evenness, and diversity are critical for drawing correct ecological inferences using MGAS data.</p>

opencc-zeroNov 2021View details →
zenodo40/100

Supplementary Data for "Sequencing the Pandemic: Rapid and High-Throughput Processing and Analysis of COVID-19 Clinical Samples for 21st Century Public Health"

<p>Supplementary material for F1000 methods manuscript. Includes raw sequencing metrics for two COVID sequencing methodologies, as well as a complete cost breakdown for each methodology.</p>

opencc-by-4.0Jan 2022View details →
dryad40/100

Genome-wide sequence data show no evidence of hybridization and introgression among pollinator wasps associated with a community of Panamanian strangler figs

<p>The specificity of pollinator host choice influences opportunities for reproductive isolation in their host plants. Similarly, host plants can influence opportunities for reproductive isolation in their pollinators. For example, in the fig and fig wasp mutualism, offspring of fig pollinator wasps mate inside the inflorescence that the mothers pollinate. Although often host specific, multiple fig pollinator species are sometimes associated with the same fig species, potentially enabling hybridization between wasp species. Here we study the 19 pollinator species (<em>Pegoscapus</em> spp.) associated with an entire community of 16 Panamanian strangler fig species (<em>Ficus</em> subgenus <em>Urostigma</em>, section <em>Americanae</em>) to determine whether the previously documented history of pollinator host switching and current host sharing predicts genetic admixture among the pollinator species, as has been observed in their host figs. Specifically, we use genome-wide ultraconserved element (UCE) loci to estimate phylogenetic relationships and test for hybridization and introgression among the pollinator species. In all cases, we recover well-delimited pollinator species that contain high interspecific divergence. Even among pairs of pollinator species that currently reproduce within syconia of shared host fig species, we found no evidence of hybridization or introgression. This is in contrast to their host figs, where hybridization and introgression have been detected within this community, and more generally, within figs worldwide. Consistent with general patterns recovered among other obligate pollination mutualisms (<em>e.g.</em>, yucca moths and yuccas), our results suggest that while hybridization and introgression are processes operating within the host plants, these processes are relatively unimportant within their associated insect pollinators.<br>  </p>

opencc-zeroFeb 2022View details →
dryad40/100

Data from: Restriction site-associated DNA sequencing reveals local adaptation despite high levels of gene flow in Sardinella lemuru (Bleeker, 1853) along the northern coast of Mindanao, Philippines

<p>Stock identification and delineation are important in the management and conservation of marine resources. These were highlighted as priority research areas for Bali sardinella (<em>Sardinella lemuru</em>) which is among the most commercially important fishery resources in the Philippines. Previous studies have already assessed the stocks of <em>S. lemuru</em> between Northern Mindanao Region (NMR) and Northern Zamboanga Peninsula (NZP), yielding conflicting results. Phenotypic variation suggests distinct stocks between the two regions, while mitochondrial DNA did not detect evidence of genetic differentiation for this high gene flow species. This paper tested the hypothesis of regional structuring using genome-wide single nucleotide polymorphisms (SNPs) acquired through restriction-site associated DNA sequencing (RADseq). We examined patterns of population genomic structure using a full panel of 3,573 loci, which was then partitioned into a neutral panel of 3,348 loci and an outlier panel of 31 loci. Similar inferences were obtained from the full and neutral panels, which were contrary to the inferences from the outlier panel. While the full and neutral panels suggested a panmictic population (global F<sub>ST</sub> ~ 0, p &gt; 0.05), the outlier panel revealed genetic differentiation between the two regions (global F<sub>ST</sub> = 0.161, p = 0.001; F<sub>CT</sub> = 0.263, p &lt; 0.05). This indicated that while gene flow is apparent, selective forces due to environmental heterogeneity between the two regions play a role in maintaining adaptive variation. Annotation of the outlier loci returned five genes that were mostly involved in organismal development. Meanwhile, three unannotated loci had allele frequencies that correlated with sea surface temperature. Overall, our results provided support for local adaptation despite high levels of gene flow in <em>S. lemuru</em>. Management therefore should not only focus on demographic parameters (e.g., stock size, catch volume), but also consider the preservation of adaptive variation.</p>

opencc-zeroFeb 2022View details →
zenodo40/100

Integrated Data of Single cell RNA sequencing for Human Pancreatic Adenocarcinoma

<p>These data are collected and integrated from five available deposit data and one original data of single cell RNA sequencing from human pancreatic adenocarcinoma. Further analyses data for bulk transcriptomics (such as TCGA )using scRNAseq data and re-clustering for ductal epithelial cells and fibroblasts are also stored in step by step. Moreover, all R code is uploaded.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

PredictIO Sequencing Data

<p># Data for our paper titled &quot;Leveraging Big Data of Immune Checkpoint Blockade Response Identifies Novel Potential Targets&quot;.</p> <p>------------------------------------------------<br> ------------------------------------------------</p> <p><br> *Yacine Bareche, Deirdre Kelly, Farnoosh Abbas-Aghababazadeh, Minoru Nakano, Parinaz Nasr Esfahani, Denis Tkachuk, Hassan Mohammad, Robert Samstein, Chung-Han Lee, Luc G. T. Morris, Philippe L. Bedard, Benjamin Haibe-Kains &amp; John Stagg*</p> <p>## Summary<br> Several genomics and gene expression signatures have been proposed as predictive biomarkers of clinical response to immune checkpoint blockade (ICB), with questionable reproducibility. We here report a large-scale comparative analysis of candidate biomarkers of ICB responses in a pan-cancer meta-analysis of over 3,500 ICB-treated patients representing 12 different tumor types. A web-application (predictIO.ca) was developed to allow researchers to further interrogate this data compendium. We tested the hypothesis that a de novo pan-cancer gene expression analysis would bring forth critical ICB resistance pathways and novel therapeutic targets. At the genomic level, we confirmed that non-synonymous tumour mutational burden (nsTMB) was significantly associated with ICB responses across tumor types, with the exception of kidney cancer. At the transcriptional level, 21 out of 39 published gene expression signatures were significantly associated with pan-cancer ICB responses. Strikingly, the predictive value of a de novo gene expression signature (referred to as PredictIO_100) composed of the top 100 genes most significantly associated with ICB responses at pan-cancer level was superior to nsTMB and other gene expression biomarkers. Within PredictIO_100, two genes, F2RL1 (encoding protease-activated receptor-2) and RBFOX2 (encoding RNA binding motif protein 9), were concomitantly associated with worse ICB clinical outcomes, T cell dysfunction in ICB-naive patients and resistance to dual PD-1/CTLA-4 blockade in preclinical mouse cancer models. Taken together, our study underlined the relative impact of candidate biomarkers of ICB responses in a large pan-cancer cohort, demonstrated the potential of de novo pan-cancer gene expression signatures and identified F2RL1, previously involved in tumor immune regulation, and RBFOX2, a critical regulator of epithelial-to-mesenchymal transition, as potential therapeutic targets to overcome ICB resistance.</p> <p>------------------------------------------------<br> ------------------------------------------------</p> <p>## mouseModel_Zemek<br> Expression data of the mouse model study from Zemek et al. (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE117358)</p> <p>## process_data<br> Expression and SNV data of the discovery cohort</p> <p>## validation_cohort<br> Expression and SNV data of the validation cohort</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

A volumetric model of rabbit heart and torso including ECG data and ventricular activation sequence

<p>Data generated and analyzed of our work titled &quot;A computational model of rabbit geometry and ECG: Optimizing ventricular activation sequence and APD distribution&quot;. Please see the respective publication for more context.</p> <p>&nbsp;</p> <ul> <li>BSPM_filtered.dat <ul> <li>Contains the filtered ECG Data</li> </ul> </li> <li>BSPM_original.bdf <ul> <li>Contains the originally recorded signal.<br> Information on the file format itself can be found here: <a href="https://www.biosemi.com/faq/file_format.htm">https://www.biosemi.com/faq/file_format.htm</a><br> Links to various toolboxes to open the file can be found here: <a href="https://www.biosemi.com/download.htm">https://www.biosemi.com/download.htm</a></li> </ul> </li> <li>CT_DataDCM.zip <ul> <li>Contains the recorded CT images of heart and torso in DCM file format</li> </ul> </li> <li>ECG_NodeIndices.txt <ul> <li>Contains the the node IDs of the torso mesh corresponding to the electrode positions of the ECG Vest</li> </ul> </li> <li>Endocardial_Surface_Papillary.stl <ul> <li>Segmented endocardial surface including papillary muscles</li> </ul> </li> <li>Mat_LeadField.dat <ul> <li>Contains the calculated lead field matrix to be used in combination with the provided Mesh_Ven.vtu. Make sure to keep the node order</li> <li> <pre><code class="language-python"># Python example of usage # Define a read_vm_vec function which reads your calculated data beforehand import numpy as np t_begin = 0 t_end = 400 LF_mat = np.loadtxt('Mat_LeadField.dat') times = np.linspace(t_begin, t_end, t_end-t_begin) result = np.zeros((len(times), 31)) for i,t in enumerate(times): vm_vec = read_vm_vec(t) result[i, :] = LF_mat.dot(vm_vec)[0:31] result = np.insert(result, 0, times, axis=1) np.savetxt('BSPM.dat', result)</code></pre> <p>&nbsp;</p> </li> </ul> </li> <li>Mesh_PurkinjeTree.vtp <ul> <li>The resulting optimized Purkinje Node tree. We recommend using <a href="https://www.paraview.org/">ParaView</a> for visualization</li> </ul> </li> <li>Mesh_StimPoints.vtp <ul> <li>The resulting points of stimulation.</li> </ul> </li> <li>Stim_IndexTime.dat <ul> <li>Contains the stimulation pattern in terms the node index of Mesh_Ven.vtu and the respective stimulation time</li> </ul> </li> <li>Mesh_Ven.vtu <ul> <li>Contains the ventricular mesh as well as the repective lead field matrix values for each surface node.</li> <li>Material:<br> Right Ventricle 30<br> Left Ventricle 31</li> </ul> </li> <li>Mesh_Torso.vtu <ul> <li>Contains the whole torso mesh.</li> <li>Material:<br> Fat 2<br> Bones 3<br> Blood 9<br> Cartilage 14<br> Liver 20<br> Lungs 17<br> Right Ventricle 30<br> Left Ventricle 31<br> Right Atrium 32<br> Left Atrium 33<br> Aorta 60<br> Pulmonary artery 61<br> Left Vena Jugularis 62<br> Right Vena Jugularis 62<br> Post Vena Cava 62</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data

<p>To date, ancient genome analyses have been largely confined to the study of single nucleotide polymorphisms (SNPs). Copy number variants (CNVs) are a major contributor of disease and of evolutionary adaptation, but identifying CNVs in ancient shotgun-sequenced genomes is hampered by (i) most published genomes being &lt;1x&nbsp;coverage, (ii) ancient DNA fragments being typically &lt;80 bps. These characteristics preclude state-of-the-art CNV detection software to be effectively applied to ancient genomes. Here we present CONGA, an algorithm tailored for genotyping deletion and duplication events in genomes with low depths of coverage. Simulations and down-sampling experiments show that CONGA can genotype deletions &gt;1 kbps with F-scores &gt;0.75 at &gt;=1x, and distinguish between heterozygous and homozygous states. Using CONGA, we analyse deletion events at 10,018 loci in 56 ancient human genomes spanning the last 50,000 years, with coverages 0.4x-26x. We show that inter-individual genetic diversity measured using deletions and SNPs are highly correlated, as in modern-day genomes, confirming that deletion frequencies broadly reflect demographic history. We also identify signatures of strong purifying selection on deletions in ancient-genomes, such as an excess of singletons compared to those in SNPs. CONGA paves the way for systematic studies of drift, mutation load, and adaptation in ancient and modern-day gene pools through the lens of CNVs.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Illumina sequencing data of Agro-mediated gene-edited apple lines

<p>FastaQ pair-end Illumina sequencing data of the Dipm1/4, Hipm1, and Mlo19 genes&nbsp;from the different apple Agro-edited lines obtained in the project. GA = Gala; GD = Golden Delicious</p> <p>Lines included are:</p> <p>GA1, GA3, GA4, GA5, and GA WT</p> <p>GD1, 2, 6, 10, 11, 12, 15, 17, 18, 19, 20, 21, 23, 24, 26, 27, 29, 31, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and WT</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Sequencing data of RNA editing, RNA modifications, and transcriptional units in Listeria monocytogenes

<p>Sequencing data for &quot;RNA editing, RNA modifications, and transcriptional units in <em>Listeria monocytogenes</em>&quot; manuscript, which is submitted to BMC genomics.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Data in support of "An exploration of linkage fine-mapping on sequences from case-control studies"

<p>These data were simulated for an exploration of linkage fine-mapping on sequences from case-control studies. The code to generate and analyze the data is available on GitHub in the scripts at&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.&nbsp;Queries may be directed to Payman Nickchi at&nbsp;<a href="mailto:pnickchi@sfu.ca">pnickchi@sfu.ca</a>&nbsp;or Charith (Bhagya) Karunarathna at&nbsp;<a href="mailto:ch757276@dal.ca">ch757276@dal.ca</a>.</p>

opencc-by-4.0May 2022View details →
dryad40/100

Morphological and DNA sequence data generated by Sanger sequencing and target capture methods for moss plants in the genus Fissidens from herbarium specimens

<p><span>Morphological evolution in mosses has long been hypothesized to accompany shifts in microhabitats and can be tested using comparative phylogenetics. These lines of inquiry have developed substantially, in part, by target capture sequencing allowing for phylogenomic scale data generated from herbarium specimens. In the present study, we test the relationship between taxonomically important morphological characters in the moss genus <em>Fissidens</em>, using both a 400-locus dataset generated using a target-capture approach as well as a three-locus phylogeny generated using sanger sequencing. Phylogenetic trees were generated using ASTRAL and Bayesian Inference and used to test the monophyly of subgenera/sections and provided the basis for ancestral character reconstruction and phylogenetic correlation analyses among five morphological characters as well as habitat moisture scored from literature. The characters <em>axillary hyaline nodules</em>, <em>limbidium</em>, <em>costa</em>, and <em>peristome morphology</em> as well as <em>sexual system</em>, <em>minimum habitat moisture</em>, <em>average habitat moisture</em>, <em>maximum habitat moisture</em>, and <em>habitat moisture niche breadth</em> each exhibit statistically significant phylogenetic signal. Significant correlations were found between the limbidium (phyllid/leaf border) and habitat moisture niche breadth, which could be interpreted as a more extensive <em>limbidium</em> enabling species to survive across a wider variety of habitats. Correlations were also found between <em>costa anatomy</em> and the <em>limbidum</em> of the gametophyte and sporophyte <em>peristome</em> <em>morphology</em>, as well as <em>average habitat moisture</em> and <em>sexual system</em>. Continued exploration of the relationships between morphological evolution, life history, and habitat will enable us to expand our understanding of functional morphology in mosses.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Data in support of "An exploration of linkage fine-mapping on sequences from case-control studies"

<p>These data were simulated for an exploration of linkage fine-mapping on sequences from case-control studies. The scripts&nbsp;to generate and analyze the data are&nbsp;available&nbsp;at&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.&nbsp;Queries may be directed to Payman Nickchi at&nbsp;<a href="mailto:pnickchi@sfu.ca">pnickchi@sfu.ca</a>&nbsp;or Charith (Bhagya) Karunarathna at&nbsp;<a href="mailto:ch757276@dal.ca">ch757276@dal.ca</a>.</p> <p><strong>README file for All_data directory</strong></p> <p><strong>Directory structure</strong></p> <p>The&nbsp;All_data&nbsp;directory consists of this README file and 500 sub-directories named&nbsp;DatasetX, for&nbsp;X=1 to 500. Within each&nbsp;DatasetX&nbsp;sub-directory are further sub-directories named&nbsp;alt&nbsp;and&nbsp;null&nbsp;containing files named&nbsp;pop_data.RData&nbsp;and&nbsp;sample_data.RData.</p> <p><strong>alt&nbsp;<em>versus</em>&nbsp;null&nbsp;directories</strong></p> <p>The files in the&nbsp;alt&nbsp;and&nbsp;null&nbsp;directories contain the same variant data but different phenotype data. In particular, under the null hypothesis, disease status is simulated at random according to a 5% prevalence in the population, whereas under the alternative hypothesis disease status is simulated according to a penetrance model that depends on causal SNVs. The R script to simulate data<br> under the alternative hypothesis is in the file&nbsp;1_SimulateData.R&nbsp;in the Github repository&nbsp;<a href="https://github.com/SFUStatgen/PBJ0">https://github.com/SFUStatgen/PBJ0</a>.</p> <p><strong>pop_data.RData&nbsp;and&nbsp;sample_data.RData&nbsp;files</strong></p> <p>The data structures contained in the&nbsp;pop_data.RData&nbsp;and&nbsp;sample_data.RData&nbsp;files are described below. The structure is the same under both the null and alternative hypothesis.</p> <p><strong>pop_data.RData</strong></p> <p>From R,&nbsp;load(&quot;pop_data.RData&quot;)&nbsp;loads a list named&nbsp;pop_data&nbsp;whose elements describe the population&rsquo;s haplotype and phenotype data. The list elements are as follows.</p> <ul> <li>Variants: a matrix of variants for the population of 6200 haplotypes <ul> <li>rows are SNVs,</li> <li>columns are sequences</li> </ul> </li> <li>Positions: a data frame of SNV positions <ul> <li>rows are SNVs,</li> <li>column 1 is the SNV name and column 2 is the SNV position in base pairs</li> </ul> </li> <li>Population.Mapping: a data frame telling us how the sequences are paired into individuals <ul> <li>rows are individuals</li> <li>First column 1 is an individual ID from 1,&hellip;,3100; columns 2 and 3 are the sequence IDs of the first and second sequence for that individual where the sequence IDs are the column names of the&nbsp;Variants&nbsp;matrix.</li> </ul> </li> <li>Genotype.Matrix: a matrix of genotypes (i.e.&nbsp;variant counts) for the 3100 individuals <ul> <li>rows are SNVs</li> <li>columns are the individuals</li> </ul> </li> <li>causal_region: a vector containing the lower- and upper-limit of the causal region in base pairs.</li> <li>cSNV: a vector containing the IDs of the causal SNVs, where the SNV IDs are the row names of the&nbsp;Variants&nbsp;matrix.</li> <li>DISCRETE: a list with the following elements. <ul> <li>CaseIndividuals: vector of IDs of the affected individuals in the population.</li> <li>ControlIndividuals: vector of IDs of the unaffected in the population.</li> <li>BinaryTrait: a vector of trait status (0=unaffected, 1=affected) for each individual.</li> </ul> </li> </ul> <p><strong>Note:</strong>&nbsp;Within the same&nbsp;DatasetX&nbsp;directory, the only difference between the&nbsp;pop_data&nbsp;data structures under the null and alternative hypothesis is the phenotype information contained in their respective&nbsp;DISCRETE&nbsp;list elements. Both the null and alternative pop_data data structure share&nbsp;list elements: Variants,&nbsp;Positions,&nbsp;Population.Mapping,&nbsp;Genotype.Matrix,&nbsp;causal_region&nbsp;and&nbsp;cSNV.</p> <p><strong>sample_data.RData</strong></p> <p>From R,&nbsp;load(&quot;sample_data.RData&quot;)&nbsp;loads a list whose elements describe the sequences and phenotypes of the sample of 50 affected individuals (cases) and 50 unaffected individuals (controls) from the population.</p> <ul> <li>Haps: a list with two elements. <ul> <li>sample_haps: a matrix of 200 sequences for the 50 cases and 50 controls. Rows are SNVs and columns are sequences, with the sequences of sampled cases appearing first (i.e.&nbsp;first 100 columns), followed by the sequences of sampled controls (i.e.&nbsp;last 100 columns). Sequences include only those SNVs that are polymorphic in the sample.</li> <li>ccStatus: a vector indicating the case/control status of the individual to which the sequence belongs, with case=1 and control=0.</li> </ul> </li> <li>Genos: a list with two elements. <ul> <li>sample_genos: a matrix of 100 genotypes for the 50 cases and 50 controls. Rows are SNVs and columns are genotypes, with genotypes of cases appearing first, followed by genotypes of controls.</li> <li>ccStatus: a vector indicating the case/control status of each individual, with case=1 and control=0.</li> </ul> </li> <li>Posn: a data frame of SNV positions for each SNV that is polymorphic in the sample. The first column is the SNV name and the second is the SNV position in base pairs.&nbsp;Posn&nbsp;is a subset of&nbsp;pop_data$Positions.</li> <li>poly_cSNV: a vector of IDs for causal SNVs that are polymorphic in the sample.</li> <li>CaseIND: a vector of individual IDs for the case individuals (see&nbsp;pop_data$Population.Mapping).</li> <li>ControlIND: a vector of individual IDs for the control individuals (see&nbsp;pop_data$Population.Mapping).</li> <li>CaseHapID: a vector of IDs for the sequences that belong to cases (see the sequence IDs in the column names of the matrix&nbsp;pop_data$Variants).</li> <li>ControlHapID: a vector of IDs for the sequences that belong to controls (see the sequence IDs in the column names of the matrix&nbsp;pop_data$Variants).</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Data file: Amino acid sequence Cm28; a scorpion toxin

<p>The Data set contains the amino acid sequence of Cm28, a peptide toxin found in the venom of <em>Centruroides margaritatus</em>.&nbsp;The peptide sequence and functional&nbsp;data are&nbsp;reported in an article in the Journal of General Physiology (JGP) under this DOI&nbsp;10.1085/jgp.202213146&nbsp;and this data&nbsp;will appear in the UniProt Knowledgebase under the accession<br> number C0HM22.</p>

opencc-by-4.0Jun 2022View details →
dryad40/100

Data for: Range and niche expansion through multiple interspecific hybridization - a genotyping by sequencing analysis of Cherleria (Caryophyllaceae)

<p><b>Background:</b> <i>Cherleria</i> (Caryophyllaceae) is a circumboreal genus that also occurs in the high mountains of the northern hemisphere. In this study, we focus on a clade that diversified in the European High Mountains, which was identified using nuclear ribosomal (nrDNA) sequence data in a previous study. With the nrDNA data, all but one species was monophyletic, with little sequence variation within most species. Here, we use genotyping by sequencing (GBS) data to determine whether the nrDNA data showed the full picture of the evolution in the genomes of these species.</p> <p><b>Results:</b> The overall relationships found with the GBS data were congruent with those from the nrDNA study. Most of the species were still monophyletic and many of the same subclades were recovered, including a clade of three narrow endemic species from Greece and a clade of largely calcifuge species. The GBS data provided additional resolution within the two species with the best sampling, <i>C. langii</i> and <i>C. laricifolia</i>, with structure that was congruent with geography. In addition, the GBS data showed significant hybridization between several species, including species whose ranges did not currently overlap.</p> <p><b>Conclusions:</b> The hybridization led us to hypothesize that lineages came in contact on the Balkan Peninsula after they diverged, even when those lineages are no longer present on the Balkan Peninsula. Hybridization may also have helped lineages expand their niches to colonize new substrates and different areas. Not only do genome-wide data provide increased phylogenetic resolution of difficult nodes, they also give evidence for a more complex evolutionary history than what can be depicted by a simple, branching phylogeny.</p>

opencc-zeroJun 2022View details →
zenodo40/100

OTU level 16s sequence data for "Algae drive convergent bacterial community assembly when nutrients are scarce"

<p>16s sequence data at the OTU level for the experiments conducted in&nbsp;&quot;Algae drive convergent bacterial community assembly when nutrients are scarce&quot;</p> <p>The file is in fasta format, which can be read by many software packages including biopython, R, and SILVA&#39;s alignment, classification and tree service.</p> <p>The sequence ids can be used to find the phylogeny and OTU ID of the sequences from Supplementary dataset 4.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Training data for 'Functional annotation of protein sequences' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for functional annotation of protein sequences.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record