Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941

<p><strong>Data Processing</strong></p> <p>&nbsp;</p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p><strong>software_versions</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p><strong>paired_reads_assembly</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p><strong>consensus_building</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p><strong>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT</p> <p>&nbsp;</p> <p><strong>Format</strong></p> <p>&nbsp;</p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>C_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Isotype subclass</p> <p><strong>SEQUENCE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sequence identifier</p> <p><strong>V_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene and allele</p> <p><strong>D_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene and allele</p> <p><strong>J_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene and allele</p> <p><strong>JUNCTION_LENGTH&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction length</p> <p><strong>CONSCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Total number of mutations in V gene&nbsp;</p> <p><strong>SEQUENCE_INPUT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Full length sequence</p> <p><strong>SEQUENCE_IMGT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>Run&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>ID of sequencing run</p> <p><strong>Sample_type&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..)</p> <p><strong>Sex&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sex of the Subject</p> <p><strong>Age&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Age of the subject</p> <p><strong>UNIQUE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Subject identifier&nbsp;</p> <p><strong>SAMPLE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sample identifier, linking back to raw data</p> <p><strong>Subset&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell subset&nbsp;</p> <p><strong>Repertoire&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in CDR region</p> <p><strong>R_SFWR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in FWR region</p> <p><strong>V_FAM&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V family gene</p> <p><strong>V_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene</p> <p><strong>D_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene</p> <p><strong>J_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene</p> <p><strong>Clust_Rank&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster rank</p> <p><strong>Clust_REPRES&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster representative</p> <p><strong>Clust_SIZE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster size</p> <p><strong>Clust_MAXFREQ&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster maximum frequency</p> <p><strong>Clust_SHAREDNESS&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster sharedness</p> <p><strong>CDR3_AA_GRAVY&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_CHARGE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 charge</p> <p><strong>CDRH3PDB&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDRH3 PDB (Structure) code</p> <p><strong>H1Canon&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H1 Canonical class</p> <p><strong>H2Canon&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H2 Canonical class</p> <p><strong>H1_GERMLINE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H1 Germline Canonical class</p> <p><strong>H2_GERMLINE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H2 Germline Canonical class</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Apr 2019View details →
zenodo36/100

Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon" - PART I

<p>Raw Fast5 data for &quot;Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon&quot;. See Supplementary Table 2 for associating each sample to its barcode.</p> <p>- FC1_1 includes data for the HM mock community from BEI resources and skin microbiome of the chin in dogs.</p> <p>- FC1_2 includes data for the dorsal skin samples</p> <p>- FC2 includes data for the Zymobiomics mock community&nbsp;and Staphylococcus pseudintermedius isolate</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

Dataset used for "Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing data"

<p>Each simulated dataset was generated using the AbSim R package (version 0.2.6) in a B cell single-lineage fashion. Each B cell clone simulation begins with a random selection from sets of IGHV, IGHD, and IGHJ germline sequences to produce a unique V(D)J recombination event. Then, clones are made by introducing mutations using a local nucleotide context-dependent model (S5F model) along a phylogenetic tree in which branching events occur stochastically.&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Fig. 12 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)

Fig. 12. Known distribution of the species of the T. opinatus subgroup.

opencc-by-3.0May 2016View details →
dryad36/100

Characterizing and classifying neuroendocrine neoplasms through microRNA sequencing and data mining

<p>Neuroendocrine neoplasms (NENs) are clinically diverse and incompletely characterized cancers that are challenging to classify. MicroRNAs (miRNAs) are small regulatory RNAs that can be used to classify cancers. Recently, a morphology-based classification framework for evaluating NENs from different anatomic sites was proposed by experts, with the requirement of improved molecular data integration. Here, we compiled 378 miRNA expression profiles to examine NEN classification through comprehensive miRNA profiling and data mining. Following data preprocessing, our final study cohort included 221 NEN and 114 non-NEN samples, representing 15 NEN pathological types and five site-matched non-NEN control groups. Unsupervised hierarchical clustering of miRNA expression profiles clearly separated NENs from non-NENs. Comparative analyses showed that miR-375 and miR-7 expression is substantially higher in NEN cases than non-NEN controls. Correlation analyses showed that NENs from diverse anatomic sites have convergent miRNA expression programs, likely reflecting morphologic and functional similarities. Using machine learning approaches, we identified 17 miRNAs to discriminate 15 NEN pathological types and subsequently constructed a multi-layer classifier, correctly identifying 217 (98%) of 221 samples and overturning one histologic diagnosis. Through our research, we have identified common and type-specific miRNA tissue markers and constructed an accurate miRNA-based classifier, advancing our understanding of NEN diversity.</p>

opencc-zeroJun 2020View details →
dryad36/100

Data from: Accumulation curves of environmental DNA sequences predict coastal fish diversity in the Coral Triangle

Environmental DNA (eDNA) has the potential to provide more comprehensive biodiversity assessments particularly for vertebrates in species-rich regions. Yet, this method requires the completeness of a reference database, i.e. a list of DNA sequences attached to each species, which is never met. As an alternative, a diversity of Operational Taxonomic Units (OTUs) can be extracted from eDNA metabarcoding. However, the extent to which the diversity of OTUs provided by a limited eDNA sampling effort can predict regional species diversity is unknown. Here, by modelling OTU accumulation curves of eDNA seawater samples across the Coral Triangle, we obtained an asymptote reaching 1,531 fish OTUs while 1,611 fish species are recorded in the region. Besides, we also accurately predict (R² = 0.92) the distribution of species richness among fish families from OTU-based asymptotes. Thus, the multi-model framework of OTU accumulation curves extends the use of eDNA metabarcoding in ecology, biogeography and conservation.

opencc-zeroJul 2020View details →
dryad36/100

Data from: RapidRat: development, validation and application of a genotyping-by-sequencing panel for rapid biosecurity and invasive species management

<p>Invasive alien species (IAS) are among the main causes of global biodiversity loss. Invasive brown (Rattus norvegicus) and black (R. rattus) rats, in particular, are leading drivers of extinction on islands, especially in the case of seabirds where &gt;50% of all extinctions have been attributed to rat predation. Eradication is the primary form of invasive rat management, yet this strategy has resulted in a ~10-38% failure rate on islands globally. Genetic tools can help inform IAS management, but such applications to date have been largely reactive, time-consuming, and costly. Here, we developed a Genotyping-in-Thousands by sequencing (GT-seq) panel for rapid species identification and population assignment of invasive brown and black rats (RapidRat) in Haida Gwaii, an archipelago comprising ~150 islands off the central coast of British Columbia, Canada. We constructed an optimized panel of 443 single nucleotide polymorphisms (SNPs) using previously generated double-digest restriction-site associated DNA (ddRAD) genotypic data (27,686 SNPs) from brown (n=295) and black rats (n=241) sampled throughout Haida Gwaii. The informativeness of this panel for identifying individuals to species and island of origin was validated relative to the ddRAD results; in all comparisons, admixture coefficients and population assignments estimated using RapidRat were consistent. To demonstrate application, 20 individuals from novel invasions of three islands (Agglomerate, Hotspring, Ramsay) were genotyped using RapidRat, all of which were confidently assigned (&gt;98.5% probability) to Faraday and Murchison Islands as putative source populations. These results indicated that a previous eradication on Hotspring Island was conducted at an inappropriate geographic scale; future management should expand the eradication unit to include neighboring islands to prevent re-invasion. Overall, we demonstrated that RapidRat is an effective tool for managing invasive rat populations in Haida Gwaii and provided a clear framework for GT-seq panel development for informing biodiversity conservation in other systems.</p>

opencc-zeroDec 2019View details →
zenodo36/100

A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula

<p>This genomic dataset provides highly variable single-nucleotide polymorphism&nbsp;(SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>

opencc-by-4.0Jun 2020View details →
dryad36/100

Data from: Evolutionary and phylogenetic insights from a nuclear genome sequence of the extinct, giant subfossil koala lemur Megaladapis edwardsi

<p><span>No endemic Madagascar animal with body mass &gt;10 kg survived a relatively recent wave of extinction on the island. From morphological and isotopic analyses of skeletal 'subfossil' remains we can reconstruct some of the biology and behavioral ecology of giant lemurs (primates; up to ~160 kg), elephant birds (up to ~860 kg), and other extraordinary Malagasy megafauna that survived well into the past millennium. Yet much about the evolutionary biology of these now extinct species remains unknown, along with persistent phylogenetic uncertainty in some cases. Thankfully, despite the challenges of DNA preservation in tropical and sub-tropical environments, technical advances have enabled the recovery of ancient DNA from some Malagasy subfossil specimens. Here we present a nuclear genome sequence (~2X coverage) for one of the largest extinct lemurs, the koala lemur <i>Megaladapis edwardsi </i>(~85kg). To support the testing of key phylogenetic and evolutionary hypotheses we also generated new high-coverage complete nuclear genomes for two extant lemur species, <i>Eulemur rufifrons</i> and <i>Lepilemur mustelinus</i>, and we aligned these sequences with previously published genomes for three other extant lemur species and 47 non-lemur vertebrates. Our phylogenetic results confirm that <i>Megaladapis</i> is most closely related to the extant Lemuridae (typified in our analysis by <i>E. rufifrons</i>) to the exclusion of <i>L. mustelinus</i>, which contradicts morphology-based phylogenies. Our evolutionary analyses identified significant convergent evolution between <i>M. edwardsi</i> and extant folivorous primates (colobine monkeys) and ungulate herbivores (horses) in genes encoding protein products that function in the biodegradation of plant toxins and nutrient absorption. These results suggest that koala lemurs were highly adapted to a leaf-based diet, which may also explain their convergent craniodental morphology with the small-bodied folivore <i>Lepilemur</i>.</span></p>

opencc-zeroOct 2020View details →
dryad36/100

Paired human macrophage RNA sequencing data

<p>Allele-specific expression (ASE) analysis, which quantifies the relative expression of two alleles in a diploid individual, is a powerful tool for identifying <em>cis</em>-regulated gene expression variations that underlie phenotypic differences among individuals. Existing methods for gene-level ASE detection analyze one individual at a time, therefore failing to account for shared information across individuals. Failure to accommodate such shared information not only reduces power, but also makes it difficult to interpret results across individuals. However, when only RNA sequencing (RNA-seq) data are available, ASE detection across individuals is challenging because the data often include individuals that are either heterozygous or homozygous for the unobserved <em>cis</em>-regulatory SNP, leading to sample heterogeneity as only those heterozygous individuals are informative for ASE, whereas those homozygous individuals have balanced expression. To simultaneously model multi-individual information and account for such heterogeneity, we developed ASEP, a mixture model with subject-specific random effect to account for multi-SNP correlations within the same gene. ASEP only requires RNA-seq data, and is able to detect gene-level ASE under one condition and differential ASE between two conditions (e.g., pre- versus post- treatment). Extensive simulations demonstrated the convincing performance of ASEP under a wide range of scenarios. We applied ASEP to a human kidney RNA-seq dataset, identified ASE genes and validated our results with two published eQTL studies. We further applied ASEP to a human macrophage RNA-seq dataset, identified genes showing evidence of differential ASE between M0 and M1 macrophages, and confirmed our findings by results from cardiometabolic trait-relevant genome-wide association studies. To the best of our knowledge, ASEP is the first method for gene-level ASE detection at the population level that only requires the use of RNA-seq data. With the growing adoption of RNA-seq, we believe ASEP will be well-suited for various ASE studies for human diseases.</p>

opencc-zeroApr 2020View details →
dryad36/100

Data from: The genome sequence and insights into the immunogenetics of the bananaquit (Passeriformes: Coereba flaveola)

Avian genomics, especially of non-model species, is in its infancy relative to mammalian genomics. Here, we describe the sequencing, assembly, and annotation of a new avian genome, that of the bananaquit Coereba flaveola (Passeriformes: Thraupidae). We produced ∼30-fold coverage of the genome with an assembly size of ca. 1.2 Gb, including approximately 16,500 annotated genes. Passerine birds, such as the bananaquit, are commonly infected by avian malarial parasites (Haemosporida), which presumably drive adaptive evolution of immunogenetic loci within the host genome. In the context of our research on the distribution of avian Haemosporida, we specifically characterized immune loci, including toll-like receptor (TLR) and major histocompatibility complex (MHC) genes. Additionally, we identified novel molecular markers in the form of single nucleotide polymorphisms (SNPs), both genome-wide and within identified immune loci. We discovered nine TLR genes and four MHC genes and identified five other TLR- or MHC- associated genes. Genome-wide, over 6 million high-quality SNPs were annotated, including 568 within TLR genes and 102 in MHC genes. This newly described genome and immune characterization expands the knowledge base for avian genomics and phylogenetics and allows for immune genotyping in the bananaquit, providing tools for the investigation of host-parasite coevolution.

opencc-zeroDec 2015View details →
dryad36/100

Data from: Laying sequence interacts with incubation temperature to influence rate of embryonic development and hatching synchrony in a precocial bird

Incubation starts during egg laying for many bird species and causes developmental asynchrony within clutches. Faster development of late-laid eggs can help reduce developmental differences and synchronize hatching, which is important for precocial species whose young must leave the nest soon after hatching. In this study, we examined the effect of egg laying sequence on length of the incubation period in Wood Ducks (Aix sponsa). Because incubation temperature strongly influences embryonic development rates, we tested the interactive effects of laying sequence and incubation temperature on the ability of late-laid eggs to accelerate development and synchronize hatching. We also examined the potential cost of faster development on duckling body condition. Fresh eggs were collected and incubated at three biologically relevant temperatures (Low: 34.9°C, Medium: 35.8°C, and High: 37.6°C), and egg laying sequences from 1 to 12 were used. Length of the incubation period declined linearly as laying sequence advanced, but the relationship was strongest at medium temperatures followed by low temperatures and high temperatures. There was little support for including fresh egg mass in models of incubation period. Estimated differences in length of the incubation period between eggs 1 and 12 were 2.7 d, 1.2 d, and 0.7 d at medium, low and high temperatures, respectively. Only at intermediate incubation temperatures did development rates of late-laid eggs increase sufficiently to completely compensate for natural levels of developmental asynchrony that have been reported in Wood Duck clutches at the start of full incubation. Body condition of ducklings was strongly affected by fresh egg mass and incubation temperature but declined only slightly as laying sequence progressed. Our findings show that laying sequence and incubation temperature play important roles in helping to shape embryo development and hatching synchrony in a precocial bird.

opencc-zeroDec 2017View details →
dryad36/100

Data from: Sectional relationships in the Eurasian bearded iris (subgen. Iris) based on phylogenetic analyses of sequence data

Subgenus Iris is wholly Eurasian, distributed in temperate regions from northeastern China to eastern and southern Europe where they occur in mountainous and/or dry rocky sites from near sea level to elevations of 4,500 m. These species have an easily discerned synapomorphy, a multicellular beard on each petaloid sepal. Currently two large and relatively well known and four smaller and less known sections are recognized in the subgenus. This study investigated the monophyly of circumscribed sections and relationships among these sections. Seventy-one taxa, representing each of the six sections and about 80% of the recognized species in subgen. Iris, and 11 outgroup taxa were included in the study. Also included were five Asian species that share some morphological characteristics with subgen. Iris but are typically considered in other subgenera. Phylogenetic analyses of sequence data recovered six major clades but sects. Psammiris, Pseudoregelia, and Regelia, are not monophyletic as currently circumscribed. The sister clade to subgen. Iris is comprised of I. domestica and I. dichotoma, two beardless species that occur in eastern Asia. Iris verna, a species from the eastern United States is sister to I. domestica + I. dichotoma + subgen. Iris. The sepal of I. verna has pubescence but not a beard of multicellular trichomes.

opencc-zeroDec 2016View details →
dryad36/100

Data from: A novel method to analyze social transmission in chronologically sequenced assemblages, implemented on cultural inheritance of the art of cooking

Here we present an analytical technique for the measurement and evaluation of changes in chronologically sequenced assemblages. To illustrate the method, we studied the cultural evolution of European cooking as revealed in seven cook books dispersed over the past 800 years. We investigated if changes in the set of commonly used ingredients were mainly gradual or subject to fashion fluctuations. Applying our method to the data from the cook books revealed that overall, there is a clear continuity in cooking over the ages – cooking is knowledge that is passed down through generations, not something (re-)invented by each generation on its own. Looking at three main categories of ingredients separately (spices, animal products and vegetables), however, disclosed that all ingredients do not change according to the same pattern. While choice of animal products was very conservative, changing completely sequentially, changes in the choices of spices, but also of vegetables, were more unbounded. We hypothesize that this may be due a combination of fashion fluctuations and changes in availability due to contact with the Americas during our study time period. The presented method is also usable on other assemblage type data, and can thus be of utility for analyzing sequential archaeological data from the same area or other similarly organized material.

opencc-zeroDec 2014View details →
dryad36/100

Data from: Intraspecific DNA contamination distorts subtle population structure in a marine fish: decontamination of herring samples before restriction-site associated (RAD) sequencing and its effects on population genetic statistics

Wild specimens are often collected in challenging field conditions, where samples may be contaminated with the DNA of conspecific individuals. This contamination can result in false genotype calls, which are difficult to detect, but may also cause inaccurate estimates of heterozygosity, allele frequencies, and genetic differentiation. Marine broadcast spawners are especially problematic, because population genetic differentiation is low and samples are often collected in bulk and sometimes from active spawning aggregations. Here, we used contaminated and clean Pacific herring (Clupea pallasi) samples to test (i) the efficacy of bleach decontamination, (ii) the effect of decontamination on RAD genotypes, and (iii) the consequences of contaminated samples on population genetic analyses. We collected fin tissue samples from actively spawning (and thus contaminated) wild herring and non-spawning (uncontaminated) herring. Samples were soaked for 10 minutes in bleach or left untreated, and extracted DNA was used to prepare DNA libraries using a restriction-site associated DNA (RAD) approach. Our results demonstrate that intraspecific DNA contamination affects patterns of individual and population variability, causes an excess of heterozygotes, and biases estimates of population structure. Bleach decontamination was effective at removing intraspecific DNA contamination and compatible with RAD sequencing, producing high-quality sequences, reproducible genotypes, and low levels of missing data. Although sperm contamination may be specific to broadcast spawners, intraspecific contamination of samples may be common and difficult to detect from high-throughput sequencing data, and can impact downstream analyses.

opencc-zeroDec 2017View details →
dryad36/100

Data from: A stable phylogenomic classification of Travunioidea (Arachnida, Opiliones, Laniatores) based on sequence capture of ultraconserved elements

Molecular phylogenetics has transitioned into the phylogenomic era, with data derived from next-generation sequencing technologies allowing unprecedented phylogenetic resolution in all animal groups, including understudied invertebrate taxa. Within the most diverse harvestmen suborder, Laniatores, most relationships at all taxonomic levels have yet to be explored from a phylogenomics perspective. Travunioidea is an early-diverging lineage of laniatorean harvestmen with a Laurasian distribution, with species distributed in eastern Asia, eastern and western North America, and south-central Europe. This clade has had a challenging taxonomic history, but the current classification consists of ~77 species in three families, the Travuniidae, Paranonychidae, and Nippononychidae. Travunioidea classification has traditionally been based on structure of the tarsal claws of the hind legs. However, it is now clear that tarsal claw structure is a poor taxonomic character due to homoplasy at all taxonomic levels. Here, we utilize DNA sequences derived from capture of ultraconserved elements (UCEs) to reconstruct travunioid relationships. Data matrices consisting of 317–677 loci were used in maximum likelihood, Bayesian, and species tree analyses. Resulting phylogenies recover four consistent and highly supported clades; the phylogenetic position and taxonomic status of the enigmatic genus Yuria is less certain. Based on the resulting phylogenies, a revision of Travunioidea is proposed, now consisting of the Travuniidae, Cladonychiidae, Paranonychidae (Nippononychidae is synonymized), and the new family Cryptomastridae Derkarabetian &amp; Hedin, fam. n., diagnosed here. The phylogenetic utility and diagnostic features of the intestinal complex and male genitalia are discussed in light of phylogenomic results, and the inappropriateness of the tarsal claw in diagnosing higher-level taxa is further corroborated.

opencc-zeroDec 2017View details →
dryad36/100

Data from: Concealed by darkness: interactions between predatory bats and nocturnally migrating songbirds illuminated by DNA sequencing

Recently, several species of aerial-hawking bats have been found to prey on migrating songbirds, but details on this behaviour and its relevance for bird migration are still unclear. We sequenced avian DNA in feather-containing scats of the bird-feeding bat Nyctalus lasiopterus from Spain collected during bird migration seasons. We found very high prey diversity, with 31 bird species from eight families of Passeriformes, almost all of which were nocturnally flying sub-Saharan migrants. Moreover, species using tree hollows or nest boxes in the study area during migration periods were not present in the bats' diet, indicating that birds are solely captured on the wing during night-time passage. Additional to a generalist feeding strategy, we found that bats selected medium-sized bird species, thereby assumingly optimizing their energetic cost-benefit balance and injury risk. Surprisingly, bats preyed upon birds half their own body mass. This shows that the 5% prey to predator body mass ratio traditionally assumed for aerial hunting bats does not apply to this hunting strategy or even underestimates these animals' behavioural and mechanical abilities. Considering the bats' generalist feeding strategy and their large prey size range, we suggest that nocturnal bat predation may have influenced the evolution of bird migration strategies and behaviour.

opencc-zeroDec 2015View details →
dryad36/100

Data from: Genotyping by sequencing and genome–environment associations in wild common bean predict widespread divergent adaptation to drought

Drought will reduce global crop production by &gt;10% in 2050 substantially worsening global malnutrition. Breeding for resistance to drought will require accessing crop genetic diversity found in the wild accessions from the driest high stress ecosystems. Genome–environment associations in crop wild relatives reveal natural adaptation, and therefore can be used to identify adaptive variation. We explored this approach in the food crop Phaseolus vulgaris L., characterizing 86 geo-referenced wild accessions using Genotyping by Sequencing (GBS) to discover single-nucleotide-polymorphisms (SNPs). The wild beans represented Mesoamerica, Guatemala, Colombia, Ecuador/Northern Peru and Andean groupings. We found high polymorphism with a total of 22,845 SNPs across the 86 accessions loci that confirmed genetic relationships for the groups. As a second objective, we quantified allelic associations with a bioclimatic-based drought index using 10 different statistical models that accounted for population structure. Based on the optimum model, 115 SNPs in 90 regions, widespread in all 11 common bean chromosomes, were associated with the bioclimatic-based drought index. A gene coding for an Ankyrin repeat-containing protein and a phototropic-responsive NPH3 gene were identified as potential candidates. Genomic windows of 1Mb containing associated SNPs had more positive Tajima's D scores than windows without associated markers. This indicates that adaptation to drought, as estimated by bioclimatic variables, has been under natural divergent selection, suggesting that drought tolerance may be favorable under dry conditions but harmful in humid conditions. Our work exemplifies that genomic signatures of adaptation are useful for germplasm characterization, potentially enhancing future marker-assisted selection and crop improvement.

opencc-zeroDec 2017View details →
zenodo36/100

16S rRNA Sequence Data, Brazilian Coffee Soils

<p>16S sequencing data for DNA extracted from soils from Brazillian coffee farms. Sequenced on Illumina MiSeq with primers from Caporaso (2011, 2012).</p>

opencc-zeroAug 2014View details →
zenodo36/100

A Bayesian Approach to Detect Pedestrian Destination-Sequences from WiFi Signatures: Data (tech. report 2013)

<p>This dataset contains the data used in:</p> <p>Danalet, A., Farooq, B. and Bierlaire, M. (2013). A Bayesian Approach to Detect Pedestrian Destination-Sequences from WiFi Signatures, Technical report, Transport and Mobility Laboratory, ENAC, Ecole Polytechnique Fédérale de Lausanne, Lausanne. URL: http://infoscience.epfl.ch/record/189759 (full text available)</p> <p>It contains data and a technical report describing</p> <ul> <li>WiFi traces</li> <li>Pedestrian Semantically-Enriched Routing Graph (SERG), and</li> <li>Potential Attractivity measure (PAM).</li> </ul>

opencc-by-sa-4.0Mar 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record