Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
244
datasets available to search
ShareScore release 0.9.0
Dataset results
244 results for “Model Organisms”
Data from: genetic resources of macroalgae: development of an efficient method using microsatellite markers in non-model organisms
<p><span>Red and brown seaweeds are species with high ecological and economic importance. Here we report the feasibility of cost-effective molecular marker development in 6 species from different clades. Microsatellites markers of two brown seaweed species <em>Alaria esculenta</em>, <em>Pylaiella littoralis</em>, and of four red seaweed species <em>Calliblepharis jubata</em>, <em>Gracilaria gracilis</em>, <em>Gracilaria dura </em>and <em>Palmaria palmata</em> were identified and characterized using genomic sequences of Double-Digest Restriction site Associated DNA (ddRAD). A total of 64,623,186 reads were generated from two runs of multiplexed Illumina Miseq sequencing for which 30,636 reads containing microsatellites and 15,443 microsatellite loci with primers pairs were found. Five hundred seventy-six primers pairs were selected for amplification trials and levels of polymorphism. From the 338 that gave a positive amplification, 142 primers pairs were polymorphic. For genetic analyses two or three populations per species from 13 different geographic locations were used. A total of 28 usable polymorphic markers for <em>A. esculenta</em>, 18 for <em>P. littoralis</em>, 11 for <em>C. jubata</em>, 14 for <em>G. gracilis</em>, 21 for <em>G. dura </em>and 13 for <em>P. palmata </em>were developed. The overall number of alleles per locus ranged from 2 to 22. These 105 new microsatellite markers will be useful for further studies of population genetics, breeding programs and conservation genetics of these species. Compared with traditional approaches, our study yielded thousands of microsatellite loci in a short tim</span><span>e with affordable costs in six different species. This study based on ddRAD-sequencing for the development of microsatellite markers provides preliminary data u</span><span>sing a few individuals from two distinct populations on the genetic structure and reproduction mode of a non-model species as shown </span>with the detection of clonality for the two red algae, <em>C. jubata </em>and <em>G. dura</em> and the detection of highly genetically divergent populations corresponding probably to different cryptic species under the name of<em> P. littoralis</em>.</p>
Target enrichment of long open reading frames and ultraconserved elements to link microevolution and macroevolution in non-model organisms
Open the record for dataset details and reuse information.
Data from: Microplastic exposure is associated with epigenomic effects in the model organism Pimephales promelas (fathead minnow)
Open the record for dataset details and reuse information.
Can short-term data accurately model long-term environmental exposures? Investigating the multigenerational adaptation potential of Daphnia magna to environmental concentrations of organic ultraviolet filters
Open the record for dataset details and reuse information.
Data from: genetic resources of macroalgae: development of an efficient method using microsatellite markers in non-model organisms
Open the record for dataset details and reuse information.
Detailed global modelling of soil organic carbon in cropland, grassland and forest soils
<p>Supporting information of the paper: Morais, T.G., Teixeira, R.F.M., Domingos, T. 2019. Detailed global modelling of soil organic carbon in cropland, grassland and forest soils. PloS One.</p> <p>Version 2 includes raster files (.tif) for each land use class (including: Attainable SOC stock, mineralization rate, and fator K).</p>
Data from: Citizen science reveals unexpected continental-scale evolutionary change in a model organism
Organisms provide some of the most sensitive indicators of climate change and evolutionary responses are becoming apparent in species with short generation times. Large datasets on genetic polymorphism that can provide an historical benchmark against which to test for recent evolutionary responses are very rare, but an exception is found in the brown-lipped banded snail (Cepaea nemoralis). This species is sensitive to its thermal environment and exhibits several polymorphisms of shell colour and banding pattern affecting shell albedo in the majority of populations within its native range in Europe. We tested for evolutionary changes in shell albedo that might have been driven by the warming of the climate in Europe over the last half century by compiling an historical dataset for 6,515 native populations of C. nemoralis and comparing this with new data on nearly 3,000 populations. The new data were sampled mainly in 2009 through the Evolution MegaLab, a citizen science project that engaged thousands of volunteers in 15 countries throughout Europe in the biggest such exercise ever undertaken. A known geographic cline in the frequency of the colour phenotype with the highest albedo (yellow) was shown to have persisted and a difference in colour frequency between woodland and more open habitats was confirmed, but there was no general increase in the frequency of yellow shells. This may have been because snails adapted to a warming climate through behavioural thermoregulation. By contrast, we detected an unexpected decrease in the frequency of Unbanded shells and an increase in the Mid-banded morph. Neither of these evolutionary changes appears to be a direct response to climate change, indicating that the influence of other selective agents, possibly related to changing predation pressure and habitat change with effects on micro-climate.
Data from: Pan-African phylogeography of a model organism, the African clawed frog "Xenopus laevis"
The African clawed frog Xenopus laevis has a large native distribution over much of sub-Saharan Africa and is a model organism for research, a proposed disease vector, and an invasive species. Despite its prominent role in research and abundance in nature, surprisingly little is known about the phylogeography and evolutionary history of this group. Here we report an analysis of molecular variation of this clade based on 17 loci (one mitochondrial, 16 nuclear) in up to 159 individuals sampled throughout its native distribution. Phylogenetic relationships among mitochondrial DNA haplotypes were incongruent with those among alleles of the putatively female-specific sex-determining gene DM-W, in contrast to the expectation of strict matrilineal inheritance of both loci. Population structure and evolutionarily diverged lineages were evidenced by analyses of molecular variation in these data. These results further contextualize the chronology, and evolutionary relationships within this group, support the recognition of X. laevis sensu stricto, X. petersii, X. victorianus, and herein re-validated X. poweri as separate species. We also propose that portions of the currently recognized distributions of X. laevis (north of the Congo Basin) and X. petersii (south of the Congo Basin) be reassigned to X. poweri.
Data from: Utility of pooled sequencing for association mapping in non-model organisms
High density genome-wide sequencing increases the likelihood of discovering genes of major effect and genomic structural variation in organisms. While there is an increasing availability of reference genomes across broad taxa, the greatest limitation to whole-genome sequencing of multiple individuals continues to be the costs associated with sequencing. To alleviate excessive costs, pooling multiple individuals with similar phenotypes and sequencing the homogenized DNA (Pool-Seq) can achieve high genome coverage, but at the loss of individual genotypes. Although Pool-Seq has been an effective method for association mapping in model organisms, it has not been frequently utilized in natural populations. To extend bioinformatic tools for rapid implementation of Pool-Seq data in non-model organisms, we developed a pipeline called PoolParty and illustrate its effectiveness in genetic association mapping. Alignment expectations based on five pooled Chinook salmon (Oncorhynchus tshawytscha) libraries showed that approximately 48% genome coverage per library could be achieved with reasonable sequencing effort. We additionally examined male and female O. tshawytscha libraries to illustrate how Pool-Seq techniques can successfully map known genes associated with functional differences among sexes such as growth hormone 2. Finally, we compared pools of individuals of different spawning ages for each sex to discover novel genes involved with age at maturity in O. tshawytscha such as opsin4 and transmembrane protein19. While not appropriate for every system, Pool-Seq data processed by the PoolParty pipeline is a practical method for identifying genes of major effect in non-model organisms when high genome coverage is necessary and cost is a limiting factor.
Data from: Diversification in wild populations of the model organism Anolis carolinensis: a genome-wide phylogeographic investigation
The green anole (Anolis carolinensis) is a lizard widespread throughout the southeastern United States and is a model organism for the study of reproductive behavior, physiology, neural biology, and genomics. Previous phylogeographic studies of A. carolinensis using mitochondrial DNA and small numbers of nuclear loci identified conflicting and poorly supported relationships among geographically structured clades; these inconsistencies preclude confident use of A. carolinensis evolutionary history in association with morphological, physiological, or reproductive biology studies among sampling localities and necessitate increased effort to resolve evolutionary relationships among natural populations. Here, we used anchored hybrid enrichment of hundreds of genetic markers across the genome of A. carolinensis and identified five strongly supported phylogeographic groups. Using multiple analyses, we produced a fully resolved species tree, investigated relative support for each lineage across all gene trees, and identified mito-nuclear discordance when comparing our results to previous studies. We found fixed differences in only one clade—southern Florida restricted to the Everglades region—while most polymorphisms were shared between lineages. The southern Florida group likely diverged from other populations during the Pliocene, with all other diversification during the Pleistocene. Multiple lines of support, including phylogenetic relationships, a latitudinal gradient in genetic diversity, and relatively more stable long-term population sizes in southern phylogeographic groups, indicate that diversification in A. carolinensis occurred northward from southern Florida.
Data from: Genotyping-by-sequencing for estimating relatedness in non-model organisms: avoiding the trap of precise bias
There has been remarkably little attention to using the high resolution provided by genotyping-by-sequencing (i.e. RADseq and similar methods) datasets for assessing relatedness in wildlife populations. A major hurdle is the genotyping error, especially allelic dropout, often found in this type of dataset that could lead to downward-biased, yet precise, estimates of relatedness. Here we assess the applicability of genotyping-by-sequencing datasets for relatedness inferences given their relatively high genotyping error rates. Individuals of known relatedness were simulated under genotyping error, allelic dropout, and missing data scenarios based on an empirical ddRAD dataset, and their true relatedness was compared to that estimated by seven relatedness estimators. We found that an estimator chosen through such analyses can circumvent the influence of genotyping error, with the estimator of Ritland (1996) shown to be unaffected by allelic dropout and to be the most accurate when there is genotyping error. We also found that the choice of estimator should not rely solely on the strength of correlation between estimated and true relatedness as a strong correlation does not necessarily mean estimates are close to true relatedness. We also demonstrated how even a large SNP dataset with genotyping error (allelic dropout or otherwise) or missing data still performs better than a perfectly genotyped microsatellite dataset of tens of markers. The simulation-based approach used here can be easily implemented by others on their own genotyping-by-sequencing datasets to confirm the most appropriate and powerful estimator for their dataset.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
This study aimed to develop a set of SNP markers with high resolution and accuracy within the African buffalo. Such a set can be used, among others, to depict subtle population genetic structure for a better understanding of buffalo population dynamics. In total, 18.5 million DNA sequences of 76 bp were generated by next generation sequencing on an Illumina Genome Analyzer II from a reduced representation library using DNA from a panel of 13 African buffalo representative of the four subspecies. We identified 2534 SNPs with high confidence within the panel by aligning the short sequences to the cattle genome (Bos taurus). The average sequencing depth of the complete aligned set of reads was estimated at 5x, and at 13x when only considering the final set of putative SNPs that passed the filtering criterion. Our set of SNPs was validated by PCR amplification and Sanger sequencing of 15 SNPs. Of these 15 SNPs, 14 amplified successfully and 13 were shown to be polymorphic (success rate: 87%). The fidelity of the identified set of SNPs and potential future applications are finally discussed.
Data from: Molecular Inversion Probes for targeted resequencing in non-model organisms
Applications that require resequencing of hundreds or thousands of predefined genomic regions in numerous samples are common in studies of non-model organisms. However few approaches at the scale intermediate between multiplex PCR and sequence capture methods are available. Here we explored the utility of Molecular Inversion Probes (MIPs) for the medium-scale targeted resequencing in a non-model system. Markers targeting 112 bp of exonic sequence were designed from transcriptome of Lissotriton newts. We assessed performance of 248 MIP markers in a sample of 85 individuals. Among the 234 (94.4%) successfully amplified markers 80% had median coverage within one order of magnitude, indicating relatively uniform performance; coverage uniformity across individuals was also high. In the analysis of polymorphism and segregation within family, 77% of 248 tested MIPs were confirmed as single copy Mendelian markers. Genotyping concordance assessed using replicate samples exceeded 99%. MIP markers for targeted resequencing have a number of advantages: high specificity, high multiplexing level, low sample requirement, straightforward laboratory protocol, no need for preparation of genomic libraries and no ascertainment bias. We conclude that MIP markers provide an effective solution for resequencing targets of tens or hundreds of kb in any organism and in a large number of samples.
Soil dissolved organic carbon (DOC) machine learning model code
Open the record for dataset details and reuse information.
Archetypes of Blockchain-Based Business Models - Analyzed Organizations
<p>Here we present a sample with the names of the analyzed organizations or projects and their websites.</p>
Convergence in simulating global soil organic carbon by structurally different models after data assimilation
<p>This is the data for results shown in the article accepted by Global Change Biology: Convergence in simulating global soil organic carbon by structurally different models after data assimilation</p>
Model data for "A model study on investigating the sensitivity of aerosol forcing on the volatilities of semi-volatile organic compounds" by Irfan et al
<p><span>Abstract: </span>This dataset contains simulation results from global aerosol-climate model ECHAM-SALSA. These simulations were performed to study the sensitivity of simulated SOA mass, CCN and radiative forcings to the assumed volatility distributions of biogenic SOA precursor species. The study employed volatility basis set (VBS) approach to represent and simulate SOA in the atmosphere. The study involved a comparative analysis between finely resolved 9-bin VBS setup with a simplified 3-bin VBS setup. It also included how the SOA mass, CCN and radiative forcing are sensitive to the volatitility of individual VBS bins.</p> <p><span>Methods: </span>Global scale aerosol-climate model simulations were performed using ECHAM-SALSA. We performed three diferent simulations each using 9-bin and 3-bin VBS setups with the volatilities increased (VBSx10) and decreased (VBSx0.1) by one-order of magnitude with respect to the original volatility (VBSx1). Another set of six different simuations were performed by increasing and decreasing the volatilities of one VBS bin at a time while keeping the original volatilities of other bins.</p> <p><span>TechnicalInfo: </span>In this study, all the ECHAM-SALSA simulations used T63 spectral truncation and 47 hybrid sigma pressure levels in horizontal and vertical resolution respectively. Simulations were performed for the year 2010 with half a year spin-up. The data was simulated with 3-hourly output for the simulation period. We then calculated monthly means of the summer months from 3-hourly data for all the model values except CDNC. We used the the 3-hourly data to analyse CDNC from grids with cloud fraction ≥ 0.95. A detailed description of model simulations is given in the setup file of each simulation.</p> <p><span>TechnicalInfo: </span>The external URL leads to a bucket containing setup files, the complete dataset (post-processed NETCDF files) presented in the manuscript, and the python scripts used for data analysis for each of the simulations.</p>
MOFSimplify: Machine Learning Models with Extracted Stability Data of Three Thousand Metal-Organic Frameworks
<p>Solvent removal stability and thermal stability associated with structurally characterized metal organic frameworks.</p>
Data sets used in "Neural network emulation of the formation of organic aerosols based on the explicit GECKO-A chemistry model"
<p>The training, validation, and testing data sets for toluene, dodecane, and alpha-pinene models described in the manuscript. A link to the manuscript will be added here when it becomes available. All trajectories in the data sets were generated using GECKO-A. The source code for using the data sets can be found at https://github.com/NCAR/gecko-ml </p>
Modeling the Organic Carbon Oxidation and Redox Sequence under the Partial- Equilibrium Approach: a Discussion by means of a Semi-Analytical Solution
<p>Filename: Analytical_solution.xls<br> File format: Excel<br> Description: It contains the calculation of the analytical model, programmed by means of excel. The sheet "advection" contains the calculation of the advection model and the sheet "diffusion" that of the diffusion model. The sheet "Phreeqc results" contains the results of the Phreeqc batch model (see Batch.pqi). Input data are highlighted in yellow.</p> <p>Filename: Batch.pqi<br> File format: Phreeqc<br> Description: It contains the input file for the calculation of the batch model by Phreeqc. It requires the Phreeqc database file Phreeqc.dat</p> <p>Filename: PEA_om_adv.pqi<br> File format: Phreeqc<br> Description: It contains the input file for the numerical calculation of the advection model by Phreeqc using the PEA (Partial Equilibrium Assumption). It requires the Phreeqc database file Phreeqc.dat</p> <p>Filename: PEA_om_dif.pqi<br> File format: Phreeqc<br> Description: It contains the input file for the numerical calculation of the diffusion model by Phreeqc using the PEA (Partial Equilibrium Assumption). It requires the Phreeqc database file Phreeqc.dat</p> <p>Filename: KIN_om_adv.pqi<br> File format: Phreeqc<br> Description: It contains the input file for the numerical calculation of the advection model by Phreeqc using the fully kinetic approach. It requires the Phreeqc database file PhreeqcKIN.dat</p> <p>Filename: KIN_om_dif.pqi<br> File format: Phreeqc<br> Description: It contains the input file for the numerical calculation of the diffusion model by Phreeqc using the fully kinetic approach. It requires the Phreeqc database file PhreeqcKIN.dat</p> <p>Filename: phreeqc.dat<br> File format: Phreeqc database<br> Description: It contains thermodynamic data for the batch model (see Batch.pqi) and the PEA models (see PEA_om_adv.pqi and PEA_om_dif.pqi). It is the default thermodynamic database of Phreeqc, dat), from which we removed ammonium to avoid the unrealistic reduction of NO3- and N2 to NH4+</p> <p>Filename: phreeqcKIN.dat<br> File format: Phreeqc database<br> Description: It contains thermodynamic data for the fully kinetic models (see KIN_om_adv.pqi and KIN_om_dif.pqi). It is the default thermodynamic database of Phreeqc, dat), from which we removed the relevant equilibrium redox reaction, so that they can be treated kinetically.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.