Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,663
datasets available to search
ShareScore release 0.9.0
Dataset results
1,663 results for “BIAS”
Data used in "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S."
<p>This dataset contains all relevant data used in the manuscript "Chemically specific sampling bias: the ratio of PM2.5 to surface AOD on average and peak days in the U.S." published in Environmental Science: Atmospheres.</p>
Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches
<p><strong>Background</strong></p> <p>The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.</p> <p><strong>Results</strong></p> <p>Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.</p> <p><strong>Conclusions</strong></p> <p>We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.</p>
Stim Circuits and Data for "Tailoring Dynamical Codes for Biased Noise: The X3Z3 Floquet code"
<p>Example stim circuits used for numerical simulations (for P6, XYZ2, CSS and X3Z3 Floquet codes) and threshold data obtained. See arXiv:2411.04974 for the associated manuscript.</p>
Data and R code for Reddin et al. 'Marine species and assemblage change foreshadowed by their thermal bias over Early Jurassic warming''
<p>This repository holds the raw and prepared datasets and R code to handle them for the manuscript Reddin et al. 'Marine species and assemblage change foreshadowed by their thermal bias over Early Jurassic warming'. Nature Communications</p>
Archived Model Output and Code for "Marine Boundary Layer Cloud Condensation Nuclei Bias over the Southern Ocean: Comparisons between the Community Atmosphere Model 6 and Field Observations "
<div> <p>This is an archive of CAM6 simulation output used in the paper Marine Boundary Layer Cloud Condensation Nuclei Bias over the Southern Ocean: Comparisons between the Community Atmosphere Model 6 and Field Observations, submitted to the AGU Journal. Codes used to read the nc file is also attached.</p> </div>
Experimental Data for the Paper "Frequency Fitness Assignment: Optimization without a Bias for Good Solutions can be Efficient"
<p><strong><em>The data for the paper "Frequency Fitness Assignment: Optimization without a Bias for Good Solutions can be Efficient"</em></strong></p> <p>This is the data set with the experimental results for our paper "Frequency Fitness Assignment: Optimization without a Bias for Good Solutions can be Efficient." We conduct more than 56 million runs, consuming more than 6.5*10<sup>14</sup> FEs as well as 150 processor years, ensuring that our results are statistically sound and rigorous. Here, you can find all the results, all the program codes used for obtaining the results, and all the tables and figures produced from the results, and the program codes used to produce them.</p> <p><strong><em>Included Files</em></strong></p> <ul> <li><code>ffa-empirical-complexity_results.tar.xz</code> (size packed 11.7 GiB, unpacked 335.2 GiB): The complete set of log files. For each run of each experiment, one distinct text-based log file is created. The log file contains every improving step of the algorithm, the final result, and the system configuration. This archive is very large and unpacked it will occupy more than 335 GiB of hard disk space.</li> <li><code>ffa-empirical-complexity_end_of_run_results_and_stats.tar.xz</code> (size packed 861.4 MiB, unpacked 6,295.9 MiB): The end-of-run result qualities and consumed runtime as well as statistics thereof. These information have been extracted from the log files and are provided in form of semicolon-separated values text files. These files are much easier to consume. They do not reflect the progress of the single runs, but only their end results.</li> <li><code>ffa-empirical-complexity_sources.tar.xz</code> (size packed 65.5 MiB, unpacked 173.6 MiB): The complete set of <code>Java</code> sources that was used to perform the experiments. Since the random seeds of the random number generators are created in a deterministic way, you could execute this code and obtain the exactly same log files in terms of consumed FEs, improving steps, and end results as we provide in <code>ffa-empirical-complexity_results.tar.xz</code>. (Of course, your systems configuration and measured runtime in milliseconds would probably be different.)</li> <li><code>ffa-empirical-complexity_evaluator.tar.xz</code> (size packed 4,015.8 KiB, unpacked 4,510 KiB): The <code>Java</code> and <code>R</code> source codes that are used to extract the end results from the log files, compute all relevant statistics, and produce the graphics in our article.</li> <li><code>ffa-empirical-complexity_evaluation.tar.xz</code> (size packed 417.3 KiB, unpacked 590.0 KiB): The high-level conclusions produced by the evaluator from the raw data, including tables and figures.</li> <li><code>ffa-empirical-complexity_saga_ffa_on_plateau.tar.xz</code> (size packed 7,238.9 KiB, unpacked 26.8 MiB): We also conducted an additional experiment to better understand the behavior of the SAGA algorithm variants using FFA on the Plateau problem. Here we provide source codes and result log files of this experiment. The result log files are much more comprehensive, as we tried to figure out why these algortihms were able to solve the Plateau problems (and they ultimately helped us to successfully do so).</li> </ul> <p><strong><em>License</em></strong></p> <p>The copyright holder of this dataset is Prof. Dr. Thomas Weise (see <a href="#contact">Contact</a>). The dataset is licensed under the <a href="https://creativecommons.org/licenses/by/4.0/en/legalcode">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong><em>Contact</em></strong></p> <p>If you have any questions or suggestions, please contact the corresponding author of this dataset, Prof. Dr. <a href="http://iao.hfuu.edu.cn/team/director">Thomas Weise</a> of the Institute of Applied Optimization (<a href="http://iao.hfuu.edu.cn/">IAO</a>) at <a href="http://www.hfuu.edu.cn/english/main.htm">Hefei University</a> [<a href="http://www.hfuu.edu.cn">合肥学院</a>] in Hefei, Anhui, China via email to <a href="mailto:tweise@hfuu.edu.cn">tweise@hfuu.edu.cn</a> with CC to <a href="mailto:tweise@ustc.edu.cn">tweise@ustc.edu.cn</a>.</p>
High rates of evolution preceded shifts to sex-biased gene expression in Leucadendron, the most sexually dimorphic angiosperms
<p>Differences between males and females are usually more subtle in dioecious plants than animals, but strong sexual dimorphism has evolved convergently in the South African Cape plant genus <i>Leucadendron</i>. Such sexual dimorphism in leaf size is expected largely to be due to differential gene expression between the sexes. We compared patterns of gene expression in leaves among ten <i>Leucadendron </i>species across the genus. Surprisingly, we found no positive association between sexual dimorphism in morphology and the number or the percentage of sex-biased genes. Sex bias in most sex-biased genes evolved recently and was species-specific. We compared rates of evolutionary change in expression for genes that were sex-biased in one species but unbiased in others and found that sex-biased genes evolved faster in expression than un-biased genes. This greater rate of expression evolution of sex-biased genes, also documented in animals, might suggest the possible role of sexual selection in the evolution of gene expression. However, our comparative analysis clearly indicates that the more rapid rate of expression evolution of sex-biased genes predated the origin of bias, and shifts towards bias were depleted in signatures of adaptation. Our results are thus more consistent with the view that sex bias is simply freer to evolve in genes less subject to constraints in expression level.</p>
Microsatellite genotype data from: Male-biased dispersal in a fungus-gardening ant symbiosis (Matthews et al, Ecology and Evolution)
<p>For nearly all organisms, dispersal is a fundamental life history trait that can shape their ecology and evolution. Variation in dispersal capabilities within a species exists and can influence population genetic structure and ecological interactions. In fungus-gardening (attine) ants, co-dispersal of ants and mutualistic fungi is crucial to the success of this obligate symbiosis. Female-biased dispersal (and gene flow) may be favored in attines because virgin queens carry the responsibility of dispersing the fungi, but a paucity of research has made this conclusion difficult. Here, we investigate dispersal of the fungus-gardening ant <i>Trachymyrmex septentrionalis</i> using a combination of maternally- (mitochondrial DNA) and biparentally-inherited (microsatellites) markers. We found three distinct, spatially isolated mitochondrial DNA haplotypes; two were found in the Florida panhandle and the other in the Florida peninsula. In contrast, biparental markers illustrated significant gene flow across this region and minimal spatial structure. The differential patterns uncovered from mitochondrial DNA and microsatellite markers suggest that most long-distance ant dispersal is male-biased and that females (and concomitantly the fungus) have more limited dispersal capabilities. Consequently, the limited female dispersal is likely an important bottleneck for the fungal symbiont. This bottleneck could slow fungal genetic diversification, which has significant implications for both ant hosts and fungal symbionts regarding population genetics, species distributions, adaptive responses to environmental change, and coevolutionary patterns.</p>
What's in a name? Taxonomic and gender biases in the etymology of new species names
<p>As our inventory of Earth's biodiversity progresses, the number of species given a Latin binomial name is also growing. While the coining of species names is bound by rules, the sources of inspiration used by taxonomists are an eclectic mix. We investigated naming trends for nearly 2900 new species of parasitic helminths described in the past two decades. Our analysis indicates that the likelihood of new species being given names that convey some information about them (name derived from morphology, host, or locality of origin) or not (named after an eminent scientist, or for something else) depends on the higher taxonomic group to which the parasite or its host belongs. We also found a consistent gender bias among species named after eminent scientists, with male scientists being immortalised disproportionately more frequently than female scientists. Finally, we found that the tendency for taxonomists to name new species after a family member or close friend has increased over the past twenty years. We end by formulating recommendations for future species naming, aimed at honouring the diverse scientific community regardless of gender or ethnicity and avoiding etymological nepotism and cronyism, while still allowing for creativity in crafting new Latin species names.</p>
Dataset on Off-Axis holography images of the MINEON device at different potential bias values
<p>Dataset of Off-Axis holography images that show to us how the electron beam's phase is modified aquiring an azimuthally changing phase profile, confirming the presence of an electron vortex beam. In this dataset we recorded phase images of the electron beam in at different values of the potential bias between the two main tips of the MINEON/chopstic electrostatic device. It is possible to see how the phase scales linearly with increasing potential bias difference, i.e., the electron vortex beam's OAM increases as the bias increases</p> <p>A description of this dataset can and similar ones are reported in:https://arxiv.org/abs/2203.00477</p>
Ignoring species availability biases occupancy estimates in single-scale occupancy models
<p>1. Most applications of single-scale occupancy models do not differentiate between availability and detectability, even though species availability is rarely equal to one. Species availability can be estimated using multi-scale occupancy models, and the availability process includes elements of species movement, behavior, and phenology. However, for the practical application of multi-scale occupancy models, it can be unclear what a robust sampling design looks like and what the statistical properties of the multi-scale and single-scale occupancy models are when availability is less than one.</p> <p>2. Using simulations, we explore the following common questions asked by ecologists during the design phase of a field study: (Q1) what is a robust sampling design for the multi-scale occupancy model when there are <i>a priori</i> expectations of parameter estimates?, (Q2) what is a robust sampling design when we have no expectations of parameter estimates?, and (Q3) can a single-scale occupancy model with a random effects term adequately absorb the extra heterogeneity produced when availability is less than one and provide reliable estimates of occupancy probability?.</p> <p>3. Our results show that there is a tradeoff between the number of sites and surveys needed to achieve a specified level of acceptable error for occupancy estimates using the multi-scale occupancy model. We also document that when species availability is low (< 0.40 on the probability scale), then single-scale occupancy models underestimate occupancy by as much as 0.40 on the probability scale, produce overly precise estimates, and provide poor parameter coverage. This pattern was observed when a random effects term was and was not included in the single-scale occupancy model, suggesting that adding a random-effects term does not adequately absorb the extra heterogeneity produced by the availability process. In contrast, when species availability was high (> 0.60), single-scale occupancy models performed similarly to the multi-scale occupancy model.</p> <p>4. As a companion, we provide an RShiny app that allows users to further explore our results and sampling designs across a number of different scenarios <a href="https://gdirenzo.shinyapps.io/multi-scale-occ/"><span>https://gdirenzo.shinyapps.io/multi-scale-occ/</span></a>. Our results suggest that unaccounted for availability can lead to underestimating species distributions when using single-scale occupancy models, which can have large implications on ecological inference and predictions for practitioners, such as those working at the front lines of invasion ecology, disease emergence, and species conservation. </p>
Herbarium specimens may provide biased flowering phenology estimates for dioecious species
<p>Dataset for manuscript on Lindera obtusiloba phenology using herbarium specimens. Contains two files: 1) CSV file with dataset "Yang_ea_2022 LinderaSpecimenData.csv" and 2) XLSX file with metadata "Yang_ea_2022 LinderaSpecimensREADME.xlsx"</p>
Robust biases in the estimation of passive yaw rotations
<p>We investigated the ability to estimate passive self-motion perception with and without external auditory sources of information. In the first experiment (1a), auditory cues were automatically delivered, while in the second experiment (1b) participants themselves generated the auditory cues. Here we share the raw dataset from the two experiments.</p>
Data from: Correcting a bias in the computation of behavioral time budgets that are based on supervised learning
<p>Supervised learning of behavioral modes from body-acceleration data has become a widely used research tool in Behavioral Ecology over the past decade. One of the primary usages of this tool is to estimate behavioral time budgets from the distribution of behaviors as predicted by the model. These serve as the key parameters to test predictions about the variation in animal behavior. In this paper we show that the widespread computation of behavioral time budgets is biased, due to ignoring the classification model confusion probabilities. Next, we introduce <em>the confusion matrix correction for time budgets</em> -- a simple correction method for adjusting the computed time budgets based on the model's confusion matrix. Finally, we show that the proposed correction is able to eliminate the bias, both theoretically and empirically in a series of data simulations on body acceleration data of a fossorial rodent species (Damaraland mole-rat, <em>Fukomys damarensis</em>). Our paper provides a simple implementation of <em>the confusion matrix correction for time budgets</em>, and we encourage researchers to use it to improve accuracy of behavioral time budget calculations.</p>
Data for 'Normalization procedure for obtaining the local density of states from high-bias scanning tunneling spectroscopy'
<p>This folder contains all the raw data needed to generate the figures in the paper '<em>Normalization procedure for obtaining the local density of states from high-bias scanning tunneling spectroscopy.</em>' The data are seperated by the figures in which they appear, with a text folder in each folder that contains any relevant additional information. </p>
Risk of bias assessments and support for judgement with ROB 2 tool for the Cochrane Review-Update: Systemic corticosteroids for the treatment of COVID-19.
<p>Risk of bias assessments and support for judgement with ROB 2 tool for the Cochrane Review-Update: Systemic corticosteroids for the treatment of COVID-19.</p>
Sperm limitation produces male biased family sex ratios
<p>Haplo-diploid sex determination in the parasitoid wasp, <em>Nasonia vitripennis</em> (Walker), allows females to adjust their brood sex ratios. Females influence whether ova are fertilized, producing diploid females, or remain unfertilized, producing haploid males. Females appear to adjust their brood sex ratios to minimize "local mate competition," <em>i.e.</em>, competition among sons for mates. Because mating occurs between siblings, females may optimize mating opportunities for their offspring by producing only enough sons to inseminate daughters when ovipositing alone and producing more sons when superparasitism is likely. Although widely accepted, this hypothesis makes no assumptions about gamete limitation in either sex. Because sperm is used to produce daughters, repeated oviposition could reduce sperm supplies, causing females to produce more sons. In contrast, if egg-limited females produce smaller broods, they might use fewer sperm, making sperm limitation less likely. To investigate whether repeated oviposition and female fertility influence gamete limitation within females, we created two treatments of six mated female wasps, which received a series of six hosts at intervals of 24 or 48 hrs. All females produced at least one mixed-sex brood (63 total broods; 3,696 offspring). As expected, if females became sperm limited, in both treatments, brood sex ratios became increasingly male-biased with increasing host number. The interhost interval did not affect brood size, total offspring number or sex ratio, indicating females did not become egg limited. Our results support earlier studies showing sperm depletion affects sex allocation in <em>N. vitripennis</em>¸ and could limit adaptive sex ratio manipulation in these parasitoid wasps.</p>
A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias
<p>Tropical ecosystems are often biodiversity hotspots, and invertebrates represent the main underrepresented component of diversity in large-scale analyses. This problem is partly related to the scarcity of data widely available to conduct these studies and the lack of systematic organization of knowledge about invertebrates' distributions in biodiversity hotspots. Here, we introduce and analyze a comprehensive data compilation of Amazonian ant diversity. Using records from 1817 to 2020 from both published and unpublished sources, we describe the diversity and distribution of ant species in the Brazilian Amazon Basin. Further, using high-definition images and data from taxonomic publications, we build a comprehensive database of morphological traits for the ant species that occur in the region. In total, we recorded 1,067 nominal species in the Brazilian Amazon Basin, with sampling locations strongly biased by access routes, urban centers, research institutions, and major infrastructure projects. Large areas where ant sampling is non-existent represent about 52% of the basin and are concentrated mainly in the North, Southeastern, and Western Brazilian Amazon. We found that distance to roads is the main driver of ant sampling in the Amazon. Contrary to our expectations, morphological traits had lower predictive power in predicting sample bias than purely geographic variables. However, when geographic predictors were controlled, habitat stratum and traits contribute to explain the remaining variance. More species were recorded in better-sampled areas, but species richness estimation models suggest that areas in South Amazonian edge forests are associated with especially high species richness. Our results represent the first trait-based, large-scale study for insects in Amazonian forests and a starting point for macroecological studies focusing on insect diversity in the Amazon Basin.</p>
Sex-biased admixture and assortative mating shape genetic variation and influence demographic inference in admixed Cabo Verdeans
<p>Inferred ROH and IBD calls from Korunes et al (2022). bioRxiv DOI: https://doi.org/10.1101/2020.12.14.422766</p> <p>Samples originally collected and analyzed in Beleza et al. 2013, PLoS Genetics. Inferred local ancestry information can be found at <a href="https://doi.org/10.5281/zenodo.4021277">https://doi.org/10.5281/zenodo.4021277</a></p> <p>See README.txt in upload for more detailed information.</p>
Exploring Gender Bias in Remote Pair Programming among Software Engineering Students: The Twincode Original Study and First External Replication (datasets)
<p>This repository contains the datasets of the original experiment (University of Seville, December 2021) and its first external replication (University of California, Berkeley, May 2022) of the Twincode exploratory study on the effects of gender bias in remote pair programming among software engineering students.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.