Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,663
datasets available to search
ShareScore release 0.7.1
Dataset results
1,663 results for “biases”
Widespread Sampling Biases in Herbaria Revealed from Large-Scale Digitization 1656-2016
Non-random collecting practices may bias conclusions drawn from analyses of herbarium records. Recent efforts to fully digitize and mobilize regional floras offer a timely opportunity to assess commonalities and differences in herbarium sampling biases. We determined spatial, temporal, trait, phylogenetic, and collector biases in ~5 million herbarium records, representing three of the most complete digitized floras of the world: Australia (AU), South Africa (SA), and New England, USA (NE) We identified numerous shared and unique biases among these regions. Shared biases included specimens i) collected close to roads and herbaria; ii) collected more frequently during spring; iii) of threatened species collected less frequently; and iv) of close relatives collected in similar numbers. Regional differences included i) over-representation of graminoids in SA and AU and of annuals in AU; and ii) peak collection during the 1910s in NE, 1980s in SA, and 1990s in AU. Finally, in all regions, a disproportionately large percentage of specimens were collected by a few individuals. These mega-collectors, and their associated preferences and idiosyncrasies, may have shaped patterns of collection bias via ‘founder effects’. Studies using herbarium collections should account for sampling biases and future collecting efforts should avoid compounding these biases.
Reward biases spontaneous neural reactivation during sleep
Open the record for dataset details and reuse information.
Sparse observations induce large biases in estimates of the global ocean CO2 sink: an ocean model subsampling experiment
<p>Dataset underlying the analysis in Hauck et al., 2023: Sparse observations induce large biases in estimates of the global ocean CO<sub>2</sub> sink - an ocean model subsampling experiment, Philosophical Transactions A</p> <p>Surface ocean partial pressure of CO<sub>2 </sub>(pCO<sub>2</sub>) and air-sea CO<sub>2</sub> flux reconstructions, using two mapping methods (MPI-SOM-FFN, CarboScope) three different sampling masks: SOCAT, SOCAT+SOCCOM, IDEAL (based on bgcArgo, Roemmich et al., 2019).</p> <p>Also, all FESOM-REcoM output fields that were used in the reconstructions are provided.</p> <p>We further provide the three masks that were used for subsampling: SOCAT, SOCAT+SOCCOM, IDEAL (bgcArgo).</p> <p> </p>
High-frequency measurements of chlorophyll fluorescence, characterizing F.I.Z. bias.
This dataset consists of high-frequency measurements of chlorophyll a fluorescence (Fchl), collected as part of manipulative experiments conducted in Lake George, NY. Experiments were set up to test for the potential of phototactic zooplankton to interfere with Fchl measurements. To test for any bias associated with fluorometer interference by zooplankton (FIZ), fluorometers were placed in the shallows of Lake George, and collected data under various treatment conditions. It was found that excitation light from fluorometers triggered a positive phototactic response during nighttime hours, biasing Fchl data by as much as 31x. Full results gleaned from this dataset can be found in: Moriarty, V.W., Lucius, M.A., Johnston, K.E., Borrelli, J.J., Mattes, B.M., Pezzuoli, A.R., Watson, C.D., Eichler, L.W. and Relyea, R.A. (2021), Fluorometer optical path interference via zooplankton phototaxis: Implications for high‐frequency data collection. Limnol Oceanogr Methods. https://doi.org/10.1002/lom3.10411
Dataset-Gender bias in magazines oriented to men and women: a computational approach
<p>This is the dataset associated with the research article 'Gender bias in magazines oriented to men and women: a computational approach' https://arxiv.org/abs/2011.12096</p>
RemoTeC full-physics retrieval GOSAT/TANSO-FTS Level 2 bias-corrected XCO2 version 2.4.0 operated at Heidelberg University
<p>The data set contains bias-corrected column averaged dry air mole fractions (XCO2) retrieved with the RemoTeCv2.4.0 full-physics algorithm (Butz et al. 2011, Guerlet et al. 2013) applied on GOSAT TANSO-FTS Level 1B (L1B) data from 2009-04-18 to 2019-06-30. The GOSAT TANSO-FTS L1B data product is produced by JAXA/NOIES/MOE and provided by ESA. The XCO2 data together with related variables are aggregated as daily files, only good quality retrievals are included.</p> <p> </p> <p>If the data is used for publications, please contact andre.butz@uni-heidelberg.de to discuss potential co-authorship and technical details.</p> <p>To cite the data in publications:</p> <p>André Butz (2019), RemoTeC full-physics retrieval GOSAT/TANSO-FTS Level 2 bias-corrected XCO2 version 2.4.0, Institute of Environmental Physics, Heidelberg University, Heidelberg, Germany, Accessed: [Date], 10.5281/zenodo.5886662</p> <p> </p> <p>Summary:</p> <p>Shortname: REMOTEC_L2_CO2_GOSAT</p> <p>Longname: RemoTeC full-physics retrieval GOSAT/TANSO-FTS Level 2 bias-corrected XCO2 version 2.4.0</p> <p>DOI: 10.5281/zenodo.5886662</p> <p>Version: 2.4.0</p> <p>Format: netCDF</p> <p>Spatial Coverage: -180.0,-90.0,180.0,90.0</p> <p>Temporal Coverage: 2009-04-18 to 2019-06-30</p>
RoCliB - Bias corrected CORDEX RCM dataset over Romania
<p>This dataset contains a set of four climate variables from 10 General Circulation Models (GCMs), dynamically downscaled in the EURO-CORDEX initiative by several Regional Climate Models (RCMs) and adjusted (bias-corrected) over Romania for the period 1971–2100. The climate models data were obtained from the <a href="https://cordex.org/data-access/">EURO-CORDEX archive</a>. Two climate change scenarios were selected, namely the moderate (RCP4.5) and business-as-usual scenario (RCP8.5). The multivariate bias correction by the N-dimensional probability density method (MBCn) was used to bias correct the RCMs outputs [1], using as reference the ROCADA gridded dataset [2].</p> <p>Characteristic:</p> <ul> <li><strong>Climate variables</strong>: air temperature (tasAdjust - Celsius degree), maximum air temperature (tasmaxAdjust - Celsius degree), minimum air temperature (tasminAdjust - Celsius degree) and precipitation (prAdjust - mm)</li> <li><strong>Bias-correction method:</strong> multivariate bias correction (N-pdft)</li> <li><strong>The reference period used for bias correction: </strong>1971-2005</li> <li><strong>The observational dataset used as a reference for bias correction: </strong>ROCADAv1</li> <li><strong>Temporal resolution:</strong> daily</li> <li><strong>Temporal extent</strong>: <ul> <li>Historical: 1971-2005;</li> <li>RCP4.5 and RCP8.5: 2006-2100.</li> </ul> </li> <li><strong>Spatial resolution:</strong> 0.1 degrees (~10km)</li> <li><strong>Spatial extent:</strong> from 20.1 to 29.8°E and 43.5 to 48.4°N</li> <li><strong>File format: n</strong>etCDF, CF-1.4-compliant format using netCDF4 compression</li> <li><strong>Coordinate system: </strong>WGS 1984 (EPSG:4326)</li> <li><strong>Naming conventions: </strong><em>variablename</em>_ROU-11_<em>cmip5experiment</em>_<em>globalmodel</em>_<em>run</em>_r<em>egionalmodel</em>_<em>rcmversionid</em>_<em>timefrequency</em>_<em>starttime-endtime</em><em>.</em>nc</li> <li><strong>RMCs</strong> (Institution or working group, RCM Model, GCM Institute, GCM Driving): <ul> <li>Climate Limited-area Modelling Community (CLMcom) CCLM4-8-17 CNRM-CERFACSCNRM-CM5</li> <li>Royal Netherlands Meteorological Institute (KNMI) RACMO22E CNRM-CERFACS CNRM-CM5</li> <li>Swedish Meteorological and Hydrological Institute (SMHI) RCA4CNRM-CERFACS CNRM-CM5</li> <li>Climate Limited-area Modelling Community (CLMcom) CCLM4-8-17 ICHECEC-EARTH</li> <li>Swedish Meteorological and Hydrological Institute (SMHI) RCA4I CHECEC-EARTH</li> <li>Royal Netherlands Meteorological Institute (KNMI) RACMO22E ICHECEC-EARTH</li> <li>Danish Meteorological Institute (DMI) HIRHAM5 ICHECEC-EARTH</li> <li>Climate Limited-area Modelling Community (CLMcom) CCLM4-8-17 MPI-MMPI-ESM-LR</li> <li>Swedish Meteorological and Hydrological Institute (SMHI) RCA4 MPI-MMPI-ESM-LR</li> <li>Climate Service Center Germany (GERICS) REMO2015 NCC NorESM1-M</li> </ul> </li> </ul> <p><strong>The terms of use</strong> for RoCliB datasets are the same as those from the original EURO-CORDEX simulations obtained from ESGF servers: <a href="https://is-enes-data.github.io/cordex_terms_of_use.pdf">https://is-enes-data.github.io/cordex_terms_of_use.pdf</a>.</p> <p><strong>To access and visualize</strong> relevant facts and statistics about climate change based on the RoCliB datasets use <a href="http://suscap.meteoromania.ro/en/roclib">http://suscap.meteoromania.ro/en/roclib</a>.</p> <p><strong>Acknowledgement</strong><br> This work was supported by a grant from the Romanian National Authority for Scientific Research and Innovation, CCCDI-UEFISCDI, project number COFUND-SUSCROP-SUSCAP-2, within PNCDI III. We also acknowledge the World Climate Research Programme's Working Group on Regional Climate, and the Working Group on Coupled Modelling, former coordinating body of CORDEX and responsible panel for CMIP5.</p>
Data to "Object visibility, not energy expenditure, accounts for spatial biases in human grasp selection"
<p>This record contains experimental and analysis scripts (written in Matlab) as well as raw and processed data to reproduce the results shown in:</p> <p><strong>Maiello, G</strong>.<sup> †</sup>, Paulun, V. C.<sup> †</sup>, Klein, L. K. , & Fleming, R. W. (2018) Object visibility, not energy expenditure, accounts for spatial biases in human grasp selection. <em>i-Perception,10</em>(1), 1–5. doi:10.1177/2041669519827608.</p> <p><sup>†</sup>co-first authors</p>
Variability and bias in measurements of metals mass fractions in automobile shredder residue
<p>Measured mass fractions of various metals in individually digested test samples of automobile shredder light fraction (single_digestions_ppm.csv) and the calculated means and standard deviations of these (mean_sd_ppm.csv). For all metadata see accompanying readme file Loevik2019_metal_mass_fractions_in_automobile_SLF_Readme.txt.</p>
Data and software supporting the manuscript 'The population frequency of human mitochondrial DNA variants is highly dependent upon mutational bias'
<p>Next-generation sequencing can quickly reveal genetic variation potentially linked to heritable disease. As databases encompassing human variation continue to expand, rare variants have been of high interest, since the frequency of a variant is expected to be low if the genetic change leads to a loss of fitness or fecundity. However, the use of variant frequency when seeking genomic changes linked to disease remains very challenging. Here, we explore the role of selection in controlling human variant frequency using the HelixMT database, which encompasses hundreds of thousands of mitochondrial DNA (mtDNA) samples. We find that a substantial number of synonymous substitutions, which have no effect on protein sequence, were never encountered in this large study, while many other synonymous changes are found at very low frequencies. Further analyses of human and mammalian mtDNA datasets indicate that the population frequency of synonymous variants is predominantly determined by mutational biases rather than by strong selection acting upon nucleotide choice. Our work has important implications that extend to the interpretation of variant frequency for non-synonymous substitutions. </p> <p> </p>
First Steps towards a Risk of Bias Corpus of Randomized Controlled Trials
<p><strong>Abstract</strong></p> <p>Risk of bias (RoB) assessment of randomized clinical trials (RCTs) is vital to conducting systematic reviews. Manual RoB assessment for hundreds of RCTs is a cognitively demanding, lengthy process and is prone to subjective judgment. Supervised machine learning (ML) can help to accelerate this process but requires a hand-labelled corpus. There are currently no RoB annotation guidelines for randomized clinical trials or annotated corpora. In this pilot project, we test the practicality of directly using the revised Cochrane RoB 2.0 guidelines for developing an RoB annotated corpus using a novel multi-level annotation scheme. We report inter-annotator agreement among four annotators who used Cochrane RoB 2.0 guidelines. The agreement ranges between 0% for some bias classes and 76% for others. Finally, we discuss the shortcomings of this direct translation of annotation guidelines and scheme and suggest approaches to improve them to obtain an RoB annotated corpus suitable for ML.</p> <p> </p> <p><strong>Methods</strong></p> <p>The upload contains two zip files and a .json file.</p> <ul> <li>plain.html.zip</li> </ul> <p>Original corpus (n = 10) in .html format. The corpus was generated using the methodology described in the paper. Each .html file could be opened in any default text editor in any operating system or browser. A .html contains full text divided into several annotatable text parts. </p> <p> </p> <ul> <li>ann.json.zip</li> </ul> <p>The .zip contains RoB annotations conducted by the authors (R.H., M.S., K.G., R.C.). The annotation files are in .json format. Each .json is divided into two JSON objects and three JSON arrays. </p> <ol> <li>annotatable (object): Parts from the full-text document corresponding to the text parts from the plain .html files. </li> <li>metas (object): full-text document label</li> <li>entities (array): contains labelled entities. Each entity is linked to which part of the full-text it is linked to.</li> <li>relations (array)</li> <li>sources (array)</li> </ol> <p> </p> <ul> <li>annotations-legend.json</li> </ul> <p>This .json file contains entity and entity labels encoded to text legends. For example, entity class label "1_2_Yes_Good" is encoded as "e_113".</p> <p> </p> <p><strong>Resources</strong></p> <p>The code to parse annotations can be found on <a href="http:// https://github.com/anjani-dhrangadhariya">GitHub</a>.</p> <p> </p> <p><strong>Funding</strong></p> <p>HES-SO Valais-Wallis, Sierre, Switzerland</p>
Temperature logger deployment methods and irradiance-biased temperature data, King Abdullah University of Science and Technology, Red Sea, 2023.
Solar irradiance can offset the temperature recorded by underwater sensing instruments (aka "loggers"). We collected temperature and PAR (photosynthetic active radiation) data during two short-term in situ deployments on a shallow fringing reef adjacent to the King Abdullah University of Science and Technology (KAUST) in the Red Sea. The first deployment quantified the measurement bias due to solar heating over five days in February 2023 while the second compared the effect of different shading methods on logger performance over 24 hours in June 2023. We also recorded temperature in a controlled calibration bath in the lab with ten of the most widely used loggers to further assess their accuracy, response time, and intra-logger variation. Finally, to understand current practices of measuring temperature on coral reefs, we summarized logger deployment method details from a literature review of coral reef studies published from 2013 to 2022. Such details included how often loggers recorded the temperature, the depth where loggers were deployed, and whether the authors reported shading or protecting their loggers. This data package is complete and part of a larger project that aims to develop an instrument deployment framework for restoration-based reef monitoring, which includes instrument recommendations and deployment guidelines.
MiRoR7-P1- Disagreements in risk of bias assessment for randomised controlled trials included in more than one Cochrane systematic reviews: a research on research study using cross-sectional design
<p>dataset referring to </p> <p><strong>Disagreements in risk of bias assessment for randomised controlled trials included in more than one Cochrane systematic reviews: a research on research study using cross-sectional design</strong></p> <p> </p> <p> </p> <p>Lorenzo Bertizzolo<sup>1</sup>, Patrick M Bossuyt<sup>2</sup>, Ignacio Atal<sup>1, 5</sup>, Philippe Ravaud<sup>1, 3-6</sup>, Agnès Dechartres<sup>7</sup></p> <p> </p> <p><sup>1</sup> INSERM, U1153 Epidemiology and Biostatistics Sorbonne Paris Cité Research Center (CRESS), Methods of therapeutic evaluation of chronic diseases Team (METHODS), Paris, F-75004 France; Paris Descartes University, Sorbonne Paris Cité, France.</p> <p><sup>2</sup> Department of Clinical Epidemiology, Biostatistics and Bioinformatics, Academic Medical Center, University of Amsterdam, Netherlands.</p> <p><sup>3</sup> Centre d’Épidémiologie Clinique, Hôpital Hôtel Dieu, AP-HP (Assistance Publique des Hôpitaux de Paris), Paris, France.</p> <p><sup>4</sup> Faculté de Médecine, Université Paris Descartes, Sorbonne Paris Cité, Paris, France.</p> <p><sup>5</sup> Cochrane France, Paris, France</p> <p><sup>6</sup> Columbia University, Mailman School of Public Health, Department of Epidemiology, New York, USA</p> <p><sup>7</sup> Sorbonne Université, INSERM, Institut Pierre Louis de Santé Publique, Département Biostatistique, Santé Publique et Information Médicale, AP-HP, Hôpitaux Universitaires Pitié Salpêtrière – Charles Foix, Paris, France</p>
Data for "How Array Design Creates SNP Ascertainement Bias"
<p>The repository contains the raw SNP data in vcf format for the publication "How Array Design Creates SNP Ascertainment Bias". Note that the variants are <strong>not</strong> filtered at this timepoint. Samples starting with pl_ are pooled sequences of ~10 individuals while samples starting with i_ were individually sequenced. Please find detailed information about samples, raw sequencing data and SNP calling pipeline in the linked preprint (<a href="https://doi.org/10.1101/833541">https://doi.org/10.1101/833541</a>)/ publication (<a href="https://doi.org/10.1371/journal.pone.0245178">https://doi.org/10.1371/journal.pone.0245178</a>). In case you need additional information, please contact<a href="mailto:johannes.geibel@uni-goettingen.de"> johannes.geibel@uni-goettingen.de</a></p>
Raw and processed GO term data to support running GCEA analyses using ensemble-based nulls, as described in the manuscript, 'Overcoming bias in gene category enrichment analyses of brain-wide transcriptomic data'.
<p>Data to support a toolbox for performing gene category enrichment analyses, including against ensembles of null phenotypes.</p> <p>Descriptions of how these data files can be used for this purpose are in the documentation for the toolbox, at https://github.com/benfulcher/GCEA_FalsePositives</p>
Bias-corrected monthly precipitation data over South Siberia for 1979-2019
<p>Bias-<strong>C</strong>orrected <strong>P</strong>recipitation data over <strong>S</strong>outh <strong>S</strong>iberia (<strong>CPSS 1.2</strong>) contains monthly precipitation data for the area within the coordinates 50–65 N, 60–120 E for the period from January 1979 to December 2019. CPSS data were combined from monthly total precipitation data from ERA5 reanalysis European Centre for Medium-Range Weather Forecasts (Copernicus Climate Change…, 2017) and precipitation data records from ground weather stations (Il’in et al., 2013). The ERA5 data were scaled according to the derived scale coefficient. The linear scaling coefficient for each month and weather station were calculated and extrapolated to the study area using the ordinary kriging method. Data spatial resolution is 0.25° in the latitude and 0.25° in the longitude. CPSS reproduces the spatial variability of precipitation more precisely than can be done from the weather station observation network. The CPSS dataset will be useful for the study of extreme precipitation events and allow for more accurate hydrologic risk assessment at a regional level based on climate model results. Data provided in NetCDF (Network Common Data Form) format.</p> <p>Copernicus Climate Change Service (C3S), 2017. <em>ERA5: Fifth generation of ECMWF atmospheric reanalyses of the global climate.</em> Copernicus Climate Change Service Climate Data Store (CDS), Available at <a href="https://cds.climate.copernicus.eu/cdsapp#!/home"><em>https://cds.climate.copernicus.eu/cdsapp#!/home</em></a></p> <p>Il’yin, B.M., Bulygina, O.N., Bogdanova, E.G, Veselov, V.M. and Gavrilova, S.Y., 2013. <em>Dataset of monthly precipitation totals, with the elimination of systematic errors of precipitation gauges</em>. Available at <a href="http://meteo.ru/data/506-mesyachnye-summy-osadkov-s-ustraneniem-sistematicheskikh-pogreshnostej-osadkomernykh-priborov"><em>http://meteo.ru/data/506-mesyachnye-summy-osadkov-s-ustraneniem-sistematicheskikh-pogreshnostej-osadkomernykh-priborov</em></a></p>
Dataset supplementing the article Einhäuser, W., Methfessel, P., & Bendixen, A. (2017). Newly acquired audio-visual associations bias perception in binocular rivalry. Vision Research, 133, 121-129.
<p>This dataset supplements the publication<br> Einhäuser, W., Methfessel, P., & Bendixen, A. (2017). Newly acquired audio-visual associations bias perception in binocular rivalry. Vision Research, 133, 121-129. doi: 10.1016/j.visres.2017.02.001</p> <p>Use is free for scientific purposes, provided the aforementioned reference is appropriately cited.<br> Description of files<br> - conditionsByObserver.csv<br> contains for each of the 16 observers the color and grating direction that had been coupled to either the low-pitch or the high-pitch tone<br> column 1: observer number<br> column 2: color associated with low-pitch tone<br> column 3: color associated with high-pitch tone<br> column 4: drift direction associated with low-pitch tone<br> column 5: drift direction associated with high-pitch tone</p> <p>- conditionsByObserver.mat contains the same information as matlab variables (as four vectors/cell arrays with one entry per observer)</p> <p>- toneByBlockAndTrial.csv<br> contains the conditions for all 18 rivalry trials (6 rivalry blocks with 3 trials each) for each observer<br> column 1: observer number<br> column 2: block number<br> column 3: trial number<br> column 4: tone (low [pitch], high [pitch], none) played in this trial<br> Note that due to a technical error for observer #16, block 6 was presented first, followed by 1,2,3,4,5; for all other observers blocks were used in the order given (1,2,3,4,5,6).</p> <p>- toneByBlockAndTrial.mat contains the same information as a 16x6x3 matrix named toneByBlockAndTrial ; tones are coded numerically (1-low pitch,2-high pitch,3-none)</p> <p>- eyeTraces.mat contains three cell arrays of dimensions 16x6x3 (observer x rivalry block x rivalry trial) called xEye, oknGain, and timeSinceTrialStart;</p> <p>o each entry of xEye contains the horizontal eye position for<br> the respective trial in eye-tracker coordinates (which correspond to screen pixels, except that (1/1) is the upper right rather than the upper left and values increase from right to left due to the setup configuration)</p> <p>o oknGain contains the gain computed from these eye positions.</p> <p>o timeSinceTrialStart contains the time in seconds since onset of the trial</p> <p><br> For all variables, the sampling rate is 500 Hz, in eye-tracker coordinates the speed of the grating is 240 units/ms. Blinks were removed from both eye-data variables, fast-phases were removed from the gain data. Removed data were set to NaN in eye-data variables.</p> <p>- Matlab functions figure1d.m, figure 2.m, figure3.m and figure4.m compute raw versions of the aforementioned paper's figures from the datafiles to exemplify their usage.</p> <p>[Note: In the originally published version of the article, the first two means and their standard errors of section 3.3 were stated incorrectly. All figures and statistical analyses are based on the correct data].</p>
Dataset for the paper "Historical model biases in monthly high temperature anomalies indicate under-projection of future temperature extremes"
<div> <div>This repository holds data and scripts related to the revision of the paper entitled: <span>"Historical model biases in monthly high temperature anomalies indicate under-projection of future temperature extremes" </span>by Lei Duan, Lyssa M. Freese, Govindasamy Bala, and Ken Caldeira. <span>The paper is currently submitted for peer review. </span>Any questions regarding the data and paper could be sent to the corresponding author: Lei Duan (leiduan@carnegiescience.edu). </div> </div>
Navigating News Narratives: A Media Bias Analysis Dataset
<p>The prevalence of bias in the news media has become a critical issue, affecting public perception on a range of important topics such as political views, health, insurance, resource distributions, religion, race, age, gender, occupation, and climate change. The media has a moral responsibility to ensure accurate information dissemination and to increase awareness about important issues and the potential risks associated with them. This highlights the need for a solution that can help mitigate against the spread of false or misleading information and restore public trust in the media.</p><p><strong>Data description: </strong>This is a dataset for news media bias covering different dimensions of the biases: political, hate speech, political, toxicity, sexism, ageism, gender identity, gender discrimination, race/ethnicity, climate change, occupation, spirituality, which makes it a unique contribution. The dataset used for this project does not contain any personally identifiable information (PII).</p><p><strong>Data Format: </strong>The format of data is:</p><ul><li>ID: Numeric unique identifier.</li><li>Text: Main content.</li><li>Dimension: Categorical descriptor of the text.</li><li>Biased_Words: List of words considered biased.</li><li>Aspect: Specific topic within the text.</li><li>Label: Neutral, Slightly Biased , Highly Biased</li></ul><p><br><strong>Annotation Scheme: </strong>The annotation scheme is based on Active learning, which is Manual Labeling --> Semi-Supervised Learning --> Human Verifications (iterative process)</p><ul><li>Bias Label: Indicate the presence/absence of bias (e.g., no bias, mild, strong).</li><li>Words/Phrases Level Biases: Identify specific biased words/phrases.</li><li>Subjective Bias (Aspect): Capture biases related to content aspects.</li></ul><p><br><strong>List of datasets used : </strong>We curated different news categories like Climate crisis news summaries , occupational, spiritual/faith/ general using RSS to capture different dimensions of the news media biases. The annotation is performed using active learning to label the sentence (either neural/ slightly biased/ highly biased) and to pick biased words from the news.</p><p>We also utilize publicly available data from the following links. Our Attribution to others.</p><p> <strong>MBIC (media bias): </strong>Spinde, Timo, Lada Rudnitckaia, Kanishka Sinha, Felix Hamborg, Bela Gipp, and Karsten Donnay. "MBIC--A Media Bias Annotation Dataset Including Annotator Characteristics." arXiv preprint arXiv:2105.11910 (2021). <a href="https://zenodo.org/records/4474336">https://zenodo.org/records/4474336</a> </p><p><strong>Hyperpartisan news: </strong>Kiesel, Johannes, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast. "Semeval-2019 task 4: Hyperpartisan news detection." In Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 829-839. 2019. <a href="https://huggingface.co/datasets/hyperpartisan_news_detection">https://huggingface.co/datasets/hyperpartisan_news_detection</a> </p><p><strong>Toxic comment classification: </strong>Adams, C.J., Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski. 2017. "Toxic Comment Classification Challenge." Kaggle. <a href="https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge">https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge</a>.</p><p><strong>Jigsaw Unintended Bias: </strong>Adams, C.J., Daniel Borkan, Inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and Nithum. 2019. "Jigsaw Unintended Bias in Toxicity Classification." Kaggle. <a href="https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification">https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification</a>.</p><p><strong>Age Bias : </strong>Díaz, Mark, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. "Addressing age-related bias in sentiment analysis." In Proceedings of the 2018 chi conference on human factors in computing systems, pp. 1-14. 2018. <a href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/F6EMTS">Age Bias Training and Testing Data - Age Bias and Sentiment Analysis Dataverse (harvard.edu)</a></p><p><strong>Multi-dimensional news Ukraine: </strong>Färber, Michael, Victoria Burkard, Adam Jatowt, and Sora Lim. "A multidimensional dataset based on crowdsourcing for analyzing and detecting news bias." In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pp. 3007-3014. 2020. <a href="https://zenodo.org/records/3885351#.ZF0KoxHMLtV">https://zenodo.org/records/3885351#.ZF0KoxHMLtV</a> </p><p><strong>Social biases: </strong>Sap, Maarten, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. "Social bias frames: Reasoning about social and power implications of language." arXiv preprint arXiv:1911.03891 (2019). <a href="https://maartensap.com/social-bias-frames/">https://maartensap.com/social-bias-frames/</a> </p><p> </p><p><strong>Goal of this dataset :</strong>We want to offer open and free access to dataset, ensuring a wide reach to researchers and AI practitioners across the world. The dataset should be user-friendly to use and uploading and accessing data should be straightforward, to facilitate usage.</p><p><strong>If you use this dataset, please cite us.</strong></p><p>Navigating News Narratives: A Media Bias Analysis Dataset © 2023 by <a href="https://www.linkedin.com/in/shainaraza/">Shaina Raza, Vector Institute </a>is licensed under <a href="http://creativecommons.org/licenses/by-nc/4.0/?ref=chooser-v1">CC BY-NC 4.0 </a></p><p> </p>
User study data: Nudges to Mitigate Confirmation Bias during Web Search for Opinion Formation, automatic vs. reflective study
<p>Data of two user studies (282 and 307 participants), investigating the risks and benefits of warning labels with and without obfuscations to mitigate confirmation bias during web search on debated topics.</p> <p> </p> <p>Study Variables (study 1 and study 2)</p> <p> </p> <p> display_con: Search result display<br> - Study 1<br> - 1: targeted warning label with obfuscation<br> - 2: random warning label with obfuscation<br> - 3: regular (no intervention)<br> - Study 2<br> - 1: targeted warning label with obfuscation<br> - 2: targeted warning label without obfuscation<br> - 3: random warning label with obfuscation<br> - 4: random warning label without obfuscation<br> - 5: regular (no intervention)<br>- CRT_cat: Cognitive reflection<br> - 1: intuitive<br> - 2: analytic<br>- topic: Assigned debated topic<br> - 1: Is drinking milk healthy for humans? <br> - 2: Is homework beneficial?<br> - 3: Should people become vegetarian?<br> - 4: Should students have to wear school uniforms?<br>- clicksup_prop: Clicks on attitude-confirming (AC) search results (proportion of all clicks)<br>- clickwarn_prop: Clicks on warning label (WL) search results (proportion of all clicks)<br>- show_clicked: Clicks on show-button (number of clicks, only in conditions with obfuscation)<br>- accuracy_bias: Accuracy bias estimation (Difference between a) observed bias (as the proportion of attitude-confirming clicks) and b) perceived bias (reported in the post-interaction questionnaire and re-coded into values from 0 to 1), positive values indicate an overestimation of bias)<br>- att_change: Attitude change (Difference between attitude reported in the pre-interaction questionnaire and the post-interaction questionnaire. Negative values indicate an attitude change in the attitude-opposing direction, while positive values indicate an attitude strengthening in the attitude-supporting direction.)<br>- knowledge_1: Self-reported prior knowledge (Reported on a seven-point Likert scale ranging from non-existent to excellent as a response to how they would describe their knowledge on the topic they were assigned to)<br>- N_clicks: Cumulative clicks (Number of all clicks on search results)<br>- NFC: Need for Cognition (Mean response to 4-item subset of the NFC questionnaire)<br>- UX_usability: Usability (Mean of responses on a seven-point Likert scale to the module "usability"from the meCUE 2.0 questionnaire)<br>- UX_usefulness: Usefulness (Mean of responses on a seven-point Likert scale to the module "usefulness"from the meCUE 2.0 questionnaire)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.