Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,663

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,663 results for “BIAS”

Learn how ShareScore rates datasets ↗
dryad40/100

Population size differences can lead to biases in phylogenetic inference and introgression detection in the presence of purifying selection

<p>Phylogenetic reconstruction and introgression detection rely on an assumption about the probability distribution of gene tree topologies. Recently, evidence has emerged that population size differences can affect the probability distribution of gene tree topologies in the presence of purifying selection. Here, using the population genetic simulator SLiM, we provide evidence that in the presence of purifying selection, population size differences can lead to biases in phylogenetic inference. We also provide evidence that in the presence of purifying selection, population size differences can cause statistics used for introgression detection to exhibit patterns resembling those caused by introgression. In addition, we present a theoretical analysis showing that the occurrence of population size–dependent gene tree distributions is an inherent consequence of purifying selection. Our work underscores the importance of considering the potential confounding effect of purifying selection on phylogenetic inference and introgression detection.</p>

opencc-zeroFeb 2024View details →
zenodo40/100

Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News

<p>Data supporting "Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News".&nbsp;<br><br></p> <p>&nbsp;If you use this dataset in your own research, please cite this paper:</p> <p>```<br>@misc{abels2024mitigating,<br>&nbsp; &nbsp; &nbsp; title={Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News},&nbsp;<br>&nbsp; &nbsp; &nbsp; author={Axel Abels and Elias Fernandez Domingos and Ann Now&eacute; and Tom Lenaerts},<br>&nbsp; &nbsp; &nbsp; year={2024},<br>&nbsp; &nbsp; &nbsp; eprint={2403.08829},<br>&nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br>&nbsp; &nbsp; &nbsp; primaryClass={cs.HC}<br>}<br>```</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>column name</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>treatment</td> <td>identifier for the set of headlines presented to the participant</td> </tr> <tr> <td>trial</td> <td>trial/round in which the headline was presented&nbsp;</td> </tr> <tr> <td>arm</td> <td>which "arm" the headline was presented as (0=left, 1=middle, 2=right)</td> </tr> <tr> <td>advice</td> <td>the participant's response (0=very unlikely, 0.25=unlikely, 0.5=undecided, 0.75=likely, 1=very likely)</td> </tr> <tr> <td>genuine</td> <td>whether the headline was genuine (1) or altered (0)</td> </tr> <tr> <td>headline</td> <td>the headline as shown to the participant</td> </tr> <tr> <td>original</td> <td>the headline before a possible alteration</td> </tr> <tr> <td>expert_id</td> <td>participant's identifier</td> </tr> <tr> <td>sentiment</td> <td>whether the headline reported a negative (-1) or positive (1) outcome</td> </tr> <tr> <td>expert:ethnicity</td> <td>the participant's ethnicity</td> </tr> <tr> <td>expert:sex</td> <td>the participant's sex</td> </tr> <tr> <td>expert:age</td> <td>the participant's age</td> </tr> <tr> <td>outcome:white, outcome:black, outcome:young, outcome:old, outcome:male, outcome:female</td> <td>whether the headline reported a negative (-1) or positive (1) or neutral (0) outcome for the specified group</td> </tr> <tr> <td>trial_time</td> <td>how long the participant took to respond to the trial/round</td> </tr> </tbody> </table> <p><strong>abstract</strong><br>Individual and social biases undermine the effectiveness of human advisers by inducing judgment errors which can disadvantage protected groups. In this paper, we study the influence these biases can have in the pervasive problem of fake news by evaluating human participants' capacity to identify false headlines. By focusing on headlines involving sensitive characteristics, we gather a comprehensive dataset to explore how human responses are shaped by their biases. Our analysis reveals recurring individual biases and their permeation into collective decisions. We show that demographic factors, headline categories, and the manner in which information is presented significantly influence errors in human judgment. We then use our collected data as a benchmark problem on which we evaluate the efficacy of adaptive aggregation algorithms. In addition to their improved accuracy, our results highlight the interactions between the emergence of collective intelligence and the mitigation of participant biases.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Testing strengths, limitations and biases of current Pulsar Timing Arrays detection analyses on realistic data

<p>In this project, we carried out an extensive investigation of the performance of current Pulsar Timing Arrays (PTA) analyses on simulated PTA datasets where we modeled the gravitational waves (GW) signal as the incoherent superposition of sinusoidal signals from a cosmic population of super-massive black hole binaries (SMBHBs). Here we publish the dataset referred to in the paper as the <em><strong>SMBHB_set</strong></em>: 100 realisations of PTA datasets with 25 pulsars and a GWB signal of nominal amplitude 2.4e-15. For each realisation, listed from 1001 to 1100, this repository contains the pulsars .par and .tim files, the mcmc chain (sampling over 66 parameters: 60 for pulsars intrinsic noise parameters and log amplitude and slope of a common red noise process), and the SMBHB details for that specific realisation of the GWB. The columns of the <em>MBHB_list.dat</em> files contain (from left to right): log10 MBH chirp mass (in solar masses in the source frame), mass ratio, source redshift, log10 of GW fundamental mode (n=2) in the observer frame, phi and theta angles in the sky, source inclination angle, source polarisation angle, source initial orbital phase, source initial direction of periastron and source initial eccentricity.</p> <p>If you make use of any of this data, please cite:<br><br></p> <div>@article{ refId0,</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {{Valtolina, Serena} and {Shaifullah, Golam} and {Samajdar, Anuradha} and {Sesana, Alberto}},</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; title = {Testing strengths, limitations, and biases of current pulsar timing arrays&rsquo; detection analyses on realistic data},</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;DOI= "10.1051/0004-6361/202348084",</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url= "https://doi.org/10.1051/0004-6361/202348084",</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;journal = {A&amp;A},</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;year = 2024,</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;volume = 683,</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;pages = "A201",</div> <div>}</div>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Fig. 9 in Cryptic Speciation And Characteristics Of The Transition Bias Following An Example Of The Cytb Gene In Palearctic Mammals

Fig. 9. Variation of summarized tv/ts-index in micro- (1) and macromammals (2) depending on nucleotide substitution level. Thick lines illustrate exponential approximation.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Fig. 8 in Cryptic Speciation And Characteristics Of The Transition Bias Following An Example Of The Cytb Gene In Palearctic Mammals

Fig. 8. Variation of transition (upper lines) and transversion (lower lines) frequencies accordingly to substitution frequencies level. Think lines are empirical data for each subfamily/family, thick ones — polynomial approximations of averaged data.

opencc-by-4.0Mar 2024View details →
zenodo40/100

Fig. 2 in Cryptic Speciation And Characteristics Of The Transition Bias Following An Example Of The Cytb Gene In Palearctic Mammals

Fig. 2. Average frequencies of nucleotide substitutions frequencies (sub) and its standard errors of the three taxonomical levels in micromammals (black) and macromammals (gray).

opencc-by-4.0Mar 2024View details →
zenodo40/100

T a b l e 1 in Cryptic Speciation And Characteristics Of The Transition Bias Following An Example Of The Cytb Gene In Palearctic Mammals

T a b l e 1. Average (M), sample deviations (SD) of nucleotide substitution, ts/tv and F indexes of different taxonomical levels within 15 Palearctic mammal families/subfamilies

opencc-by-4.0Mar 2024View details →
zenodo40/100

Electrochemical data plotted in A. Fasano, V. Fourmond and C. Léger, « Outer-sphere effects on the O2 sensitivity, catalytic bias and catalytic reversibility of hydrogenases », Chem. Sc. (2024). doi: 10.1039/D4SC00691G

<p>Text files of the electrochemical data plotted in A. Fasano, V. Fourmond and C. L&eacute;ger, &laquo; Outer-sphere effects on the O2 sensitivity, catalytic bias and catalytic reversibility of hydrogenases &raquo;, Chem. Sc. (2024) (<a href="http://dx.doi.org/10.1039/D4SC00691G" target="_blank" rel="noopener">doi: 10.1039/D4SC00691G)</a></p>

opencc-by-4.0Mar 2024View details →
dryad40/100

Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior

<p>Capturing qualitative features of animal behavior requires recording occurrences of behavior over time. Continuous sampling is best for capturing brief behaviors, but can be very time consuming. Instantaneous sampling can reduce the amount of labor required, but can miss short-duration behaviors. We therefore synthesized these techniques by continuously sampling during randomly scattered time intervals; a technique we call piecewise continuous sampling. To optimize and test the efficacy of this technique, we collected a continuous behavioral dataset of harvester ant workers, and then we developed a protocol to estimate the amount of sampling time necessary to reconstruct the proportion of time animals spend in different behavioral states. This protocol finds the sample size needed for the variance of the sample to converge on the variation of the population. We then divided this estimated time into equal-duration intervals that were randomly distributed across the entire continuous dataset. Finally, we calculated both time-dependent and time-independent error from this sample. We found that 4 to 16 sampling intervals minimize both types of error simultaneously. This finding was robust to differences in underlying behavior and was validated with simulations, implying that this method could be used for many types of organisms.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Electrochemical data shown in A. Fasano, A. Jacq-Bailly, J. Wozniak, V. Fourmond, and C. Léger, « Catalytic Bias and Redox-Driven Inactivation of the Group B FeFe Hydrogenase CpIII », ACS Catalysis (2024). doi: 10.1021/acscatal.4c01352

<p>Text file of all the electrochemical data shown in the following paper: A. Fasano, A. Jacq-Bailly, J. Wozniak, V. Fourmond, and C. L&eacute;ger, &laquo; Catalytic Bias and Redox-Driven Inactivation of the Group B FeFe Hydrogenase CpIII &raquo;, ACS Catalysis (2024). <a href="dx.doi.org/10.1021/acscatal.4c01352" target="_blank" rel="noopener">doi: 10.1021/acscatal.4c01352</a></p>

opencc-by-4.0Apr 2024View details →
dryad40/100

Sampling bias exaggerates a textbook example of a trophic cascade

<p>Understanding trophic cascades in terrestrial wildlife communities is a major challenge because these systems are difficult to sample properly. We show how a tradition of nonrandom sampling has confounded this understanding in a textbook system (Yellowstone National Park) where carnivore [<em>Canis lupus</em> (wolf)] recovery is associated with a trophic cascade involving changes in herbivore [<em>Cervus canadensis</em> (elk)] behavior and density that promote plant regeneration. Long-term data indicate a practice of sampling only the tallest young plants overestimated regeneration of overstory aspen (<em>Populus tremuloides</em>) by a factor of 3-8 compared to random sampling because it favored plants taller than the preferred browsing height of elk and overlooked non-regenerating aspen stands. Random sampling described a trophic cascade, but it was weaker than the one that nonrandom sampling described. Our findings highlight the critical importance of basic sampling principles (e.g., randomization) for achieving an accurate understanding of trophic cascades in terrestrial wildlife systems.</p>

opencc-zeroNov 2021View details →
dryad40/100

The new normal? Redaction bias in biomedical science

<p>A concerning amount of biomedical research is not reproducible. Unreliable results impede empirical progress in medical science, ultimately putting patients at risk. Many proximal causes of this irreproducibility have been identified, a major one being inappropriate statistical methods and analytical choices by investigators. Within this, we formally quantify the impact inappropriate redaction beyond a threshold value in biomedical science. This is effectively truncation of a data-set by removing extreme data points, and we elucidate its potential to accidentally or deliberately engineer a spurious result in significance testing. We demonstrate that the removal of a surprisingly small number of data points can be used to dramatically alter a result. It is unknown how often redaction bias occurs in the broader literature, but given the risk of distortion to the literature involved, we suggest that it must be studiously avoided, and mitigated with approaches to counteract any potential malign effects to the research quality of medical science.</p>

opencc-zeroDec 2021View details →
zenodo40/100

State of biodiversity documentation in the Philippines: Metadata gaps, taxonomic biases, and spatial biases in the DNA barcode data of animal and plant taxa in the context of species occurrence data

<p>These files can be categorized into three groups: (1) raw datasets obtained from public databases (i.e., GBIF, BOLD, and GenBank), (2) manually edited files needed for parsing and analysis, and (3) supplementary files for spatial analysis. All are used in the examination of gaps and biases present in Philippine biodiversity data,&nbsp;which can direct research on the taxa and spatial regions that need more sampling.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Projected changes in droughts and extreme droughts in Great Britain are strongly influenced by the choice of drought index: UKCP18-based bias adjusted potential evapotranspiration

<p>Potential evapotranspiration calculated from the UKCP18 RCM ensemble using the Penman-Monteith method as implemented by Robinson et al. (2017) and bias adjusted using Lange et al. (2019). This dataset was used for analysis of future drought characteristics in Reyniers et al. (2022). Details on the bias adjustment of this potential evapotranspiration dataset, as well as bias-adjusted precipitation and temperature from the same climate model ensemble, can be found in Reyniers et al. (2025).</p> <p>---</p> <p>Reyniers, N., Osborn, T. J., Addor, N., and Darch, G.: Projected changes in droughts and extreme droughts in Great Britain strongly influenced by the choice of drought index, Hydrol. Earth Syst. Sci., 27, 1151&ndash;1171, https://doi.org/10.5194/hess-27-1151-2023, 2023.</p> <p>Reyniers, N., Zha, Q., Addor, N., Osborn, T. J., Forstenh&auml;usler, N., and He, Y.: Two sets of bias-corrected regional UK Climate Projections 2018 (UKCP18) of temperature, precipitation and potential evapotranspiration for Great Britain, Earth Syst. Sci. Data, 17, 2113&ndash;2133, https://doi.org/10.5194/essd-17-2113-2025, 2025.&nbsp;</p> <p>Robinson, E. L., Blyth, E. M., Clark, D. B., Finch, J., Rudd, A. C. (2017). Trends in atmospheric evaporative demand in Great Britain using high-resolution meteorological data. HESS, <em>21</em>(2), 1189-1224.</p> <p>Lange, S. (2019). Trend-preserving bias adjustment and statistical downscaling with ISIMIP3BASD (v1. 0). <em>GMD,</em> <em>12</em>(7), 3055-3070.</p>

opencc-by-4.0Feb 2022View details →
dryad40/100

Once an optimist, always an optimist? Studying cognitive judgment bias in mice

<p>This repository contains raw data and analysis code for the manuscript entitled "Once an Optimist, Always an Optimist? Studying Cognitive Judgment Bias in Mice" from Marko Bračić, Lena Bohn, Viktoria Siewert, Vanessa von Kortzfleisch, Holger Schielzeth, Sylvia Kaiser, Norbert Sachser, S. Helene Richter, accepted for publication in the journal Behavioral Ecology.</p> <p>The aim of the study was to investigate the causes and stability of cognitive judgment bias (aka "optimism").</p> <p>Individuals differ in the way they judge ambiguous information: some individuals interpret ambiguous information in a more optimistic, and others in a more pessimistic way. Over the past two decades, such "optimistic" and "pessimistic" cognitive judgement biases (CJBs) have been utilized in animal welfare science as indicators of animals' emotional states. However, empirical studies on their ecological and evolutionary relevance are still lacking.</p> <p>We, therefore, aimed at transferring the concept of "optimism" and "pessimism" to behavioral ecology and investigated the role of genetic and environmental factors in modulating CJB in mice, using an automated, touchscreen-based active choice paradigm. In addition, we assessed the temporal stability of individual differences in CJB.</p> <p><span></span></p> <p>We show that the chosen genotypes (C57BL/6J and B6D2F1N) and environments ("scarce" and "complex") did not have a statistically significant influence on the responses in the CJB test. By contrast, they influenced anxiety-like behavior (assessed in the elevated plus maze (EPM), an open field test (OFT), and a free exploration test (FET)) with C57BL/6J mice and mice from the "complex" environment displaying less anxiety-like behavior than B6D2F1N mice and mice from the "scarce" environment. As the selected genotypes and environments did not explain the existing differences in CJB, future studies might investigate the impact of other genotypes and environmental conditions on CJB, and additionally, elucidate the role of other potential causes like endocrine profiles and epigenetic modifications. Furthermore, we show that individual differences in CJB were repeatable over a period of seven weeks, suggesting that CJB represents a temporally stable trait in laboratory mice. Therefore, we encourage the further study of CJB within an animal personality framework.</p>

opencc-zeroApr 2022View details →
zenodo40/100

Bias-corrected CORDEX daily precipitation dataset for the Carpathian Region

<p>This dataset contains bias-corrected regional climate model (RCM) daily outputs for daily precipitation under the RCP8.5 scenario.</p> <p><br> The reference dataset is CARPATCLIM (Szalai et al., 2013) which covers the the Carpathian Region for the period 1961-2010.</p> <p>The dataset contains bias corrected daily outputs of the following high-resolution (0.11o) RCMs from the framework of EURO-CORDEX (Jacob et al., 2014) and Med-CORDEX (Ruti et al., 2016):<br> - ALADIN<br> - CCLM<br> - HIRHAM<br> - RACMO<br> - RCA4<br> - RegCM<br> - REMO<br> - WRF</p> <p>&nbsp;</p> <p>The dataset covers the following periods with grid spacing of 0.11o on a regular lon/lat grid (between latitudes 44&deg;N and 50&deg;N, and longitudes 17&deg;E and 27&deg;E):</p> <p>- 1976-2005</p> <p>- 2021-2050</p> <p>- 2070-2099</p> <p>&nbsp;</p> <p>File format: NetCDF</p> <p>All data have been created following the work of Mezghani et al. (2017).</p> <p>Using the dataset please cite the following reference paper (also further details are given there): <a href="https://doi.org/10.28974/idojaras.2020.1.2">https://doi.org/10.28974/idojaras.2020.1.2</a>.</p> <p>&nbsp;</p> <p>References:<br> Jacob, D., Petersen, J., Eggert, B., Alias, A., Christensen, O.B., Bouwer, L.M., Braun, A., Colette, A., D&eacute;qu&eacute;, M., Georgievski, G., Georgopoulou, E., Gobiet, A., Menut, L., Nikulin, G., Haensler, A., Hempelmann, N., Jones, C., Keuler, K., Kovats, S., Kr&ouml;ner, N., Kotlarski, S., Kriegsmann, A., Martin, E., van Meijgaard, E., Moseley, C., Pfeifer, S., Preuschmann, S., Radermacher, C., Radtke, K., Rechid, D., Rounsevel, M., Samuelsson, P., Somot, S., Soussana, J.-F., Teichmann, C., Valentini, R., Vautard, R., Weber, B. and Yiou, P. (2014) EURO-CORDEX New high resolution climate change projections for European impact research. Reg. Environ. Change, 14, 563&ndash;578. https://doi.org/10.1007/s10113-013-0499-2</p> <p><br> Mezghani, A., Dobler, A., Haugen, J.E., Benestad, R.E., Parding, K.M., Piniewski, M., Kardel, I. and Kundzewicz, Z.W. (2017) CHASE-PL Climate Projection dataset over Poland &ndash; bias adjustment of EURO-CORDEX simulations. Earth Syst. Sci. Data, 9, 905&ndash;925. https://doi.org/10.5194/essd-9-905-2017</p> <p><br> Ruti, P.M., Somot, S., Giorgi, F., Dubois, C., Flaounas, E., Obermann, A., Dell&#39;Aquila, A., Pisacane, G., Harzallah, A., Lombardi, E., Ahrens, B., Akhtar, N., Alias, A., Arsouze, T., Aznar, R., Bastin, S., Bartholy, J., B&eacute;ranger, K., Beuvier, J., Bouffies-Cloch&eacute;, S., Brauch, J., Cabos, W., Calmanti, S., Calvet, J.-C., Carillo, A., Conte, D., Coppola, E., Djurdjevic, V., Drobinski, P., Elizalde-Arellano, A., Gaertner, M., Gal&aacute;n, P., Gallardo, C., Gualdi, S., Goncalves, M., Jorba, O., Jordi, G., L&#39;Heveder, B., Lebeaupin-Brossier, C., Li, L., Liguori, G., Lionello, P., Maci&aacute;s, D., Nabat, P., Onol, B., Raikovic, B., Ramage, K., Sevault, F., Sannino, G., Struglia, M.V., Sanna, A., Torma, C. and Vervatis, V. (2016) MED-CORDEX initiative for Mediterranean climate studies. Bulletin of the American Meteorological Society, 97, 1187&ndash;1208. https://doi.org/10.1175/BAMS-D-14-00176.1</p> <p><br> Szalai, S., Auer, I., Hiebl, J., Milkovich, J., Radim, T., Stepanek, P., Zahradnicek, P., Bihari, Z., Lakatos, M., Szentimrey, T., Limanowka, D., Kilar, P., Cheval, S., Deak, Gy., Mihic, D., Antolovic, I., Mihajlovic, V., Nejedlik, P., Stastny, P., Mikulova, K., Nabyvanets, I., Skyryk, O., Krakovskaya, S.,Vogt, J., Antofie, T. and Spinoni, J. (2013) Climate of the Greater Carpathian Region. Final Technical Report. http://www.carpatclim-eu.org<br> &nbsp;</p>

opencc-by-nc-4.0Apr 2022View details →
zenodo40/100

Supplementary Material of "How Does Author Affiliation Affect Preprint Citation Count? Analyzing Citation Bias at the Institution and Country Level"

<p>The source code and dataset for the following paper:</p> <p>Nishioka, C., F&auml;rber, M., and Saier, T. How Does Author Affiliation Affect Preprint Citation Count? Analyzing Citation Bias at the Institution and Country Level. In Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2022 (JCDL &#39;22), 2022.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
dryad40/100

Temporal trends in the spatial bias of species occurrence records

<p>Large-scale biodiversity databases have great potential for quantifying long-term trends of species, but they also bring many methodological challenges. Spatial bias of species occurrence records is well recognized. Yet, the dynamic nature of this spatial bias - how spatial bias has changed over time - has been largely overlooked. We examined the spatial sampling bias of species occurrence records within multiple biodiversity databases in Germany and tested whether spatial bias in relation to land cover or land use (urban and protected areas) has changed over time. We focused our analyses on urban and protected areas as these represent two well-known correlates of sampling bias in biodiversity datasets. We found that the proportion of annual records from urban areas has increased over time while the proportion of annual records within protected areas has not consistently changed. Using simulations, we examined the implications of this changing sampling bias for estimation of long-term trends of species' distributions. When assessing biodiversity change, our findings suggest that the effects of spatial bias depend on how it affects sampling of the underlying land-use change drivers affecting species. Oversampling of regions undergoing the greatest degree of change, for instance near human settlements, might lead to overestimation of the trends of specialist species. For robust estimation of the long-term trends in species' distributions, analyses using species occurrence records may need to consider not only spatial bias, but also changes in the strength of spatial bias through time.</p>

opencc-zeroMay 2022View details →
zenodo40/100

Downscaled and bias corrected 10 km water withdrawal of China, 1981-2010

<p>This dataset contains</p> <p>(1) MonthlyDomIndcon_CN_0.1deg_1981-2010.nc -- the 0.1&deg;&nbsp;gridded&nbsp;monthly domestic &amp; industrial&nbsp;water consumption of China</p> <p>(2) MonthlyDomInduse_CN_0.1deg_1981-2010.nc -- the 0.1&deg;&nbsp;gridded&nbsp;monthly domestic &amp; industrial&nbsp;water withdrawal of China</p> <p>(3)&nbsp;MonthlyIrrigation_CN_0.1deg_1981-2010.nc --&nbsp; the 0.1&deg;&nbsp;gridded&nbsp;monthly irrigation water withdrawal of China</p> <p>These data are&nbsp;first spatially downscaled from Huang et al. (2018) (https://doi.org/10.5281/zenodo.1209296) based on population density, GDP and irrigation area, and then bias corrected against provincial-level statistics published by local water agencies. Refer to Huang et al. (2018) and Dong et al. (2022) for more details.</p> <p>Dong, N.,&nbsp;Wei, J.,&nbsp;Yang, M.,&nbsp;Yan, D.,&nbsp;Yang, C.,&nbsp;Gao, H., et al. (2022).&nbsp;Model estimates of China&#39;s terrestrial water storage variation due to reservoir operation.&nbsp;Water Resources Research,&nbsp;58, e2021WR031787.&nbsp;<a href="https://doi.org/10.1029/2021WR031787">https://doi.org/10.1029/2021WR031787</a></p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record