Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,663

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,663 results for “bias”

Learn how ShareScore rates datasets ↗
zenodo44/100

Collider Bias Correction for Multiple Covariates in GWAS Using Robust Multivariable Mendelian Randomization

<p>This repository contains the data underlying the figures in paper "Collider Bias Correction for Multiple Covariates in GWAS<br>Using Robust Multivariable Mendelian Randomization".</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The file names and sheet names in the xlsx file indicate the corresponding figures of data.&nbsp;</p> <p><br>The underlying data of manhattan plots and QQ plots are in text file. For other figures, the underlying data are in the spreadsheet.</p> <p>In each file, column names indicate the MVMR method used to obtain the result.&nbsp;</p> <p>For example:&nbsp;</p> <p>In text files:</p> <p>The abbreviation "mPC" refers to metabolomic principle components.</p> <p>beta_no_correction: the SNP effect estimate without bias correction.</p> <p>beta_cml or beta_MVMR_cml: the standard error of SNP effect estimate after the bias correction of MVMR-cML.</p> <p>SE_UVMR_cml: the standard error of SNP effect estimate after the bias correction of UVMR-cML.</p> <p>p_value_Egger or p_value_MVMR_Egger: the p-value of SNP effect estimate after the bias correction of MVMR-Egger regression.</p> <p><br>In the spreadsheet, column names follow the same style.&nbsp;</p> <p>The GWAS data is also available. The column names follows the plink output file. The detailed explanations are available at https://www.cog-genomics.org/plink/2.0/formats#glm_linear</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Supplementary Data Files for the paper "Intrinsically disordered compositional bias in proteins: Sequence traits, region clustering, and generation of hypothetical functional associations"

<div> <div> <div> <div> <p><strong>Supplementary data files relating to <a href="https://doi.org/10.1177/11779322241287485">https://doi.org/10.1177/11779322241287485.&nbsp;</a></strong></p> <p><strong><span>Suppl. File 1: Protein Family Clusters.</span></strong></p> <p><strong><span>Suppl. File 2: Cluster GO enrichments/depletions. </span></strong></p> <p><strong><span>Suppl. File 3: The raw ID-CBR data with annotations. </span></strong></p> <p><strong><span>Suppl. File 4: &shy;ID-CBR Cluster membership.</span></strong></p> <p><strong><span>Each file has an explanatory header.&nbsp;</span></strong></p> <p>&nbsp;</p> </div> </div> </div> </div>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Biased ensembles of pulsating active matter: figure data

<p>The following is a zip file containing figure data for all figures published in the manuscript tittled 'Biased ensembles of pulsating active matter', available in archive: https://arxiv.org/abs/2403.16961</p> <p>v2 includes the updated data for figure 2. Otherwise all remain as before.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Most bacterial gene families are biased toward specific chromosomal positions

<p><span>The arrangement of genes along bacterial chromosomes influences their expression through growth rate-dependent gene copy number changes during DNA replication. While translation and transcription genes often cluster near the origin of replication, the extent of positional biases across gene families remains unclear. We hypothesized that natural selection broadly favors specific chromosomal positions to optimize growth rate-dependent expression. Analyzing 910 bacterial species and proteomics data from <em>Escherichia coli</em> and <em>Bacillus subtilis</em>, we find that about two-thirds of bacterial gene families are positionally biased, mainly near the origin or terminus of replication, with the strongest natural selection in fast-growing species. Our findings reveal chromosomal positioning as a fundamental mechanism for coordinating gene expression with growth rate, highlighting evolutionary constraints on bacterial genome architecture.</span></p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Global Surface Ozone Concentration Dataset 1990-2017 Generated by Bayesian Maximum Entropy Data Fusion With RAMP Bias Correction

<p>This dataset reports estimates of surface ozone concentration at fine spatial resolution for 1990 to 2017, at 0.5 degree horizontal resolution.&nbsp; Also reported is the variance.&nbsp; Estimates correspond to this paper:</p> <p><span>Becker, J. S.</span><span>, DeLang, M. N., K.-L. Chang, M. L. Serre, O. R. Cooper, <u>H. Wang</u>, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, M. Brauer, and <span>J. J. West</span> (2023) Using Regionalized Air Quality Model Performance and Bayesian Maximum Entropy data fusion to map global surface ozone concentration, <em>Elementa Science of the Anthropocene</em>, 11: 1, doi: 10.1525/elementa.2022.00025.</span></p> <p>The dataset reports estimates of surface ozone for the OSDMA8 metric (the 6-month ozone-season average of the daily maximum 8-hr concentration), estimated through a data fusion of ozone observations from the Tropospheric Ozone Assessment Report (TOAR) database, and output from multiple global atmospheric models.&nbsp; Estimates are created in each year by a combination of M3Fusion to create a multi-model composite, Regional Air Quality Model Performance (RAMP) regional and nonlinear bias correction, and Bayesian Maximum Entropy (BME) data fusion in space and time.&nbsp; The estimates here are the final results using a weighted RAMP bias correction.&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Bias-corrected EURO-CORDEX RCM simulations for the OPTAIN case studies

<p>Bias-corrected EURO-CORDEX RCM simulations are available on a daily timescale for:</p> <p>-period 1981-2099/2100,</p> <p>-6 RCM,</p> <p>-3 scenarios (RCPs 2.6, 4.5 and 8.5),</p> <p>-7 variables (mean, minimum and maximum temperature, precipitation, solar radiation, wind speed at 2 m and relative humidity) and</p> <p>-18 domains and 23 locations within these domains.</p> <p>Bias correction and further downscaling to 0.1&deg; was done using <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-land?tab=overview">ERA5-Land</a> reanalysis data with non-parametric empirical quantile mapping. Moreover, the interpolation of gridded bias-corrected climate model simulations to the locations was made using universal kriging.</p> <p><strong>Organization of the data</strong></p> <p>The name of the files are <em>domain</em>-<em>type</em>.zip, where <em>type</em> is gridded (NetCDF) or point (csv). Each zip file contains multiple files, organized in subfolders: <em>experiment</em>/<em>modelNumber</em>/<em>variable</em>.nc for gridded and <em>experiment</em>/<em>modelNumber</em>/<em>variable-pilotFieldNumber</em>.txt for point data, where <em>experiment </em>is rcp26, rcp45 or rcp85.</p> <p><em>domain and pilotFieldNumber</em></p> <table> <tbody> <tr> <td> <p><strong>domain</strong></p> </td> <td> <p><strong>domain </strong><strong>location (min and max. Longitude, min and max latitude</strong><strong>)</strong></p> </td> <td> <p><strong>pilotFieldNumber</strong></p> </td> <td> <p><strong>pilot field </strong><strong>location (longitude, latitude)</strong></p> </td> <td> <p><strong>case study</strong><strong> number</strong></p> </td> <td> <p><strong>country</strong></p> </td> <td> <p><strong>Name (OPTAIN case study)</strong></p> </td> </tr> <tr> <td> <p>01</p> </td> <td> <p>50.95 51.45 14.55 15.05</p> </td> <td>&nbsp;</td> <td> <p>&nbsp;</p> </td> <td> <p>1</p> </td> <td> <p>DEU</p> </td> <td> <p>Schoeps</p> </td> </tr> <tr> <td> <p>02</p> </td> <td> <p>46.35 47.05 6.55 7.15</p> </td> <td> <p>2</p> </td> <td> <p>46.816667 6.95</p> </td> <td> <p>2</p> </td> <td> <p>CHE</p> </td> <td> <p>Petite Glane</p> </td> </tr> <tr> <td> <p>02_1</p> </td> <td> <p>46.75 47.25 7.25 7.75</p> </td> <td> <p>1</p> </td> <td> <p>46.983333 7.466667</p> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>02_34</p> </td> <td> <p>47.35 47.85 8.35</p> </td> <td> <p>3</p> <p>4</p> </td> <td> <p>47.433333 8.516667</p> <p>47.683333 8.616667</p> </td> </tr> <tr> <td> <p>02_5</p> </td> <td> <p>46.15 46.65 5.95 6.45</p> </td> <td> <p>5</p> </td> <td> <p>46.4 6.233333</p> </td> </tr> <tr> <td> <p>03a</p> </td> <td> <p>46.65 47.15 17.45 17.95</p> </td> <td> <p>1</p> <p>2</p> <p>3</p> <p>4</p> </td> <td> <p>46.92649 17.68246</p> <p>46.9166 17.68976</p> <p>46.91283 17.69754</p> <p>46.91283 17.69723</p> </td> <td> <p>3a</p> </td> <td> <p>HUN</p> </td> <td> <p>Csorsza</p> </td> </tr> <tr> <td> <p>03b</p> </td> <td> <p>46.45 46.95 16.65 17.15</p> </td> <td>&nbsp;</td> <td> <p>&nbsp;</p> </td> <td> <p>3b</p> </td> <td> <p>HUN</p> </td> <td> <p>Felso Valicka</p> </td> </tr> <tr> <td> <p>04</p> </td> <td> <p>52.35 52.85 18.45 18.95</p> </td> <td> <p>1</p> </td> <td> <p>52.597469 18.728617</p> </td> <td> <p>4</p> </td> <td> <p>POL</p> </td> <td> <p>Upper Zglowiaczka</p> </td> </tr> <tr> <td> <p>05</p> </td> <td> <p>46.35 46.85 15.35 15.85</p> </td> <td>&nbsp;</td> <td> <p>&nbsp;</p> </td> <td> <p>5</p> </td> <td> <p>SVN</p> </td> <td> <p>Pesnica</p> </td> </tr> <tr> <td> <p>06</p> </td> <td> <p>46.45 46.95 16.15 16.65</p> </td> <td>&nbsp;</td> <td> <p>&nbsp;</p> </td> <td> <p>6</p> </td> <td> <p>HUN/SVN</p> </td> <td> <p>Kebele/Kobiljski</p> </td> </tr> <tr> <td> <p>07</p> </td> <td> <p>49.85 50.35 4.75 5.25</p> </td> <td>&nbsp;</td> <td> <p>&nbsp;</p> </td> <td> <p>7</p> </td> <td> <p>BEL</p> </td> <td> <p>La Wimbe</p> </td> </tr> <tr> <td> <p>08</p> </td> <td> <p>55.15 55.75 23.55 24.05</p> </td> <td> <p>1</p> <p>2</p> </td> <td> <p>55.522057 23.799235</p> <p>55.42233194 23.82580339</p> </td> <td> <p>8</p> </td> <td> <p>LTU</p> </td> <td> <p>Dotnuvele</p> </td> </tr> <tr> <td> <p>09</p> </td> <td> <p>45.45 45.95 9.65 10.15</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>9</p> </td> <td> <p>ITA</p> </td> <td> <p>Cherio</p> </td> </tr> <tr> <td> <p>10</p> </td> <td> <p>59.45 59.95 10.75 11.25</p> </td> <td> <p>1</p> <p>2</p> <p>3</p> <p>4</p> <p>5</p> <p>6</p> <p>7</p> <p>8</p> </td> <td> <p>59.71949 10.83576</p> <p>59.6833306 10.8833298</p> <p>59.6833306 10.8833298</p> <p>59.665 10.9475</p> <p>59.665 10.9475</p> <p>59.841012 10.903597</p> <p>59.757631 11.072031</p> <p>59.539623 10.856447</p> </td> <td> <p>10</p> </td> <td> <p>NOR</p> </td> <td> <p>Krogstad</p> </td> </tr> <tr> <td> <p>11</p> </td> <td> <p>46.45 46.95 17.55 18.05</p> </td> <td> <p>1</p> <p>2</p> </td> <td> <p>46.658333 17.75583</p> <p>46.656944 17.75833</p> </td> <td> <p>11</p> </td> <td> <p>HUN</p> </td> <td> <p>Tetves</p> </td> </tr> <tr> <td> <p>12</p> </td> <td> <p>49.35 49.85 14.75 15.25</p> </td> <td> <p>1</p> </td> <td> <p>49.616837 15.078266</p> </td> <td> <p>12</p> </td> <td> <p>CZE</p> </td> <td> <p>Cechticky</p> </td> </tr> <tr> <td> <p>13</p> </td> <td> <p>55.85 56.35 25.85 26.45</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>13</p> </td> <td> <p>LVA</p> </td> <td> <p>Dviete</p> </td> </tr> <tr> <td> <p>14</p> </td> <td> <p>59.75 60.25 17.55 18.05</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>&nbsp;</p> </td> <td> <p>14</p> </td> <td> <p>SWE</p> </td> <td> <p>Ingvastaan Lehstaan</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em>modelNumber</em></p> <table> <tbody> <tr> <td> <p><strong>modelNumber</strong></p> </td> <td> <p><strong>Driving Model (GCM)</strong></p> </td> <td> <p><strong>Ensemble</strong></p> </td> <td> <p><strong>RCM </strong></p> </td> <td> <p><strong>End date</strong></p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>EC-EARTH</p> </td> <td> <p>r12i1p1</p> </td> <td> <p>CCLM4-8-17</p> </td> <td> <p>31.12.2100</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>EC-EARTH</p> </td> <td> <p>r3i1p1</p> </td> <td> <p>HIRHAM5</p> </td> <td> <p>31.12.2100</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>HadGEM2-ES</p> </td> <td> <p>r1i1p1</p> </td> <td> <p>HIRHAM5</p> </td> <td> <p>30.12.2099</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>HadGEM2-ES</p> </td> <td> <p>r1i1p1</p> </td> <td> <p>RACMO22E</p> </td> <td> <p>30.12.2099</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p>HadGEM2-ES</p> </td> <td> <p>r1i1p1</p> </td> <td> <p>RCA4</p> </td> <td> <p>30.12.2099</p> </td> </tr> <tr> <td> <p>6</p> </td> <td> <p>MPI-ESM-LR</p> </td> <td> <p>r2i1p1</p> </td> <td> <p>REMO2009</p> </td> <td> <p>31.12.2100</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em>variable</em></p> <table> <tbody> <tr> <td> <p><strong>variable</strong></p> </td> <td> <p><strong>description</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>Tmean</p> </td> <td> <p>Mean temperature</p> </td> <td> <p>&deg;C</p> </td> </tr> <tr> <td> <p>Tmin</p> </td> <td> <p>Min temperature</p> </td> <td> <p>&deg;C</p> </td> </tr> <tr> <td> <p>Tmax</p> </td> <td> <p>Max temperature</p> </td> <td> <p>&deg;C</p> </td> </tr> <tr> <td> <p>prec</p> </td> <td> <p>Precipitation</p> </td> <td> <p>mm</p> </td> </tr> <tr> <td> <p>solarRad</p> </td> <td> <p>Solar radiation</p> </td> <td> <p>MJ/m2</p> </td> </tr> <tr> <td> <p>windSpeed</p> </td> <td> <p>Wind speed at 2m</p> </td> <td> <p>m/s</p> </td> </tr> <tr> <td> <p>relHum</p> </td> <td> <p>Relative humidity</p> </td> <td> <p>%</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Methodolody</strong></p> <p>Bias correction was done using non-parametric empirical quantile mapping with modified method from R package <a href="https://cran.r-project.org/web/packages/qmap/index.html">qmap</a>. Parameters selected were: corrections for each day of the year using a moving windows for a 31 days; 100 quantiles; wet days corrections for precipitation. The reference period is 1981-2010.</p> <p>The interpolation of gridded bias-corrected climate model simulations to the location was made using universal kriging&nbsp; with R packages <a href="https://cran.r-project.org/web/packages/automap/index.html">automap</a> and <a href="https://cran.r-project.org/web/packages/gstat/index.html">gstat</a> with (external) variables x, y, x2, y2, x*y, z, where x is latitude, y is longitude, and z is elevation. For Digital Elevation Model <a href="https://webmap.ornl.gov/wcsdown/dataset.jsp?dg_id=10008_1">Shuttle Radar Topography Mission</a> was used. If there was an error using above mentioned variables, the number of variables was reduced to x, y, x*y, z and if there was still an error to x, y, z.</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 862756.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Staff survey on awareness of gender bias in ATHENA RPOs and RFOs

<p>Task 2.3.1 from WP2 required the collection of data on awareness of gender bias to staff from the ATHENA RPOs and RFOs. An online survey was distributed among the staff of the ATHENA consortium institutions developing a gender equality plan (GEP). ATHENA institutions were requested that the sample was representative as much as possible, meaning this to consist proportionally of female/male, junior/senior positions and staff by occupations depending of the total staff composition in each institution.</p> <p>The aim of the staff survey is:</p> <ul> <li>To identify how aware are the respondents on gender equality in science and research institutions.&nbsp;</li> <li>To identify the biases/stereotypes related to the women&acute;s and men&acute;s role in science and research institutions.</li> <li>To identify gender imbalances and disadvantages in: <ul> <li>Recruitment and promotion,</li> <li>Gaining academic/scientific degree,</li> <li>Participation in decision making,</li> <li>Working conditions and workload,</li> <li>Work-life balance</li> <li>Experiences in harassment</li> </ul> </li> <li>The staff survey results will complement the results of the interviews and focus groups to provide a comprehensive picture of gender equality imbalances in the particular institution.</li> </ul> <p>The staff survey was realised through a standardised questionnaire developed by the WP2 coordinator (UVSK SAV) and approved by the Project Coordinator.</p> <p>The questionnaire consisted of 8 sections devoted to the particular gender areas- dimensions of interest:</p> <ul> <li>Introduction</li> <li>Information on the respondents &acute;current job</li> <li>Information of respondents&acute; background</li> <li>Opinions and perception of gender equality in research</li> <li>Recruitment and career development</li> <li>Striving for scientific/academic degree</li> <li>Gender balance in decision-making positions</li> <li>Workload and work-life balance</li> <li>Bulling and harassment</li> </ul> <p>Each section contained 4 &ndash; 8 closed and open questions.&nbsp;</p> <p>The results of the data served as support for the WP2 gender equality audit and assessment of procedures and practices at organizational level in D2.3 &ndash; Gender equality reports.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Male biased sex ratio in the offspring of roe deer

<p>The file&nbsp;&quot;metafor_roe_sex.csv&quot; was used for the meta-analysis regarding the sex ratio of roe deer juveniles. It contains the columns:&nbsp;</p> <p>Author (author(s) of the publication, Country (country were the study was publised), Location (specific location, in case that the data set contains several different locations the term &quot;divers&quot; is used),&nbsp;&nbsp;Year (year(s) were the sex of roe deer offspring were documented), Pub-Year (Year of publication)&nbsp;, N (number of offspring), N_female (number of female offspring), Sex ratio (primary - P or secondary - S sex ratio), Habitat conditions (F - free ranging, I - Island conditions, C - captive), Proportion female (proportion of female juveniles), Low_95 (lower boundary for the proportion of females using&nbsp;an exact binomial test with a&nbsp;95% confidence interval),&nbsp;High_95 (upper&nbsp;boundary for the proportion of females using&nbsp;an exact binomial test with a&nbsp;95% confidence interval), d (effect size), d_se (standard error of the effect size), North (Latitude), East (Longitude)</p> <p>The file &quot;sex_bw.csv&quot;&nbsp;contains information about roe deer juveniles tacked in Baden-W&uuml;rttemberg. The columns read as: Year (year were the juvenile was tacked), sex (the sex of the juvenile, m- male; f- female), and&nbsp;Hasl (elevation in m)</p> <p>The file &quot;temp_data_comma_sep.csv&quot; contains the montly mean values for temperature (Temp) and precipitation (NDS) for the months January, February, March, ..., December for the German federal state&nbsp;Baden-W&uuml;rttemberg (Source German Weather Service).</p> <p>All data sets were used for analysis presented in:&nbsp;Evidence for a male-biased sex ratio in the offspring of a large herbivore: the role of environmental conditions in the sex ratio variation.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Accurately modeling biased random walks on weighted networks using node2vec+ - Additional data

<p>Human gene interaction network data used to reproduce gene classification experiments&nbsp;https://github.com/krishnanlab/node2vecplus_benchmarks</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Data and code for: Precipitation Biases and Snow Physics Limitations Drive the Uncertainties in Macroscale Modeled Snow Water Equivalent

<p>Code and data&nbsp;to reproduce figures in manuscript entitled &quot;Precipitation Biases and Snow Physics Limitations Drive the Uncertainties in Macroscale Modeled Snow Water Equivalent&quot;&nbsp;published in&nbsp;Hydrology and Earth System Sciences (https://hess.copernicus.org/preprints/hess-2022-136/).</p> <p>The contents include three folders, &quot;Codes&quot;, &quot;Data&quot;,&nbsp;and &quot;Figures&quot;. In &quot;Codes&quot; folder, R scripts are listed in the order needed to reproduce the figures.&nbsp;All code is written in R version 4.2.0. Data sets needed to reproduce figures are provided in &quot;Data&quot; folder (Rdata format).&nbsp;The pdf files in &quot;Figures&quot; folder are outputs generated from the corresponding R scripts. Note that final figures&nbsp;in the article were produced by&nbsp;combining multiple&nbsp;figures&nbsp;using a&nbsp;vector graphics software (Inkscape) or PowerPoint. Please contact Eunsang Cho (<a href="mailto:eunsang.cho@nasa.gov">eunsang.cho@nasa.gov</a>) with any questions.&nbsp;</p> <p>Preferred citation:&nbsp;Cho, E., Vuyovich, C. M., Kumar, S. V., Wrzesien, M. L., Kim, R. S., and Jacobs, J. M. (2022). Precipitation Biases and Snow Physics Limitations Drive the Uncertainties in Macroscale Modeled Snow Water Equivalent, Hydrol. Earth Syst. Sci., https://doi.org/10.5194/hess-2022-136.</p> <p>Corresponding author: Eunsang Cho (<a href="mailto:eunsang.cho@nasa.gov">eunsang.cho@nasa.gov</a>;&nbsp;<a href="mailto:escho@umd.edu">escho@umd.edu</a>)</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

RemoTeC full-physics retrieval GOSAT/TANSO-FTS Level 2 bias-corrected XCO2 2009-2023, version 2.4.1 operated at Heidelberg University

<p>The data set contains bias-corrected column averaged dry air mole fractions (XCO2) retrieved with the RemoTeCv2.4.1 full-physics algorithm (Butz et al. 2011, Guerlet et al. 2013) applied on GOSAT TANSO-FTS Level 1B (L1B) data from 2009-04-18 to 2023-08-29. The GOSAT TANSO-FTS L1B data product is produced by JAXA/NOIES/MOE and provided by ESA. The XCO2 data together with related variables are aggregated as daily files, only good quality retrievals are included.</p> <p>If the data is used for publications, please contact andre.butz@uni-heidelberg.de to discuss potential co-authorship and technical details.</p> <p>&nbsp;</p> <p>Summary:</p> <p>Shortname: REMOTEC_L2_CO2_GOSAT</p> <p>Longname: RemoTeC full-physics retrieval GOSAT/TANSO-FTS Level 2 bias-corrected XCO2 version 2.4.1</p> <p>DOI: 10.5281/zenodo.12773070</p> <p>Version: 2.4.1</p> <p>Format: netCDF</p> <p>Spatial Coverage: -180.0,-90.0,180.0,90.0</p> <p>Temporal Coverage: 2009-04-18 to 2023-08-29</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data Potential Bias in Peer Review of Grant Applications at the Swiss National Science Foundation

<p>Potential biases in the peer review of grant applications at the Swiss National Science Foundation.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

CLDF dataset derived from Joo's "Phonosemantic Biases" from 2020

<p>Cite the source of the dataset as:</p> <blockquote> <p>Joo, I. (2020). Phonosemantic biases found in Leipzig-Jakarta lists of 66 languages. Linguistic Typology, 24(1), 1–12. https://doi.org/10.1515/lingty-2019-0030</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Phantom measurement data for 'Fast bias-corrected conductivity mapping using stimulated echoes', Iyyakkunnel et al. (2024)

<p>This dataset contains the phantom measurement data used in the article by Iyyakkunnel et al., titled "Fast Bias-Corrected Conductivity Mapping Using Stimulated Echoes," published in MAGMA, 2024 (doi: 10.1007/s10334-024-01194-3). In this study, the feasibility of using a stimulated echo sequence for electrical properties tomography (EPT) is demonstrated. The data were acquired with a 3T MRI system (Magnetom Prisma; Siemens Healthcare, Erlangen, Germany) using a dual-tuned 1H/23Na quadrature head coil for transmission and reception (Rapid Biomedical, Rimpar, Germany).<br>The dataset includes magnitude and phase measurements for the proposed Double-Angle Stimulated Echo (DA-STE) sequence, as well as reference measurements, including Double Angle measurements using a Gradient Echo sequence (GRE-DAM) for the B1+ magnitude, and a Single Echo Spin Echo sequence (SE) for the transceive phase.<br>For both the DA-STE and SE sequences, each measurement was repeated with inverted readout gradient polarities, denoted as LR (left-right) and RL (right-left) in the respective measurement folders. For each measurement, magnitude and phase data are provided in separate folders (in dicom (.dcm) format). Please note that for DA-STE, the two echo acquisitions are sequentially stored in the same measurement folder.<br>For further measurement details, please refer to the mentioned original article.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Supporting data for "Raising awareness of potential biases in medical machine learning: Experience from a Datathon"

<p>This archive contains files from a Datathon held virtually in February<br>2024 to introduce clinicians and data scientists to the challenge of<br>reviewing a clinical dataset for potential biases.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials

<p>Archive with data supporting the paper &quot;Improving the Efficiency of Variationally Enhanced Sampling with Wavelet-Based Bias Potentials&quot; and the related PhD thesis by B. Pampel</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Precipitation oxygen isoscape for mainland China from 1870 to 2017 generated based on data fusion and bias correction of iGCMs simulations

<p>The dataset includes the stable oxygen isotope of precipitation for the mainland of China over the 1870-2017 period, at a spatial resolution of 50-60 km and a monthly temporal resolution. In order to make&nbsp;full use of observations to integrate the advantages of various iGCMs, the combination of data fusion and bias correction methods are used.&nbsp;Some physical-based ancillary data are introduced in the fusion methods, including elevation and meteorological data, to enrich the climate and terrain information in the process of data fusion.&nbsp;Specifically,</p><p>(1) for the 1979-2001 period, nine simulations from six iGCMs (CAM2, GISS E, HadAM3, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused with observations by using the CNN fusion method;</p><p>(2) for the 2002-2007 period, seven simulations from four iGCMs (GISS E, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused by using the CNN fusion method;</p><p>(3) for the 1969-1978 period, four simulations from three iGCMs (CAM2, GISS E, and HadAM3) and ancillary data are fused by using the CNN fusion method;</p><p>(4) for the 1958-1968 and 2008-2017 periods, two iGCM simulations (CAM2 and HadAM3 for 1958-1968 and IsoGSM2 and LMDZ4 zoomed for 2008-2017) are corrected by using two BCMs, and ensemble mean (mean of four simulations) is then calculated;</p><p>(5) for the 1870-1957 period, one iGCM simulation (HadAM3) is corrected by using two BCMs, and the ensemble mean (mean of two simulations) is then calculated.</p><p>Compared with the existing iGCMs, the isoscape has high quality and stability for a large region in China at the monthly scale.&nbsp;However, it should be noted that the isoscape may be more reliable for the common periods of most iGCMs (1969-2007), but mediocre for other periods.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Data to support Whitney JL, Coleman RR, Deakos MH "Genomic evidence indicates small island-resident populations and sex-biased behaviors of Hawaiian Reef Manta Rays"

<p>Datasets supporting the manuscript: Whitney JL, Coleman RR, Deakos MH &quot;Genomic evidence indicates small island-resident populations and sex-biased behaviors of Hawaiian Reef Manta Rays&quot;. <em>BMC Ecology and Evolution&nbsp;</em><strong>23</strong>, 31 (2023). https://doi.org/10.1186/s12862-023-02130-0</p> <p>Nuclear data:</p> <p>&quot;Mobula-alfredi_nuclear_reference_RAD_contigs.fasta&quot; is a fasta of 359,751 contigs that serve as the reference for nuclear alignment of genotypes to RAD loci. Contigs begin and end with GATC cut site.</p> <p>Mobula-alfredi_nuclear_all_2048snps_38genotypes.vcf is a VCF file with all 2048 nuclear SNPs in final filtered SNP dataset. 38 genotypes are included from Maui Nui and Hawaii Island. This 2048 SNPs includes both 2038 neutral and 10 outlier SNPs.&nbsp;</p> <p>Mobula-alfredi_nuclear_neutral_2038snps_38genotypes.vcf&nbsp;is a VCF file with 2038 neutral nuclear SNPs genotyped in 38&nbsp;individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Mobula-alfredi_nuclear_outliers_10snps_38genotypes.vcf is a VCF file with 10 outlier SNPs genotyped in 38&nbsp;individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Structure (.str) files are also provided in addition to VCFs.&nbsp;In all files Population prefixes M=Maui Nui and K=Hawaii Island.&nbsp;</p> <p>Mitochondrial data:</p> <p>Mobula-alfredi_mitogenome_34haplotypes_9sites_min4x.vcf is a VCF file with 9 variant sites across the mitogenome haplotyped in 34 individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Mobula-alfredi_mitogenome_34haplotypes_allsites_min4x.fasta is a FASTA file with whole mitogenomes aligned to OP562409 [https://www.ncbi.nlm.nih.gov/nuccore/OP562409]. Sites with less than 4x coverage&nbsp;were masked with Ns.&nbsp;</p> <p>Mobula-alfredi_mitogenome_reference_OP562409.fasta is a FASTA file containing the <em>Mobula alfredi</em> reference mitogenome&nbsp;OP562409 [https://www.ncbi.nlm.nih.gov/nuccore/OP562409].</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Qbias – A Dataset on Media Bias in Search Queries and Query Suggestions

<p>We present Qbias, two novel datasets&nbsp;that promote the investigation of bias in online news search as described in</p> <blockquote> <p>Fabian Haak and Philipp Schaer. 2023. 𝑄𝑏𝑖𝑎𝑠 - A Dataset on Media Bias in Search Queries and Query Suggestions. In Proceedings of ACM Web Science Conference (WebSci&rsquo;23). ACM, New York, NY, USA, 6 pages.&nbsp;<a href="https://doi.org/10.1145/3578503.3583628">https://doi.org/10.1145/3578503.3583628</a>.</p> </blockquote> <p><strong>Dataset 1: AllSides Balanced News Dataset (allsides_balanced_news_headlines-texts.csv)</strong></p> <p>The dataset contains 21,747 news articles collected from <a href="https://www.allsides.com/headline-roundups">AllSides balanced news headline</a> roundups in November 2022 as presented in our publication. The AllSides balanced news feature three expert-selected U.S. news articles from sources of different political views (left, right, center), often featuring spin bias, and slant other forms of non-neutral reporting on political news. All articles are tagged with a bias label by four expert annotators based on the expressed political partisanship, left, right, or neutral. The AllSides balanced news aims to offer multiple political perspectives on important news stories, educate users on biases, and provide multiple viewpoints. Collected data further includes headlines, dates, news texts, topic tags (e.g., &quot;Republican party&quot;, &quot;coronavirus&quot;, &quot;federal jobs&quot;), and the publishing news outlet. We also include AllSides&#39; neutral description of the topic of the articles.<br> Overall, the dataset contains 10,273 articles tagged as left, 7,222 as right, and 4,252 as center.</p> <p>To provide easier access to the most recent and complete version of the dataset for future research, we provide a scraping tool and a regularly&nbsp;updated version of the dataset at <a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>. The repository also contains regularly updated more recent versions of the dataset with additional tags (such as the URL to the article). We chose to publish the version used for fine-tuning the models on Zenodo to enable the reproduction of the results of our study.&nbsp;</p> <p>&nbsp;</p> <p><strong>Dataset 2: Search Query Suggestions&nbsp;(suggestions.csv)</strong></p> <p>The second dataset we provide consists of 671,669 search query suggestions for root queries based on tags of the AllSides biased news dataset. We collected search query suggestions from Google and Bing for the 1,431 topic tags, that have been used for tagging AllSides news at least five times, approximately half of the total number of topics.&nbsp;The topic tags include names, a wide range of political terms, agendas, and topics (e.g., &quot;communism&quot;, &quot;libertarian party&quot;, &quot;same-sex marriage&quot;), cultural and religious terms (e.g., &quot;Ramadan&quot;, &quot;pope Francis&quot;), locations and other news-relevant terms.&nbsp;On average, the dataset contains 469 search queries for each topic.&nbsp;In total, 318,185 suggestions have been retrieved from Google and 353,484 from Bing.</p> <p>The file contains a &quot;root_term&quot; column based on the AllSides topic tags. The &quot;query_input&quot; column contains the search term submitted to the search engine (&quot;search_engine&quot;). &quot;query_suggestion&quot; and &quot;rank&quot; represents the search query suggestions at the respective positions returned by the search engines at the given time of search &quot;datetime&quot;. We scraped our data from a US server saved in &quot;location&quot;.</p> <p>We retrieved ten search query suggestions provided by the Google and Bing search autocomplete systems for the input of each of these root queries, without&nbsp;performing a search. Furthermore, we extended the root queries by the letters a to z (e.g., &quot;democrats&quot; (root term) &gt;&gt; &quot;democrats a&quot; (query input) &gt;&gt;&nbsp;&quot;democrats and recession&quot; (query suggestion)) to simulate a user&#39;s input during information search and generate a total of up to 270 query suggestions per topic and search engine. The dataset we provide contains columns for root term, query input, and query suggestion for each suggested query. The location from which the search is performed is the location of the Google servers running Colab, in our case Iowa in the United States of America, which is added to the dataset.&nbsp;</p> <p><strong>AllSides Scraper</strong></p> <p>At&nbsp;<a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>, we provide a scraping tool, that allows for the automatic retrieval of all available articles at the AllSides balanced news headlines.&nbsp;</p> <p>We want to provide an easy means of retrieving the news and all corresponding information. For many tasks it is relevant to have the most recent documents available. Thus, we provide this Python-based scraper, that scrapes all available AllSides news articles and gathers available information. By providing the scraper we facilitate access to a recent version of the dataset for other researchers.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Bengali Identity Bias Evaluation Dataset (BIBED)

<p>Critical studies found NLP systems to bias based on gender and racial identities. However, few studies focused on identities defined by cultural factors like religion and nationality. Compared to English, such research efforts are even further limited in major languages like Bengali due to the unavailability of labeled datasets. Our paper (see the reference) describes a process for developing a bias evaluation dataset highlighting cultural influences on identity. We also provide this&nbsp;Bengali dataset as an artifact outcome that can contribute to future critical research.</p> <p>If you find this dataset useful, please cite the associated paper:</p> <p>Das, D., Guha, S., &amp; Semaan, B. (2023, May). Toward Cultural Bias Evaluation Datasets: The Case of Bengali Gender, Religious, and National Identity. In&nbsp;<em>Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)</em>&nbsp;(pp. 68-83).</p> <p>BibTeX:</p> <pre>@inproceedings{das-etal-2023-toward, title = &quot;Toward Cultural Bias Evaluation Datasets: The Case of {B}engali Gender, Religious, and National Identity&quot;, author = &quot;Das, Dipto and Guha, Shion and Semaan, Bryan&quot;, booktitle = &quot;Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)&quot;, month = may, year = &quot;2023&quot;, address = &quot;Dubrovnik, Croatia&quot;, publisher = &quot;Association for Computational Linguistics&quot;, url = &quot;https://aclanthology.org/2023.c3nlp-1.8&quot;, pages = &quot;68--83&quot;, }</pre>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record