Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,663

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,663 results for “BIAS”

Learn how ShareScore rates datasets ↗
zenodo36/100

Dataset of the paper: "How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study"

<p>This replication package contains datasets and scripts related to the paper: "*How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study*"</p><p>&nbsp;</p><p>## Root directory</p><p>- `statistics.r`: R script used to compute the correlation between usage and downloads, and the RQ1/RQ2 inter-rater agreements</p><p>- `modelsInfo.zip`: zip file containing all the downloaded model cards (in JSON format)</p><p>- `script`: directory containing all the scripts used to collect and process data. For further details, see README file inside the script directory.</p><p>&nbsp;</p><p>## Dataset</p><p>- `Dataset/Dataset_HF-models-list.csv`: list of HF models analyzed</p><p>- `Dataset/Dataset_github-prj-list.txt`: list of GitHub projects using the *transformers* library</p><p>- `Dataset/Dataset_github-Prj_model-Used.csv`: contains usage pairs: project, model</p><p>- `Dataset/Dataset_prj-num-models-reused.csv`: number of models used by each GitHub project</p><p>- `Dataset/Dataset_model-download_num-prj_correlation.csv` contains, for each model used by GitHub projects: the name, the task, the number of reusing projects, and the number of downloads</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ1</p><p>- `RQ1/RQ1_dataset-list.txt`: list of HF datasets</p><p>- `RQ1/RQ1_datasetSample.csv`: sample set of models used for the manual analysis of datasets</p><p>- `RQ1/RQ1_analyzeDatasetTags.py`: Python script to analyze model tags for the presence of datasets. it requires to unzip the `modelsInfo.zip` in a directory with the same name (`modelsInfo`) at the root of the replication package folder. Produces the output to stdout. To redirect in a file fo be analyzed by the `RQ2/countDataset.py` script</p><p>- `RQ1/RQ1_countDataset.py`: given the output of `RQ2/analyzeDatasetTags.py` (passed as argument) produces, for each model, a list of Booleans indicating whether (i) the model only declares HF datasets, (ii) the model only declares external datasets, (iii) the model declares both, and (iv) the model is part of the sample for the manual analysis</p><p>- `RQ1/RQ1_datasetTags.csv`: output of `RQ2/analyzeDatasetTags.py`</p><p>- `RQ1/RQ1_dataset_usage_count.csv`: output of `RQ2/countDataset.py`</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ2</p><p>- `RQ2/tableBias.pdf`: table detailing the number of occurrences of different types of bias by model Task</p><p>- `RQ2/RQ2_bias_classification_sheet.csv`: &nbsp;results of the manual labeling</p><p>- `RQ2/RQ2_isBiased.csv`: file to compute the inter-rater agreement of whether or not a model documents Bias</p><p>- `RQ2/RQ2_biasAgrLabels.csv`: &nbsp;file to compute the inter-rater agreement related to bias categories</p><p>- `RQ2/RQ2_final_bias_categories_with_levels.csv`: for each model in the sample, this file lists (i) the bias leaf category, (ii) the first-level category, and (iii) the intermediate category</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ3</p><p>- `RQ3/RQ3_LicenseValidation.csv`: manual validation of a sample of licenses</p><p>- `RQ3/RQ3_{NETWORK-RESTRICTIVE|RESTRICTIVE|WEAK-RESTRICTIVE|PERMISSIVE}-license-list.txt`: lists of licenses with different permissiveness</p><p>- `RQ3/RQ3_prjs_license.csv`: for each project linked to models, among other fields it indicates the license tag and name</p><p>- `RQ3/RQ3_models_license.csv`: for each model, indicates among other pieces of info, whether the model has a license, and if yes what kind of license</p><p>- `RQ3/RQ3_model-prj-license_contingency_table.csv`: usage contingency table between projects' licenses (columns) and models' licenses (rows)</p><p>- `RQ3/RQ3_models_prjs_licenses_with_type.csv`: pairs project-model, with their respective licenses and permissiveness level</p><p>&nbsp;</p><p>## scripts</p><p>Contains the scripts used to mine Hugging Face and GitHub. Details are in the enclosed README</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

A simple method to estimate capture height biases at landbird banding stations: opportunities and limitations

<p>Mist-nets are one of the most important tools for the capture of wild birds in ornithological research. The probability of capturing birds may vary by net height, which may drive capture biases. Such biases are rarely estimated, likely because of the relatively high cost and effort associated with constructing and operating elevated mist-net rigs where multiple mist-nets are stacked above one another. Therefore, a low-cost and -effort method to collect capture height data may allow broader investigation and better accounting of potential bias in existing banding protocols. Here, we investigate whether recording net panel of capture (with net panels indicating capture height, e.g., "upper panel") in ground-level mist-nets provides sufficient information to estimate capture height biases and compare these estimations to those obtained with traditional elevated mist-net rigs. Of the 29 taxa analyzed, we detected elevated capture biases for 11 (37.9%) and ground-level capture biases for seven (24.1%). When compared to estimates derived from elevated mist-net rigs at the same study site, we found high agreement with ground-level biases (75.0%) and low agreement with elevated biases (23.1%). These results suggest panel height of ground-level nets is a reliable method to estimate ground-level biases; however, scale of sampling may influence elevated biases, particularly for species that center their activity at the mid-story. Recording panel height may be quickly integrated into a station's processing protocols and broader application may improve our understanding of these biases.</p>

opencc-zeroNov 2023View details →
dryad36/100

Data from: Evaluation of different bias correction methods for dynamical downscaled future projections of the California Current Upwelling System

<p class="Abstract">Biases in global Earth System Models (ESMs) are an important source of errors when used to obtain boundary conditions for regional models. Here we examine historical and future conditions in the California Current System (CCS) using three different methods to force the regional model: (1) interpolation of ESM output to the regional grid with no bias correction; (2) a "seasonally-varying" delta method that obtains a season-dependent mean climate change signal from the ESM for a 30-year future period; and (3) a "time-varying" delta method that includes the interannual variability of the ESM over the 1980–2100 period. To compare these methods, we use a high-resolution (0.1˚) physical-biogeochemical regional model to dynamically downscale an ESM projection under the RCP8.5 emission scenario. Using different downscaling methods, the sign of future changes agrees for most of the physical and ecosystem variables, but the spatial patterns and magnitudes of these changes differ, with the seasonal- and time-varying delta simulations showing more similar changes. Not correcting the ESM forcing leads to amplification of biases in some ecosystem variables as well as misrepresentation of the California Undercurrent and CCS source waters. In the non-bias corrected and time-varying delta simulations, most of the ecosystem variables inherit trends and decadal variability from the ESM, while in the seasonally-varying delta simulation, the future variability reflects the observed historical variability (1980–2010). Our results demonstrate that bias correcting the forcing prior to downscaling improves historical simulations and that the bias correction method may impact the spatial and temporal variability of future projections. </p>

opencc-zeroNov 2023View details →
zenodo36/100

Telemetry data from: Realized thermal niche approach eliminates temperature bias in 3 bioenergetic model estimates

<h4>Raw data for Ivanova et al paper in Ecology and Evolution</h4><p>Datafile is an .rds file.</p>

opencc-by-4.0Dec 2023View details →
dryad36/100

Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models

<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>

opencc-zeroDec 2023View details →
dryad36/100

Data from: Multi-generation genetic contributions of immigrants reveal cryptic elevated and sex-biased effective gene flow within a natural meta-population

<p>Impacts of immigration on micro-evolution and population dynamics fundamentally depend on net rates and forms of resulting gene flow into recipient populations. Yet, the degrees to which observed rates and sex ratios of physical immigration translate into multi-generational genetic legacies have not been explicitly quantified in natural meta-populations, precluding inference on how movements translate into effective gene flow and eco-evolutionary outcomes. Our analyses of three decades of complete song sparrow (<em>Melospiza melodia</em>) pedigree data show that multi-generational genetic contributions from regular natural immigrants substantially exceeded those from contemporary natives, consistent with heterosis-enhanced introgression. Further, while contributions from female immigrants exceeded those from female natives by up to three-fold, male immigrants' lineages typically went locally extinct soon after arriving. Both the overall magnitude, and the degree of female bias, of effective gene flow therefore greatly exceeded those which would be inferred from observed physical arrivals, reshaping the eco-evolutionary implications of immigration.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Polski frontendu and American Back-end: GitHub Profile Recruitment Bias in GPT-4

<p>This repository serves as the online appendix for the paper "Polski frontendu and American Back-end: GitHub Profile Recruitment Bias in GPT-4".</p>

opencc-by-4.0Dec 2023View details →
dryad36/100

Sex-biased gene content associates with sex chromosome turnover in Danaini butterflies

<p>Sex chromosomes play an outsized role in adaptation and speciation, and thus deserve particular attention in evolutionary genomics. In particular, fusions between sex chromosomes and autosomes can produce neo-sex chromosomes, which offer important insights into the evolutionary dynamics of sex chromosomes. Here we investigate the evolutionary origin of the previously reported <em>Danaus</em> neo-sex chromosome within the tribe Danaini. We assembled and annotated genomes of <em>Tirumala septentrionis</em> (subtribe Danaina), <em>Ideopsis similis</em> (Amaurina), <em>Idea leuconoe</em> (Euploeina), and <em>Lycorea halia</em> (Itunina) and identified their Z-linked scaffolds. We found that the <em>Danaus </em>neo-sex chromosome resulting from the fusion between a Z chromosome and an autosome corresponding to the <em>Melitaea cinxia</em> chromosome (McChr) 21 arose in a common ancestor of Danaina, Amaurina, and Euploina. We also identified two additional fusions as the W chromosome further fused with the synteny block McChr31 in <em>I. similis</em> and independent fusion occurred between the ancestral Z chromosome and McChr12 in <em>L. halia</em>. We further tested a possible role of sexually antagonistic selection in sex chromosome turnover by analyzing the genomic distribution of sex-biased genes in <em>I. leuconoe</em> and <em>L. halia</em>. The autosomes corresponding to McChr21 and McChr31 involved in the fusions are significantly enriched in female- and male-biased genes, respectively, which could have hypothetically facilitated fixation of the neo-sex chromosomes. This suggests a role of sexual antagonism in sex chromosome turnover in Lepidoptera. The neo-Z chromosomes of both <em>I. leuconoe</em> and <em>L. halia</em> appear fully compensated in somatic tissues, but the extent of dosage compensation for the ancestral Z is variable across tissues and species.</p>

opencc-zeroDec 2023View details →
dryad36/100

Moss functional trait ecology: Trends, gaps, and biases in the current literature

<p>Functional traits are critical tools in plant ecology for capturing organism-environment interactions based on trade-offs as well as making links between organismal and ecosystem processes. While broad frameworks for functional traits have been developed for vascular plants, we lack the same for bryophytes, despite an escalation in the number of bryophyte functional trait studies conducted in the last 45 years and an increased recognition of the ecological roles bryophytes play across ecosystems. In this review, we compiled data from 282 published articles (10005 records) focusing on functional traits measured in mosses, and sought to (a) examine trends in types of traits measured, (b) capture taxonomic and geographic breadth of trait coverage, (c) reveal biases in coverage in the current literature, and (d) develop a bryophyte-function index (BFI) to describe completeness of current trait coverage and identify global gaps to focus research efforts. The most commonly measured response traits (those related to growth/reproduction in individual organisms) and effect traits (those that directly affect community/ecosystem scale processes) fell into the categories of morphology (e.g. leaf area, shoot height) and nutrient storage/cycling, and our BFI revealed that these data were most commonly collected from temperate and boreal regions of Europe, North America and east Asia. However, fewer than 10% of known moss species have available functional trait information. Our synthesis revealed that there is a need for research on traits related to ontogeny, sex, and intraspecific plasticity, and on co-measurement of traits related to water-relations and bryophyte-mediated soil processes. </p>

opencc-zeroJan 2024View details →
zenodo36/100

Precision and bias in dynamic light scattering optical coherence tomography measurements of diffusion and flow

<p>This repository contains raw data and analysis routines of the publication <strong>&ldquo;<em>Precision and bias in dynamic light scattering optical coherence tomography measurements of diffusion and flow</em>&rdquo;</strong> in Biomedical Optics Express (doi.org/10.1364/BOE.505847<em>).&nbsp;</em>The reader is free to use the scripts and data in this depository if the manuscript is correctly cited in their work. For further questions, feel free to contact the corresponding author. Python 3.7 was used for programming. Kindly note that simulating autocorrelation functions from extensive time series data, especially with a high repetition rate, can be time-consuming, often requiring more than 5-10 minutes. Despite parallelized processing routines for the measurement data, the full analysis may still take up to an hour. Please restart the kernel and run the code again if the parallelization fails.</p> <p>For the diffusion measurement under static conditions, there is only one file. However, for experiments involving both flowing and diffusing particles, the dataset comprises diffusion calibration, focus (beam shape) calibration, and flow measurement files. Due to the upload size limitations of the Zenodo repository, only the flow measurements corresponding to one discharge rate have been uploaded. Furthermore, only the non-dilute flow dataset has been uploaded for the same reason. However, for the dilute flow, the analysis logic remains the same, but users will need to utilize the complete g2 formula outlined in Section 2.2 of our article. All file names are sufficiently descriptive, showing whether it is diffusion, focus (waist) calibration or flow measurement. To conduct the analysis, it's essential to have information regarding the time series length (number of A-scans), the number of repeats (B-scans), and the acquisition rate.</p> <p>The results are plotted at the end of our analysis routines. The parameters are displayed as a function of depth. Users can readily compute the Signal-to-Noise Ratio (SNR) at each depth by utilizing the fitted autocorrelation amplitudes. Occasionally, the fitted amplitudes may surpass unity. In such instances, users can assume an extremely high (even infinite) SNR.</p> <div> <table> <tbody> <tr> <td> <p><strong>Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Parameters</strong></p> </td> </tr> <tr> <td> <p>Diffusion_03032023.oct</p> </td> <td> <p>Diffusion measurement file.</p> </td> <td> <p>Na=4096, Nb=1100, 5.5 kHz</p> </td> </tr> <tr> <td> <p>Diffusion_07032023.oct</p> </td> <td> <p>Diffusion calibration file for flow measurement.</p> </td> <td> <p>Na=4096,&nbsp;Nb=10, 36 kHz</p> </td> </tr> <tr> <td> <p>Waist_07032023.oct</p> </td> <td> <p>Beam waist calibration file for flow measurement.</p> </td> <td> <p>Na=4096,&nbsp;Nb=40, 36 kHz</p> </td> </tr> <tr> <td> <p>Q=2_07032023.oct</p> </td> <td> <p>Flow measurement file for a discharge rate of 2 ml/min.</p> </td> <td> <p>Na=4096, Nb=1000, 36 kHz</p> </td> </tr> <tr> <td> <p>Chirp.data</p> </td> <td> <p>File containing k-interpolation data.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>ReadOCTFile.py</p> </td> <td> <p>Written by Jos de Wit, this module reads and imports spectra from raw OCT files.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>Data_processing.py</p> </td> <td> <p>This module contains all analysis, simulation and processing routines.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>Simulation_diffusion.py</p> </td> <td> <p>This script is for simulating and fitting g1 and g2 from diffusive particles.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>Simulation_flow.py</p> </td> <td> <p>This script is for simulating and fitting g1 and g2 from flowing and diffusive particles.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>Diffusion_parallel.py</p> </td> <td> <p>This script is for analyzing static diffusion measurements performed using Thorlabs Ganymede OCT system.</p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>Flow_parallel.py</p> </td> <td> <p>This script is for analyzing flow measurements performed using Thorlabs Ganymede OCT system.</p> </td> <td> <p>&nbsp;</p> </td> </tr> </tbody> </table> </div> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
dryad36/100

Data from: Sea-surface temperature pattern effects have slowed global warming and biased warming-based constraints on climate sensitivity

<p>The observed rate of global warming since the 1970s has been proposed as a strong constraint on equilibrium climate sensitivity (ECS) and transient climate response (TCR) – key metrics of the global climate response to greenhouse-gas forcing. Using CMIP5/6 models, we show that the inter-model relationship between warming and these climate sensitivity metrics (the basis for the constraint) arises from a similarity in transient and equilibrium warming patterns within the models, producing an effective climate sensitivity (EffCS) governing recent warming that is comparable to the value of ECS governing long-term warming under CO<sub>2</sub> forcing. However, CMIP5/6 historical simulations do not reproduce observed warming patterns. When driven by observed patterns, even high ECS models produce low EffCS values consistent with the observed global warming rate. The inability of CMIP5/6 models to reproduce observed warming patterns thus results in a bias in the modeled relationship between recent global warming and climate sensitivity. Correcting for this bias means that observed warming is consistent with wide ranges of ECS and TCR extending to higher values than previously recognized. These findings are corroborated by energy balance model simulations and coupled model (CESM1-CAM5) simulations that better replicate observed patterns via tropospheric wind nudging or Antarctic meltwater fluxes. Because CMIP5/6 models fail to simulate observed warming patterns, proposed warming-based constraints on ECS, TCR, and projected global warming are biased low. The results reinforce recent findings that the unique pattern of observed warming has slowed global-mean warming over recent decades, and that how the pattern will evolve in the future represents a major source of uncertainty in climate projections.</p>

opencc-zeroFeb 2024View details →
dryad36/100

Biogeography of the world's worst invasive species has spatially-biased knowledge gaps but is predictable

<p>The world's "100 worst invasive species" were listed in 2000. The list is taxonomically diverse and often cited (typically for single-species studies), and its species are frequently reported in global biodiversity databases. We acted on the principle that these notorious species should be well-reported to help answer two questions about global biogeography of invasive species (i.e., not just their invaded ranges): (1) "how are data distributed globally?" and (2) "what predicts diversity?" We collected location data for each of the 100 species from multiple databases; 95 had sufficient data for analyses. For question (1), we mapped global species richness and cumulative occurrences since 2000 in (0.5 degree)<sup>2</sup> grids. For question (2) we compared alternative regression models representing non-exclusive hypotheses for geography (i.e., spatial autocorrelation), sampling effort, climate, and anthropocentric effects.</p> <p>Reported locations of the invasive species were spatially-biased, leaving large gaps on multiple continents. Accordingly, species richness was best explained by both anthropocentric effects not often used in biogeographic models (Government Effectiveness, Voice &amp; Accountability, human population size) and typical natural factors (climate, geography; R<sup>2</sup> = 0.87). Cumulative occurrence was strongly related to anthropocentric effects (R<sup>2</sup> = 0.62). We extract five lessons for invasive species biogeography; foremost is the importance of anthropocentric measures for understanding invasive species diversity patterns and large lacunae in their known global distributions. Despite those knowledge gaps, advanced models here predict well the biogeography of the world's worst invasive species for much of the world.</p>

opencc-zeroFeb 2024View details →
zenodo36/100

Data from: Casimir repulsion with biased semiconductors

<p>This repository provides data underlying the figures in the manuscript titled 'Casimir repulsion with biased semiconductors'.</p> <p>The data is given in the CSV-file format and it is organized into subfolders with the respective figure name. For more information see the README files in those subfolders.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Weather station temperature exposure metadata and exposure bias estimates

<p>Thermometer exposure metadata and monthly estimates of the exposure biases arising from the transition to Stevenson screens at mid-latitude weather stations in an extended version of the CRUTEM5 station database (which starts in 1781; Osborn et al., 2021).</p> <p>This dataset accompanies Wallis et al. (2024). Please see the attached readme file and Wallis et al. (2024) for further information about the dataset format and its creation.</p> <p>---</p> <p><strong>References</strong></p> <p>Osborn, T.J., Jones, P.D., Lister, D.H., Morice, C.P., Simpson, I.R., Winn, J., Hogan, E. &amp; Harris, I.C. (2021) Land surface air temperature variations across the globe updated to 2019: the CRUTEM5 dataset. <em>Journal of Geophysical Research: Atmospheres</em>,&nbsp;<strong>126</strong>, e2019JD032352, https://doi.org/10.1029/2019JD032352</p> <p>Wallis, E.J.,&nbsp;Osborn, T.J., Taylor, M., Jones, P.D., Joshi, M. &amp; Hawkins, E. (2024) Quantifying exposure biases in early instrumental land surface air temperature observations. <em>International Journal of Climatology, </em>https://doi.org/10.1002/joc.8401</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Replication Data for the Paper "Is there a secular decline in disruptive patents? Correcting for measurement bias"

<p><strong>Working Paper Title: The Illusive Slump of Disruptive Patents</strong></p> <p>The repository contains replication data and scripts for the paper:&nbsp;</p> <p>Jeffrey T. Macher, Christian Rutzer, Rolf Weder,<br>Is there a secular decline in disruptive patents? Correcting for measurement bias,<br>Research Policy,<br>Volume 53, Issue 5,<br>2024,<br>104992,<br>ISSN 0048-7333,<br>https://doi.org/10.1016/j.respol.2024.104992</p> <p>The core component of the repository is the file `replication_file.R` which is an R script to replicate all figures and tables of the paper.</p> <p>The script relies on multiple data files. The datasets contain CD-values for granted USPTO utility patents for the years 1976-2016. All data is available in two formats, `.fst' (from the R package fst) and `.csv'.</p> <p>&nbsp;</p> <p><strong>Dataset Descriptions</strong></p> <p><strong>Main Datasets:</strong></p> <p>1. `dat_cd5_no_trunc`: This dataset contains the CD5 index of patents based on a method without truncation.&nbsp;</p> <p>2. `dat_cd5_trunc_1975`: This dataset contains the CD5 index of patents based on a truncation method as in Park et al. (Nature, 2023).</p> <p>3. `dat_cd5_no_trunc_app_adj`: This dataset contains the CD5 index of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>&nbsp;</p> <p><strong>Additional Datasets:</strong></p> <p>4. `dat_cd5_trunc_1985`: This dataset contains the CD5 index of patents based on a truncation of all backward citations to patents published before 1985.</p> <p>5. `dat_cd5_trunc_1995`: This dataset contains the CD5 index of patents based on a truncation of all backward citations to patents published before 1995.</p> <p>6. `dat_cd5_no_trunc_app`: This dataset contains the CD5 index of patents based on a method without truncation and including citations to patent applications granted by 2021, as well as those not yet granted.&nbsp;</p> <p>7. `age_bwc_untrunc`: This dataset contains the age of backward citations using untruncated data.</p> <p>8. age_bwc_trunc: This dataset contains the age of backward citations using truncated data as in Park et al. (Nature, 2023).</p> <p>9. `age_bwc_untrunc_app_adj`: This dataset contains the age of backward citations using untruncated data and considering citations of patent applications granted until 2021.</p> <p>10. `dat_cd10_trunc_1975`: This dataset contains the CD10 index of patents based on a truncation method as in Park et al. (Nature, 2023).&nbsp;</p> <p>11. `dat_cd10_no_trunc_app_adj`: This dataset contains the CD10 index of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>12. `dat_cd2021_trunc_1975`: This dataset contains the CD index as of 2021 of patents based on a truncation method as in Park et al. (Nature, 2023).&nbsp;</p> <p>13. `dat_cd2021_no_trunc_app_adj`: This dataset contains the CD index as of 2021 of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>&nbsp;</p> <p>To successfully run the `replication_file.R' script, make sure all data files are in the directory and the `mainDir1' variable at the beginning of the script is set to the correct path to where the data is stored. In addition, set the `mainDir2' variable to the folder where you want to store the figures created by the script.&nbsp;</p> <p>If you have any questions, please contact christian.rutzer@unibas.ch</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Optical Bias and Cryogenic Laser Readout of a Multipixel Superconducting Nanowire Single Photon Detector

<p>The complete datasets of the Publication are updloaded in this repository.</p> <p>Have a look into the ReadMe files for further information.</p> <p>Info about the Spice Model: ReadME LTSPICE</p> <p>Info about the Detector Efficiency Analysis: ReadMe Countrate</p> <p>Info about the PNR Extraction: ReadMe PNR Extraction</p> <p>Info about the SNSPD traces: ReadMe Traces&nbsp;</p> <p>Dataset of the Laser Diode Characterisation: IVP_Laserdiode.txt</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

Data from: Limited evidence of biased offspring sex allocation in a cavity-nesting conspecific brood parasite

<p>Sex allocation theory predicts that mothers should bias investment in offspring toward the sex that yields higher fitness returns; one such bias may be a skewed offspring-sex ratio. Sex allocation is well-studied in birds with cooperative breeding systems, with theory on local resource enhancement and production of helpers at the nest, but little theoretical or empirical work has focused on birds with brood parasitic breeding systems. Wood ducks (<em>Aix sponsa</em>) are conspecific brood parasites, and rates of parasitism appear to increase with density. Because female wood ducks show high natal philopatry and nest sites are often limiting, local resource competition (LRC) theory predicts that females should overproduce male offspring—the dispersing sex—when competition (density) is high. However, the unique features of conspecific brood parasitism generate alternative predictions from other sex allocation theories, which we develop and test here. We experimentally manipulated the nesting density of female wood ducks in four populations from 2013-2016 and analyzed the resulting sex allocation of &gt;2000 ducklings.  In contrast to predictions we did not find overproduction of male offspring by females in high-density populations, females in better condition, or parasitic females; modest support for LRC was found in overproduction of only female parasitic offspring with higher nest box availability. The lack of evidence for sex ratio biases, as expected for LRC and some aspects of brood parasitism, could reflect conflicting selection pressures from nest competition and brood parasitism, or that mechanisms of adaptive sex ratio bias are not possible.</p>

opencc-zeroMar 2024View details →
zenodo36/100

Jupyter Notebook and comprising data for GRL2023GL106264R: Understanding the Cascade: Removing GCM biases improves dynamically downscaled climate projections

<p>This notebook and attendant files allows users to interface with a small subset of the data used to create the data in GRL2023GL106264R. Also feel free to check out the overall description of the non-bias corrected dynamically downscaled GCMs in WUS-D3 here: https://zenodo.org/records/10635867. This DOI also contains version of WRF 4.1.3 allowing for yearly CH4, CO2, and N2O updates, as well as a 360-day calendar version.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

AskDoc Chemical Entity NER Dataset for Gender Bias Analysis

<p>This repository contains the dataset for the NAACL 2024 paper titled, "A Comprehensive Study of Gender Bias in Chemical Named Entity Recognition Models."</p> <p>The dataset contains two files:</p> <p>askdoc_female.conll - Female AskDoc dataset<br>askdoc_male.conll - Male AskDoc dataset</p> <p>These files contain annotated AskDoc data with chemical mentions from r/AskDocs on Reddit. The female file contains examples from people that self-identify as female and the male file contains data for people that self-identify as male.</p> <p>A comprehensive datasheet is also provided.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Bias-corrected EURO-CORDEX daily temperature and precipitation dataset for Hungary

<p>These datasets contain bias-corrected regional climate model outputs for daily precipitation, near surface mean-, minimum- and maximum temperature for the historical period 1993-2005 on a regular 0.11&deg;x0.11&deg; lon/lat grid (between latitudes 45.6825&deg;N and 48.6525&deg;N, and longitudes 15.9075&deg;E and 22.9475&deg;E).</p> <p>&nbsp;</p> <p>The datasets contain daily outputs of the following high-resolution (0.11&deg;) regional climate models from the framework of EURO-CORDEX (Jacob et al., 2014):</p> <p>-CCLM</p> <p>-HIRHAM</p> <p>-RACMO</p> <p>-RCA</p> <p>-REMO</p> <p>&nbsp;</p> <p>The reference dataset is HUCLIM, which covers Hungary for the period 1971-2022 (downloaded in 2023) produced by the HungaroMet Hungarian Meteorological Service.</p> <p>File format: NetCDF</p> <p>All bias-corrected data produced by the use of HUCLIM have been created following the work of Mezghani et al. (2017).</p> <p>&nbsp;</p> <p>References:</p> <p>Jacob, D., Petersen, J., Eggert, B., Alias, A., Christensen, O.B., Bouwer, L.M., Braun, A., Colette, A., D&eacute;qu&eacute;, M., Georgievski, G., Georgopoulou, E., Gobiet, A., Menut, L., Nikulin, G., Haensler, A., Hempelmann, N., Jones, C., Keuler, K., Kovats, S., Kr&ouml;ner, N., Kotlarski, S., Kriegsmann, A., Martin, E., van Meijgaard, E., Moseley, C., Pfeifer, S., Preuschmann, S., Radermacher, C., Radtke, K., Rechid, D., Rounsevel, M., Samuelsson, P., Somot, S., Soussana, J.-F., Teichmann, C., Valentini, R., Vautard, R., Weber, B. and Yiou, P. (2014) EURO-CORDEX New high resolution climate change projections for European impact research. Reg. Environ. Change, 14, 563&ndash;578.&nbsp;<a href="https://doi.org/10.1007/s10113-013-0499-2" target="_blank" rel="noopener">https://doi.org/10.1007/s10113-013-0499-2</a></p> <p>Mezghani, A., Dobler, A., Haugen, J.E., Benestad, R.E., Parding, K.M., Piniewski, M., Kardel, I. and Kundzewicz, Z.W. (2017) CHASE-PL Climate Projection dataset over Poland &ndash; bias adjustment of EURO-CORDEX simulations. Earth Syst. Sci. Data, 9, 905&ndash;925.&nbsp;<a href="https://doi.org/10.5194/essd-9-905-2017" target="_blank" rel="noopener">https://doi.org/10.5194/essd-9-905-2017</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record