Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

317

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

317 results for “Experts”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Assessing cumulative impacts of forest development on the distribution of furbearers using expert-based habitat modeling

Cumulative impacts of anthropogenic landscape change must be considered when managing and conserving wildlife habitat. Across the central-interior of British Columbia, Canada, industrial activities are altering the habitat of furbearer species. This region has witnessed unprecedented levels of anthropogenic landscape change following rapid development in a number of resource sectors, particularly forestry. Our objective was to create expert-based habitat models for three furbearer species: fisher (Pekania pennanti), Canada lynx (Lynx canadensis), and American marten (Martes americana) and quantify habitat change for those species. We recruited 10 biologist and 10 trapper experts and then used the analytical hierarchy process to elicit expert knowledge of habitat variables important to each species. We applied the models to reference landscapes (i.e., registered traplines) in two distinct study areas and then quantified the change in habitat availability from 1990 to 2013. There was strong agreement between expert groups in the choice of habitat variables and associated scores. Where anthropogenic impacts had increased considerably over the study period, the habitat models showed substantial declines in habitat availability for each focal species (78% decline in optimal fisher habitat, 83% decline in optimal lynx habitat, and 79% decline in optimal marten habitat). For those traplines with relatively little forest harvesting, the habitat models showed no substantial change in the availability of habitat over time. The results suggest that habitat for these three furbearer species declined significantly as a result of the cumulative impacts of forest harvesting. Results of this study illustrate the utility of expert knowledge for understanding large-scale patterns of habitat change over long time periods.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Crowds replicate performance of scientific experts scoring phylogenetic matrices of phenotypes

Scientists building the Tree of Life face an overwhelming challenge to categorize phenotypes (e.g., anatomy, physiology) from millions of living and fossil species. This biodiversity challenge far outstrips the capacities of trained scientific experts. Here we explore whether crowdsourcing can be used to collect matrix data on a large scale with the participation of the non-expert students, or "citizen scientists." Crowdsourcing, or data collection by non-experts, frequently via the internet, has enabled scientists to tackle some large-scale data collection challenges too massive for individuals or scientific teams alone. The quality of work by non-expert crowds is, however, often questioned and little data has been collected on how such crowds perform on complex tasks such as phylogenetic character coding. We studied a crowd of over 600 non-experts, and found that they could use images to identify anatomical similarity (hypotheses of homology) with an average accuracy of 82% compared to scores provided by experts in the field. This performance pattern held across the Tree of Life, from protists to vertebrates. We introduce a procedure that predicts the difficulty of each character and that can be used to assign harder characters to experts and easier characters to a non-expert crowd for scoring. We test this procedure in a controlled experiment comparing crowd scores to those of experts and show that crowds can produce matrices with over 90% of cells scored correctly while reducing the number of cells to be scored by experts by 50%. Preparation time, including image collection and processing, for a crowdsourcing experiment is significant, and does not currently save time of scientific experts overall. However, if innovations in automation or robotics can reduce such effort, then large-scale implementation of our method could greatly increase the collective scientific knowledge of species phenotypes for phylogenetic tree building. For the field of crowdsourcing, we provide a rare study with ground truth, or an experimental control that many studies lack, and contribute new methods on how to coordinate the work of experts and non-experts. We show that there are important instances in which crowd consensus is not a good proxy for correctness.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Development and validation of an expert-based habitat suitability model to support boreal caribou conservation

The long-term persistence of boreal caribou (Rangifer tarandus caribou) is threatened by the negative impacts of human activities, including industrial development. The vast geographical distribution and behavioural variability of caribou warranted the development and validation of an appropriate conservation tool for managers in eastern Canada. We developed a habitat suitability model for boreal caribou in Québec by integrating expert knowledge into a hierarchical analysis. We elicited responses from 14 experts on caribou ecology to determine the best spatiotemporal scales of application, the relative importance of habitat variables, the zone of influence of human infrastructure, and the parameterization of the model. Based on this input, we built the model using 8 habitat categories and 3 human infrastructure variables. Experts identified mature conifer-dominated forests and open lichen woodlands as the most important vegetation categories for boreal caribou, whereas density of and proximity to paved roads, forest roads, and mines decreased habitat quality. We mapped the resulting model over the entire province of Québec (up to the northern forest allocation limit), and validated it using independent GPS telemetry datasets acquired in 3 distinct regions. Our model predicted that habitat suitability was highest in the northeastern part of our study area, where timber harvest activities and roads were virtually absent. Conversely, southern parts of Québec were generally unsuitable for boreal caribou. Our habitat suitability model is among the first tools for boreal caribou conservation available to wildlife and land managers in Québec.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Adult experts' perceptions of telemental health for youth: a Delphi study

Objectives: Our objectives were to measure experts' opinions and develop consensus via the Delphi process on the barriers, applications, and concerns associated with telemental health (TMH) for youth. Materials and methods: We delivered 3 online surveys over 2 months in Summer, 2016–2025 adult experts, including adults who experienced youth depression or suicidality, parents of youth with lived experience, and professionals (ie youth mental health researchers, clinicians/staff, or educators). We used the Delphi method to construct Likert and open-ended questions, developing expert consensus over 3 iterative surveys on the barriers and benefits of TMH for youth. Results: Adult experts identified stigma and knowledge barriers to youth mental health care. Although TMH is perceived as beneficial for screening, education, follow-up, and emotional support, no single delivery method (eg websites or instant messaging) was deemed universally beneficial. Discussion: Adults are the developers, administrators, and gatekeepers of youth mental health care. Although adult experts see potential for TMH to supplement traditional therapy via familiar technologies, there is no consensus on the technologies by which TMH should be delivered. However, there is consensus that family members and friends provide potential pathways to care; thus, an online TMH toolkit for youth would be beneficial for both caretakers and practitioners. Conclusion: Telemental health may not overcome barriers for crisis management but adult experts agreed that TMH had potential benefits for youth. Health care organizations should conduct research and provide training and education to youth caretakers and practitioners on potential barriers and benefits of TMH technologies for youth.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Use of opportunistic sightings and expert knowledge to predict and compare Whooping Crane stopover habitat

Predicting a species' distribution can be helpful for evaluating management actions such as critical habitat designations under the U.S. Endangered Species Act or habitat acquisition and rehabilitation. Whooping Cranes (Grus americana) are one of the rarest birds in the world, and conservation and management of habitat is required to ensure their survival. We developed a species distribution model (SDM) that could be used to inform habitat management actions for Whooping Cranes within the state of Nebraska (U.S.A.). We collated 407 opportunistic Whooping Crane group records reported from 1988 to 2012. Most records of Whooping Cranes were contributed by the public; therefore, developing an SDM that accounted for sampling bias was essential because observations at some migration stopover locations may be under represented. An auxiliary data set, required to explore the influence of sampling bias, was derived with expert elicitation. Using our SDM, we compared an intensively managed area in the Central Platte River Valley with the Niobrara National Scenic River in northern Nebraska. Our results suggest, during the peak of migration, Whooping Crane abundance was 262.2 (90% CI 40.2−3144.2) times higher per unit area in the Central Platte River Valley relative to the Niobrara National Scenic River. Although we compared only 2 areas, our model could be used to evaluate any region within the state of Nebraska. Furthermore, our expert-informed modeling approach could be applied to opportunistic presence-only data when sampling bias is a concern and expert knowledge is available.

opencc-zeroDec 2014View details →
zenodo32/100

Photonics4All - Expert Interview

<p>General video about Photonics and the Photonics4All Project with different expert interviews.</p>

opencc-by-nc-nd-4.0Jun 2016View details →
zenodo32/100

Expert and AI-generated annotations of the tissue types for the RMS-Mutation-Prediction microscopy images

<div> <p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute <a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.. You can use the manifests included in this Zenodo record to download the content of the collection following the&nbsp;<strong>Download instructions</strong>&nbsp;below.</p> <h3>Collection description</h3> </div> <div> <div> <p>This dataset contains 2 components:</p> <ol> <li>Annotations of multiple&nbsp; regions of interest performed by an expert pathologist with eight years of experience for a subset of hematoxylin and eosin (H&amp;E) stained images from the RMS-Mutation-Prediction image collection [1,2]. Annotations were generated manually, using the Aperio ImageScope tool, to delineate regions of alveolar rhabdomyosarcoma (ARMS), embryonal rhabdomyosarcoma (ERMS), stroma, and necrosis [3]. The resulting planar contour annotations were originally stored in ImageScope-specific XML format, and subsequently converted into Digital Imaging and Communications in Medicine (DICOM) Structured Report (SR) representation using the open source conversion tool [4].</li> <li>AI-generated annotations stored as probabilistic segmentations.</li> </ol> <p><strong>WARNING</strong>: After the release of IDC v20 (v2 of this data record), it was discovered that a mistake had been made during data conversion that affected the newly-released segmentations accompanying the "RMS-Mutation-Prediction" collection. Segmentations released in v20 for this collection have the segment labels for alveolar rhabdomyosarcoma (ARMS) and embryonal rhabdomyosarcoma (ERMS) switched in the metadata relative to the correct labels. Thus segment 3 in the released files is labelled in the metadata (the SegmentSequence) as ARMS but should correctly be interpreted as ERMS, and conversely segment 4 in the released files is labelled as ERMS but should be correctly interpreted as ARMS. This mistake was fixed in the version v3 of this record (IDC data release v21).</p> <p>Many pixels from the whole slide images annotated by this dataset are not contained inside any annotation contours and are considered to belong to the background class. Other pixels are contained inside only one annotation contour and are assigned to a single class.&nbsp; However,&nbsp; cases also exist in this dataset where annotation contours overlap.&nbsp; In these cases, the pixels contained in multiple contours could be assigned membership in multiple classes.&nbsp; One example is a necrotic tissue contour overlapping an internal subregion of an area designated by a larger ARMS or ERMS annotation.&nbsp; The ordering of annotations in this DICOM dataset preserves the order in the original XML generated using ImageScope.&nbsp; These annotations were converted, in sequence, into segmentation masks and used in the training of several machine learning models. Details on the training methods and model results&nbsp; are presented in [1].&nbsp; In the case of overlapping contours, the order in which annotations are processed may affect the generated segmentation mask if prior contours are overwritten by later contours in the sequence.&nbsp; It is up to the application consuming this data to decide how to interpret tissues regions annotated with multiple classes. The annotations included in this dataset are available for visualization and exploration from the National Cancer Institute Imaging Data Commons (IDC) [5] (also see IDC Portal at <a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov</a>) as of data release v18.&nbsp;Direct link to open the collection in IDC Portal: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.</p> </div> <div> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example,&nbsp;<code>pan_cancer_nuclei_seg_dicom-collection_id-idc_v19-aws.s5cmd</code> corresponds to the annotations for th eimages in the <code>collection_id</code> collection introduced in IDC data release v19. DICOM Binary segmentations were introduced in IDC v20. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <p>For each of the collections, the following manifest files are provided:</p> <ol> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-dcf.dcf</code>: Gen3 manifest (for details see&nbsp;<a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in&nbsp;<code>-aws.s5cmd</code>&nbsp;reference files stored in Amazon Web Services (AWS) buckets, while&nbsp;<code>-gcs.s5cmd</code>&nbsp;reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using&nbsp;<code>.s5cmd</code>&nbsp;manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the&nbsp;<code>.s5cmd</code>&nbsp;manifest file:&nbsp;<code>idc download&nbsp;manifest.s5cmd</code></li> </ol> <p>To download the files using&nbsp;<code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <p>If you use the files referenced in the attached manifests, we ask you to cite this dataset, as well as the publication describing the original dataset&nbsp;<a href="https://paperpile.com/c/NHiBXI/njdR">[2]</a>&nbsp;and publication acknowledging IDC&nbsp;<a href="https://paperpile.com/c/NHiBXI/uJJZ">[5]</a>.</p> <h3>References</h3> </div> </div> <div> <p>[1] D. Milewski et al., "Predicting molecular subtype and survival of rhabdomyosarcoma patients using deep learning of H&amp;E images: A report from the Children's Oncology Group," Clin. Cancer Res., vol. 29, no. 2, pp. 364&ndash;378, Jan. 2023, doi: 10.1158/1078-0432.CCR-22-1663.</p> <p>[2] Clunie, D., Khan, J., Milewski, D., Jung, H., Bowen, J., Lisle, C., Brown, T., Liu, Y., Collins, J., Linardic, C. M., Hawkins, D. S., Venkatramani, R., Clifford, W., Pot, D., Wagner, U., Farahani, K., Kim, E., &amp; Fedorov, A. (2023). DICOM converted whole slide hematoxylin and eosin images of rhabdomyosarcoma from Children's Oncology Group trials [Data set]. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.8225132" rel="noopener">https://doi.org/10.5281/zenodo.8225132</a></p> <p>[3] Agaram NP. Evolving classification of rhabdomyosarcoma. Histopathology. 2022 Jan;80(1):98-108. doi: 10.1111/his.14449. PMID: 34958505; PMCID: PMC9425116,https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9425116/</p> <p>[4] Chris Bridge. (2024). ImagingDataCommons/idc-sm-annotations-conversion: v1.0.0 (v1.0.0). Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.10632182" rel="noopener">https://doi.org/10.5281/zenodo.10632182</a></p> <p>[5] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. &amp; Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence. Radiographics 43, (2023).</p> </div>

opencc-by-4.0Nov 2024View details →
zenodo32/100

PRETEST AND POSTEST OF EXPERT SYSTEM FOR VOCATIONAL GUIDANCE

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

AGIMUS Interviews with experts in AI ethics

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo32/100

Underlying dataset of Experts and Machines against Bullies: A Hybrid Approach to Detect Cyberbullies

<p>YouTube data collection for cyberbullying studies.&nbsp;</p><p>Citation:</p><p>M. Dadvar, R.B. Trieschnigg and F.M.G. de Jong, &nbsp;Experts and Machines Against Bullies: A Hybrid Approach to Detect Cyberbullies. In 27th Canadian Conference on Artificial Intelligence, &nbsp;University of Waterloo, Montréal, Canada, 2014</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Experts fail to reliably detect AI-generated histological data

<p>This repository contains material related to the paper "<em>Experts fail to reliably detect AI-generated histological data</em>":</p> <ul> <li>Dreambooth parameters (<em>dreambooth_parameters.zip</em>)</li> <li>Images displayed during the study <em>(images.zip)</em></li> <li>Data collected during the survey (<em>results-survey.xlsx</em>)</li> <li>R Code to reproduce results and figures (<em>analysis_code.zip</em>)</li> </ul> <p>Please find our associated work here:</p> <p>Hartung, J., Reuter, S., Kulow, V.A., F&auml;hling, M., Spreckelsen, C., and Mrowka, R. (2024). Experts fail to reliably detect AI-generated histological data. Sci Rep&nbsp;<em>14</em>, 28677. https://doi.org/10.1038/s41598-024-73913-8.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

A global expert elicitation on present-day human-fire interactions: data & analysis code

<p>These files support "A global expert elicitation on present-day human-fire interactions", accepted for publication in the Philosophical Transactions of the Royal Society B. doi:&nbsp;<span>10.1098/rstb.2023-0463</span></p> <p>The files contain data from a survey of experts on human-fire interactions as well as code used to process and analyse it.</p> <p>An overview of the files is given below.&nbsp;</p> <h2><strong>1) Data</strong></h2> <p>Raw survey data are provided by geographic region ("GFUS_region"). Data processed to produce analyses in the associated paper are provided as "GFUS_Processed".</p> <p>Columns in data files are named by the question of the survey that generated them (Q1b through Q90). The contents of the associated questions that were asked are given in the data_dictionary.xlsx file.</p> <p>- GFUS_spatial provides the data merged with a shapefile of the survey regions.&nbsp;</p> <p>- "Gov_compare" and "LIFE_SH_fire" provide files to compare survey data with the DAFI and LIFE literature meta-analyses (see Supplementary 3 to the main text).</p> <p>- Question meta-data and question-summary provides an overview of the questions, as well as a topline overview of the numbers of responses and internal coherence (entropy) of survey responses.&nbsp;</p> <h2>2) Code</h2> <p>4 scripts are provided.&nbsp;</p> <p>- Firstly the script that was used to summarise survey responses by geographic region (Summarise_by_region)<br>- Secondly the code used to conduct statistical tests presented in the paper<br>- Thirdly the code used to produce plots in the paper<br>- Fourthly the code used to produce topline descriptive statistics presented in the paper</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Expert annotations for the Catalan Common Voice (v13)

<h2>Dataset Description</h2> <p>- Homepage:&nbsp;<a href="https://projecteaina.cat/tech/">https://projecteaina.cat/tech/</a>]<br>- Point of Contact: langech@bsc.es</p> <h3>Dataset Summary</h3> <p>These are the annotations made by a team of experts on the speakers with more than 1200 seconds recorded in the Catalan set of the <a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_13_0/tree/main/transcript/ca">Common Voice dataset (v13)</a>.</p> <p>The annotators were initially tasked with evaluating all recordings associated with the same individual. Following that, they were instructed to annotate the speaker's accent, gender, and the overall quality of the recordings.</p> <p>The accents and genders taken into account are the ones used until version 8 of the Common Voice corpus.</p> <p>See annotations for more details.</p> <h3>Supported Tasks and Leaderboards</h3> <p>Gender classification, Accent classification.</p> <h3>Languages</h3> <p>The dataset is in Catalan (ca).</p> <h2>Dataset Structure</h2> <h3>Instances</h3> <p>Two xlsx documents are published, one for each round of annotations.</p> <p>The following information is available in each of the documents:</p> <p><br><code>{</code><br><code>&nbsp; 'speaker ID': '1b7fc0c4e437188bdf1b03ed21d45b780b525fd0dc3900b9759d0755e34bc25e31d64e69c5bd547ed0eda67d104fc0d658b8ec78277810830167c53ef8ced24b',&nbsp;</code><br><code>&nbsp; 'idx': '31',&nbsp;</code><br><code>&nbsp; 'same speaker': {'AN1': 'SI',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN2': 'SI',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN3': 'SI',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'agreed': 'SI',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'percentage': '100'},&nbsp;</code><br><code>&nbsp; 'gender': {'AN1': 'H',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN2': 'H',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN3': 'H',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'agreed': 'H',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'percentage': '100'},&nbsp;</code><br><code>&nbsp; 'accent': {'AN1': 'Central',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN2': 'Central',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN3': 'Central',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'agreed': 'Central',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'percentage': '100'},&nbsp;</code><br><code>&nbsp; 'audio quality': {'AN1': '4.0',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN2': '3.0',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN3': '3.0',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'agreed': '3.0',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'percentage': '66',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'mean quality': '3.33',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'stdev quality': '0.58'},&nbsp;</code><br><code>&nbsp; 'comments': {'AN1': '',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN2': 'pujades i baixades de volum',</code><br><code>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 'AN3': 'Deu ser d'alguna zona de transici&oacute; amb el central, perqu&egrave; no fa una reducci&oacute; total voc&agrave;lica, per&ograve; hi t&eacute; molta tend&egrave;ncia'}, &nbsp;&nbsp;</code><br><code>}</code></p> <p>&nbsp;</p> <p>We also publish the document Guia anotaci&oacute; parlants.pdf, with the guidelines the annotators recieved.</p> <h3>Data Fields</h3> <ul> <li>speaker ID (string): An id for which client (voice) made the recording in the Common Voice corpus</li> <li>&nbsp;idx (int): Id in this corpus</li> <li>&nbsp;AN1 (string): Annotations from Annotator 1</li> <li>&nbsp;AN2 (string): Annotations from Annotator 2</li> <li>&nbsp;AN3 (string): Annotations from Annotator 3</li> <li>&nbsp;agreed (string): Annotation from the majority of the annotators</li> <li>percentage (int): Percentage of annotators that agree with the agreed annotation</li> <li>mean quality (float): Mean of the quality annotation</li> <li>stdev quality (float): Standard deviation of the mean quality</li> </ul> <h3>Data Splits</h3> <p>The corpus remains undivided into splits, as its purpose does not involve training models.</p> <h2>Dataset Creation</h2> <h3>Curation Rationale</h3> <p>During 2022, a campaign was launched to promote the Common Voice corpus within the Catalan-speaking community, achieving remarkable success. However, not all participants provided their demographic details such as age, gender, and accent. Additionally, some individuals faced difficulty in self-defining their accent using the standard classifications established by specialists.</p> <p>In order to obtain a balanced corpus with reliable information, we have seen the the necessity of enlisting a group of experts from the University of Barcelona to provide accurate annotations.</p> <p>We release the complete annotations because transparency is fundamental to our project. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <h3>Source Data</h3> <p>The original data comes from the [Catalan sentences of the Common Voice corpus](https://commonvoice.mozilla.org/en/datasets).</p> <p><strong>Initial Data Collection and Normalization</strong></p> <p>We have selected speakers who have recorded more than 1200 seconds of speech in the Catalan set of the <a href="https://commonvoice.mozilla.org/en/datasets">version 13 of the Common Voice corpus</a>.</p> <p><strong>Who are the source language producers?</strong></p> <p>The original data comes from the&nbsp;<a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main/transcript/ca">Catalan sentences of the Common Voice corpus</a>.</p> <h3>Annotations</h3> <p><strong>Annotation process</strong></p> <p>Starting with <a href="https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main/transcript/ca">version 13 of the Common Voice corpus</a> we identified the speakers (273) who have recorded more than 1200 seconds of speech.&nbsp;</p> <p>A team of three annotators was tasked with annotating:</p> <ul> <li>if all the recordings correspond to the same person</li> <li>the gender of the speaker</li> <li>the accent of the speaker</li> <li>the quality of the recording</li> </ul> <p>They conducted an initial round of annotation, discussed their varying opinions, and subsequently conducted a second round.</p> <p>We release the complete annotations because transparency is fundamental to our project. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <p><strong>Who are the annotators?</strong></p> <p>The annotation was entrusted to the [CLiC (Centre de Llenguatge i Computaci&oacute;)](https://clic.ub.edu/en/que-es-clic) team from the University of Barcelona.&nbsp;<br>They selected a group of three annotators (two men and one woman), who received a scholarship to do this work.&nbsp;</p> <p>The annotation team was composed of:</p> <ul> <li>Annotator 1: 1 female annotator, aged 18-25, L1 Catalan, student in the Modern Languages and Literatures degree, with a focus on Catalan.</li> <li>Annotators 2 &amp; 3: 2 male annotators, aged 18-25, L1 Catalan, students in the Catalan Philology degree.</li> <li>1 female supervisor, aged 40-50, L1 Catalan, graduate in Physics and in Linguistics, Ph.D. in Signal Theory and Communications.</li> </ul> <p>To do the annotation they used a Google Drive spreadsheet</p> <h3>Personal and Sensitive Information</h3> <p>The Common Voice dataset consists of people who have donated their voice online. We don't share here their voices, but their gender and accent.&nbsp;<br>You agree to not attempt to determine the identity of speakers in the Common Voice dataset.</p> <h2>Considerations for Using the Data</h2> <h3>Social Impact of Dataset</h3> <p>The ID come from the Common Voice dataset, that consists of people who have donated their voice online.</p> <p><em>You agree to not attempt to determine the identity of speakers in the Common Voice dataset.</em></p> <p>The information from this corpus will allow us to train and evaluate well balanced Catalan ASR models. Furthermore, we believe they hold philological value for studying dialectal and gender variants.</p> <h3>Discussion of Biases</h3> <p>Most of the voices of the common voice in Catalan correspond to men with a central accent between 40 and 60 years old. The aim of this dataset is to provide information that allows to minimize the biases that this could cause.</p> <p>For the gender annotation, we have only considered "H" (male) and "D" (female).</p> <h3>Other Known Limitations</h3> <p>[N/A]</p> <h2>Additional Information</h2> <h3>Dataset Curators</h3> <p>Language Technologies Unit at the Barcelona Supercomputing Center (langtech@bsc.es)</p> <p>This work has been promoted and financed by the Generalitat de Catalunya through the <a href="https://projecteaina.cat/">Aina project</a>.</p> <h3>Licensing Information</h3> <p>This dataset is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a> license.</p> <p>It can be used for any purpose, whether academic or commercial, under the terms of the license.&nbsp;<br>Give appropriate credit, provide a link to the license, and indicate if changes were made.</p> <h3>Citation Information</h3> <p><a href="../badge/DOI/10.5281/zenodo.11104388.svg">DOI</a></p> <h3>Contributions</h3> <p>The annotation was entrusted to the <a href="https://stel2.ub.edu/el-servei">STeL</a> team from the University of Barcelona.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Expert Interview Description of Concepts

<p>All the identified codes from the coding process in the qualitative analysis of the biodesign expert interview transcriptions is being described or defined based on evidence from the interview transcriptions.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Transcript of FGD and Interview of participant and expert in survey and health PR

<p>Transcript in bahasa Indonesia about FGD expert panel and interview of survei participant on AI and health PR</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Domain expert readability dataset

<p>Judgments gathered from 10 experts through a web-based survey on the readability of publication abstracts. The abstracts used were a subset of the AMiner&#39;s DBLP citation nework v10 dataset (<a href="https://aminer.org/citation">https://aminer.org/citation</a>) in the discipline of data and knowledge management. In particular, abstracts containing the following keywords were used: &quot;database&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;data management&quot;, &quot;cloud computing&quot;, &quot;data mining&quot;, &quot;algorithms&quot;, &quot;classification&quot;, &quot;query processing&quot;, &quot;networks&quot;, &quot;indexing&quot;, &quot;distributed systems&quot;.</p> <p>After reading the abstract, each expert had to answer&nbsp;the following questions on a 5 point scale.</p> <ul> <li>Q1:&nbsp;Please rate how well-written the abstract is.</li> <li>Q2: Does the abstract contain linguistic errors?</li> <li>Q3: Please rate how clear the contribution of the paper is (based on the abstract).</li> </ul> <p>For each question, the interpretation of the extreme scale values (i.e., 1 and 5) were&nbsp;provided. In particular, 1 = &ldquo;very poorly written&rdquo; / &ldquo;so many ling. errors that make abstract incomprehensible&rdquo; / &ldquo;not clear at all&rdquo; (Q1/Q2/Q3) and 5 = &ldquo;excellently written&rdquo; / &ldquo;no errors&rdquo; / &ldquo;completely clear&rdquo; (Q1/Q2/Q3).</p> <p>The pairwise correlations (Kendall&rsquo;s &tau;) of expert judgments on questions Q1-Q3 are presented in this&nbsp;<a href="http://andrea.imis.athena-innovation.gr/readability/table6.png">table</a>.</p> <p>The contained dataset is a tsv file that includes the following fields:</p> <ul> <li>user_id:&nbsp;expert identifier</li> <li>paper_id:&nbsp;AMiner&#39;s identifier from DBLP citation nework v10 dataset</li> <li>rating_1:&nbsp;answer for Q1</li> <li>rating_2:&nbsp;answer for Q2</li> <li>rating_3:&nbsp;answer fro Q3&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Please cite:</strong><br> Thanasis Vergoulis, Ilias Kanellos, Anargiros Tzerefos, Serafeim Chatzopoulos, Theodore Dalamagas, Spiros Skiadopoulos.&nbsp;A study on the readability of scientific publications.&nbsp;<em>23<sup>rd</sup> International Conference on Theory and Practice of Digital Libraries</em>. Oslo, Norway 2019 (to appear)</p>

opencc-by-4.0Apr 2019View details →
zenodo32/100

Data From: ChatGPT versus expert feedback on clinical reasoning questions and their effect on learning: a randomized controlled trial

<p>Dataset Info</p> <p><strong>1) Immediate Test</strong><br>- &nbsp; &nbsp;The first row of the dataset identifies the columns.<br>- &nbsp; &nbsp;Column A represents the participants&rsquo; iDs.<br>- &nbsp; &nbsp;Column B represents the participants&rsquo; assigned group [0: Control (ExpertFeedback) 1: Intervention (ChatGPTFeedback)].<br>- &nbsp; &nbsp;Column C represents the genders of the participants (1: Female, 2: Male).<br>- &nbsp; &nbsp;Column D represents the first-year repetition status of the participants. (0: No, 1: Yes)<br>- &nbsp; &nbsp;Column E to H represent the scores in four different uncomplicated urinary tract infection (UTI) Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column I represents the total scores in uncomplicated UTI Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column J to M represent the scores in four different complicated UTI Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column N represents the total scores in complicated UTI Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column O to R represent the scores in four different pyelonephritis Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column S represents the total scores in pyelonephritis Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column T represents the total scores in immediate test.&nbsp;</p> <p><strong>2) Delayed Test</strong><br>- &nbsp; &nbsp;The first row of the dataset identifies the columns.<br>- &nbsp; &nbsp;Column A represents the participants iDs.<br>- &nbsp; &nbsp;Column B represents the participants&rsquo; assigned group [0: Control (ExpertFeedback) 1: Intervention (ChatGPTFeedback)].<br>- &nbsp; &nbsp;Column C represents the genders of the participants (1: Female, 2: Male).<br>- &nbsp; &nbsp;Column D represents the first-year repetition status of the participants. (0: No, 1: Yes)<br>- &nbsp; &nbsp;Column E to H represent the scores in four different uncomplicated urinary tract infection (UTI) Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column I represents the total scores in uncomplicated UTI Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column J to M represent the scores in four different complicated UTI Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column N represents the total scores in complicated UTI Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column O to R represent the scores in four different pyelonephritis Key-Features Questions Items separately.&nbsp;<br>- &nbsp; &nbsp;Column S represents the total scores in pyelonephritis Key-Features Questions Items.&nbsp;<br>- &nbsp; &nbsp;Column T represents the total scores in delayed test.&nbsp;</p> <p><strong>3) Pre-Intervention Survey on Critical Approach to AI</strong><br>- &nbsp; &nbsp;The first row of the dataset identifies the columns.<br>- &nbsp; &nbsp;Column A represents the participants iDs.<br>- &nbsp; &nbsp;Column B represents the participants&rsquo; assigned group [0: Control (ExpertFeedback) 1: Intervention (ChatGPTFeedback)].<br>- &nbsp; &nbsp;Column C represents the genders of the participants (1: Female, 2: Male).<br>- &nbsp; &nbsp;Column D represents the first-year repetition status of the participants (0: No, 1: Yes).<br>- &nbsp; &nbsp;Column E to J represent the responses of the participants to survey questions before the intervention. Each column is evaluated on a scale from 1 to 7. As it progresses from 1 to 7, the agreement status of participants to survey questions increases. (1: No agreement at all, 7: completely agree)</p> <p><strong>4) Post-Intervention Survey on Critical Approach to AI</strong><br>- &nbsp; &nbsp;The first row of the dataset identifies the columns.<br>- &nbsp; &nbsp;Column A represents the participants iDs.<br>- &nbsp; &nbsp;Column B represents the participants&rsquo; assigned group [0: Control (ExpertFeedback) 1: Intervention (ChatGPTFeedback)].<br>- &nbsp; &nbsp;Column C represents the genders of the participants (1: Female, 2: Male).<br>- &nbsp; &nbsp;Column D represents the first-year repetition status of the participants (0: No, 1: Yes).<br>- &nbsp; &nbsp;Column E to J represent the evaluation of the participants to survey questions after intervention. Each column is evaluated on a scale from 1 to 7. As it progresses from 1 to 7, the agreement status of participants to survey questions increases. &nbsp;(1: No agreement at all, 7: completely agree)</p>

opencc-by-4.0Sep 2024View details →
dryad32/100

Expert-based assessment of rewilding indicates progress at site-level, yet challenges for upscaling

Rewilding is gaining importance across Europe, as agricultural abandonment trajectories provide opportunities for large-scale ecosystem restoration. However, its effective implementation is hitherto limited, in part due to a lack of monitoring of rewilding interventions and their interactions. Here, we provide a first assessment of rewilding progress across seven European sites. Using an iterative and participatory Delphi technique to standardize and analyze expert-based knowledge of these sites, we 1) map rewilding interventions onto the three central components of the rewilding framework (i.e., stochastic disturbances, trophic complexity and dispersal), 2) assess rewilding progress by quantifying 19 indicators spanning human forcing and ecological integrity, and 3) compile key success and threat factors for rewilding progress. We find that the most common interventions were keystone species reintroductions, whereas the least common targeted stochastic disturbances. We find that rewilding scores have improved in five sites, but declined in two, partly due to competing socio-economic trends. Major threats for rewilding progress are related to land-use intensification policies and persecution of keystone species. Major determinants of rewilding success are its societal appeal and socio-economic benefits to local people. We provide an assessment of rewilding that is crucial in improving its restoration outcomes and informed implementation at scale across Europe in this decade of ecosystem restoration.

opencc-zeroJul 2021View details →
zenodo32/100

video games recommendations dataset crafted by human experts

<p>video games recommendations dataset crafted by human experts</p>

opencc-by-4.0Jan 2023View details →
dryad32/100

Fishing effort data, fishing fleet segmentation, and statistical details used in the expert knowledge elicitation experiment

<p><span>Based on an explorative</span> <span>but rigorous elicitation framework, we obtained the bycatch fishing probability</span> <span>at the fishing fleet segment level using expert estimates. Based on the knowledge of three scientific experts, we developed a new and creative structured method for smart and fast fishery-related risk assessments for </span><span>species of high conservation concern. In order to test the method here propose, we applied it to 76 cartilaginous</span><span> fish</span><span> species </span><span>(</span><span>included in the IUCN Red Lists) and on five different fishing segments at both Italian and Mediterranean scale. The method produced qualitative results specific to the threat posed by fishing for each species and each segment with information between and within the segments. Based on the interpretation of resilience-disturbance interactions developed for ecological systems, the quantitative results provided reliable cumulative metrics, measuring the extinction risk due to fishing and the response to overfishing for the species considered. Additionally, the results highlight that the method performs best on a small geographic scale. Therefore, the application of this new method on other subregional</span> <span>or local scales where very few data are available (e.g. fishing effort) could be a valuable tool for the preliminary assessment for species of conservation concern. In fact, despite the absence of detailed catch data at local geographic scales, the flexibility of this method could help to highlight potential fishery-related conservation problems and thus redirect conservation strategies for threatened marine species such as many sharks and rays species.</span></p>

opencc-zeroFeb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record