Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,184
datasets available to search
ShareScore release 0.9.0
Dataset results
1,184 results for “conversion”
Modeling Word Importance in Conversational Transcripts from the Perspective of Deaf and Hard of Hearing Viewers
<ol> <li> <p><strong>MaskedPOSAugmentedData.csv</strong></p> </li> </ol> <p><strong>This file contains word embeddings of 6659 tokens augmented with POS tagging and word importance score. The embedding size for each token is a one dimensional vector of length 768 by 1. Due to augmenting POS tagging, the new feature vector becomes a size of 769 by 1. So now the total feature matrix size is 6659 by 769. In the dataset the final column represents the word importance score. </strong></p> <p><br> </p> <ol> <li> <p><strong>MaskedSentenceWithImportanceTag.csv</strong></p> </li> </ol> <p><strong>This dataset contains the newly generated sentence using standard Masking technique and their corresponding token-wise importance score.</strong><br> </p> <ol> <li> <p><strong>Data cleaning procedure</strong></p> </li> </ol> <p><strong>Using several steps, we have cleaned the “switchboard corpus” that we are using for this particular study. Here are the steps we followed:</strong></p> <p><strong>– Convert all the letter into lower case</strong></p> <p><strong>– Punctuation has be removed</strong></p> <p><strong>– Numbers or Cardinal values have been converted to text representation </strong></p> <p><strong>– We use lemmatization techniques to eradicate the possibility of multiple versions of the same word token.</strong><br> <br> </p> <p> </p> <ol> <li> <p><strong>Annotation Instruction</strong></p> </li> </ol> <p><strong>If researchers intend to produce additional masked text for further the size of the dataset, we recommend using the method we described in the paper.</strong></p> <p><strong>However, if someone wants to manually annotate the words or tokens in a dataset, it is important to remember that annotators need to put a score on each word based on its relative importance within a sentence or the information available around that text. It may not be appropriate to allow annotators to read the whole document first and then conduct annotation. Also using multiple annotators is recommended otherwise interrater agreement may not work well.</strong></p> <p> </p> <ol> <li> <p><strong>Test and training data splitting</strong></p> </li> </ol> <p><strong>As described in the paper, during our experiment, we have retained 10% of data as test data and use 90% data to train the models. For proper replication, we recommend reading our paper thoroughly. </strong></p> <p> </p> <ol> <li> <p><strong>We understand that the dataset size is relatively small for training and testing a model that might be reliable. It is important to remember that data annotation with this particular user group might be challenging. </strong></p> </li> </ol> <p> </p> <p><strong>N.B: While augmenting the new feature within the dataset, a portion of data has been excluded during the data curation phase. For validation, we have replicated the previous models so that we can measure how the dataset can perform with this newly formed dataset. </strong></p>
The visualization of the collected data corresponding to the following paper: "The development of a Self-Rated ICF-based questionnaire (HEAR-COMMAND Tool) to evaluate Hearing, Communication, and Conversation disability: multinational experts' and patients' perspectives"
<p>These two PDF files include the data collected for a study conducted by Afghah et.al, 2022. They include the responses of the participant in this study to a newly developed self-rated ICF-based questionnaire. One file includes the responses to 30 demographic questions and the other one 88 ICF-based questions. The presented data were collected in Germany, the USA, and Egypt as well as overall data.</p> <p>The design of the questionnaire is described here:<br> Afghah, T., Alfakir, R., Meis, M., van Leeuwen, L. M., Kramer, S. E., Hammady, M., Youssif, M., & Wagener, K. C. (2021). The development of a Self-Rated ICF-based questionnaire (HEAR-COMMAND Tool) to evaluate hearing, communication, and conversation disability: multinational experts' and patients' perspectives. Zenodo. https://doi.org/10.5281/zenodo.5534360.</p> <p>This study was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – Project ID: 352015383 – SFB 1330, C4.</p>
INSCIT: Information-Seeking Conversations with Mixed-Initiative Interactions
<p>Data, model checkpoints and output files for the paper INSCIT: Information-Seeking Conversations with Mixed-Initiative Interactions</p>
Original data for publications: Direct solution–based synthesis of the Na4(B12H12)(B10H10) solid electrolyte, Experimental investigation of Mg(B3H8)2 dimensionality, materials for energy storage applications, Thermal Conversion of Unsolvated Mg(B3H8)2 to BH4– in the Presence of MgH2
<p>Original data for the following papers (Only part of University of Geneva, synthesis part):</p> <p>Direct solution–based synthesis of the Na4(B12H12)(B10H10) solid electrolyte<br> Gigante, A.; Duchêne, L.; Moury, R.; Pupier, M.; Remhof, A.; Hagemann, H.<br> ChemSusChem 2019, 12, 4832-4837</p> <p>Experimental investigation of Mg(B3H8)2 dimensionality, materials for energy storage applications<br> Romain Moury, Angelina Gigante, Arndt Remhof, Elsa Roedern and Hans Hagemann<br> Dalton Trans., 2020,49, 12168-12173, https://doi.org/10.1039/D0DT02170A</p> <p>Thermal Conversion of Unsolvated Mg(B3H8)2 to BH4– in the Presence of MgH2<br> Angelina Gigante, Noemi Leick, Andrew S. Lipton, Ba Tran, Nicholas A. Strange, Mark Bowden, Madison B. Martinez, Romain Moury, Thomas Gennett, Hans Hagemann, and Tom S. Autrey<br> ACS Appl. Energy Mater. 2021, 4,3737-3747. https://doi.org/10.1021/acsaem.1c00159</p> <p> </p>
Data for "Ultrafast Spin-Charge Conversion at SnBi2Te4/Co Topological Insulator Interfaces Probed by Terahertz Emission Spectroscopy"
<p>Data for "Ultrafast Spin-Charge Conversion at SnBi2Te4/Co Topological Insulator Interfaces Probed by Terahertz Emission Spectroscopy"</p> <p>(<a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/adom.202102061">https://onlinelibrary.wiley.com/doi/abs/10.1002/adom.202102061</a> and <a href="https://arxiv.org/pdf/2203.08756.pdf">https://arxiv.org/pdf/2203.08756.pdf</a>)</p> <p> </p> <p>E Rongione, S Fragkos, L Baringthon, J Hawecker, E Xenogiannopoulou, P Tsipas, C Song, M Mičica, J Mangeney, J Tignon, T Boulier, N Reyren, R Lebrun, J‐M George, P Le Fèvre, S Dhillon, A Dimoulas, H Jaffrès</p>
Data from: Ultra-low-noise Microwave to Optics Conversion in Gallium Phosphide
<p>Source data for Figures.</p>
Computed data for "Unveiling the Multielectron Acceptor Properties of π‐Expanded Pyracylene: Reversible Boat to Chair Conversion"
<p>Computed structures and TD-DFT raw data as discussed in the article "Unveiling the Multielectron Acceptor Properties of π‐Expanded Pyracylene: Reversible Boat to Chair Conversion" published in JACS, <span>https://doi.org/10.1021/jacs.4c02314</span></p>
Artifacts for paper "Code Refinement with the Assistance of Conversation-Aided Large Language Models"
<p>All the conversation data collected and the source code for our method are inside the uploaded zip file.</p>
Probing conversion-driven freeze-out at the LHC - Code and Data
<p>The file LLP-CDFO-main.zip contains all the code and processed data for reproducing the results in the <a href="https://arxiv.org/abs/2404.16086">Probing conversion-driven freeze-out at the LHC</a> paper.</p> <p>The raw data is provided as a separate tarball.</p> <p>Additional instructions can be found in the README file contained in LLP-CDFO.zip or in the <a href="https://github.com/andlessa/LLP-CDFO">GitHub repository</a>.</p> <p> </p>
Data from: Evolutionary variation in gene conversion at the avian MHC is explained by fluctuating selection, gene copy numbers, and life history
<p>The Major Histocompatibility Complex (MHC) multigene family encodes key pathogen-recognition molecules of the vertebrate adaptive immune system. Hyper-polymorphism of MHC genes is <em>de novo</em> generated by point mutations, but new haplotypes may also arise by re-shuffling of existing variation through intra- and inter-locus gene conversion. Although the occurrence of gene conversion at the MHC has been known for decades, we still have limited understanding of its functional importance. Here, I took advantage of extensive genetic resources (~9000 sequences) to investigate a broad scale macroevolutionary patterns in gene conversion processes at the MHC across nearly 200 avian species. Gene conversion was found to constitute a universal mechanism in birds, as 83% of species showed footprints of gene conversion at either MHC class and 25% of all allelic variants were attributed to gene conversion. Gene conversion processes were stronger at MHC-II than MHC-I, but inter-specific variation at both MHC classes was explained by similar evolutionary scenarios, reflecting fluctuating selection towards different optima and drift. Gene conversion showed uneven phylogenetic distribution across birds and was driven by gene copy number variation, supporting significant role of inter-locus gene conversion processes in the evolution of the avian MHC. Finally, MHC gene conversion was stronger in species with fast life histories (high fecundity) and in long-distance migrants, likely reflecting variation in population sizes and host-pathogen coevolutionary dynamics. The results provide a robust comparative framework for understanding macroevolutionary variation in gene conversion at the avian MHC and reinforce important contribution of this mechanism to functional MHC diversity.</p>
Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES) Dataset - Anonymized
<p>Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES) and this corresponding dataset aim to provide tools for measuring user enjoyment from an external perspective to supplement self-reported user enjoyment responses in human-robot interaction research, with future potential application for autonomous detection of user enjoyment in real-time in robots and agents for adapting conversations contingently to provide enjoyable and long-lasting interactions.</p> <p>The dataset consists of 25 older adults' (12 men, 13 women) open-domain dialogue with an autonomous companion robot with an integrated large language model (GPT-3.5, text-davinci-003) from participatory design workshops conducted in March 2023. The conversations are annotated for user enjoyment based on HRI CUES by 3 expert annotators, as described in the paper (arXiv:2405.01354). Robot architecture and participatory design workshops are described in DOI: 10.21203/rs.3.rs-2884789/v1.</p> <p><strong>Exchanges</strong> file contains the participant ID, the number of the turn (conversation exchange by Robot-Participant response), the start and end of the turn, the anonymized transcript for the turn, and three annotator scores for the user enjoyment in the exchange. </p> <p><strong>Overall </strong>file contains the participant ID, self-reported user perception scores from the questionnaire ("I was satisfied with my conversation with the robot", "It was fun talking to the robot", "The conversation with the robot was interesting", "It felt strange talking to the robot") and three annotator scores for the user enjoyment in the overall interaction.</p> <p>The conversations are in Swedish. Participants' mean age is 74.6 (SD=5.8). 20 participants had no prior interaction with a robot, and only one had previously talked with a robot. The average interaction duration is 7.4 min (SD=1.5) with 12 to 29 turns. Each turn lasts 5 to 61 seconds (M=17.7, SD=7.2). The total duration of the interactions is 174 min, corresponding to 590 turns. </p> <p><em>Videos of the interactions are available upon request, contingent upon a signed agreement to maintain data confidentiality in accordance with GDPR regulations.</em></p> <p>Anonymization macros:</p> <p>[P_NAME]: Participant's name (may include surname). The robot always uses the first name even when the surname is given.</p> <p>[NAME_REMOVED]: A name of another person mentioned by the participant.</p> <p>[LOCATION_REMOVED]: Small town/village/area where the participant lives or lived.</p> <p>[MEDICAL_INFO_REMOVED]: Medical information shared by the participant.</p> <p>[AGE_REMOVED]: Participant's or other person's age.</p> <p>[INFORMATION_REMOVED]: Sensitive information shared by the participant.</p> <p>[MISTAKEN_NAME]: Speech recognition error resulted in the name being misunderstood.</p>
Dataset for the manuscript "Osmotic Energy Conversion in Serpentinite-Hosted Deep-Sea Hydrothermal Vents"
Open the record for dataset details and reuse information.
WangchanThaiInstruct Multi-turn Conversation Dataset
<h1>WangchanThaiInstruct Multi-turn Conversation Dataset</h1> <p>We create a Thai multi-turn conversation dataset from <a href="https://huggingface.co/datasets/airesearch/WangchanThaiInstruct" target="_blank" rel="noopener">airesearch/WangchanThaiInstruct (Batch 1)</a> by LLM. It was created from synthetic method using open source LLM in Thai language.</p>
Audiovisual recordings of acted casual conversations between four speakers in German
<p>This dataset was used in a study by Hendrikse et al. (2018). It contains 7 audiovisual recordings of acted casual conversations between 4 actors (two males and two females), as well as 2-3 questions about the content of each of the conversations asked by one of the speakers. The conversations last between 1min 24s to 1min 39s and the topics are food, holidays/travelling, weather, work, plans, movies and anecdotes. Two of the actors are native speakers and two are non-native speakers with a C1 level of German. The actors memorized the script and acted as in a casual conversation. The audiovisual material was recorded in a circular space with uniform purple background using two cameras (Canon EOS 700D) and one microphone for each speaker (Neumann KM 184). <br> Additionally the visual stimuli used in the aforementioned study with virtual characters is attached: a screencast of a simulation of each conversation with virtual characters (Makehuman) with gaze to target speaker and lip-syncing (Llorach et al. 2016). The 3D scene is rendered with Blender3D with a camera with a 120º field-of-view.</p> <p> </p> <p>References:</p> <p>M.M.E. Hendrikse, G. Llorach, G. Grimm, V. Hohmann. Influence of visual cues on head and eye movements during listening tasks in multi-talker audiovisual environments with animated characters, Accepted in Special Issue on Realism in Robust Speech and Language Processing of Speech Communication, 2018.</p> <p>Llorach G, Evans A, Blat J, Grimm G, Hohmann V. Web-based live speech-driven lip-sync. In Games and Virtual Worlds for Serious Applications (VS-Games), 2016 8th International Conference on 2016 Sep 7 (pp. 1-4). IEEE.</p>
Identifier Refinery Conversion Matrixes
<p><strong>identifier-refinery</strong></p> <p>Tools and assets for easy and reproducable gene identifier conversion.</p> <p><strong>Methods</strong></p> <p>This repository is used to build matricies which can convert between different gene identifiers.</p> <p>These conversion matricies are built by:</p> <ul> <li>Randomly choosing raw CEL files from NCBI GEO for a given platform accession code (in <code>/cels</code>)</li> <li>Reading the CEL header and joining Brainarray (e.g., <code>hgu133plus2hsensgprobe</code>) and Bioconductor (e.g., <code>hgu133plus2.db</code>) (x, y) coordinates</li> <li>Finding intersecting probe identifiers</li> <li>Extracting supported identifiers and probe IDs from the Bioconductor package</li> <li>Filtering on probe IDs and Ensembl Gene IDs in Brainarray</li> <li>Writing the output to a conversion TSV file</li> <li>Check that all output conversion TSV files have a shared SHA1</li> </ul> <p><strong>Repository Contents</strong></p> <p><strong>Source Files</strong></p> <p>The <code>cels</code> directory contains raw CEL files taken from GEO. The list of supported platforms is in <code>supported_microarray_platforms.csv</code>. Source files can be acquired by running the <code>acquire_cels.py</code> script.</p> <p><strong>Docker Image</strong></p> <p>The conversion scripts are run on custom Docker images.</p> <p>Two Dockerfiles are provided in this repository - <code>base</code> Docker image, which is used to install the quire R dependancies, and the <code>pd</code> image, which is used to build the required databases for a given platform.</p> <p><strong>Conversion Scripts</strong></p> <p>A <code>build_and_convert.py</code> script is provided, which build a unique Docker image for each package, mount the downloaded CEL files as a volume, and then run the gene conversion script <code>R/gene_convert.R</code> inside the image and output the master conversion matrix. Output TSV files live in <code>cels/out/</code>.</p> <p><strong>Reproducing</strong></p> <p>The entire process can be reproduced by running the following command script from a fresh checkout of this repository. It will take some time:</p> <pre><code>$ ./generate_matricies_from_scratch.sh </code></pre> <p>You can also choose to only build a specific platform, ex.,:</p> <pre><code>$ ./generate_matricies_from_scratch.sh celegans </code></pre> <p><strong>Identifiers</strong></p> <p>Released assets in this repository are availble under the DOI, <code>xyz:1.2.3.4</code>, which can be seen on Zenodo <a href="https://link.todo">here</a>.</p> <p><strong>Related Projects</strong></p> <ul> <li><a href="https://github.com/AlexsLemonade/refinebio">AlexsLemonade/refinebio</a></li> </ul> <p><strong>Copyright</strong></p> <p><code>identifier-refinery</code> output assets are released under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal</a> license. All code is released under the BSD 3-clause license. Input assets are property of the original providers to NCBI GEO, but may be <a href="https://www.ncbi.nlm.nih.gov/geo/info/disclaimer.html">freely downloaded and redistributed</a> unless otherwise noted.</p> <p> </p> <p>https://github.com/AlexsLemonade/identifier-refinery</p>
Replication data for Nematzadeh et al. "Information Overload in Group Communication: From Conversation to Cacophony in the Twitch Chat"
<p>A subset of the chat logs dump from Twitch used in this work is provided, to help replicate the central findings of this work (<a href="https://doi.org/10.5281/zenodo.1182793">https://doi.org/10.5281/zenodo.1182793</a>). Data are aggregated and include the number of messages posted in each channel and the number of users posting them, sampled at intervals of 5 minutes. To protect the identity of the users in this data collection, message contents and user names are not included in this dataset. Stream names have been replaced with numeric IDs. <br> No additional filtering or data cleaning operation has been applied to this data. Replication code is available on Github (<a href="https://github.com/glciampaglia/twitch-overload-replication">https://github.com/glciampaglia/twitch-overload-replication</a>).</p>
Quantifying the spatial heterogeneity of forest conversion costs and how it relates to biodiversity, conservation and land use history
<b>Description: </b><p>Start and end dates of salvage logging activity at the SAFE Project experimental site</p><p><b>Project: </b>This dataset was collected as part of the following SAFE research project: <a href="https://www.safeproject.net/projects/project_view/6"><b>Quantifying the spatial heterogeneity of forest conversion costs and how it relates to biodiversity, conservation and land use history</b></a></p><p><b>Funding: </b>These data were collected as part of research funded by: </p><ul><li>Sime Darby (grant)</li></ul><p>This dataset is released under the CC-BY 4.0 licence, requiring that you cite the dataset in any outputs, but has the additional condition that you acknowledge the contribution of these funders in any outputs.</p><p></p><p><b>Permits: </b>These data were collected under permit from the following authorities:</p><ul><li>Sabah Biodiversity Council (Research licence na)</li></ul><p></p><p><b>XML metadata: </b>GEMINI compliant metadata for this dataset is available <a href="https://www.safeproject.net/datasets/xml_metadata?id=3266827">here</a></p><p><b>Files: </b>This dataset consists of 2 files: template_Symes.xlsx, SAFE_COUPE.zip</p><p><b>template_Symes.xlsx</b></p><p>This file contains dataset metadata and 1 data tables:</p><ol><li><p><b>Salvage logging records</b> (described in worksheet Data)</p><p>Description: Dates of earliest and latest known salvage logging activity in logging coupes</p><p>Number of fields: 10</p><p>Number of data rows: 187</p><p>Fields: </p><ul><li><b>CoupeNumber</b>: Coupe number (Field type: Location)</li><li><b>StartDateTrack</b>: Earliest date of logging activity recorded through GPS loggers on bulldozers (Field type: Date)</li><li><b>EndDateTrack</b>: Last date of logging activity recorded through GPS loggers on bulldozers (Field type: Date)</li><li><b>StartDateLocation</b>: Earliest date of logging activity recorded through 'Location' method (Field type: Date)</li><li><b>EndDateLocation</b>: Last date of logging activity recorded through 'Location' method (Field type: Date)</li><li><b>StartDateMeasurement</b>: Earliest date of logging activity recorded through 'Measurement' method (Field type: Date)</li><li><b>EndDateMeasurement</b>: Last date of logging activity recorded through 'Measurement' method (Field type: Date)</li><li><b>Contractor</b>: Name of the contractor responsible for the coupe (Field type: ID)</li><li><b>GlobalStart</b>: Earliest date of any recorded salvage logging activity (Field type: Date)</li><li><b>GlobalEnd</b>: Last date of any recorded salvage logging activity (Field type: Date)</li></ul></li></ol><p><b>SAFE_COUPE.zip</b></p><p>Description: SAFE coupe data</p><p><b>Date range: </b>2013-01-02 to 2015-12-31</p><p><b>Latitudinal extent: </b>4.5000 to 5.0700</p><p><b>Longitudinal extent: </b>116.7500 to 117.8200</p>
Logging of rainforest and conversion to oil palm reduces bioturbator diversity but not levels of bioturbation.
<p>This dataset contains data about bioturbation performed by different soil animals in three types of differently degraded habitats in Sabah, Borneo. This include not only levels of bioturbation, but also data about the numbers, growth and turnover of termite mounds plus litter depth for each sampling point. </p>
Mono-to-Stereo Conversion Evaluation Dataset
<p>MONO-TO-STEREO CONVERSION EVALUATION DATASET.</p> <p>This dataset has been constructed from audio excerpts taken from the Bach10 dataset by Duan et al. [1]. This database has been used to evaluate the effectiveness of a mono-to-stereo converter by means of a listening test. Results were presented in my PhD thesis [2], in Chapter 6, on pages 167-176.</p> <p>[1] Z. Duan, B. Pardo, and C. Zhang, "Multiple fundamental frequency estimation by modelling spectral peaks and non-peak regions," IEEE Transactions on Audio, Speech and Language Processing, vol. 18, no. 8, pp. 2121-2133, 2010.</p> <p>[2] Delgado Castro, A. "Iterative separation of note events from single-channel Polyphonic Recordings," Ph.D. University of York. 2019.</p>
Dwelling conversion and energy retrofit modify building anthropogenic heat emission under past and future climates: a case study of London terraced houses
<p>This archive includes the data used (e.g. Time use survey (UK-TUS) data), model files (idf files for running EnergyPlus) and codes for analysis in the paper (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.enbuild.2024.114668" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.enbuild.2024.114668</a>).</p> <p>Files in this archive should include:</p> <ul> <li>Time use survey data analysis</li> </ul> <p>o Main dataset: TUS_activity.zip</p> <p>o Code: TUS_clustering_code.zip</p> <p>o Output: InternalHeatProfile.zip</p> <ul> <li>Building energy modeling </li> </ul> <p>o Main dataset (run in EnergPlus 9.4): IDFfiles.zip</p> <p>o Output: Eplus_output.zip</p> <ul> <li>PostProcess analysis</li> </ul> <p>o Code: QF_analysis_code.zip</p> <p>o Output: QF_output.zip</p> <p> </p> <p>Note: this version currently only includes the outputs of all processes, the main dataset and code will be updated later.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.