Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,184

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,184 results for “conversations”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: A pantropical analysis of the impacts of forest degradation and conversion on local temperature

Temperature is a core component of a species' fundamental niche. At the fine scale over which most organisms experience climate (mm to ha), temperature depends upon the amount of radiation reaching the Earth's surface, which is principally governed by vegetation. Tropical regions have undergone widespread and extreme changes to vegetation, particularly through the degradation and conversion of rainforests. Since most terrestrial biodiversity is in the tropics, and many of these species possess narrow thermal limits, it is important to identify local thermal impacts of rainforest degradation and conversion. We collected pan-tropical, site-level (< 1 ha) temperature data from the literature to quantify impacts of land-use change on local temperatures, and to examine whether this relationship differed above-ground relative to below-ground and between wet and dry seasons. We found that local temperature in our sample sites was higher than primary forest in all human-impacted land-use types (N = 113,894 day-time temperature measurements from 25 studies). Warming was pronounced following conversion of forest to agricultural land (minimum +1.6°C, maximum +13.6°C), but minimal and non-significant when compared to forest degradation (e.g. by selective logging; minimum +1°C, maximum +1.1°C). The effect was buffered below-ground (minimum buffering 0°C, maximum buffering 11.4°C), whereas seasonality had minimal impact (maximum buffering 1.9°C). We conclude that forest-dependent species that persist following conversion of rainforest have experienced substantial local warming. Deforestation pushes these species closer to their thermal limits, making it more likely that compounding effects of future perturbations, such as severe droughts and global warming, will exceed species' tolerances. By contrast, degraded forests and below-ground habitats may provide important refugia for thermally-restricted species in landscapes dominated by agricultural land.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure

PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Tropical rainforest conversion and land-use intensification reduce understory plant phylogenetic diversity

1. Conversion of rainforest into agricultural land affects multiple facets of tropical plant diversity. While the effects of tropical land use change and intensification on species diversity are comparatively well studied, the effects on phylogenetic diversity and structure of plant communities are largely unknown. Furthermore, it is not clear how the loss of native species and addition of alien species collectively affect phylogenetic diversity and structure. 2. We investigated the phylogenetic diversity and structure of understorey plants; a diverse and ecologically important, yet poorly studied group. We studied four prominent land use systems (tropical lowland rainforest, jungle rubber agroforest, rubber plantations and oil palm plantations) in the lowlands of Sumatra (Indonesia), a region experiencing dramatic land use changes. 3. Across the four systems, we investigated differences in four metrics of phylogenetic community structure (phylogenetic diversity, mean pairwise distance, mean nearest taxon distance and their abundance-weighted variants). Our analyses were based on a comprehensive vegetation survey consisting of 32 plots, 1,197 species of vascular plants, and 146,599 plant individuals. 4. Our results showed that forest conversion into agricultural systems leads to a pronounced loss of phylogenetic diversity. Furthermore, the standard effect size of mean pairwise distance indicated a gradual change from clustered to overdispersed phylogenetic community structure with increasing land use intensity from forest over jungle rubber to the monoculture plantations. In most land use systems, the presence or absence of alien plant species did not affect phylogenetic structure. Only in oil palm plantations, removing alien species from the data led to a more overdispersed structure. In conclusion, conserving the phylogenetic diversity and structure requires efficient protection of the last remaining rainforests. 5. Synthesis and applications. Forest conversion into agricultural areas negatively affects phylogenetic understorey plant diversity and leads to a shift from clustered to overdispersed phylogenetic community structure. These trends are partly driven by alien species particularly in oil palm plantations. Protecting the remaining rainforests, and considering multi-species agroforestry systems in favour of intensive monoculture plantations are thus imperative to conserve phylogenetic plant diversity and community structure.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Gene conversion yields novel gene combinations in paralogs of GOT1 in the copepod Tigriopus californicus

Background: Gene conversion of duplicated genes can slow the divergence of paralogous copies over time but can also result in other interesting evolutionary patterns. Islands of genetic divergence that persist in the face of gene conversion can point to gene regions undergoing selection for new functions. Novel combinations of genetic variation that differ greatly from the original sequence can result from the transfer of genetic variation between paralogous genes by rare gene conversion events. Genetically divergent populations of the copepod Tigriopus californicus provide an excellent model to look at the patterns of divergence among paralogs across multiple independent evolutionary lineages. Results: In this study the evolution of a set of paralogous genes encoding putative aspartate transaminase proteins (called GOT1 here) are examined in populations of the copepod T. californicus. One pair of duplicated genes, GOT1p1 and GOT1p2, has regions of high divergence between the copies in the face of apparent on-going gene conversion. The GOT1p2 gene also has unique haplotypes in two populations that appear to have resulted from a transfer of genetic variation via inter-paralog gene conversion. A second pair of duplicated genes GOT1Sr and GOT1Sd also shows evidence of gene conversion, but this gene conversion does not appear to have maintained each as a functional copy in all populations. Conclusions: The patterns of conservation and sequence divergence across this set of paralogous genes among populations of T. californicus suggest that some interesting evolutionary patterns are occurring at these loci. The results for the GOT1p1/GOT1p2 paralogs illustrate how gene conversion can factor in the creation of a mosaic pattern of regions of high divergence and low divergence. When coupled with rare gene conversion events of divergent regions, this pattern can result in the formation of novel proteins differing substantially from either original protein. The evolutionary patterns across these paralogs show how gene conversion can both constrain and facilitate diversification of genetic sequences.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Disparity in preemptive end-of-life conversation experience caused by subjective economic status among general Japanese elderly people: a cross-sectional study with stratified random sampling

Objectives: Preemptive conversations (PCs) about end-of-life (EOL) preferences are beneficial for both elderly people and their families to understand and share the preferences. However, the factors which promote/inhibit PCs have yet to be clarified. We therefore aimed to determine the factors related to having PCs with hypothesis that age, subjective economic status and subjective health status are associated with having PC experience. Design: A cross-sectional study administering a questionnaire and using stratified random sampling by gender and region. Setting: Residents aged 65 years or older who were not receiving nursing care as of November 1, 2016, were extracted from the Japanese long-term care insurance system registry in Koriyama City, Fukushima Prefecture, Japan. Participants: 1,575 participants (717 males and 858 females). Outcome: Presence or absence of PC experience with family or friends (yes/no). Results: The mean age of the participants was 74.0 years. A multivariable logistic-regression analysis revealed that having PC experience was significantly associated with gender (OR = 1.907; 95% CI = 1.556, 2.337; p < 0.001), subjective economic status (OR = 0.832; 95% CI = 0.716, 0.966; p=0.016), and subjective happiness (OR = 0.926; 95% CI = 0.880, 0.973; p=0.003). Conclusions: Poor subjective economic status of elderly people may result in absence of EOL conversation experience with their families and friends, hindering the elderly from sharing and understanding the EOL preferences. To promote PCs about EOL, gerontology and public health professionals should give special consideration to the subjective economic status of elderly people.

opencc-zeroSep 2019View details →
zenodo32/100

Final geometries and energies, statistical analysis and estimated errors of single metals and bimetallics for CO2 to methanol conversion

<p>The dataset accommodate all the extra data discussed in:<br>Pisal, P., Krejč&iacute;, O. &amp; Rinke, P. Machine learning accelerated descriptor design for catalyst discovery in CO<sub>2</sub> to methanol conversion. <em>npj Comput Mater</em> <strong>11</strong>, 213 (2025). https://doi.org/10.1038/s41524-025-01664-9&nbsp;</p> <p>The datased contains four types of data:</p> <ol> <li>All the final geometries and energies of adsorbated (*H, *O, *OCHO &amp; *OCH3) and all the 158 single metals and bimetallic alloys on all the surfaces with Miller indices in {-2, -1, ... 2} optimized with Open Catalyst Project (OCP) 20 <em>equiformer_V2</em> machine-learned force-field model. These are in the <a href="https://zenodo.org/api/records/15587232/draft/files/geometries_and_energies.zip/content" target="_blank" rel="noopener noreferrer">geometries_and_energies.zip</a> file organized by the metal/alloys name, with the final geometries and enerigies in a json file, using a json ASE format.</li> <li>All the estimated mean absolute errors (MAE) of predicted adsorption energies for all the considered metals and bimetallic alloys in&nbsp;<a href="https://zenodo.org/api/records/15587232/draft/files/Estimated_MAEs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Estimated_MAEs_metals_bimetallics.csv</a> and xlsx file. The data content is identical, files differs only by a format.</li> <li>All the adsorption energy disctibutions (AEDs) for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> <li>All the statistical information of the adsorption energies for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/Statistics_AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Statistics_AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> </ol>

opencc-by-4.0Aug 2024View details →
zenodo32/100

DevGPT: Studying Developer-ChatGPT Conversations

<p>DevGPT is a curated dataset which encompasses 17,913 prompts and ChatGPT's responses including 11,751 code snippets, coupled with the corresponding software development artifacts—ranging from source code, commits, issues, pull requests, to discussions and Hacker News threads—to enable the analysis of the context and implications of these developer interactions with ChatGPT.</p><p><strong>Important</strong><br>Version 9 (2023-11-09) resolves the empty list of conversations attribute: https://github.com/NAIST-SE/DevGPT/issues/8</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Workflow Run RO-Crate capturing provenance from WSI conversion

<p>Example of <a href="https://www.researchobject.org/workflow-run-crate/profiles/">Workflow Run RO-Crate</a> capturing provenance data from an execution of the <a href="https://github.com/crs4/fair-crcc-img-convert/tree/main">fair-crcc-img-convert</a> workflow on a whole-slide image from the <a href="https://doi.org/10.7937/25T7-6Y12">Cancer Moonshot Biobank - Prostate Cancer Collection (CMB-PCA)</a>.</p><ul><li>Slide ID: MSB-02917-01-02, generated by Natasha Honomichl</li><li>Image License: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></li></ul><p>Note that the license for the RO-Crate is CC BY 4.0, except for the workflow, which is licensed under the <a href="https://www.gnu.org/licenses/gpl-3.0.en.html">GPL-3.0</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Dataset of conversation graphs from 2019 Anti CAA protest in India

<p>To facilitate future research in online social movements, &nbsp;we are publishing this dataset of conversation graphs in the from the 2019 anti-CAA protest tweets. &nbsp;The conversation graph dataset has two types of conversations:</p><p>• <strong>Mention Graphs</strong>: In this, vertices are users and (directed) edges represent any interaction between the users. Within the context of our dataset, these interactions are limited to mentions alone when a user refers to another user directly by username.</p><p>• <strong>Reply Graphs:</strong> In this case, vertices are users, and a (directed) edge exists if one user replies to another.</p><p><strong>The reply graph is built as follows:</strong></p><ul><li>Since replies often form chains, find the root 'destination' user nodes that have been replied to at least once.</li><li>&nbsp;Find all the immediate next-level users who directly replied to the root nodes (these are 'source' nodes)</li><li>Iteratively map the immediate next level of users who reply to the previous level of corresponding users, stopping when all of the users in the next level have not been replied to (leaf users, so to say).</li></ul><p><strong>The mention graph is built as follows:</strong></p><ul><li>Since replies often form chains, find the root 'destination' user nodes that have been replied to at least once.</li><li>Find all the immediate next-level users who directly replied to the root nodes (these are 'source' nodes)</li><li>Iteratively map the immediate next level of users who reply to the previous level of corresponding users, stopping when all of the users in the next level have not been replied to (leaf users, so to say).</li></ul><p>Additionally, the dataset has the users and conversation graphs labeled for emotion and toxicity. &nbsp;Motif count has been provided for each of the graph.&nbsp;</p><p>To maintain confidentiality as per Twitter guidelines, user ids are anonymized &nbsp;and Tweet IDs and Tweet texts are not part of the dataset.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Non-crystalline Zeolitic Imidazolate Frameworks Tethered with Ionic Liquids as Catalysts for CO2 Conversion into Cyclic Carbonates

<p>This folder /final_logs/ contains the DFT-optimized geometries (in .xyz format together with the gas-phase energy, E) accompanying the paper</p> <p>"Design of Non-crystalline Zeolitic Imidazolate Frameworks Tethered with Ionic Liquids as Highly Active and Stable Catalysts for Mild CO2 Conversion into Cyclic Carbonates"</p> <p>where conformers occur, they are always named from the lowest Gibbs energy to the highest in ascending order from c1 (sometimes omitted), c2, c3, ...</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Replication Package for ML-EUP Conversational Agent Study

<p>This is the replication package of the paper <a href="https://conf.researchr.org/details/icse-2024/icse-2024-research-track/5/How-to-Support-ML-End-User-Programmers-through-a-Conversational-Agent">How to Support ML End-User Programmers through a Conversational Agent</a>, published at ICSE 2024.</p> <p><strong>Replication Package Files</strong></p> <ul> <li><strong>Readme.pdf:</strong> document that describes the replication package and indicates how to use it.&nbsp;</li> <li> <p><strong>1. Forms.zip: </strong>contains the forms used to collect data for the experiment.</p> </li> <li> <p><strong>2. Experiments.zip: </strong>contains the participants&rsquo; and sandboxers&rsquo; experimental task workflow with Newton.</p> </li> <li> <p><strong>3. Responses.zip: </strong>contains the responses collected from participants during the experiments.</p> </li> <li> <p><strong>4. Analysis.zip:</strong> contains the data analysis scripts and results of the experiments.</p> </li> <li> <p><strong>5. newton.zip: </strong>contains the tool we used for the WoZ experiment.</p> </li> <li> <p><strong>Interactions.pdf: </strong>explains Figure 4 of the paper in detail by depicting the interactions of P4.</p> </li> <li> <p><strong>TutorialStudy.pdf:</strong> script used in the experiment with and without Newton to be consistent with all participants.</p> </li> <li> <p><strong>Woz_Script.pdf:</strong> script wizard used to maintain consistent Newton responses among the participants.</p> </li> <li><strong>Dockerfile</strong>: docker definition of newton-docker.tar.gz.</li> <li> <p><strong>newton-docker.tar.gz:</strong> docker image that contains both the tool and the analysis files.</p> </li> <li><strong>LICENSE:&nbsp;</strong>license file describing the license of data and code files.</li> </ul> <p>&nbsp;</p> <p><strong>1. Forms.zip</strong></p> <p>The forms zip contains the following files:</p> <ul> <li> <p><strong>Demographics.pdf: </strong>a PDF form used to collect demographic information from participants before the experiments</p> </li> <li> <p><strong>Post-Task Control (without the tool).pdf:</strong> a PDF form used to collect data from participants about challenges and interactions when performing the task without Newton&nbsp;</p> </li> <li> <p><strong>Post-Task Newton (with the tool).pdf:</strong> a PDF form used to collect data from participants after the task with Newton.</p> </li> <li> <p><strong>Post-Study Questionnaire.pdf:</strong> a PDF form used to collect data from the participant after the experiment.</p> </li> </ul> <p>&nbsp;</p> <p><strong>2. Experiments.zip</strong></p> <p>The experiments zip contains two types of folders:</p> <ul> <li> <p><strong>exp[participant&rsquo;s number]-c[number of dataset used for control task]e[number of dataset used for experimental task]</strong>. Example: exp1-c2e1 (experiment participant 1 - control used dataset 2, experimental used dataset 1)</p> </li> <li> <p><strong>sandboxing[sandboxer&rsquo;s number].</strong> Example: sandboxing1 (experiment with sandboxer 1)</p> </li> </ul> <p>&nbsp;</p> <p>Every experiment subfolder contains:</p> <ul> <li> <p><strong>warmup.json: </strong>a JSON file with the results of Newton-Participant interactions in the chat for the warmup task.</p> </li> <li> <p><strong>warmup.ipynb: </strong>a Jupyter notebook file with the participant&rsquo;s results from the code provided by Newton in the warmup task.</p> </li> <li> <p><strong>sample1.csv: </strong>Death Event dataset.</p> </li> <li> <p><strong>sample2.csv: </strong>Heart Disease dataset.</p> </li> <li> <p><strong>tool.ipynb: </strong>a Jupyter notebook file with the participant&rsquo;s results from the code provided by Newton in the experimental task.</p> </li> <li> <p><strong>python.ipynb: </strong>a Jupyter notebook file with the participant&rsquo;s results from the code they tried during the control task.</p> </li> <li> <p><strong>results.json:</strong> a JSON file with the results of Newton-Participant interactions in the chat for the task with Newton.</p> </li> </ul> <p>&nbsp;</p> <p>To load an experiment chat log into Newton, add the following code to the notebook:</p> <pre><code>import anachat import json with open("result.json", "r") as f: anachat.comm.COMM.history = json.load(f) </code></pre> <p>Then, click on the notebook name&nbsp;inside Newton chat</p> <p>Note 1: the subfolder for P6 is exp6-e2c1-serverdied because the experiment server died before we were able to save the logs. We reconstructed them using the notebook newton_remake.ipynb based on the video recording.</p> <p>Note 2: The sandboxing occurred during the development of Newton. We did not collect all the files, and the format of JSON files is different than the one supported by the attached version of Newton.</p> <p>&nbsp;</p> <p><strong>3. Responses.zip</strong></p> <p>The responses zip contains the following files:</p> <ul> <li> <p><strong>demographics.csv: </strong>a CSV file containing the responses collected from participants using the demographics form</p> </li> <li> <p><strong>task_newton.csv: </strong>a CSV file containing the responses collected from participants using the post-task newton form.</p> </li> <li> <p><strong>task_control.csv:</strong> a CSV file containing the responses collected from participants using the post-task control form.</p> </li> <li> <p><strong>post_study.csv:</strong> a CSV file containing the responses collected from participants using the post-study control form.</p> </li> </ul> <p>&nbsp;</p> <p><strong>4. Analysis.zip</strong></p> <p>The analysis zip contains the following files:</p> <ul> <li> <p><strong>1.Challenge.ipynb:</strong> a Jupyter notebook file that performs the statistical tests and creates the perceptions of challenges figure.</p> </li> <li> <p><strong>2.Interactions.py: </strong>a Python file that creates the&nbsp;participants&rsquo; JSON files.</p> </li> <li> <p><strong>3.Interactions.Graph.ipynb: </strong>a Jupyter notebook file that creates the participant&rsquo;s interaction figure.</p> </li> <li> <p><strong>4.Interactions.Count.ipynb:</strong> a Jupyter notebook file that counts participants&rsquo; interaction with each figure.</p> </li> <li> <p><strong>config_interactions.py:</strong> this file contains the definitions of interaction colors and grouping</p> </li> <li> <p><strong>interactions.json:&nbsp;</strong>a JSON file with the interactions during the Newton task of each participant based on the categorization.</p> </li> <li> <p><strong>requirements.txt: </strong>dependencies required to run the code to generate the graphs and json analysis.</p> </li> </ul> <p>&nbsp;</p> <p>To run the analyses, please follow the steps:</p> <p>1- Extract Analysis.zip and cd into the directory</p> <p>2- Install Python 3.10, and then the analysis dependencies with the following command:</p> <pre><code>pip install -r requirements.txt</code></pre> <p>3- Run Jupyter Notebook/Lab and execute all cells of <strong>1.Challenge.ipynb</strong>. It will generate the challenges figure.</p> <p>4- Run <strong>2.Interactions.py</strong> using the following command:</p> <pre><code>python 2.Interactions.py</code></pre> <p>This file was created manually by individually categorizing each interaction of the participants. The execution will generate the file interactions.json with the graph definitions of the interactions.</p> <p>5- Run Jupyter Notebook/Lab and execute all cells of <strong>3.Interactions.Graph.ipynb</strong>. It will create the interactions graph visualization.</p> <p>6- Run Jupyter Notebook/Lab and execute all cells of <strong>4.Interactions.Count.ipynb</strong>. It will create the interactions table.</p> <p>&nbsp;</p> <p><strong>5. newton.zip</strong></p> <p>The newton zip contains the <strong>source code of the Jupyter Lab extension</strong> we used in the experiments. Read the <strong>README.md</strong> file inside it for instructions on how to install and run it.</p> <p>&nbsp;</p> <p><strong>6. newton-docker.tar.gz</strong></p> <p>This file contains the Docker image with the replication package in a configured environment for both the tool and the dat analyses.</p> <p>To import the image, run:</p> <pre><code><span>docker load &lt;</span> <span>newton-docker.tar.gz</span></code></pre> <p>Then, start the container by running:</p> <pre><code><span>docker run -p 8888:8888 -it newton</span></code></pre> <p>Finally, start Jupyter Lab:</p> <pre><code><span>jupyter lab --collaborative --ip="*" --port=8888 --allow-root</span></code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Search Protocol for "Conversational Systems for AI-Augmented Business Process Management"

<p>Results obtained from implementing the search protocol devised for the literature survey "<em>Conversational Systems for AI-Augmented Business Process Management</em>".</p> <p>The dataset comprises four spreadsheets, each corresponding to one of the four BPM areas identified in the paper, namely:</p> <ul> <li><em>Descriptive Process Analytics</em>;</li> <li><em>Predictive Process Analytics</em>;</li> <li><em>Prescriptive Process Optimization</em>;</li> <li><em>Augmented Process Execution</em>.</li> </ul> <p>Each spreadsheet consists of multiple sheets:</p> <ul> <li>The first four sheets document the papers collected from each data source (<em>Google Scholar</em>, <em>Scopus</em>, <em>ACM Digital Library</em>, <em>IEEE Xplore</em>) by applying the search strings defined in the paper.</li> <li>"All" reports all the papers obtained in the search.</li> <li>"All(-duplicates)" lists all publications, excluding duplicates.</li> <li>"Inclusion" applies the inclusion criteria defined in the paper to select the works considered in this survey.</li> <li>"Final" comprises the selected papers, representing the outcomes of the search protocol's application.</li> <li>"Results" provides statistical insights into the application of the search protocol for the specific BPM area under analysis."</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Understanding rate and capacity limitations in Li-S batteries based on solid-state sulfur conversion in confinement

<p>Raw data sets of the publication A. Senol G&uuml;ng&ouml;r et al. Understanding rate and capacity limitations in Li-S batteries based on solid-state sulfur conversion in confinement, 2024.</p> <p>The dataset was generated within the ALISA project (project number 9359) provided by the m-ERA.NET network (part of the European Union&rsquo;s Horizon 2020 research and innovation program, and the ERC Starting-Grant project ERC-2022-STG, SOLIDCON (101078271).</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

ChatGPT conversation examples of 6 tasks

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

Evaluating Large Language Models in Summarizing Developer Chat Conversations: A Linguistic Perspective

<p>This is a replication package that includes:</p> <ul> <li>GoldenSet: contains the best summaries by participants for each conversation and the corresponding LLM generated summaries)</li> <li>LinguisticAnalysis_HumanGenerated: linguistic analysis such as speech tags, entities, etc. for summaries created by the Mturk participants (golden set)</li> <li>LinguisticAnalysis_LLMGenerated: linguistic analysis such as speech tags, entities etc. for summaries generated by the large language models</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo32/100

Conversational Agents in Software Engineering - Replication Package

<p>Replication Package for the research method conducted in the paper &quot;Conversational Agents in Software Engineering&quot;</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Impact of gigahertz and terahertz transport regimes on spin propagation and conversion in the antiferromagnet IrMn

<p>Data for the publication &quot;Impact of gigahertz and terahertz transport regimes on spin propagation and conversion in the antiferromagnet IrMn&quot; published in Applied Physics Letters. The following datasets are provided: GHz and THz raw data as function of the IrMn thickness, THz raw data spectra of the Pt|AF|F sample set, GHz and THz charge currents for forward (N|AF|F) and reversly-grown (F|AF|N) trilayer samples for N= Pt, W and Ta with varying AF thickness as well as the frequency-dependence of the charge current in the THz regime. Moreover, the temporal dynamics of the spin current densities from IrMn thicknesses 0nm, 3nm and 6nm are supplied.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

WAC Corpus - Wikipedia Abusive Conversations

<p>This repository contains conversations between Wikipedia editors, which are annotated in terms of various types of abuse, at the level of messages. This corpus is described in the following publication:</p> <p>N. C&eacute;cillon, V. Labatut, R. Dufour, and G. Linar&egrave;s, &ldquo;WAC: A Corpus of Wikipedia Conversations for Online Abuse Detection,&rdquo; in <em>12th Language Resources and Evaluation Conference</em>, 2020, pp. 1375&ndash;1383.&nbsp;⟨<a href="https://hal.archives-ouvertes.fr/hal-02497514">hal-02497514</a>⟩</p> <p>The repository also contains the figures shown in this article.</p> <p><strong>Sources. </strong>Our corpus aligns two existing corpora:</p> <ul> <li>Messages and conversation structures of&nbsp;<em>WikiConv</em> (<a href="https://github.com/conversationai/wikidetox/tree/master/wikiconv">https://github.com/conversationai/wikidetox/tree/master/wikiconv</a>)</li> <li>Manual annotations in toxicity of <em>Wikipedia Comment Corpus</em> (WCC -- <a href="https://doi.org/10.6084/m9.figshare.4054689">https://doi.org/10.6084/m9.figshare.4054689</a>)</li> </ul> <p><strong>Citation. </strong>If you use this dataset, please cite the above article.</p> <p><br><code>@InProceedings{Cecillon2020,</code><br><code>&nbsp; author &nbsp; &nbsp;= {C&eacute;cillon, No&eacute; and Labatut, Vincent and Dufour, Richard and Linar&egrave;s, Georges},</code><br><code>&nbsp; title &nbsp; &nbsp; = {{WAC}: A Corpus of {W}ikipedia Conversations for Online Abuse Detection},</code><br><code>&nbsp; booktitle = {12\textsuperscript{th} Language Resources and Evaluation Conference},</code><br><code>&nbsp; year &nbsp; &nbsp; &nbsp;= {2020},</code><br><code>&nbsp; pages &nbsp; &nbsp; = {1375-1383},</code><br><code>&nbsp; address &nbsp; = {Marseille, FR},</code><br><code>&nbsp; url&nbsp; &nbsp; &nbsp; &nbsp;= {http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.172.pdf},</code><br><code>}</code></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo32/100

Twitter Conversations about the COVID-19 Omicron Variant: A Large Scale Dataset of more than 500,000 Tweets

<p><strong>Please cite the following paper when using this dataset:</strong></p> <p>N. Thakur and C.Y. Han, &ldquo;An Exploratory Study of Tweets about the SARS-CoV-2 Omicron Variant: Insights from Sentiment Analysis, Language Interpretation, Source Tracking, Type Classification, and Embedded URL Detection,&rdquo; Journal of COVID, 2022, Volume 5, Issue 3, pp. 1026-1049</p> <p><strong>Abstract</strong></p> <p>This open-access dataset is one of the salient contributions of the above-mentioned paper. It presents a total of&nbsp;<strong>522,886</strong> Tweet IDs of the same number of <strong>Tweets about the SARS-CoV-2 Omicron Variant</strong> posted on Twitter since the first detected case of this variant on November 24, 2021. The dataset is compliant with the privacy policy, developer agreement, and guidelines for content redistribution of Twitter, as well as with the FAIR principles (Findability, Accessibility, Interoperability, and Reusability) principles for scientific data management.</p> <p><strong>Data Description</strong></p> <p>The Tweet IDs&nbsp;are presented in 7&nbsp;different .txt files based on the timelines of the associated tweets. The&nbsp;data collection followed a keyword-based&nbsp;approach and tweets comprising the &quot;omicron&quot; keyword were filtered, collected, and added to this dataset.&nbsp;The following is the description of these dataset files.</p> <ul> <li>Filename: TweetIDs_November.txt (No. of Tweet IDs: 16471, Date Range of the Tweet IDs: November 24, 2021 to November 30, 2021)</li> <li>Filename:&nbsp;TweetIDs_December.txt&nbsp;(No. of Tweet IDs:&nbsp;99288, Date Range of the Tweet IDs:&nbsp;December 1, 2021 to December 31, 2021)</li> <li>Filename:&nbsp;TweetIDs_January.txt&nbsp;(No. of Tweet IDs:&nbsp;92860, Date Range of the Tweet IDs:&nbsp;January 1, 2022 to January 31, 2022)</li> <li>Filename:&nbsp;TweetIDs_February.txt&nbsp;(No. of Tweet IDs:&nbsp;89080, Date Range of the Tweet IDs:&nbsp;February 1, 2022 to February 28, 2022)</li> <li>Filename:&nbsp;TweetIDs_March.txt&nbsp;(No. of Tweet IDs:&nbsp;97844, Date Range of the Tweet IDs:&nbsp;March 1, 2022 to March 31, 2022)</li> <li>Filename:&nbsp;TweetIDs_April.txt&nbsp;(No. of Tweet IDs:&nbsp;91587, Date Range of the Tweet IDs:&nbsp;April 1, 2022 to April 20, 2022)</li> <li>Filename:&nbsp;TweetIDs_May.txt&nbsp;(No. of Tweet IDs:&nbsp;35756, Date Range of the Tweet IDs:&nbsp;May 1, 2022 to May 12, 2022)</li> </ul> <p>In the above table, the last date for May is May 12 as it was the most recent date at the time of data collection and dataset upload. The dataset would be updated soon to incorporate more recent tweets.</p> <p>The dataset contains&nbsp;only Tweet IDs&nbsp;in compliance with the terms and conditions mentioned in the privacy policy, developer agreement, and guidelines for content redistribution of Twitter. The Tweet IDs&nbsp;need to be hydrated to be used.&nbsp;For hydrating this dataset the Hydrator application (<a href="https://github.com/DocNow/hydrator/releases">link to download</a>&nbsp;and a&nbsp;<a href="https://towardsdatascience.com/learn-how-to-easily-hydrate-tweets-a0f393ed340e#:~:text=Hydrating%20Tweets">step-by-step tutorial</a>&nbsp;on how to use Hydrator)&nbsp;may be used.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Everyday Conversation Videos Dataset

<p>This dataset contains short (all about 1 minute long) videos (at least 1080p)&nbsp;from four different actors (A1-A4) acting out a set of eight scenarios (S1-S8). The scenarios cover a range of speech styles in professional and personal contexts. Each simulates a conversational situation and are filmed from the perspective of he listener. The monologues for these scenarios were written by an experienced script writer. The script writer and all actors were recruited on fiverr.</p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record