Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,916
datasets available to search
ShareScore release 0.7.1
Dataset results
1,916 results for “software,”
FIGURE 7 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 7. Screenshot of the GUI-based version of ABGD, a program that delimits species by detecting the barcoding gap from pairwise single-locus sequence distances (Puillandre et al. 2012). For this tool, the original ABGD code written in C was wrapped with a Python GUI and compiled as standalone executable. The different output files produced by ABGD (text and graphs) can be selected and pre-viewed within the GUI.
FIGURE 6 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 6. Screenshot of the GUI-based version of PTP, a program that delimits species from non-ultrametric trees. The original Python code of PTP was written by Zhang et al. (2013); iTaxoTools adds the GUI, as well as functionality to export species partition in the SPART format (Miralles et al. 2021).
FIGURE 4 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 4. Screenshot of spartmapper, a tool that plots distribution records from geographical coordinates on a map and categorizes the records based on a species partition provided as SPART file (Miralles et al. 2021). The program allows live view and produces a kml file to visualize the records in Google Earth or Google Maps.
FIGURE 2 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 2. Main launcher window of iTaxoTools 0.1 with the option to start various species delimitation tools (additional tools can be started from the other tabs).
FIGURE 3 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 3. Screenshot of one of the newly programmed quick conversion tools, dnaconvert, which implements numerous autocorrect options to avoid sequence output files generating errors in downstream programs. dnaconvert also supports tabdelimited table input and its conversion to common sequence formats such as FASTA, NEXUS, or PHYLIP, to facilitate storage and management of sequences and sequence metadata in spreadsheet editors such as Microsoft Excel.
FIGURE 1 in iTaxoTools 0.1: Kickstarting a specimen-based software toolkit for taxonomists
FIGURE 1. Overview of the various tools implemented in iTaxoTools, and their scope. In the present version a focus is on molecular data analysis, but more functionalities to analyze and visualize morphological and geographic data will be implemented in the near future, while data integration remains the main focus for long-term implementation.
Supplementary data for: Collaborative Program Comprehension via Software Visualization in Extended Reality
<p>Supplementary videos and images for: Collaborative Program Comprehension via Software Visualization in Extended Reality</p> <p>Videos are additionally hosted on: <a href="https://www.youtube.com/channel/UCijDvGaoZqH0LiFW3DSNH1w">https://www.youtube.com/channel/UCijDvGaoZqH0LiFW3DSNH1w</a></p>
The Clarity Software Documentation Dataset
<p>This repository holds the Clarity Dataset which is a companion to the SANER'22 entitled "An Empirical Investigation into the Use of Image Captioning for Automated Software Documentation". The dataset consists of 45,998 captions 10,204 GUI screenshots and xml metadata files (akin to the "html" for stipulating GUIs) of Android applications. The NL captions were obtained from human labelers, underwent several quality control mechanisms, and contain both high- (screen-level) and low-(component) level descriptions of screen functionality. This dataset is meant as a new source of data to augment techniques for software documentation that can take advantage of the rich pixel-based information contained within screenshots.</p>
Extra Testing Data for paper "OC_Finder: A deep learning-based software for osteoclast segmentation, classification, and counting"
<pre>Here we have 9 datasets we used to validate OC_Finder's performance on various imaging settings. The 9 datasets are inside the folder named "9 datasets for validation experiment". Each dataset is composed of image files and csv files for the coordination of osteoclasts and non-osteoclasts that were manually labelled by human examiner. csv files ending "_posi" has coordination of osteoclasts and "_nega" has coordination of non-osteoclasts. Images in dataset #4, #5, #6, #7, #8, and #9 were resized so the scale of the images matched to the OC_Finder's training dataset. Images in original size before resizing are also provided in "Original images before resizing". Detailed capture setting and resizing information of images in each dataset can be found in "capture setting.xlsx". The number of images in each dataset are as following: #1: 18 #2: 18 #3: 18 #4: 36 #5: 36 #6: 36 #7: 16 #8: 16 #9: 16</pre>
Uncovering structural ensembles from single particle cryo-EM data using cryoDRGN | Software, datasets, and results
<p>Software, datasets, and results referenced in "Uncovering structural ensembles from single particle cryo-EM data using cryoDRGN"</p>
Supporting data and software for: Low-temperature open-air synthesis of PVP-coated NaYF4:Yb,Er,Mn upconversion nanoparticles with strong red emission
<p>Upconversion nanoparticles (UCNPs) have unique photonic properties that make them ideally suited for many applications. They are excited by low-energy near-infrared photons and emit at higher energy (typically visible) wavebands. However, synthesis of UCNPs requires either high pressure reaction chambers or inert atmospheres. Combined with the requirements for high-temperatures (200 to 400 °C) and long reaction times (e.g. up to 24 hours), these place barriers to entry for UCNP research, in terms of both financial barriers and knowledge/"know how". These constraints may also limit the scale of UCNP production for end-user applications.</p> <p>We adapted and further developed a method for producing UCNPs with simple laboratory equipment, i.e. a hot-plate and beakers. No pressure vessel or inert atmosphere is required. The UCNPs produced have a<span> polyvinylpyrrolidone (PVP) polymer coating, with strong red emission due to Mn<sup>2+</sup> co-doping within the UCNP crystal lattice. It was found that UCNPs of composition NaYF<sub>4</sub>:Yb,Er,Mn (Yb = 20 mol %, Er = 2 mol%, Mn = 35 mol%) maximised the red emission whilst also minimising the diameter of the UCNPs to </span> 36 ± 15 nm. These combination of optical and physical properties should make these UCNPs ideal for further development and exploitation, particularly for biological applications where red emission can penetrate over a centimetre of tissue.</p> <p>This dataset and software accompanies the manuscript <em>'Low-temperature open-air synthesis of PVP-coated NaYF<sub>4:</sub>Yb,Er,Mn upconversion nanoparticles with strong red emission</em>', which was published in Royal Society Open Science on 19th January 2022. https://doi.org/10.1098/rsos.211508</p>
Gamification in Software Engineering: The Mediating Role of Developer Engagement and Job Satisfaction
<p>Replication package with covariance matrices (instead of original dataset) and R script.</p>
Mining Fork-Including Software Development Traces
<p>This dataset relates to the paper: Mining Fork-Including Development Traces (abstract below)<br> Authors: Iris Reinhartz-Berger and Amir Tomer<br> Starting point: readme.txt</p> <p>Open-source software development is a common practice that encourages collaborative development and reuse across projects. Forking is a way to make a copy of an existing project and explore it for different purposes. Two types of forks are commonly mentioned in the literature: <em>contributing forks</em> which continue the development lines of the forked projects and aim at merging the contribution back to the forked projects; and <em>independently developed forks</em> which open new lines of development deviating from the forked projects. In this study, we aim to explore characteristics of fork-involving software development traces. Analyzing 880 Java projects and their related action and observation events, with process mining and statistical techniques, we found that the occurrence of certain event types may predict the fork type, while the creation of certain fork types increase the involvement of users in the forked projects.</p>
Data and Software for "Complex Dynamics in a Synchronized Cell-Free Genetic Clock"
<p>Contains raw data, analysis scripts, simulation scripts, device operation software for the publication: "Complex Dynamics in a Synchronized Cell-Free Genetic Clock".</p>
Reports about the software behaviour delivered within the GeoBIM benchmark 2019 - Task 3
<p>Answers delivered through the online forms, about the tests performed by participants within the Task 3 - support for CityGML of the GeoBIM benchmark 2019, funded as a Scientific Initiative 2019 by the International Society of Photogrammetry and Remote Sensing (ISPRS) and co-funded by the European association for Spatial Data Research (EuroSDR).</p> <p>Full details and additional resources about the project are available in the project website: https://3d.bk.tudelft.nl/projects/geobim-benchmark/</p> <p>The dataset results from the collaboration of the authors with all the participants to the benchmark, listed at https://3d.bk.tudelft.nl/projects/geobim-benchmark/participants.html</p> <p>It is composed by 2 files:</p> <p>- the answers to the delivered online forms, organised in excel sheet;</p> <p>- the same answers organised in a more human-readable PDF, with reliable images and links.</p>
Supplementary Data & Software for "Balancing the marine sulfur cycle", Johnson & Adkins (submitted)
<p>These files are the supplementary data and software for the manuscript "Balancing the marine sulfur cycle" by Johnson & Adkins (submitted).</p> <p>The MATLAB workspace file "DeepSeaSO4_ClusterAnalysisData.mat" contains organized data structures with DSDP, ODP, and IODP site metadata, porosity data, pore water sulfate concentration data, methane concentration data, total organic carbon wt% data, and calcium carbonate wt% data + their corresponding depths. These data served as the basis for the cluster analysis performed by the included MATLAB script "SO4_conc_cluster_analysis.m" in the study.</p> <p>The "DatasetS1.xlsx" spreadsheet file contains all compiled pyrite sulfur isotopic data, pyrite abundance data, and total sulfur abundance data included within the literature compilation discussed in the manuscript. Pyrite sulfur isotopic data references are listed and are cited in the main text. DOI URLs for the total sulfur abundance data from the PANGAEA repository are also included in the spreadsheet. These data have additionally been saved into the MATLAB workspace files "Pyrite_d34S_Comp.mat" and "Pyrite Abundance_Comp.mat" for import into MATLAB by other scripts.</p> <p>The MATLAB workspace file "POROSITY.mat" includes vectors of latitude (in degrees), longitude (in degrees), water depth (in mbsl), sediment depth (in mbsf), porosity (in vol%), and corresponding PANGAEA DOI URLs along with the data tables from which this information was extracted. This file is called by "PorosityPlotter.m" to extract initial porosity values, compile coring methods, and make plots of initial porosity as a function of water depth for each coring method.</p> <p>The MATLAB script "hypsometry.m" loads the global topographic data included in "topography.mat", calculates associated surface areas across different ocean depth intervals, and determines average porosity and sedimentary pyrite parameters across those intervals in tandem with "S_Iso_Mass.m". "S_Iso_Mass.m" uses the averages to calculate the steady-state sulfur isotopic composition of seawater.</p> <p>The script "phi_waterdepth_analysis.m" loads the cluster analysis data and uses the corresponding porosity data to calculate linear regressions between initial porosity and water depth for use by other scripts; the workspace files "phi_waterdepth_fit.mat" and "phi_waterdepth_trimmed_fit.mat" are saved versions of the regressions used in the manuscript and called by "hypsometry.m" for calculating averages. </p> <p>The additional included files "PyriteSAbundanceCompilationSources.pdf" and "PorosityCompilationSources.pdf" files include full citations to the data included in the pyrite S abundance and porosity compilations (respectively).</p>
Reports about the software behaviour delivered within the GeoBIM benchmark 2019 - Task 1
<p><strong>Answers delivered through the online forms</strong>, about the tests performed by participants within the <strong>Task 1 - support for IFC</strong> of the <strong>GeoBIM benchmark 2019</strong>, funded as a Scientific Initiative 2019 by the International Society of Photogrammetry and Remote Sensing (ISPRS) and co-funded by the European association for Spatial Data Research (EuroSDR).</p> <p>Full details and additional resources about the project are available in the project website: https://3d.bk.tudelft.nl/projects/geobim-benchmark/</p> <p>The dataset results from the collaboration of the authors with all the participants to the benchmark, listed at https://3d.bk.tudelft.nl/projects/geobim-benchmark/participants.html</p> <p>It is composed by 4 files:</p> <p>- the answers to the delivered online forms, organised in excel sheet;</p> <p>- the same answers organised in a more human-readable PDF, with reliable images;</p> <p>- the same answers organised in a more human-readable PDF, with low-resolution images and working links;</p> <p>- the answers regarding the IFCgeometries dataset in IFC 2x3 format, in excel;</p> <p>- the answers regarding the IFCgeometries dataset in IFC 4 format, in excel.</p>
Dataset of the paper "An Empirical Characterization of Software Bugs in Open-Source Cyber-Physical Systems"
<p><br> #Dataset Package for the paper "An Empirical Characterization of Software Bugs in Open-Source Cyber-Physical Systems"</p> <p><br> Description of the content:</p> <p><br> 1) "1_RQ-CPS-bugs-Taxonomy" folder contains all the main experimental data concerning the issues sampled and analyzed from all the Projects considered in the study,<br> including row-data on the taxonomy validtion steps.<br> <br> - Under "the sub-folder "1_Taxonomy-Raw-data" are reported the row-data concerning the taxonomy validtion steps </p> <p><br> 2) "2_Scripts" contains all scripts used to generate the issue data and sampled issue raw-data in the previous folders: </p> <p><br> - "setup.md" file in the folder describes how to set=up and run the script used for collecting and sampling the issues for the validation steps:<br> <br> - runJSONtoCSV.sh<br> - JSONtoCSV.py<br> - generateListOfAllSamples.py<br> - generateAllSamples.r<br> <br> Under "the sub-folder "1_Scripts/1_Data_Collection":<br> <br> <br> 3) "3_Final Taxonomy" folder contains the final Table representation (also reported in the previous folder) and main figures of the CPSs Bugs Taxonomy.</p>
DASP: A Framework for Driving the Adoption of Software Security Practices
<p>Online appendix for the publication:</p> <p>DASP: A Framework for Driving the Adoption of Software Security Practices.</p>
Tutorial videos of MEFisTo software and other UVic ECE 340 course content
<p>An assemblage of tutorial videos and content that document the usage of MEFisTo software and detail example questions from the <em>Fundamentals of Applied Electromagnetism</em> textbook by Ulaby <strong>2014</strong>. These materials were used for the ECE 340 course at the University of Victoria in Winter 2021.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.