Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “UML”
UML class diagrams for layout quality checking
<p>We offer a dataset of labelled UML class diagrams. In this dataset, we supply for every diagrams the following information: 1) a manually established ground truth of the quality of the layout, 2) a value for the layout-quality of the diagram as predicted by our classifier, and 3) the values of key features of the layout of the diagram. These features are extracted automatically via image processing. This dataset can be used for replication of our study and for others to build on and improve on our work.</p>
Teaching UML Models with FLOSS Projects: A study carried out during the period of social isolation imposed by the COVID-19 pandemic
<p>Software Engineering Education must include professional practice experiences so that students are able to specify, design, implement, maintain and evaluate computer systems, using appropriate theories, practices and tools. However, teaching Software Engineering (SE) principles, concepts and practices and relating them to real-world scenarios are challenging tasks. The adoption of Free/Libre/Open-Source Software (FLOSS) projects can assist in facing these challenges. This paper presents an experience report on the pedagogical use of FLOSS projects in a SE course, in the context of social isolation imposed by the COVID-19 pandemic. The course was fully virtualized, with supporting tools such as Google Classroom, Google Meet, Github, Padlet and Trello, and modeled after an instructional design based on Bloom’s taxonomy of cognitive learning objectives. FLOSS projects were used in the context of software modeling activities with UML class, sequence and activity diagrams. An online survey with twenty-six (26) students and an interview with the instructor were conducted to collect their perceptions about such experience. Results showed that most students (i) agreed that the instructional methods addressed the knowledge, comprehension and application levels in a satisfactory way; (ii) were satisfied in working with real software projects developed by others; and (iii) liked the developed activities and the tools used in the classes. The instructor and students (iv) missed face-to-face activities; and (v) agreed that the instructional design could be used in future face-to-face SE classes.</p>
Open Science materials of the paper "Automatically Recognizing the Semantic Elements from UML Class Diagram Images"
<p>This submission contains such files:</p> <p>1. "questionnaire.docx": The questionnaire for the survey. The file contains all the questions and answers.<br> 2. "raw data collected from participants.xlsx": The raw data collected from participants. Each row in the file represents a participant's answers to all questions, including the date, source, IP, and answers.<br> 3. "raw data collected from open-source community.xlsx": The raw data collected from open-source community (the UML diagram usage). It contains the repositories and GitHub URLs, the UML diagrams and the corresponding links, and some statistics about the UML diagram usage.<br> 4. "an implementation of ReSECDI.zip", "utility source code.zip", "utility compiled JAR.zip", and ".m2.zip": An implementation of ReSECDI, and its dependencies. The implementation is in Java, and it requires JDK11 or higher. It depends on a project named "utility", in addition to other dependencies. The source code of "utility" is provided in "utility source code.zip", the compiled JAR file is in "utility compiled JAR.zip", and the maven dependency files are provided in ".m2.zip". The ways to add the "utility" to the implementation's dependencies are explained in the "readme.txt".<br> 5. "instructions for how to use the artifacts.docx": The instructions for how to use the implementation of ReSECDI. It mainly explains the key components of the implementation, and how to set the parameters.<br> 6. "diagrams used for its evaluation.zip": The diagrams used for the evaluation. There are 50 diagrams collected from the open-source community. Each diagram's name represents its belonging repository.<br> 7. "raw data collected during the evaluation.xlsx": The raw data collected during the evaluation. It contains the statistics of the classes and relationships for each diagram, and the recognition results.<br> 8. "Manuscript.pdf": The manuscript explaining our approach.<br> 9. "readme.txt": The readme file explaining details about each file.</p>
Modeling and Configuring UML-based Software Product Lines with SMartyModeling
<p>Vídeo de Apresentação para o Fórum de Pós-Graduação na IV Escola Regional de Engenharia de Software (ERES) - Sessão Técnica 9. Online Streaming, 11 a 13 de Novembro de 2020.</p>
UMLS Heading Sequences in Spanish
<p>UMLS Heading Sequences in Spanish used to compute <a href="https://zenodo.org/record/6647060">Word embeddings for the Spanish clinical language</a></p>
UML Sourcing Domain Predictions
<p><strong>WHAT: </strong>This data accompanies Glick, H.B., Ament, J.M., Dallinga, J.S., Torres-Batlló, J., Verma, M., Clinton, N., and Wilcox, A. (2023). Model-based prediction and ascription of deforestation risk within commodity sourcing domains: Improving traceability in the palm oil supply chain.<br> <br> The data represents a collection of geographically referenced raster images (GeoTIFF) capturing the predicted sourcing domains of 1,570 palm oil processing facilities (mills) in Indonesia and Malaysia. There is one image per facility, delivered in WGS84 (EPSG:4326) with a nominal equatorial spatial resolution of 250 m. With respect to the models discussed in Glick et al (2023), each image file captures the results of our most accurate model, which was an ensemble of predictions from MaxEnt, random forest, and gradient boosted regression tree-based machine learning models, trained on passive geolocational traceability data (n = 3,355,437 cellular pings). The individual pixel values in each image are the probability of containing aggregated cellular ping data from individuals that have been spatio-temporally linked to the given processing facility. Functionally, these values represent the probability that a given pixelated location is part of a facility's sourcing domain, where a sourcing domain encompasses both harvesting locations and the intermediary transportation and sub-processing space. Please refer to the parent manuscript for details.</p> <p>Please note that our predictions were derived from a processing chain that used a modified World Mollweide projected coordinate system (essentially ESRI:54009), where the central meridian (longitude of origin) was set to 109.5 degrees. The images are delivered here in WGS84 (EPSG:4326). Users can access the data in its original coordinate reference system using a Google Earth Engine ImageCollection: ee.ImageCollection('projects/ul-gs-d-901791-09-prj/assets/users/hglick/Glick_et_al_2023/Ensemble).</p> <p><strong>WHEN:</strong> Passive geolocational training data was gathered in 2020 and 2021. Palm oil processing facilities were derived from the Universal Mill List in November 2021.</p> <p><strong>WHERE:</strong> All images contain predictions for palm oil processing facilities located in Indonesia and Malaysia, with predictions made to a maximum Euclidean distance of 100 km from each facility.</p> <p><strong>WHY:</strong> Palm oil accounts for approximately 50% of global vegetable oil production, and trends in consumption have driven large-scale expansion of oil palm (<em>Elaeis guineensis</em>) plantations in Southeast Asia. This expansion has led to deforestation and other socio-environmental concerns that challenge consumer goods companies to meet no deforestation and sustainability commitments. In support of these commitments and supply chain traceability, we seek to improve on the current industry standard sourcing model for ascribing social and environmental risks to particular actors. Among other uses, this data can support the ascription or allocation of deforestation, carbon loss, and biodiversity risk to relevant actors, permitting targeted outreach, contract negotiation, and mitigation of large-scale resource degradation.</p> <p><strong>HOW:</strong> Passive geolocational training data was gathered by Orbital Insights. All modeling was conducted in Python; Google Earth Engine served as the primary distributed computing platform. Details are presented in Glick et al (2023).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.