Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.7.1
Dataset results
166 results for “open-source”
ThermoCyte: an inexpensive open-source temperature control system for in vitro live cell imaging
Open the record for dataset details and reuse information.
Cardio PyMEA: A user-friendly, open-source Python application for cardiomyocyte microelectrode array analysis
Open the record for dataset details and reuse information.
SOils DAta Harmonization database (SoDaH): an open-source synthesis of soil data from research networks
This SOils DAta Harmonization (SoDaH) database is designed to bring together soil carbon data from diverse research networks into a harmonized dataset that can be used for synthesis activities and model development. The research network sources for SoDaH span different biomes and climates, encompass multiple ecosystem types, and have collected data across a range of spatial, temporal, and depth gradients. The rich data sets assembled in SoDaH consist of observations from monitoring efforts and long-term ecological experiments. The SoDaH database also incorporates related environmental covariate data pertaining to climate, vegetation, soil chemistry, and soil physical properties. The data are harmonized and aggregated using open-source code that enables a scripted, repeatable approach for soil data synthesis.
Refactoring Test Smells: A Perspective from Open-Source Developers
<p>Presentation video for the <strong>5th Brazilian Symposium on Systematic and Automated Software Testing (SAST)</strong>, during the <strong>11th Brazilian Conference on Software: Practice and Theory (CBSoft 2020)</strong></p>
Quantitative analysis of subcellular distributions with an open-source, object-based tool
<p>The subcellular localization of objects, such as organelles, proteins, or other molecules, instructs cellular form and function. Understanding the underlying spatial relationships between objects through colocalization analysis of microscopy images is a fundamental approach used to inform biological mechanisms. We generated an automated and customizable computational tool, the SubcellularDistribution pipeline, to facilitate object-based image analysis from 3D fluorescence microcopy images. To test the utility of the SubcellularDistribution pipeline, we examined the subcellular distribution of mRNA relative to centrosomes within <i>Drosophila</i> embryos. Centrosomes are microtubule-organizing centers, and RNA enrichments at centrosomes are of emerging importance. Our open-source and freely available software detected RNA distributions comparably to commercially available image analysis software. The SubcellularDistribution pipeline is designed to guide the user through the complete process of preparing image analysis data for publication, from image segmentation and data processing to visualization.</p>
Open-source software collaboration network mining dataset
<p>The resulting dataset of the <a href="https://github.com/gotec/git2net">git2net </a>and <a href="https://github.com/wschuell/repo_tools">repo_tools </a>mining process for randomly selected large open-source repositories.</p>
Mandelbugs in Open-Source Software
<p>This dataset contains a list bugs from four open-source projects (the Linux kernel, the MySQL DBMS, the Apache HTTPD web server, and the Apache AXIS WS framework). The bugs have been classified into Mandelbugs, Bohrbugs, or Aging-Related Bugs, by analyzing the conditions that exercise the bug (i.e., the "fault trigger"). This classification is useful to get insights into bugs and failures that can occur in OSS projects, and to tune testing and fault-tolerance strategies according to the distribution of bug types in a project.</p> <p>The dataset contains an ARFF file for each subsystem of the four open-source projects. Each row of the ARFF file contains:</p> <p>- An IDs of the bug, which can be used to retrieve more information about the bug from the issue tracker of the project;</p> <p>- A string that represents the class of the bug (BOH = Bohrbug; NAM = Mandelbug; ARB = Aging-Related Bug; UNK = Unknown class);</p> <p>- A string that represents the sub-class of the bug (for Bohrbugs, the sub-class is not available; the subclasses for Mandelbugs are LAG, ENV, TIM, SEQ; the subclasses for Aging-Related bugs are MEM, STO, LOG, NUM, TOT).<br> </p>
Open-Source Cardiac MR Fingerprinting
<p>Magnetic Resonance (MR) raw data acquired with an open-source cardiac MR Fingerprinting (cMRF) sequence of a phantom at four different MR scanners. More details can be found here: https://github.com/PTB-MR/cMRF. The colormaps are taken from https://zenodo.org/records/11185704 because zenodo_get failed on trying to download this record in a jupyter notebook.</p> <p>Additionally cMRF data was acquired in three volunteers who were scanned at two different scanners. Cartesian and golden radial cine data was acquired to verify the anatomical features seen in the quantitative maps.</p>
Sample Stripped Pre-supernova Progenitors for open-source code CHIPS (Complete History for Interaction-Powered Supernovae)
<p>Inlists, mainly based on the test suite "example_make_pre_ccsn" in r12778, with slight amendments for removal of hydrogen (and helium, for Ic progenitors) envelope at core hydrogen (helium) exhausion.</p><p>For details: https://ui.adsabs.harvard.edu/abs/2023arXiv230810785T/abstract</p>
Artifact for "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software"
<p>This is the artifact for the paper "Inside Bug Report Templates: An Empirical Study on Bug Report Templates in Open-Source Software".</p> <p><strong>What the artifact does:</strong><br>1) a questionnaire that we used for our online survey (PDF);<br>2) the valid responses of our online survey (CSV).</p> <p>3) the code of preprocessing (.py).</p> <p>4) the dataset of preprocessing and labeling (CSV).</p>
Chinese Reference Population: open-source age-dependent computational phantoms of reference Chinese population
<p>The<strong> Chinese Reference Population (CRP)</strong> phantoms dataset encompass <strong>30 phantoms</strong> available in both voxel and NURBS formats, with age in 0, 1, 2, 3, 4, 5, 6, 8, 10, 12, 15, 18 years and adult male and female, as well as 4 pregnant women and fetus in early pregnancy, first trimester, second trimester and third trimester.</p> <ul> <li><strong>Voxelized phantoms</strong> are accessible in NII format :<strong> <em>"XXX.nii", which could be opened in AMIDE software.</em></strong></li> <li>Excel file<strong> </strong>containing<strong> organ masses and other descriptive information </strong>:<strong> <em>"CRP_descriptive_Info.xlsx"</em></strong></li> <li>In the application of F18−FDG dose calculation, <strong>organ absorbed doses per unit activity administered </strong>is provided in an Excel file :<strong> <em>"Application_F18-FDG.xlsx"</em></strong></li> </ul> <p>All data are stored on Zenodo and can be publicly accessed.</p>
GloHydroRes - a global dataset combining open-source hydropower plant and reservoir data
<div> <div> <div> <div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>Analyzing the impacts of drought and climate change on hydropower requires detailed data not only on hydropower attributes such as plant type, head, and installed capacity, but also on reservoir characteristics like area, depth, and volume. Current open-source hydropower datasets typically lack information on reservoirs, while reservoir datasets often omit hydropower details. GloHydroRes is a global dataset that integrates open-source hydropower and reservoir data, offering 29 attributes, including key information such as installed capacity, plant type, dam height, reservoir depth, area, volume, and river name. Overall, GloHydroRes provides data on 7,775 hydropower plants across 128 countries.</p> </div> </div> </div> </div> <div> <div> <div> </div> </div> </div> </div> </div>
An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects - Dataset
<p>This repository contains the dataset of the manuscript:</p> <p>"An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects"</p>
SoK: Taxonomy of Attacks on Open-Source Software Supply Chains - Visualization Tool Screenshots & Selected Papers
<p>This artifact complements the paper "SoK: Taxonomy of Attacks on Open-Source Software Supply Chains", submitted at IEEE S&P 2023.</p> <p>The papers selected during the Systematic Literature Review (SLR) are presented in the CSV file.</p> <p>This screenshots display the main features of the visualization tool that allows to explore the taxonomy of attacks on OSS supply chains, as well as the related safeguards and the selected references.</p>
TrainRuns.jl: an Open-Source Tool for Running Time Estimation - Supplement Data
<p>This additional data contains the initial data and the calculated results for comparing FBS and TrainRuns.jl.</p> <p><strong>File description</strong></p> <ul> <li><em>local.yaml</em>: input parameters for the local train</li> <li><em>freight.yaml</em>: input parameters for the freight train</li> <li><em>running_path.yaml</em>: input parameters for the path</li> <li><em>freight_FBS.csv</em>: export of calculation from FBS for the freight train</li> <li><em>freight_TrainRuns.csv</em>: export of calculation from TrainRun.jl converted in FBS units for the freight train</li> <li><em>freight_diff.csv</em>: the calculated difference between FBS.csv and TrainRuns.csv for the freight train</li> <li><em>local_FBS.csv</em>: export of calculation from FBS for the local train</li> <li><em>local_TrainRuns.csv</em>: export of calculation from TrainRun.jl converted in FBS units for the local train</li> <li><em>local_diff.csv</em>: the calculated difference between FBS.csv and TrainRuns.csv for the local train</li> <li><em>running_path.csv</em>: converted running_path.yaml for displaying</li> <li><em>comparison.tex</em>: LaTeX code for the graph in comparison.pdf</li> </ul> <p><strong>Sources</strong></p> <p>The calculations in FBS were done with the file 'Ostsachsen_V220.railml'. FBS needs a commercial license, which can be purchased. License for 'Ostsachsen_V220.railml' is Attribution-NonCommercial-NoDerivs 3.0 Unported (CC BY-NC-ND 3.0).<br> The file 'Ostsachsen_V220.railml' can be found at:<br> https://www.railml.org/en/user/exampledata.html (last accessed 2022-06-06 with login) -> "Real world railway examples from professional tools" -> "East Saxony railway network by FBS" -> "Ostsachsen_V220.railml"</p> <p>Other sources are mentioned in the files.</p>
Reproducible Evaluation of Open-Source Tools for Prostate Segmentation on Public Datasets
<p>Segmentation of the prostate and surrounding regions is important for a variety of clinical and research applications. Our goal is to evaluate the generalizability of publicly available state-of-the-art AI models on publicly available datasets. To compare the AI generated segmentations to the available manually annotated ground-truth, quantitative measures such as Dice Coefficient and Hausdorff distance, along with shape radiomics features, were analyzed. Our study also aims to show how cloud-based tools can be used to analyze, store, and visualize evaluation results.<strong> </strong></p> <p>Three open-source pre-trained AI prostate segmentation tools were evaluated against expert annotations, on three publicly available MRI prostate collections, available in NCI Imaging Data Commons[1]. Two pre-trained models originate from the nnU-Net framework[2], the last pre-trained model originates from Prostate158 paper[4]. ProstateX[5], QIN-Prostate-Repeatability[6] and PROSTATE-MRI-US-Biopsy[7]. Expert annotations of the the whole prostate gland, peripheral zone (PZ) and transition zone (TZ) of the prostate are available for ProstateX collection, whole prostate gland and PZ for QIN-Prostate-Repeatability collection, and whole prostate gland for PROSTATE-MRI-US-Biopsy collection.</p> <p>We rely on the DICOM standard to encode our segmentation and radiomics results. The DICOM standard aims to achieve interoperability and FAIR[10] principles. Encoding our results in DICOM representation allows us to leverage DICOM-reliant tools, such as Google Cloud Computing tools for storage,computation, analysis and visualization. Open-source DICOM-based visualization tools such as OHIF[8] viewer can also be used to look qualitatively at the AI and expert annotations and the referenced images.. DICOM Segmentation objects are used to encode the AI models predictions, using dcmqi[11], DICOM Structured Reports on the other hand are used to encode radiomics features[3] extracted from the AI and expert annotations, using dcmqi and highdicom[12]. </p> <p>This dataset is organized in three parts: </p> <p>AI_SEGMENTATIONS_DICOM.zip, AI_STRUCTURED_REPORTS_DICOM.zip and EXPERT_SRUCTURED_REPORTS_DICOM..zip. All zip files contain DICOM objects only, sorted based on DICOM attributes, following this pattern:</p> <p>PatientID/<br> └───Modality-%StudyInstanceUID/<br> └───%SeriesInstanceUID-%SeriesDescription.dcm.</p> <p>AI_SEGMENTATIONS_DICOM.zip contains all the pre-trained AI models evaluated segmentation results, encoded as DICOM Segmentation objects. AI_STRUCTURED_REPORTS_DICOM..zip contains firstorder and shape radiomics features extracted for the AI segmentation results, such as Segmentation Volume, encoded as DICOM Structured Reports. EXPERT_SRUCTURED_REPORTS_DICOM.zip contains firstorder and shape radiomics features extracted for the expert annotations (for ProstateX, QIN-Prostate-Repeatability and PROSTATE-MRI-US-Biopsy collections) stored a DICOM Structured Reports objects.</p> <p>Code repository containing evaluation cloud-based notebooks and results/metadata .csv tables is available here:<br><a href="https://github.com/ImagingDataCommons/idc-prostate-mri-analysis">https://github.com/ImagingDataCommons/idc-prostate-mri-analysis</a></p> <h2>Additional Notes</h2> <p><strong> </strong>This project has been funded in whole or in part with Federal funds from the NCI, NIH, under task order no. HHSN26110071 under contract no. HHSN261201500003l.<br>https://portal.imaging.datacommons.cancer.gov/</p> <p>nnU-Net: <a href="https://github.com/MIC-DKFZ/nnUNet">https://github.com/MIC-DKFZ/nnUNet</a> <br>Prostate158: <a href="https://github.com/Project-MONAI/model-zoo/tree/dev/models/prostate_mri_anatomy">https://github.com/Project-MONAI/model-zoo/tree/dev/models/prostate_mri_anatomy</a><br>Pyradiomics: <a href="https://github.com/AIM-Harvard/pyradiomics">https://github.com/AIM-Harvard/pyradiomics</a> <br>Highdicom: <a href="https://github.com/herrmannlab/highdicom">https://github.com/herrmannlab/highdicom</a> <br>DCMQI: <a href="https://github.com/QIICR/dcmqi">https://github.com/QIICR/dcmqi</a> <br>Github repo: <a href="https://github.com/ImagingDataCommons/idc-prostate-mri-analysis">https://github.com/ImagingDataCommons/idc-prostate-mri-analysis</a></p> <h2>Related information</h2> <p>ProstateX - <a href="https://doi.org/10.7937/K9TCIA.2017.MURS5CL">https://doi.org/10.7937/K9TCIA.2017.MURS5CL </a><br><br>QIN-Prostate-Repeatability - <a href="https://doi.org/10.7937/K9/TCIA.2018.MR1CKGND">https://doi.org/10.7937/K9/TCIA.2018.MR1CKGND</a><br><br>PROSTATE-MRI-US-BIOPSY - <a href="https://doi.org/10.7937/TCIA.2020.A61IOC1A">https://doi.org/10.7937/TCIA.2020.A61IOC1A</a></p> <h2>References </h2> <p>[1] Fedorov A, Longabaugh WJ, Pot D, Clunie DA, Pieper S, Aerts HJ, Homeyer A, Lewis R, Akbarzadeh A, Bontempi D, Clifford W. NCI imaging data commons. Cancer research. 2021 Aug 8;81(16):4188.</p> <p>[2] Isensee F, Jaeger PF, Kohl SA, Petersen J, Maier-Hein KH. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods. 2021 Feb;18(2):203-11.</p> <p>[3] Van Griethuysen JJ, Fedorov A, Parmar C, Hosny A, Aucoin N, Narayan V, Beets-Tan RG, Fillion-Robin JC, Pieper S, Aerts HJ. Computational radiomics system to decode the radiographic phenotype. Cancer research. 2017 Nov 1;77(21):e104-7.</p> <p>[4] Adams, Lisa C., Marcus R. Makowski, Günther Engel, Maximilian Rattunde, Felix Busch, Patrick Asbach, Stefan M. Niehues, et al. 2022. “Prostate158 - An Expert-Annotated 3T MRI Dataset and Algorithm for Prostate Cancer Detection.” Computers in Biology and Medicine 148 (September): 105817.</p> <p>[5] Natarajan, S., Priester, A., Margolis, D., Huang, J., & Marks, L. (2020). Prostate MRI and Ultrasound With Pathology and Coordinates of Tracked Biopsy (Prostate-MRI-US-Biopsy) (version 2) [Data set]. The Cancer Imaging Archive. DOI: 10.7937/TCIA.2020.A61IOC1A</p> <p>[6] Fedorov, A; Schwier, M; Clunie, D; Herz, C; Pieper, S; Kikinis, R; Tempany, C; Fennessy, F. (2018). Data From QIN-PROSTATE-Repeatability. The Cancer Imaging Archive. DOI: 10.7937/K9/TCIA.2018.MR1CKGND</p> <p>[7] Natarajan, S., Priester, A., Margolis, D., Huang, J., & Marks, L. (2020). Prostate MRI and Ultrasound With Pathology and Coordinates of Tracked Biopsy (Prostate-MRI-US-Biopsy) (version 2) [Data set]. The Cancer Imaging Archive. DOI: 10.7937/TCIA.2020.A61IOC1A</p> <p>[8] Open Health Imaging Foundation Viewer: An Extensible Open-Source Framework for Building Web-Based Imaging Applications to Support Cancer Research. Erik Ziegler, Trinity Urban, Danny Brown, James Petts, Steve D. Pieper, Rob Lewis, Chris Hafey, and Gordon J. Harris</p> <p>[9] Clark K, Vendt B, Smith K, Freymann J, Kirby J, Koppel P, Moore S, Phillips S, Maffitt D, Pringle M, Tarbox L. The Cancer Imaging Archive (TCIA): maintaining and operating a public information repository. Journal of digital imaging. 2013 Dec;26(6):1045-57.</p> <p>[10] Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, Blomberg N, Boiten JW, da Silva Santos LB, Bourne PE, Bouwman J. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data. 2016 Mar 15;3(1):1-9.</p> <p>[11] Herz C, Fillion-Robin JC, Onken M, Riesmeier J, Lasso A, Pinter C, Fichtinger G, Pieper S, Clunie D, Kikinis R, Fedorov A. DCMQI: an open source library for standardized communication of quantitative image analysis results using DICOM. Cancer research. 2017 Nov 1;77(21):e87-90.</p> <p>[12] Bridge CP, Gorman C, Pieper S, Doyle SW, Lennerz JK, Kalpathy-Cramer J, Clunie DA, Fedorov AY, Herrmann MD. Highdicom: A python library for standardized encoding of image annotations and machine learning model outputs in pathology and radiology. Journal of Digital Imaging. 2022 Aug 22:1-9.</p> <p> </p>
Figure 1 in Use of open-source object detection algorithms to detect Palmer amaranth (Amoronthus polmeri) in soybean
Figure 1. Examples of images collected for soybean in the VE-VC (A) and R2 (B) growth stages.
eELib: Open-Source Model Library for Prosumer Power Systems and Energy Management Strategies (data)
<p>Dataset and results used for the simulations in following publication:</p> <p>Carsten Wegkamp, Henrik Wagner, Eike Niehs, Julien Essers, Marcel Lüdecke, Mattias Hadlak, Bernd Engel:<br>"<strong>eELib: Open-Source Model Library for Prosumer Power Systems and Energy Management Strategies</strong>",<br>Open Source Modelling and Simulation of Energy Systems (OSMSES) 2024, Vienna, Austria, 2024</p> <p> </p> <p>This contains the input (scenario) files for the building & grid scenario and the results of the two simulations.<br>It uses the elenia Energy Library (eELib) with release version 1.0.0: https://gitlab.com/elenia1/elenia-energy-library</p>
Why is my community reacting like this? Understanding reactions in open-source communities
<p>In 2016, GitHub introduced the Reactions feature to facilitate the expression of sentiments and reduce noise in communications on its platform. Recent studies indicated that developers has been adopting the feature and was observed a reduction of noise on conversations inside the platform. However, the patterns of usage and profiles of users expressing these reactions in inside their communities remain underexplored. Identifying these patterns may help maintainers to better understand members' behaviors in their communities, and researchers to build supporting tools focused on users' reactions. This paper presents an initial study to (i) understand these interactions on open-source software communities, (ii) identify types of resources that receive the most reactions, (iii) analyzing seasonal factors influencing usage, and (iv) correlating the provided reactions with the roles of developers within the community. Preliminary results indicate that users primarily react to comments in Issues, with notable periods of heightened activity. Additionally, significant differences were observed between the reactions of maintainers and other members of the community.</p>
An Economical Open-Source Lagrangian Drifter Design to Measure Deep Currents in Lakes
<p>An economical, open-source Lagrangian drifter designed to collect current data on lakes<200km2 was evaluated against existing designs. The new design was tested in deep inland lakes in the Finger Lakes region of New York, USA and is effective at tracking deep currents. The ease and low-cost of fabrication and launch/recovery should facilitate use of this design by less-advantaged communities & researchers.</p> <p>This project includes data and code for preparation of graphs and charts to illustrate Lagrangian drifter experiments in Seneca Lake and Keuka Lake, New York, USA.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.