Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
62
datasets available to search
ShareScore release 0.9.0
Dataset results
62 results for “ACM”
GRIME AI Water Segmentation Model for the USGS Monitoring Site at Pecos River near Acme, NM, 2022-2024
Ground-based observations from fixed-mount cameras have the potential to fill an important role in environmental sensing, including direct measurement of water levels and qualitative observation of ecohydrological research sites. All of this is theoretically possible for anyone who can install a trail camera. Easy acquisition of ground-based imagery has resulted in millions of environmental images stored, some of which are public data, and many of which contain information that has yet to be used for scientific purposes. The goal of this project was to develop and document key image processing and machine learning workflows, primarily related to semi-automated image labeling, to increase the use and value of existing and emerging archives of imagery that is relevant to ecohydrological processes. This data package includes imagery, annotation files, water segmentation model and model performance plots, and model test results (overlay images and masks) USGS Monitoring Site at Pecos River near Acme, NM, 2022-2024. All imagery was acquired from the USGS Hydrologic Imagery Visualization and Information System (HIVIS; see https://apps.usgs.gov/hivis/camera/NM_Pecos_River_near_Acme for this specific data set) and/or the National Imagery Management System (NIMS) API. Water segmentation models were created by tuning the open-source Segment Anything Model 2 (SAM2, https://github.com/facebookresearch/sam2) using images that were annotated by team members on this project. The models were trained on the "water" annotations, but annotation files may include additional labels, such as "snow", "sky", and "unknown". Image annotation was done in Computer Vision Annotation Tool (CVAT) and exported in COCO format (.json). All model training and testing was completed in GaugeCam Remote Image Manager Educational Artificial Intelligence (GRIME AI, https://gaugecam.org/) software (Version: Beta 16). Model performance plots were automatically generated during this process. This project was
CE-MS data for Autonomous CE Mass-Spectra Examination (ACME) for the Ocean Worlds Life Surveyor (OWLS)
<p>These folders contain the original data used to develop the ACME software [1].</p> <p>The Golden and Silver dataset come from simulations. They underrepresent the complexity in the CE-MS observations but provide additional data with known peak locations and peak properties. For more information see [1]</p> <p>The Dev-, Train-, and Test-set contain CE-MS [2] observations of Mix25 (a standard set of 25 organic compounds relevant to astrobiology) and labels for peak locations from subject matter experts. </p> <p>The ACME software is available at: <br> https://github.com/JPLMLIA/OWLS-Autonomy </p> <p> </p> <p>When using the data please cite this dataset [3] and the two papers below. </p> <p>For further questions please reach out to:<br> Steffen Mauceri, Steffen.Mauceri@jpl.nasa.gov</p> <p> </p> <p>References:<br> [1] Mauceri, S., Lee, J., Wronkiewicz, M., et.al. (2022). Autonomous CE Mass-Spectra Examination (ACME) for the Ocean Worlds Life Surveyor (OWLS). (submitted) Earth and Space Science</p> <p>[2] Mora et al., F.(2021). Detection of biosignatures by capillary electrophoresis and mass spectrometry in the presence of salts relevant to missions to ocean worlds (submitted). Astrobiology.</p> <p>[3] 10.5281/zenodo.5849873</p> <p><br> © 2022. California Institute of Technology. Government sponsorship acknowledged</p>
Enhanced 3D radiative transfer ACM-RT calculations and output for Cole et al., 2022
<p>For the Earth Cloud, Aerosol, Radiation Explorer (EarthCARE) satellite mission there are a number of algorithms used to process the observations. One of these algorithms, called ACM-RT, is designed to use retrieved geophysical properties to perform forward radiative transfer calculations using 1D and 3D solar and thermal radiative transfer models. The ACM-RT algorithm is documented in an Atmospheric Measurements and Techniques (AMT) article. </p> <p>To illustrate outputs from ACM-RT, including the benefits of 3D radiative transfer, “enhanced” radiative transfer calculations were performed, relative to calculations that would be performed operationally during the EarthCARE mission. In particular, the 3D Monte Carlo radiative transfer calculations used an increased number of samples to reduce the Monte Carlo uncertainty and calculations were performed for more of the input data.</p> <p>The relevant publication for these calculations is:</p> <p>Cole, J. N. S., H. W. Barker, Z. Qu, N. Villefranque and M. W. Shephard : Broadband Radiative Quantities for the EarthCARE Mission: The ACM-COM and ACM-RT Products. Submitted to AMT, November 2022.</p>
ACM-Digital Library (ACM)
<p>ACM-Digital Library (ACM): a subset of the ACM Digital Library with 24, 897 documents containing articles related to Computer Science. We considered only the first level of the taxonomy adopted by ACM, where each document is assigned to one of 11 classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>
Artifacts supplementing the ACM DTRAP 2020 article "Will You Trust This TLS Certificate? Perceptions of People Working in IT (extended version)"
<p>These research artifacts supplement the following two publications:</p> <ul> <li>Will You Trust This TLS Certificate? Perceptions of People Working in IT [ACSAC 2019], DOI 10.1145/3359789.3359800, more details at https://crocs.fi.muni.cz/public/papers/acsac2019</li> <li>Will You Trust This TLS Certificate? Perceptions of People Working in IT (extended version) [ACM DTRAP 2020], DOI 10.1145/3419472, more details at https://crocs.fi.muni.cz/public/papers/dtrap2020</li> </ul> <p>The artifacts contain the full experimental setup (as described in Section 2.1 of the paper) and the complete anonymized dataset underlying the evaluation presented in Sections 3 and 4.</p> <p>The experimental setup contains the documents accompanying the task: the informed consent, pre-task questionnaire, task description, trust scales, and the list of questions posed during the post-task interview (all in PDFs). We further include the custom website with certificate validation documentation for the “redesigned” condition (static HTML). While working on the task, participants in the “redesigned” condition could access this website via a link that was in the redesigned error messages. Furthermore, we provide the software with which the participants interacted. It contains the displayed error messages and validated certificates. These things are available both individually and incorporated in a snapshot of a virtual machine used at the experiment (importable directly into VirtualBox).</p> <p>The collected data is presented in a single dataset (SPSS format; you can use PSPP as a free alternative). It includes the analysis syntax files to obtain the numerical results presented in the paper. For each participant, the dataset contains: 1) pre-task questionnaire answers, 2) reported trust ratings, 3) sub-task timing, 4) information on whether they browsed the Internet and 5) the interview codes assigned. Note that we do not publish the interview transcripts to preserve participant privacy.</p>
Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal
<p>This Research Artifact contains supplemental material from the study: "Supplemental Material: Research Artifacts for Human-Oriented Experiments in Software Engineering: An ACM Badges–driven Structure Proposal". The supplementary material contains:</p> <ol> <li>The list of the 106 primary studies classified by journals.</li> <li>The list of the 12 research artifacts classified by conferences.</li> <li>The .xlsx file of the dataset used to analyze the RQs.</li> <li>The .xlsx file of the dataset used to analyze the research artifacts problems (Table 3). </li> <li>The list of the figures published in the scientific article.</li> </ol>
Dataset: ACM Research, Inc. (ACMR) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: ACM Research, Inc. (ACMR) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
ACM MMSys ODS The Unobtrusive Group Interaction (UGI) Corpus
<p>Studying group dynamics requires fine-grained spatial and temporal understanding of human behavior. Social psychologists studying human interaction patterns in face-to-face group meetings often find themselves struggling with huge volumes of data that require many hours of tedious manual coding. There are only a few publicly available multi-modal datasets of face-to-face group meetings that enable the development of automated methods to study verbal and non-verbal human behavior. In this paper, we present a new, publicly available multi-modal dataset for group dynamics study that differs from previous datasets in its use of ceiling-mounted, unobtrusive depth sensors. These can be used for fine-grained analysis of head and body pose and gestures, without any concerns about participants' privacy or inhibited behavior. The dataset is complemented by synchronized and time-stamped meeting transcripts that allow analysis of spoken content. The dataset comprises 22 group meetings in which participants perform a standard collaborative group task designed to measure leadership and productivity. Participants' post-task questionnaires, including demographic information, are also provided as part of the dataset. We show the utility of the dataset in analyzing perceived leadership, contribution, and performance, by presenting results of multi-modal analysis using our sensor-fusion algorithms designed to automatically understand audio-visual interactions.</p>
Dataset to ACM MMSys'20 paper entitled "Comparing Fixed and Variable Segment Durations for Adaptive Video Streaming – A Holistic Analysis"
<p>Dataset for the ACM MMSys'20 paper entitled "Comparing Fixed and Variable Segment Durations for Adaptive<br> Video Streaming – A Holistic Analysis".<br> The dataset includes</p> <ul> <li>Results from video encoding (using variable and fixed segment durations)</li> <li>Video sequences used for streaming evaluations</li> </ul>
Artifact for ACM CSUR article: "A Meta-Study of Software-Change Intentions"
<p>This archive includes the CSV files with the raw and aggregated data we used for our meta study:</p> <p>"A Meta-Study of Software-Change Intentions" published at the ACM Computing Surveys journal.</p> <p> </p> <p>Please refer to the readme for details on the files.</p>
Climate model (CM2.6) and regional model (ACM) processed output used to investigate the physical drivers and biogeochemical effects of the weakening of the northwest North Atlantic Shelfbreak Jet (Garcia-Suarez & Fennel., 2024; JAMES)
<p>Key processed output from the climate model GFDL CM2.6 and the regional Atlantic Canada model (ACM) used to investigate the physical drivers and the biogeochemical effects of the weakening of the shelfbreak jet in the northwest North Atlantic Ocean. The dataset includes all model variables required to reproduce the key results in <em>Garcia-Suarez & Fennel (2024, JAMES)</em>. See <em>GarciaSuarezandFennel_JAMES_CM26_ACM_data_README_v2.txt</em> for more details.</p>
Tutorial for the 2022 ACM SIGMOD Conference: Spatial Data Quality in the IoT Era: Management and Exploitation
<p>Within the rapidly expanding Internet of Things (IoT), growing amounts of spatially referenced data are being generated. Due to the dynamic, decentralized, and heterogeneous nature of the IoT, spatial IoT data (SID) quality has attracted considerable attention in academia and industry. How to invent and use technologies for managing spatial data quality and exploiting low-quality spatial data are key challenges in the IoT. In this tutorial, we highlight the SID consumption requirements in applications and offer an overview of spatial data quality in the IoT setting. In addition, we review pertinent technologies for quality management and low-quality data exploitation, and we identify trends and future directions for quality-aware SID management and utilization. The tutorial aims to not only help researchers and practitioners to better comprehend SID quality challenges and solutions, but also offer insights that may enable innovative research and applications.</p>
BBR and ACM-RT inputs for figures in ACMB-DF documenting paper (Barker et al., 2024).
<p><span>For the Earth Cloud, Aerosol, Radiation Explorer (EarthCARE) satellite mission there are a number of algorithms used to process the observations. One of these algorithms, called ACMB-DF, is designed to perform continuous radiative closure assessment of EarthCARE observations. This is done using radiance and flux measurements from the Broad-Band Radiometer (BBR) and forward solar and thermal radiative transfer calculations applied to retrieved geophysical properties. The ACMB-DF algorithm is documented in an Atmospheric Measurements and Techniques (AMT) article. </span></p> <p><span>To illustrate the methodology used for ACMB-DF, detailed calculations were performed for the “Hawaii” scene.<span> </span>These include detailed 3D Monte Carlo calculations of the BBR and Multi-Spectral Imager (MSI) radiances applied to select sections of the Hawaii test scene.<span> </span>These where then used in the chain of EarthCARE retrievals to produce geophysical retrievals used as input for the ACM-RT processor to compute radiative quantities needed for closure assessment and as input to generate BBR radiances and fluxes.</span></p> <p><span>The relevant publication for these calculations is:</span></p> <p><span>Barker, H. W., J. N. S. Cole, N. Villefranque, Z. Qu, A. Velazquez-Blazquez, C. Domenech, S. L. Mason, and R. J. Hogan : Radiative Closure Assessment of Retrieved Cloud and Aerosol Properties for the EarthCARE Mission: The ACMB-DF Product. Submitted to AMT, May 2024.</span></p>
CVoiceFake-Full ("SafeEar: Content Privacy-Preserving Audio Deepfake Detection", ACM CCS 2024)
<h1><strong>Introduction:</strong></h1> <p>CVoiceFake (Full) encompasses <strong>five common languages (English, Chinese, German, French, and Italian)</strong> and utilizes <strong>multi-advanced and classical voice cloning techniques</strong> (Parallel WaveGAN, Multi-band MelGAN, Style MelGAN, Griffin-Lim, WORLD, and DiffWave) to produce audio samples that bear a high resemblance to authentic audio.</p> <ol> <li><strong>Parallel WaveGAN</strong>: As a non-autoregressive vocoder-based model, Parallel WaveGAN produces high-fidelity audio rapidly, ideal for efficient and quality deepfake generation.</li> <li><strong>Multi-band MelGAN</strong>: Multi-band MelGAN is a variant of MelGAN that divides the frequency spectrum into sub-bands for faster and more stable multi-lingual vocoder training, enhancing the robustness and scalability of the dataset.</li> <li><strong>Style MelGAN</strong>: Style MelGAN is designed to capture fine prosodic and stylistic nuances of speech, making it particularly compelling for deepfake applications that require high levels of expressivity and variation in speech synthesis.</li> <li><strong>Griffin-Lim</strong>: This algorithm reconstructs waveforms from spectrograms using an iterative phase estimation method. Though less high-fidelity than neural vocoders, it serves as a traditional baseline for comparing deepfake generation.</li> <li><strong>WORLD</strong>: WORLD is a statistical parameter-based voice synthesis system that offers fine control over the spectral and prosodic features of the synthesized audio. Its fine manipulation is useful for crafting the nuanced variations needed in deepfake datasets.</li> <li>We have also built the SOTA diffusion-based deepfake audio (DiffWave); please contact the author at <code>xinfengli@zju.edu.cn</code> if you are interested in the dataset, particularly the DiffWave portion. Furthermore, any additional discussions are welcomed.<br><strong>DiffWave</strong>: DiffWave is a diffusion probability model for waveform generation. It converts the white noise signal into structured waveform through a Markov chain, capable of both conditional and unconditional generation tasks. DiffWave represents the advanced synthesis method for its fast synthesis speed and high synthesis quality.</li> </ol> <h1><strong>🔥 News:</strong></h1> <p>Please note that we recently released our DiffWave subset in Version 2 in comparison to Version 1. You can download the file named CVoiceFake_Large_diffwave_update.tar.gz.xx, and after unzipping it, you will find it retains the same file structure as before.<br> <strong>| CVoiceFake_Large_diffwave_update.tar.gz.00 |<br> | CVoiceFake_Large_diffwave_update.tar.gz.01 |</strong></p> <p> </p> <h1><strong>Full Dataset & Project Page:</strong></h1> <p>The sampled small dataset is available on <a href="11124319" target="_blank" rel="noopener">CVoiceFake Small</a> as well. Please kindly also refer to the project page: <a title="SafeEar Website" href="https://safeearweb.github.io/Project/" target="_blank" rel="noopener">SafeEar Website</a>.</p> <p> </p> <h1><strong>Citation:</strong></h1> <p>If you find our paper/code/dataset helpful, please kindly consider citing this work with the following reference:</p> <pre>@inproceedings{li2024safeear,<br> author = {Li, Xinfeng and Li, Kai and Zheng, Yifan and Yan, Chen and Ji, Xiaoyu, and Xu, Wenyuan},<br> title = {{SafeEar: Content Privacy-Preserving Audio Deepfake Detection}},<br> booktitle = {Proceedings of the 2024 {ACM} {SIGSAC} Conference on Computer and Communications Security (CCS)}<br> year = {2024},<br>}</pre>
arXiv-ACM Dataset
<p>arXiv-ACM Dataset used in: Exploiting Label Dependencies for Multi-Label Text Classification Using Transformers </p>
Supplemental Material: Laboratory Packages for Human– Oriented Experiments in Software Engineering: An ACM Badging Approach
<p>This laboratory package contains supplemental material from the study: "Laboratory Packages for Human–Oriented Experiments in Software Engineering: An ACM Badging Approach". The supplementary material contains:<br> 1. The list of the 118 primary studies (Labpacks) classified by journals and conferences.<br> 2. The .xlsx file of the dataset used to analyze the RQs.<br> 3. The list of figures published in the scientific article.</p>
smer_acm_bcb_20
<p>Structural representations of DNA regulatory substrates can enhance sequence-based algorithms by associating functional sequence variants</p>
Artfiact for the Proc. ACM Softw. Eng. article "Sharing Software-Evolution Datasets: Practices, Challenges, and Recommendations."
<p>This dataset is a collection of all notes taken for the article:</p> <p> David Broneske, Sebastian Kittan, and Jacob Krüger:<br> Sharing Software-Evolution Datasets: Practices, Challenges, and Recommendations. <br> Proc. ACM Softw. Eng. 1, FSE, 2024.<br> https://doi.org/10.1145/3660798</p> <p>Please refer to the readme for a description of the files involved in the zip file.</p>
Study of Comparative Bioavailability and Pharmacokinetics of ACM-001.1) and Pindolol in Healthy Volunteers (HV)
ClinicalTrials.gov study NCT06028321. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.