Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
42
datasets available to search
ShareScore release 0.7.1
Dataset results
42 results for “data artifacts”
data artifact for "Keeping It Real: Why HPC Data Services Don't Achieve I/O Microbenchmark Performance"
<p>This is the data artifact associated with the PDSW 20 workshop paper submission entitled "Keeping It Real: Why HPC Data Services Don’t Achieve I/O Microbenchmark Performance" in accordance with the PDSW 20 artifact submission guidelines.</p>
Data from: Fine-scale landscape genetics of the American badger (Taxidea taxus): disentangling landscape effects and sampling artifacts in a poorly understood species
Landscape genetics is a powerful tool for conservation because it identifies landscape features that are important for maintaining genetic connectivity between populations within heterogeneous landscapes. However, using landscape genetics in poorly understood species presents a number of challenges, namely, limited life history information for the focal population and spatially biased sampling. Both obstacles can reduce power in statistics, particularly in individual-based studies. In this study, we genotyped 233 American badgers in Wisconsin at 12 microsatellite loci to identify alternative statistical approaches that can be applied to poorly understood species in an individual-based framework. Badgers are protected in Wisconsin owing to an overall lack in life history information, so our study utilized partial redundancy analysis (RDA) and spatially lagged regressions to quantify how three landscape factors (Wisconsin River, Ecoregions and land cover) impacted gene flow. We also performed simulations to quantify errors created by spatially biased sampling. Statistical analyses first found that geographic distance was an important influence on gene flow, mainly driven by fine-scale positive spatial autocorrelations. After controlling for geographic distance, both RDA and regressions found that Wisconsin River and Agriculture were correlated with genetic differentiation. However, only Agriculture had an acceptable type I error rate (3–5%) to be considered biologically relevant. Collectively, this study highlights the benefits of combining robust statistics and error assessment via simulations and provides a method for hypothesis testing in individual-based landscape genetics.
Artifact: Quick Theory Exploration for Algebraic Data Types via Program Transformations
<p>This is the repeatability package, including the tool (called LemmaCalc) in the paper, the benchmark files, a TheSy binary, Z3 binaries for Linux, and scripts to repeat the experiments.</p> <p>Parts of the paper supported by this artifact: Main results in Fig 1, implementation of Algs 1--3.</p>
Data and code for 'Fast and artifact-free excitation multiplexing using synchronized image scanning'
<p>Data and source code for the publication 'Fast and artifact-free excitation multiplexing using synchronized image scanning'.</p>
Data Artifact for "Gaze into the Pattern: Characterizing Spatial Patterns with Internal Temporal Correlations for Hardware Prefetching"
<p>This dataset contains the Ligra, PARSEC, GAP, and QMM traces used in our paper "Gaze into the Pattern: Characterizing Spatial Patterns with Internal Temporal Correlations for Hardware Prefetching", which is accepted by HPCA'25. These traces are provided as part of the AE proces.</p>
Raw data and R code for the stone artifact analysis in "Lithic miniaturization in South China since the terminal Pleistocene: a multivariate analysis of lithic reduction from Fodongdi, Fulin and Xiqiaoshan"
Open the record for dataset details and reuse information.
Fulfilling Industrial Needs for Consistency Among Engineering Artifacts - Evaluation Data
<p>This is a dataset with the evaluation data from the paper entitled "Fulfilling Industrial Needs for Consistency Among Engineering Artifacts". The dataset also contains two demo videos showing the approach being used in practice.</p>
Artifact Evaluation for "Optimizing Remote Data Transfers in X10" - PACT'18
<p>Artifact Evaluation Reproduction for "Optimizing Remote Data Transfers in X10", PACT 2018.</p> <p>--------------------------------------------------------------------------------------------------------------------</p> <p>This repository contains artifacts and source codes<br> to reproduce experiments from the PACT 2018 research paper<br> titled "Optimizing Remote Data Transfers in X10"</p> <p>Our evaluations are done on Linux based Operation system.</p> <p><br> Hardware pre-requisities:<br> ------------------------------<br> We recommend the following architectures any of the following architectures:<br> * An Intel System with two nodes, each node having 32 cores (32 CPUs). Total 64 cores.<br> * An AMD System, with two nodes, each node having 16 cores (16 CPUs). Total 32 cores.<br> </p> <p><br> Software pre-requisites:<br> ----------------------------<br> 1) For Evaluation using Virtual Machine required software are already installed in the provided image.<br> A VirtualBox (from Oracle) is required to install the image and do evaluation.</p> <p>2) For Manual evaluation, the below required software needs to be installed in the system.<br> * g++ (preferred version 5.4.0)<br> * Apache ant software (preferred version 1.9.6 )<br> * Java software<br> - preferred jvm jdk1.8.0_151 (or) java-8-oracle<br> - After installation, set JAVA_HOME path in ~/.bashrc to point to the installed JAVA bin</p> <p><br> Benchmarks:<br> ---------------<br> 1) Taken from IMSuite Benchmark kernels (http://www.cse.iitm.ac.in/~krishna/imsuite/). Already, included in the package.</p> <p><br> Installation, Execution and Validation of results:<br> ---------------------------------------------------------<br> 1) For Evaluation using Virtual Machine see README.md file in ./AE/VM folder</p> <p>2) For Evaluation using Manual method see README.md file in ./AE/Manual folder</p> <p>3) For Evaluating AT-Opt technique againt varing input sizes (not discussed in the paper), look at<br> web page http://www.cse.iitm.ac.in/~krishna/imsuite/. In the ./AE/Manual folder, "x10-base" folder is<br> the baseline compiler (after bulid) and "x10-atOpt" (after build) folder is the AT-Opt compiler (has our techniques implemented). </p> <p><br> If anything in unclear, or any unexpected results occur, please report it to the authors.</p> <p> </p>
Data for Mathematical Expressions in Software Engineering Artifacts
<p>Data for the experiments in the paper, Mathematical Expressions in Software Engineering Artifacts.</p> <p>The dataset contains the following sub-directories:</p> <ol> <li>Bug_data: Bug reports from 10 open-sourced projects.</li> <li>MathyB_A: Bug reports annotated by MEDSEA.</li> <li>MathyB_H: Bug reports annotated manually by humans.</li> <li>NNGen_modified_data: Modified log messages of <a href="https://github.com/Tbabm/nngen">NNGen</a> dataset (<a href="https://dl.acm.org/doi/10.1145/3238147.3238190">Neural-machine-translation-based commit message generation: how far are we?</a>) based on the annotations made by MEDSEA.</li> </ol>
Data from: Fine-scale landscape genetics of the American badger (Taxidea taxus): disentangling landscape effects and sampling artifacts in a poorly understood species
Open the record for dataset details and reuse information.
Data from: Quantitative comparison of commercial and non-commercial metal artifact reduction techniques in computed tomography
Objectives: Typical streak artifacts known as metal artifacts occur in the presence of strongly attenuating materials in computed tomography (CT). Recently, vendors have started offering metal artifact reduction (MAR) techniques. In addition, a MAR technique called the metal deletion technique (MDT) is freely available and able to reduce metal artifacts using reconstructed images. Although a comparison of the MDT to other MAR techniques exists, a comparison of commercially available MAR techniques is lacking. The aim of this study was therefore to quantify the difference in effectiveness of the currently available MAR techniques of different scanners and the MDT technique. Materials and Methods: Three vendors were asked to use their preferential CT scanner for applying their MAR techniques. The scans were performed on a Philips Brilliance ICT 256 (S1), a GE Discovery CT 750 HD (S2) and a Siemens Somatom Definition AS Open (S3). The scans were made using an anthropomorphic head and neck phantom (Kyoto Kagaku, Japan). Three amalgam dental implants were constructed and inserted between the phantom's teeth. The average absolute error (AAE) was calculated for all reconstructions in the proximity of the amalgam implants. Results: The commercial techniques reduced the AAE by 22.0±1.6%, 16.2±2.6% and 3.3±0.7% for S1 to S3 respectively. After applying the MDT to uncorrected scans of each scanner the AAE was reduced by 26.1±2.3%, 27.9±1.0% and 28.8±0.5% respectively. The difference in efficiency between the commercial techniques and the MDT was statistically significant for S2 (p=0.004) and S3 (p<0.001), but not for S1 (p=0.63). Conclusions: The effectiveness of MAR differs between vendors. S1 performed slightly better than S2 and both performed better than S3. Furthermore, for our phantom and outcome measure the MDT was more effective than the commercial MAR technique on all scanners.
Data from: Suppression of overlearning in independent component analysis used for removal of muscular artifacts from electroencephalographic records
This paper addresses the overlearning problem in the independent component analysis (ICA) used for the removal of muscular artifacts from electroencephalographic (EEG) records. We note that for short EEG records with high number of channels the ICA fails to separate artifact-free EEG and muscular artifacts, which has been previously attributed to the phenomenon called overlearning. We address this problem by projecting an EEG record into several subspaces with a lower dimension, and perform the ICA on each subspace separately. Due to a reduced dimension of the subspaces, the overlearning is suppressed, and muscular artifacts are better separated. Once the muscular artifacts are removed, the signals in the individual subspaces are combined to provide an artifact free EEG record. We show that for short signals and high number of EEG channels our approach outperforms the currently available ICA based algorithms for muscular artifact removal. The proposed technique can efficiently suppress ICA overlearning for short signal segments of high density EEG signals.
VigIA: data and artifacts
<p>These data and artifacts support the VigIA tool and paper currently under review.</p>
Raw Data and Artifact
<p>Raw Data and Artifact</p>
Data from: Quantitative comparison of commercial and non-commercial metal artifact reduction techniques in computed tomography
Open the record for dataset details and reuse information.
Data from: Suppression of overlearning in independent component analysis used for removal of muscular artifacts from electroencephalographic records
Open the record for dataset details and reuse information.
Benchmarking dataset for measurement and removal of index hopping artifacts in multiplexed droplet-based single-cell RNA-seq data
GEO Series GSE149087. Homo sapiens. 4 samples. Type: Other.
Artifact for Counterfactual Explanations for Machine Learning on Multivariate HPC Time Series Data
<p>This includes the data sets used in the SC'20 submission "Counterfactual Explanations for Machine Learning on Multivariate HPC Time Series Data".</p>
The lead isotope data of unalloyed copper artifacts of the Shang and Eastern Zhou dynasties, and copper ore samples of modern copper deposits in Huili.
<p>The lead isotope data of unalloyed copper artifacts of the Shang and Eastern Zhou dynasties, and copper ore samples of modern copper deposits in Huili. And the results of four runs for the SRM981 determination and published values.</p>
MRI raw data to publication: Fuzzy ripple artifact in high resolution fMRI: identification, cause, and mitigation.
<p>These Data are refering to the maunscript "Fuzzy ripple artifact in high resolution fMRI: identification, cause, and mitigation" authored by </p> <p>Renzo Huber1, Rüdiger Stirnberg2, A Tyler Morgan1, David A Feinberg3,4-5, Samantha J Ma6, Philipp Ehses3, Omer Faruk Gulban2,7, Kenshu Koiso2, Isabel Gephart1, Stephanie Swegle1, Susan Wardle1, Emily Ma2, Andrew Persichetti1, Alexander JS Beckett4-5, Tony Stöcker3, Nicolas Boulant8, Benedikt A Poser2, Peter Bandettini1</p> <p><strong> </strong></p> <p>1 NIMH, NIH, Bethesda, United States,</p> <p>2 German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany,</p> <p>3 Helen Wills Neuroscience Institute, University of California, Berkeley, Berkeley, CA, United States, </p> <p>4 Advanced MRI Technologies, Sebastopol, CA, United States,</p> <p>5 CN, FPN, University of Maastricht, The Netherlands,</p> <p>6 Siemens Medical Solutions USA, Inc., Berkeley, CA, USA,</p> <p>7 Brain Innovation, Maastricht, The Netherlands,</p> <p>8 CEA, NeuroSpin, University Paris Saclay, France.</p> <p> </p> <h1>Abstract</h1> <p><strong>Purpose:</strong> High resolution fMRI is an emerging research field focused on capturing functional signal changes across cortical layers. However, the data acquisition is limited by low spatial frequency EPI artifacts; termed as Fuzzy Ripples. These artifacts limit the practical applicability of acquisition protocols with higher spatial resolution, faster acquisition speed, and they challenge imaging in lower brain areas. </p> <p><strong>Methods: </strong>We characterize Fuzzy Ripple artifacts across commonly used sequences and distinguish them from conventional EPI Nyquist ghosts, off-resonance effects, and GRAPPA artifacts. To investigate their origin, we employ dual polarity readouts.</p> <p><strong>Results: </strong>Our findings indicate that Fuzzy Ripples are primarily caused by kx-specific imperfections in gradient trajectories, which can be exacerbated by inductive coupling between third-order shims and readout gradients. We also find that these artifacts can be mitigated through complex-valued averaging of dual polarity EPI or by disconnecting the third-order shim.</p> <p><strong>Conclusion:</strong> The proposed mitigation strategies allow for overcoming current limitations in layer-fMRI protocols: </p> <p>(1) Achieving resolutions beyond 0.8mm is feasible, and even at 3T, we achieved 0.53mm voxel functional connectivity mapping. </p> <p>(2) Temporal acquisition speed can be increased to GRAPPA 8. </p> <p>(3) Sub-millimeter fMRI is achievable in lower brain areas, including the cerebellum.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.