Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,956
datasets available to search
ShareScore release 0.9.0
Dataset results
1,956 results for “test data”
TCTracer: Establishing Test-to-Code Traceability Links Using Dynamic and Static Techniques - Evaluation Data - Empirical Software Engineering 2021
<p>This repository provides the data artefacts for the experiments conducted using our tool TCTracer for the journal paper "TCTracer: Establishing Test-to-Code Traceability links Using Dynamic and Static Techniques" as submitted to the Empirical Software Engineering journal in 2021.</p>
Data from: Computer-aided X-ray screening for tuberculosis and HIV testing among adults with cough in Malawi (the PROSPECT study): a randomized trial and cost-effectiveness analysis
<p>Suboptimal tuberculosis (TB) diagnostics and HIV contribute to the high global burden of TB. We investigated costs and yield from systematic HIV-TB screening, including computer-aided digital chest X-ray (DCXR-CAD). Suboptimal tuberculosis (TB) diagnostics and HIV contribute to the high global burden of TB. We investigated costs and yield from systematic HIV-TB screening, including computer-aided digital chest X-ray (DCXR-CAD).</p> <p>In this open, three-arm randomised trial, adults (≥18 years) with cough attending acute primary services in Malawi were randomised (1:1:1) to standard-of-care (SOC); oral HIV testing (HIV screening) and linkage to care; or HIV testing and linkage to care plus DCXR-CAD with sputum Xpert for high CAD4TBv5 scores (HIV-TB screening). Participants and study staff were not blinded to intervention allocation, but investigator blinding was maintained until final analysis. The primary outcome was time to TB treatment. Secondary outcomes included proportion with same-day TB treatment; prevalence of undiagnosed/untreated bacteriologically-confirmed TB on day 56; and undiagnosed/untreated HIV. Analysis was done on an intention to treat basis. Cost-effectiveness analysis used a health-provider perspective. Between 15/11/2018-27/11/2019, 8236 were screened for eligibility, with 473, 492, and 497 randomly allocated to SOC, HIV, and HIV-TB screening arms; 53 (11%), 52 (9%), and 47 (9%) were lost to follow-up, respectively. At 56 days, TB treatment had been started in 5 (1.1%) SOC, 8 (1.6%) HIV-screening, and 15 (3.0%) HIV-TB screening participants. Median (IQR) time to TB treatment was 11 (6.5-38), 6 (1-22) and 1 (0-3) days (hazard ratio for HIV-TB vs. SOC: 2.86, 1.04-7.87), with same-day treatment of 0/5 (0%) SOC, 1/8 (12.5%) HIV, and 6/15 (40.0%) HIV-TB screening arm TB patients (p=0.03). At day 56, 2 SOC (0.5%), 4 HIV (1.0%), and 2 HIV-TB (0.5%) participants had undiagnosed microbiologically-confirmed TB. HIV screening reduced the proportion with undiagnosed or untreated HIV from 10 (2.7%) in the SOC arm to 2 (0.5%) in the HIV-screening arm (risk ratio [RR]: 0.18, 0.04-0.83), and 1 (0.2%) in the HIV-TB screening arm (RR: 0.09, 0.01-0.71). Incremental costs were US$3.58 and US$19.92 per participant screened for HIV and HIV-TB; the probability of cost-effectiveness at a US$1200/quality-adjusted life-year (QALY) threshold were 83.9% and 0%. Main limitations were the lower than anticipated prevalence of tuberculosis and short participant follow-up period; cost and quality of life benefits of this screening approach may accrue over a longer time horizon.</p> <p>DCXR-CAD with universal HIV screening significantly increased the timeliness and completeness of HIV and TB diagnosis. If implemented at scale this has potential to rapidly and efficiently improve TB and HIV diagnosis and treatment.</p>
IWC : Test data for VGP workflows v2.0
<p>Test data for VGP workflows in <a href="https://github.com/galaxyproject/iwc">iwc</a>. Pipeline VGP assembly v2.0</p>
Processed data for the "Property-Based Testing of Web APIs" paper
<p>Processed data for the "Property-Based Testing of Web APIs" paper. Each directory in the archive consists of:</p> <p>- metadata.json. Metadata about a test run - tested fuzzer name, run duration, etc</p> <p>- fuzzer.json - Structured fuzzer output</p> <p>- deduplicated_cases.json - Deduplicated reported failures, when fuzzers provide it</p> <p>- sentry.json - Cleaned Sentry events for this run</p> <p>- target.json - Parsed stdout for Gitlab & Disease.sh targets that were tested without Sentry integration</p>
Data for: Are immigrants outbred and unrelated? Testing standard assumptions in a wild metapopulation
<p><span>Immigration into small recipient populations is expected to alleviate inbreeding and increase genetic variation, and hence facilitate population persistence through genetic and/or evolutionary rescue. Such expectations depend on three standard assumptions: that immigrants are outbred, unrelated to existing natives at arrival, and unrelated to each other. These assumptions are rarely explicitly verified, including in key field systems in evolutionary ecology. Yet, they could be violated due to non-random or repeated immigration from adjacent small populations. We combined molecular genetic marker data for 150-160 microsatellite loci with comprehensive pedigree data to test the three assumptions for a song sparrow (<i>Melospiza melodia)</i> population that is a model system for quantifying effects of inbreeding and immigration in the wild. Immigrants were less homozygous than existing natives on average, with mean homozygosity that closely resembled outbred natives. Immigrants can therefore be considered outbred on the focal population scale. Comparisons of homozygosity of real or hypothetical offspring of immigrant-native, native-native and immigrant-immigrant pairings implied that immigrants were typically unrelated to existing natives and to each other. Indeed, immigrants' offspring would be even less homozygous than outbred individuals on the focal population scale. The three standard assumptions of population genetic and evolutionary theory were consequently largely validated. Yet, our analyses revealed some deviations that should be accounted for in future analyses of heterosis and inbreeding depression, implying that the three assumptions should be verified in other systems to probe patterns of non-random or repeated dispersal and facilitate precise and unbiased estimation of key evolutionary parameters.</span></p>
SEED-G: Simulated EEG Data Generator for testing connectivity algorithms
<p>SEED-G toolbox was developed in MATLAB environment (tested on version R2017a and R2020b) and released on the GitHub page <a href="https://github.com/aanzolin/SEED-G-toolbox">https://github.com/aanzolin/SEED-G-toolbox</a> (accessed date 12 April 2021). It is organized in the following subfolders:</p> <ul> <li> <p><strong>main</strong>: it is the core of the toolbox and contains all the functions for the generation of EEG data according to a predefined ground-truth network.</p> </li> <li> <p><strong>dependencies</strong>: containing parts of other toolboxes required to successfully run SEED-G functions. The links to the full packages can be found in the documentation on the GitHub page. The additional packages are Brain Connectivity Toolbox (BCT) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B40-sensors-21-03632">40</a>], FieldTrip [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B41-sensors-21-03632">41</a>], Multivariate Granger Causality Toolbox (MVGC) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B24-sensors-21-03632">24</a>], and AsympPDC Package (PDC_AsympSt) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B42-sensors-21-03632">42</a>,<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B43-sensors-21-03632">43</a>]. Additionally, the implemented forward model is solved according to the New York Head (NYH) model, whose parameters are contained in the structure available on the ICBM-NY platform [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B28-sensors-21-03632">28</a>].</p> </li> <li> <p><strong>real data</strong>: containing real EEG data acquired from one healthy subject during resting state at scalp level (‘EEG_real_sources.mat’) and its reconstructed version in source domain (‘sLOR_cortical_sources.mat’). These signals can be employed to extract the AR components to be included in the model to generate data with the same spectral properties of the real ones.</p> </li> <li> <p><strong>demo</strong>: containing examples of MATLAB scripts to be used to learn the different functionalities of the toolbox. For example, the code ‘run_generation.m’ allows to specify the directory containing the real sources and each specific input of the function ‘simulatedData_generation.m’.</p> </li> <li> <p><strong>auxiliary functions</strong>: containing either original MATLAB functions or modified version of free available functions.</p> </li> </ul>
Hydrogeophysical data Schillerslage test site, joint MRT ERT GPR
<p>Dataset of hydrogeophysical Survey at Schillerslage test site including magnetic resonance tomography (MRT), ground-penetrating Radar (GPR) and electrical resistivity tomography (ERT).</p> <p> </p>
Data in support to the manuscript: Testing a novel sensor design to jointly measure cosmic-ray neutrons, muons and gamma rays for non-invasive soil moisture estimation by Gianessi et al. (2024)
<p>The files contain data presented and discussed in the manuscript: Testing a novel sensor design to jointly measure cosmic-ray neutrons, muons and gamma rays for non-invasive soil moisture estimation by Gianessi et al. (2024).</p> <div> <div>Gianessi, Stefano, Matteo Polo, Luca Stevanato, Marcello Lunardon, Till Francke, Sascha E. Oswald, Hami Said Ahmed, et al. “Testing a Novel Sensor Design to Jointly Measure Cosmic-Ray Neutrons, Muons and Gamma Rays for Non-Invasive Soil Moisture Estimation.” <em>Geoscientific Instrumentation, Methods and Data Systems</em> 13, no. 1 (January 16, 2024): 9–25. <a href="https://doi.org/10.5194/gi-13-9-2024">https://doi.org/10.5194/gi-13-9-2024</a>.</div> </div> <p> </p>
Data from: A test of the competitive ability – cold tolerance trade-off hypothesis in seasonally breeding beetles
<p><span>Closely related species that use similar resources often differ in their seasonal patterns of activity, but the factors that limit their distributions across seasons are unknown for most species. One hypothesis to explain seasonal variation in the distributions of species involves a trade-off between competitive ability and cold tolerance, where tolerance to the cold compromises competitive ability in warmer (benign) temperatures, either at the level of the individual or population.</span></p> <p><span>We tested both individual-level and population-level mechanisms of this hypothesis in two co-occurring species of temperate burying beetles (Silphidae: <em>Nicrophorus sayi</em>, <em>N. orbicollis</em>) that differ in their seasonal patterns of activity.</span></p> <p><span>We measured cold tolerance, breeding activity as a function of temperature, and competitive ability as a function of temperature and season.</span></p> <p><span>Consistent with our hypothesis, the mid-season <em>N. orbicollis</em> was less able to function at the cold temperatures that characterize early spring, when the early-season <em>N. sayi</em> is most active. The larger beetle, however, always won one-on-one competitive trials at warm temperatures, regardless of species, inconsistent with an individual-level trade-off. <em>N. orbicollis</em> was usually larger and successful when competing for the same carrion later in the season, mostly because of its larger population size, consistent with a trade-off between competitive ability and cold tolerance acting at the population level.</span></p> <p><span>Our findings suggest that cold temperatures limit the mid-season <em>N. orbicollis</em> from earlier spring emergence, while competitive pressure from the more abundant, larger <em>N. orbicollis</em> constrains the early-season <em>N. sayi</em> from remaining active through the summer.</span></p>
Data for: Testing for fitness epistasis in a transplant experiment identifies a candidate adaptive locus in Timema stick insects
<p>Identifying the genetic basis of adaptation is a central goal of evolutionary biology. However, identifying genes and mutations affecting fitness remains challenging because a large number of traits and variants can influence fitness. Selected phenotypes can also be difficult to know <em>a priori</em>, complicating top-down genetic approaches for trait mapping that involve crosses or genome-wide association studies. In such cases, experimental genetic approaches, where one maps fitness directly and attempts to infer the traits involved afterward, can be valuable. Here, we re-analyse data from a transplant experiment involving <em>Timema</em> stick insects, where five physically clustered SNPs associated with cryptic body colouration were shown to interact to affect survival. Our analysis covers a larger genomic region than past work and revealed a locus previously not identified as associated with survival. This locus resides near a gene, <em>Punch</em> (<em>Pu</em>), involved in pteridine pigments production, implying that it could be associated with an unmeasured colouration trait. However, by combining previous and newly obtained phenotypic data, we show that this trait is not eye or body colouration. We discuss the implications of our results for the discovery of traits, genes, and mutations associated with fitness in other systems, as well as for supergene evolution.</p>
Data presented in "An ultracold molecular beam for testing fundamental physics"
<p>Data presented in "An ultracold molecular beam for testing fundamental physics". The original paper can be found at https://doi.org/10.1088/2058-9565/ac107e. This is the data underlying the simulations and the experimental results.</p>
Data from a test of female defense in male collared lizards
<p>A widely held principle in behavioral ecological research is that polygynous social systems evolve either by direct male defense of females or male defense of resources, although which of these mechanisms applies in particular species is rarely examined experimentally. We tested the relative importance of female versus resource defense in polygynous territorial male collared lizards (<em>Crotaphytus</em> <em>collaris</em>). Using a novel experimental design, we temporarily removed some of the resident females from male territories to create a female-free removal zone, whereas resident females were left intact within a non-removal zone. We then compared activity of males within each zone during three experimental phases: before we removed females, for two days when females were absent, and the day following return of females. If males defend females directly, we expected them to adjust the location of their patrol and display within removal and non-removal zones depending on the presence/absence of females, whereas we expected no such change if males defend resources. Male activity in the removal zone generally decreased when females were removed but then increased when females were replaced, whereas we observed the opposite pattern in the non-removal zone. The observed shifts in the location of patrol and display in response to the presence/absence of females, while resources remained constant, indicate that polygynous male collared lizards defend females directly. Our results suggest that male collared lizards take advantage of strong female philopatry to relatively small areas by focusing their patrol and display activities where potential mates reside.</p>
Dental mesowear and microwear raw data for Cervus elaphus, Rupicapra pyrenaica and Sus scrofa from Balma del Gai; and the ANOVA - test for equal means
<p>Quantitative data for the dental microwear and mesowear analyses on <em>Cervus elaphus</em>, <em>Rupicapra pyrenaica </em>and <em>Sus scrofa</em> from the Epipalaeolithic sequence of Balma del Gai (Moià, Spain). And the ANOVA - Test for equal means.</p>
Code and data for: Nitrogen-induced hysteresis in grassland biodiversity: A theoretical test of litter-mediated mechanisms
<p>The global rise in anthropogenic reactive nitrogen (N) and the negative impacts of N deposition on terrestrial plant diversity are well-documented. The R* theory of resource competition predicts reversible decreases in plant diversity in response to N loading. However, empirical evidence for the reversibility of N-induced biodiversity loss is mixed. In a long-term N-enrichment experiment in Minnesota, a low-diversity state that emerged during N addition has persisted for decades after additions ceased. Hypothesized mechanisms preventing recovery of biodiversity include nutrient recycling, insufficient external seed supply, and litter inhibition of plant growth. Here we present an ODE model that unifies these mechanisms, produces bistability at intermediate N inputs, and qualitatively matches the observed hysteresis at Cedar Creek. Key features of the model, including native species' growth advantage in low-N conditions and limitation by litter accumulation, generalize from Cedar Creek to North American grasslands. Our results suggest that effective biodiversity restoration in these systems may require management beyond reducing N inputs, such as burning, grazing, haying, and seed additions. By coupling resource competition with an additional inter-specific inhibitory process, the model also illustrates a general mechanism for bistability and hysteresis that may occur in multiple ecosystem types. </p>
Data from: Testing the success of palaeontological methods in the delimitation of clam shrimp (Crustacea, Branchiopoda) on extant species
<p><span>Fossil spinicaudatan taxonomy heavily relies on carapace features (size, shape, ornamentation), and palaeontologists have greatly refined methods to study and describe carapace variability. Whether carapace features alone are sufficient for distinguishing between species of a single genus has remained untested. In our study, we tested common palaeontological methods on 481 individuals of the extant Australian genus <em>Ozestheria</em> that have been previously assigned to ten species based on genetic analysis. All species are morphologically distinct based on geometric morphometrics (p </span><span>≤ </span><span>0.001), but they occupy overlapping regions in <em>Ozestheria</em> morphospace. Linear discriminant analysis of Fourier shape coefficients reaches a mean model performance of 93.8% correctly classified individuals over all possible 45 pairwise species comparisons. This can be further increased by combining the size and shape datasets. Nine of the ten examined species are clearly sexually dimorphic but male and female morphologies strongly overlap within species with little influence on model performance. Ornamentation is commonly species-diagnostic; seven ornamentation types are distinguished of which six are species-specific while one is shared by four species. A transformation of main ornamental features (e.g. from punctate to smooth) can occur among closely related species suggesting short evolutionary timescales. Our overall results support the taxonomic value of carapace features, which should also receive greater attention in the taxonomy of extant species. The extensive variation in carapace shape and ornamentation is noteworthy and several species would probably have been assigned to different genera or families if these had been fossils, bearing implications for the systematics of fossil Spinicaudata.</span></p>
FijiRelax plugin test data
<p>Test dataset and experimental results of the paper: <strong>"FijiRelax: Fast and noise-corrected estimation of NMR relaxation maps in 3D + t".</strong></p> <p> </p> <p> </p>
Semi-implicit barotropic mode solver using ForTrilinos in MPAS-O and its test data
<p>To run the ForTrilinos-enabled MPAS-O,</p> <ol> <li>Download all files</li> <li>Install Trilinos <ul> <li>Unzip: tar -xzvf Trilinos.tar.gz</li> <li>cd Trilinos ; mkdir build ; cd build ; cp ../do-configure_gnu ./</li> <li>Check install directories and options in 'do-configure_gnu'</li> <li>Run 'do-configure_gnu'</li> <li>Trilinos information & installation refer to <a href="https://trilinos.github.io/">https://trilinos.github.io/</a></li> </ul> </li> <li>Install ForTrilinos (inside Trilinos) <ul> <li>cd Trilinos/ForTrilinos ; mkdir build ; cd build ; cp ../do-configure_gnu ./</li> <li>Check Trilinos and ForTrilinos install directories and options in 'do-configure_gnu'</li> <li>Run 'do-configure_gnu'</li> <li>ForTrilinos information & installation refer to <a href="https://fortrilinos.readthedocs.io/en/latest/">https://fortrilinos.readthedocs.io/en/latest/</a></li> </ul> </li> <li>Install ForTrilinos-enabled MPAS-O <ul> <li>Unzip: tar xzvf MPAS-Model_fortrilinos.tar.gz</li> <li>cd MPAS-Model_fortrilinos</li> <li>Check ForTrilinos directories at line 478 (FORTRILINOS_ROOT) in 'Makefile' </li> <li>Install PIO (refer to <a href="https://ncar.github.io/ParallelIO/">https://ncar.github.io/ParallelIO/</a>)</li> <li>MAPS-O information & installation refer to <a href="https://mpas-dev.github.io/ocean/ocean.html">https://mpas-dev.github.io/ocean/ocean.html</a></li> <li>For GNU compiler: make gnu-nersc USE_PIO2=true FORTRILINOS=true</li> </ul> </li> <li>Run test cases <ul> <li>Example <ul> <li>cd MPAS-O_Initial_data/baroclinicEddies/strong_scaling</li> </ul> </li> <li>Link a MPAS-O compiled executable to a test case directory: <ul> <li>ln -fs MPAS-Model_fortrilinos/ocean_model MPAS-O_Initial_data/baroclinicEddies/strong_scaling/</li> </ul> </li> <li>Link a XML deck (solver configurations) for Trilinos to a test case directory <ul> <li>ln -fs xml_decks/no_precond/stratimikos.xml_SCG MPAS-O_Initial_data/baroclinicEddies/strong_scaling/stratimikos.xml</li> </ul> </li> <li>Run <ul> <li>mpirun -n $N ocean_model</li> <li> <p>If the simulation was successful, you will see:</p> <pre><code>tail -n 1 log.ocean.0000.out Logging complete. Closing file at ...</code></pre> <p> </p> </li> </ul> </li> <li>For the global test case, please download here: <a href="https://doi.org/10.5281/zenodo.1252425">MPAS-O_V6.0_RRS30to10.tar</a>. Please see <a href="http://mpas-dev.github.io/">http://mpas-dev.github.io</a> for User's Guide, github release page, description of each test case, and more.</li> </ul> </li> </ol>
Data used in: Born rule as a test of the accuracy of a public quantum computer
<p>A data and scripts used during the preparation of <em>Born rule as a test of the accuracy of a public quantum computer</em>.</p>
Training and test data for: Not getting in too deep: A practical deep learning approach to routine crystallisation image classification
<p>These data were used to classify crystallisation experiments in Milne et al., (<a href="https://doi.org/10.1101/2022.09.28.509868">https://doi.org/10.1101/2022.09.28.509868</a>). Here, four of the most widely-used convolutional deep-learning network architectures that can be implemented without the need for extensive computational resources were compared. It was shown that the classifiers have different strengths that can be combined to provide an ensemble classifier achieving a classification accuracy comparable to that obtained by a large consortium initiative (Bruno et al. PLOS one, 13(6), 2018). Eight classes were used to rank the experimental outcomes, thereby providing detailed information that can be used with routine crystallography experiments to automatically identify crystal formation for drug discovery and pave the way for further exploration of the relationship between crystal formation and crystallisation conditions.</p>
Data from: Summary tests of introgression are highly sensitive to rate variation across lineages
<p>The evolutionary implications and frequency of hybridization and introgression are increasingly being recognized across the tree of life. To detect hybridization from multi-locus and genome-wide sequence data, a popular class of methods is based on summary statistics from subsets of 3 or 4 taxa. However, these methods often carry the assumption of a constant substitution rate across lineages and genes, which is commonly violated in many groups. In this work, we quantify the effects of rate variation on the <em>D </em>test (also known as ABBA-BABA test), the <em>D</em><sub>3</sub> test, and HyDe. All three tests are used widely across a range of taxonomic groups, in part because they are very fast to compute. We consider rate variation across species lineages, across genes, their lineage-by-gene interaction, and residual variation across gene-tree edges. We do so by simulating gene trees within species networks according to a birth-death-hybridization process so as to capture a range of realistic species phylogenies. For all three methods tested, we found a marked increase in the false discovery of reticulation (type-1 error rate) when there is rate variation across species lineages. The <em>D</em><sub>3</sub> test was the most sensitive, with around 80% type-1 error, such that <em>D</em><sub>3</sub> appears to be more sensitive to a departure from the clock than to the presence of reticulation. For all three tests, the power to detect hybridization events decreased as the number of hybridization events increased, indicating that multiple hybridization events can obscure one another if they occur within a small subset of taxa. Our study highlights the need to consider rate variation when using site-based summary statistics and points to the advantages of methods that do not require assumptions on evolutionary rates across lineages or across genes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.