Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “generalized codes”
Data and code from: Insect biomass decline scaled to species diversity: General patterns derived from a hoverfly community
<p>To study changes in flying insect communities, and hoverflies in particular, malaise trap samples from a German site were compared between two years (Hallmann et al. 2020). The data files deposited here contain data obtained from six malaise traps in the Wahnbachtal (North Rhine-Westphalia, Germany, 50.851944N, 7.320833E) that were deployed in 1989 and again in 2014, at the exact same locations. Traps were situated in wet meadows as well as tall perennial meadows, in close proximity to shrub corridors, to forest–grassland borders, and to the Wahnbach River and surrounded by agricultural land, essentially a rather heterogeneous habitat. The Wahnbach River and the greater part of the valley are protected for watershed purposes and are subject to nature conservation management by the Wahnbach Talperrenverband. Hence, several restrictions apply to safeguard against water contamination.</p> <p>Total insect biomass collected with these traps was already included in Hallmann et al. (2017), but here we focus on additional information: the abundance and richness of hoverflies (Syrphidae) in each of the collected samples (pots). Methodologies of collection are described in Sorg (1990), Schwan et al. (1993), Sorg et al. (2013), Hallmann et al. (2017), and Ssymank et al. (2018). In brief, malaise traps were deployed throughout the growing season and operated continuously (day and night). Malaise trap construction (e.g., size, material, colouring, and ground sealing) and placing (e.g., positioning, orientation, and slope of the locations) were standardised in all aspects. Insect samples were preserved in 80% ethanol solution. Catches of the six traps investigated in the present study were emptied regularly: On average exposure intervals were 7.0 d (SD = 0.5) in 1989 and 16.7 d (SD = 5.6) in 2014. Across the six traps in 2014 the total exposure time (in number of days) was 42% higher compared to 1989. All collected samples (n = 196) were used in the present analysis with in total 19,604 individual hoverflies counted, distributed over 162 species and 59 genera.</p> <p>To assess how environmental conditions have changed over the 25 year, several additional datasets were assembled. Climatic<br> data were obtained from 169 climatic stations and were used to interpolate daily weather variables to each trap location, using spatiotemporal kriging. These steps are described in detail in Hallmann et al. (2017).</p> <p>Our analysis (see R code) consists of three components. First, we considered total abundance, species richness, and species diversity, at two temporal scales: pooled per year, i.e., across the sampling season, and seasonally (i.e., per day), and we compared these metrics between 1989 and 2014. Second, we examined how total flying biomass (i.e., the weight of all trapped insects, of which hoverflies are only a small proportion) related to total abundance as well as species richness of hoverflies. Third, we derived persistence probabilities and population growth rate trends per species, to examine interspecific variation in these parameters.</p> <p>Descriptions of the deposited files:</p> <p><strong>Groups.csv</strong><br> MF_NR = identifier of each of the six malaise trap locations<br> yrf = year of sampling<br> pot = sample identifier<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> mean.daynr = mean day-of-the-year of the sampling period<br> Nspec = number of different hoverfly species found in a pot<br> Nind = number of hoverfly individuals found in a pot</p> <p><strong>Counts.csv</strong><br> A matrix of counts of individual hoverflies per pot per species. The 196 rows represent the pots in the same order as in the file 'Groups.csv'. The columns represent the 162 different hoverfly species found. The scientific species names are indicated in the column headers.</p> <p><strong>PairedData.csv</strong><br> pot = sample identifier<br> JAHR = year of sampling<br> MF_NR = identifier of each of the six malaise trap locations<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> NI = number of hoverfly individuals found in a potbiomass.daily<br> NSP = number of different hoverfly species found in a pot<br> biomass.daily = daily fresh weight [gram] of flying insects: total fresh weight in a pot divided by the number of sampling days.</p> <p><strong>ModelFrame.csv</strong><br> MF_NR = identifier of each of the six malaise trap locations<br> yrf = year of sampling<br> pot = sample identifier<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> mean.daynr = mean day-of-the-year of the sampling period<br> plot = identifier of each of the six malaise trap locations<br> date = date for which the weather variables are interpolated<br> daynr = day-of-the-year for which the weather variables are interpolated<br> altitude = altitude [m] of the malaise trap locations<br> year = year of sampling<br> temperature = interpolated temperature [degrees Celsius]<br> precipitation = interpolated precipitation [mm per day]<br> wind.speed = interpolated wind speed [m/s]</p> <p><strong>Data_Rcode.pdf</strong><br> This pdf provides the R-code behind the analysis of the Hoverfly data. Three datasets are provided along with this R-code document, namely "Counts.csv", "Groups.csv", "PairedData.csv" and "ModelFrame.csv". Additionally, the BUGS-code ""syrphidModel.jag" is required for running the daily-activity model in JAGS.</p> <p><strong>syrphidModel.jag</strong><br> This BUGS-code is required for running the daily-activity model in JAGS.</p>
Data and Code for 'Increased generalization in a peak procedure after delayed reinforcement'
<p>This dataset comes from the paper</p> <p>Buritica, J., & Alcala, E. (2019). Increased generalization in a peak procedure after delayed reinforcement. <em>Behavioural processes</em>, 169, 103978.</p> <p>The data comes in raw files and processed files. Commented scripts are also provided. For information about its structure see the Readme file.</p> <p><strong>License</strong></p> <p>CC-BY-4.0</p>
Data and original code for: A generalized approach to characterise optical properties of natural objects
<p>To understand the diversity of ways in which natural materials interact with light, it is important to consider how their reflectance changes with the angle of illumination or viewing and to consider wavelengths beyond the visible. We chose a set of existing measurements and parameters that are generalisable to any wavelength range and spectral shape and we highlight which subsets of measures are relevant to different biological questions. As a case study, we applied these measures to 30 species of Christmas beetles. Here we provide the raw spectral data of angle integrated and angle-dependent reflection by the beetle elytra. We also provide the original code used for our analysis and figures.</p>
Data and code: Evaluation of the General Practice Pharmacist (GPP) intervention to optimise prescribing in Irish primary care: a non‐randomised pilot study
<p>This is a dataset and Stata analytical code relating to prescribing issues identified in the GPP pilot feasibility study. The paper reporting this study has been published as follows: </p> <p>Cardwell K, Smith SM, Clyne B on behalf of the General Practice Pharmacist (GPP) Study Group, et al. Evaluation of the General Practice Pharmacist (GPP) intervention to optimise prescribing in Irish primary care: a non-randomised pilot study. BMJ Open 2020;10:e035087. doi: 10.1136/bmjopen-2019-035087</p> <p>The abstract of the study is included below:</p> <p><strong>Objective:</strong> Limited evidence suggests integration of pharmacists into the general practice team could improve medicines management for patients, particularly those with multimorbidity and polypharmacy. This study aimed to develop and assess the feasibility of an intervention involving pharmacists, working within general practices, to optimise prescribing in Ireland.</p> <p><strong>Design:</strong> Non-randomised pilot study</p> <p><strong>Setting:</strong> Primary care in Ireland</p> <p><strong>Participants:</strong> Four general practices, purposively sampled and recruited to reflect a range of practice sizes and demographic profiles.</p> <p><strong>Intervention:</strong> A pharmacist joined the practice team for six months (10 hours/week) and undertook medication reviews (face-to-face or chart-based) for adult patients, provided prescribing advice, supported clinical audits, and facilitated practice-based education.</p> <p><strong>Outcome measures:</strong> Anonymised practice-level medication (e.g. medication changes) and cost data were collected. Patient-Reported Outcome Measure (PROM) data were collected on a subset of older adults (aged ≥65 years) with polypharmacy using patient questionnaires, before and six weeks after medication review by the pharmacist.</p> <p><strong>Results:</strong> Across four practices, 787 patients were identified as having 1,521 prescribing issues by the pharmacists. Issues relating to potentially inappropriate or high-risk prescribing were addressed most often by the prescriber (51.8%), compared to cost-related issues (7.5%). Medication changes made during the study equated to approximately €57,000 in cost savings assuming they persisted for 12 months. Ninety-six patients aged ≥65 years with polypharmacy were recruited from the four practices for PROM data collection and 64 (66.7%) were followed up. There were no changes in patients’ treatment burden or attitudes to deprescribing following medication review, and there were conflicting changes in patients' self-reported quality of life.</p> <p><strong>Conclusions:</strong> This non-randomised pilot study demonstrated that an intervention involving pharmacists, working within general practices is feasible to implement and has potential to improve prescribing quality. This study provides rationale to conduct a randomised controlled trial to evaluate the clinical and cost-effectiveness of this intervention.</p>
Assets (code, scripts and datasets) for the manuscript "Correction of the Air-Sea Heat Fluxes in Ocean General Circulation Models Using Neural Networks"
<p>This dataset contains all relevant software and data related to the manuscript "Correction of the Air-Sea Heat Fluxes in Ocean General Circulation Models Using Neural Networks", submitted to AGU journals.</p>
Two-Streams Revisited: General Equations, Exact Coefficients, and Optimized Closures (data and code)
<p>Code and data necessary to reproduce figures presented in our paper "Two-Streams Revisited: General Equations, Exact Coefficients, and Optimized Closures" (Ho, and Pincus 2024). The figures and code to reproduce them are in the Jupyter Notebook `Two-stream_revisited_code.ipynb` which includes 3D interactive versions of Figures 6, 7, 10. We also provided code to further explore the dataset of coupling coefficients `twostreams_revisited_data.npz`.</p> <p>Identical GitHub repository: <a href="https://github.com/LDEO-CREW/Two-streams_revisited_code">https://github.com/LDEO-CREW/Two-streams_revisited_code</a>.</p> <p>This record is for the post-review (final) version of the paper, with a slight code change to accomodate `PythonicDISORT > 0.9.1`.</p> <p>You may contact me, Dion, through dh3065@columbia.edu if you have any questions.</p> <p>Requirements</p> <ul> <li>Python 3.8+</li> <li>`PythonicDISORT >= 0.8.0` (A minor code change will be required for `PythonicDISORT <= 0.9.1`)</li> <li>`numpy >= 1.8.0`</li> <li>`scipy >= 1.8.0`</li> <li>`matplotlib >= 3.6.0`</li> <li>`jupyter > 1.0.0`</li> <li>`notebook > 6.5.2`</li> <li>`autograd >= 1.5`</li> <li>`plotly >= 5.22.0`</li> </ul> <p> </p>
CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code Models
<p>The raw datasets and tasks of the paper "CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code Models". Source code is available at https://anonymous.4open.science/r/CrossCodeBench-C538/.</p>
Dataset and R Code for Species-level Avian Influenza Phylogenetic Generalized Least Squares Regression
<p>Dataset for Species-level Avian Influenza Phylogenetic Generalized Least Squares (PGLS) Regression:<br> Variables include taxonomic information for each species, # of IAV-positive individuals, # of IAV-tested individuals, the prevalence of IAV, the proportion of diet made up of different food types, the proportion of foraging time spent in different strata (below water, water surface, ground, understory, etc), sampling-related variables (mean latitude, mean date, the proportion of hatch year individuals), migration and territoriality category, climatologic variables, mean clutch size, and mating system.</p> <p>R Code for PGLS and Avian Influenza Prevalence ContMap.</p>
Data and original code for: A generalized approach to characterise optical properties of natural objects
Open the record for dataset details and reuse information.
Data and code associated with the publication "Emotional states elicited by wolf videos are diverse and explain general attitudes towards wolves", Arbieu et al., People and Nature 2024
<p>This folder contains the data and R scripts needed to replicate the analysis of the publication entitled "<span>Emotional states elicited by wolf videos are diverse and explain general attitudes towards wolves". <span>This dataset represents a social survey in rural populations of 24 randomly selected cities in France (n=795) to (i) quantify emotional diversity and (ii)<span> test the relationship between emotional states and attitudes towards wolves, accounting for individual and regional factors. <span>All </span>data were collected between November 2018 and May 2019. </span></span></span></p>
Code and Data for the Study "Exact Algorithms in Bar Nesting: How to Cut General Items from Linear Stocks so that Wastage is Minimised"
<p>This resource contains the code and results used in the paper:</p> <p>Lewis, R. and L. Bonnet (2025) '<a href="https://www.sciencedirect.com/science/article/pii/S0360835224009604" target="_blank" rel="noopener">Exact Algorithms in Bar Nesting: How to Cut General Items from Linear Stocks so that Wastage is Minimised</a>'. Computers & Industrial Engineering, vol. 200, 110838.</p> <p>The paper can be found <a href="https://www.sciencedirect.com/science/article/pii/S0360835224009604" target="_blank" rel="noopener">here</a>.</p> <p>Please consult <strong>UserGuide.pdf</strong> for further information. </p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
<p>Diet analysis integrates a wide variety of visual, chemical and biological identification of prey. Samples are often treated as compositional data, where each prey is analyzed as a continuous percentage of the total. However, analyzing compositional data results in analytical challenges, e.g., highly parameterized models or prior transformation of data. Here, we present a novel approximation involving a Tweedie generalized linear model (GLM). We first review how this approximation emerges from considering predator foraging as a thinned and marked point process (with marks representing prey species and individual prey size). This derivation can motivate future theoretical and applied developments. We then provide a practical tutorial for the Tweedie GLM using new package <i>mvtweedie</i> that extends capabilities of widely used packages in R (<i>mgcv</i> and <i>ggplot2</i>) by transforming output to calculate prey compositions. We demonstrate this approach and software using two examples. Tufted puffins (<i>Fratercula cirrhata</i>) provisioning their chicks on a colony in the northern Gulf of Alaska show decadal prey switching among sand lance and prowfish (1980-2000) and then Pacific herring and capelin (2000-2020), while wolves (<i>Canis lupus ligoni</i>) in Southeast Alaska forage on mountain goats and marmots in northern uplands and marine mammals in seaward island coastlines. </p>
Dynamic targeting enables domain-general inhibitory control over action and thought by the prefrontal cortex (data & code)
<p><strong>Data and code for:</strong></p> <p>Apšvalka, D., Ferreira, C. S., Schmitz, T. W., Rowe, J. B., & Anderson, M. C. (2022). Dynamic targeting enables domain-general inhibitory control over action and thought by the prefrontal cortex. <em>Nature Communications, </em> <strong>13, </strong>274<em>.</em> <a href="https://doi.org/10.1038/s41467-021-27926-w"> https://doi.org/10.1038/s41467-021-27926-w</a></p> <blockquote> <p>Over the last two decades, inhibitory control has featured prominently in accounts of how humans and other organisms regulate their behaviour and thought. Previous work on how the brain stops actions and thoughts, however, has emphasised distinct prefrontal regions supporting these functions, suggesting domain-specific mechanisms. Here we show that stopping actions and thoughts recruits common regions in the right dorsolateral and ventrolateral prefrontal cortex to suppress diverse content, via dynamic targeting. Within each region, classifiers trained to distinguish action-stopping from action-execution also identify when people are suppressing their thoughts (and vice versa). Effective connectivity analysis reveals that both prefrontal regions contribute to action and thought stopping by targeting the motor cortex or the hippocampus, depending on the goal, to suppress their task-specific activity. These findings support the existence of a domain-general system that underlies inhibitory control and establish Dynamic Targeting as a mechanism enabling this ability.</p> </blockquote>
Data and Code for "No general support of functional diversity enhancing resilience across terrestrial plant communities"
<p>The data and code provided here is to support the study "No general support of functional diversity enhancing resilience across terrestrial plant communities" </p> <p>This repository contains the following files:</p> <p> The code to reproduce main analysisi and graphs in R and HTML format</p> <ul> <li>RcodeNoGeneralSupport.R</li> <li>RcodeNoGeneralSupport.html</li> </ul> <p>The data to be used for the different analyses</p> <ul> <li>ResilienceFDIndices.csv</li> <li>ResilienceFDIndices-BiomassH.csv</li> <li>ResilienceFDIndices-BiomassW.csv</li> <li>ResilienceFDIndices-CompositionW.csv</li> <li>ResilienceFDIndices-Herbaceous.csv</li> <li>ResilienceFDIndices-woody.csv</li> <li>ResilienceSR.csv</li> <li>ResilienceSR-BiomassH.csv</li> <li>ResilienceSR-BiomassW.csv</li> <li>ResilienceSR-CompositionW.csv</li> <li>ResilienceSR-herbaceous.csv</li> <li>ResilienceSR-woody.csv</li> </ul> <p>The detailed information for each variable in each data set</p> <ul> <li>README.txt</li> </ul>
Quasi-Newton methods for partitioned simulation of fluid-structure interaction reviewed in the generalized Broyden framework: code and data
<p>These files accompany the publication</p><p>N. Delaissé, T. Demeester, R. Haelterman and J. Degroote. Quasi-Newton methods for partitioned simulation of fluid-structure interaction reviewed in the generalized Broyden framework.<i> Archives of Computational Methods in Engineering</i>, Vol.<strong> </strong>30, 3271-3300, 2023. doi: <a href="https://doi.org/10.1007/s11831-023-09907-y">10.1007/s11831-023-09907-y</a></p><p>In this work, the performance of multiple quasi-Newton methods are compared in terms of memory requirements and computational time. The results are generated for the well-known flexible tube example case, using the open-source code <a href="http://github.com/pyfsi/coconut">CoCoNuT</a>. This code, developed at Ghent University, is Python-based and has the capability to couple existing solvers, both open-source and commercial solvers.</p><p>This archive consists of the following files.</p><ul><li><strong>coconut.tar.gz: </strong>the specific CoCoNuT version used (sep-2022), including the Python flow and structure solvers for the flexible tube and modifications for monitoring memory requirements</li><li><strong>compare_coupling_algorithms.tar.gz: </strong>the scripts to set up the cases and perform the calculations and post-processing</li><li><strong>results.tar.gz:</strong> the generated result data</li></ul><p>For requirements to run CoCoNuT, refer to the <a href="http://pyfsi.github.io/coconut/">documentation</a>. Additionally, the Python package guppy3 is required for monitoring the memory use. In this work the data were generated with Andaconda3-2022.05 and the package guppy3-3.1.2.</p><p>Before running the provided scripts, make sure the parent directory of the "coconut" folder is added to the PYTHONPATH. The calculations can be started with "python run.py". For the cases which names contain "_m" followed by a number, e.g. "_m100", the number refers to the number of discretization points on the interface. The cases with suffix "_c" are distinct from those without, as they don't perform the time consuming memory monitoring and are therefore used for measuring computational time.</p>
Generalized LDPC codes for ultra reliable low latency communication in 5G and beyond
<p>Fifth-generation (5G) systems aim to increase the capacity of existing mobile networks by a factor of 1000, supporting an extremely high user density, as well as numerous device- to-device and machine communications. Ultra Reliable Low Latency Communication (URLLC) constitutes one of the critical operating regimes in 5G, since it will enable low-cost and power-efficient anywhere and anytime signalling services </p> <div> <div> <p>Generalized low-density parity-check (GLDPC) codes, where single parity-check constraints on the code bits are replaced with generalized constraints (an arbitrary linear code), are a promising class of codes for low-latency communication. We have constructed quasi-cyclic GLDPC codes, where the proportion of generalized constraints is determined by an asymptotic analysis. We have analyzed the complexity and performance of the message passing decoder with various update rules (including standard full-precision sum-product and min-sum algorithms) and quantization schemes for a GLDPC code over the additive white Gaussian noise (AWGN) channel and determined a constraint-to-variable update rule based on the specific codewords of the component codes. This data set includes the simulated GLDPC code constructions and the block error rate performance, which is shown to outperform a variety of state- of-the-art code and decoder designs with suitable lengths and rates for the 5G ultra-reliable low-latency communication regime over an AWGN channel with quadrature PSK modulation.</p> </div> </div>
Data and Code Supplement for "A Mountain-Induced Moist Baroclinic Wave Test Case for the Dynamical Cores of Atmospheric General Circulation Models"
<p>Code and Data Supplement for "A Mountain-Induced Moist Baroclinic Wave Test Case for the Dynamical Cores of Atmospheric General Circulation Models"<br> ===========================================================</p> <p>This directory contains the data and scripts used to create the plots from our publication as well as the source<br> code modifications necessary to run this test case within the CESM and MPAS models.</p> <p>Generating Plots<br> ---------------</p> <p>The `netcdf` directory contains the nominal half-degree runs necessary to generate nearly all of the plots from the paper. The one plot which is not reproducible from these data is the volume-integrated Eddy Kinetic Energy in the Spectral Element model. Storing high-resolution 4D wind fields requires a prohibitive amount of space. These data can be provided by the corresponding author, O.K. Hughes (owhughes@umich.edu). However, because this is several hundred GB of data I would strongly recommend generating these high-resolution runs yourself on your local system if you need them. Using 288 Intel Skylake cores (that is, 8 nodes each with two 18C processors) ran on the order of an hour.</p> <p><em>In order to generate the plots from the paper, you need only install NCL and then run</em> run.bash. Instructions for installing NCL<br> can be found in the `run.bash` script.</p> <p>Source Code Modifications<br> ----------------</p> <p><strong>CESM</strong><br> The `src` subdirectory contains the files `user_nl_cam` and `ic_baroclinic.F90`. Create a case using `--compset=FKESSLER` and `--run-unsupported` options when running `create_newcase`. If your case is located at `${CASE_DIR}`, then from within the directory containing this README, run `cp user_nl_cam ${CASE_DIR}/user_nl_cam`, and then run `cp ic_baroclinic.F90 ${CASE_DIR}/SourceMods/src.cam/`. Then build and run the model using the usual workflow.</p> <p><strong>MPAS</strong></p> <p>The MPAS code was run using a branch of the MPAS model provided by the model developers to the authors. While the source code modifications are provided in the `src` directory, I would strongly recommend contacting the corresponding author if you wish to run this test case in the MPAS codebase.</p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
Open the record for dataset details and reuse information.
Code from: Costs of parasite generalism revealed by abundance patterns across mammalian hosts
Open the record for dataset details and reuse information.
Generalized LDPC codes for ultra reliable low latency communication in 5G and beyond
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.