Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
764
datasets available to search
ShareScore release 0.7.1
Dataset results
764 results for “Reproducible”
IrrigationMapV0, initial files and scripts to reproduce the simulation of Druel et al, 2022
<p>DATA supplementary to the corresponding article published in Geosci. Model Dev.: Druel, A., Munier, S., Mucia, A., Albergel, C., and Calvet, J.-C.: Implementation of a new crop phenology and irrigation scheme in<br> the ISBA land surface model using SURFEX_v8.1, 2022.</p> <p>==> Start by read the "README.txt" file, it explains how to use all files and be able to reproduce the simulations describe in the article.</p> <p>==> The map of irrigation can be downloaded and used for other applications. In this case, please refer to the article above.</p> <p><strong>Liste of files:</strong></p> <p>irrigcover_v0.tar.gz: Irrigation map 300 m x 300 m used for irrigation in ECOCLIMAP-SG. From a spatial resampling from the 1 km x 1km map of Meier et al, 2017.</p> <p>PGD.nc contains the initial characteristics of the simulation, such as the fraction of irrigation. It's computed at the beginning of the simulation and ensures the accuracy of the initial configuration.</p> <p>PREP.nc contains the result of the 20 year spinup (by repeating year 1979). It is used as initial condition of the published simulation.</p> <p>forcingScript.tar.gz contains scripts to help to produce forcing files (code made for LDAS launch is used, https://www.umr-cnrm.fr/spip.php?article1022)</p> <p>runLauncherFiles.tar.gz contains scripts and specific config files to be able to launch the simulations of the article with the exact same configurations.</p> <p> </p> <p> </p>
Code to reproduce the figures in the paper 'Listener Preference for Different Reproduction Systems and Mixes in Popular Music'
<p>In this upload you find all the scripts and data you need in order to reproduce<br> the figure from the paper Wierstorf et al., "Listener Preference for Different<br> Reproduction Systems and Mixes in Popular Music" [1].</p> <p>Software Requirements<br> ---------------------</p> <p>For the statistic analysis you will need [python](https://www.python.org) and<br> [R](https://www.r-project.org). I have used python 3.5.2 and R 3.2.3 for<br> published analysis.</p> <p>Under R you need to install the [eba](https://cran.r-project.org/package=eba)<br> package, which implements the Bradley-Terry-Luce model. You can install it in R<br> by running `install.packages("eba")`.</p> <p>Under python you have to install pandas and numpy.</p> <p>Reproduce figures<br> -----------------</p> <p>All figures were plotted using gnuplot 5.0. Every figure folder has an<br> ``figXX.plt`` (replace ``XX`` by the figure number) file that you can execute<br> and you will get the resulting pdf file. For Fig. 5 up to Fig. 9, also a<br> ``figXX.sh`` file is provided, that will rerun the statistical analysis of the<br> data presented in the figures.</p> <p>References<br> ----------</p> <p>[1] H. Wierstorf, C. Hold, A. Raake, "Listener Preference for Different<br> Reproduction Systems and Mixes in Popular Music," J. Audio. Eng. Soc, submitted. <br> </p>
Supporting code and data to reproduce analysis for: Genomic signatures of past megafrugivore-mediated dispersal in Malagasy palms
<p>Seed dispersal affects gene flow and hence genetic differentiation of plant populations. During the Late Quaternary, most fruit-eating and seed-dispersing megafauna went extinct, but whether these animals have left signatures in the population genetics of their food plants, particularly those with large, 'megafaunal' fruits (i.e., > 4 cm – megafruits), remains unclear.</p> <p>Here, we assessed the population history, genetic differentiation, and recent migration among populations of four animal-dispersed palm (Arecaceae) species with large (<em>Borassus madagascariensis</em>), medium-sized (<em>Hyphaene coriacea,</em> <em>Bismarckia nobilis</em>), and small (<em>Chrysalidocarpus madagascariensis</em>) fruits on Madagascar. We integrated double-digest restriction-site-associated DNA sequencing (ddRAD) of 167 individuals from 25 populations with (past) distribution ranges for extinct and extant seed-dispersing animals (e.g., giant lemurs, elephant birds), landscape and human impact data, and applied linear mixed-effects models to explore the drivers of genetic variation in Malagasy palms.</p> <p>Palm populations that shared more megafrugivore species in the past had lower genetic differentiation than populations that shared fewer megafrugivore species. This suggests that megafrugivore-mediated seed dispersal in the past may have led to frequent gene flow among populations. In comparison, extant frugivore diversity only decreased genetic differentiation in the small-fruited palm. Furthermore, genetic differentiation decreased with landscape connectivity (i.e., environmental suitability, forest cover and river density), and human impact (i.e., road density) has decreased genetic differentiation among populations.</p> <p><em>Synthesis: </em>Our results suggest that the legacy of megafrugivores regularly achieving long dispersal distances is still reflected in the population genetics of palms that were formerly dispersed by such animals. Furthermore, low genetic differentiation was possibly maintained after the megafauna extinctions through alternative dispersal (e.g., human- or river-mediated), long generation times, and long lifespans of these megafruit palms. Our study illustrates how species interactions that happened >1000 years ago can leave imprints in population genetics.</p>
Comparability and Reproducibility in HPC Applications' Energy Consumption Characterization
<p>The computational power of HPC systems continues to grow, and improving their energy efficiency is a critical issue for the field in the face of climate change and energy crises. One major aspect of energy optimization lies in the applications run on the systems themselves. In this work, we are looking into comparing energy consumption between different systems using a characterization process based on the recent energy characterization paper as a reference and starting point for other data centers to assess their application’s energy patterns. We demonstrated that we could use the methods from the starting paper, replicate the findings, and extend the work to more applications and more systems. Our work acts as a proof of concept for a repository of HPC applications’ energy patterns in our future work.<br><br>This is the collection of jobscripts, data, and python scripts used in the paper.</p>
STATA code to reproduce results in the manuscript "Low birth weight risk during COVID-19: Evidence from a nationwide study in India"
<p>This STATA code will reproduce results in the manuscript "Low birth weight risk during COVID-19: Evidence from a nationwide study in India" The users will have to register and access the data from www.dhsprogram.com to run the analysis code. </p>
Source Code Archiving to the Rescue of Reproducible Deployment — Replication Package
<p>Replication package for the paper:</p> <p>Ludovic Courtès, Timothy Sample, Simon Tournier, Stefano Zacchiroli.<br><em>Source Code Archiving to the Rescue of Reproducible Deployment</em><br><a href="https://acm-rep.github.io/2024/">ACM REP'24</a>, June 18-20, 2024, Rennes, France<br><a href="https://doi.org/10.1145/3641525.3663622">https://doi.org/10.1145/3641525.3663622</a></p> <h2>Generating the paper</h2> <p>The paper can be generated using the following command:</p> <pre><code>guix time-machine -C channels.scm \ -- shell -C -m manifest.scm \ -- make </code></pre> <p>This uses GNU Guix to run <code>make</code> in the exact same computational environment used when preparing the paper. The computational environment is described by two files. The <code>channels.scm</code> file specifies the exact version of the Guix package collection to use. The <code>manifest.scm</code> file selects a subset of those packages to include in the environment.</p> <p>It may be possible to generate the paper without Guix. To do so, you will need the following software (on top of a Unix-like environment):</p> <ul> <li>GNU Make</li> <li>SQLite 3</li> <li>GNU AWK</li> <li>Rubber</li> <li>Graphviz</li> <li>TeXLive</li> </ul> <h2>Structure</h2> <ul> <li><code>data/</code> contains the data examined in the paper</li> <li><code>scripts/</code> contains dedicated code for the paper</li> <li><code>logs/</code> contains logs generated during certain computations</li> </ul> <h2>Preservation of Guix</h2> <p>Some of the claims in the paper come from analyzing the Preservation of Guix (PoG) database as published on January 26, 2024. This database is the result of years of monitoring the extent to which the source code referenced by Guix packages is archived. This monitoring has been carried out by Timothy Sample who occasionally publishes reports on his personal website: <a href="https://ngyro.com/pog-reports/latest/">https://ngyro.com/pog-reports/latest/</a>. The database included in this package (<code>data/pog.sql</code>) was downloaded from <a href="https://ngyro.com/pog-reports/2024-01-26/pog.db">https://ngyro.com/pog-reports/2024-01-26/pog.db</a> and then exported to SQL format. In addition to the SQL file, the database schema is also included in this package as <code>data/schema.sql</code>.</p> <p>The database itself is largely the result of scripts, but also of manual adjustments (where necessary or convenient). The scripts are available at <a href="https://git.ngyro.com/preservation-of-guix/">https://git.ngyro.com/preservation-of-guix/</a>, which is preserved in the Software Heritage archive as well: <a href="https://archive.softwareheritage.org/swh:1:snp:efba3456a4aff0bc25b271e128aa8340ae2bc816;origin=https://git.ngyro.com/preservation-of-guix">https://archive.softwareheritage.org/swh:1:snp:efba3456a4aff0bc25b271e128aa8340ae2bc816;origin=https://git.ngyro.com/preservation-of-guix</a>. These scripts rely on the availability of source code in certain locations on the Internet, and therefore will not yield exactly the same result when run again.</p> <h3>Analysis</h3> <p>Here is an overview of how we use the PoG database in the paper. The exact way it is queried to produce graphs and tables for the paper is laid out in the Makefile.</p> <p>The <code>pog-types.sql</code> query gives the counts of each source type (e.g. “git” or “tar-gz”) for each commit covered by the database.</p> <p>The <code>pog-status.sql</code> query gives the archival status of the sources by commit. For each commit, it produces a count of how many sources are <em>stored</em> in the Software Heritage archive, <em>missing</em> from it, or <em>unknown</em> if stored or missing. The <code>pog-status-total.sql</code> query does the same thing but over all sources without sorting them into individual commits.</p> <p>The <code>disarchive-ratio.sql</code> query estimates the success rate of Disarchive disassembly.</p> <p>Finally, the <code>swhid-ratio.sql</code> query gives the proportion of sources for which the PoG database has an SWHID.</p> <h3>Estimating missing sources</h3> <p>The Preservation of Guix database only covers sources from a sample of commits to the Guix repository. This greatly simplifies the process of collecting the sources at the risk of missing a few. We estimate how many are missed by searching Guix’s Git history for Nix-style base-32 hashes. The result of this search is compared to the hashes in the PoG database.</p> <p>A naïve search of Git history results in an over estimate due to Guix’s branch development model. We find hashes that were never exposed to users of ‘guix pull’. To work around this, we also approximate the history of commits available to ‘guix pull’. We do this by scraping push events from the guix-commits mailing list archives (<code>data/guix-commits.mbox</code>). Unfortunately, those archives are not quite complete. Missing history is reconstructed in the <code>data/missing-links.txt</code> file.</p> <p>This estimate requires a copy of the Guix Git repository (not included in this package). The repository can be obtained from GNU at <a href="https://git.savannah.gnu.org/git/guix.git">https://git.savannah.gnu.org/git/guix.git</a> or from the Software Heritage archive: <a href="https://archive.softwareheritage.org/swh:1:snp:9d7b8dcf5625c17e42d51357848baa226b70e4bb;origin=https://git.savannah.gnu.org/git/guix.git">https://archive.softwareheritage.org/swh:1:snp:9d7b8dcf5625c17e42d51357848baa226b70e4bb;origin=https://git.savannah.gnu.org/git/guix.git</a>. Once obtained, its location must be specified in the Makefile.</p> <p>To generate the estimate, use:</p> <pre><code>guix time-machine -C channels.scm \ -- shell -C -m manifest.scm \ -- make data/missing-sources.txt </code></pre> <p>If not using Guix, you will need additional software beyond what is used to generate the paper:</p> <ul> <li>GNU Guile</li> <li>GNU Bash</li> <li>GNU Mailutils</li> <li>GNU Parallel</li> </ul> <h2>Measuring link rot</h2> <p>In order to measure link rot, we ran Guix Scheme scripts, i.e., scripts that exploit Guix as a Scheme library. The scripts depend on the state of world at the very specific moment when they ran. Hence, it is not possible to reproduce the exact same outputs. However, their tendency over the passing of time should be very similar. For running them, you need an installation of <a href="https://guix.gnu.org/manual/deve/en/html_node/Installation.html">Guix</a>. For instance,</p> <pre><code>guix repl -q scripts/table-per-origin.scm </code></pre> <p>When running these scripts for the paper, we tracked their output and saved it inside the <code>logs</code> directory.</p>
Data and Reproducible Analysis For: "Fine-Scale Associations Between Land Cover Composition and the Oviposition Activity of Native and Invasive Aedes Vectors of La Crosse Virus"
<h1><strong>Data and Reproducible Analysis For: "Fine-Scale Associations Between Land Cover Composition and the Oviposition Activity of Native and Invasive Aedes Vectors of La Crosse Virus"</strong></h1> <p>This repository contains pre-processed data sets and code scripts to reproduce the data processing and analyses that are presented in the corresponding manuscript. Some minor pre-processing was completed before presenting this -- namely, the land cover raster was clipped to the study area of Knox County, Tennessee, USA, prior to placing in the repository to reduce the file size. </p> <h2><strong>How to use this repository to reproduce results </strong></h2> <p>This repository is designed to support the reproduction of analyses in the associated manuscript. The entire project can be downloaded and stored anywhere on your computer, as long as the file structure is not altered. The project contains folders with all data sets and code scripts necessary for analysis.</p> <p><strong>What you will need: </strong><br> - Installed R and RStudio for purely spatial cluster and global model analyses<br> - Basic understanding of how to open R and run code </p> <p><strong> You do NOT need:</strong><br> - To download or install R packages on your own; that is taken care of within this environment<br> - To write any code <br> - To set up any working directories in R </p> <h3><strong>Important: Using `renv`</strong></h3> <p>Short Version: When you open the R project, run `renv::restore()` and follow the prompts to install the necessary R packages. </p> <p>The R package `renv` was used to create a <strong>project library</strong>, which contains all R packages that are used by the project. The packages in the project library are <strong>the versions used during the original analysis</strong>. This means that if any packages are updated by developers in ways that would change the results of the analysis, this project can still produce the original results because of `renv`. When you open this project for the first time, `renv` will automatically download and install itself and ask you to run `renv::restore()`. <strong>You should run `renv::restore()` to automatically download and install all of the packages within this reproducible environment</strong>. </p> <h2><strong>## Basic step-by-step guide:</strong></h2> <p>- 1. Download the entire repository by clicking "Code -> Download ZIP" on GitHub or by downloading the ZIP file in Zenodo<br>- 2. Extract the ZIP file anywhere on your computer (do not change the structure of the files once extracted)<br>- 3. In RStudio, click *File -> Open Project* and browse to the location where you extracted the repository; in the repository file, open the knoxaedeslandcover R Project file <br>- 4. Open any of the R scripts in the `analysis/` folder<br>- 5. Run the code `renv::restore()` in the script or in the console and follow the prompt to install the packages <br> - Now you can run the R Scripts; start from the top with loading the packages and data, then work your way down line-by-line</p> <h3><strong># `analysis/` Folder</strong></h3> <p>The `analysis/` folder contains scripts for processing data and conducting analyses. Each file is an R script that should be opened in R studio. The first shows how to process and aggregate the various raw data files; if you are only interested in reproducing analyses from the manuscript, you can skip to the second file and work from there. </p> <p><strong><em>## Files within the `analysis/` folder</em></strong></p> <p>The files are numbered in the order that they were run for the original analysis. In this case, none of the analyses are dependent on the others, so they can technically be used in any order. The numbers associated with each file describe the order that the analyses would normally be run. </p> <p> - `(1)dataprep.R` contains the code for cleaning and combining the land cover, climate, and mosquito data -- this includes calculating the land cover percentages at different scales and calculating weekly and timelagged climate values<br> - `(2)summary_analysis.R` contains code for reproducing summary data and creating graphs from the manuscript<br> - `(3)variable_selection.R` contains code for asssessing collinearity and fitting models to identify the best fitting variables for each speceis<br> - `(4)finalmodels.R` contains code for fitting the final models using the selected variables for each species </p> <h3><strong># `data/` Folder</strong></h3> <p>This folder contains several datasets, including one that compiles them all for analyses (`knox_joined`). The raw data are included to show how the data was processed and aggregated, but the individual raw data files are not needed for analyses. See `data dictionary.txt` for a description of all attributes contained within each file. </p> <p><strong><em>## Files within the `data/` folder</em></strong></p> <p> - `knox22_joined.RDS` contains a cleaned and joined version of land cover, climate, and mosquito data in R Data Serialization format, which maintains predefined factor and numeric designations for columns. <br> - `knox22_joined.csv` contains a cleaned and joined version of land cover, climate, and mosquito data in CSV format -- identical to 'knox22_joined.RDS'<br> - `sites22.csv` contains the names, site codes, and coordinates of the study sites<br> - `aedes22_clean.csv` contains the raw mosquito collection data for the study without any climate or land cover information <br> - `NLCD_2019_landcover_clippedtoKnox.tif` contains the NLCD land cover data, already clipped to Knox County, TN, USA<br> - `knox22_temperature.csv` contains raw daily temperatures for the city of Knoxville in 2022<br> - `knox22_rainfall.csv` contains raw daily precipitation for the city of Knoxville watersheds in 2022<br> - `rainfall_stations.csv` contains the descriptions, approximate street addresses, and geographic coordinates for rainfall monitoring sites <br> - `data dictionary.txt` file that defines column names and other data attributes for every dataset </p> <h3><strong># `renv/` Folder</strong></h3> <p>The `renv/` folder contains bits and pieces needed for the `renv` package. Nothing should be altered in this folder. </p> <p> </p> <h2><strong>References for source data </strong></h2> <p> - Some of the data in this repository were originally obtained from open access sources. </p> <p> - Land cover data was obtained from the National Land Cover Database (NLCD) 2019 data product, specifically the "NLCD 2019 Land Cover (CONUS)" product. The original, unclipped raster can be freely downloaded here: https://www.mrlc.gov/data/nlcd-2019-land-cover-conus</p> <p> - Temperature data was downloaded from the United States National Oceanic and Atmospheric Administration (NOAA) weather station for Knoxville, Tennessee. The source data can be downloaded from this site: https://www.weather.gov/mrx/tysclimate</p> <p> - Rainfall data was obtained from the City of Knoxville rainfall data website, located here: https://www.knoxvilletn.gov/government/city_departments_offices/engineering/stormwater_engineering_division/rainfall_data</p> <p> - All mosquito collection data was collected directly by the manuscript authors</p>
Solutions for Reproducibility in Empirical Research: Virtual Machines, Containers, Environment Management Packages, and Cloud Platforms
<p>This image provides a comprehensive overview of various technologies and platforms used to enhance the reproducibility of empirical research. It is divided into several sections:</p> <ol> <li><strong>Virtual Machines (VMs): </strong>the left section of the image illustrates the architecture of VMs with Type 1 and Type 2 hypervisors. <br> - <em>Type 1 Hypervisor </em>runs directly on the hardware, providing high efficiency and performance. Examples include VMware ESXi, <strong>Microsoft Hyper-v</strong>, and Xen Project.<br> - <em>Type 2 Hypervisor</em> runs on an existing operating system, offering flexibility at the cost of some performance. Examples include <strong>Oracle VirtualBox</strong>, VMware Workstation, and Parallels.</li> <li><strong>Containers: </strong>the middle section of the image explains the containerization concept, which shares the host operating system's kernel, making containers more lightweight than VMs. Technologies like <strong>Docker</strong> and <strong>Kubernetes</strong> are shown as popular solutions for container orchestration.</li> <li><strong>Environment Management Packages: </strong>the top right section focuses on tools for managing software dependencies and environments. <strong>renv</strong> (for R) and <strong>Conda</strong> (for Python and other languages) are highlighted as key tools for creating reproducible research environments.</li> <li>Cloud Platforms: the bottom right section features various cloud-based platforms that facilitate reproducible research by providing scalable and shareable computational environments. Platforms include <strong>Google Colab</strong>, <strong>Posit Cloud</strong>, JupyterHub, <strong>Binder</strong>, Nextjournal, OpenShift, and <strong>Code Ocean</strong>.</li> </ol> <p>Together, these solutions provide a robust framework for ensuring that empirical research can be reliably reproduced and validated by others, addressing the challenges of dependency management, environment consistency, and computational resource availability.</p>
Data and code to reproduce: Host and parasite intervality in differentially human-modified habitats
<p>Data and code in:</p> <p>Llopis-Belenguer, Feijen, Morand, Chaisiri, Ribas and Jokela (2024) Host and parasite intervality in differentially human-modified habitats. Oikos. DOI: 10.1111/oik.10446</p>
Figure 2 in Bendima: a database for marine macro-invertebrate bycatch data designed to improve reproducibility in benthic ecology
Figure 2. – The proportion of the various phyla represented in the Bendima database, based on the number of single organisms or colonies.
Figure 1 in Bendima: a database for marine macro-invertebrate bycatch data designed to improve reproducibility in benthic ecology
Figure 1. – Images and samples collection; sorting of the caught organisms (A), full photographing (B), taxa identification/counting/measurement and storage in the Bendima database in the form of cropped images (C), conservation of representative sub-samples for taxonomy and DNA barcoding (D).
Figure 3 in Bendima: a database for marine macro-invertebrate bycatch data designed to improve reproducibility in benthic ecology
Figure 3. – The geographical coverage of the Bendima database; the location of the zoomed area is in provided in the inset world map; stations are aggregated into presence data according to a grid of cells of 1°; the letters and num- bers relate to the name of the French Economic Exclusive Zones (EEZ) and various geomorphic structures located in international waters: A: Kerguelen EEZ, B: Crozet EEZ, C: Saint-Paul et Amsterdam EEZ, 1: Del Cano ridge, 2: Elan Bank, 3: Ob et Lena Bank, 4 and 5: Antarctica shelf.
Reproducibility package for: Badges for Geoscience Containers
<p>This repository contains demo data and software for research in using badges to improve discovery for scholarly publications in the geosciences. The web service "badger" provides an HTTP API for retrieving several types of badges based on a publications digital object identifier (DOI). The "extender" client integrates with the web browser and enhances existing platforms for discovery of publications with badges independent of the platforms code.</p> <p>Included in the archive are the source code of the required web services and client as well as ready to use Docker containers (including database and required web service with test data) to reproduce the results with minimal system requirements.</p> <p>Background on the project is provided in a peer-reviewed short paper from the AGILE conference 2019, <a href="https://doi.org/10.31223/osf.io/xtsqh">https://doi.org/10.31223/osf.io/xtsqh</a> , and in the blog post <a href="https://o2r.info/2017/09/12/reproducible-research-badges/">https://o2r.info/2017/09/12/reproducible-research-badges/</a>.</p> <p>For reproduction, follow the instructions in README.md and Makefile.and means to download the required software to enhance research websites with badges for reproducible research papers.</p>
Data for: A systematic review and meta-analysis of Drosophila short-term-memory genetics: robust reproducibility, but little independent replication
<p>All the data, code, analyses, and figures used in the study entitled: "A systematic review and meta-analysis of Drosophila short-term-memory genetics: robust reproducibility, but little independent replication" <em>(doi: https://doi.org/<a href="http://bb2sz3ek3z.search.serialssolutions.com/?url_ver=Z39.88-2004&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&__char_set=utf8&rft_id=info:doi/10.1101/247650&rfr_id=info:sid/libx&rft.genre=article">10.1101/247650</a>)</em></p> <p><strong>Abstract</strong></p> <p>Geneticists have long used olfactory conditioning techniques in <em>Drosophila</em> to identify the neurons and genes that mediate learning. While this method has characterized an abundance of memory-related genes, little is known about how these genes induce short-term memory (STM) via signaling pathways; characterizing these networks will be essential to developing mechanistic models of memory formation. Here, we investigated why elucidating the STM pathways has been relatively slow. One possibility is that the STM evidence base is weak due to publication of poorly reproducible results, as has been observed in other fields. We examined this hypothesis by performing a systematic review and subsequent meta-analysis of the STM genetics field. Using several metrics to quantify the variation between discovery articles and follow-up studies, we found that seven genes were highly replicated, showed no publication bias, and had generally high reproducibility. However, the remaining ~80% memory genes have not been replicated since their initial discovery. Although we observed only a few studies that investigated gene interactions, the reviewed genes could together account for >1000% memory. This large summed effect size indicates either that some of the gene findings are not reproducible, that many memory genes participate in shared pathways, or that current protocols lack the specificity needed to identify core plasticity memory genes. Mechanistic theories of memory and cognition will require the convergence of evidence from system, circuit, cellular, molecular, and genetic experiments. As this study demonstrates, systematic data synthesis is an essential tool for this integrated brain science.</p>
Reproducibility Package for "Reproducible research and GIScience: an evaluation using AGILE conference papers"
<p>Data and code for analysis and plots used in the manuscript "Reproducible research and GIScience: an evaluation using AGILE conference papers": <a href="https://doi.org/10.7287/peerj.preprints.26561v1">https://doi.org/10.7287/peerj.preprints.26561v1</a></p> <p>The deposited archived includes a <a href="https://en.wikipedia.org/wiki/Docker_(software)">Dockerfile</a> and an <a href="http://rmarkdown.rstudio.com/">R Markdown</a> document suitable for use with <a href="http://mybinder.org/">Binder</a>: <a href="https://mybinder.org/v2/gh/nuest/reproducible-research-and-giscience/6">https://mybinder.org/v2/gh/nuest/reproducible-research-and-giscience/6</a></p> <p>The version tag of this repository matches the <a href="https://git-scm.com/book/en/v2/Git-Basics-Tagging">git tag</a> on the code repository at <a href="https://github.com/nuest/reproducible-research-and-giscience">https://github.com/nuest/reproducible-research-and-giscience</a>, except version <code>6-fixed</code> which matches the tag <code>6</code>.</p> <p> </p>
Data to reproduce FAME results
<p>Synthetic WGBS data set with ground truth and processed EPIC bead data set to reproduce results obtained for FAME and other methods appearing in [coming soon].</p> <p>The original WGBS data complementing the EPIC bead calls are published by Pidsley et al. and should be downloaded using the IDs provided in the Online Appendix [coming soon].</p>
Data for figures in "Reproducibility in Benchmarking Parallel Fast Fourier Transform based Applications"
<p>FFT benchmark data and Python plotting programs</p>
The effect of numerical aperture on quantitative use-wear studies and its implication on reproducibility [complement to Supplementary Material 3]
<p>Raw data, and RStudio project, R markdown scripts and HTML outputs of the statistical procedures.</p> <p>Instructions to download all files at once are given here: <a href="https://doi.org/10.5281/zenodo.4011952">https://doi.org/10.5281/zenodo.4011952</a></p>
Pelvic arcade and vertebral structure of the second chief specimen of Tyrannosaurus rex, Amer. Mus. 5027, discovered in 1908. An orthogonal projection executed on a very large scale and reproduced one-twelfth natural size. C 1-C 10 cervical series, D 1-D 13 dorsal or thoracic series, S 1-S 5 sacral series, Cd 1- Cd 53 caudal series. The caudals actually preserved are shaded; those drawn in outline are conjectural and restored. The total number of caudals is conjectural. in Skeletal Adaptations of Ornitholestes, Struthiomimus, Tyrannosaurus
Pelvic arcade and vertebral structure of the second chief specimen of Tyrannosaurus rex, Amer. Mus. 5027, discovered in 1908. An orthogonal projection executed on a very large scale and reproduced one-twelfth natural size. C 1-C 10 cervical series, D 1-D 13 dorsal or thoracic series, S 1-S 5 sacral series, Cd 1- Cd 53 caudal series. The caudals actually preserved are shaded; those drawn in outline are conjectural and restored. The total number of caudals is conjectural.
Reproducibility package for the KDD 2019 paper "Pairwise Comparisons with Flexible Time-Dynamics"
<p>This archive contains the data necessary to reproduce most experiments and all plots presented in the following paper:</p> <p>Lucas Maystre, Victor Kristof, Matthias Grossglauser, <a href="https://arxiv.org/abs/1903.07746">Pairwise Comparisons with Flexible Time-dynamics</a>, KDD 2019.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.