Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15,520
datasets available to search
ShareScore release 0.7.1
Dataset results
15,520 results for “R”
R Code and Images for Developing a Spatial Concordance Coefficient at Harvard Forest 2010
Concordance correlation coefficients have been developed in a variety of different contexts. This problem has been widely addressed in a non-spatial context, but here we consider a coefficient that for a fixed spatial lag allows the comparison of two spatial sequences (e.g., images). We define a spatial concordance coefficient for second-order stationary processes.
Evaluation of Mask R-CNN Model for Counting Reproductive Structures of Six Plant Species 1895-2018
Phenology––the timing of life-history events––is a key trait for understanding responses of organisms to climate. The digitization and online mobilization of herbarium specimens is rapidly advancing our understanding of plant phenological response to climate and climatic change. The current common practice of manually harvesting data from individual specimens greatly restricts our ability to scale data collection to entire collections. Recent investigations have demonstrated that machine-learning models can facilitate data collection from herbarium specimens. However, present attempts have focused largely on simplistic binary coding of reproductive phenology (e.g., flowering or not). Here, we use crowd-sourced phenological data of numbers of buds, flowers, and fruits of more than 3000 specimens of six common wildflower species of the eastern United States (Anemone canadensis, A. hepatica, A. quinquefolia, Trillium erectum, T. grandiflorum, and T. undulatum} to train a model using Mask R-CNN to segment and count phenological features. A single global model was able to automate the binary coding of reproductive stage with greater than 90% accuracy. Segmenting and counting features were also successful, but accuracy varied with phenological stage and taxon. Counting buds was significantly more accurate than flowers or fruits. Moreover, botanical experts provided more reliable data than either crowd-sourcers or our Mask R-CNN model, highlighting the importance of high-quality human training data. Finally, we also demonstrated the transferability of our model to automated phenophase detection and counting of the three Trillium species, which have large and conspicuously-shaped reproductive organs. These results highlight the promise of our two-phase crowd-sourcing and machine-learning pipeline to segment and count reproductive features of herbarium specimens, providing high-quality data with which to study responses of plants to ongoing climatic change.
AMOC reconstruction between 1981 and 2016 from hydrographic data using an empirical linear regression model from Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, https://doi.org/10.5194/os-17-285-2021, 2021.
<p>Dataset used to create Figure 8 in Worthington et al., 2021 (https://doi.org/10.5194/os-17-285-2021). Details of the data and methods can be found in the journal article.<br> <br> Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, <a href="https://doi.org/10.5194/os-17-285-2021">https://doi.org/10.5194/os-17-285-2021</a>, 2021.</p>
Software Tools to Collect and Use Provenance in R
The software tools that scientists use to process and analyze data are typically optimized for performance and ease of use. Few if any such tools are designed to capture and record the details of what happens as the tool performs its task. This detailed information, and more generally the history of an item of data from its creation to its present state, is known as provenance. Provenance has the potential to make science more transparent, reliable, and reproducible. This project focused on collecting and using provenance for scripts written in the R statistical language, which is widely used by ecologists and environmental scientists for data analysis and visualization. Our tools include a provenance collector (rdtLite), which collects provenance as an R script executes (or during a console session), as well as other tools that use the collected provenance to document and visualize the execution or to support activites such as script debugging. The R packages included here are also available on CRAN. For more details, see the project website on GitHub (https://end-to-end-provenance.github.io).
Datasets and R-scripts used for the revision of the Dibrachys cavus complex by Peters & Baur, 2011, Zootaxa 2937.1
<p>In 2011 we published a revision on the Dibrachys cavus complex in Zootaxa (Peters and Baur, 2011, here a link to our <a href="https://doi.org/10.11646/zootaxa.2937.1.1">open access paper</a>). It was our wish to also publish two versions of the dataset (one with missing values, one with missing values imputed) as supplementary files. Unfortunately, the data files seem to be no longer available on the publishers webpage. Hence, we publish the data files herewith again in CSV format.</p> <p>We take the opportunity to also publish the R-scripts that we used for calculating multivariate analyses, tests, and the multiple imputation of missing values.</p> <p>All files are available individually and with an own link. For convenience, we have compiled all files also in a ZIP file.</p> <p><a href="https://doi.org/10.5281/zenodo.4256704">Baur (2020)</a> used the dataset for further exploration in a Multivariate Ratio Analysis (MRA).</p> <p>Papers quoted above you may find in the section <em>References</em> of the Zenodo package.</p> <p><strong>Citation of this package</strong><br> Peters, Ralph S., & Baur, Hannes (2020, November 9) Datasets and R-scripts used for the revision of the Dibrachys cavus complex by Peters & Baur, 2011, Zootaxa 2937.1. Zenodo. https://doi.org/10.5281/zenodo.4264539 (directs to the newest version of the package).</p>
Detailed abundances based on different nuclear physics for theoretical r-process scenarios
<p>This data set contains detailed abundances (at a time t=10^6 years after the event) for individual trajectories for seven different simulations of potential r-process sites, and based on nine different combinations of nuclear mass models and fission fragment distribution models. The data have been used and are discussed in Cote, Eichler, Yagüe, et al. (https://ui.adsabs.harvard.edu/abs/2020arXiv200604833C/abstract) to determine the isotopic ratios of I129/Cm247 and compare them to meteoritic data.</p> <p>Furthermore, a code is included which samples a subset of trajectories reproducing the measured meteoritic I129/Cm247 abundance ratio of 438 +- 92. See the README file and the publication (https://ui.adsabs.harvard.edu/abs/2020arXiv200604833C/abstract) for more details.</p>
Data and R code for Tansley review New Phytologist 2021: "An integrated framework of plant form and function: The belowground perspective"
<p>The files in this archive are related to the paper of Weigelt, Mommer, Andraczek et al. (2021) An integrated framework of plant form and function: The belowground perspective. Tansley Review New Phytologist. The paper developed and tested a new conceptual framework of plant form and function linking above and belowground traits of 2510 species. We found that an integrated, whole-plant trait space required as much as four axes. The two main axes represented the fast-slow ‘conservation’ gradient on which leaf and fine-root traits were well aligned, and the ‘collaboration’ gradient in roots. The two additional axes were separate, orthogonal plant size axes for height and rooting depth.</p> <p>This archives contains four files:</p> <ol> <li><strong>Weigelt et al.2021RCode.DataCleaning.txt</strong> - RCode for the complete data processing starting with the downloaded database files from the Plant Trait Database version 5.0 (TRY, Kattge et al. 2020), the Global Root Trait database (GRooT, Guerrero-Ramirez et al. 2020) and a small number of additional data files listed in Table S2 of the original paper. Additional information was later incorporated using FungalRoot Database (Soudzilovkaia et al. 2020), nodDB Database (Tedersoo et al. 2018) and a compiled dataset on rooting depth (Fan et al. 2017). The code processes, cleans and merges the data and produces a final table for PCA analysis of species specific mean traits. This final table is provided as a second file in this archive (Weigelt_et_al_2021_Main.PCA.Matrix.xlsx). A second part of the RCode.DataCleaning extracts species-specific individual trait data where root and shoot traits were measured on the same plant individual or plot. This data was compiled from 43 studies identified in Table S2 of the original publication. The final table for individual trait data is the third file in this archive (Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx).</li> <li><strong>Weigelt_et_al_2021_Main.PCA.Matrix.xlsx</strong> – Datafile with species-specific global mean trait data for 17 traits of 2510 species with at least one root and one shoot trait available. Meta-data is provided in the data file.</li> <li><strong>Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx</strong> – Datafile with species-specific trait data where root and shoot traits were measured on the same individual or plot for 6 traits of 455 species. Meta-data is provided in the data file.</li> <li><strong>Weigelt et al.2021RCode.Analysis.txt – </strong>RCode for all analyses and figures provided in the paper for both the species mean and individual based dataset. The Code is annotated to help reproducibility of the analysis.</li> </ol>
Hail Event on 2022-06-28 in Locarno-Monti (TI), Switzerland: Drone Photogrammetry Imagery, Mask R-CNN Model and Analysis Data of Hailstones
<p>This hail data collection belongs to a drone hail survey performed on 2022-06-28 in Locarno-Monti (TI, Switzerland). The supercell reached the location around 07:50 UTC in the morning. Only one photogrammetry flight could be performed and thus no estimation of the hail melting process is available. The orthophoto is masked to ignore parts where detection of hail is unwanted.</p> <p> </p> <p>Expert 1 (lai, mlainer), Expert 2 (jtm), Expert 3 (por, jportmann)</p>
Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to R and the Tidyverse
<p>Data and conda software environment file for the chapter 'Introduction to R and the Tidyverse' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>
R-CAUSTIC: Rippling CAUSTICs underwater Image dataset
<p><strong>Description</strong></p><p>Rippling caustics seem to be the main factor degrading the underwater RGB image quality and affecting the image- based 3D reconstruction process in very shallow waters. These effects are adversely affecting image matching algorithms by throwing off most of them, leading to less accurate matches and causing issues in the Simultaneous Localization and Mapping (SLAM) based navigation of the Remotely Operated Vehicles (ROV) and Autonomous Underwater Vehicles (AUV) on shallow waters. Also, they are the main cause for dissimilarities in the generated textures and orthoimages. In order to fill the gap in the literature regading underwater rippling caustics imagery with real ground truth and reference images, the first real-world underwater caustics benchmark dataset which contains 1465 underwater images is presented. Together with the RGB imagery, the corresponding generated ground truth images are delivered for facilitating the training and testing of machine learning and deep learning methods for image classification. R-CAUSTIC dataset also provides the necessary data to evaluate, at least to some extent, the performance of 3D reconstruction approaches. Data were acquired using a GoPro Hero 4 Black action camera with image dimensions of 4000 x 3000 pixels, focal length of 2.77mm and pixel size of 1.55μm and a tripod. Action cameras are widely used for underwater image acquisition. The dataset was captured in near-shore underwater sites at depths varying from 0.5 to 2m. No artificial light sources were used. Due to the wind, the turbulent surface of the water created dynamic rippling caustics on the seabed. In total 1465 RGB images were collected, separated in 7 different datasets; five of them containing stereo images, one of them tri-stereo images and one consists of multi-stereo imagery acquired in 7 different camera poses.</p><p> </p><p><strong>Publication</strong></p><p>The paper is availbale in Open Access here: https://ieeexplore.ieee.org/document/10172291</p><p><strong>If you use this dataset please cite it as R-CAUSTIC</strong> [Reference].<br>[Reference]: <strong>P. Agrafiotis, K. Karantzalos and A. Georgopoulos, "Seafloor-Invariant Caustics Removal From Underwater Imagery," in </strong><i><strong>IEEE Journal of Oceanic Engineering</strong></i><strong>, vol. 48, no. 4, pp. 1300-1321, Oct. 2023, doi: 10.1109/JOE.2023.3277168.</strong></p><p>BibTeX:</p><p>@ARTICLE{10172291, author={Agrafiotis, Panagiotis and Karantzalos, Konstantinos and Georgopoulos, Andreas}, journal={IEEE Journal of Oceanic Engineering}, title={Seafloor-Invariant Caustics Removal From Underwater Imagery}, year={2023}, volume={48}, number={4}, pages={1300-1321}, doi={10.1109/JOE.2023.3277168}}</p><p> </p>
Supporting Information for 'forceX and forceR: a mobile setup and R package to measure and analyse a wide range of animal closing forces'
<p><strong>Supporting Information of 'forceX and forceR: a mobile setup and R package to measure and analyse a wide range of animal closing forces'</strong></p> <p>This dataset contains the Supporting Information of the publication </p> <p>Rühr PT & Blanke A <strong>(2022)</strong>: 'forceX and forceR: a mobile setup and R package to measure and analyse a wide range of animal closing forces'. doi: <a href="https://doi.org/10.1111/2041-210X.13909">10.1111/2041-210X.13909</a>.</p> <p>It includes</p> <ul> <li>validation measurements the forceX setups (1 Ruehr Blanke 2022 validation measurements.zip)</li> <li>all CAD files to build the forceX setup (3D-printed or metal-turned) (2 Ruehr Blanke 2022 forceX CAD files.zip)</li> <li>forceX assembly instructions in HTML format, including schematics of custom electronics (3 Ruehr Blanke 2022 forceX Assembly instructions.html)</li> <li>forceX assembly instructions as video (4 Ruehr Blanke 2022 forceX assembly video 03.mp4)</li> <li>R code that produced all validation-related figures used in the original publication and that functions as a forceR v.1.0.13 example workflow (5 Ruehr Blanke 2022 forceR_workflow_example.R)</li> <li>Python code to take videos of force measurements using the forceX camera module (6 Ruehr Blanke 2022 forceX_RPi_camera_code.py)</li> <li>bundled version of forceR v.1.0.15 (forceR_1.0.15.tar.gz)</li> </ul> <p>The CAD files and assembly instructions are also available on <a href="https://www.thingiverse.com/thing:4961834">Thingiverse</a>. The forceR package is available on <a href="https://cran.r-project.org/web/packages/forceR/index.html">CRAN</a> (stable version) and <a href="https://github.com/Peter-T-Ruehr/forceR">GitHub</a> (development version).</p>
Datasets and R codes for Prokkola et al. 2022 adipose tissue samples
<p>Data and R codes for the analyses reported in Prokkola et al. (pre-print, 2022) <em>Adipose tissue mitochondrial respiration in Atlantic salmon: implications for sex-dependent life-history variation.</em></p> <p>Overview of files can be found in the README file.</p> <p>The file "Cell size data.zip" contains TIFF-images of adipose tissue cryosections, a README file, the result files for each image file and an R code for parsing the results files.</p> <p>To skip the data parsing steps and get the final data, download the AdiposeTissue_data_all.txt file (tab-separated).</p>
Data, scripts, and R Notebook for Carneiro et al 2023. Flight performance and wing morphology in the bat Carollia perspicillata: biophysical models and energetics. Integrative Zoology DOI:10.1111/1749-4877.12707
<p>Files provided as supporting information for the paper by Carneiro et al. 2023. Flight performance and wing morphology in the bat <em>Carollia perspicillata</em>: biophysical models and energetics. Integrative Zoology. DOI:10.1111/1749-4877.12707</p> <p>File descriptions</p> <p>ArmTA.txt - Temperature and surface areas for arms of <em>C. perspicillata</em> after flight experiment<br> BodyTA.txt - Temperature and surface areas for body of <em>C. perspicillata</em> after flight experiment<br> HeadTA.txt - Temperature and surface areas for head of <em>C. perspicillata</em> after flight experiment<br> WingTA.txt - Temperature and surface areas for wings (patagium) of <em>C. perspicillata</em> after flight experiment<br> WingMorph.txt - Morphological variables measured in the body and wings of <em>C. perspicillata</em><br> HeatLoss.R - Function to estimate heat loss (Qt)<br> PowFlight.R - Function to estimate minimum power required to fly<br> Script-HeatLoss-FlightPerformance.R - R script with set of analyses performed<br> SupportingInformationFile.docx - R notebook with set of analyses performed, word format<br> SupportingInformationFile.nb.html - R notebook with set of analyses performed, html format<br> SupportingInformationFile.Rmd - R notebook with set of analyses performed (R markdown)</p> <p>For the R scripts (Script-HeatLoss-FlightPerformance.R) and notebook (<br> SupportingInformationFile.Rmd) to work and be compiled, all files need to be copied to the same folder.</p>
R script and data files for Oakley et al (2017) Journal of Proteome Research. DOI: 10.1021/acs.jproteome.6b00797
<p>This R script and data replicates the analysis of Oakley et al (2017) Thermal shock induces host proteostasis disruption and endoplasmic reticulum stress in the model symbiotic Cnidarian <em>Aiptasia</em>. <em>Journal of Proteome Research</em>. 16:2121-2134. DOI: 10.1021/acs.jproteome.6b00797. </p>
metapsyData: R Package to Access the Metapsy Databases
<p>The <code>metapsyData</code> package allows to access the Metapsy meta-analytic psychotherapy databases direct in your <code>R</code> environment. Once installed, simply run the <code>data</code> function (e.g. <code>data(DepPsychDB)</code>) to save the data locally. The documentation of the package is also hosted by <a href="https://rdrr.io/github/metapsy-project/metapsyData/">rdrr.io</a>.</p> <p>The interactive Metapsy web application (<a href="https://www.metapsy.org/">metapsy.org</a>) uses <code>metapsyData</code> in the background. You can open the Metapsy website in <code>R</code> by running <code>open_app()</code>.</p> <p>The raw data files can be accessed in the associated GitHub repository under <code>data</code>. To search for available databases in <code>metapsyData</code>, type in <code>metapsyData::</code> in your RStudio console.</p>
Dataset and R script for the analysis in the article "Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy", Journal of Cleaner Production
<p>We hereby publish the dataset (with metadata) and the R script (R Core team, 2018) used for implementing the analysis presented in the paper "Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy", <em>Journal of Cleaner Production </em>(Piras et al., 2023). The dataset is provided in csv format with semicolons as separators and "NA" for missing data. The dataset includes all the variables used in at least one of the models presented in the paper, either in the main text or in the Supplementary Material. Other variables gathered by means of the questionnaires included as Supplementary Material of the paper have been removed. The dataset includes inputted values for missing data on independent variables. These were inputted using two approaches: last observation carried forward (LOCF) - preferred when possible - and last observation carried backward (LOCB). The metadata are presented as a PDF file.</p>
Example data set for the R package riversCentralAsia
<p>This data set contains example data for demonstrating the functionality of the R package riversCentralAsia. riversCentralAsia (https://github.com/hydrosolutions/riversCentralAsia) includes several functions for pre-processing hydrological data to facilitate hydrological modelling with RS MINERVE (https://crealp.github.io/rsminerve-releases/). The package is used extensively in the open-source teaching course Modeling of Hydrological Systems in Semi-Arid Central Asia (https://hydrosolutions.github.io/caham_book/). </p>
Datasets and R source code of manuscript "Parasites make hosts more profitable but less available to predators"
<p>Data about experimentations of DIV-1 (virus) infection on Daphnia magna.</p> <p>Linked article: Parasites make hosts more profitable but less available to predators</p>
HRV-ACC: a dataset with R-R intervals and accelerometer data for the diagnosis of psychotic disorders using a Polar H10 wearable sensor
<p><strong>ABSTRACT</strong></p> <p>The issue of diagnosing psychotic diseases, including schizophrenia and bipolar disorder, in particular, the objectification of symptom severity assessment, is still a problem requiring the attention of researchers. Two measures that can be helpful in patient diagnosis are heart rate variability calculated based on electrocardiographic signal and accelerometer mobility data. The following dataset contains data from 30 psychiatric ward patients having schizophrenia or bipolar disorder and 30 healthy persons. The duration of the measurements for individuals was usually between 1.5 and 2 hours. R-R intervals necessary for heart rate variability calculation were collected simultaneously with accelerometer data using a wearable Polar H10 device. The Positive and Negative Syndrome Scale (PANSS) test was performed for each patient participating in the experiment, and its results were attached to the dataset. Furthermore, the code for loading and preprocessing data, as well as for statistical analysis, was included on the corresponding GitHub repository.</p> <p><strong>BACKGROUND</strong></p> <p>Heart rate variability (HRV), calculated based on electrocardiographic (ECG) recordings of R-R intervals stemming from the heart's electrical activity, may be used as a biomarker of mental illnesses, including schizophrenia and bipolar disorder (BD) [Benjamin et al]. The variations of R-R interval values correspond to the heart's autonomic regulation changes [Berntson et al, Stogios et al]. Moreover, the HRV measure reflects the activity of the sympathetic and parasympathetic parts of the autonomous nervous system (ANS) [Task Force of the European Society of Cardiology the North American Society of Pacing Electrophysiology, Matusik et al]. Patients with psychotic mental disorders show a tendency for a change in the centrally regulated ANS balance in the direction of less dynamic changes in the ANS activity in response to different environmental conditions [Stogios et al]. Larger sympathetic activity relative to the parasympathetic one leads to lower HRV, while, on the other hand, higher parasympathetic activity translates to higher HRV. This loss of dynamic response may be an indicator of mental health. Additional benefits may come from measuring the daily activity of patients using accelerometry. This may be used to register periods of physical activity and inactivity or withdrawal for further correlation with HRV values recorded at the same time.</p> <p><strong>EXPERIMENTS</strong></p> <p>In our experiment, the participants were 30 psychiatric ward patients with schizophrenia or BD and 30 healthy people. All measurements were performed using a Polar H10 wearable device. The sensor collects ECG recordings and accelerometer data and, additionally, prepares a detection of R wave peaks. Participants of the experiment had to wear the sensor for a given time. Basically, it was between 1.5 and 2 hours, but the shortest recording was 70 minutes. During this time, evaluated persons could perform any activity a few minutes after starting the measurement. Participants were encouraged to undertake physical activity and, more specifically, to take a walk. Due to patients being in the medical ward, they received instruction to take a walk in the corridors at the beginning of the experiment. They were to repeat the walk 30 minutes and 1 hour after the first walk. The subsequent walks were to be slightly longer (about 3, 5 and 7 minutes, respectively). We did not remind or supervise the command during the experiment, both in the treatment and the control group. Seven persons from the control group did not receive this order and their measurements correspond to freely selected activities with rest periods but at least three of them performed physical activities during this time. Nevertheless, at the start of the experiment, all participants were requested to rest in a sitting position for 5 minutes. Moreover, for each patient, the disease severity was assessed using the PANSS test and its scores are attached to the dataset.</p> <p>The data from sensors were collected using Polar Sensor Logger application [Happonen]. Such extracted measurements were then preprocessed and analyzed using the code prepared by the authors of the experiment. It is publicly available on the GitHub repository [Książek et al].</p> <p>Firstly, we performed a manual artifact detection to remove abnormal heartbeats due to non-sinus beats and technical issues of the device (e.g. temporary disconnections and inappropriate electrode readings). We also performed anomaly detection using Daubechies wavelet transform. Nevertheless, the dataset includes raw data, while a full code necessary to reproduce our anomaly detection approach is available in the repository. Optionally, it is also possible to perform cubic spline data interpolation. After that step, rolling windows of a particular size and time intervals between them are created. Then, a statistical analysis is prepared, e.g. mean HRV calculation using the RMSSD (Root Mean Square of Successive Differences) approach, measuring a relationship between mean HRV and PANSS scores, mobility coefficient calculation based on accelerometer data and verification of dependencies between HRV and mobility scores.</p> <p><strong>DATA DESCRIPTION</strong></p> <p>The structure of the dataset is as follows. One folder, called <em>HRV_anonymized_data</em> contains values of R-R intervals together with timestamps for each experiment participant. The data was properly anonymized, i.e. the day of the measurement was removed to prevent person identification. Files concerned with patients have the name <em>treatment_X.csv</em>, where <em>X</em> is the number of the person, while files related to the healthy controls are named <em>control_Y.csv</em>, where <em>Y</em> is the identification number of the person. Furthermore, for visualization purposes, an image of the raw RR intervals for each participant is presented. Its name is <em>raw_RR_{control,treatment}_N.png</em>, where <em>N</em> is the number of the person from the control/treatment group. The collected data are raw, i.e. before the anomaly removal. The code enabling reproducing the anomaly detection stage and removing suspicious heartbeats is publicly available in the repository [Książek et al]. The structure of consecutive files collecting R-R intervals is following:</p> <table> <tbody> <tr> <td><strong>Phone timestamp</strong></td> <td><strong>RR-interval [ms]</strong></td> </tr> <tr> <td>12:43:26.538000</td> <td>651</td> </tr> <tr> <td>12:43:27.189000</td> <td>632</td> </tr> <tr> <td>12:43:27.821000</td> <td>618</td> </tr> <tr> <td>12:43:28.439000</td> <td>621</td> </tr> <tr> <td>12:43:29.060000</td> <td>661</td> </tr> <tr> <td>...</td> <td>...</td> </tr> </tbody> </table> <p>The first column contains the timestamp for which the distance between two consecutive R peaks was registered. The corresponding R-R interval is presented in the second column of the file and is expressed in milliseconds. <br> The second folder, called <em>accelerometer_anonymized_data</em> contains values of accelerometer data collected at the same time as R-R intervals. The naming convention is similar to that of the R-R interval data: <em>treatment_X.csv </em>and <em>control_X.csv</em> represent the data coming from the persons from the treatment and control group, respectively, while <em>X </em>is the identification number of the selected participant. The numbers are exactly the same as for R-R intervals. The structure of the files with accelerometer recordings is as follows:</p> <table> <tbody> <tr> <td><strong>Phone timestamp</strong></td> <td><strong>X [mg]</strong></td> <td><strong>Y [mg]</strong></td> <td><strong>Z [mg]</strong></td> </tr> <tr> <td>13:00:17.196000</td> <td>-961</td> <td>-23</td> <td>182</td> </tr> <tr> <td>13:00:17.205000</td> <td>-965</td> <td>-21</td> <td>181</td> </tr> <tr> <td>13:00:17.215000</td> <td>-966</td> <td>-22</td> <td>187</td> </tr> <tr> <td>13:00:17.225000</td> <td>-967</td> <td>-26</td> <td>193</td> </tr> <tr> <td>13:00:17.235000</td> <td>-965</td> <td>-27</td> <td>191</td> </tr> <tr> <td>...</td> <td>...</td> <td>...</td> <td>...</td> </tr> </tbody> </table> <p>The first column contains a timestamp, while the next three columns correspond to the currently registered acceleration in three axes: X, Y and Z, in milli-g unit.</p> <p>We also attached a file with the PANSS test scores (<em>PANSS.csv</em>) for all patients participating in the measurement. The structure of this file is as follows:</p> <table> <tbody> <tr> <td><strong>no_of_person</strong></td> <td><strong>PANSS_P</strong></td> <td><strong>PANSS_N</strong></td> <td><strong>PANSS_G</strong></td> <td><strong>PANSS_total</strong></td> </tr> <tr> <td>1</td> <td>8</td> <td>13</td> <td>22</td> <td>43</td> </tr> <tr> <td>2</td> <td>11</td> <td>7</td> <td>18</td> <td>36</td> </tr> <tr> <td>3</td> <td>14</td> <td>30</td> <td>44</td> <td>88</td> </tr> <tr> <td>4</td> <td>18</td> <td>13</td> <td>27</td> <td>58</td> </tr> <tr> <td>...</td> <td>...</td> <td>...</td> <td>...</td> <td>..</td> </tr> </tbody> </table> <p><br> The first column contains the identification number of the patient, while the three following columns refer to the PANSS scores related to positive, negative and general symptoms, respectively.</p> <p><strong>USAGE NOTES</strong></p> <p>All the files necessary to run the HRV and/or accelerometer data analysis are available on the GitHub repository [Książek et al]. HRV data loading, preprocessing (i.e. anomaly detection and removal), as well as the calculation of mean HRV values in terms of the RMSSD, is performed in the <em>main.py</em> file. Also, Pearson's correlation coefficients between HRV values and PANSS scores and the statistical tests (Levene's and Mann-Whitney U tests) comparing the treatment and control groups are computed. By default, a sensitivity analysis is made, i.e. running the full pipeline for different settings of the window size for which the HRV is calculated and various time intervals between consecutive windows. Preparing the heatmaps of correlation coefficients and corresponding p-values can be done by running the <em>utils_advanced_plots.py</em> file after performing the sensitivity analysis. Furthermore, a detailed analysis for the one selected set of hyperparameters may be prepared (by setting <em>sensitivity_analysis = False</em>), i.e. for 15-minute window sizes, 1-minute time intervals between consecutive windows and without data interpolation method. Also, patients taking quetiapine may be excluded from further calculations by setting <em>exclude_quetiapine = True</em> because this medicine can have a strong impact on HRV [Hattori et al].</p> <p>The accelerometer data processing may be performed using the <em>utils_accelerometer.py</em> file. In this case, accelerometer recordings are downsampled to ensure the same timestamps as for R-R intervals and, for each participant, the mobility coefficient is calculated. Then, a correlation coefficient between mean HRV values and mobility coefficient is computed. The plotting of the pure accelerometer signal may be done by running the <em>utils_loading.py </em>file.</p> <p>The comparison of age distribution between the tested groups can be made by the histogram plotted with the use of the <em>utils_basic_plots.py</em> file.</p>
Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN
<p>This repository contains the data released in the paper 'Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN' <em>(DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>).</em></p> <p>We release a detailed catalogue of Giant Star-forming Clumps (GSFCs), detected for the full set of Galaxy Zoo: Clump Scout galaxies observed by SDSS using the Faster R-CNN architecture with the Zoobot classification-CNN as a feature extraction backbone.</p> <p>The final models and code are made publicly available via Github: <a href="https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout">https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout</a>.</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>) when using the data in this repository.</p> <p>The csv-file <em>FRCNN_Zoobot_SDSS_GZCS_detections.csv</em> has the following columns. Alternatively, the file <em>FRCNN_Zoobot_SDSS_GZCS_detections.gzip</em> contains the same data but stored as a parquet-file.</p> <table> <tbody><tr> <th>Column name</th> <th>Description</th> </tr> </tbody><tbody> <tr> <td>specobjid</td> <td>SDSS spec object ID</td> </tr> <tr> <td>dr7objid</td> <td>SDSS DR7 object ID</td> </tr> <tr> <td>clump_id</td> <td>Clump index</td> </tr> <tr> <td>clump_label_id</td> <td>Clump label ID (1 or 2)</td> </tr> <tr> <td>clump_label_name</td> <td>Clump label name</td> </tr> <tr> <td>clump_score</td> <td>Detection score for the clump</td> </tr> <tr> <td>clump_centre_ra</td> <td>Clump centroid RA in degrees</td> </tr> <tr> <td>clump_centre_dec</td> <td>Clump centroid dec in degrees</td> </tr> <tr> <td>clump_flux_u</td> <td>Clump u-band flux in Jy</td> </tr> <tr> <td>clump_flux_g</td> <td>Clump g-band flux in Jy</td> </tr> <tr> <td>clump_flux_r</td> <td>Clump r-band flux in Jy</td> </tr> <tr> <td>clump_flux_i</td> <td>Clump i-band flux in Jy</td> </tr> <tr> <td>clump_flux_z</td> <td>Clump z-band flux in Jy</td> </tr> <tr> <td>clump_flux_err_u</td> <td>Clump u-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_g</td> <td>Clump g-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_r</td> <td>Clump r-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_i</td> <td>Clump i-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_z</td> <td>Clump z-band flux error in Jy</td> </tr> <tr> <td>clump_mag_u</td> <td>Clump u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_g</td> <td>Clump g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_r</td> <td>Clump r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_i</td> <td>Clump i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_z</td> <td>Clump z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_ext_mag_u</td> <td>Clump u-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_g</td> <td>Clump g-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_r</td> <td>Clump r-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_i</td> <td>Clump i-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_z</td> <td>Clump z-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u</td> <td>Clump corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_g</td> <td>Clump corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_r</td> <td>Clump corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_i</td> <td>Clump corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_z</td> <td>Clump corrected z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u_g</td> <td>Clump colour (u-g)</td> </tr> <tr> <td>clump_mag_corr_g_r</td> <td>Clump colour (g-r)</td> </tr> <tr> <td>clump_mag_corr_r_i</td> <td>Clump colour (r-i)</td> </tr> <tr> <td>clump_mag_corr_i_z</td> <td>Clump colour (i-z)</td> </tr> <tr> <td>clump_flux_ratio</td> <td>Est. clump/galaxy near-UV flux ratio (u-band)</td> </tr> <tr> <td>is_clump_3pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is >3%</td> </tr> <tr> <td>is_clump_8pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is >8%</td> </tr> <tr> <td>galaxy_ra</td> <td>Host galaxy RA in degrees</td> </tr> <tr> <td>galaxy_dec</td> <td>Host galaxy dec in degrees</td> </tr> <tr> <td>galaxy_z</td> <td>Host galaxy redshift</td> </tr> <tr> <td>galaxy_mag_u</td> <td>Host galaxy u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_g</td> <td>Host galaxy g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_r</td> <td>Host galaxy r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_i</td> <td>Host galaxy i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_z</td> <td>Host galaxy z-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_u</td> <td>Host galaxy u-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_g</td> <td>Host galaxy g-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_r</td> <td>Host galaxy r-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_i</td> <td>Host galaxy i-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_z</td> <td>Host galaxy z-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_flux_u</td> <td>Host galaxy u-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_g</td> <td>Host galaxy g-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_r</td> <td>Host galaxy r-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_i</td> <td>Host galaxy i-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_z</td> <td>Host galaxy z-band flux in Jy</td> </tr> <tr> <td>galaxy_expAB_r</td> <td>Host galaxy axis ratio from SDSS</td> </tr> <tr> <td>galaxy_expRad_r</td> <td>Host galaxy exponential fit scale radius from SDSS</td> </tr> <tr> <td>galaxy_lmass</td> <td>Host galaxy log mass in MSun</td> </tr> <tr> <td>galaxy_lssfr</td> <td>Host galaxy log specific SFR</td> </tr> <tr> <td>galaxy_mag_corr_u</td> <td>Host galaxy corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_g</td> <td>Host galaxy corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_r</td> <td>Host galaxy corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_i</td> <td>Host galaxy corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_z</td> <td>Host galaxy corrected z-band magnitude (AB-mag)</td> </tr> </tbody> </table> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.