Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,330
datasets available to search
ShareScore release 0.9.0
Dataset results
9,330 results for “Approach”
Seasonal hindcast of temperature and precipitation at a local scale by using TeWA approach
<p><strong>Methodology</strong></p> <p>Data set of simulated time-series of temperature and precipitation for the 1982-2020 period. Our statistical seasonal prediction model have two main components: a) the ocean-atmosphere coupling represented by correlations between surface variables with delayed teleconnections and b) the self-predictability of the residual anomalies by trends or cycles (quasi-oscillations).</p> <p>The approach has three stages approach with two main predictor components, as mentioned above. The first two stages consist of separate predictions, one per each component, and the third stage is a combination of both predictions (Fig. 2): Teleconnection-based approach (Redolat et al. 2019, 2020) and a self-predictability by using Wavelet-ARIMA models (Conejo et al. 2005; Joo and Kim 2015). Therefore, the total method is a Teleconnection+Wavelet+ ARIMA (TeWA) approach.</p> <p><strong>References</strong></p> <p>Conejo, A.J., M.A. Plazas, R. Espinola, A.B. Molina, 2005: Day-ahead electricity price forecasting using the wavelet transform and ARIMA models. IEEE Trans. Power Syst., 20, 1035-1042, https://doi.org/10.1109/TPWRS.2005.846054.</p> <p>Joo, T., S. Kim, 2015: Time series forecasting based on wavelet filtering. Expert Syst. Appl. 42, 3868-3874. https://doi.org/10.1016/j.eswa.2015.01.026</p> <p>Redolat, D., R. Monjo, C. Paradinas, J. Pórtoles, E. Gaitán, C. Prado-López, and J. Ribalaygua, 2020: Local decadal prediction according to statistical/dynamical approaches. Int. J. Climatol., 40: 5671–5687. https://doi.org/10.1002/joc.6543.</p> <p>Redolat, D.; R. Monjo, J.A. Lopez-Bustins, and J. Martin-Vide, 2019: Upper-Level Mediterranean Oscillation index and seasonal variability of rainfall and temperature. Theor. Appl. Climatol., 135: 1059–1077. https://doi.org/10.1007/s00704-018-2424-6.</p>
Dataset: Does vendor breeding colony influence sign- and goal-tracking in Pavlovian conditioned approach?
<p>Vendor differences are thought to affect Pavlovian conditioning in rats. After observing possible differences in sign-tracking and goal-tracking behaviour with rats from different breeding colonies, we performed an empirical replication of the effect. 40 male Long-Evans rats from Charles River colonies ‘K72’ and ‘R06’ received 11 Pavlovian conditioned approach training sessions (or “autoshaping”), with a lever as the conditioned stimulus (CS) and 10% sucrose as the unconditioned stimulus (US). Each 58-min session consisted of 12 CS-US trials. Paired rats (n = 15/colony) received the US following lever retraction. Unpaired control rats (n = 5/colony) received sucrose during the inter-trial interval. Next, we evaluated the conditioned reinforcing properties of the CS, by determining whether rats would learn to nose-poke into a new, active (vs. inactive) port to receive CS presentations alone (no sucrose). Preregistered confirmatory analyses showed that during autoshaping sessions, Paired rats made significantly more CS-triggered entries into the sucrose port (i.e., goal-tracking) and lever activations (sign-tracking) than Unpaired rats did, demonstrating acquisition of the CS-US association. Confirmatory analyses showed no effects of breeding colony on autoshaping. During conditioned reinforcement testing, analysis of data from Paired rats alone showed significantly more active vs. inactive nosepokes, suggesting that in these rats, the lever CS acquired incentive motivational properties. Analysing Paired rats alone also showed that K72 rats had higher Pavlovian Conditioned Approach scores than R06 rats did. Thus, breeding colony can affect outcome in Pavlovian conditioned approach studies, and animal breeding source should be considered as a covariate in such work.Vendor differences are thought to affect Pavlovian conditioning in rats. After observing possible differences in sign-tracking and goal-tracking behaviour with rats from different breeding colonies, we performed an empirical replication of the effect. 40 male Long-Evans rats from Charles River colonies ‘K72’ and ‘R06’ received 11 Pavlovian conditioned approach training sessions (or “autoshaping”), with a lever as the conditioned stimulus (CS) and 10% sucrose as the unconditioned stimulus (US). Each 58-min session consisted of 12 CS-US trials. Paired rats (n = 15/colony) received the US following lever retraction. Unpaired control rats (n = 5/colony) received sucrose during the inter-trial interval. Next, we evaluated the conditioned reinforcing properties of the CS, by determining whether rats would learn to nose-poke into a new, active (vs. inactive) port to receive CS presentations alone (no sucrose). Preregistered confirmatory analyses showed that during autoshaping sessions, Paired rats made significantly more CS-triggered entries into the sucrose port (i.e., goal-tracking) and lever activations (sign-tracking) than Unpaired rats did, demonstrating acquisition of the CS-US association. Confirmatory analyses showed no effects of breeding colony on autoshaping. During conditioned reinforcement testing, analysis of data from Paired rats alone showed significantly more active vs. inactive nosepokes, suggesting that in these rats, the lever CS acquired incentive motivational properties. Analysing Paired rats alone also showed that K72 rats had higher Pavlovian Conditioned Approach scores than R06 rats did. Thus, breeding colony can affect outcome in Pavlovian conditioned approach studies, and animal breeding source should be considered as a covariate in such work.</p>
An External Replication on the Effects of Test-driven Development Using a Multi-site Blind Analysis Approach
<p>This dataset contains the <strong>unblinded </strong>version of the data collected and analyzed for the experiment reported in the paper. </p> <p>The semantics of the data can be found in the spreadsheet. For the formulas on how to obtain this data from the raw data, please see the paper. </p>
Local explanation SHAP approach applied to MIROC5,RCP8.5-forced multi-model ensemble study of GrIS future sea-level contributions
<p>The repository contains materials for analysing the results of the Local explanation named SHAP-CTREE (Redelmeier et al., 2020) approach applied to the MIROC5,RCP8.5-forced multi-model ensemble study of GrIS future sea-level contributions from Goelzer et al. (2020).</p> <p>The available files are:<br> - run_SupplMat.R: the main R script to perform the diagnostics and the different analyses (levels 1 - 3)<br> - utilsPLOT.R: functions for plotting<br> - Diagnostics.zip: the zip file with the png figures, named 'GrIS_CaseXXX_yYYY.png', that depict the diagnostic for case XXX for prediction time YYY<br> - SupplementaryMaterials.zip<br> - RData files for each prediction time YYY "Shapley_yYYY" with:<br> S: matrix N=55 cases x d+1: SHAP values for the d inputs (+ average sea level value at time YYY)<br> YHAT: ML-based predictions of the sea level for the 55 cases<br> YTRUE: true values for the 55 cases<br> mae: mean absolute error<br> - RData file containing the design of experiments "DOE_GrIS_MIROC5-RCP85.RData"<br> doe: matrix with values of the d=9 inputs</p> <p>These constitute the supplementary materials of Rohmer et al. (2022, The Cryosphere). All technical details are provided in this reference.</p>
Supplementary files for Vertical Displacements and Sea-Level Changes in Eastern North America Driven by Glacial Isostatic Adjustment: an Ensemble Modeling Approach
<p>Model input and output files associated with the manuscript entitled "Vertical Displacements and Sea-Level Changes in Eastern North America Driven by Glacial Isostatic Adjustment: an Ensemble Modeling Approach" that will be submitted to Journal of Geophysical Research.</p>
Elevating Cybersecurity for Smart Grid Systems—A Container-Based Approach Enhanced by Machine Learning
<p>README<br>Title<br>Elevating Cybersecurity for Smart Grid Systems—A Container-Based Approach Enhanced by Machine Learning</p> <p>Authors<br>Mays Abukeshek, School of Computer Science, Faculty of Technology, University of Sunderland, University of Huddersfield, UK<br>Email: mays.abukeshek@sunderland.ac.uk, Mays.abukeshek@hud.ac.uk<br>Basel Barakat, School of Computer Science, Faculty of Technology, University of Sunderland, UK<br>Email: basel.barakat@sunderland.ac.uk<br>Bamidele Ajayi, School of Computer Science, Faculty of Technology, University of Sunderland, UK<br>Email: bamidele.ajayi@research.sunderland.ac.uk<br>Abstract<br>This dataset supports the paper "Elevating Cybersecurity for Smart Grid Systems—A Container-Based Approach Enhanced by Machine Learning," which presents a comprehensive implementation of a cybersecurity solution for smart grid network containers. The methodology utilizes:</p> <p>Qualys API-based vulnerability scanning and reporting system for vulnerability identification<br>Docker deployment for security and isolation<br>Advanced load balancing techniques for resource optimization<br>Machine learning-powered anomaly detection for threat identification and vulnerability prioritization.<br>The dataset contains details of several simulated attacks enabling effective training and evaluation of a robust machine-learning model.</p> <p>Data Description<br>The dataset includes logs from conducted attacks on containerized nodes, generated to reflect real-world scenarios. The simulated attacks include:</p> <p>Denial of Service (DoS)<br>Remote-to-Local (R2L)<br>User-to-Root (U2R)<br>Probes<br>Contents<br>Csv_file.csv: This file contains the dataset used for training and evaluating the machine learning models. The columns in the dataset represent various features and results of the simulated attacks.<br>Data Columns and Rows<br>Timestamp:</p> <p>Description: The exact date and time when the data was recorded.<br>time: 2023-06-01 12:00:00</p> <p>Attack_Type:</p> <p>Description: The type of cyber-attack conducted.<br>Possible Values: DoS, R2L, U2R, Probe<br>Example: DoS<br>Notes: Categorizes the type of attack, crucial for training classification models.<br>CPU_Utilization (%):</p> <p>Description: The percentage of CPU resources used during the attack.<br>Example: 52.3<br>Notes: Indicates the load on the CPU during the attack, useful for assessing the impact of attacks on system performance.<br>Memory_Utilization (%):</p> <p>Description: The percentage of memory resources used during the attack.<br>Example: 63.4<br>Notes: Shows memory usage which can be a critical factor in understanding system performance under attack conditions.<br>Network_Bandwidth (Mbps):</p> <p>Description: The bandwidth of the network in Megabits per second.<br>Example: 100<br>Notes: Reflects the network load and is essential for analyzing the impact on network performance.<br>Vulnerabilities_Detected:</p> <p>Description: The number of vulnerabilities detected during the attack.<br>Example: 289<br>Notes: Indicates the effectiveness of the vulnerability scanning process and the system's exposure to threats.<br>Mean_Response_Time (ms):</p> <p>Description: The average response time in milliseconds during the attack.<br>Example: 87<br>Notes: Important for evaluating the responsiveness of the system under attack conditions.<br>Throughput (requests/second):</p> <p>Description: The number of requests the system can handle per second during the attack.<br>Example: 1068<br>Notes: Measures the capacity and efficiency of the system under load.<br>Example Row<br>Timestamp Attack_Type CPU_Utilization (%) Memory_Utilization (%) Network_Bandwidth (Mbps) Vulnerabilities_Detected Mean_Response_Time (ms) Throughput (requests/second)<br>2023-06-01 12:00:00 DoS 52.3 63.4 100 289 87 1068<br>Usage<br>This dataset can be used to:</p> <p>Train and evaluate machine learning models for cybersecurity applications in smart grid systems.<br>Analyze the performance of different machine learning models in detecting and prioritizing vulnerabilities.<br>Understand the impact of various types of cyber-attacks on containerized environments.<br>Methodology<br>The dataset was created using a combination of Qualys API-based vulnerability scanning and Docker containerization. Multiple container clusters were subjected to various simulated attacks, and the performance of machine learning models was evaluated based on accuracy, precision, recall, and F1-scores.</p> <p>Acknowledgments<br>This research was supported by the University of Sunderland and the University of Huddersfield.</p> <p>References<br>Please refer to the full paper for detailed methodology, implementation, and analysis:<br>IEEE</p>
Experimental determination of the sulfur K-shell fundamental parameters employing the holistic approach
<p>This dataset contains the experimentally determined fundamental parameters for the sulfur K-subshells from as shown in the publication with the title "Experimental determination of the sulfur K-shell fundamental parameters employing the holistic approach". The paper will be published soon in a peer-reviewd journal.</p> <p>This file contains the following fundamental parameters for sulfur: K-subshell fluorescence yield, K-shell Auger yield, Ka and Kb transition probabilities, K-subshell photo ionization cross sections up to 10 keV, K-subshell fluorescence prodution cross sections up to 10 keV</p>
Source Code Accompanying the Paper "More on network approaches in Historical Chinese Phonology (音韻學)"
<p>First version of the source code and data accompanying the paper "More on Network Approaches in Historical Chinese Phonology".</p> <p>This paper is available here:</p> <ul> <li>List, Johann-Mattis (2018): <strong>More on network approaches in Historical Chinese Phonology (音韻學)</strong>. Paper prepared for the <em>LFK Society Young Scholars Symposium</em>. Taibei: Li Fang-Kuei Society ofr Chinese Linguistics. URL: <a href="https://hal.archives-ouvertes.fr/hal-01706927">https://hal.archives-ouvertes.fr/hal-01706927</a>.</li> </ul> <pre><code>@InProceedings{List2018a, author = {List, Johann-Mattis}, title = {{More on Network Approaches in Historical Chinese Phonology (音韻學)}}, booktitle = {{LFK Society Young Scholars Symposium}}, year = {2018}, publisher = {Li Fang-Kuei Society for Chinese Linguistics}, pdf = {https://hal.archives-ouvertes.fr/hal-01706927/file/main.pdf}, url = {https://hal.archives-ouvertes.fr/hal-01706927}, address = {Taipei}, hal_id = {hal-01706927}, } </code></pre> <p>See the README.md for mor information.</p> <ul> <li> </li> </ul>
Comparison and practical review of segmentation approaches for label-free microscopy
<p>This dataset contains microscopic images of PNT1A cell line captured by multiple microcopic without use of any labeling and a manually annotated ground truth for subsequent use in segmentation algorithms. Dataset also includes images reconstructed according to the methods described below in order to ease further segmentation. </p> <p>See Vicar et al. Cell segmentation methods for label-free contrast microscopy: review and comprehensive comparison. BMC Bioinformatics (2019) 20:360. DOI <a href="https://doi.org/10.1186/s12859-019-2880-8">10.1186/s12859-019-2880-8</a></p> <p>Code using this dataset is available at <a href="https://github.com/tomasvicar/Cell-segmentation-methods-comparison">https://github.com/tomasvicar/Cell-segmentation-methods-comparison</a></p> <p><strong>Materials and methods </strong></p> <p>Cells were cultured in RPMI-1640 medium supplemented with antibiotics (penicillin 100 U/ml and streptomycin 0.1 mg/ml) with 10% fetal bovine serum. Prior microscopy acquisition, cells were maintained at 37 cenigrade in a humidified incubator with 5% CO2. Intentionally, high passage number of cells was used (>30) in order to describe distinct morphological heterogeneity of cells (rounded and spindle-shaped, relatively small to large polyploid cells). For acquisition purposes, cells were cultivated in Flow chambers µ-Slide I Luer Family (Ibidi, Martinsried, Germany).</p> <p>Quantitative phase imaging (QPI) microscopy was performed on Tescan Q-PHASE (Tescan, Brno, Czech republic), with objective Nikon CFI Plan Fluor 10x/0.30 captured by Ximea MR4021MC (Ximea, Münster, Germany). Imaging is based on the original concept of coherence-controlled holographic microscope \cite{Kolman:10,Slaby:13}, images are shown in grayscale with units of pg/µm2.</p> <p>DIC microscopy was performed on microscope Nikon A1R (Nikon, Tokyo, Japan), with objective Nikon CFI Plan Apo VC 20x/0.75 captured by CCD camera Jenoptik ProgRes MF (Jenoptik, Jena, Germany). </p> <p>HMC microscopy was performed on microscope Olympus IX71 (Olympus, Tokyo, Japan), with objective Olympus CplanFL N 10x/0.3 RC1 captured by CCD camera Hamamatsu Photonics ORCA-R2 (Hamamatsu Photonics K.K., Hamamatsu, Japan).</p> <p>PC microscopy was performed on a Nikon Eclipse TS100-F microscope, with a Nikon CFI Achro ADL 10x/0.25 objective captured by CCD camera Jenoptik ProgRes MF.</p> <p><strong>Folder structure and file and filename description</strong><br> <br> <em>folder "source data+groundtruth"</em><br> - includes raw microscopic data <br> (uncompressed 16-bit for DIC, HMC and PC, 32-bit for QPI)<br> - includes manualy annotated groundtruth (zip file - imageJ ROI file, 1bit png mask)</p> <p>e.g. <br> DIC_01_raw.tif<br> DIC_01_groundtruth_imagejROI.zip<br> DIC_01_groundtruth_mask.png</p> <p><br> <em>folder "reconstructions"</em></p> <p>includes reconstructed images using reconstructions with highest dice coefficient achieved. </p> <p>for DIC and HMC: rDIC-Koos, rDIC-Yin, and rWeka<br> for PC: rPC-Top-Hat, rDIC-Yin, and rWeka<br> for QPI: rWeka</p> <p>note that for rWeka images numbered 01 for DIC, HMC and PC and 01-03 for QPI were used for learning.</p> <p><strong>Abbreviations</strong><br> DIC, differential image contrast<br> HMC, Hoffman modulation contrast<br> PC, phase contrast<br> QPI, quantitative phase imaging<br> rDIC-Koos, DIC/HMC image reconstruction according to Koos et al, Sci Rep. 2016;6:30420<br> rDIC-Yin, DIC/HMC image reconstruction according to Yin et al, Inf Process Med Imaging. 2011;22:384-97.<br> rPC-Yin, PC image reconstruction according to Yin et al, Med Im Anal. 2012; 16(5):1047<br> rPC-Top-Hat, Top-Hat filter according to Dewan et al, IEEE Transactions on Biomedical Circuits and<br> Systems.2014;8(5):716-728<br> rWeka, probability map using Trainable Weka segmentation according to Arganda-Carreras et al. Bioinformatics. 2017</p>
Data From: The Oyster River Protocol: A multi assembler and kmer approach for de novo transcriptome assembly.
<p>Characterizing transcriptomes in non-model organisms has resulted in a massive increase in our understanding of biological phenomena. This boon, largely made possible via high-throughput sequencing, means that studies of functional, evolutionary and population genomics are now being done by hundreds or even thousands of labs around the world. For many, these studies begin with a <em>de novo</em> transcriptome assembly, which is a technically complicated process involving several discrete steps. The Oyster River Protocol (ORP), described here, implements a standardized and benchmarked set of bioinformatic processes, resulting in an assembly with enhanced qualities over other standard assembly methods. Specifically, ORP produced assemblies have higher Detonate and TransRate scores and mapping rates, which is largely a product of the fact that it leverages a multi-assembler and kmer assembly process, thereby bypassing the shortcomings of any one approach. These improvements are important, as previously unassembled transcripts are included in ORP assemblies, resulting in a significant enhancement of the power of downstream analysis. Further, as part of this study, I show that assembly quality is unrelated with the number of reads generated, above 30 million reads. Code Availability: The version controlled open-source code is available at <a href="https://github.com/macmanes-lab/Oyster_River_Protocol">https://github.com/macmanes-lab/Oyster_River_Protocol</a>. Instructions for software installation and use, and other details are available at <a href="http://oyster-river-protocol.rtfd.org/">http://oyster-river-protocol.rtfd.org/</a>.</p>
Hybrid Approaches to Detect Comments Violating Macro Norms on Reddit
<p>[<strong>Content warning: </strong><em>Files may contain instances of highly inflammatory and offensive content.]</em></p> <p><br> This dataset was generated as an extension of our <a href="https://www.cc.gatech.edu/~eshwar3/uploads/3/8/0/4/38043045/eshwar-norms-cscw2018.pdf">CSCW 2018 paper</a>:</p> <p><em>Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The Internet’s Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro Scales. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 32.</em></p> <p><strong>Description:</strong></p> <p>Working with over 2M removed comments collected from 100 different communities on Reddit (subreddit names listed in data/study-subreddits.csv), we identified <strong>8 macro norms</strong>, i.e., norms that are widely enforced on most parts of Reddit. We extracted these macro norms by employing a hybrid approach—classification, topic modeling, and open-coding—on comments identified to be norm violations within at least 85 out of the 100 study subreddits. Finally, we labelled over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated, and make this dataset publicly available (also available on <a href="https://github.com/ceshwar/reddit-norm-violations">Github</a>).</p> <p>For each of the labeled topics, we identified the top 5000 removed comments that were best fit by the LDA topic model. In this way, we identified over 5000 removed comments that are examples of each type of macro norm violation described in the paper. The removed comments were sorted by their topic fit, stored into respective files based on the type of norm violation they represent, and are made available on this repo.</p> <p>Here we make the following datasets publicly available:</p> <p>* <strong>1 file</strong> containing the log of over 2M removed comments obtained from the top 100 subreddits between May 2016 to March 2017, after filtering out the following comments: 1) comments by u/AutoModerator, 2) replies to removed comments (i.e., children of the poisoned tree - refer to the paper for more information), and 3) non-readable comments (not utf-8 encoded).</p> <p>* <strong>8 files</strong>, each containing 5000+ removed comments obtained from Reddit, are stored in: data/macro-norm-violations/ , and they are split into different files based on the macro norm they violated. Each new line in the files represent a comment that was posted on Reddit between May 2016 to March 2017, and subsequently removed by subreddit moderators for violating community norms. All comments were preprocessed using the script in code/preprocessing-reddit-comments.py , in order to do the following: 1. remove new lines, 2. convert text to lowercase, and 3. strip numbers and punctuations from comments.</p> <p><strong>Description of 1 file</strong> containing over<em> 2M removed comments </em>from <em>100 subreddits.</em></p> <ul> <li>"reddit-removal-log.csv" - all comments that were removed from the 100 study subreddits during the study period described above (post-filtering).</li> </ul> <p><strong>Descriptions of each file</strong> containing <em>5059 comments</em> (that were removed from Reddit, and preprocessed)<strong> violating macro norms </strong>present in data/macro-norm-violations/:</p> <ul> <li>"macro-norm-violations-n10-t0-misogynistic-slurs.csv" - Comments that use misogynistic slurs.</li> <li>"macro-norm-violations-n15-t2-hatespeech-racist-homophobic.csv" - Comments containing hate speech that is racist or homophobic.</li> <li>"macro-norm-violations-n10-t3-opposing-political-views-trump.csv", "macro-norm-violations-n15-t10-opposing-political-views-trump.csv" - Comments with opposing political views around Trump (depends on originating sub).</li> <li>"macro-norm-violations-n10-t4-verbal-attacks-on-Reddit.csv" - Comments containing verbal attacks on Reddit or specific subreddits.</li> <li>"macro-norm-violations-n10-t5-porno-links.csv" - Comments with pornographic links.</li> <li>"macro-norm-violations-n10-t8-personal-attacks.csv", "macro-norm-violations-n10-t9-personal-attacks.csv"- Comments containing personal attacks.</li> <li>"macro-norm-violations-n15-t3-abusing-and-criticisizing-mods.csv" - Comments abusing and criticisizng moderators.</li> <li>"macro-norm-violations-n15-t9-namecalling-claiming-other-too-sensitive.csv" - Comments with name-calling, or claiming that the other person is too sensitive.</li> </ul> <p>More details about the dataset can be found on arXiv: <a href="https://arxiv.org/abs/1904.03596">https://arxiv.org/abs/1904.03596</a></p>
Identifying Coronal Mass Ejection Active Region Sources: An automated approach - Catalogue results
<p>Catalogue of Coronal Mass Ejection (CME) active region sources. Includes a database version and a simplified .csv version. For full details, refer to the source code at <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. We include a README file for each describing each column.</p> <p>We also include the raw data used to generate the catalogue so that results may be reproduced following the steps detailed in <a href="https://github.com/JulioHC00/cmesrc">https://github.com/JulioHC00/cmesrc</a>. This is a collection of data from other works and we provide it only to allow the results to be reproduced</p> <p>Below, we detail the data sources for the raw_data folders</p> <p>==============================<br><strong>RAW DATA SOURCES</strong><br>==============================</p> <p><strong>DIMMINGS FOLDER</strong></p> <p>Data is from Solar Demon, .csv was provided by Emil Kraaikamp through private communication.</p> <blockquote> <p>Solar Demon – an approach to detecting flares, dimmings, and EUV waves on SDO/AIA images<br>Emil Kraaikamp, Cis Verbeeck<br>J. Space Weather Space Clim. 5 A18 (2015)<br>DOI: 10.1051/swsc/2015019</p> </blockquote> <p><strong>HARPNUM_TO_NOAA FOLDER</strong></p> <p>Obtained from http://jsoc.stanford.edu/doc/data/hmi/harpnum_to_noaa/all_harps_with_noaa_ars.txt</p> <p><strong>LASCO FOLDER</strong></p> <p>This CME catalog is generated and maintained at the CDAW Data Center by NASA and The Catholic University of America in cooperation with the Naval Research Laboratory. SOHO is a project of international cooperation between ESA and NASA.</p> <p>Downloaded from https://cdaw.gsfc.nasa.gov/CME_list/</p> <p><strong>MVTS FOLDER</strong></p> <p>Data from</p> <blockquote> <p>Angryk, R.A., Martens, P.C., Aydin, B. et al. Multivariate time series dataset for space weather data analytics. Sci Data 7, 227 (2020). https://doi.org/10.1038/s41597-020-0548-x</p> </blockquote> <p>Available at the Harvard Dataverse</p> <blockquote> <p>Angryk, Rafal; Martens, Petrus; Aydin, Berkay; Kempton, Dustin; Mahajan, Sushant; Basodi, Sunitha; Ahmadzadeh, Azim; Xumin Cai; Filali Boubrahimi, Soukaina; Hamdi, Shah Muhammad; Schuh, Micheal; Georgoulis, Manolis, 2020, "SWAN-SF", https://doi.org/10.7910/DVN/EBCFKM, Harvard Dataverse, V1</p> </blockquote> <p>The DT_SWAN folder contains the same data but with extra columns obtained directly from the Joint Science Operations Center (JSOC) through the python package drms.</p>
Supplementary data for the article Linguistic system and sociolinguistic environment as competing factors in linguistic variation: A typological approach
<p>This material contains the dataset and the R-scripts from the <a href="https://version.helsinki.fi/gramadapt/linguistic-system-and-sociolinguistic-environment">gitlab repository</a> of the following article. Please cite the article when using the data.</p> <p>Sinnemäki, Kaius 2020. Linguistic system and sociolinguistic environment as competing factors in linguistic variation: A typological approach. <em>Journal of Historical Sociolinguistics</em> 6(2): 20190101. <a href="https://doi.org/10.1515/jhsl-2019-1010">https://doi.org/10.1515/jhsl-2019-1010</a></p>
Dataset for "Best organic farming deployment scenarios for pest control: a modeling approach" V3
<p>Organic Farming (OF) has been expanding recently in response to growing consumer demand and as a response to environmental concerns. The area under OF is expected to further increase in the future. The effect of OF expansion on pest densities in organic and conventional crops remains difficult to predict because OF expansion impacts Conservation Biological Control (CBC), which depends on the surrounding landscape context. In order to understand and forecast how pests and their biological control may vary during OF expansion, we modeled the effect of spatial changes in farming practices on population dynamics of a pest and its natural enemy. We investigated the impact on pest density and on predator to pest ratio of three contrasted scenarios aiming at 50% organic fields through the progressive conversion of conventional fields. Scenarios were 1) conversion of Isolated conventional fields first (IP), 2) conversion of conventional fields within Groups of conventional fields first (GP), and 3) Random conversion of conventional field (RD). We coupled a neutral spatially explicit landscape model to a predator-prey model to simulate pest dynamics in interaction with natural enemy predators. The three OF expansion scenarios were applied to nine landscape types differing in their proportion and fragmentation of semi-natural habitat. We further investigated if the ranking of scenarios was robust to pest control methods in OF fields and pest and predator dispersal abilities.</p> <p>We found that organic farming expansion affected more predator densities than pest densities for most landscape types. The impact of OF expansion on final pest and predator densities was also stronger in organic than conventional fields and in landscapes with large proportions of highly fragmented semi-natural habitats. Based on pest densities and the predator to pest ratio, our results suggest that a progressive organic conversion with a focus on isolated conventional fields (scenario IP) could help promote CBC. Careful landscape planning of OF expansion appeared most necessary when pest management was substantially less efficient in organic than in conventional crops, and in landscapes with low proportion of semi-natural habitats.</p> <p><strong>This dataset contains simulation outputs and the R script that was used to describe, display and analyse data. The model itself can be found at <a href="https://doi.org/10.17605/OSF.IO/Z2QCX">https://doi.org/10.17605/OSF.IO/Z2QCX</a></strong></p> <p><strong>Please note that this is the third version of this dataset, following recommendations from the PCI Ecology reviewers and editor.</strong></p>
A novel and holistic approach for experimental X-ray fundamental parameter determination - the Ru L-shell
<p>This dataset contains the experimentally determined fundamental parameters for the ruthenium L-subshells from our paper with the title "A novel and holistic approach for experimental X-ray fundamental parameter determination - the Ru L-shell". The paper will be published soon in a peer-reviewd journal.</p> <p>This file contains L-subshell fluorescence yields, L-shell Coster-Kronig factors, L-shell Auger yields, <br> Mass attenuation coefficients in energy range from 2.41 keV to 8 keV, L-subshell photo ionization cross sections up to 8 keV and <br> L-subshell fluorescence prodution cross sections of Ru.</p>
Land Productivity Dynamics (LPD) maps based on approaches recommended by the United Nations Convention to Combat Desertification (UNCCD) for the Brazilian Semiarid Region (BSR).
<p>Maps of the land productivity dynamic (LPD) based on approaches recommended by the United Nations Convention to Combat Desertification (UNCCD) for the Brazilian Semiarid Region for the period 2001-2015. LPD from Trends.Earth (TE), LPD from the Joint Research Center (JRC-LPD), and LPD from the Food and Agriculture Organization – World Overview of Conservation Approaches and Technologies (FAO-WOCAT LPD). In addition, TE-based LPD with climate correction based on Rain Use Efficiency (RUE), Residual Trend Analysis (RESTREND), and Water Use Efficiency (WUE) calculated using the procedures described in the second version of the Good Practice Guidance for SDG Indicator 15.3.1. The annual LCLU maps from the MapBiomas project at 30 m spatial resolution for 2001 and 2015, and the 16-day MOD13Q1 NDVI dataset from 2001 to 2015 were used as inputs.</p> <p>**********************</p> <p>A total of seven GeoTIFF files in Geographic Tagged Image File Format (GeoTIFF) format are provided at 250 m spatial resolution.</p> <p>Coding for the LPDs.</p> <p>Value Meaning</p> <p>-32768 No data</p> <p>1 Declining</p> <p>2 Moderate decline</p> <p>3 Stressed</p> <p>4 Stable</p> <p>5 Increasing</p> <p>Funding: this study was undertaken as part of the Satellite Desertification Monitoring Program in the Brazilian Semiarid Region [Grant Number 403223/2021-0] supported by the CNPq. It also had the support of Capes, through Notice no. 28/2022 – PDPG Social Vulnerability & Human Rights [Grant Number 88881.705050/2022-01].</p>
Image sets used in the development of a connected auto-encoders based approach to separate mixed X-radiographs from double-sided paintings
<p>The following sets of images were used during the development of an algorithm (described in the publication detailed below) designed to separate the mixed X-radiographs from double-sided paintings into two hypothetical X-ray images corresponding to each side of the painting, when visible images of the two sides of the painting are available.</p> <p>The images sets are taken from a painting that is only painted on one side and were used to assess the regularization parameters associated with the separation approach. The details are taken from the visible image and the X-radiograph of Anthony van Dyck’s painting <em>Lady Elizabeth Thimbelby and Dorothy, Viscountess Andover</em> dated to about 1635 and now in the collection of the National Gallery in London (NG6437). See <a href="https://www.nationalgallery.org.uk/paintings/anthony-van-dyck-lady-elizabeth-thimbelby-and-her-sister">https://www.nationalgallery.org.uk/paintings/anthony-van-dyck-lady-elizabeth-thimbelby-and-her-sister</a> for further details of the painting.</p> <p>The code can be downloaded from: <a href="https://github.com/ART-ICT/Xray_Separation_2RGB">https://github.com/ART-ICT/Xray_Separation_2RGB</a> and the algorithm is described in W. Pu, B. Sober, N. Daly, C. Zhou, Z. Sabetsarvestani, C. Higgitt, I. Daubechies and M. Rodrigues, ‘Image Separation with Side Information: A Connected Auto-Encoders Based Approach’, <em>Transactions on Image Processing, </em>2023 </p> <p><strong>All images © The National Gallery, London</strong></p> <p> </p> <p><strong><em>Datasets available: </em></strong></p> <p><strong>NG6437_vis_800pixel_230502.tif</strong>: 800 pixel thumbnail visible image of the entire painting showing the location of the two details used for the algorithm development. This image is derived from a visible image of the whole painting acquired 25 November 2019 (Original file: N-6437-00-000041.tif; 6272 x 5940 pixels).</p> <p><strong>NG6437_xray_800pixel_230502.tif</strong>: 800 pixel thumbnail image of the X-radiograph of the entire painting showing the location of the two details used for the algorithm development. This image is derived from the composite X-radiography of the whole painting created by mosaicking digital scans of the individual sheets of film and then registering the resulting image to the high resolution visible image described above (Original file: N-6437-00-000049.tif; 36847 x 32516 pixels).</p> <p><strong>NG6437_vis_crop_01_230502.tif</strong>: 1543 x 2078 pixel detail taken from the high resolution visible image described above.</p> <p><strong>NG6437_xray_crop_01_230502.tif</strong>: 1543 x 2078 pixel detail of the X-radiograph corresponding to NG6437_vis_crop_01_230502.tif. The X-ray images were acquired using sheets of film (27 November 2019) and 16-bit digital scans were then produced (original files: N-6437-00-000047-009 and -014 (each 9539 x 7199 pixels), processed 28 January 2020). This crop is an 8-bit composite image of 2 X-ray plates that had been manually registered to the high resolution visible image described above using Adobe Photoshop.</p> <p><strong>NG6437_vis_crop_02_230502.tif</strong>: 1562 x 2023 pixel detail taken from the high resolution visible image described above.</p> <p><strong>NG6437_xray_crop_02_230502.tif</strong>: 1543 x 2078 pixel detail of the X-radiograph corresponding to NG6437_vis_crop_02_230502.tif. The X-ray images were acquired using sheets of film (27 November 2019) and 16-bit digital scans were then produced (original files: N-6437-00-000047-002 and -007 (each 9539 x 7199 pixels), processed 28 January 2020). This crop is an 8-bit composite image of 2 X-ray plates that had been manually registered to the high resolution visible image described above using Adobe Photoshop.</p> <p><strong>NG6437_vis_crop_03_230502.tif</strong>: 2088 x 2088 pixel detail taken from the high resolution visible image described above.</p> <p><strong>NG6437_xray_crop_03_230502.tif</strong>: 2088 x 2088 pixel detail taken from the composite X-radiograph described above corresponding to NG6437_vis_crop_03_230502.tif. </p> <p><strong>NG6437_vis_crop_04_230502.tif</strong>: 2088 x 2088 pixel detail taken from the high resolution visible image described above.</p> <p><strong>NG6437_xray_crop_04_230502.tif</strong>: 2088 x 2088 pixel detail taken from the composite X-radiograph described above corresponding to NG6437_vis_crop_04_230502.tif. </p> <p> </p>
Datasets and codes for the peer review article "Human and natural impacts on the U.S. freshwater salinization and alkalinization: A machine learning approach"
<p>Ongoing salinization and alkalinization in U.S. rivers have been attributed to inputs of road salt and effects of human-accelerated weathering in previous studies. Salinization poses a severe threat to human and ecosystem health, while human derived alkalinization implies increasing uncertainty in the dynamics of terrestrial sequestration of atmospheric carbon dioxide. A mechanistic understanding of whether and how human activities accelerate weathering and contribute to the geochemical changes in U.S. rivers is lacking. To address this uncertainty, we compiled dissolved sodium (salinity proxy) and alkalinity values along with 32 watershed properties ranging from hydrology, climate, geomorphology, geology, soil chemistry, land use, and land cover for 226 river monitoring sites across the coterminous U.S. Using these data, we built two machine-learning models to predict monthly-aggregated sodium and alkalinity fluxes at these sites. The sodium-prediction model detected human activities (represented by population density and impervious surface area) as major contributors to the salinity of U.S. rivers. In contrast, the alkalinity-prediction model identified natural processes as predominantly contributing to variation in riverine alkalinity flux, including runoff, carbonate sediment or siliciclastic sediment, soil pH and soil moisture. Unlike prior studies, our analysis suggests that the alkalinization in U.S. rivers is largely governed by local climatic and hydrogeological conditions.</p>
SeaFlux v2023: harmonised sea-air CO2 fluxes from surface pCO2 data products using a standardised approach
<p><strong>BE SURE TO DOWNLOAD 2023.02</strong></p> <p>See the additional notes for updates on the products. </p> <p>Fluxes calculated using the standardized approach:</p> <p> \(F\text{CO}_2=K_0 \cdot K_w \cdot (p\text{CO}_2^\text{sea} - p\text{CO}_2^\text{atm})\ \cdot (1 - [ice])\).</p> <p>We provide each of the components to this equation to reduce the potential for errors in fluxes due to methodological differences.</p> <p>The netCDF files contain the following data (<strong>note that only bold names have been updated in v2023</strong>): </p> <ul> <li>fgco2_all_winds_products: the sea-air CO2 flux for all spCO2 products (6) and <em>kw</em> from all wind products (5). </li> <li>fgco2_global:<strong> </strong>the globally integrated sea-air CO2 fluxes for all spCO2 products (6) and <em>kw</em> from all wind products (6)</li> <li><strong>sol:</strong> \(K_0\) is calculated using the Weiss (1974) parameterization with EN4 salinity and OISST temperatures </li> <li><strong>kw:</strong> \(k_w\) is calculated for winds with each being scaled independently to a 14-C bomb flux estimate of 16.5 cm/hr using the quadratic formulation by Wanninkhof (1992). <ul> <li>CCMPv2</li> <li>ERA5</li> <li>JRA55</li> <li>NCEP1</li> <li>NCEP2</li> </ul> </li> <li>spco2_SOCOM_unfilled<em>: </em>\(p\text{CO}_2^\text{sea}\) downloaded from various sources contains the following products: <ul> <li>CMEMS_FFNN</li> <li>CSIR_ML6</li> <li>JENA_MLS</li> <li>JMA_MLR</li> <li>MPI_SOMFFN</li> <li>NIES_FNN</li> </ul> </li> <li>spco2_filler<em>: </em>scaled version of the Landschützer et al. (2020) climatology used to fill missing regions of <em>spco2_SOCOM_unfilled</em></li> <li><strong>fco2atm: </strong>\(p\text{CO}_2^\text{atm}\) is calculated from NOAA's marine boundary layer product with ERA5 mean sea level pressure corrected for pH2O. The virial coefficient is then applied to pCO2atm</li> <li><strong>ice: </strong>\([ice]\) is the ice fraction from the OISST product</li> <li><strong>area_ocean:</strong><em> </em>the surface area of the ocean including the fractional area of the coastal regions</li> <li><strong>seafrac: </strong>the fraction of a pixel that is ocean</li> </ul> <p><strong><em>Units are listed in the metadata of each of the netCDF variables. </em></strong></p>
Data for paper "Stratocumulus adjustments to aerosol perturbations disentangled with a causal approach"
<p>Timeseries data used for the causal effect estimation of the paper "Stratocumulus adjustments to aerosol perturbations disentangled with a causal approach".</p> <p>This dataset contains several cloud parameters and meteorological co-variates corresponding to the evolution of the South-East Atlantic stratocumulus deck for the time period January 2016 to December 2017 and the spatial domain [lon1,lon2,lat1,lat2]=[0, 10, -20, -10]. </p> <p>The processing code used to generate the timeseries data, as well as the analysis code are uploaded separately on Zenodo. The input raw satellite and reanalysis data for the processing code are from EUMETSAT (Copyright (c) (2020) EUMETSAT), NASA and COPERNICUS data (generated using Copernicus Climate Change Service information [2022]). </p> <p> </p> <p>Citations for the raw data sources: </p> <p>Finkensieper, S., Meirink, J.-F., van Zadelhoff, G.-J., Hanschmann, T., Benas, N., Stengel, M., Fuchs, P., Hollmann, R., Kaiser, J., Werscheck, M.: CLAAS-2.1: CM SAF CLoud property dAtAset using SEVIRI - Edition 2.1. Satellite Application Facility on Climate Monitoring (2020). <a href="https://doi.org/10.5676/EUM_SAF_CM/CLAAS/V002_01">https://doi.org/10.5676/EUM_SAF_CM/CLAAS/V002_01</a></p> <p>Huffman, G.J., Stocker, E.F., Bolvin, D.T., Nelkin, E.J., Tan, J.: GPMIMERG Final Precipitation L3 Half Hourly 0.1 degree x 0.1 degree V06. MD, Goddard Earth Sciences Data and Information Services Center (GES DISC) (2019). <a href="https://doi.org/10.5067/GPM/IMERG/3B-HH/06">https://doi.org/10.5067/GPM/IMERG/3B-HH/06</a>.</p> <p>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horanyi, A., Munoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., Th ́ebaut, J.-N.: ERA5 hourly data on single levels from 1959 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS) (2018). <a href="https://doi.org/10.24381/cds.adbb2d47">https://doi.org/10.24381/cds.adbb2d47</a></p> <p>Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horanyi, A., Munoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., Th ́ebaut, J.-N.: ERA5 hourly data on pressure levels from 1959 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS) (2018). <a href="https://doi.org/10.24381/cds.bd0915c6">https://doi.org/10.24381/cds.bd0915c6</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.