Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
97
datasets available to search
ShareScore release 0.9.0
Dataset results
97 results for “affective dataset”
Datasets for sandboxing use case SUC3 corresponding to cyber attacks affecting the differential protection scheme of a HV transformer
<p><span>These datasets reflect two main scenarios (S1-S2) associated to the operation of a sandboxing use case SUC3 corresponding to cyber attacks affecting the differential protection scheme of a HV transformer. Details about are illustrated in Section 1.3 of the supporting document. These scenarios analyse the operation of the digital twin of the IEEE 9-bus system and the differential protection scheme under healthy conditions, cyber-attack on communication channels of IEC 61850 Sample Values (SVs) protocol, and a fault in HV side of a transformer in the power system. The scenarios are presented with selected time-series plots in Section 1.3, accompanied a detailed analysis of the processes included and an impact assessment. Thus, d</span><span>uring execution of each scenario, data such as electrical measurements were captured and are collected</span> in the form of the datasets presented here.</p> <p>Specifically, </p> <ul> <li>SUC3/S1 <strong>Differential protection operation during transformer fault</strong> corresponds to the dataset of first scenario (S1) of the third sandboxing use case (SUC3) of the KIOS CoE Sandboxing for cyber-physical analysis of EPES, which examines the operation of differential protection scheme (implemented in Typhoon controller) for a HV/MV transformer. The protection scheme receives data sent through IEC 61850 SVs from the two sides of the transformer. Specifically, this dataset corresponds to the first scenario (S1) of SUC3, where a short-circuit occurred on the HV side of a HV/MV transformer of the system. More details about the scenario related to this dataset can be found in Section 1.3.1 of the supporting document. This dataset includes electrical measurements of the current flow, in RMS and sinusoidal format, from the HV and MV sides of HV/MV transformer of the digital twin of the IEEE 9-bus system. The dataset is provided in the form of time-series measurements available as MATLAB (.mat) and CSV files which were recorded with a 30-second and 40-second time resolution, respectively. The measurements of RMS values were recorded by the Typhoon controller, while the sinusoidal measurements were recorder by OPAL-RT.</li> <li>SUC3/S2 <strong>MITM with FDI cyber-attack in the SVs of HV transformer side</strong> corresponds to the dataset of the second scanario (S2) of the third sandboxing use case (SUC3) of the KIOS CoE Sandboxing for cyber-physical analysis of EPES, which examines a MITM with FDI cyber-attack is conducted on the measurements of the HV side of the transformer, virtually implemented within the sandboxing, and introduces a multiplicative change to the current measurements before they are received by the differential protection scheme via IEC 61850 protocol. Section 1.3.1 of the supporting document provides more details about the scenario related to this<br>dataset. This dataset includes electrical measurements of the current flow, in RMS and sinusoidal format, from the HV and MV sides of HV/MV transformer of the digital twin of the IEEE 9-bus system. The dataset is provided in the form of time-series measurements available as MATLAB (.mat) and CSV files which were recorded with a 30-second and 40-second time resolution, respectively. The measurements of RMS values were recorded by the Typhoon controller, while the measurements from the sine waves were recorder by OPAL-RT.</li> </ul>
Main dataset 'Environmental specificity in Drosophila-bacteria symbiosis affects host developmental plasticity'
<p>Main dataset from the manuscript 'Environmental specificity in <em>Drosophila</em>-bacteria symbiosis affects host developmental plasticity' (2019)</p>
MANOVA dataset 'Environmental specificity in Drosophila-bacteria symbiosis affects host developmental plasticity'
<p>MANOVA dataset from the manuscript 'Environmental specificity in <em>Drosophila</em>-bacteria symbiosis affects host developmental plasticity' (2019)</p>
Dataset associated with Senf et al. (2023): "How the extreme 2019-2020 Australian wildfire affected global circulation and adjustments"
<p>This contains data aggregates derived from global ECHAM-HAM simulation for the study for effects due to the extreme Australian wildfire event 2019/2020. This data build the basis for analysis and figures in Senf et al. (2023) submitted to ACP.</p> <p> </p> <p>Simulation Data are</p> <ul> <li>available for freely running ensembles (36 member) and nudged simulations</li> <li>conducted for fire emissions artificially scaled with factors 0, 1, 2, 3, 5.</li> <li>stored for Jan - Mar 2020</li> </ul>
Dataset related to article: "Congenital insensitivity to pain a novel mutation affecting a U12-type intron causes multiple aberrant splicing of SCN9A"
<p>raw data related to article reported at title</p>
Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset
<p><strong>Licensing</strong></p> <p>This dataset is adapted from the Group Affect and Performance dataset which is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. <a href="https://creativecommons.org/licenses/by-nc/4.0/">https://creativecommons.org/licenses/by-nc/4.0/</a></p> <p> </p> <p><strong>Description</strong></p> <p>This dataset contains the audio files containing manually annotated cases of overlapped utterances, classified into True Interruptions and False Interruptions. It is derived from the <a href="https://github.com/gmfraser/gap-corpus/tree/master">Group Affect and Performance</a> dataset created by the University of the Fraser Valley, Canada. Original conversation transcripts and audio files have been supplied for context. The Group Affect and Performance dataset provides a rich source of interruptions and overlapped utterances in general, yielding 200 True Interruptions from 355 instances of overlapped utterances in the 14 Group meetings which were annotated.</p> <p> </p> <p><strong>Structure</strong></p> <p>This dataset is structured into three parts:</p> <p>1. data.json contains a list of all instances of overlapped utterances, classified into ‘interruption’ and ‘non-interruption’ corresponding to True and False Interruptions respectively. Each instance is uniquely identified by the Group in which it occurred, the speaker and the starting time of the utterance.</p> <p>2. The 'audio' directory contains the audio of each instance of overlapped utterances corresponding to those found in data.json. The naming convention of the files is as such: ‘Group [group number]: [utterance start time] - [utterance end time].wav’.</p> <p>3. Also included is a copy of the original dataset which includes the full audio and transcript. This allows the full meeting to be heard and any context for interruptions to be evaluated.</p> <p>Note that directories 2. and 3. can be accessed by unzipping audio-and-transcripts.zip.</p> <p> </p> <p><strong>Data Collection Protocol</strong></p> <p>Of paramount importance to our process are the definitions of an overlapped utterance and a True Interruption. A False Interruption is simply an overlapped utterance which is not a True Interruption. These definitions directly impact the dataset; for overlapped utterance it informs which data points are included in our dataset and for True Interruption it informs the classes assigned to each sample.</p> <p>In defining an overlapped utterance, our primary aim is to create an overarching class encompassing interruptions and all instances that could be deemed a True Interruption. For this reason, we omit cases where the timing misplaced speech and early-onset responses.</p> <p>An overlapped utterance is defined as an instance where one interlocutor provides speech or noise during another interlocutor’s speech, creating an overlap that may be deemed a possible interruption when considering its timing alone. For this reason we omit cases of where the timing indicates misplaced speech or early-onset responses.</p> <p>Our definition of True Interruption is an instance where an interrupting party intentionally attempts to take over a turn of the conversation from an interruptee and, in doing so, creates an overlap in speech.</p> <p>As previously mentioned, due to the ‘intent’ part of this definition, we avoid cases of misplaced speech and early-onset responses. The former is enforced by not considering cases of overlapped speech which begin within 300ms of each other since this is an estimate for the average human reaction time of articulating a vowel in response to a speech stimuli. The latter is enforced by not considering speech starting within the last 10% of first utterance in the overlapped speech. Note that this approach fails to filter out all cases of misplaced speech, so we manually remove the remaining instances.</p> <p> </p> <p><strong>Methodology</strong></p> <p>Three main steps were taken to produce this dataset:</p> <p>1. Parsing the transcripts for cases of overlapping speech</p> <p>2. Manually annotating these cases per our protocol and adding them to data.json</p> <p>3. Extracting audio samples from data.json and adding them to the audio folder</p> <p> </p> <p>If you use this dataset, please cite the following paper:</p> <p> </p> <p>Doyle, D.; Şerban, O. Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset. <em>Data</em> <strong>2024</strong>, <em>9</em>, 104. https://doi.org/10.3390/data9090104</p> <p> </p> <p>@article{data9090104,</p> <p>AUTHOR = {Doyle, Daniel and Şerban, Ovidiu},</p> <p>TITLE = {Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset},</p> <p>JOURNAL = {Data},</p> <p>VOLUME = {9},</p> <p>YEAR = {2024},</p> <p>NUMBER = {9},</p> <p>ARTICLE-NUMBER = {104},</p> <p>URL = {https://www.mdpi.com/2306-5729/9/9/104},</p> <p>ISSN = {2306-5729},</p> <p>DOI = {10.3390/data9090104}</p> <p>}</p>
Dataset for the collected responses for the items measuring constructs affecting eHS non-acceptance behavior in Nigeria
<p>This dataset is a collection of responses from the questionnaire distributed to study the e-health service non-acceptance behavior prominent in Nigeria. A total of 543 valid responses were collected. This research used an integration model based on TPB and SOR theory The dataset were analysed using PLS-SEM. Refer to the article for the results of this study.<br><br></p> <p><strong>Note:</strong> CO = Communication overload; CHO = Choice overload; PR = Perceived oisk; HL = Health literacy; NA = Negative attitude; SN = Subjective norms; PBC = Perceived behavioral control; INTU = Intention not to use eHS; NAB = Non-acceptance behavior</p>
Dataset - How do you propose your code changes? Empirical Analysis of Affect Metrics of Pull Requests on GitHub
<p>This package contains the raw open data for the study </p> <p>Marco Ortu, Giuseppe Destefanis, Daniel Graziotin, Michele Marchesi, Roberto Tonelli. 2020. How do you propose your code changes? Empirical Analysis of Affect Metrics of Pull Requests on GitHub. Under Review.</p> <p>The dataset is based on GHTorrent dataset:</p> <p>Georgios Gousios. 2013. The GHTorent dataset and tool suite. In Proceedings of the 10th Working Conference on Mining Software Repositories (MSR ’13). IEEE Press, 233–236</p> <p>And released with the same license (CC BY-SA 4.0).</p>
Source code and datasets used to link new waves of plague outbreaks in medieval Europe to climate fluctuations affecting the reservoirs of the disease in Asia.
<p>The zipfile contains the project directory which includes the source code and datasets used in the paper on <strong>Climate-driven introduction of the Black Death and successive plague reintroductions into Europe</strong>, as published in <em>Proceedings of the National Academy of Sciences</em> (PNAS). Access the paper at http://www.doi.org/pnas.1412887112</p> <p>If you are not familiar with Clojure, Leiningen, and its project directory format, see http://clojure.org/getting_started for one of the IDE's to run the clojure code in, and use http://leiningen.org/ as the project / dependency manager.</p>
ReCANVo: A Dataset of Real-World Communicative and Affective Nonverbal Vocalizations
<p>A dataset of 7077 labeled vocalizations made by non-speaking individuals. Each vocalization lasts approximately 0.5-4 seconds and is labeled with its affective or communicative meaning. Data were acquired in real-world settings (homes, schools, etc.) and were labeled in real-time by parents or caregivers who knew the non-speaking communicator well. </p> <p>dataset_file_directory.csv provides the name of each vocalization file, the corresponding participant ID, and the vocalization meaning or label (delighted, frustrated, request, etc.).</p> <p>If you use this dataset, please <strong>cite Johnson & Narain et al., "ReCANVo: A Database of Real-World Communicative and Affective Nonverbal Vocalizations"</strong>. The authors are Jaya Narain, Kristina T. Johnson, Thomas Quatieri, Pattie Maes, and Rosalind Picard. This paper provides more information about the dataset, including data acquisition methodology, pre-processing procedures, and participant demographics. </p> <p>**J.N. and K.T.J. are joint first authors on this project. Please include both names in attribution when possible (e.g., Johnson & Narain et al.).</p>
Dataset for: Sensory environment affects Icelandic threespine stickleback's anti-predator escape behaviour
<p>Human-induced changes in climate and habitats push populations to adapt to novel environments, including new sensory conditions, such as reduced visibility. We studied how colonizing newly formed glacial lakes with turbidity-induced low visibility affects anti-predator behaviour in Icelandic threespine sticklebacks. We tested nearly 400 fish from 15 populations and four habitat types varying in visibility and colonization history in their reaction to two predator cues (mechano-visual versus olfactory) in high versus low visibility light treatments. Fish reacted differently to the cues and were affected by lighting environment, confirming that cue modality and light levels are important for predator detection and evasion. Spring-fed fish, especially from the highlands (likely more diverged from marine fish than lowland fish) reacted fastest to mechano-visual cues and were generally most active. Highland glacial fish showed strong responses to olfactory cues and, counter to predictions from the flexible stem hypothesis, the greatest plasticity in response to light levels. This study, leveraging natural, repeated invasions of novel sensory habitats, 1) illustrates rapid changes in antipredator behaviour that follow due to adaptation, early life experience, or both, and 2) suggests an additional role for behavioural plasticity enabling population persistence in the face of frequent changes in environmental conditions.</p>
Dataset for manuscript entitled: Switchgrass cropping systems affect soil carbon and nitrogen and microbial diversity and activity on marginal lands
<p class="MsoListParagraph">Switchgrass (<em>Panicum virgatum</em> L.),<span> </span>as a dedicated bioenergy crop, can provide cellulosic feedstock for biofuel production while improving or maintaining soil quality. However, comprehensive evaluations of how switchgrass cultivation and nitrogen (N) management impact soil and plant parameters remain incomplete. We conducted<span> </span>field trials in three years (2016–2018) at six locations in the North Central Great Lakes Region to evaluate the effects of cropping systems (switchgrass, restored prairie, undisturbed control) and N rates (0, 56 kg N ha<sup>-1</sup> yr<sup>-1</sup>) on biomass yield and soil physicochemical, microbial, and enzymatic parameters. Switchgrass cropping system yielded an aboveground biomass 2.9–3.3 times higher than the other two systems (Jayawardena et al., In submission) but our study found that this biomass accumulation didn't reduce soil dissolved organic C (DOC), total dissolved N (TDN), or bacterial diversity. The annual aboveground biomass removal for bioenergy feedstock, however, reduced soil microbial biomass C (MBC) and N (MBN) and bacterial richness in the 2<sup>nd</sup> and 3<sup>rd</sup> years; despite this, continuous monocropping of switchgrass improved soil TDN, inorganic N, bacterial diversity, and shoot biomass in the 2<sup>nd</sup> and/or 3<sup>rd</sup> years when compared to the 1<sup>st</sup> year. N fertilization increased aboveground biomass yield by 1.2 times and significantly increased soil TDN, MBN, and the shoot biomass of switchgrass when compared to the unfertilized control. Locations with higher C and N contents and lower C:N ratio had higher aboveground biomass, MBC, MBN, and the activity of BG, CBH, and UREA enzymes; by contrast, locations with higher pH had higher soil TDN and activity of NAG and LAP enzymes. Our research demonstrates that switchgrass cultivation could improve or maintain soil N content and N fertilization can increase plant biomass yield. The comprehensive data also can inform future biogeochemical models to successfully implement switchgrass for bioenergy production.</p>
Dataset of the research article "A semi-automated analysis of displacement-to-length scaling of the grabens affecting lunar Floor-Fractured craters"
<p>This dataset includes all the raster data (DEM, orthoimage, slope) and vectors (shapefiles) used in our research, investigating the relationships between displacement and length of the faults affecting Lunar Floor-Fractured craters.</p>
WorkingAge In-lab Facial Affect Dataset
<p>The automatically detected frame-level facial AU features for WorkingAge In-lab test visual dataset.</p>
Dataset to adjusted spectral correction method for calculating extreme winds in tropical cyclone affected water areas
<p>This is a dataset of the 50-year wind of an effective temporal resolution of 10 min over three areas with the presence of tropical cyclones at 10 m, 50 m, 100 m and 150 m.</p> <p>There are 18 files in total, with 12 files for the 50-year winds:</p> <p>'XXcfsru50atYYmcorr_revision.dat'</p> <p>and 6 files for the corresponding latitudes and longitudes:</p> <p>'TC_XX_lat.dat' and 'TC_XX_lon.dat'</p> <p>The three areas are indexed as E1, W1, W2 (as in the file names 'XX').</p> <p>The heights are 10 m, 50 m, 100 m and 150 m (as in the filenames 'YY').</p> <p> </p> <p> </p>
Autophagosomal/autolysosomal/lysosomal dynamics of FaDu and HGFb cells affected by autophagy modulators (dataset of dual labeling of autophagosomes and labeling lysosomes, confocal microscopy)
<p>In this dataset, the impact of autophagy modulators on the auophagosomes and lysosomes in FaDu and HGF cells was investigated. Experimental details are described in the accompanying paper. This dataset contains total 702 CZI confocal image Z-stacks.</p> <p><strong>Relevant paper</strong>: HANELOVA, Klara, RAUDENSKA, Martina, KRATOCHVILOVA, Monika, NAVRATIL, Jiri, VICAR, Tomas, BUGAJOVA, Maria, GUMULEC, Jaromir, MASARIK, Michal and BALVAN, Jan. Autophagy modulators influence the content of important signalling molecules in PS-positive extracellular vesicles. <em>Cell Communication and Signaling</em>. 24 May 2023. Vol. 21, no. 1, p. 120. DOI <a href="https://doi.org/10.1186/s12964-023-01126-z">10.1186/s12964-023-01126-z</a>.</p> <div> <div> <div> </div> </div> </div> <p><strong>Model cell lines</strong></p> <p> </p> <p>The cell line FaDu (HTB-43TM), derived from a squamous cell carcinoma (SCC) of the hypopharynx, and the human gingival fibroblast cell line HGF (derived from the histologically normal gingival biopsy) were used in this study. The authenticated cell lines were purchased from the American Type Culture Collection (ATCC; Manassas, Virginia, USA) within the last five years. </p> <p><strong>Autophagy modulation</strong></p> <p>For autophagy modulation, FaDu cells were treated for 24 h with 5nM bafilomycin A1 (Sigma-Aldrich, B1793), 50 µM of hydroxychloroquine sulphate (Sigma-Aldrich, H0915), 100 µM of Cpd18 (Calbiochem), 50 nM of autophinib (Sigma, SML2632), 10 µM of EACC (MedChemExpress), 200 nM of rapamycin (Sigma-Aldrich, R0395), 3 nM of Torin-1 (MedChemExpress), or 30 nM of NVP-BEZ235 (MedChemExpress). To induce starvation, cells were cultured in DMEM F12 without glutamine and without FBS (Biosera). Modulation of autophagy did not reduce the viability of FaDu cells.</p> <p><strong>Conditioned media preparation</strong></p> <p>see details in the accompanying paper</p> <p><strong>Fluorescence Microscopy</strong><br>The autophagosomal/autolysosomal/lysosomal dynamics of affected cells were observed using the combination of PremoTM Autophagy Tandem Sensor (P36239, Invitrogen) with the far-red emitting LysoTracker® Deep Red (L12492, Invitrogen) (Ex 647 nm/Em 668 nm). By combining acid-sensitive Emerald GFP (Ex 488 nm/Em 509 nm) with acid-insensitive TagRFP (Ex 555 nm/Em 584 nm) in the PremoTM kit, autophagosomes and autolysosomes labelling (yellow and red, respectively) is possible.<br>Immediately after transduction with 12 µl PremoTM Autophagy Tandem Sensor/2ml cell suspension, cells were seeded at 5 × 10<sup>5</sup> into 35-mm glass-bottomed gelatin-coated dishes (Ibidi, μ-Dish 35 mm, high Glass Bottom) and cultured for 48 h to equilibrate expression levels. Subsequently, cells were exposed to the selected agents for 6, 12, 24 and 48 hours before imaging. LysoTracker® Deep Red staining was performed 1 h before imaging. <br>To monitor the uptake of isolated PS-EVs by fibroblasts, we stained EVs with PKH67 (Sigma, PKH67GL) and then removed the remaining dye using Exosome Spin Columns (MW 3000) (Thermo Scientific, #4484449). The stained EVs were then suspended in 400 ul of cultivation medium and added to HGFB cells growing for 24 h in Ibidi µ-Slide I Luer (Ibidi, 80176). Image acquisition was performed 24 hours after EVs addition. 1 µl of 1 µg/ml of Hoechst 33342 (Enzo) (Ex 350 nm/ Em 461 nm) nuclear stain was added 1 hour before imaging.<br>To determine the viability of the cell population prior to isolation of EVs, cells were left in a culture dish with 1 ml of culture medium to which propidium iodide (Sigma-Aldrich) and Hoechst 33342 were added 45 minutes before capturing. <br>To maximize the possibility of comparison between samples, all samples (from a single cell line) were captured on the same day in a single run using the same microscope settings. For each time and each treatment, 10–12 fields of view were captured from randomized sites of the culture dish. Epifluorescent microscopy images and confocal microscopy images were acquired using Laser scanning confocal microscope Zeiss LSM 880 with AiryscanFast module (Carl Zeiss Inc.) using a C-apochromat 40x/1.20 W and C-Apochromat 63 /1.20 W. LysoTracker® Deep Red was excited HeNe 633 nm solid-state laser and emitted light was detected at 638–759 nm. Emerald GFP was excited 488 nm ArgonRemote laser, and emitted light was detected at 493–576 nm. TagRFP was excited DPSS 561 nm laser, and emitted light was detected at 570–650 nm. Hoechst 33342 was excited with a 405 nm solid-state laser, and emitted light was detected at 410–508 nm. Fluorescence images were acquired by the transmitted light detector. <br>Images were analyzed using ImageJ software and custom MATLAB software developed in our laboratory. The analysis process consists of the segmentation of cells from the background and extraction of the intensity of fluorescence channels (TagRFP, GFP, LysoTracker) inside cells. For segmentation, a thresholding-based method was used, where a manually selected threshold was applied. Segmentation was applied to the image created as the sum of all fluorescence channels to achieve segmentation independent of the intensity of individual channels. To achieve better segmentation without noisy pixels, fluorescence images were preprocessed with median filter (7x7) and Gaussian filter (standard deviation 1); additionally, binary segmentation was post-processed with morphological closing and removal of small binary connected components (<5000px). For intensity extraction, the mean value of segmented cell pixels was used for each field of view. Besides individual fluorescence channels, the mean colocalization of TagRFP and GFP was calculated with a pixel-wise multiplication of TagRFP and GFP channels.</p> <p><strong>File Naming</strong></p> <p>CZI images are organised in folders according to cell line, treatment time, and treatment. Names include magnification, cell line measured, treatment, and time of treatment,</p> <blockquote> <p>40x_FADU_BAF_12h_BF_4.czi</p> </blockquote> <p>The abbreviations used for treatments are as follows: BAF, bafilomycin A1; HCQ, hydroxychloroquine sulphate; Cpd18; APB, autophinib; EACC; RAPA, rapamycin; TOR1, Torin-1; BEZ, NVP-BEZ235; starv, starvation; </p>
Dataset from Competition and drought affect cleistogamy in a non-additive way in the annual ruderal Lamium amplexicaule
<p>Data associated with the manuscript 23147-TS1R1 published in AoB Plants. For detailed information about files and their content see ReadMe</p>
Dry ECG Dataset: Normal and Interference-Affected Records (NSIR Dataset)
<p><strong>The dataset consists of 149 ECG signals that were sampled at a frequency of 500 Hz using a single lead. </strong></p> <p><strong>The main objective</strong> is to identify and differentiate normal-recorded ECG signals from those affected by two common types of interference:</p> <ul> <li>50/60 Hz power line interference .</li> <li> electrode contact noise caused by unstable movement.</li> </ul> <p>The data was recorded at the LINS Laboratory within the University of USTHB in Algeria using <strong>the Orbital 90 dry electrode</strong>, which has a 25 mm conductive area diameter based on the three electrodes configurations to measure the potential difference between the two arms, with the right leg serving as the reference. In order to improve signal quality, a fourth-order Butterworth low-pass filter was utilized for its smooth frequency response and minimal phase distortion, followed by a notch filter to eliminate 50 Hz power line noise. These filtering processes significantly enhanced the ECG signal quality by reducing noise and interference while preserving important cardiac information. The recorded data was stored in <strong>CSV</strong> format using UART communication via a USB connected to a computer.</p> <ul> <li><strong>Dataset Summary</strong></li> </ul> <table> <tbody> <tr> <td><strong>Type of Record </strong></td> <td><strong>Number of Records</strong></td> </tr> <tr> <td>well-recorded ECG</td> <td>42</td> </tr> <tr> <td>50/60 Hz power line interference</td> <td>44</td> </tr> <tr> <td>electrode contact noise</td> <td>63</td> </tr> <tr> <td><strong>Total</strong></td> <td>149</td> </tr> </tbody> </table> <p> </p> <p><strong>Note :</strong></p> <ol> <li><strong>You can find in the attachment a set of scalogram representations of the data with different wavelets, converted using the CWT technique.</strong></li> <li><strong>These data were used in transfer learning and gave a remarkable result, highlighting the unnecessary need for a huge dataset to train a transfer learning algorithm when the data is well recorded and the classes are distinguished.</strong></li> </ol> <p> </p> <ul> <li>For more details please feel free to contact us :</li> </ul> <ol> <li>Kharziwisseme@hotmail.com</li> <li>kerdjidjoussama@gmail.com</li> <li>malikakedir@gmail.com</li> <li>nac.meziane@gmail.com</li> </ol>
Ecosystem-Level Factors Affecting the Survival of Open-Source Projects: A Case Study of the PyPI Ecosystem - the dataset
<pre><em>Replication pack, FSE2018 submission #164: </em><em>------------------------------------------ </em></pre> <pre><strong>**</strong>Working title:<strong>** </strong>Ecosystem-Level Factors Affecting the Survival of Open-Source Projects: A Case Study of the PyPI Ecosystem <strong>**</strong>Note:<strong>** </strong>link to data artifacts is already included in the paper. Link to the code will be included in the Camera Ready version as well. <em>Content description </em><em>=================== </em> <strong>- **</strong>ghd-0.1.0.zip<strong>** </strong>- the code archive. This code produces the dataset files described below <strong>- **</strong>settings.py<strong>** </strong>- settings template for the code archive. <strong>- **</strong>dataset_minimal_Jan_2018.zip<strong>** </strong>- the minimally sufficient version of the dataset. This dataset only includes stats aggregated by the ecosystem (PyPI) <strong>- **</strong>dataset_full_Jan_2018.tgz<strong>** </strong>- full version of the dataset, including project-level statistics. It is ~34Gb unpacked. This dataset still doesn't include PyPI packages themselves, which take around 2TB. <strong>- **</strong>build_model.r, helpers.r<strong>** </strong>- R files to process the survival data (`survival_data.csv` in <strong>**</strong>dataset_minimal_Jan_2018.zip<strong>**</strong>, `common.cache/survival_data.pypi_2008_2017-12_6.csv` in <strong>**</strong>dataset_full_Jan_2018.tgz<strong>**</strong>) <strong>- **</strong>Interview protocol.pdf<strong>** </strong>- approximate protocol used for semistructured interviews. <strong>- </strong>LICENSE - text of GPL v3, under which this dataset is published <strong>- </strong>INSTALL.md - replication guide (~2 pages)</pre> <pre><em>Replication guide </em><em>================= </em> <em>Step 0 - prerequisites </em><em>---------------------- </em> <strong>- </strong>Unix-compatible OS (Linux or OS X) <strong>- </strong>Python interpreter (2.7 was used; Python 3 compatibility is highly likely) <strong>- </strong>R 3.4 or higher (3.4.4 was used, 3.2 is known to be incompatible) Depending on detalization level (see Step 2 for more details): <strong>- </strong>up to 2Tb of disk space (see Step 2 detalization levels) <strong>- </strong>at least 16Gb of RAM (64 preferable) <strong>- </strong>few hours to few month of processing time <em>Step 1 - software </em><em>---------------- </em> <strong>- </strong>unpack <strong>**</strong>ghd-0.1.0.zip<strong>**</strong>, or clone from gitlab: git clone https://gitlab.com/user2589/ghd.git git checkout 0.1.0 `cd` into the extracted folder. All commands below assume it as a current directory. <strong>- </strong>copy `settings.py` into the extracted folder. Edit the file: <strong> * </strong>set `DATASET_PATH` to some newly created folder path <strong> * </strong>add at least one GitHub API token to `SCRAPER_GITHUB_API_TOKENS` <strong>- </strong>install docker. For Ubuntu Linux, the command is `sudo apt-get install docker-compose` <strong>- </strong>install libarchive and headers: `sudo apt-get install libarchive-dev` <strong>- </strong>(optional) to replicate on NPM, install yajl: `sudo apt-get install yajl-tools` Without this dependency, you might get an error on the next step, but it's safe to ignore. <strong>- </strong>install Python libraries: `pip install --user -r requirements.txt` . <strong>- </strong>disable all APIs except GitHub (Bitbucket and Gitlab support were not yet implemented when this study was in progress): edit `scraper/init.py`, comment out everything except GitHub support in `PROVIDERS`. <em>Step 2 - obtaining the dataset </em><em>----------------------------- </em> The ultimate goal of this step is to get output of the Python function `common.utils.survival_data()` and save it into a CSV file: # copy and paste into a Python console from common import utils survival_data = utils.survival_data('pypi', '2008', smoothing=6) survival_data.to_csv('survival_data.csv') Since full replication will take several months, here are some ways to speedup the process: <em>####Option 2.a, difficulty level: easiest </em> Just use the precomputed data. Step 1 is not necessary under this scenario. <strong>- </strong>extract <strong>**</strong>dataset_minimal_Jan_2018.zip<strong>** </strong><strong>- </strong>get `survival_data.csv`, go to the next step <em>####Option 2.b, difficulty level: easy </em> Use precomputed longitudinal feature values to build the final table. The whole process will take 15..30 minutes. <strong>- </strong>create a folder `<DATASET_PATH>/common.cache`, where `<DATASET_PATH>` is the value of the variable `DATASET_PATH` in `settings.py` <strong>- </strong>extract <strong>**</strong>dataset_minimal_Jan_2018<strong>** </strong>to the newly created folder <strong>- </strong>rename files: mv backporting.csv monthly_data.pypi_backporting.csv mv cc_degree.csv monthly_data.pypi_cc_degree.csv mv commercial.csv monthly_data.pypi_commercial.csv mv commits.csv monthly_data.pypi_commits.csv mv contributors.csv monthly_data.pypi_contributors.csv mv dc_katz.csv monthly_data.pypi_dc_katz.csv mv downstreams.csv monthly_data.pypi_downstreams.csv mv d_upstreams.csv monthly_data.pypi_d_upstreams.csv mv github_user_info.csv user_info.pypi.csv mv issues.csv monthly_data.pypi_issues.csv mv non_dev_issues.csv monthly_data.pypi_non_dev_issues.csv mv non_dev_submitters.csv monthly_data.pypi_non_dev_submitters mv package_urls.csv package_urls.pypi.csv mv q90.csv monthly_data.pypi_q90.csv # raw_dependencies.csv is not required # raw_packages_info.csv is not required # Feel free to read README.md for more details about the data mv submitters.csv monthly_data.pypi_submitters.csv # In this scenario we'll generate a new survival_data.csv mv university.csv monthly_data.pypi_university.csv mv upstreams.csv monthly_data.pypi_upstreams.csv <strong>- </strong>edit `common/decorators.py`, set `DEFAULT_EXPIRY` to some higher value, e.g. `DEFAULT_EXPIRY = float('inf') # cache never expires` Then, use the Python code above to obtain `survival_data.csv`. <em>####Option 2.c, difficulty level: medium </em> Use predownloaded raw data to build longitudinal feature values, and then the dataset. Despite most of the data is cached, some functions will pull up updates which might take anywhere from days to couple weeks to run. <strong>- </strong>Download <strong>**</strong>dataset_full_Jan_2018.tgz<strong>** </strong>from http://k.soberi.us/dataset_full_Jan_2018.tgz . This file is not included in this archive because of its size (5.4Gb compressed, 34Gb unpacked). <strong>- </strong>edit `common/decorators.py`, set `DEFAULT_EXPIRY` to some higher value, e.g. `DEFAULT_EXPIRY = float('inf') # cache never expires` <strong>- </strong>extract the content of this archive into `<DATASET_PATH>`. <strong>- </strong>clean up `<DATASET_PATH>/common.cache` (otherwise you'll get Step 2.a. You can reproduce Step 2.b by deleting only `survival_data.pypi_2008_2017-12_6.csv`) Run the Python code above to obtain `survival_data.csv`. <em>####Option 2.d, difficulty level: hard </em> Build the dataset from scratch. Although most of the processing is parallelized, it will take at least couple months on a reasonably powerful server (32 cores, 512G of RAM, 2Tb+ of HDD space in our setup). <strong>- </strong>ensure the `<DATASET_PATH>` is empty <strong>- </strong>add more GitHub tokens (borrow from your coworkers) to `settings.py`. Run the Python code above to obtain `survival_data.csv`. <em>Step 3 - run the regression </em><em>--------------------------- </em> install R libraries: install.packages(c("htmlTable", "OIsurv", "survival", "car", "survminer", "ggplot2", "sqldf", "pscl", "texreg", "xtable")) Use `build_model.r` (e.g. in RStudio) and produced `survival_data.csv` to build the regressions used in the paper. This process takes at least 16Gb of RAM and takes few hours to run due to the gigantic size of the dataset. </pre>
Dataset of the article: "The social dimension of equine welfare: social contact positively affects the emotional state of stalled horses"
<p>Here are the data used in the analyses of the manuscript "The social dimension of equine welfare: social contact positively affects the emotional state of stalled horses.".<br>Experiment 1: Observations in the individual stall<br>Experiment 2: Observations during horse grooming<br>Experiment 3: Judgement bias test</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.