Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
419
datasets available to search
ShareScore release 0.9.0
Dataset results
419 results for “large dataset”
Dataset for algorithmic thinking skills assessment: Results from the virtual CAT large-scale study in Swiss compulsory education
<p><strong>Overview</strong><br>This dataset was collected during a main study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.<br>As algorithmic thinking becomes increasingly vital in our digital age, this study bridges the gap between traditional assessments and the needs of today's learners by introducing a digital platform. The virtual CAT, a digital adaptation of an unplugged assessment activity, offers scalable, automated assessments with reduced human intervention.</p> <p><strong>Study Context, Location and Participants</strong><br>To comprehensively investigate algorithmic competencies within compulsory education, exploring their variations and determining the factors influencing them, in Spring 2023 we conducted an experimental study with the virtual CAT's.<br>The sample comprises 129 students (65 girls and 64 boys), selected from nine classes across five public schools in Ticino and Solothurn cantons.</p> <p><strong>Data Collection</strong><br>During the data collection process, session and participant details were manually recorded by the administrator. <br>Each session has been assigned a unique identifier, and specific details, such as the date, canton, school name and type, and the students’ HarmoS grade (HG) level, have been recorded. <br>Student information are limited to sex and date of birth, with birth dates used to calculate ages, a significant factor in our demographic analysis. <br>To protect student privacy, unique identifiers have been assigned to each participant, keeping the data anonymous and secure. <br>The assessment tool automatically tracked all user interaction within the platform.<br>All data collected have been pseudonymised, aligning with prevailing open science practices in Switzerland (SNSF, 2021). <br>Data collection was integrated into a validation module of the app. </p> <p><strong>Data Features</strong><br>The dataset comprises the following files:</p> <ul> <li>STUDENTS_SESSIONS.csv</li> <li>RESULTS.csv</li> <li>LOGS.csv</li> <li>CANTONS.csv</li> <li>ALGORITHMS.csv</li> </ul> <p>These files collectively provide insights into the algorithmic actions of the students, demographic details, session logs, results, and more.</p> <p><strong>Usage & Ethics</strong><br>In the spirit of open science, this dataset is made available to the public after meticulous anonymisation to ensure all participants' privacy and ethical treatment. <br>Initial authorisations were secured from school administrators, teachers, and parents. <br>Detailed communication regarding the study's nature, data handling, and objectives was transparently shared with all stakeholders.</p> <p><strong>REFERENCES</strong></p> <p><strong>[1]</strong> A. Piatti, G. Adorni, L. El-Hamamsy, L. Negrini, D. Assaf, L. Gambardella & F. Mondada. (2022). The CT-cube: A framework for the design and the assessment of computational thinking activities. Computers in Human Behavior Reports, 5, 100166. <a href="https://doi.org/10.1016/j.chbr.2021.100166">https://doi.org/10.1016/j.chbr.2021.100166</a></p> <p><strong>[2]</strong> Adorni, G., & Piatti, S., & Karpenko, V. (2023). virtual CAT: An app for algorithmic thinking assessment within Swiss compulsory education. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10027851">https://doi.org/10.5281/zenodo.10027851</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/</a></p> <p><strong>[3]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT programming language interpreter. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10016535">https://doi.org/10.5281/zenodo.10016535</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/</a></p> <p><strong>[4]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT data infrastructure. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10015011">https://doi.org/10.5281/zenodo.10015011</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure</a></p> <p> </p>
nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in french watercourses
<p><strong>nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in French watercourses</strong></p> <p>Antoine Casquin, Marie Silvestre, Vincent Thieu</p> <p>10.5281/zenodo.10830852</p> <p>v0.1, 18<sup>th</sup> March 2024</p> <p><strong>Brief overview of data: </strong></p> <p>· Carbon and nutrients data for 5470 continental French catchments</p> <p>· Modelled discharge for 5128 of catchments out of 5470</p> <p>· Geopackages with catchment delineations and outlets</p> <p>· DEM conditioned to delimit additional catchments</p> <p>· Land-use and climatic data for 5470 continental French catchments</p> <p><strong>Citation of this work<br></strong></p> <p>A data paper is currently being submitted with details of methods and results. Once published, it will be the preferential source to cite. The data paper will be link to the new version of the dataset that will be updated on doi.org/10.5281/zenodo.10830852. If you use this dataset in your research or report, you must cite it.</p> <p><strong>Motivations</strong></p> <p>Data was collected and curated for the nuts-STeauRY project (<a href="http://nuts-steaury.cnrs.fr">http://nuts-steaury.cnrs.fr</a>), which deployed a national generic land to sea modelling chain.</p> <p>Data was primarily used (see related works):</p> <ol> <li>To calibrate concentrations of dissolved organic carbon and dissolve silica in headwaters</li> <li>To validate spatially and temporally the modelling chain (DOC, NO3-, NH4+, TP, SRP, DSi)</li> </ol> <p>Hydrochemical large sample datasets have numerous other uses: trends computations elucidate transfer mechanisms, machine learning, retrospective studies etc.</p> <p>The objective here is to provide a large sample curated dataset of carbon and nutrients concentrations along with modelled discharges, catchment characteristics and delimitations for the continental France. Such large sample dataset aims at easing the large sample studies over France and/or Europe. Although part of the data gathered here is obtainable via public sources, the catchments delineations, their characteristics and modelled hydrology were note not publicly available yet. Moreover, a unification of units and detection and removal of outliers was performed on carbon and nutrients data.</p> <p><strong>Data sources & processing</strong></p> <p>Sampling points where snapped on the CCM database v2.1 (<a href="http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221">http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221</a>)(Vogt et al., 2007) and catchments were delineated using a 100m resolution Digital Elevation Model (DEM) conditioned by the hydrographic network and elementary catchments’ delineations of the CCM data v2.1. <strong>More than 6000 catchments were delineated and screened manually</strong> to check consistency: 5470 were retained<strong>.</strong></p> <p>Nutrient data was collected mainly through the Naiades portal (<a href="https://naiades.eaufrance.fr/">https://naiades.eaufrance.fr/</a>), a database collecting water quality data produced by different water related actors across France. Nutrient data was also collected directly with regional water agencies (<a href="https://www.eau-seine-normandie.fr/">https://www.eau-seine-normandie.fr/</a>, <a href="https://eau-grandsudouest.fr/">https://eau-grandsudouest.fr/</a>, <a href="https://www.eaurmc.fr/">https://www.eaurmc.fr/</a>, <a href="https://www.eau-artois-picardie.fr/">https://www.eau-artois-picardie.fr/</a>, <a href="https://www.eau-rhin-meuse.fr/">https://www.eau-rhin-meuse.fr/</a> and <a href="https://agence.eau-loire-bretagne.fr/home.html">https://agence.eau-loire-bretagne.fr/home.html</a>), and pre-processed using a database management system relying on PostgreSQL with PostGIS extension (Thieu & Silvestre, 2015). A three-pass strategy was used to curate raw carbon and nutrients data: 1. Removal of “obvious outliers”, 2. Detection of baseline change and correction if possible (or removal of data) 3. Removal of outliers using a quantile based approach by element and temporal series.</p> <p>Hydrological time series are interpolation trough hydrograph transfer (de Lavenne et al., 2023) of 1664 time series of discharge completed with GR4J model (Pelletier & Andréassian, 2020; Pelletier 2021).</p> <p>Land cover data was extracted from Corine Land Cover dataset for years 2000, 2006, 2012, and 2018 (EEA, 2020). Raw CLC typology contains 44 classes. Results of percent cover per year per class were computed for each catchment. An aggregated typology of 8 classes is also proposed.</p> <p>Climatological data was extracted from daily reconstruction at 5 arcmin for temperatures and 1 arcmin for precipitation over Europe (Thiemig et al., 2022). Mean by catchment for min&max daily temperature and precipitation were computed for each catchment for the 1990-2019 period.</p> <p><strong>Nuts-STeauRY dataset</strong></p> <p><strong>Carbon and nutrients time series</strong></p> <p>Time series of carbon and nutrients within the 1962-2019 period on 5470 stations: Dissolved Organic Carbon (DOC), Total Organic Carbon (TOC) Nitrates (NO3-), Nitrites (NO2-), Ammonia (NH4+), Soluble Reactive Phosphorus (SRP), Total Phosphorus (TP) and Dissolved Silica (DSi).</p> <p><code>|var | n_unique_station| n_total_meas| mean_duration_y| mean_frequency_y|</code></p> <p><code>|:---|----------------:|------------:|---------------:|----------------:|</code></p> <p><code>|DOC | 4 992| 658 147| 14.3| 9.0|</code></p> <p><code>|DSi | 3 299| 333 866| 12.9| 8.3|</code></p> <p><code>|NH4 | 5 318| 907 343| 19.3| 8.7|</code></p> <p><code>|NO2 | 5 264| 891 886| 19.2| 8.6|</code></p> <p><code>|NO3 | 5 465| 939 279| 19.0| 9.0|</code></p> <p><code>|SRP | 5 361| 910 107| 19.1| 8.7|</code></p> <p><code>|TOC | 935| 111 993| 13.6| 9.6|</code></p> <p><code>|TP | 5 199| 802 841| 17.1| 8.8|</code></p> <p>Note that some SRP and DSi measurements were declared as realized on raw water. A thorough analysis of time series show no evidence of difference on baselines. For more accuracy, it is advised to filter out those analyses using the “fraction” attribute of each measurement.</p> <p><strong>Discharge modelled daily time series</strong></p> <p>Modelled naturalized discharge through hydrograph transfer and interpolated measured discharges when available for the 1980-2019 period.</p> <p>A daily discharge was computed for 5128 catchments. For small catchments (< 1000 km<sup>2</sup>, n = 4530), hydrograph transfer was used, while for big catchments, a direct interpolation of measured/completed discharges was performed. The direct interpolation was only possible for 598 catchments > 1000 km<sup>2</sup>. The criteria retained for a direct interpolation is 0.8*area_discharge_station < area_quality < 1.2*area_discharge_station when discharge and quality stations were nested.</p> <p>Hydrological time series uncertainties varies a lot depending on: quality of data source, distance from pseudo-gauged outlets, land cover of the catchments, natural spatial and temporal variability of discharge, size of the catchment (de Lavenne et al., 2016). We advise a cautious use of those modelled discharges as uncertainties could not be computed.</p> <p><strong>Catchments, outlets and conditioned DEM</strong></p> <p>5470 catchments and outlets are delivered as geopackages (EPSG: 3035).</p> <p>The DEM, conditioned by CCM 2.1 is also delivered as a GeoTIFF (EPSG: 3035) as way to delimit new catchment for the area that are consistent with the dataset.</p> <p><strong>Catchments characteristics and climate</strong></p> <p>Refer to Data sources & processing and File descriptions.</p> <p><strong> </strong></p> <p><strong>File and attributes descriptions: </strong></p> <p>The key “sta_code” is present across all files. For time varying records, “date” can be a secondary key. </p> <p><strong>Description of CNPSi.csv data attributes</strong></p> <p>Each line is a couple measurement/parameter/station</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· var: Abbreviation of parameter name</p> <p>· fraction: "water_filtrated" or "water_raw"</p> <p>· date: date of sampling</p> <p>· hour: hour of sampling</p> <p>· value: analytical result (concentration)</p> <p>· provider: provider of the data</p> <p>· producer: producer of the data</p> <p>· from_db: "Naiades2022" (https://naiades.eaufrance.fr/france-entiere#/ dump from 2022) or "DoNuts" (Thieu, V., Silvestre, M., 2015. DoNuts: un système d’information sur les observations environnementales. Présentation Séminaire UMR Métis)</p> <p>· n_meas: number of observations for a given parameter / station</p> <p>· unit: unit of concentration</p> <p>· element: "C" "N" "P" or "Si"</p> <p>· year: year of observation</p> <p>· month: month of observation</p> <p>· day: day of observation</p> <p>· julian_day: julian day observation (1-366)</p> <p>· decade: decade of observation (one of "1961-1970", "1971-1980", "1981-1990", "1991-2000", "2001-2010", "2011-2020")</p> <p> </p> <p> </p> <p><strong>Description of CNPSi_stats.csv data attributes</strong></p> <p>Each line is a couple parameter / station</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· var: Abbreviation of parameter name</p> <p>· n_meas: number of observations for a given parameter / station</p> <p>· start_year: year of first observation for a given parameter / station</p> <p>· end_year: year of last observation for a given parameter / station</p> <p>· duration_y_tot: total duration of observation in years for a given parameter / station</p> <p>· duration_y_tot: duration of observation in years for a given parameter / station for years with at least 1 meas</p> <p>· mean_nmeas_per_y_tot: mean number of observations per year considering total duration</p> <p>· mean_nmeas_per_y_meas: mean number of observations per year considering years with measurements</p> <p>· is_fully_continuous: TRUE if at least one measurement per year for a given parameter / station</p> <p>· start_cont_seq: year in which starts the longest continuous sequence for a given parameter / station</p> <p>· end_cont_seq: year in which ends the longest continuous sequence for a given parameter / station</p> <p>· duration_y_cont_seq: duration in years for the longest continuous sequence for a given parameter / station</p> <p>· nmeas_cont_seq: number of measurements for the longest continuous sequence for a given parameter / station</p> <p>· mean_nmeas_per_y_cont_seq: mean number of observations per year for the longest continuous sequence for a given parameter / station</p> <p>· mean: mean value (concentration) for a given parameter / station</p> <p>· median: median value (concentration) for a given parameter / station</p> <p>· sd: standard deviation (concentration) for a given parameter / station</p> <p>· cv: coeficient of variation (concentration) for a given parameter / station</p> <p>· c05,c25,c50,c75,c95: centiles 5, 25, 50, 75 & 95 for a given parameter / station</p> <p><strong>Description of catchments.gpkg and outlets.gpkg data attributes</strong></p> <p>Each line is a catchment or an outlet (sampling point)</p> <p>File is a .gpkg (EPSG = 3035)</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· watercourse: Name of the water course (from spatial join on IGN BD Topo)</p> <p>· mun_name: Name of the municipality of the outlet (from spatial join on IGN BD Admin Express)</p> <p>· ccm_wso_id: Seaoutlet id from CCM v2.1 database</p> <p>· ccm_wso1_id: Elementary catchment id from CCM v2.1 database</p> <p>· ccm_strahler: Strahler order of the catchment from CCM v2.1 database</p> <p>· area_km2: Computed area in km2 of the catchment</p> <p><strong>Description of daily discharges data attributes</strong></p> <p>Each line corresponds to a daily modelled discharge at a quality station from 1980 to 2019</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· date: Date in format yyyy-mm-dd</p> <p>· flow_mm: Discharge expressed in mm.d-1</p> <p>· flow_m3s: Discharge expressed in m3.s-1</p> <p><strong>Description of climate data attributes</strong></p> <p>Each line in the pr_tmin_tmax_1990-2019_lt_mean.csv corresponds to a mean value within a catchment for the 1990-2019 period.</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· period: 1990-2019</p> <p>· source: EMO-1 (pr) & EMO-5 (tmin, tmax)</p> <p>· pr: mean yearly precipitation (mm)</p> <p>· tmin: mean daily minimal temperature (°C)</p> <p>· tmin: mean daily maximal temperature (°C)</p> <p><strong>Description of land cover data attributes</strong></p> <p>Each line in the clc_8class.csv and clc_44class.csv corresponds to Corine Land Cover (CLC) class for a year (1990, 2000, 2006, 2012, or 2018) and a catchment. Raw CLC typology describes 44 classes that were aggregated to 8 classes (see clc_44class_to_8class.csv).</p> <p>· clc_44class.csv</p> <p>o sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o year: Year as stated in CLC product</p> <p>o clc_name: Description of land cover class in CLC product</p> <p>o clc_code: Code for land cover class in CLC product</p> <p>o percent_cover: Percent cover by CLC class in the catchment (0-100)</p> <p>· clc_8class.csv</p> <p>o sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o year: Year as stated in CLC product</p> <p>o label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o code_clc_8class: Code for land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o percent_cover: Percent cover by aggregated CLC class in the catchment (0-100)</p> <p>· clc_44class_to_8class.csv</p> <p>o code_clc: Code for land cover class in CLC product (44 classes)</p> <p>o code_clc_8class: Code for land cover class in aggregated CLC product (8classes)</p> <p>o label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p> </p> <p> </p> <p><strong>Acknowledgement</strong></p> <p>This publication has been prepared using European Union's Copernicus Land Monitoring Service information; <a href="https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac">https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac</a></p> <p>The authors thank Vasken Andréassian for communicating the discharge data and discharge station data and Alban de Lavenne for its help in using the transfr package, both for INRAE UR HYCAR.</p> <p> </p> <p><strong>References</strong></p> <p>de Lavenne, A., Skøien, J. O., Cudennec, C., Curie, F., & Moatar, F. (2016). Transferring measured discharge time series: Large-scale comparison of Top-kriging to geomorphology-based inverse modeling: transferring measured discharge time series. Water Resources Research, 52(7), 5555–5576. https://doi.org/10.1002/2016WR018716</p> <p>de Lavenne, A., Loree, T., Squividant, H., & Cudennec, C. (2023). The transfR toolbox for transferring observed streamflow series to ungauged basins based on their hydrogeomorphology. Environmental Modelling & Software, 159, 105562. <a href="https://doi.org/10.1016/j.envsoft.2022.105562">https://doi.org/10.1016/j.envsoft.2022.105562</a></p> <p>EEA. (2020). Corine Land Cover édition 2018. CLC 2018. <a href="https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine">https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine</a></p> <p>Pelletier, A., & Andréassian, V. (2020). Hydrograph separation: An impartial parametrisation for an imperfect method. Hydrology and Earth System Sciences, 24(3), 1171–1187. <a href="https://doi.org/10.5194/hess-24-1171-2020">https://doi.org/10.5194/hess-24-1171-2020</a></p> <p>Pelletier, A. (2021). Complétion d'hydrogrammes avec le modèle GR4J - Note méthodologique. INRAE, UR HYCAR.</p> <p>Thiemig, V., Gomes, G. N., Skøien, J. O., Ziese, M., Rauthe-Schöch, A., Rustemeier, E., Rehfeldt, K., Walawender, J. P., Kolbe, C., Pichon, D., Schweim, C., and Salamon, P.: EMO-5: a high-resolution multi-variable gridded meteorological dataset for Europe, Earth Syst. Sci. Data, 14, 3249–3272, https://doi.org/10.5194/essd-14-3249-2022, 2022</p> <p>Thieu, V., Silvestre, M., 2015. DoNuts : un système d'information sur les observations environnementales. Présentation Séminaire UMR Métis</p> <p>Vogt, J., A. de Jager, E. Rimaviciute, W. Mehl, S. Foisneau, K. Bódis, J. Dusart, M.L. Paracchini, P. Haastrup, & C. Bamps. (2007). A pan-European river and catchment database. (European Commission. Joint Research Centre. Institute for Environment and Sustainability.). Publications Office. https://data.europa.eu/doi/10.2788/35907</p>
Mappings for "Developing a Scalable Annotation Method for Large Datasets That Enhances Alarms With Actionability Data to Increase Informativeness: Mixed Methods Approach"
<p>Studies identified false and non-actionnable alarms as a factor for alarm fatigue in intensive care units.</p> <p>To annotate patient alarms, and analyse the alarm situation in intensive care units, we conceptualized and performed data mappings related to airway management and medication interventions. The mappings were based on information retrieved from the patient data management system (PDMS) and clinical expertise. For the airway management mappings, we used additional resources such as ISO 19223:2019 or ventilator instruction manuals. The mappings do not include patient data.</p> <p>As the mappings are generic, they could be used in other contexts than alarm annotation and research.</p> <p><strong>1. Respiratory Management Mappings:</strong></p> <ul> <li>General tables summarizing the 1) categories based on ISO 19223:2019 to describe respiratory support therapies (RSTs), 2) defining the invasiveness level of a RST and 3) listing the abbreviations used in the mappings</li> <li> <p>Tables including PDMS entries for airway devices (ADs), ventilation devices (VDs), and ventilation modes (VMs)</p> </li> <li> <p>Mapping of AD entries (from the PDMS) to defined categories</p> </li> <li> <p>Mapping of VDs, VMs, and ADs to defined RSTs, including information on invasiveness</p> </li> <li> <p>Table specifying suitable ventilation parameters in the context of each RST</p> </li> </ul> <p><strong>2. Medication Mappings:</strong></p> <ul> <li> <p>General tables providing information on physiological alarm conditions (PACs), interventions, routes, and techniques of administration of interest</p> </li> <li> <p>Mapping of routes of administration to techniques of administration including PDMS entries</p> </li> <li> <p>Mapping of active ingredients (including SNOMED CT Fully Specified Names and Identifiers), related PDMS information, and routes and techniques of administration to defined PAC and interventions</p> </li> </ul>
Dataset for "Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections"
<p>This dataset is used to reproduce the results presented in the following publication:</p> <p>Goelzer, H., Noel, B. P. Y., Edwards, T. L., Fettweis, X., Gregory, J. M., Lipscomb, W. H., van de Wal, R. S. W., and van den Broeke, M. R.: Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections, The Cryosphere Discuss., https://doi.org/10.5194/tc-2019-188, in review, 2019.</p> <p> </p>
Large-Scale Dataset for Radio Frequency based Device-Free Crowd Estimation
<p>This dataset serves to estimate the status, in particular the size, of a crowd given the impact on radio frequency communication links within a wireless sensor network. To quantify this relation, signal strengths across sub-GHz communication links are collected at the premises of the Tomorrowland music festival. The communication links are formed between the network nodes of wireless sensor networks deployed in three of the festival's stage environments. </p> <p>The table below lists the eighteen dataset files. They are collected at the music festival's 2017 and 2018 editions. There are three environments, labeled: ‘Freedom Stage 2017’, ‘Freedom Stage 2018’, and ‘Main Comfort 2018’. Each environment has both 433 MHz and 868 MHz data. The measurements at each environment were collected over a period of three festival days. The dataset files are formatted as Comma-Separated Values (CSV).</p> <pre><code class="language-markdown">| Dataset file | Reference file | Number of messages | |-------------------- |------------------------- |-------------------- | | free17_433_fri.csv | None | 393 852 | | free17_868_fri.csv | None | 472 202 | | free17_433_sat.csv | free17_transactions.csv | 996 033 | | free17_868_sat.csv | free17_transactions.csv | 1 023 059 | | free17_433_sun.csv | free17_transactions.csv | 1 007 066 | | free17_868_sun.csv | free17_transactions.csv | 1 036 456 | | free18_433_fri.csv | None | 765 024 | | free18_868_fri.csv | None | 757 657 | | free18_433_sat.csv | free18_transactions.csv | 711 438 | | free18_868_sat.csv | free18_transactions.csv | 714 390 | | free18_433_sun.csv | free18_transactions.csv | 648 329 | | free18_868_sun.csv | free18_transactions.csv | 656 290 | | main18_433_fri.csv | None | 791 462 | | main18_868_fri.csv | None | 908 407 | | main18_433_sat.csv | main18_counts.csv | 863 666 | | main18_868_sat.csv | main18_counts.csv | 884 682 | | main18_433_sun.csv | main18_counts.csv | 903 862 | | main18_868_sun.csv | main18_counts.csv | 894 496 |</code></pre> <p>In addition to the datasets and reference files, a software example is provided to illustrate the data use and visualise the initial findings and relation between crowd size and network signal strength impact.</p> <p>In order to use the software, please retain the following file structure: </p> <pre><code class="language-markdown">. ├── data ├── data_reference ├── graphs └── software</code></pre> <p>The peer-reviewed data descriptor for this dataset has now been published in MDPI Data - an open access journal aiming at enhancing data transparency and reusability, and can be accessed here: <a href="https://doi.org/10.3390/data5020052">https://doi.org/10.3390/data5020052</a>.<br> Please cite this when using the dataset.</p>
Dataset for the paper "Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset"
<p>We present a large-scale anomaly detection dataset collected from IBM Cloud's Console over approximately 4.5 months. This high-dimensional dataset captures telemetry data from multiple data centers, specifically designed to aid researchers in developing and benchmarking anomaly detection methods in large-scale cloud environments. It contains 39,365 entries, each representing a 5-minute interval, with 117,448 features/attributes, as interval_start is used as the index. The dataset includes detailed information on request counts, HTTP response codes, and various aggregated statistics. The dataset also includes labeled anomaly events identified through IBM's internal monitoring tools, providing a comprehensive resource for real-world anomaly detection research and evaluation.</p> <p><strong>File Descriptions</strong></p> <ul> <li><code>location_downtime.csv</code> - Details planned and unplanned downtimes for IBM Cloud data centers, including start and end times in ISO 8601 format.</li> <li><code>unpivoted_data.parquet</code> - Contains raw telemetry data with 413 million+ rows, covering details like location, HTTP status codes, request types, and aggregated statistics (min, max, median response times).</li> <li><code>anomaly_windows.csv</code> - Ground truth for anomalies, listing start and end times of recorded anomalies, categorized by source (Issue Tracker, Instant Messenger, Test Log).</li> <li><code>pivoted_data_all.parquet</code> - Pivoted version of the telemetry dataset with 39,365 rows and 117,449 columns, including aggregated statistics across multiple metrics and intervals.</li> <li><code>demo/demo.[ipynb|html]</code>: This demo file provides examples of how to access data in the Parquet files, available in Jupyter Notebook (<code>.ipynb</code>) and HTML (<code>.html</code>) formats, respectively.</li> </ul> <p>Further details of the dataset can be found in <strong>Appendix B: Dataset Characteristics</strong> of the <a href="https://arxiv.org/abs/2411.09047">paper</a> titled <strong><em>"Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset."</em></strong> Sample code for training anomaly detectors using this data is provided in <a href="https://doi.org/10.5281/zenodo.14598119" target="_blank" rel="noopener">this package</a>.</p> <p> </p> <p>When using the dataset, please cite it as follows:</p> <pre><code>@misc{islam2024anomaly,</code><br><code> title={Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset}, </code><br><code> author={Mohammad Saiful Islam and Mohamed Sami Rakha and William Pourmajidi and Janakan Sivaloganathan and John Steinbacher and Andriy Miranskyy},</code><br><code> year={2024},</code><br><code> eprint={2411.09047},</code><br><code> archivePrefix={arXiv},</code><br><code> url={https://arxiv.org/abs/2411.09047}</code><br><code>}</code></pre> <p> </p>
BioVars - bioclimatic datasets for Europe based on a large regional climate ensemble for periods between 1971 to 2098
<p>We present 26 bio-climatic variables that are calculated based on a large ensemble consisting of 70 bias-adjusted GCM-RCM (Global Climate Model – Regional Climate Model) simulations for 1971 to 2098. Both, the historic and the projection periods were calculated using the same models to ensure consistency between the periods. The variables are validated against E-OBS observations from which we calculated the same bio-climatic variables. For projection periods we chose 20 year ranges between 2021 to 2098. Here, we offer two versions of them 1) variables separated into RCP 2.6, 4.5 and 8.5 including the 5th, 50th and 95th percentiles among the realisations and within the RCPS. And 2) variables per realisation separately. We then extracted the temporal 5th, 50th and 95th percentile per period as representing values. Each zipped file contains these 26 bio-climatic variables according to their aggregation. The variables and the units are explained within the data descriptor publication. </p> <p> </p> <p><strong>File descriptions</strong></p> <ul> <li>bioVars_1971-2000_met.tar.gz >> Projections per realisations for period 1971-2000</li> <li>bioVars_2021-2040_met.tar.gz >> Projections per realisations for period 2021-2040</li> <li>bioVars_2041-2060_met.tar.gz >> Projections per realisations for period 2041-2060</li> <li>bioVars_2061-2080_met.tar.gz >> Projections per realisations for period 2061-2080</li> <li>bioVars_2079-2098_met.tar.gz >> Projections per realisations for period 2079-2098</li> <li>bioVars_2021-2040_rcp.tar.gz >> Projections per RCP for period 2021-2040</li> <li>bioVars_2041-2060_rcp.tar.gz >> Projections per RCP for period 2041-2060</li> <li>bioVars_2061-2080_rcp.tar.gz >> Projections per RCP for period 2061-2080</li> <li>bioVars_2079-2098_rcp.tar.gz >> Projections per RCP for period 2079-2098</li> <li>validation.tar.gz >> Validation using E-OBS (v20.0) and Worldclim (version 2.1)</li> </ul> <p> </p> <p><strong>References</strong> <br>Reichmuth, A., Rakovec, O., Boeing, F. <em>et al.</em> BioVars - A bioclimatic dataset for Europe based on a large regional climate ensemble for periods in 1971–2098. <em>Sci Data</em> <strong>12</strong>, 217 (2025). https://doi.org/10.1038/s41597-025-04507-w</p>
Dia-Pol: A large scale BlackLivesMatter and MeToo Twitter dataset
<p>This dataset (tweets_id_list.json) contains 258609 number of tweets sent in English extracted from Twitter API using the query word “#blacklivesmatter.” The dataset spans the period from 2020-01-01 to 2021-12-31 and was retrieved on 2022-06-10.</p>
Large Dataset of Nigeria Covid-19 Tweets for Sentiment Analysis and Opinion Mining Tasks
<p><strong>Background</strong></p> <p>Information is essential for growth; without it, little can be accomplished. Data gathering has seen significant changes throughout the previous few centuries because of certain transitory medium. The look and style of information transference are affected by the employment of new and emerging technologies, some of which are efficient, others are reliable, and many more are quick and effective, but a few were disappointing for various reasons.</p> <p><strong>Aims</strong></p> <p>This study aims at using TextBlob and VADER analyser with historical tweets, to analyse emotional responses to the corona virus pandemic (covid-19). It shows us how much of a sociological, environmental, and economic impact it has in Nigeria, among other things. This study would be a tremendous step forward for students, researchers, and scholars who want to advance in fields like data science, machine learning, and deep learning.</p> <p><strong>Methodology</strong></p> <p>The hashtag ‘covid-19' was used to collect 1,048,575 tweets from Twitter. The tweets were pre-processed with a twitter tokenizer, and Valence Aware Dictionary for Sentiment Reasoning (VADER) and TextBlob were used for sentiment and text mining, respectively. Topic modelling was done with Latent Dirichlet Allocation (LDA). The simulated subjects, on the other hand, were visualized using Multidimensional scaling (MDS).</p> <p><strong>Results</strong></p> <p>The result of the VADER sentiment returned 39.8%, 31.3% and 28.9%, positive, neutral, and negative sentiment respectively while the result of the TextBlob sentiment returned 46.0%, 36.7% and 17.3%, neutral, positive, and negative sentiment, respectively.</p> <p><strong>Conclusion</strong></p> <p>With all of this, information from social media may be used to help organizations, governments, and nations around the world make smart and effective decisions about how to restrict and limit the negative effects of covid-19. Also know the opinion and challenges of people, then deal with problem of misinformation.</p> <p>It is concluded that with popular belief a significant number of the populace regards covid-19 as a virus that has come to stay, some believe it will eventually be conquered.</p>
Dataset for "Large Language Models as molecular design engines"
<ol> <li><strong>claude-gpt-paper.zip :</strong><br><br>This dataset contains data and results associated with the paper "Large Language Models as molecular design<br>engines" The paper investigates the use of large language models, specifically Claude 3 Opus, for generating and analyzing chemical structures based on various prompts from A-H (as mentioned in the manuscript), and guided design related to electron-withdrawing groups (EWG), electron-donating groups (EDG).</li> </ol> <p>The dataset includes:</p> <ol> <li>PM7 MOPAC energy calculations for generated molecules, along with their SMILES representations and molecule IDs.</li> <li>PM7-calculated charges for the generated molecules.</li> <li>Output files from the Claude 3 Opus language model for each prompt category along.</li> <li>Original dataset (subset of ZINC database) used to build common keys and the initial design space.</li> <li>JSON file containing common keys for featurizing unknown SMILES.</li> <li>PCA object to convert molecule embeddings to 3-dimensional embeddings.</li> </ol> <p>The data is organized into the following folders:</p> <ul> <li><code>pm7_charge_results</code>: Contains HOMO-LUMO energy differences for plotting.</li> <li><code>pm7_charge_calculation</code>: Contains PM7 MOPAC energy calculations and charges.</li> <li><code>out</code>: Contains output files from the Claude 3 Opus language model.</li> <li><code>fact-dropbox</code>: Contains the original dataset, common keys, and PCA object file.</li> </ul> <p>The data can be used to reproduce the results presented in the paper and serve as a foundation for further research in this area.</p> <p>For a detailed description of the folder structure and contents, please refer to the File_descriptions.md file included in the dataset.<br><br><br>2. llm-visulizer-dashapp.zip<br><br>This is the code for the visualizer app for viewing the molecules generated by the LLM. The README.md file has details about running the app.</p> <p>3. claude-gpt-paper-codes.zip </p> <p>This contains the notebook GPT_modification_just_plots.ipynb for plotting, and other codes. The README.md file has details about running the main notebook for getting the plots.</p>
ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics
<p>ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics</p> <p>Supplementary file (full dataset download)</p> <p>- v2 with SMILES errors fixed (19.01.2019)</p>
Dataset for Clock drift corrections for large aperture ocean bottom seismometer arrays: application to the UPFLOW array in the mid-Atlantic Ocean
<p>Dataset from Clock drift corrections for large aperture ocean bottom seismometer arrays: application to the UPFLOW array in the mid-Atlantic Ocean DOI: 10.1093/gji/ggae354.</p> <p>This dataset includes the clock drift polynoms refered to jthe deployment date (jul day from 2022) and consecutive days up to the recovery date. Two types of formats.</p> <ol> <li>Txt files</li> <li>Python Pickle files with a Dictionary containing the NumPY polynom1D and additional information.</li> </ol>
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
TBPos: Dataset for Large-Scale Precision Visual Localization (database files)
<p>Large-scale dataset for visual localization, provided in the format of the well-known InLoc dataset (Taira et al, 2018). Contains co-registered RGB point clouds and a script for generating the rest of the 'database' files for visual localization by the InLoc algorithm. Note: query images are provided in a separate repository.</p>
Large Movie Review Dataset
<p>IMDB dataset having 50K movie reviews for natural language processing or Text analytics.<br> This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training and 25,000 for testing. So, predict the number of positive and negative reviews using either classification or deep learning algorithms.<br> For more dataset information, please go through the following link,<br> <a href="http://ai.stanford.edu/~amaas/data/sentiment/">http://ai.stanford.edu/~amaas/data/sentiment/</a>.</p>
Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
<p>We present a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability's root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle's continuous integration and deployment (CI/CD) process. </p> <p> </p> <p>The uploaded dataset contains 1000 codes, each comprising 20 to 50 lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p> </p> <p> </p>
Large Ensemble Dataset for Discovering Global Peak Water Limit of Future Groundwater Withdrawals Using 900 GCAM Runs
<h2><strong>Global Groundwater Withdrawals Peak Over the 21st Century </strong></h2> <p>The large ensemble dataset contains groundwater related model outputs from 900 scenarios modeled using <a href="http://jgcri.github.io/gcam-doc/toc.html">Global Change Analysis Model (GCAM)</a>. The scenario ensemble members include five Shared Socioeconomic Pathways (SSPs), four Representative Concentration Pathways (RCPs), five global climate model outputs, three groundwater depletion limits, two surface water storage expansion regimes, and two historical groundwater depletion trends.</p> <h3><strong>Journal Article</strong></h3> <p>Niazi, H., Wild, T.B., Turner, S.W.D., Graham, N.T., Hejazi, M., Msangi, S., Kim, S., Lamontagne, J.R., & Zhao, M. (2024). <a href="https://rdcu.be/dFpb5">Global peak water limit of future groundwater withdrawals</a>. <em>Nature Sustainability, 7</em>(4), 413–422. <a href="https://doi.org/10.1038/s41893-024-01306-w" rel="nofollow">https://doi.org/10.1038/s41893-024-01306-w</a></p> <p>Read full-text here: <a href="https://rdcu.be/dFpb5">https://rdcu.be/dFpb5</a> </p> <h3><strong>Data Repository </strong></h3> <p>This <em><strong>data</strong></em> repository is to be used in combination with the <em><strong>main</strong></em> <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a> containing all scripts and files for reproducing the experiment as well as the analysis and post-processing of the model outputs.</p> <p>Scripts and smaller files are provided in the <a href="https://github.com/JGCRI/niazi-etal_202X_xyz">GitHub meta-repository</a> whereas larger files are provided in this data repository. Please complete the repository by placing the files as described hereunder. Please find the GitHub meta-repository here: <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">https://github.com/JGCRI/niazi-etal_2024_nature-sustainability</a></p> <p>Descriptions of files:</p> <ol> <li><em><strong>gcam-5.7z</strong></em> contains the GCAM version used to simulate 900 scenarios of plausible futures. The model folder contains all necessary input files to reproduce the simulations. <ul> <li>The model is to be used in combination with the <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a> to setup batch runs on cluster.</li> <li>Please navigate to <a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model">model/</a> folder for other scenario-specific and model setup folders and files. <em><strong>gcam-5</strong></em> is to be extracted in the same directory (./<em>model/gcam-5/</em>). </li> <li>For the first-time users of GCAM, please follow guidance on <a href="http://jgcri.github.io/gcam-doc/toc.html">GCAM wiki</a> to setup GCAM or for background knowledge. </li> </ul> </li> <li><em><strong>crop_yeild.7z</strong></em>: This file contains inputs related to climate impacts on crop yields. This is to be downloaded and extracted in the <a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model/combined_impacts">model/combined_impacts/</a> folder. </li> <li><em><strong>outputs-all.7z: </strong></em>Key model outputs queried and collated from 900 GCAM runs are explained hereunder. The files could be downloaded individually (.csv files) or all at once in .7z format (<a href="../api/files/80b237d3-b22f-499f-8b8e-76c3846720a0/outputs-all.7z">outputs-all.7z</a>). These files are to be placed in the <a href="https://github.com/JGCRI/niazi-etal_202X_xyz/tree/main/model/outputs">model/outputs</a> folder of the <a href="https://github.com/JGCRI/niazi-etal_2024_nature-sustainability">meta-repository</a>. <ul> <li><em><strong>ag_prod_all_GW_scenarios.csv</strong></em> - Agricultural production across all scenario for 2050 and 2100 (tonnes)</li> <li><em><strong>prices_water_withdrawal_all.csv</strong> - </em>Water prices across all scenarios and years ($/km<sup>3</sup>)</li> <li><em><strong>global_irrigated_prod_by_crop.csv</strong></em> - All irrigated agricultural production for each crop across and scenarios all years (tonnes)</li> <li><em><strong>surface_water_production_all.csv</strong></em> - Runoff across all scenarios and years (km<sup>3</sup>)</li> <li><em><strong>groundwater_production_FINAL.csv</strong></em> - Groundwater withdrawals across all scenarios and years (km<sup>3</sup>)</li> <li><em><strong>water_withdrawals_desal_all.csv</strong></em> - Water withdrawals from desalination plants across all scenarios and years (km<sup>3</sup>)</li> </ul> </li> </ol> <h3><strong>Short introduction to the study</strong></h3> <p>Using 900 GCAM runs, this study finds that global groundwater withdrawals are expected to peak around mid-century, followed by a decline through 21st century, exposing about half of the population living in one-third of basins to groundwater stress, with cost and availability of surface water storage being the most significant driver of future groundwater withdrawals. This first-ever robust, quantitative confirmation of the peak-and-decline pattern for groundwater, previously only known for fossil fuels and minerals, raises concerns for basins heavily dependent on groundwater.</p> <p>Niazi, H., Wild, T.B., Turner, S.W.D., Graham, N.T., Hejazi, M., Msangi, S., Kim, S., Lamontagne, J.R., & Zhao, M. (2024). <a href="https://rdcu.be/dFpb5">Global peak water limit of future groundwater withdrawals</a>. <em>Nature Sustainability, 7</em>(4), 413–422. <a href="https://doi.org/10.1038/s41893-024-01306-w" rel="nofollow">https://doi.org/10.1038/s41893-024-01306-w</a></p> <p>Read full-text here: <a href="https://rdcu.be/dFpb5">https://rdcu.be/dFpb5</a></p> <h3><strong>Contact </strong></h3> <p>Please reach out to Hassan Niazi at <a href="mailto:hassan.niazi@pnnl.gov">hassan.niazi@pnnl.gov</a> for any questions. </p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for 1000 Genomes
<p>Tree sequences inferred for the 1000 Genomes phase 3 autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip 1kg_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("1kg_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("1kg_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for Simons Genome Diversity Project
<p>Tree sequences inferred for the SGDP autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip sgdp_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("sgdp_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("sgdp_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Public Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"
<p>Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found <a href="https://arxiv.org/pdf/1802.00393.pdf">here</a>. </p> <p>The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: </p> <ul> <li> <p>hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. </p> </li> <li> <p>hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> <li> <p>retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> </ul> <p> </p> <p>UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide <a href="https://zenodo.org/record/3706866#.YYLG6S8RqjQ">here </a>the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). </p> <p>Please cite the paper in any published work that uses any of these resources. </p> <p>@inproceedings{founta2018large, <br> title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior}, <br> author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas}, <br> booktitle={11th International Conference on Web and Social Media, ICWSM 2018}, <br> year={2018}, <br> organization={AAAI Press} <br> } </p> <p>For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.