Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “Repositories”
EPSRC HEED Data Repository: Footfall Monitoring System
<p>The dataset deposited here was prepared under the EPSRC-funded <a href="http://heed-refugee.coventry.ac.uk/">Humanitarian Engineering and Energy for Displacement</a> research project (EP/P029531/1). The project aimed to understand energy needs of displaced communities, create an evidence base on the usage of different energy interventions and provide recommendations for improved design of future energy interventions to better meet the needs of people. </p> <p>As part of the project, we deployed a Footfall Monitoring System in the Uttargaya settlement in Nepal. Footfall monitors are designed to measure the step count of passers-by with the aim to: Evaluate the level of activity in an area by measuring footfall count and Evaluate the effect of streetlights on the level of activity.</p> <p>For the purpose of this study, the 7 footfall monitors are deployed beside 7 streetlights. The footfall monitors were deployed prior to commissioning of streetlights to gather baseline data and evaluate the impact of streetlights on the footfall count. The key constituents of footfall monitors are: Raspberry Pi 3B and Case; PiFace Real Time Clock and CAM008 70º night vision camera. The total cost of a monitor is £92.88. The Raspberry Pi is the central unit of the system that runs a program to sense the footfall count as measured by the IR sensor. The IR sensor counts footfall by tracking the number of times a horizontal beam of light is “broken” when a person crosses a threshold. If new data is recorded by the sensor, the updated footfall count along with the direction of movement and the current time (measured from PiFace RTC) is stored onto an SD card. A packet containing the updated values is also transmitted to the heed-data server hosted at Coventry University.</p> <p>Post Deployment Challenges:</p> <ul> <li><strong>Damage to footfall</strong>: In April 2019, footfall monitor 7 was damaged due to a gust of storm and heavy rains in the camp. This monitor was replaced in May 2019.</li> <li><strong>Power outages:</strong> These are common in the camp. Data is lost during this time as the devices have no access to power.</li> <li><strong>Internet connectivity: </strong>The availability and reliability of Wi-Fi continue to be an issue for the transmission of data to heed-data server.</li> </ul>
Malware Repositories and Their Authors on GitHub
<p>This dataset is rooted in a study aimed at unveiling the origins and motivations behind the creation of malware repositories on GitHub. Our research embarks on an innovative journey to dissect the profiles and intentions of GitHub users who have been involved in this dubious activity. </p> <p>Employing a robust methodology, we meticulously identified 14,000 GitHub users linked to malware repositories. By leveraging advanced large language model (LLM) analytics, we classified these individuals into distinct categories based on their perceived intent: 3,339 were deemed Malicious, 3,354 Likely Malicious, and 7,574 Benign, offering a nuanced perspective on the community behind these repositories. </p> <p>Our analysis penetrates the veil of anonymity and obscurity often associated with these GitHub profiles, revealing stark contrasts in their characteristics. Malicious authors were found to typically possess sparse profiles focused on nefarious activities, while Benign authors presented well-rounded profiles, actively contributing to cybersecurity education and research. Those labeled as Likely Malicious exhibited a spectrum of engagement levels, underlining the complexity and diversity within this digital ecosystem.</p> <p> </p> <p>We are offering two datasets in this paper. First, a list of malware repositories - we have collected and extended the malware repositories on the GitHub in 2022 following the original papers. Second, a csv file with the github users information with their maliciousness classfication label. </p> <ol> <li> <p><strong>malware_repos.txt</strong></p> <ul> <li><strong>Purpose</strong>: This file contains a curated list of GitHub repositories identified as containing malware. These repositories were identified following the methodology outlined in the research paper <a href="https://www.usenix.org/conference/raid2020/presentation/omar">"SourceFinder: Finding Malware Source-Code from Publicly Available Repositories in GitHub."</a></li> <li><strong>Contents</strong>: The file is structured as a simple text file, with each line representing a unique repository in the format <code>username/reponame</code>. This format allows for easy identification and access to each repository on GitHub for further analysis or review.</li> <li><strong>Usage</strong>: The list serves as a critical resource for researchers and cybersecurity professionals interested in studying malware, understanding its distribution on platforms like GitHub, or developing defense mechanisms against such malicious content.</li> </ul> </li> <li> <p><strong>obfuscated_github_user_dataset.csv</strong></p> <ul> <li><strong>Purpose</strong>: Accompanying the list of malware repositories, this CSV file contains detailed, albeit obfuscated, profile information of the GitHub users who authored these repositories. The obfuscation process has been applied to protect user privacy and comply with ethical standards, especially given the sensitive nature of associating individuals with potentially malicious activities.</li> <li><strong>Contents</strong>: The dataset includes several columns representing different aspects of user profiles, such as obfuscated identifiers (e.g., ID, login, name), contact information (e.g., email, blog), and GitHub-specific metrics (e.g., followers count, number of public repositories). Notably, sensitive information has been masked or replaced with generic placeholders to prevent user identification.</li> <li><strong>Usage</strong>: This dataset can be instrumental for researchers analyzing behaviors, patterns, or characteristics of users involved in creating malware repositories on GitHub. It provides a basis for statistical analysis, trend identification, or the development of predictive models, all while upholding the necessary ethical considerations.</li> </ul> </li> </ol>
Data quality assurance at research data repositories: Survey data
<p>This dataset documents findings form a survey on the status quo of data quality assurance practices at research data repositories.</p> <p>The personalized online survey was conducted among repositories indexed in re3data in 2021. It covered the scope of the repository, types of data quality assessment, quality criteria, responsibilities, details of the review process, and data quality information, and yielded 332 complete responses.</p> <p>The dataset comprises a documentation file, the data file, a codebook, and the survey instrument.</p> <p>The <strong>documentation file</strong> (documentation.pdf) outlines details of the survey design and administration, survey response, and data processing. The <strong>data file</strong> (01_survey_data.csv) contains all 332 complete responses to 19 survey questions, fully anonymized. The <strong>codebook</strong> (02_codebook.csv) describes the variables, and the <strong>survey instrument</strong> (03_survey_instrument.pdf) comprises the questionnaire that was distributed to survey participants.</p>
Data from calculated radial neutron flux distributions in a KBS-3 type geological repository
<p>Data from calculations of radial distribution of neutron flux per emitted neutron from rods of spent nuclear fuel in a KBS-3 type geological repository. Reference (<em>Jansson, 2022</em>) contain a summary of the calculations and a description of the structure of this data.</p> <p>This data was computed on resources provided by Swedish National Infrastructure for Computing (SNIC) at Uppsala Multidisciplinary Center for Advanced Computational Science (UPPMAX), National Supercomputer Centre at Linköping University (NSC) and the SNIC Cloud, partially funded by the Swedish Research Council through grant agreement no. 2018-05973, under projects SNIC 2021/5-299 and SNIC 2021/18-12.</p>
Raw Data for Mapping Repositories and their Institutional Open Science Policies in Asia
<p>Persistent Identifiers (PIDs), particularly Digital Object Identifiers (DOIs), are crucial for establishing a robust and globally accessible research infrastructure. In Asia, a diverse array of research outputs and resources are produced and published in repositories. However, a significant number of these repositories, and outputs remain undiscoverable in global registries and aggregators. <br><br>These three datasets provides comprehensive information on the adoption of repositories, Open Access mandates, and DOIs adoption in Asian countries. It includes detailed records from different registry sources and repository platforms.<br><br>You can read the full report titled 'Mapping Repositories and their Institutional Open Science Policies in Asia' at <a href="https://doi.org/10.5281/zenodo.12566244">https://doi.org/10.5281/zenodo.12566244</a></p>
Data Repository: Land surface modelling activities at Weierbach catchment.
<p>The data in this repository comes from the modelling activities with the Community Land Model version 5.0 (CLM5) carried out at the Weierbach catchment, Luxembourg. The repository contains:</p> <ol> <li>A list of matric potentials of <em>Fagus sylvatica </em>at which it experiences a specific loss of conductivity (i.e., 12%, 50%, 88%) obtained from published data [File: additional_PHT_Fagus_sylvatica_Europe.csv].</li> <li>The hourly atmospheric forcing used during the simulations with CLM 5.0 in a NetCDF format [File: atmospheric_forcing.zip].</li> <li>All model results per experiment [model_results.zip].</li> <li>The R scripts for processing the model results for obtaining the information required for each figure [Files: manuscript_figure_#.R].</li> <li>A daily summary of the tree water deficit calculated per PFT, individual tree species, and the whole ecosystem [File: twd.csv].</li> <li>A daily summary of tree transpiration scaled at the catchment level per PFT, individual tree species, and the whole ecosystem [File: et_mm_wei.csv]. This daily summary is based on the hourly data available on: Klaus, J., Fabiani, G., Schoppach, R., Chun, K. P., Iffly, J. F., Penna, D., & Juilleret, J. (2024). Detailed sap flow monitoring data at Weierbach catchment, Luxembourg (Version v01) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.11381618" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.11381618</a></li> </ol>
Listing of data repositories that embed schema.org metadata in dataset landing pages
<p>Machine-readable metadata available from landing pages for datasets facilitate data citation by enabling easy integration with reference managers and other tools used in a data citation workflow. Embedding these metadata using the schema.org standard with the JSON-LD is emerging as the community standard. This dataset is a listing of data repositories that have implemented this approach or are in the progress of doing so.</p> <p>This is the first version of this dataset and was generated via community consultation. We expect to update this dataset, as an increasing number of data repositories adopt this approach, and we hope to see this information added to registries of data repositories such as re3data and FAIRsharing.</p> <p>In addition to the listing of data repositories we provide information of the schema.org properties supported by these data repositories, focussing on the required and recommended properties from the "Data Citation Roadmap for Scholarly Data Repositories".</p>
The Red Queen in the Repository: metadata quality in an ever-changing environment (preprint of paper, presentation slides and dataset collection with validation schemas to IDCC2019 conference paper)
<p>This fileset contains a preprint version of the conference paper (.pdf), presentation slides (as .pptx) and the dataset(s) and validation schema(s) for the IDCC 2019 (Melbourne) conference paper: <em>The Red Queen in the Repository: metadata quality in an ever-changing environment. </em>Datasets and schemas are in .xml, .xsd , Excel (.xlsx) and .csv (two files representing two different sheets in the .xslx -file). The <em>validationSchemas.zip</em> holds the additional validation schemas (.xsd), that were not found in the schemaLocations of the metadata xml-files to be validated. The schemas must all be placed in the same folder, and are to be used for validating the Dataverse <em>dcterms</em> records (with <em>metadataDCT.xsd</em>) and the Zenodo <em>oai_datacite</em> feeds respectively (<em>schema.datacite.org_oai_oai-1.0_oai.xsd</em>). In the latter case, a simpler way of doing it might be to replace the incorrect URL "<em>http://schema.datacite.org/oai/oai-1.0/ oai_datacite.xsd</em>" in the <em>schemaLocation </em>of these xml-files by the CORRECT: <em>schemaLocation="http://schema.datacite.org/oai/oai-1.0/ http://schema.datacite.org/oai/oai-1.0/oai.xsd"</em> as has been done already in the sample files here. The sample file folders <em>testDVNcoll.zip </em>(Dataverse), <em>testFigColl.zip </em>(Figshare)<em> </em>and <em>testZenColl.zip </em>(Zenodo)<em> </em>contain all the metadata files tested and validated that are registered in the spreadsheet with objectIDs.<br> In the case of Zenodo, one original file feed,<br> <em>zen2018oai_datacite3orig-https%20_zenodo.org_oai2d%20verb=ListRecords%26metadata<br> Prefix=oai_datacite%26from=2018-11-29%26until=2018-11-30.xml</em> ,<br> is also supplied to show what was necessary to change in order to perform validation as indicated in the paper.</p> <p>For Dataverse, a corrected version of a file,<br> <em>dvn2014ddi-27595<strong>Corr</strong>_https%20_dataverse.harvard.edu_api_datasets_export%20<br> exporter=ddi%26persistentId=doi%253A10.7910_DVN_27595<strong>Corr</strong>.xml</em> ,<br> is also supplied in order to show the changes it would take to make the file validate without error.</p>
Reflection Spectra Repository for Cool Giant Planets
<p>Supplementary material for <a href="http://iopscience.iop.org/article/10.3847/1538-4357/aabb05"><em>Exploring H2O Prominence in Reflection Spectra of Cool Giant Planets</em></a> - ApJ 858, 69 (2018).</p> <p>This repository contains 65520 model reflection spectra of cool giant planets. The grid explores the influence of metallicity, gravity, effective temperature, and sedimentation efficiency on H<sub>2</sub>O absorption signatures in giant planet atmospheres. We also include two animations to visualise how the prominence of H<sub>2</sub>O absorption evolves over this parameter space. The included models range over:</p> <p>*m => 1-100 x solar (log(m) @ 0.0, 0.5, 1.0, 1.5, 1.7, 2.0 dex) <-- log(m) = 1.7 new for V2 of the database.<br> *g => 1-100 m/s<sup>2</sup> (evenly over log(g) in steps of 0.1 dex)<br> *T<sub>eff</sub> => 150-400 K (linearly in steps of 10 K)<br> *f<sub>sed</sub> => 1-10 (linearly in steps of 1)</p> <p>(V 1.0, March 30th 2018):</p> <blockquote> <p>Initial release of the reflection spectra repository. </p> </blockquote> <p>(V 2.0, Oct 1st 2019): </p> <blockquote> <p>The cool giant reflection spectra grid has been re-computed using the latest version of the PICASO albedo code (doi: <a href="https://arxiv.org/ct?url=https%3A%2F%2Fdx.doi.org%2F10.3847%2F1538-4357%2Fab1b51&v=77076c4a">10.3847/1538-4357/ab1b51</a>). This fixes a few bugs and adds new model features (e.g. Raman scattering, see Batalha+2019).</p> <p>The new grid is packaged as a HDF5 file with an accompanying python script 'Open_Albedo_Database.py'. The python script is provided to show how to open the albedo database, plot the spectra, and save spectra as a .txt file. The user need only change 4 lines (specifying log(m), log(g), T<sub>eff</sub>, f<sub>sed</sub>) and run the python script to produce a plot of the albedo spectra (both with and without H<sub>2</sub>O absorption).</p> </blockquote> <p><strong>NEW</strong>: (V 2.1, Oct 3rd 2019): </p> <blockquote> <p>Fixed a bug causing models with log(g) = 3.4 or 3.9 to not display cloud opacity.</p> </blockquote>
GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.
<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of <a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username </td> </tr> <tr> <td>html_url </td> <td>URL of the GitHub profile </td> </tr> <tr> <td>type </td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization </td> </tr> </tbody> </table> <p><strong>GitHub repositories <em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name </strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>html_url </td> <td>URL link to the repository </td> </tr> <tr> <td>description</td> <td>GitHub project description </td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project </td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated <em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name </strong></td> <td><strong>Description </strong></td> <td><strong>Collection Method </strong></td> </tr> <tr> <td>html_url </td> <td>Repository URL </td> <td> </td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license </td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) </td> <td> </td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis. </td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) </td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages. </td> </tr> <tr> <td>project_information </td> <td>Presence of basic project information in README (True/False) </td> <td>(Manual) Checked if the readme have basic information about the project. </td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. </td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit </td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory. </td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool </td> <td>Name of the continuous integration tool used</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule </td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (linters) Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) </td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language </td> <td>Programming language used in the repository </td> <td> </td> </tr> <tr> <td>type </td> <td>Specifies if the profile is a user or organization </td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation </td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group </td> <td> </td> </tr> </tbody> </table> <p> </p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>
Github commit data for the article "Beyond Zipf's law: Exploring the discrete generalized beta distribution in open-source repositories"
<p><span>This dataframe corresponds to the data used in the Nowak's et al. 2024 article "Beyond Zipf’s law: Exploring the discrete generalized beta distribution in open-source repositories" (see reference below).</span></p> <p><span>It consists of the distirbutions of number of commits per user across a number of GitHub repositories. <br><br>There are three columns:</span></p> <ul> <li><span>repository: the repository name</span></li> <li><span># of commits: the number of commits of a given individual</span></li> <li><span>rank: the user rank in the repository (by decreasing number of commits)<br><br></span></li> </ul> <p><strong><span>Reference:</span></strong></p> <p><span>Nowak, P., Santolini, M., Singh, C., Siudem, G., & Tupikina, L. (2024). Beyond Zipf’s law: Exploring the discrete generalized beta distribution in open-source repositories. <em>Physica A: Statistical Mechanics and Its Applications</em>, <em>649</em>, 129927. <a href="https://doi.org/10.1016/j.physa.2024.129927">https://doi.org/10.1016/j.physa.2024.129927</a></span></p>
Raw Data for Mapping Repositories and their Institutional Open Science Policies in the Middle East and North Africa (MENA)
<div> <p>Persistent Identifiers (PIDs), particularly Digital Object Identifiers (DOIs), are crucial for establishing a robust and globally accessible research infrastructure. In the Middle East and North Africa (MENA) region, a diverse array of research outputs and resources are produced and published in repositories. However, a significant number of these repositories, and outputs remain undiscoverable in global registries and aggregators. <br><br>These three datasets provides comprehensive information on the adoption of repositories, Open Access mandates, and DOIs adoption in MENA countries. It includes detailed records from different registry sources and repository platforms.<br><br>You can read the full report titled 'Mapping Repositories and their Institutional Open Science Policies in MENA' at <a href="https://doi.org/10.5281/zenodo.11370031">https://doi.org/10.5281/zenodo.11370031</a></p> </div>
Repository of Thoracolumbar Spine Triangulated Meshes
<p><em><strong>Repository of Thoracolumbar Spine Triangulated Meshes:</strong></em></p> <ul> <li>42 stereolithography (stl) files (.stl extension) representing patient-specific thoracolumbar spine triangulated meshes ("<a href="../record/7715658/files/42%20Patient-Specific%20stl%20files.rar?download=1">42 Patient-Specific stl files.rar</a>"), including point coordinates and triangulated mesh connectivity IDs. These stl files are reconstructed using sterEOS software (IRCCS, Milan), thanks to non-operated bi-planar images. Each stl file includes triangulated meshes of vertebras, pelvis, sacrum, and the femoral head. The stl files can be opened by many image viewers, 3D modelling, and CAD programs, like: Microsoft 3D Viewer (windows), and MeshLab (multiplatform). Data can be filtered, visualized, and downloaded using the data visualisation platform (<a href="https://thc.spineview.upf.edu/">https://thc.spineview.upf.edu</a>/).</li> <li>16807 stl files representing the thoracolumbar spine triangulated models ("<a href="../api/files/a36be640-9968-46cc-a0f5-b4aa610dfcfa/stl.part01.rar?versionId=9aa98aa1-9e2d-4638-809e-2358f406de2d">stl.part01.rar</a>" to "<a href="../api/files/a36be640-9968-46cc-a0f5-b4aa610dfcfa/stl.part09.rar?versionId=f9cb0431-4db1-4dd0-9313-6c782c626280">stl.part09.rar</a>"), including point coordinates and triangulated mesh connectivity IDs. These stl files are sampled by combining the first 5 shape modes of the SSM, in which each shape mode is discretized into 7 Standard Deviations (SD): -3, -2, -1, 0, 1, 2, 3. Each stl file includes triangulated meshes of vertebras, pelvis, sacrum, and the femoral head. The stl files can be opened by many image viewers, 3D modelling, and CAD programs, like: Microsoft 3D Viewer (windows), and MeshLab (multiplatform). Data can be filtered, visualized, and downloaded using the data visualisation platform (<a href="https://thc.spineview.upf.edu/">https://thc.spineview.upf.edu</a>/).</li> <li>An excel file "<a href="../records/8108354/files/Descriptive_List.xlsx?download=1">Descriptive_List.xlsx</a>" reporting measured spinopelvic parameters for 42 patient-specific triangulated models, and virtual cohort 16807 triangulated models. Spinopelvic parameters are: Pelvic-Incidence (PI), Pelvic Tilt (PT), Sacral Slope (SS), Lumbar Lordosis (LL), LL-PI, Global Tilt (GT), Relative Pelvic Version (RPV), Relative Lumbar Lordosis (RLL), Lumbar Distribution Index (LDI), Relative Spinopelvic Alignment (RSA), T1 pelvic Angle (TPA), and scoliosis cobb angle. Global Alignment and Proportion (GAP) score is further measured as a complementary assessment. The Excel file has two sheets: sheet1 for 16807 virtual cohort, and sheet2 for 42 patient-specific models. Model ID in the excel file is correspondent to the model's name in both virtual cohort and patient-specific models. Model number in "Descriptive_List.xlsx" is correspondent to the same model number in 42 and 16807 FE input files, respectively (42 FE input files (DOI: <a href="../records/10994164">10.5281/zenodo.10994164</a>), 16807 virtual FE input files (DOI: <a href="https://doi.org/10.5281/zenodo.8107354">10.5281/zenodo.8107354</a>)).</li> </ul> <p>Developed by: Morteza Rasouligandomani (Ph.D. in biomedical engineering, Pompeu Fabra university, BCN Med-Tech group, DTIC department, Barcelona, Spain).</p> <p>Email contact: jerome.noailly@upf.edu</p>
A2.2a Digital repositories data citation practices. Supplementary material
<p>Data to complement the quantitative analysis of data citation practices in digital repositories based on metadata records from the re3data.org repositories registry.</p> <p>Data was retrieved using re3data.org API on 23-02-2023 and 06-03-2023 and processed using the OpenRefine software.</p> <p>Part of "A FAIR-enabling citation model for Cultural Heritage Objects" project activities.</p>
Saskatchewan Glacier Basal Icequake Event Repository
<p>EPSL-D-23-00491 REV1 Repository</p> <p>Supplementary repository for manuscript:<br> Stevens, N.T., Zoet, L.K., Hansen, D.D., Alley, R.B., Roland, C.J., Schwans, E., and Shepherd, C.S. (In Review) Icequake insights on transient glacier slip mechanics near channelized subglacial drainage. Earth Planet. Sci. Lett.</p> <p>This repository contains data and metadata for interested parties to reproduce seismic event analyses for 21915 events classified as basal icequakes recorded between August 1st and 19th 2019 at Saskatchewan Glacier, Alberta, Canada.</p>
INNAG XML Epidoc Inscription Repository
<p>VIDARBHA EPIGRAPHY - Merded Epidoc Editions of Brahmi inscriptions from the Nagpur Area (Maharashtra).</p>
A Dataset for GitHub Repository Deduplication
<p>GitHub projects can be easily replicated through the site's fork process or through a Git clone-push sequence. This is a problem for empirical software engineering, because it can lead to skewed results or mistrained machine learning models. We provide a dataset of 10.6 million GitHub projects that are copies of others, and link each record with the project's ultimate parent. The ultimate parents were derived from a ranking along six metrics. The related projects were calculated as the connected components of an 18.2 million node and 12 million edge denoised graph created by directing edges to ultimate parents. The graph was created by filtering out more than 30 hand-picked and 2.3 million pattern-matched clumping projects. Projects that introduced unwanted clumping were identified by repeatedly visualizing shortest path distances between unrelated important projects. Our dataset identified 30 thousand duplicate projects in an existing popular reference dataset of 1.8 million projects. An evaluation of our dataset against another created independently with different methods found a significant overlap, but also differences attributed to the operational definition of what projects are considered as related. </p> <p>The dataset is provided as two files identifying GitHub repositories using the <em>login-name/project-name</em> convention. The file <em>deduplicate_names</em> contains 10,649,348 tab-separated records mapping a duplicated <em>source project</em> to a definitive <em>target project</em>.</p> <p>The file <em>forks_clones_noise_names</em> is a 50,324,363 member superset of the source projects, containing also projects that were excluded from the mapping as noise.</p>
'MARSH' - respiratory signal repository
<p>This dataset, 'MARSH', includes respiratory signals as part of the study "Fusion enhancement for tracking of respiratory rate through intrinsic mode functions in photoplethysmography."<br> It is meant to support academic research, particularly on algorithm development tools.</p> <p>Contents:<br> - Data.txt (age, gender, height, weight, systole, diastole, [respectively])<br> - ECG.mat (raw ECG data)<br> - ECG_annot.mat (annotations for the R peaks in ECG data)<br> - IP.mat (Raw IP data)<br> - IP_annot.mat (annotations for the local maxima of IP data [end of inspiration phase])<br> - NASAL.mat (Thermistor mask data)<br> - NASAL_annot.mat (annotations for the local maxima of thermistor mask data [end of inspiration phase])<br> - PPG.mat (Raw PPG signal data)</p> <p>When referring to this dataset, please consider including the following reference:</p> <p>Mikko Pirhonen and Vehkaoja Antti, Fusion enhancement for tracking of respiratory rate through intrinsic mode functions in photoplethysmography. Biomedical Signal Processing and Control. 2020</p>
Coronavirus COVID-19 (2019-nCoV) Data Repository for Africa
<p>The purpose of this repository is to collate data on the ongoing coronavirus pandemic in Africa. Our goal is to record detailed information on each reported case in every African country. We want to build a line list – a table summarizing information about people who are infected, dead, or recovered. The table for each African country would include demographic, location, and symptom (where available) information for each reported case. The data will be obtained from official sources (e.g., WHO, departments of health, CDC etc.) and unofficial sources (e.g., news). Such a dataset has many uses, including studying the spread of COVID-19 across Africa and assessing similarities and differences to what’s being observed in other regions of the world.</p> <p>See the repo here <a href="https://github.com/dsfsi/covid19africa">https://github.com/dsfsi/covid19africa</a></p>
EPSRC HEED Data Repository: Stove Use Monitoring System
<p>The dataset deposited here was prepared under the EPSRC-funded <a href="http://heed-refugee.coventry.ac.uk/">Humanitarian Engineering and Energy for Displacement</a> research project (EP/P029531/1). The project aimed to understand energy needs of displaced communities, create an evidence base on the usage of different energy interventions and provide recommendations for improved design of future energy interventions to better meet the needs of people.</p> <p>As part of the project, we deployed Stove Use Monitoring Systems (SUM) on clay cook stoves in Kigeme camp, Rwanda. The aim was to (a) measure and evaluate temperature profiles within stove enclosure and on the surface of stoves (b) evaluate frequency and duration of stove use. The SUM consisted of 2 sensors - a thermocouple (to measure temperature within the stove) and a Si7021 sensor (to measure temperature and humidity outside the stove), connected to an Arduino MKR GSM 1400 board. The data measured by the sensors was stored only if the change in values exceeded a set threshold for either of the readings. The SUM was powered by a re-chargeable Li-Ion battery of 3.7V and a rating of 7.59Wh.</p> <p>The study was conducted in 2 phases. In phase 1 (02 July 2019 to 30 September 2019), data was collected from 15 SUM and stored locally on SD card as well as communicated to a remote server via GSM. The time of data collection was recorded using GSM functionality. However, several GSM and MQTT failures were noted leading to loss of timestamp values as well as shorter battery lifetime due to re-transmission tries. In phase 2 (02 October 2019 to 17 October 2019), data was collected from 9 SUM and only stored locally on SD cards. The time of data collection was recorded using an external RTC clock connected to the Arduino board. The data from both phases of study is deposited here along with the metadata.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.