Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “package”
Replication package and appendixes for Causal inference of server- and client-side code smells in web apps evolution
<p>-Analysis <br>--R scripts used to make the analisys, divided by folders<br>--Data folders used in the questions</p> <p>-Appendixes - used in the article to shwo extra tables and plots</p> <p>-data folders - Aggregation of data, each app has two files, CSV and xls</p> <p>-separated data folders - 5 files for each app, with lines corresponding to the each released official version<br>--serversmells<br>--clientsmells<br>--javascriptsmells<br>--Cloc(metrics)<br>--version (all oficial releases)</p> <p>-issues_bugs<br>--data -issues by app by release <br>--data_bugs_more - the same but only bugs, by app by release<br>--scripts - scrips used to aggregate issues (from daily issues to by release) anf the same for bugs</p> <p> </p>
Replication package for: "Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions"
<p>This repository contains the data and code to replicate the analyses performed in <a href="https://academic.oup.com/ej/article/135/668/1261/7918442?utm_source=authortollfreelink&utm_campaign=ej&utm_medium=email&guestAccessKey=d6c8adb1-257c-47ad-827c-79b94cf86664" target="_blank" rel="noopener">"Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions"</a> by <a href="https://people.smu.edu/kdesmet/">Klaus </a><a href="https://people.smu.edu/kdesmet/" target="_blank" rel="noopener">Desmet</a>, <a href="https://sites.google.com/view/ignacioortuno" target="_blank" rel="noopener">Ignacio Ortuño-Ortín</a>, and <a href="http://omerozak.com">Ömer </a><a href="http://omerozak.com" target="_blank" rel="noopener">Özak</a>. If you use the code or data in this repository, please cite both the original paper and the dataset.<br><br>Citation:</p> <p>Desmet, Klaus, Ortuño-Ortín, Ignacio, and Özak, Ömer. (2024) "<a href="https://academic.oup.com/ej/article/135/668/1261/7918442?utm_source=authortollfreelink&utm_campaign=ej&utm_medium=email&guestAccessKey=d6c8adb1-257c-47ad-827c-79b94cf86664" target="_blank" rel="noopener">Is Secessionism Mostly About Income or Identity? A Global Analysis of 3,153 Subnational Regions</a>", Economic Journal, Volume 135, Issue 668, May 2025, Pages 1261–1299.</p>
FAIR Charging Station data package (Normalised)
<p>FAIR and normalised dataset based on the BNetzA charging station data.</p> <p>Original source: <a href="https://www.bundesnetzagentur.de/DE/Fachthemen/ElektrizitaetundGas/E-Mobilitaet/Ladesaeulenkarte/start.html">BNetzA Ladesaeulenregister (from 01.12.2024)</a></p> <p>Cleaning and annotation scripts: <a href="https://doi.org/10.5281/zenodo.10201060">FAIR Charging station data</a></p> <p>Metadata key reference:<a href="https://github.com/OpenEnergyPlatform/oemetadata/blob/develop/metadata/latest/metadata_key_description.md"> OEMETADATA Key description</a></p> <p>The data can be loaded individually from the csv files or as a whole using <a href="https://github.com/frictionlessdata/frictionless-py">frictionless.py</a>, for example, unzipping and calling:</p> <p> </p> <blockquote> <p>import frictionless as fl</p> </blockquote> <blockquote> <p>package = fl.Package('bnetza_charging_stations_normalised_01_12_2024.json')</p> </blockquote>
R package n2khab: providing preprocessed reference data for Flemish Natura 2000 habitat analyses
The n2khab package is an R package with preprocessing functions and standard reference data, useful for analyses regarding Flemish Natura 2000 habitats and regionally important biotopes (RIBs). URL: <a href="https://inbo.github.io/n2khab">https://inbo.github.io/n2khab</a>.
Data and code for "Tweezepy: A Python package for calibrating forces in single-molecule video-tracking instruments"
<p>Data and code for "Tweezepy: A Python package for calibrating forces in single-molecule video-tracking instruments."</p> <p>Data includes representative real and simulated bead trajectories used in the manuscript.</p> <p>Code includes all simulations, analysis, and plot details for the Figures in the manuscript. </p> <p>See included README.txt for more details.</p>
R code for archaeological examples of calculating isotopic niche space and overlap using the rKIN package
<p>This R code was written to apply the tools of the rKIN package to calculate isotopic niche space and overlap for the three archaeological case studies for the manuscript Investigating Isotopic Niche Space: Using rKIN for Stable Isotope Studies in Archaeology published in the Journal of Archaeological Method and Theory. Raw data for the case studies are available in the supplemental Excel file.</p>
Reproduction package for paper "How far are we from reproducible research on code smell detection? A systematic literature review"
<p>Checklist and data extracted from publications analyzed for "How far are we from reproducible research on code smell detection? A systematic literature review" paper, together with processing scripts and calculations of Cohen's Kappa.</p> <p>Paper that describes details of the data is available here: https://doi.org/10.1016/j.infsof.2021.106783</p>
Example code and data for ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework
<p>This repository contains an R script (grouse_example.R) and data (grouse_data.csv) used to reproduce the grouse abundance analysis described in Kellner, K. F., et al. (2021) ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework. Methods in Ecology and Evolution. The R script requires installation of the ubms R package, which can be obtained from CRAN (https://cran.r-project.org/package=ubms).</p> <p>The repository also contains an additional example occupancy analysis (occupancy_example.R) using the crossbill dataset included with the unmarked R package.</p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)
<p><strong>Sources: </strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p> </p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Data for "Paris Agreement requires substantial, broad, and sustained policy efforts beyond COVID-19 recovery packages"
<p>This dataset contains the underlying data for the following publication: Tanaka, K., C. Azar, O. Boucher, P. Ciais, Y. Gaucher, D. J. A. Johansson (2022) Paris Agreement requires substantial, broad, and sustained policy efforts beyond COVID-19 public stimulus packages. <em>Climatic Change</em> <strong>172, </strong>1 (2022). https://doi.org/10.1007/s10584-022-03355-6</p> <p>Earlier manuscripts were published as a preprint. https://arxiv.org/abs/2104.08342</p>
Dipeptidyl peptidase 11 (PgDPP11); A Target Enabling Package
<p><em>Porphyromonas</em> gingivalis (<em>P. gingivalis</em>) is the main causative agent of Periodontitis, the most widespread inflammatory condition world-wide. Recently this organism has been implicated in several systemic conditions, such as Alzheimer’s disease and type 2 diabetes. <em>P. gingivalis</em> does not ferment carbohydrates, instead it uses proteases to generate energy and carbon source. Dipeptidyl peptidase 11 plays a central role in the energy metabolism of this bacterium and has been proposed as an attractive drug target. This TEP provide early tools to develop inhibitors of PgDPP11, including purification protocols of recombinant proteins, a crystal structure of the protein in complex with a dipeptide, crystallisation conditions suitable for crystallography-based fragment screening, an inhibition assay and fragment hits in the active site and an allosteric site. These molecules provide a promising starting point for the development of more specific and potent PgDPP11 inhibitors.</p>
SciKGTeX Scientific Contribution Metadata LaTeX Package User Evaluation Results & Analysis
<p>The responses and measured variables from 26 participants of the first user test of the SciKGTeX package.</p> <p><a href="https://github.com/Christof93/SciKGTeX">https://github.com/Christof93/SciKGTeX</a></p> <p>Also the raw text source for the evaluation tasks and the result analysis notebook with the results saved as tsv file.</p>
Reproduction package for the publication 'Galaxy cluster photons alter the ionisation state of the nearby warm-hot intergalactic medium'
<p>The following files can be used to reproduce the figures and data from the paper <strong>Galaxy cluster photons alter the ionisation state of the nearby warm-hot intergalactic medium</strong><strong> </strong>by L. Štofanová, A. Simionescu, N. A. Wijers, J. Schaye, and J. Kaastra to be accepted in Monthly Notices of the Royal Astronomical Society (MNRAS).</p>
Datasets associated with the publication of the "satuRn" R package
<p>On this Zenodo link, we share the data that is required to reproduce all the analyses from our publication "satuRn: Scalable Analysis of differential Transcript Usage for bulk and single-cell RNA-sequencing applications".</p> <p>This repository includes input transcript-level expression matrices and metadata for all datasets, as well as intermediate results and final outputs of the respective DTU analyses. For a more elaborate description of the data, we refer to the companion GitHub for our publications; https://github.com/statOmics/satuRnPaper. Note that this is version 1.0.3 of the data (uploaded on 2022-07-08). If any changes were to be made to the datasets in the future, this will also be communicated on our companion GitHub page. </p>
BY-COVID Work Package 2 List of Resources
<p>Work Package 2: <em>Accessing heterogeneous data across domains and jurisdictions for enabling the downstream processing of COVID-19 and future pandemic episodes data </em>has gathered a relevant list of resources for the following areas: Non-patient related, Human-patient biomolecular, Human-patient clinical and health and socio-economics. </p>
ecochange: An R-package to derive ecosystem change indicators from freely available earth observation products
<p>This release includes the R code necessary to reproduce Figures 2-4 in the Application paper entitled: "ecochange: An R-package to derive ecosystem change indicators from freely available Earth Observation products."</p>
Packaging_machine_09/29/22
Documentation material from the Mastic pilot of the Mingei project
Perceptions on the utility of community question and answer websites like Stack Overflow to software developers (Replication package)
<p>Interview Questions on the perception of the utility of CQAs like Stack Overflow to software developers. In this study, we focused on the questions highlighted in yellow.</p>
Reproduction package for the paper "Exploring the directly imaged HD 1160 system through spectroscopic characterization and high-cadence variability monitoring"
<p>This is a basic reproduction package for the paper <a href="https://doi.org/10.1093/mnras/stae1315">"Exploring the directly imaged HD 1160 system through spectroscopic characterization and high-cadence variability monitoring" by Sutlieff et al. (2024)</a>. It aims to provide the most important data products to check and reproduce the main results of the paper.</p>
Data package for "Fast event-driven simulations for soft spheres: from dynamics to Laves phase nucleation"
<p>This dataset contains supporting data for the publication:</p> <p><em>Fast event-driven simulations for soft spheres: from dynamics to Laves phase nucleation</em></p> <p>A. Castagnède, L. Filion, and F. Smallenburg, J. Chem. Phys. 160 (2024), doi:10.1063/5.0209178, arXiv:2403:12755</p> <p> </p> <p><strong>Contents:</strong></p> <p>The main folder <em>data_package</em> contains three subfolders: <em>figures</em>, <em>SLNN</em>, and <em>snapshots</em>. The <em>figures</em> subfolder contains supporting data for each of the figures found in the publication, accompanied by details on statepoints and methods in individual README files. The <em>SLNN</em> subfolder contains the trained neural network classifier used in this work for crystalline phase identification, alongside usage instructions and an exemple system to analyze. Finally, the <em>snapshots</em> subfolder contains supplementary snapshots of the crystalline clusters obtained in simulations. </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.