Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
422
datasets available to search
ShareScore release 0.9.0
Dataset results
422 results for “Authors”
Raw data for the submitted manuscript entitled "Mapping and Disposal of Irrigation Pipes for a Sustainable Management of Agricultural Plastic Waste", authors Ileana Blanco, Giuliano Vox, Fabiana Convertino, and Evelia Schettini
<p><span>The file regards the evaluation of plastic indexes and agricultural plastic waste quantities in Apulia region due to the use of irrigation pipes. The data is used to identify the critical areas for plastic waste production due to irrigation pipes.</span></p>
A Standardized Review of Bat Names Across Multiple Taxonomic Authorities
<p>The Bat Eco-Interactions Working Group, in collaboration with GBatNet and the international Bat Taxonomy Group, developed the <strong>Bat Taxonomic Alignment (BTA)</strong> to reconcile taxonomic discrepancies across currently recognized bat species. As knowledge of bat population structure and evolutionary history advances, taxonomic boundaries and species names are frequently revised. To address these changes, the BTA integrates data from ten leading taxonomic authorities and consolidates relationships, synonyms, and historic combinations for over <strong>1,480 valid bat species</strong> across <strong>1,680 taxonomic treatments</strong>.</p> <p>This open-access, searchable tool provides a time-calibrated inventory of Chiroptera taxonomy, documenting valid names, alternative names, subspecies, and synonymies. By aligning these classifications, the BTA enables users to identify unharmonized binomials and trace nomenclatural changes over time. It promotes taxonomic clarity critical for research, biodiversity assessments, and conservation planning, where misidentified or misaligned taxa can lead to gaps in knowledge, resource misallocation, or overlooked species. The BTA thus represents a foundational advancement in bat biodiversity informatics, emphasizing transparency, data provenance, and interoperability across digital taxonomic frameworks.</p> <p> </p>
Malware Repositories and Their Authors on GitHub
<p>This dataset is rooted in a study aimed at unveiling the origins and motivations behind the creation of malware repositories on GitHub. Our research embarks on an innovative journey to dissect the profiles and intentions of GitHub users who have been involved in this dubious activity. </p> <p>Employing a robust methodology, we meticulously identified 14,000 GitHub users linked to malware repositories. By leveraging advanced large language model (LLM) analytics, we classified these individuals into distinct categories based on their perceived intent: 3,339 were deemed Malicious, 3,354 Likely Malicious, and 7,574 Benign, offering a nuanced perspective on the community behind these repositories. </p> <p>Our analysis penetrates the veil of anonymity and obscurity often associated with these GitHub profiles, revealing stark contrasts in their characteristics. Malicious authors were found to typically possess sparse profiles focused on nefarious activities, while Benign authors presented well-rounded profiles, actively contributing to cybersecurity education and research. Those labeled as Likely Malicious exhibited a spectrum of engagement levels, underlining the complexity and diversity within this digital ecosystem.</p> <p> </p> <p>We are offering two datasets in this paper. First, a list of malware repositories - we have collected and extended the malware repositories on the GitHub in 2022 following the original papers. Second, a csv file with the github users information with their maliciousness classfication label. </p> <ol> <li> <p><strong>malware_repos.txt</strong></p> <ul> <li><strong>Purpose</strong>: This file contains a curated list of GitHub repositories identified as containing malware. These repositories were identified following the methodology outlined in the research paper <a href="https://www.usenix.org/conference/raid2020/presentation/omar">"SourceFinder: Finding Malware Source-Code from Publicly Available Repositories in GitHub."</a></li> <li><strong>Contents</strong>: The file is structured as a simple text file, with each line representing a unique repository in the format <code>username/reponame</code>. This format allows for easy identification and access to each repository on GitHub for further analysis or review.</li> <li><strong>Usage</strong>: The list serves as a critical resource for researchers and cybersecurity professionals interested in studying malware, understanding its distribution on platforms like GitHub, or developing defense mechanisms against such malicious content.</li> </ul> </li> <li> <p><strong>obfuscated_github_user_dataset.csv</strong></p> <ul> <li><strong>Purpose</strong>: Accompanying the list of malware repositories, this CSV file contains detailed, albeit obfuscated, profile information of the GitHub users who authored these repositories. The obfuscation process has been applied to protect user privacy and comply with ethical standards, especially given the sensitive nature of associating individuals with potentially malicious activities.</li> <li><strong>Contents</strong>: The dataset includes several columns representing different aspects of user profiles, such as obfuscated identifiers (e.g., ID, login, name), contact information (e.g., email, blog), and GitHub-specific metrics (e.g., followers count, number of public repositories). Notably, sensitive information has been masked or replaced with generic placeholders to prevent user identification.</li> <li><strong>Usage</strong>: This dataset can be instrumental for researchers analyzing behaviors, patterns, or characteristics of users involved in creating malware repositories on GitHub. It provides a basis for statistical analysis, trend identification, or the development of predictive models, all while upholding the necessary ethical considerations.</li> </ul> </li> </ol>
S54 | EFSAPRI | European Food Safety Authority Priority Substances
<p>This is the dataset associated with list S54 EFSAPRI on the NORMAN Suspect List Exchange:</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p>
Local Governance in Ukraine during the full-scale Russian invasion. – Merged data from online surveys of local self-government authorities by the Congress of Local and Regional Authorities of the Council of Europe in 2022 and Kyiv School of Economics in 2024.
The dataset includes responses from two waves of online surveys targeting local self-government representatives in Ukraine, with a focus on crisis governance during the ongoing Russian war. The first wave was conducted from August 30 to September 20, 2022, by the Congress of Local and Regional Authorities of the Council of Europe, yielding 241 responses (16% of all Ukrainian local communities). The second wave was conducted by Kyiv School of Economics from January 1 to March 12, 2024, with 181 responses (14% of government-controlled municipalities). Data formats include CSV and SAV files, along with an XSL codebook for both waves. The merged dataset comprises 442 responses from small, medium, and large municipalities under varied security conditions, with a total file size of approximately 4 MB.
Prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, number of journals per country and publisher
<p>An analysis on the prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, country and publisher according to the number of journals.</p>
CONUS-wide Balancing Authority Scale Hydropower Projections derived from 9505 Third Assessment
<p>This dataset provides historical and climate projection monthly hydropower generation timeseries for balancing authorities within the contiguous U.S. (CONUS). These data were developed as an extension to the Department of Energy Water Power Technologies Office's SECURE Water Act Section 9505 Third Assessment (9505) and include both federal and non-federal hydropower facilities. Additional modeling detail can be found in <a href="https://iopscience.iop.org/article/10.1088/1748-9326/ad6ceb" target="_blank" rel="noopener">Broman et al., 2024</a> and in the article's <a href="https://github.com/9505-PNNL/broman-etal_2024_erl">metarepository</a>. </p> <p>The dataset is provided in three separate formats to facilitate ease of use:</p> <p>1) Machine-readable csv in 'tidy' data format:</p> <table> <tbody> <tr> <td><strong>Short Name</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>class</td> <td>N/A</td> <td>simulation type; control: historical, cc: climate scenario</td> </tr> <tr> <td>forcing</td> <td>N/A</td> <td>meteorological forcing used to drive hydrology model</td> </tr> <tr> <td>model</td> <td>N/A</td> <td>hydrology model</td> </tr> <tr> <td>hp</td> <td>N/A</td> <td>hydropower model</td> </tr> <tr> <td>gcm*</td> <td>N/A</td> <td>global climate model name</td> </tr> <tr> <td>ds*</td> <td>N/A</td> <td>downscaling method; DBCCA (statistical), RegCM (dynamical)</td> </tr> <tr> <td>balancing_authority</td> <td>N/A</td> <td>balancing authority code</td> </tr> <tr> <td>year</td> <td>N/A</td> <td>year</td> </tr> <tr> <td>month</td> <td>N/A</td> <td>month</td> </tr> <tr> <td>modeled_generation_MWh</td> <td>MWh per month</td> <td>simulated generation</td> </tr> </tbody> </table> <p>* only present in the climate projection (cc) files</p> <p>2) xlsx with balancing authority data by tab</p> <p>3) csv by balancing authority:</p> <p>for historical data: year, <em>month</em>, and <em>HUC4_group</em> columns are the same as above. Data column headers are <em>class</em>_<em>forcing</em>_<em>model</em>_<em>hp</em> and with the units <em>MWh per month</em>.</p> <p>for climate projection (cc) data: year, <em>month</em>, and <em>HUC4_group</em> columns are the same as above. Data column headers are <em>class</em>_<em>forcing</em>_<em>model</em>_<em>hp_gcm_ds</em> and with the units <em>MWh per month</em>.</p>
Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis
<p>Dataset of the research paper: <strong>Impact of Software Engineering Research in Practice: A Patent and Author Survey Analysis</strong></p> <p>Existing work on the practical impact of software engineering (SE) research examines industrial relevance rather than adoption of study results, hence the question of how results have been practically applied remains open. To answer this and investigate the outcomes of impactful research, we performed a quantitative and qualitative analysis of 4 354 SE patents citing 1 690 SE papers published in four leading SE venues between 1975–2017. Moreover, we conducted a survey on 475 authors of 593 top-cited and awarded publications, achieving 26% response rate. Overall, researchers have equipped practitioners with various tools, processes, and methods, and improved many existing products. SE practice values knowledge-seeking research and is impacted by diverse cross-disciplinary SE areas. Practitioner-oriented publication venues appear more impactful than researcher-oriented ones, while industry-related tracks in conferences could enhance their impact. Some research works did not reach a wide footprint due to limited funding resources or unfavorable cost-benefit trade-off of the proposed solutions. The need for higher SE research funding could be corroborated through a dedicated empirical study. In general, the assessment of impact is subject to its definition. Therefore, academia and industry could jointly agree on a formal description to set a common ground for subsequent research on the topic.</p> <p>The following data files are included.</p> <ul> <li><em>./fields</em>: <ul> <li><strong>engi-fields.csv</strong>: Publication and PhD dissertation counts of main engineering branches</li> <li><strong>engi-fields-queries.txt</strong>: Queries applied to Elsevier's Scopus and Open Access Theses and Dissertations databases to retrieve the publication and dissertation counts</li> </ul> </li> <li><em>./patents</em>: <ul> <li><strong>sample-se-references-verified.csv</strong>: Manual verification of a random sample of references by software engineering (SE) patents to SE papers</li> <li><strong>se-cpc.tsv</strong>: Manually-identified SE-related Cooperative Patent Classification (CPC) categories</li> <li><strong>se-references-in-patents.csv</strong>: SE references made by SE patents to SE papers</li> <li><em>./patents/litigation</em>: <ul> <li><strong>case-values.csv</strong>: Manually-retrieved litigation damages of citing SE patents</li> <li><strong>lit-per-paper.csv</strong>: Litigation cases of citing SE patents</li> </ul> </li> <li><em>./patents/maintenance</em>: <ul> <li><strong>maint-code-fee-mapping.csv</strong>: Mapping of patent maintenance fee codes to their fee values</li> <li><strong>maint-fees.csv</strong>: Fee values of maintenance fee codes</li> <li><strong>maint-per-paper.csv</strong>: Maintenance fee events of citing SE patents</li> </ul> </li> <li><em>./patents/reports</em>: <ul> <li><strong>lit-sum-per-paper.csv</strong>: Counts and total damages of litigation cases of patent-cited SE papers</li> <li><strong>maint-sum-per-paper.csv</strong>: Counts and total values of maintenance fee events of patent-cited SE papers</li> <li><strong>patent-ref-counts.csv</strong>: SE patent citation counts of patent-cited SE papers</li> </ul> </li> </ul> </li> <li><em>./survey</em>: <ul> <li><strong>emse-top.csv</strong>: Most-cited papers of the Empirical Software Engineering (EMSE) journal</li> <li><strong>icse-bp.csv</strong>: Distinguished papers of the International Conference of Software Engineering (ICSE)</li> <li><strong>icse-mip.csv</strong>: Most influential ICSE papers</li> <li><strong>icse-top.csv</strong>: Most-cited ICSE papers</li> <li><strong>survey-questionnaire-emse.pdf</strong>: The EMSE survey questionnaire</li> <li><strong>survey-questionnaire.pdf</strong>: The ICSE, TSE, and TOSEM survey questionnaire</li> <li><strong>survey-responses.csv</strong>: The anonymized survey responses</li> <li><strong>tosem-top.csv</strong>: Most-cited papers of the ACM Transactions on Software Engineering and Methodology (TOSEM)</li> <li><strong>tse-top.csv</strong>: Most-cited papers of the IEEE Transactions on Software Engineering (TSE)</li> <li><em>./survey/manual-coding</em>: <ul> <li><strong>feedback.txt</strong>: Manual coding of survey feedback</li> <li><strong>practical-impact.csv</strong>: Manual coding of responses about practical impact of work</li> <li><strong>practical-impact-lack.csv</strong>: Manual coding of responses about lack of practical impact</li> <li><strong>research-methods.csv</strong>: Manual coding of additional research methods of surveyed papers</li> <li><strong>state-of-practice.csv</strong>: Manual coding of responses about changes in state of practice</li> </ul> </li> </ul> </li> <li><em>./venues</em>: <ul> <li><strong>se-venues.csv</strong>: Top SE venues according to Google Scholar Metrics</li> <li><strong>se-venues-impact.csv</strong>: SE patent citations and patent-based impact factors of SE venues</li> <li><strong>se-venues-scopus-queries.txt</strong>: Queries applied to Scopus to retrieve the publication counts of the SE venues</li> </ul> </li> </ul>
Monitoring feedback to authors on the quality of trials evaluating interventions aimed at preventing and treating COVID-19
<p>We aimed to assess transparency of reporting and risk of bias of randomized trials evaluating interventions aimed at preventing and treating COVID-19.</p> <p>This review is part of a larger project: the COVID-NMA project (Boutron 2020a). The COVID-NMA project aims to provide decision-makers with a complete, high-quality and up-to-date synthesis of evidence on interventions for the prevention and treatment of COVID 19. For this purpose, we perform a living mapping of all registered randomized controlled trials and a living evidence synthesis of data from RCTs. We developed a master protocol on the effect of all interventions for the prevention and treatment of COVID-19 (first published on April 8, 2020; an update on May 11, 2020, June 17, 2020, and September 8, 2020) (Boutron 2020b). We set-up a platform (<a href="https://covid-nma.com/">https://covid-nma.com</a>) where all our results are made available and updated weekly.</p>
LSPO: A Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation
<p>The LSPO dataset, a Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation is comprised of 554,962 NASA/ADS publications linked to 125,486 unique researchers through ORCiD identifiers. The available meta-data fields are: ORCiD identifier, author name, affiliation, title, asbtract, and name block. The dataset can be utilized to make pairs or triplets for training a author name disambiguation model. </p>
Title, Author, Publisher, Place of Publication, and Language-related Network Graphs of the Berlin State Library Main Catalog
<p>The dataset contains graphs in GML, GraphML, and a simple JSON format.</p> <p>For each of the following languages:</p> <ol> <li>cze</li> <li>dan</li> <li>dut</li> <li>eng</li> <li>fre</li> <li>fry</li> <li>ger</li> <li>gre</li> <li>ice</li> <li>ita</li> <li>lat</li> <li>nor</li> <li>pol</li> <li>por</li> <li>rum</li> <li>rus</li> <li>slo</li> <li>spa</li> <li>swe</li> </ol> <p>two graphs are made available linking</p> <ul> <li>author, publisher, and place of publication</li> <li>author, publisher, place of publication, and title</li> </ul> <p>Additionaly, a third graph links authors and publishers to the language of publication (incl. year of the publication).</p> <p>The core statistics of each graph are outlined in <em>social_analysis_statistics.csv</em>. The smallest graph (fry, author_publisher_location) has 298 nodes and 264 edges, while the largest (ger, author_publisher_location_title) has 2,499,943 nodes and 3,950,900 edges.</p> <p>The language graphs spans all languages and has 1,706,273 nodes and 1,827,759 edges.</p> <p>All graphs have been created by a Python script available <a href="https://github.com/elektrobohemian/CulturalAnalytics/blob/master/SocialAnalysisStabikat.ipynb">here.</a></p>
Bibliographic dataset based on Scientometrics, containing provenance information compliant with the OpenCitations Data Model and non disambigued authors
<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known. The data was extracted via Crossref. It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works. Journals and bibliographic resources always appear unambiguously, without duplicates. On the contrary, the authors have not been disambigued. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p>
arXiv abstracts and titles from 1,469 single-authored papers (100 unique authors) in computer science
<p>This dataset is meant to be used for experiments of Authorship Analysis. The dataset consists of abstracts of single-author papers from arXiv crawled using the arXiv's API by querying a list of computer-science-related keywords ("deep learning", "machine learning", "information retrieval", "computer science", "data mining", "support vector", "logistic regression", "artificial intelligence", "supervised learning"'). The corpus somehow follows a power-law distribution, with few prolific authors and many authors accounting for very few papers each: we retained authors with at least 10 papers, resulting in a total of 1,469 documents from 100 authors. The most prolific authors (Peter D. Turney and Subhash Kak) have 34 abstracts to their names, the 10 most prolific authors have written 22 or more articles, while 50% of the authors have no more than 12 abstracts to their names. In order to divide the corpus into a training set and a test set we perform a stratified split, with the production of each author being split into a training set (70%) and a test set (30%). We use these documents as examples of "scientific communication", characterised by a precise and compact style, with an abundance of technical terminology.</p>
Pinterest dataset for age and gender identification in author profiling
<p>This dataset was used for the experiments presented in the article "Reconstructive Classification for Age and Gender Identification in Social Networks" - IEEE Transactions on Computational Social Systems</p> <p>The dataset contains text data from 548,761 pins corresponding to 264 users of Pinterest.</p> <p>There are 7 files.</p> <p>The first 5 files correspond to the extracted textual features from the pins that are aggregated per user: ats, emojis/emoticons, hashtags, links, and words.</p> <p>There are 264 lines in each file (one per user), as the concatenation of the extracted features from all the pins corresponding to each user.</p> <p>The last 2 files are the labels for the age and gender of the users. There are also 264 lines (one per user).</p> <p>For age, there are 4 possible labels: 18-24, 25-34, 35-46, and 50+</p> <p>For gender, there are 2 possible labels: F and M</p>
Mobile Application Privacy Risk Assessments from User-authored Scenarios
<p>Mobile applications (apps) provide users valuable benefits at the risk of exposing users to privacy harms. Improving privacy in mobile apps faces several challenges, in particular, that many apps are developed by low resourced software development teams, such as end-user programmers or in startups. In addition, privacy risks are primarily known to users, which can make it difficult for developers to prioritize privacy for sensitive data. In this paper, we introduce a novel, lightweight method that allows app developers to elicit scenarios and privacy risk scores from users directly using only an app screenshot. The technique relies on named entity recognition (NER) to identify information types in user-authored scenarios, which are then fed in real-time to a privacy risk survey that users complete. The best-performing NER model predicts information types with a weighted average precision of 0.70 and recall of 0.72, after post-processing to remove false positives. The model was trained on a labeled 300-scenario corpus, and evaluated in an end-to-end evaluation using an additional 203 scenarios yielding 2,338 user-provided privacy risk scores. Finally, we discuss how developers can use the risk scores to prioritize, select and apply privacy design strategies in<br> the context of four user-authored scenarios.</p>
List of articles resulting from the Google Scholar search "graph based author name disambiguation" published after 1/1/2021
<p>This dataset contains the list of articles resulting from the Google Scholar search “graph based author name disambiguation” published after 1/1/2021. The list is provided for reproducibility of the survey article “Graph-based Methods for Author Name Disambiguation: A Survey” and it was obtained using the following Python script available at <a href="https://github.com/WittmannF/sort-google-scholar">https://github.com/WittmannF/sort-google-scholar</a>:</p> <blockquote> <p>$ python sortgs.py --kw “graph based author name disambiguation” --startyear 2021</p> </blockquote> <p>The command returned the CSV file that contains the first 94 publications matching the query (articles with corrupted metadata have been excluded), each with metadata about Title, Number of Citations, and Rank. The CSV contains a column that specified which articles have been eventually selected for the survey.</p>
Table S3. List of Locustella sound recordings included in bioacoustic analysis surrounding description of the Taliabu Grasshopper-Warbler. The table provides information on sound library sources and sampling localities of recordings as well as raw data on all 11 bioacoustic parameters measured (see Supplementary Materials section SM3 for more details on parameters). Recordings whose source is labeled as "private recording" were obtained by colleagues and are available upon demand from the corresponding author.
<p>supplement to Rheindt, Frank E., Prawiradilaga, Dewi M., Ashari, Hidayat, Suparno, Gwee, Chyi Yin, Lee, Geraldine W. X., Wu, Meng Yue, Ng, Nathaniel S. R. (2020): A lost world in Wallacea: Description of a montane archipelagic avifauna. Science 367: 167-170, DOI: 10.1126/science.aax2146</p>
Geographical distribution of co-authors of Nobel laureates 1994-2018 in Physics, Chemistry and Physiology or Medicine
<p>Geographical distribution of co-authors of Nobel laureates 1994-2018 in Physics, Chemistry and Physiology or Medicine. Appendix to the article «Quantitative analysis of the co-publications of Ukrainian scientists with the Nobel laureates 1994-2018 in Science».</p>
Datasets for The Effect of COVID-19 on AGU Journal Authors by Gender and Geographical Location
<p>These files provide anonymized source data and tabular data on gender, age, and country of corresponding authors (submitting author) of American Geophysical Union (AGU) journals from January 2018 through June 2020. These datasets supplement an iposter presented at Japan Geosciences Union- American Geophysical Union joint 2020 meeting and supplement the corresponding preprint submission to ESSOAR.</p>
Brazilian Scientific Publication Records and Author Affiliations from Lattes until Feb 2017 (Anonymized)
<p>This file contains anonymized data about researcher profiles and publication records available extracted from the Lattes Platform in in February 2017 using the LattesDataXplorer tool. Lattes is a vast repository of researchers' curriculum vitae, widely adopted in Brazil. This platform is maintained by the Brazilian National Council of Scientific and Technological Development (CNPq) and is an internationally renowned initiative.</p> <p>In the zipped file, there are two files:</p> <ul> <li>anon_authors.csv contains data about researchers. It has 3 columns <ul> <li>profile: researcher anonymized id</li> <li>instituition: researcher affiliation</li> <li>zipcode: institution zip code</li> </ul> </li> </ul> <ul> <li>anon_papers.csv contains data about researchers' publications. It has 4 columns: <ul> <li>profile: researcher anonymized id</li> <li>year: publication year</li> <li>venue: publication venue</li> <li>authors: number of authors</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.