Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
Making Data Work: A Systematic Mapping of Collaborative Data Curation Practices
<p>This file contains dataset associated with the paper Making Data Work: A Systematic Mapping of Collaborative Data Curation Practices</p>
A Curated Solidity Smart Contracts Repository of Metrics and Vulnerabilities
<p>SmarthER provides the dataset related to the full-paper accepted to PROMISE 2024 (<a href="https://promiseconf.github.io/2024/index.html" rel="nofollow">https://promiseconf.github.io/2024/index.html</a>) <strong>"A Curated Solidity Smart Contracts Repository of Metrics and Vulnerability"</strong>.</p> <p>Authored by: Giacomo Ibba, Sabrina Aufiero, Rumyana Neykova, Silvia Bartolucci, Roberto Tonelli, Marco Ortu, Giuseppe Destefanis</p> <p>This repository aims to collect a significant sample of smart contracts with associated vulnerability reports, and traditional software metrics extracted from each smart contract. The repository contains:</p> <ul> <li>Smart contracts source code.</li> <li>The vulnerability report was built with Slither for each contract.</li> <li>Traditional software metrics extracted from each contract.</li> </ul>
Dataset for training SENMAP, a automatic tool to curate LTR-retrotransposons using convolutional neural networks
<p>Transposable elements (TEs) are specific structures of the genome of species, which can move from one location to another. For that reason, they can cause mutations or changes that can be negative, such as the appearance of diseases, or beneficial, such as participating in fundamental roles in the evolution of genomes and genetic diversity. Long Terminal Repeat retrotransposons (LTR-RT) are the most abundant in plant species, hence the importance of studying these structures in particular. Over the time, these elements can suffer changes called nested insertions, which can inactivate or modify the functioning of the element, for that they are no longer consider as intact element and cannot be used for identification and classification studies. We create a dataset containing 56,442 LTR-RTs targed as "non-intact" elements and 49,215 considered as "intact". </p> <p>We formated the sequences IDs in order to keep relevant information as the superfamily and the lineage, as well as the category (Negative for "non-intact" and Positive for "intact" elements). </p> <p> This dataset (the npy files obtained from the fasta file) was used for training SENMAP, a convolutional neural network architecture to obtain intact LTR-RT sequences in plant genomes, which is composed by four convolutional layers, LeakyReLU as activation function and BinaryFocalLoss as loss function. Achieving an F1-score percentage of 91.37% with test data, identifying low quality sequences rapidly and efficiently, contributing to curate libraries of LTR retrotransposons of plants genomes published in large-scale sequencing projects due to the post-genomic era.</p>
modelforge curated dataset: QM9 1000 conformer test dataset
<h1>Curated QM9 1000 Conformer Test Dataset:</h1> <p>This provides a curated hdf5 file for a subset of the QM9 dataset to be used for testing purposes, designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains in total 1000 conformers (1 conformer per unique molecule).</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <p>This curated dataset was generated using the modelforge software at commit c5c7153:</p> <ul> <li> Link to the source code at this commit: <a href="https://github.com/choderalab/modelforge/tree/c5c7153e06172fe8e6f25015250ecb5db05655cc">https://github.com/choderalab/modelforge/tree/c5c7153e06172fe8e6f25015250ecb5db05655cc</a></li> <li>Link to the script file used to generate the dataset: <a href="https://github.com/choderalab/modelforge/blob/c5c7153e06172fe8e6f25015250ecb5db05655cc/modelforge/curation/qm9_curation.py">https://github.com/choderalab/modelforge/blob/c5c7153e06172fe8e6f25015250ecb5db05655cc/modelforge/curation/qm9_curation.py</a></li> </ul> <p> </p> <h2>Original QM9 Dataset:</h2> <p>The QM9 dataset includes 133,885 organic molecules with up to nine total heavy atoms (C,O,N,or F; excluding H) original published by Ramakrishnan, et al. Properties in the QM9 dataset were calculated at the B3LYP/6-31G(2df,p) level of quantum chemistry.</p> <h3>Citations:</h3> <p><em>Original publication:</em></p> <ul> <li>Ramakrishnan, R., Dral, P., Rupp, M. et al."Quantum chemistry structures and properties of 134 kilo molecules." Sci Data 1, 140022 (2014). <a href="https://doi.org/10.1038/sdata.2014.22">https://doi.org/10.1038/sdata.2014.22</a></li> </ul> <p><em>Source dataset, released with CCO 1.0 Universal license:</em></p> <ul> <li>Ramakrishnan, Raghunathan; Dral, Pavlo; Rupp, Matthias; Anatole von Lilienfeld, O. (2014). Quantum chemistry structures and properties of 134 kilo molecules. figshare. Collection. <a href="https://doi.org/10.6084/m9.figshare.c.978904.v5">https://doi.org/10.6084/m9.figshare.c.978904.v5</a></li> </ul>
Dataset for 'The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track'
<p>This packages comprises of analyses and evaluations of 60 datasets from the NeurIPS Datasets and Benchmarks track. It is part of a paper currently under review at the 2024 the NeurIPS Datasets and Benchmarks track, titled, "The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track". </p>
Fig. 2 in Evaluation of insecticides for curative, preventive, and rotational use on Scirtothrips dorsalis South Asia 1 (Thysanoptera: Thripidae)
Fig. 2. Mean numbers of Scirtothrips dorsalis adults per 10 leaf samples of Jalapeno pepper treated with different insecticides. Solid lines represent treatments where 1 insecticide was applied alone, and dashed lines show treatments of 2 insecticides applied in rotation. Same color solid and dashed lines represent same insecticide applied alone or in rotation with spinetoram.
Fig. 1 in Evaluation of insecticides for curative, preventive, and rotational use on Scirtothrips dorsalis South Asia 1 (Thysanoptera: Thripidae)
Fig. 1. Mean numbers of Scirtothrips dorsalis larvae per 10 leaf samples of Jalapeno pepper treated with different insecticides. Solid lines represent treatments where 1 insecticide was applied alone, and dashed lines show treatments of 2 insecticides applied in rotation. Same color solid and dashed lines represent same insecticide applied alone or in rotation with spinetoram.
Results of user research project to understand data curation practices
<p>Supporting scalable curation is a part of the mission of the Elixir Data Platform.Thus far, we have established infrastructure capable of ingesting and aggregating text-mined outputs from multiple providers and making these available via an API. This public API is used by Europe PMC to display specific entities and relationships on full text articles (via the SciLite application). To ensure that the future development of this infrastructure meets the needs of curators, we first carried out user research to understand and identify common workflow patterns and practices via an observational study. Building on these outcomes, we then devised a curator community survey to more specifically understand which entity types, sections of a paper and tools are of top priority to address. The results of the project is presented here.</p>
Figure 1 in EphemBrazil: a curated online database and dashboard to explore the distribution of mayflies (Insecta: Ephemeroptera) from Brazil
Figure 1 General view of the website and interactive map view tab showing filters on the top and family subtitles in the right corner. Note that no filter is applied and all records are shown.
Data curation materials in "Daily life in the Open Biologist's second job, as a Data Curator"
<p>This is the supplementary material accompanying the manuscript "Daily life in the Open Biologist’s second job, as a Data Curator", published in <a href="https://doi.org/10.12688/wellcomeopenres.22899.1">Wellcome Open Research</a>. </p> <p>It contains:</p> <p><strong>- Python_scripts.zip</strong>: Python scripts used for data cleaning and organization:</p> <p> -add_headers.py: adds specified headers automatically to a list of csv files, creating new output files containing a "_with_headers" suffix.</p> <p> -count_NaN_values.py: counts the total number of rows containing null values in a csv file and prints the location of null values in the (row, column) format.</p> <p> -remove_rowsNaN_file.py: removes rows containing null values in a single csv file and saves the modified file with a "_dropNaN" suffix.</p> <p> -remove_rowsNaN_list.py: removes rows containing null values in list of csv files and saves the modified files with a "_dropNaN" suffix.</p> <p><strong>- README_template.txt</strong>: a template for a README file to be used to describe and accompany a dataset. </p> <p><strong>- template_for_source_data_information.xlsx</strong>: a spreadsheet to help manuscript authors to keep track of data used for each figure (e.g., information about data location and links to dataset description).</p> <p><strong>- Supplementary_Figure_1.tif</strong>: Example of a dataset shared by us on Zenodo. The elements that make the dataset FAIR are indicated by the respective letters. Findability (F) is achieved by the dataset unique and persistent identifier (DOI), as well as by the related identifiers for the publication and dataset on GitHub. Additionally, the dataset is described with rich metadata, (e.g., keywords). Accessibility (A) is achieved by the ease of visualization and downloading using a standardised communications protocol (https). Also, the metadata are publicly accessible and licensed under the public domain. Interoperability (I) is achieved by the open formats used (CSV; R), and metadata are harvestable using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH), a low-barrier mechanism for repository interoperability. Reusability (R) is achieved by the complete description of the data with metadata in README files and links to the related publication (which contains more detailed information, as well as links to protocols on protocols.io). The dataset has a clear and accessible data usage license (CC-BY 4.0).</p>
Curation of Vibrio biofilm matrix cluster and associated proteins
<p>The supplementary data and other supporting materials for the paper titled "Comprehensive Genomic and Evolutionary Analysis of Biofilm Matrix Clusters and Proteins in the <em>Vibrio </em>Genus".</p>
Video Tutorials for Using the Solidipes Curation Tool
<h2>Contributions</h2> <ul> <li>Emmanuelle Denove realized the project, conceived the videos and proceeded to the curated export into Zenodo</li> <li>Guillaume Anciaux and Son Pham-Ba supervised the realization by attending brainstorming meetings and by providing feedbacks.</li> </ul> <h2>Data structure and information</h2> <p>This dataset contains 12 videos exported as .mp4 files, as well as their source .prproj (Premiere Pro Project) files and all media used in them. For each individual video, the name of the .mp4 file, raw file and media folder is identical. The structure of the data inside the "data" folder is as follows :</p> <ul> <li>Folder/files structure: <ul> <li><code>final_videos</code> - folder containing the final exported videos <ul> <li><code>solidipes-*.mp4</code> - the individual videos</li> </ul> </li> <li><code>raw_files</code> - folder containing the Premiere Pro source files <ul> <li><code>solidipes-*.prproj</code> - source files</li> </ul> </li> <li><code>media</code> - folder containing the media folder for every video <ul> <li><code>solidipes-*</code> - folders containing the raw media for each video</li> </ul> </li> </ul> </li> </ul> <h2>Funding</h2> <p><a href="https://ethrat.ch/en/measure-1-calls-for-field-specific-actions/">ETH-Board ORD, measure 1</a>, grant n°22945: DCSM - Cloud and web based platform for dissemination of computational solid mechanics</p> <h2>Video details</h2> <p>This section gives a short description of each video as well as its identifier in the repository.</p> <ul> <li>Creating a dataset curation: <ul> <li>identifier : <code>dcsm</code></li> <li>This video shows how to upload a dataset to be hosted on DCSM's internal servers, from which they can directly be curated with solidipes.</li> </ul> </li> <li>Initiate solidipes: <ul> <li>identifier : <code>solidipes-initialise</code></li> <li>This video shows how to initialise a directory as a solidipes curation from the terminal.</li> </ul> </li> <li>Intro and installation: <ul> <li>identifier : <code>solidipes-installation</code></li> <li>This video gives a brief intro to solidipes and shows how to install it from the terminal using pip.</li> </ul> </li> <li>What is solidipes ? : <ul> <li>identifier : <code>solidipes-intro</code></li> <li>This video gives a detailed introduction into what solidipes does and what its goals are.</li> </ul> </li> <li>Data curation from the terminal : <ul> <li>identifier : <code>solidipes-terminal-curation</code></li> <li>This video shows how a dataset can be curated on the terminal using solidipes.</li> </ul> </li> <li>Download from Zenodo : <ul> <li>identifier : <code>solidipes-terminal-download</code></li> <li>This video shows how to download a dataset from Zenodo from the terminal using solidipes.</li> </ul> </li> <li>Upload to Zenodo from the terminal : <ul> <li>identifier : <code>solidipes-terminal-export</code></li> <li>This video shows how to export a dataset curated with solidipes to Zenodo from the terminal.</li> </ul> </li> <li>Acquisition on the web service : <ul> <li>identifier : <code>solidipes-web-acquisition</code></li> <li>This video shows the aspects of the "acquisition" step of the solidipes web service.</li> </ul> </li> <li>Curation on the web service : <ul> <li>identifier : <code>solidipes-web-curation</code></li> <li>This video shows the aspects of the "curation" step of the solidipes web service.</li> </ul> </li> <li>Exporting from the web service : <ul> <li>identifier : <code>solidipes-web-export</code></li> <li>This video shows how to export a dataset to Zenodo from the solidipes web service.</li> </ul> </li> <li>Edit metadata on the web service : <ul> <li>identifier : <code>solidipes-web-metadata</code></li> <li>This video shows the aspects of the "metadata" step of the solidipes web service.</li> </ul> </li> <li>Starting the web service + intro : <ul> <li>identifier : <code>solidipes-web-overview</code></li> <li>This video shows how to start the web service from the terminal, and gives an overview of its aspects.</li> </ul> </li> </ul> <h2>Image credit</h2> <p>The following graphics were used in some of the videos :</p> <ul> <li><a href="https://www.svgrepo.com/svg/533602/arrow-narrow-left">https://www.svgrepo.com/svg/533602/arrow-narrow-left</a></li> <li><a href="https://www.svgrepo.com/svg/145033/text-file-document">https://www.svgrepo.com/svg/145033/text-file-document</a></li> <li><a href="https://www.svgrepo.com/svg/184892/scientist">https://www.svgrepo.com/svg/184892/scientist</a></li> <li><a href="https://icons8.com/icon/12245/image-file">https://icons8.com/icon/12245/image-file</a></li> </ul>
Dataset - Papyrus - A large scale curated dataset aimed at bioactivity predictions
<p><strong>Fixed version of additional_files:</strong></p> <p><strong>- In the previous version of 05.6_additional_files</strong> the data type of some descriptors was assigned incorrectly</p> <p><strong>- In this fixed version </strong>data types are correct </p> <p>This repository contains version 05.6 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" doi.org/10.1186/s13321-022-00672-x.</p> <p>Changes compared to version 05.5</p> <p>- applied small molecule filter that filters out compounds with a MW < 200 or > 800, heavy metal containing compounds and mixtures</p> <p>- include TID column which contains information on the original protein identifier</p>
BFR2: a curated benthic foraminifera ribosomal reference database
<p>The present data set provides a fasta file, a tab-separated text file, and an Excel file. The fasta file includes 5,324 18S rDNA reference sequences for benthic foraminifera. The tab-separated text file includes the following fields: BFR2 number = unique internal sequence accession number; length = sequence length; class = class to which each sequence is assigned; order/suborder/clade = order/suborder/clade to which each sequence is assigned; family = family to which each sequence is assigned; genus = genus to which each sequence is assigned; species = species to which each sequence is assigned; isolate number = unique DNA extraction number; clone/direct: indicates whether it has been directly sequenced or cloned; genbank_accession = NCBI sequence accession number; sampling site = biogeographic region where the specimen has been collected; latitude in decimal degrees, longitude in decimal degrees; collection date = year in which specimen was collected; collector = person who collected the specimen; publication = publication associated to the sequence; journal = journal associated to publication; first author = first author associated to publication/sequence; taxonomic remarks = additional taxonomic information/comment; sampling remarks = addional sampling information/comment.<br><br>List of added or updated:<br>- Clone/direct: indicates whether it has been directly sequenced or cloned.<br>- Latitude and longitude in decimal degrees.<br>- Genbank accession number of 1700 18S rRNA sequences added.</p>
Artificial Curators / CuratorBot - Questions, Reactions, Survey Responses
<p>Three datasets captured through the CuratorBot digital demo during 2022 and 2023:</p> <ol> <li>questions posed to the CuratorBot and responses generated by the system</li> <li>quick like/dislike (thumbs-up/thumbs-down) feedback given by users to specific responses</li> <li>anonymous user responses to a survey attached to the CuratorBot</li> </ol>
Linked collectors and determiners for: The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium).
Natural history specimen data linked to collectors and determiners held within, "The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/bddbc2b5-5215-4996-b553-241e06e029ee">https://bionomia.net/dataset/bddbc2b5-5215-4996-b553-241e06e029ee</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/bddbc2b5-5215-4996-b553-241e06e029ee">https://gbif.org/dataset/bddbc2b5-5215-4996-b553-241e06e029ee</a>. Formatted as a Frictionless Data package.
Facebook posts for analyzing Content Strategies for Digital Consumer Engagement: a curated dataset
<p>This database contains public data from publications made on Facebook by small and medium companies in the tourism sector operating in the Amazon region in Brazil. The collection was carried out from January to June 2018 using the public API provided by the platform.</p> <p>These data were processed and classified by different evaluators according to the content categories proposed by <a href="https://doi.org/10.1080/00913367.2017.1405751">Gavilanes</a>. Thus, there is a column in the dataset called "category", which contains the classification of each publication, with each number is associated with a category, as described below:</p> <ul> <li> <p>“0” - No category</p> </li> <li> <p>“1” - New product announcement</p> </li> <li> <p>“2” - Sweepstakes and contest</p> </li> <li> <p>“3” - Sales</p> </li> <li> <p>“4” - Consumer Feedback</p> </li> <li> <p>“5” - Infotainment</p> </li> <li> <p>“6” - Organization Branding</p> </li> <li> <p>“N\A” - Non-agreement between evaluators</p> </li> </ul> <p>Therefore, the columns with data present in the CSV file are:</p> <ul> <li> <p>“status_id” - Identification of each publication in the social network. String Textual content of each publication;</p> </li> <li> <p>“status_message” - The textual content of each publication;</p> </li> <li> <p>“link_name” - Which part of the profile the post are related;</p> </li> <li> <p>“status_type” - The type of the publication is a ’photo’, ’video’, ’status’ and/or ’link’;</p> </li> <li> <p>“status_link” - URL to the publications on the platform;</p> </li> <li> <p>“status_published” - Publication date in the social network;</p> </li> <li> <p>“num_comments” - Numbers of comments made by users;</p> </li> <li> <p>Reactions - Numbers of emoticons reactions from users on post – these numbers are presents in the columns: "num_reactions", "num_shares", "num_likes", "num_loves", "num_wows", "num_hahas", "num_sads", "num_angrys", "num_special";</p> </li> <li> <p>“category” - The category of content that we mentioned before.</p> </li> </ul>
Fig. 59. Curator Richard G in A History Of Herpetology At The American Museum Of Natural History
Fig. 59. Curator Richard G. Zweifel on his first New Guinea expedition (at Mt. Rawlinson, 1964). AMNH Dept. Herpetology Archives.
Volker Mahnert on his first zoological expedition as curator at the MHNG; together with Ivan Löbl (left) on the Ionian island of Zakynthos in 1971 (photo B. Hauser). in Volker Mahnert 3 December 1943 – 23 November 2018
Volker Mahnert on his first zoological expedition as curator at the MHNG; together with Ivan Löbl (left) on the Ionian island of Zakynthos in 1971 (photo B. Hauser).
Curated ataxia gene list
<p>A curated ataxia gene list, curated on 2022-05-31. The list contains genes from OMIM, PanelApp Australia, PanelApp UK and HPO. Genes are labelled as high (1), moderate (2) and low (3) priority genes, based on the level of evidence for causing ataxia.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.