Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

486

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

486 results for “curation”

Learn how ShareScore rates datasets ↗
zenodo40/100

PHI-Canto approved curation sessions: December 2022

<p>Approved curation sessions from the <a href="https://github.com/PHI-base/canto">PHI-Canto</a> curation tool, as of 13 December 2022. PHI-Canto is used to curate literature on pathogen&ndash;host interactions, and supplies data to <a href="http://www.phi-base.org/">PHI-base</a>, the Pathogen&shy;&ndash;Host Interactions Database.</p> <p>The curated data is exported in JSON format, and contained in a single JSON object. The object&#39;s keys are the identifiers for individual curation sessions, where each curation session corresponds to one publication. The export contains the raw data exported by PHI-Canto: no further processing has been applied.</p> <p>There is a JSON Schema file included that describes the data fields used in the export file.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

A Novel Curated Scholarly Graph Connecting Textual and Data Publications

<p>This dataset contains an open and curated scholarly graph we built&nbsp;as a training and test set for data discovery, data connection, author disambiguation, and link prediction tasks.&nbsp;This graph represents the European Marine Science community included in the OpenAIRE Graph.&nbsp;The nodes of the graph we release&nbsp;represent publications, datasets, software, and authors respectively; edges interconnecting research products always have the publication as source, and the dataset/software as target. In addition, edges are labeled with semantics that outline whether the publication is <em>referencing, citing, documenting</em>, or <em>supplementing</em> the related outcome. To curate and enrich nodes metadata and edges semantics, we relied on the information extracted from the PDF of the publications and the datasets/software webpages respectively. We curated the authors so to remove duplicated nodes representing the same person.&nbsp;</p> <p>The resource we release counts 4,047 publications, 5,488 datasets, 22 software, 21,561 authors, and 9,692 edges connect publications to datasets/software. This graph is in the <em>curated_MES</em>&nbsp;folder. We provide this resource as:</p> <ol> <li>a property graph: we provide the dump that can be imported in neo4j</li> <li>5 jsonl files containing publications, datasets, software, authors, and relationships respectively. Each line of a jsonl file contains a JSON object representing a node and contains the&nbsp;metadata of that&nbsp;node (or a relationship).</li> </ol> <p>We provide two additional scholarly graphs:</p> <ul> <li>The curated MES graph with the removed edges. During the curation we removed some edges since&nbsp;they were labeled with an inconsistent or imprecise semantics. This graph includes the same nodes and edges as the previous one, and, in addition, it contains the edges removed during the curation pipeline; these edges are marked as <em>Removed</em>.&nbsp;This graph is in the <em>curated_MES_with_removed_semantics</em> folder.<br> &nbsp;</li> <li>The original MES community of OpenAIRE. It represents the MES community extracted from the OpenAIRE Research Graph. This graph has not been curated, and the metadata and semantics are those of the OpenAIRE Research Graph. This graph is in the <em>original_MES_community</em> folder.</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Large curated dataset for drug target interaction

<p>A large data curation from the PubChem, ChEMBL and BindingDB public sources. The curated dataset includes samples of pairs of small molecules and protein targets, with information about their&nbsp;binding interactions. The data is stored in an efficient tables format, decoupling entity IDs from their string representations, to avoid redundancy. The curation also includes meaningful splits of the dataset into train, validation and test sets for the purpose of utilizing it for learning based affinity prediction models.</p>

opencc-by-sa-3.0Jul 2023View details →
dryad40/100

Genomic characterization and gene bank curation of Aegilops using genotyping-by-sequencing

<p>In this study, genotyping-by-sequencing (GBS) was performed on 1041 <em>Aegilops</em> accessions, representing 23 different species. These accessions have been maintained by the Wheat Genetics and Resource Center (WGRC) at Kansas State University. The GBS FASTQ files have been uploaded to the NCBI SRA public repository under the BioProject accession number # PRJNA985892. We have provided other files related to data analysis, such as the barcode key file, SNP matrices, and taxonomic information of the accessions in this Dryad repository, which can be accessed through the provided link.  The aim of the study was to explore the genetic and genomic characteristics of wild wheat relatives, <em>Aegilops,</em> using a larger number of SNP markers. Here, we also curated the WGRC gene bank <em>Aegilops</em> collection via the identification of misclassified accessions and genetically identical redundant accessions. Further, we explored the genomic relationship between wheat and the different <em>Aegilops</em> species. </p>

opencc-zeroJul 2023View details →
zenodo40/100

circAtlas 3.0: a gateway to 3 million curated vertebrate circular RNAs through a standardized nomenclature scheme

<p>Data and code for the manuscript &quot;circAtlas 3.0: a gateway to 3 million curated vertebrate circular RNAs through a standardized nomenclature scheme&quot;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

A large comprehensive curated dataset of small molecules and their activities covering three cardiac ion channels: hERG, Cav1.2, and Nav1.5

<p>The compressed data folder (dataset.rar) represents a data framework for researchers in the field of drug discovery to perform in depth analyses on a very large open-access unique and comprehensive hERG, Nav1.5, and Cav1.2 cardiotoxicity integrated database of small molecules and their activities. The database is organized as follows:</p> <ul> <li>Each sub-folder represents a cardiac ion channel target: hERG, Nav1.5, and Cav1.2</li> <li>Each target sub-folder consists of 3 files in CSV format: One file containing the development set (split into training and validation sets using an 80/20&nbsp;ratio&nbsp;for hyperparameter tuning). The other 2 files contain external evaluation sets. The first test dataset consists&nbsp;of compounds with a structural similarity of no more than 60% (Tanimoto similarity&nbsp; &le; 0.6) to the remaining development set, while the second test dataset comprises&nbsp;compounds with a structural similarity of no more than 70% (Tanimoto similarity &le; 0.7) to the remaining development set.</li> <li>Each file contains data with 7&nbsp;columns: "InChl Key" as a unique identifier of the chemical structure, "SMILES" as the string format of storage and exchange of the chemical structure, "Source" as the upstream data source from which the data was&nbsp;retrieved, "ChEMBL ID" as the&nbsp;ChEMBL identifier if the compound comes from&nbsp;ChEMBL database,&nbsp;"PubChem CID" as the&nbsp;PubChem compound&nbsp;identifier if the compound comes from&nbsp;PubChem database,&nbsp;"pIC50" as the&nbsp;negative logarithm of the half-maximal inhibitory concentration (IC50) to describe the potency of the compound, and "USED_AS" column specifying whether the compound was used for training or validation.</li> </ul> <p><strong>Upon usage, please cite this publication:</strong></p> <ul> <li>Issar Arab, Kristof Egghe, Kris Laukens, Ke Chen, Khaled Barakat, Wout Bittremieux, <strong>Benchmarking of Small Molecule Feature Representations for hERG, Nav1.5, and Cav1.2 Cardiotoxicity Prediction</strong>, <em>Journal of Chemical Information and Modeling</em>, (2023). doi:<a href="https://doi.org/10.1021/acs.jcim.3c01301">10.1021/acs.jcim.3c01301</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Genomic characterization and gene bank curation of Aegilops using genotyping-by-sequencing

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad40/100

Northern elephant seal tracking and diving – raw and curated data

Open the record for dataset details and reuse information.

publicMay 2025View details →
zenodo36/100

A curated dataset of aerial survey images over the central Congo Basin, 1958

<p>This dataset contains a subset data from the Belgian Science Policy Office funded &ldquo;Congo basin eco-climatological data recovery and valorisation&quot; project (COBECORE, contract BR/175/A3/COBECORE).</p> <p>The data included is curated and pre-processed aerial survey imagery as used in a a land-use land-cover change analysis &quot;Historical aerial surveys map long-term changes of forest cover and structure in the central Congo Basin&quot;.</p> <p>&nbsp;The dataset includes:</p> <ul> <li>the pre-processed images (aerial_images.tar.gz)</li> <li>the meta-data associated with the aerial images (flight_paths*)</li> <li>the final orthomosaic (yangambi_orthomosaic.tif)</li> </ul> <p>For the full methodology we refer to the full paper:</p> <p><strong>Hufkens K.</strong>, et al. (2020) Historical Aerial Surveys Map Long-Term Changes of Forest Cover and Structure in the Central Congo Basin. <strong> Remote Sensing</strong>, 12, 638.</p> <p>Please cite the work as such.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Curation and ISA representation of a SARS-Cov2/Covid-19 Proteomics Dataset - PXD107710 - ISA representation

<p>Curation and ISA representation of a SARS-Cov2/Covid-19 Proteomics Dataset deposited in PRIDE database with accession number:&nbsp;PXD107710</p> <p>ISA-Tab annotation for the &nbsp;&quot;SARS-CoV-2 infected host cell proteomics reveal potential therapy targets&quot; publication.&nbsp;</p> <p>Github repository:&nbsp;<a href="https://github.com/ISA-tools/PXD017710">https://github.com/ISA-tools/PXD017710</a></p> <p>This is part of an effort to (re-)annotate:&nbsp;<a href="https://dx.doi.org/10.21203/rs.3.rs-17218/v1">https://dx.doi.org/10.21203/rs.3.rs-17218/v1</a></p> <p>Additional work done as part of:</p> <ol> <li>&nbsp;<a href="https://github.com/virtual-biohackathons/covid-19-bh20">https://github.com/virtual-biohackathons/covid-19-bh20</a></li> <li>&nbsp;<a href="https://github.com/virtual-biohackathons/covid-19-bh20/wiki/FairData">https://github.com/virtual-biohackathons/covid-19-bh20/wiki/FairData</a></li> </ol> <p><strong>Proteomics data</strong></p> <p>Available from PRIDE at <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD017710]">https://www.ebi.ac.uk/pride/archive/projects/PXD017710</a><br> and [<a href="https://massive.ucsd.edu/ProteoSAFe/result.jsp?task=334df9b4f1af4501bca0a2aa63278a7d&amp;view=display_metadata_results&amp;file=f.RMSV000000308%2F2020-03-22_nuno_334df9b4%2Fmetadata%2FMSV000085096_SARS-CoV-2_proteome_translatome.csv#%7B%22table_sort_history%22%3A%22_dyn_%23Condition_asc%22%7D">MassIVE/CCMS Maestro+MSstats reanalysis of MSV000085096 / PXD017710</a>]</p> <p><strong>ISA-Tab representation:</strong></p> <p>Rationale: Demonstrate suitability of the ISA format for representing MS based protein profiling experiment with more granularity and details, thus providing&nbsp;a better representation of the experiment design.<br> The formatting and re-annotation are based on information extracted from:<br> - the original publication<br> - the supplementary tables available from the publishers site<br> - the &#39;filtered-results.csv&#39; helper file as supplied to @sneumann during the <a href="http://www.psidev.info/hupo-psi-meeting-2020">HUPO-PSI meeting March 2020</a></p> <p><br> Viewing the ISA-tab formatted and re-annotated PXD017710 with I<a href="https://isa-tools.org/PXD017710/isaviewer-demo.html)">SATab-Viewer</a></p> <p>Viewing the ISA-tab formatted and re-annotated PXD017710 locally, do the following:</p> <p>```bash<br> python -m http.server 8000<br> ```</p> <p>Then point your browser to `http://0.0.0.0:8000/isaviewer-demo.html`</p> <p><strong>Curation tasks performed:</strong></p> <p>* initial structure of the study design in ISA format:</p> <p>* linkage of Proteome and Translatome data (supplementary material) to ISA assay tables (via Derived Data File)</p> <p>* processing the Proteome and Translatome data (supplementary material) with python pandas library to generate the following csv files:</p> <p>&nbsp;&nbsp; &nbsp;- proteome_intensities_long_table_ggplot2.txt<br> &nbsp;&nbsp; &nbsp;- proteome_diffanal_ratio_pvalue_long_table_ggplot2.txt<br> &nbsp;&nbsp; &nbsp;- translatome_intensities_long_table_ggplot2.txt&nbsp;&nbsp; &nbsp;<br> &nbsp;&nbsp; &nbsp;- translatome_diffanal_ratio_pvalue_long_table_ggplot2<br> &nbsp;&nbsp; &nbsp;<br> &nbsp;&nbsp; &nbsp;The files are `long table` corresponding to a `melt` on the Excel file originally generated by the users and can be readily loaded in R ggplot2 library for graphical representation.<br> &nbsp;&nbsp; &nbsp;The statistical relevant elements have been annotated with the <a href="http://stato-ontology.org/">STATO ontology</a>&nbsp;and the tables comply with a Frictionless.io Data Package.<br> &nbsp;&nbsp; &nbsp;The jupyter notebook for the transformation is available.</p> <p>* conversion of raw data to mzML format: detailed in&nbsp;<a href="https://github.com/ISA-tools/PXD017710">https://github.com/ISA-tools/PXD017710</a></p> <p>install docker:&nbsp;<br> ```bash<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&gt;brew update<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&gt;brew install docker<br> ```</p> <p>sign in to docker<br> ```bash<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&gt;docker start<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&gt;docker login<br> ```</p> <p>pull docker container for ProteoWizard:<br> ```bash<br> &gt;docker pull chambm/pwiz-i-agree-to-the-vendor-licenses<br> ```</p> <p>:warning: be sure to sign-up and login to https://hub.docker.com/</p> <p>in order to be able to reach</p> <p>https://hub.docker.com/r/chambm/pwiz-skyline-i-agree-to-the-vendor-licenses</p> <p><br> run the pwiz tool from the container over the raw data:<br> ```bash<br> &nbsp;docker run -it --rm -e WINEDEBUG=-all -v /Users/Downloads/PXD017710/raw/:/data chambm/pwiz-skyline-i-agree-to-the-vendor-licenses wine msconvert /data/*.raw --mzML<br> ```</p> <p><br> * ontology markup for:<br> &nbsp;&nbsp; &nbsp;* declaration of independent variables as ISA Study Factors:{biological agent, dose, time point, replicate} -&gt;OBI<br> &nbsp;&nbsp; &nbsp;* Taxonomic information (host cells and virus) -&gt; NCBITaxonomy<br> &nbsp;&nbsp; &nbsp;* Cell line: CaCo-2 cells -&gt; Cell Line Ontology<br> &nbsp;&nbsp; &nbsp;* Disease: Colon Cancer -&gt; Human Phenotype Ontology<br> &nbsp;&nbsp; &nbsp;* MS specific aspect (TMT reagent, instrument ... ) -&gt; PSI-MS<br> &nbsp;&nbsp; &nbsp;* Statistical Tests -&gt; STATO</p> <p><br> <strong>Unresolved curatorial issues:</strong></p> <p>&nbsp;1. ambiguities related to Tandem Mass Tag labelling protocol<br> &nbsp; &nbsp; - the publication mentions TMT11 (see Figure 2 in https://www.researchsquare.com/article/rs-17218/v1)<br> &nbsp; &nbsp; - the information available from PRIDE mentions TMT6 (https://www.ebi.ac.uk/pride/archive/projects/PXD017710)<br> &nbsp; &nbsp; This may require another round of annotation on the TMT agents and fractions in the ISA a_assay representation</p> <p><br> &nbsp;2. SARS-Cov2 isolate: no clear NCBI Taxonomic anchoring and unclear origin: -&gt; the markup is made to the parent class (as of 06.04.2020)</p> <p><strong>Release and packaging as a BDBAG:</strong></p> <p>The tgz file associated with this upload has been producing using&nbsp;<a href="https://github.com/fair-research/bdbag">https://github.com/fair-research/bdbag</a>. It contains several manifest files detailing metadata and data files, providing md5 and sha256 checksums.</p> <p><strong>Github repository:</strong>&nbsp;<a href="https://github.com/ISA-tools/PXD017710">https://github.com/ISA-tools/PXD017710</a></p>

opencc-by-3.0Apr 2020View details →
zenodo36/100

Dataset with Curated Coronavirus-related R&D Outputs (1970 - March 2020)

<p><strong>The zip file includes metadata about R&amp;D outputs related to all </strong><a href="https://www.niaid.nih.gov/diseases-conditions/coronaviruses"><strong>Coronaviruses</strong></a>.</p> <p>The two main data sources used for this work include:</p> <ul> <li> <p><a href="https://docs.microsoft.com/en-us/academic-services/project-academic-knowledge/introduction"><strong>Microsoft Academic Graph (MSA) API</strong></a>: Subset of 10.000+ coronavirus R&amp;D outputs, with data from 1970, including patents and scientific publications.</p> </li> <li> <p><a href="https://www.grid.ac/"><strong>The Global Research Identifier Database (GRID)</strong></a><strong>:</strong> Used to enrich the organization data extracted from MSA.</p> </li> </ul> <p>After cleaning and enriching the data we extracted a total of <strong>1.100+ organization and</strong> <strong>26.700+ researchers</strong> spread across <strong>90+ countries</strong> and <strong>700+ cities</strong>.</p> <p>Inside the zip file&nbsp;you will find:&nbsp;</p> <ul> <li>documents.csv: Full list of documents</li> <li>topics.csv: list of topics connected to documents</li> <li>terms.csv: list of terms connected to documents</li> <li>people.csv: list of people (authors) connected to documents</li> <li>orgsplus.csv: list of organisations (incl their locations) connected to documents</li> </ul> <p>More information is available at&nbsp;<a href="https://app.gitbook.com/@dataverz/s/coronavirus-r-and-d/">https://app.gitbook.com/@dataverz/s/coronavirus-r-and-d/</a>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
dryad36/100

Data from: Curation: heat stress responses and population genetics of the kelp Laminaria digitata (Phaeophyceae) across latitudes reveal differentiation among North Atlantic populations

<p>We aim to understand the thermal plasticity of a coastal foundation species across its latitudinal distribution by assessing physiological responses to high temperature stress in the kelp <i>Laminaria digitata</i> in combination with population genetic characteristics. We <a>hypothesize</a> that Arctic and cold-temperate populations are less heat resilient than warm-temperate populations. Using meristems of natural <i>L. digitata</i> populations from six locations ranging between Kongsfjorden, Spitsbergen (79°N), and Quiberon, France (47°N), we performed a common-garden heat stress experiment applying 15°C to 23°C over eight days. We assessed growth, quantum yield, carbon and nitrogen storage, and xanthophyll pigment contents as response traits. Population connectivity and genetic diversity were analysed with microsatellite markers to relate heat resilience to genetic features and phylogeography. Microsatellite genotyping revealed all sampled populations to be genetically distinct, underlying strong hierarchical structuring between and within southern and northern clades. Genetic diversity was lowest in the isolated population of the North Sea island of Helgoland and highest in Roscoff in the English Channel. Results from the heat stress experiment suggest that the upper temperature limit of <i>L. digitata</i> is nearly identical across its distribution range, but subtle differences we<a>re </a>revealed for the two populations currently at their warm limits. They respectively <a>show</a> a significant advantage in growth at 19°C and 21°C (Quiberon) and a lack of stress responses in photosynthetic quantum yield and xanthophyll pigments at 23°C (Helgoland). In addition, quantum yield indicated the highest heat sensitivity in <i>L. digitata</i> from the northernmost population in Spitsbergen. All together, these results support the hypothesis of moderate local differentiation across <i>L. digitata</i>'s European distribution, whereas effects are likely too weak to ameliorate the species' capacity to withstand ocean warming and marine heatwaves at the southern range edge.</p>

opencc-zeroAug 2020View details →
zenodo36/100

IPBES Data Management Tutorials - Session 3.3: Data management report details: Data and metadata documentation and curation

<p>The&nbsp;<em>IPBES data management tutorials</em>&nbsp;are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The&nbsp;<em>IPBES data management reports </em>chapter&nbsp;provides an overview and discussion of specific elements of IPBES data management reports.</p> <p>This session&nbsp;<em>Data management report details: Data and metadata documentation and curation&nbsp;</em>reviews what should be included in metadata and why it should be tracked in a data management report.</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates

DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.

opencc-zeroDec 2016View details →
zenodo36/100

Replication package for How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists

<p>This dataset was used in the paper: "How the R Community Creates and Curates Knowledge: An Extended Study of Stack Overflow and Mailing Lists", Journal of Empirical Software Engineering, to appear.</p>

opencc-by-4.0Apr 2017View details →
zenodo36/100

Curated dataset of 18 computational experiments-(E1-E18)

<h2><strong>Overview</strong></h2> <p>This dataset, <strong>"Multi-Domain Experiment Dataset for Evaluating Reproducibility Tools (E1&ndash;E18),"</strong> is curated to facilitate the evaluation and benchmarking of reproducibility frameworks. It provides a structured and diverse collection of scientific experiments, enabling researchers and developers to test and compare different tools designed for computational reproducibility.</p> <h2><strong>Dataset Composition</strong></h2> <p>The dataset consists of 18<strong> experiments (E1&ndash;E18)</strong> covering multiple scientific domains, including <strong>computer science, human-computer interaction (HCI), medicine, artificial intelligence, climate change, and economics</strong>. These experiments range from simple computational scripts to complex setups requiring integrated databases, multiple programming languages, and domain-specific computational environments.</p> <h2><strong>Experiment Sources</strong></h2> <p>To ensure a well-balanced dataset, experiments were sourced from <strong>peer-reviewed scientific conferences</strong> and <strong>open-access repositories</strong>:</p> <ul> <li> <p><strong>Computer Science:</strong></p> <ul> <li><strong>Software Engineering:</strong> Experiments from the <strong>IEEE/ACM International Conference on Software Engineering (ICSE 2022)</strong>, a premier venue for software engineering research.</li> <li><strong>Databases:</strong> Selected experiments from the <strong>International Conference on Very Large Databases (VLDB 2021)</strong>, a leading database conference.</li> <li><strong>Human-Computer Interaction (HCI) and User Studies:</strong> <ul> <li>Experiments sourced from the <strong>European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2023)</strong>.</li> <li>These experiments focus on user studies related to software engineering and usability research.</li> </ul> </li> </ul> </li> <li> <p><strong>Interdisciplinary Fields (Collected from Zenodo):</strong></p> <ul> <li><strong>Artificial Intelligence (AI):</strong> Experiments covering machine learning and AI-driven methodologies.</li> <li><strong>Climate Change:</strong> Computational experiments and simulations addressing environmental research.</li> <li><strong>Medicine:</strong> Experiments focusing on medical and healthcare applications.</li> <li><strong>Economics:</strong> Computational economic models and data analysis experiments.</li> </ul> </li> <li> <p><strong>Related Work on Reproducibility Tools:</strong></p> <ul> <li>We included <strong>experiments previously used to evaluate reproducibility tools</strong> and methodologies.</li> <li>This ensures alignment with prior research and enhances comparability across tools.</li> <li>Among these, two key studies&nbsp;<a href="https://ieeexplore.ieee.org/document/9041769"><strong>SciInc</strong></a> and <strong><a href="https://doi.ieeecomputersociety.org/10.1109/eScience.2017.51">SciUnit</a> </strong>provided three fully documented experiments: <ul> <li><strong>Chicago Food Inspections Evaluation;</strong></li> <li><strong>Variable Infiltration Capacity;</strong></li> <li><strong>Incremental Query Execution.</strong></li> </ul> </li> </ul> </li> </ul> <h2><strong>Methodology for Dataset Curation</strong></h2> <p>To construct this dataset, we employed a structured selection process:</p> <ol> <li> <p><strong>Scientific Conference Selection:</strong></p> <ul> <li>We identified key research areas within <strong>computer science</strong> and selected experiments from <strong>ICSE 2022</strong>, <strong>VLDB 2021</strong>, and <strong>ESEC/FSE 2023</strong> (focusing on user studies in HCI).</li> </ul> </li> <li> <p><strong>Zenodo Repository Search:</strong></p> <ul> <li>Targeted searches were conducted using the keywords <strong>"Medical," "Artificial Intelligence," "Climate Change,"</strong> and <strong>"Economics."</strong></li> <li>We filtered results to include only <strong>software repositories</strong>.</li> <li>From the <strong>top 100</strong> search results in each category, <strong>five experiments were randomly selected per domain</strong>.</li> </ul> </li> <li> <p><strong>Reproducibility Tools &amp; Related Work:</strong></p> <ul> <li>We incorporated experiments <strong>previously used to evaluate existing reproducibility tools</strong>.</li> <li>This selection ensures <strong>comparability and continuity</strong> with past reproducibility studies.</li> </ul> </li> </ol> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Gitome: A curated dataset for GitHub README-related tasks

<h2><strong>About&nbsp;</strong></h2><p>This repository contains the source code implementation used to replicate the experimental results obtained in the submitted to the 21st International Conference on Mining Software Repositories (MSR204).</p><p><i>"Gitome: A curated dataset for GitHub README-related tasks"</i></p><p>authored by:</p><p>Claudio Di Sipio, Juri Di Rocco, Riccardo Rubei, Phuong Than Nguyen, and Davide Di Ruscio,</p><p>Università degli Studi dell'Aquila, Italy</p><h2><strong>Data description&nbsp;</strong></h2><p>The dataset is structured as follows:&nbsp;</p><ul><li><strong>emf_metamodel.zip:</strong> It contains the Ecore project with the Gitome data model</li><li><strong>existing_dumps.zip</strong>: It contains the existing datasets used to build Gitome</li><li><strong>lang_aggr_stats.csv: </strong>It contains the language data to compute the statistics presented in the paper</li><li><strong>langs.csv: </strong>It contains all the languages and their frequency</li><li><strong>output_dataset.zip:</strong> It contains the benchmarking dataset obtained by parsing the README files</li><li><strong>repository_lists.zip: </strong>It contains the list of repositories for each considered dataset (with possible duplicates)</li><li><strong>topics.csv:</strong> It contains all the topics and their frequency</li><li><strong>topics_aggr_stats.csv: &nbsp;</strong>It contains the topics data to compute the statistics presented in the paper</li><li><strong>gitome_repo.txt</strong>: It contains the list of the URLs of the considered GitHub repositories</li></ul><p>&nbsp;</p><h2><strong>How to collect Gitome</strong></h2><p>To collect all the data stored in this archive, please refer to the supporting Github repository https://github.com/MDEGroup/Gitome-MSR2024.</p><p>&nbsp;</p><p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

PPARδ dataset curated and enriched using the Enalos tools and Enalos KNIME nodes for machine learning analysis (SCENARIOS project)

<p><span>A curated and enriched dataset for PPAR</span>&delta;<span>, suitable for in silico model development, was obtained from PubChem BioAssay under the numeric identifier AID 469785 using Enalos tools and Enalos KNIME nodes. This dataset comprises 136 compounds that induce luciferase activity, serving as an indicator of agonist activity against the human PPAR</span>&delta;<span> ligand-binding domain. These molecules were tested in a human embryonic kidney cell line (293T), co-transfected with a chimeric plasmid containing the yeast GAL4 DNA-binding domain (DBD). All 136 oxazole-based compounds retrieved from the dataset are accompanied by their half-maximal effective concentration (EC50) and enriched with 777 molecular descriptors extracted from their 2D structure using EnalosMold2 KNIME nodes</span></p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Curated catalogue for IRFundusSet (Integrated Retinal Fundus Set)

<p>Integrated Retinal Fundus Set (IRFundusSet) is a dataset that consolidates and curates several public datasets to ease their use as a harmonized unit and to&nbsp;<br>have a common definition of what is normal or healthy.&nbsp;</p> <p><strong>Motivation:</strong> Existing retinal fundus public datasets do not follow a standard format or structure for organizing and archiving the datasets. Moreover, eyes/images marked as `not pathological` might represent healthy eyes or eyes with another pathology that is not relevant to the task or disease the dataset was created for.&nbsp;</p> <p>This dataset entails&nbsp;</p> <ul> <li>A Python module that that consolidates and harmonizes the public datasets and that can be used in a&nbsp; data pipeline following the PyTorch style for Dataset objects.&nbsp;</li> <li>A curated catalogue in CSV format where images have been manually reviewed to determine if the image is of a healthy or pathological eye; a common definition for what is a healthy observation&nbsp;</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes

<p>This repository contains several files describing the data discussed in Lindner et al's "Advancing source tracking: systematic review and source-specific genome database curation of fecally shed prokaryotes".&nbsp;</p> <ol> <li>"database.fna" = concatenation of the (draft or complete) genome sequences described in the paper (n=12,730 source associated prokaryotic genomes) which passed quality checks and are species-level representatives (i.e., dereplicated at 95% ANI).</li> <li>"gdef.txt" = a manifest describing which sequences belong to which genomes.</li> <li>"sources.txt" = a manifest describing which genomes belong to which sources.</li> </ol> <p>The sources this database covers:</p> <table> <tbody> <tr> <td>Source Category</td> <td>Species-level <br>genome count</td> <td>Source-specific<br>species-level&nbsp;<br>genome count</td> </tr> <tr> <td>Bird</td> <td>56</td> <td>40</td> </tr> <tr> <td>Cat</td> <td>86</td> <td>6</td> </tr> <tr> <td>Chicken</td> <td>1314</td> <td>887</td> </tr> <tr> <td>Cow</td> <td>39</td> <td>15</td> </tr> <tr> <td>Dog</td> <td>139</td> <td>56</td> </tr> <tr> <td>Pig</td> <td>2764</td> <td>2035</td> </tr> <tr> <td>Ruminant</td> <td>740</td> <td>714</td> </tr> <tr> <td>Human</td> <td>4484</td> <td>3350</td> </tr> <tr> <td>Wastewater</td> <td>3108</td> <td>3097</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>See publication for further details.&nbsp;</p>

opencc-by-4.0Feb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record