Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,549

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,549 results for “benchmarks”

Learn how ShareScore rates datasets ↗
zenodo40/100

SQANTI-SIM: a simulator of controlled transcript novelty for lrRNA-seq benchmark

<p>In this repository, we present the PacBio and ONT simulated datasets used for benchmarking transcriptome reconstruction tools, as evaluated in the manuscript titled "<i>SQANTI-SIM: a simulator of controlled transcript novelty for lrRNA-seq benchmark</i>". The dataset includes simulated long reads, short reads, CAGE peaks, and a reduced reference annotation. Additionally, we have included reconstructed transcriptomes from each method, along with SQANTI3 output files. The SQANTI-SIM software can be accessed on GitHub at the following URL: <a href="https://github.com/ConesaLab/SQANTI-SIM">https://github.com/ConesaLab/SQANTI-SIM</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Stable Vowel Corpus (SVC): a benchmark for spectral stability analyses

<p>This dataset comprises 4,860 synthesized CVC-like sequences generated using Praat's articulatory synthesizer (Boersma, 1998; Boersma &amp; Weenink, 2021). The duration and location of the spectrally stable portion of the vowel within these CVC-like sequences were randomly varied.</p><p>Each corresponding .wav file is labeled with the following format:</p><ul><li><strong>Stimulus</strong> <i>ID_TsV_TeV_C1_V_C2_speaker.wav</i></li><li><strong>e.g.</strong> <i>1000_0.2197529077064245_0.2906336685046843p_a_d_female.wav</i></li></ul><p>Here's the breakdown of the naming convention:</p><ul><li><strong>Stimulus ID</strong>: A unique identifier for each stimulus.</li><li><strong>TsV</strong>: The timecode of the start of the stable portion of the vowel.</li><li><strong>TeV</strong>: The timecode of the end of the stable portion of the vowel.</li><li><strong>C1</strong>: The first consonant category.</li><li><strong>V</strong>: The vowel category.</li><li><strong>C2</strong>: The second consonant category.</li><li><strong>Speaker</strong>: The type of artificial speaker used for the synthesis.</li></ul><p>Further details about the design and synthesis of the corpus have been published in:</p><ul><li>Genette, J., Rivera Espejo, J. M., Gillis, S., &amp; Verhoeven, J. (2023). Determining spectral stability in vowels: A comparison and assessment of different metrics. <i>Speech Communication</i>, <i>154</i>, 102984.&nbsp;<a href="https://doi.org/10.1016/j.specom.2023.102984">https://doi.org/10.1016/j.specom.2023.102984</a></li></ul><h4><strong>References</strong></h4><p>Boersma, P. (1998). <i>Functional Phonology: Formalizing the Interactions between Articulatory and Perceptual Drives</i> [PhD Thesis]. University of Amsterdam.</p><p>Boersma, P., &amp; Weenink, D. (2021).<i> Praat: doing phonetics by computer</i> [Computer program. Version 6].</p>

opencc-by-nc-nd-4.0Dec 2022View details →
zenodo40/100

tFoodL: Larger Semantic Table Annotations Benchmark for Food Domain

<p><strong>tFoodL</strong> is the successor work of <a href="https://zenodo.org/records/10048187">tFood</a> that is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata.</p><p>Similar to tFood, it is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tFoodL</strong>&nbsp;contains 43,255 entity and horizontal tables, while this repository contains only the validation fold (10%) of the entire benchmark with its ground truth data (gt).&nbsp;</p><p>The supported tasks for semantic table annotations are:&nbsp;</p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>

opencc-by-4.0Dec 2023View details →
zenodo40/100

tBiodivL: Larger Semantic Table Annotations Benchmark for Biodiversity Domain

<p><strong>tBiodivL</strong> is a dataset for tabular data to knowledge graph matching. It is derived from the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiodivL</strong> is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283015">tBiodiv</a></p><p><strong>tBiodivL&nbsp;</strong>contains <strong>222,353</strong> entity and horizontal tables, while this repository contains only a sample of <strong>1% of the total generated tables</strong> of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>312 GB</strong>. We will update this repository with the full dataset in the Future.</p><p>Please get in touch if you are interested in the full dataset,&nbsp;</p><p>The supported tasks for semantic table annotations are:&nbsp;</p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>

opencc-by-4.0Dec 2023View details →
zenodo40/100

tBiomedL: Larger Semantic Table Annotations Benchmark for Biomedical Domain

<p><strong>tBiomedL </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiomedL&nbsp;</strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using five levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283103">tBiomed</a></p><p><strong>tBiomedL&nbsp;</strong>contains <strong>860,479</strong> entity and horizontal tables, while this repository contains only <strong>a sample of 1%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>27</strong> <strong>GB</strong>. We will update this repository with the full dataset, including the test fold with its ground truth data in the Future.</p><p>Please get in touch if you are interested in the full dataset,&nbsp;</p><p>The supported tasks for semantic table annotations are:&nbsp;</p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Globe230k: A Benchmark Dense-Pixel Annotation Dataset for Global Land Cover Mapping

<p>We (Intelligent Mining and Analysis of Remote Sensing big data, IMARS) create a large-scale annotated dataset (Globe230k) for land use/land cover (LULC) mapping, which is annotated on Google Earth image of 1 m spatial resolution. Globe230k is annotated by numerous experts and students major in survey and mapping after necessary training, through visual interpretation on very high-resolution images, as well as in-situ field survey, under the guidance of the organized annotation pipeline. Globe230k has three superiorities:</p> <p>1) Large scale: the Globe230k includes 232,819&nbsp;annotated images with the size of 512x512 and spatial resolution of 1 m, with more than 3x1010 annotated pixels,&nbsp;and&nbsp;it includes&nbsp;10 first-level categories.&nbsp;</p> <p>2) Rich diversity: the annotated images are sampled from worldwide regions, with coverage area of over 60,000 km2, indicating a high variability and diversity.&nbsp;Besides, in order to ensure the category balance, we intentionally give more chance to the rare categories to be sampled, such as wetland, ice/snow, etc.</p> <p>3) Multi-modal: Globe230k not only contains RGB bands, but also include other important features for Earth system research, such as Normalized differential vegetation index (NDVI), digital elevation model (DEM), vertical-vertical polarization (VV) bands, vertical-horizontal polarization (VH) bands, which can facilitate the multi-modal data fusion research. Due to the large size of the multi-modal dataset (DEM 1.91G, NDVI 164G, VVVH 372G), these dataset are stored on Baidu Yunpan, the download link is :https://pan.baidu.com/s/12AKbiqOXSf4fnm7mYkCE0g?pwd=230k, the extraction code is 230k.</p> <p>The image patches and their corresponding annotated patches are respectively stored in "image_patch.zip" and "label_patch.zip" file. The RGB image is in forms of ".jpg", with size of 512x512, the pixel value is ranged from 0-255. The annotated patches is in forms of ".png", also with size of 512x512, the pixel value is ranged from 1-10, which respectively represent 1#cropland, 2#forest, 3#grass, 4#shrubland, 5#wetland, 6#water, 7#tundra, 8#impervious, 9#bareland, 10#ice/snow. The corresponding DEM, NDVI and VVVH patches are all in form of ".tif", with size of 512x512 (due to the different resolution of DEM, NDVI and VVVH patches, they are all uniformly resized to the same scale as the image patch).&nbsp;</p> <p>The total 232,819 pairs are officially divided into training set, validation set, and test set, based on ratio of 7:1:2, which can be find in "train_num.txt","val_num.txt","test_num.txt" file. Based on this division, the official baseline accuracy of several state-of-the-art semantic segmentation can be found in the related arcticle (https://spj.science.org/doi/10.34133/remotesensing.0078).</p> <p>We hope it can&nbsp;be used as a benchmark to promote further development of global land cover mapping and semantic segmentation algorithm development.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Benchmarking dataset for multiskilled workforce planning with uncertain demand

<p>These datasets are related to the Data Article entitled: &ldquo;A benchmark dataset for the retail multiskilled personnel planning under uncertain demand&rdquo;, submitted to the Data Science Journal. This data article describes datasets from a home improvement retailer located in Santiago, Chile. The datasets were developed to solve a multiskilled personnel assignment problem (MPAP) under uncertain demand. Notably, these datasets were used in the published article "Multiskilled personnel assignment problem under uncertain demand: A benchmarking analysis" authored by Henao et al. (2022). Moreover, the datasets were also used in the published articles authored by Henao et al. (2016) and Henao et al. (2019) to solve MPAPs.</p> <p>The datasets include real and simulated data. Regarding the real dataset, it includes information about the store size, number of employees, employment-contract characteristics, mean value of weekly hours demand in each department, and cost parameters. Regarding the simulated datasets, they include information about the random parameter of weekly hours demand in each store department. The simulated data are presented in 18 text files classified by: (i) Sample type (in-sample or out-of-sample). (ii) Truncation-type method (zero-truncated or percentile-truncated). (iii) Coefficient of variation (5, 10, 20, 30, 40, 50%).</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

The IMITATOR benchmarks library 2.1: A benchmarks library for extended parametric timed automata

<p>We present here the IMITATOR benchmarks library 2.1: A benchmarks library for extended parametric timed automata</p> <p>&nbsp;</p> <p>We present two archives:</p> <p>- one (benchmarks.zip) with the models and the properties</p> <p>- one (full.zip) with the benchmarks and all the results: the expected results, generated PDF and graphics, and a whole standalone Web page (more or less equivalent to <a href="https://www.imitator.fr/static/library.html">www.imitator.fr/static/library.html</a>) summarizing all benchmarks</p> <p>&nbsp;</p> <p>See a full description in the TAP 2021 paper ("<a href="https://link.springer.com/10.1007/978-3-030-79379-1_3">A Benchmarks Library for Extended Parametric Timed Automata</a>")</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Validation and Benchmark Dataset for Discrete Element Method Simulations

<p>Verification and Benchmark Dataset for Discrete Element Method Simulations<br>v3 (05/02/2024)<br>Authors: Jose Salomon, Fernando Patino-Ramirez, Catherine O'Sullivan<br>https://doi.org/10.5281/zenodo.10160309<br>Contact: jjs19@ic.ac.uk<br>--------------------------------------------------------------------<br>Description of the repository:</p> <p>This repository contains a collection of datafiles and scripts that can be employed to validate and benchmark new or existing DEM codes.&nbsp;<br>Two validation cases/folders are considered "FCC_packing" and "Rolling_clump". The benchmark dataset is provided in the "Toyoura_sh" folder.<br>All datafiles and scripts are in the corresponding *.zip files. A detailed description of all cases can be found in the related article.</p> <p>In each of these folders, two sub-folders can be found: (1)"Data" and (2)"Scripts". These folders contain:</p> <p>1)"Data": contains the datafiles to perform the validation or benchmark. Two types of data/folders can be found here: "Raw" and "Filtered".<br>The "Raw" folder contains raw data only. The "Filtered" data contains the post-processed data employed to generate the plots found in the related article.<br>Plots in the related article can be reproduced by using the MATLAB files found in the corresponding data folder.</p> <p>2)"Scripts": contains the LAMMPS scripts used to generate the data files contained in "Data".<br>Indications about how to run these scripts can be found in the "README.txt" file in each folder.</p> <p>In order to reproduce the simulations of this repository, LAMMPS must be built including the "GRANULAR" and "RIGID" packages. Please check the README.txt files in each folder for details.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Scout Benchmark Scenarios for U.S. Building Energy and CO2 Emissions to 2050

<p><strong>Overview and Intended Use Cases</strong></p> <p>These scenarios establish a range of futures for U.S. buildings sector energy use and CO<sub>2</sub> emissions to 2050 using <a href="https://scout-bto.readthedocs.io/en/latest/">Scout</a>, a reproducible and granular model of U.S. building energy use, emissions, and consumer costs developed by the U.S. national labs for the U.S. Department of Energy's Building Technologies Office (BTO).</p> <p>Scout benchmark scenario data are suitable for the following example use cases:</p> <ul> <li>Setting high-level policy goals for U.S. buildings sector energy use, electricity demand, and CO<sub>2</sub> emissions over both the near- and long-term (e.g., X% building CO<sub>2</sub> emissions reductions vs. 2005 levels by 2030, Y% reductions vs. 2005 levels by 2050);</li> <li>Exploring the effects of key deployment dynamics driving U.S. buildings sector energy and CO<sub>2</sub> emissions to 2050 that could be affected by policy levers (e.g., raising minimum technology performance levels; improving market penetration of commercially available technologies; accelerating electrification and/or retrofit rates; introducing breakthrough technologies to the market);</li> <li>Determining priority segments (regions, building types, and end use/technology types) and sequencing of U.S. buildings sector energy and CO<sub>2</sub> emissions reductions and/or changes in total consumption by fuel type to 2050 under a given set of assumptions;</li> <li>Identifying the energy and CO<sub>2</sub> impacts or cost effectiveness of specific technologies or operational approaches of interest&mdash;in isolation or after considering competition with other measures in a scenario portfolio; and/or</li> <li>Exploring the total cost of deploying different portfolios of building energy efficiency and end-use electrification measures, as well as the total consumer energy cost savings potential of those portfolios.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</li> </ul> <p><strong>Scenario Summary</strong></p> <p>A total of 5 scenarios explore total building energy use, CO<sub>2</sub> emissions, and technology and energy costs from 2024&ndash;2050 under varying levels of demand-side deployment of building efficiency and electrification measures and parallel decarbonization of buildings&rsquo; electricity supply. Narrative descriptions of these scenarios are as follows:</p> <ul> <li><strong>Stated Policies: </strong>Existing policies and regulations (mainly IRA for buildings) lead to modestly accelerated deployment of HPs/HPWHs but not other efficiency measures in the buildings sector. The power sector decarbonizes consistent with a &ldquo;<a href="https://www.nrel.gov/docs/fy23osti/84916.pdf">Mid-case (with tax credit phaseout)</a>&rdquo; scenario.</li> <li><strong>Mid: </strong>Policy makers rely mostly on market-based instruments to moderately increase deployment of efficient technology and fuel switching to heat pumps. The power sector decarbonizes consistent with a &ldquo;Mid-case with 95% Decarbonization by 2050 (without tax credit phaseout)&rdquo; scenario.</li> <li><strong>High: </strong>Policy makers use both regulations and market-based instruments to dramatically accelerate deployment of high efficiency technologies and fuel switching to heat pumps, though building technologies with breakthrough increases in performance at low cost do not materialize on the market. The power sector decarbonizes consistent with a &ldquo;<a href="https://www.nrel.gov/docs/fy23osti/84916.pdf">Mid-case with 100% Decarbonization by 2035 (without tax credit phaseout)</a>&rdquo; scenario.</li> <li><strong>Breakthrough: </strong>Research and innovation breakthroughs lead to market availability of cost-effective, high-performance building technologies by 2030; these, coupled with accelerated deployment of high efficiency technologies and fuel switching to heat pumps, lead to aggressive buildings sector transformation. The power sector decarbonizes consistent with a &ldquo;<a href="https://www.nrel.gov/docs/fy23osti/84916.pdf">Mid-case with 100% Decarbonization by 2035 (without tax credit phaseout)</a>&rdquo; scenario.</li> <li><strong>Inefficient Electrification Sensitivity:&nbsp;</strong>Policy makers use regulations and market-based instruments to encourage fuel switching but do not include provisions that require switching to efficient heat pumps, resulting in a substantial amount of switching to inefficient electric resistance heating and water heating technologies. The power sector decarbonizes consistent with a &ldquo;<a href="https://www.nrel.gov/docs/fy23osti/84916.pdf">Mid-case (with tax credit phaseout)</a>&rdquo; scenario.</li> </ul> <p>The key input dimensions that are varied to produce the above range of scenarios are as follows:</p> <ul> <li><u>Market-available technology performance range:</u> the energy performance levels of building technologies available for purchase by end use consumers, bounded by a minimum performance &ldquo;floor&rdquo; and maximum performance &ldquo;ceiling&rdquo;;</li> <li><u>Load electrification rate and efficiency:</u> the rate at which fossil-fired equipment is converted to electric service, and the efficiency level of the electric equipment; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li><u>Early retrofits:</u> the fraction of consumers that choose to replace existing building equipment and/or envelope components before the end of their useful lifetimes; and</li> <li><u>Power grid decarbonization:</u> the annual average CO<sub>2</sub> emissions intensity of the electricity supplied to the buildings sector across the modeled time horizon (2024&ndash;2050), resolved by grid region.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> </ul> <p>Refer to the attached &ldquo;Scenario_Guide" PDF for further scenario details and results; instructions for reproducing scenario results are available in &ldquo;Scenario_Execution&rdquo; XLSX.</p> <p>Results data are reported as an annual time series (2024&ndash;2050) at both a national and regional (<a href="https://www.eia.gov/outlooks/aeo/pdf/nerc_map.pdf">EMM grid region</a>) spatial resolution. While not reflected in this dataset, annual time series data may be further translated to a sub-annual, hourly resolution for integration with grid modeling&mdash;please contact the authors for more information.</p> <p><strong>What's New in This Version</strong></p> <p><strong><em>Note: v6.1 updates the file ./Results/Results_Summary.xlsx to reflect the latest scenario runs. Please disregard the outdated version of this file that was posted in v6.</em></strong></p> <p>This set of benchmark scenarios provides an update to <a href="../records/8087519">Version 5</a> of the Scout Benchmark Scenarios (June 2023) using the same scenario definitions but an updated set of baseline and measure input data alongside several minor methodological changes.&nbsp;&nbsp;&nbsp;&nbsp;</p> <p>The following scenario features are new in this dataset:</p> <ul> <li>Reference case data and energy use projections updated to <a href="https://www.eia.gov/outlooks/aeo/">AEO 2023</a>, including updates to energy and stock and technology cost, performance, and lifetime data; updated site-source energy conversions, CO2 emissions intensities, and energy prices; and revised peak and take period definitions that are consistent with 2023 EMM projections.</li> <li>Integration of federal and state cost incentives from AEO 2023 (see <a href="https://www.eia.gov/outlooks/aeo/IIF_IRA/pdf/IRA_IIF.pdf">AEO2023 Issues in Focus: Inflation Reduction Act Cases</a> in the AEO2023 for details); these incentives reduce the initial cost of upgrades for applicable measures.</li> <li>Revised method for allocating end use electricity baselines in AEO from census divisions to EMM regions and states by using <a href="https://www.nrel.gov/buildings/end-use-load-profiles.html">End Use Load Profiles</a> (EULP) data. EULP data now also underpin updated, EMM-resolved hourly load baseline shapes.</li> <li>Retail price projections for grid scenarios are updated to match those produced by NREL under the Department of Energy&rsquo;s DECARB Initiative (these are similar to but differ in slight ways from NREL&rsquo;s <a href="https://www.nrel.gov/analysis/standard-scenarios.html">Standard Scenarios</a>). Three scenarios are included:&nbsp; <ul> <li><em>Stated Policies</em>: includes moderate estimates for inputs such as technology costs, fuel prices, and demand growth with no nascent technologies and electric sector policies that match current federal laws and regulations (including IRA &amp; BIL); achieves an 88% reduction in building site electricity emissions <em>intensity</em> (Mt CO2/quad site) from 2005 levels by 2050.</li> <li><em>Mid</em>: consistent with<em> Stated Policies</em> except achieves 97% reduction in building site electricity emissions intensity from 2005 levels by 2050.</li> <li><em>High:</em> includes low demand growth projections with advanced inputs for technology costs and allowance of transmission expansion between regions (without limitations based on historical build rates); federal policies are consistent with implemented laws (including IRA &amp; BIL); building electricity is fully decarbonized after 2035.</li> <li>The previous version of the benchmark datasets used retail price data from EIA&rsquo;s&nbsp;<a href="https://www.eia.gov/outlooks/aeo/">Annual Energy Outlook</a> scenarios.</li> </ul> </li> <li>In contrast to <a href="https://doi.org/10.5281/zenodo.8087519">Version 5</a>, measures in the &ldquo;best available&rdquo; measure tier are not deployed with load flexibility features.&nbsp;</li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo40/100

PheKnowLator Human Disease Knowledge Graph Benchmarks Archive

<h2><strong>PKT Human Disease KG Benchmark Builds</strong></h2> <p>The PheKnowLator (PKT) Human Disease KG (PKT-KG) was built to model mechanisms of human disease, which includes the Central Dogma and represents multiple biological scales of organization including molecular, cellular, tissue, and organ. The knowledge representation was designed in collaboration with a PhD-level molecular biologist (<a href="https://user-images.githubusercontent.com/8030363/195469903-86598760-40b7-4126-857c-3d6368305a86.png">Figure</a>).&nbsp;</p> <p>The <strong>PKT Human Disease KG</strong> was constructed using 12 OBO Foundry ontologies, 31 Linked Open Data sets, and results from two large-scale experiments (<a href="https://doi.org/10.48550/arXiv.2307.05727">Supplementary Material</a>). The 12 OBO Foundry ontologies were selected to represent chemicals and vaccines (i.e., ChEBI and Vaccine Ontology), cells and cell lines (i.e., Cell Ontology, Cell Line Ontology), gene/gene product attributes (i.e., Gene Ontology), phenotypes and diseases (i.e., Human Phenotype Ontology, Mondo Disease Ontology), proteins, including complexes and isoforms (i.e., Protein Ontology), pathways (i.e., Pathway Ontology), types and attributes of biological sequences (i.e., Sequence Ontology), and anatomical entities (Uberon ontology). The RO&nbsp;is used to provide relationships between the core OBO Foundry ontologies and database entities.</p> <p>The <strong>PKT Human Disease KG</strong> contained 18 node types and 33 edge types. Note that the number of nodes and edge types reflects those that are explicitly added to the core set of OBO Foundry ontologies and does not take into account the node and edge types provided by the ontologies. These nodes and edge types were used to construct 12 different PKT Human Disease benchmark KGs by altering the Knowledge Model (i.e., class- vs. instance-based), Relation Strategy (i.e., standard vs. inverse relations), and Semantic Abstraction (i.e., OWL-NETS (yes/no) with and without Knowledge Model harmonization [OWL-NETS Only vs. OWL-NETS + Harmonization]) parameters. Benchmarks within the PheKnowLator ecosystem are different versions of a KG that can be built under alternative knowledge models, relation strategies, and with or without semantic abstraction. They provide users with the ability to evaluate different modeling decisions (based on the prior mentioned parameters) and to examine the impact of these decisions on different downstream tasks.</p> <p>The Figures and Tables explaining attributes in the builds can be found <a href="https://github.com/callahantiff/PheKnowLator/wiki/Archived-Builds">here</a>.</p> <p>&nbsp;</p> <h3><strong>Build Data Access</strong></h3> <h4><strong>Important Build Information</strong></h4> <p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed this Zenodo-based archive for the builds. While the original GCP resources contained all of the resources needed to generate the builds, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the logs associated with each build.</p> <p>🗂 For additional information on the KG file types please see the following <a href="https://github.com/callahantiff/PheKnowLator/wiki/KG-Construction#table-knowledge-graph-build-output">Wiki page</a>, which is also available as a download from this repository (PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx).&nbsp;</p> <h4><strong>v1.0.0</strong></h4> <ul> <li>KGs:&nbsp;<a href="../doi/10.5281/zenodo.7030200">https://zenodo.org/doi/10.5281/zenodo.7030200</a></li> <li>Embeddings:&nbsp;<a href="../doi/10.5281/zenodo.7030188">https://zenodo.org/doi/10.5281/zenodo.7030188</a></li> </ul> <h4><strong>All Other Build Versions</strong></h4> <p><strong>Class-based Builds</strong></p> <p><em>Standard Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029957">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180239">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180539">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180774">MAY2021</a>;<a href="../doi/10.5281/zenodo.8180825"> JUN2021</a>; <a href="../doi/10.5281/zenodo.8180972">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183987">AUG2021</a>;<a href="../doi/10.5281/zenodo.8184090"> SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184131">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184205">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029953">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180255">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180545">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180772">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180827">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180974">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183989">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184088">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184133">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184208">NOV2021</a></li> </ul> </li> </ul> <p><em>Inverse Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029893">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180269">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180550">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180766">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180829">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180976">JUL2021</a>;<a href="../doi/10.5281/zenodo.8183991"> AUG2021</a>; <a href="../doi/10.5281/zenodo.8184086">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184135">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184210">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029921">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180279">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180555">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180768">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180833">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180982">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183993">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184084">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184137">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184212">NOV2021</a></li> </ul> </li> </ul> <p><strong>Instance-based Builds</strong></p> <p><em>Standard Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029941">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180333">JAN2021</a>;<a href="../doi/10.5281/zenodo.8180558"> FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180764">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180835">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180984">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183995">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184082">SEP2021&nbsp;</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184139">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184216">NOV2021&nbsp;</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029939">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180335">JAN2021</a>;<a href="../doi/10.5281/zenodo.8180564"> FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180762">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180837">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180986">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183997">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184080">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184141">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184218">NOV2021</a></li> </ul> </li> </ul> <p><em>Inverse Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029945">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180338">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180588">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180758">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180878">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180992">JUL2021</a>; <a href="../doi/10.5281/zenodo.8184001">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184078">SEP2021&nbsp;</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184143">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184220">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029919">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180340">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180584">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180756">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180823">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180996">JUL2021</a>; <a href="../doi/10.5281/zenodo.8184003">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184076">SEP2021&nbsp;</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184145">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184222">NOV2021</a></li> </ul> </li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

CHAMMI: A benchmark for channel-adaptive models in microscopy imaging

<p>We present a cellular microscopic image dataset for investigating channel-adaptive models. We collected and pre-processed images from three publicly available sources: 1) the WTC-11 hiPSC dataset from the Allen Institute (Viana et al., 2023), 2) the Human Protein Atlas dataset (Thul et al., 2017), and 3) a combined Cell Painting dataset from the Broad Institute (Gustafsdottir et al., 2013; Bray et al., 2017; Way et al., 2021). These images contain 3, 4, or 5 channels with different cellular structures highlighted in each channel. The goal of this dataset is to facilitate the creation and evaluation of novel computer vision models that are invariant to channel numbers.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Benchmark for energy efficient obstacle detection on head mounted wearable for the vision impaired

<p>Here we present a novel benchmark dataset with the associated challenge, that is to detect obstacles based on head-mounted sensors and lightweight wearable devices to assist Blind and Visually Impaired individuals (BVIs) navigate in indoor environments. &nbsp;The challenge encompasses three objectives: (1) as accurately as possible to detect the obstacles on the pathway that likely lead to a collision; (2) as durably as possible on a given amount of battery power for the detection algorithm or model to run; (3) as reliably as possible to compensate natural head turns so nearby objects would not trigger false alarms. &nbsp;The data provided in the benchmark are collected from the following head mounted sensors: (i) nine low-cost ultrasonic sensors; (ii) one high-end ultrasonic sensor with a larger detection range but higher power consumption; (iii) a 9-Degrees of Freedom (DOF) Inertial Measurement Unit (IMU). &nbsp;The resulting dataset consists of more than 188,000 unique sequences obtained from multiple subjects walking in three different indoor scenarios. &nbsp;This benchmark is to facilitate and encourage accurate yet fast obstacle detection solutions that can really benefit BVIs. &nbsp;</p>

openmit-licenseJun 2023View details →
zenodo40/100

Benchmark-Dataset FAN-01: Low pressure Axial Fan in a short Duct

<p>The case consists of a generic axial fan for industrial applications. Provided measurement data include instationary pressure probes in the rotor's tip gap, distributions of velocity and turbulent kinetic energy gained by laser Doppler anemometry, as well as acoustic results gained by microphones and beamforming.</p> <p>A detailed description of the dataset with references can be found in the PDF-File. The rotor geometry is available as IGS or Parasolid file. The measurement data is available, including the ones (LDA-data, pressure probes, acoustic microphones, &uuml;erformance) listed in the PDF description file.</p> <p><strong>Citation of the fan and the data:</strong></p> <p>Zenger, Florian, et al. <em>A benchmark case for aerodynamics and aeroacoustics of a low pressure axial fan</em>. No. 2016-01-1805. SAE Technical Paper, 2016.</p> <p><strong>Citation of the microphone array measurements:</strong></p> <p>Kr&ouml;mer, Florian J. <em>Sound emission of low-pressure axial fans under distorted inflow conditions</em>. FAU University Press, 2018.</p> <p><strong>Citation of the python scripts:</strong></p> <p>Junger, Clemens. <em>Computational aeroacoustics for the characterization of noise sources in rotating systems</em>. Diss. Technische Universit&auml;t Wien, 2019.</p> <p><strong>Related work and existing publications:</strong></p> <p>Schoder, Stefan, Clemens Junger, and Manfred Kaltenbacher. "Computational aeroacoustics of the EAA benchmark case of an axial fan."&nbsp;<em>Acta Acustica</em>&nbsp;4.5 (2020): 22.&nbsp;<a href="https://doi.org/10.1051/aacus/2020021">https://doi.org/10.1051/aacus/2020021</a></p> <p>Schoder, Stefan, and Felix Czwielong. "Dataset fan-01: Revisiting the EAA benchmark for a low-pressure axial fan."&nbsp;<em>arXiv preprint arXiv:2211.12014</em>&nbsp;(2022).&nbsp;<a href="https://doi.org/10.48550/arXiv.2211.12014">https://doi.org/10.48550/arXiv.2211.12014</a></p> <p>Kaltenbacher, Manfred, and Stefan Schoder. "EAA Benchmark for an axial fan."&nbsp;<em>e-Forum Acusticum 2020</em>. 2020.&nbsp;<a href="https://hal.science/hal-03221387/document">https://hal.science/hal-03221387/document</a></p> <p>Tieghi, Lorenzo, et al. "Machine-learning clustering methods applied to detection of noise sources in low-speed axial fan."&nbsp;<em>Journal of Engineering for Gas Turbines and Power</em>&nbsp;145.3 (2023): 031020.&nbsp;<a href="https://doi.org/10.1115/1.4055417">https://doi.org/10.1115/1.4055417</a></p> <p>Antoniou, E., Romani, G., Jantzen, A., Czwielong, F., &amp; Schoder, S. (2023). Numerical flow noise simulation of an axial fan with a Lattice-Boltzmann solver. <em>Acta Acustica</em>, <em>7</em>, 65. <a href="https://doi.org/10.1051/aacus/2023060">https://doi.org/10.1051/aacus/2023060</a></p> <p><strong>Data curation and Questions about the Dataset</strong></p> <p>Data curated by Stefan Schoder, any questions related to the dataset to stefan.schoder@tugraz.at.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Fastq files for benchmarking somatic variant calling pipelines

<p>The <a href="https://download.imgag.de/public/validation_dataset_somatic/readme.html" target="_blank" rel="noopener">original .bam files</a> were provided by <a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/mitarbeiter/profil/3377" target="_blank" rel="noopener">Marc Sturm</a> from the&nbsp;<a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/einrichtungen/institute/medizinische-genetik-und-angewandte-genomik" target="_blank" rel="noopener">Institut f&uuml;r Medizinische Genetik und Angewandte Genomik at the University Clinic T&uuml;bingen.</a></p> <p>The files were transformed into .fq.gz files using the <a href="https://github.com/nf-core/bamtofastq" target="_blank" rel="noopener">nf-core/bamtofastq pipeline.</a> All information on the pipeline run can be found in the <a href="../api/records/10805134/draft/files/execution_report_2024-03-11_11-44-17.html/content" target="_blank" rel="noopener noreferrer">execution_report_2024-03-11_11-44-17.html</a>. The FASTQ files for the normal sample were directly uploaded to <a href="https://osf.io/cduyq/files/onedrive">the corresponding osf project</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

CD crossdock benchmark set for DiffBindFR

<p>The CD benchmark set represents a comprehensive, large-scale benchmark dataset specifically designed for cross-docking evaluations in DiffBindFR paper. It is uniquely equipped to handle a variety of cross-docking scenarios, including Apo-Holo, Holo-Holo, and AF2-Holo cross-docking. The dataset, as released, encompasses several subsets: Ensemble (including three targets: CDK2, EGFR, and FXA), ApoRef, CASF2016, GPCR-AF2, and DUDE27-HoloEns. Each cross-docking system within these subsets is comprehensively equipped with a well prepared receptor PDB file, ligand SDF file, and Mol2 file. Every system has been assigned a unique ID with the format of [Lig ID][Alt ID]_[Apo PDB ID]_[Cross-Docking ID], and Cross-Docking ID can help users identify cross-docking scenario.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

MatSeg DataSet and Benchmark For Zero-Shot Material States Segmentation From images

<h2>This is an old version for the new version see&nbsp;<a href="../records/11331618">https://zenodo.org/records/11331618</a></h2> <p>&nbsp;</p> <p>A Dataset and Benchmark for zero-shot segmentation of materials states described in: &ldquo;Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data&rdquo; Described in&nbsp;<strong><a href="https://arxiv.org/pdf/2403.03309.pdf">https://arxiv.org/pdf/2403.03309.pdf</a>&nbsp;</strong></p> <p>See ReadMe in the zip file for technical details.</p> <p>&nbsp;</p> <h2><strong>MatSeg Benchmark&nbsp;</strong></h2> <p>A benchmark for zero-shot material state segmentation. The benchmark contains 820 real-world images with a wide range of material states and settings. For example: food states (cooked/burned..), plants (infected/dry.), to rocks/soil (minerals/sediment),&nbsp; construction/metals (rusted, worn),&nbsp; liquids&nbsp; (foam/sediment), and many other states in a class-agnostic manner.&nbsp; The goal is to evaluate the segmentation of material materials without knowledge or pretraining on the material or setting. The focus is on materials with complex scattered boundaries, and gradual transition&nbsp; (like the level of wetness of the surface). The annotation of the benchmark is point-based and similarity-based. Hence, for each image, we select several points and regions (Figure 4). We group the points of the same materials into the same label, we also define a group of points that have partial similarity. For example points in group A are more similar to points in group B than to points in group C (In case materials A and B are similar to each other but not identical). This approach allows us to capture the complexity of gradual transition and partial similarities in the world. While also enabling dealing with complex scattered and blurry shapes without needing to annotate the full shape which in many cases is unclear or very hard.</p> <p>Files <a href="../api/records/10801191/draft/files/MatSegBenchmarkPart1of3.zip/content" target="_blank" rel="noopener noreferrer">MatSegBenchmark</a>*.zip</p> <h2><strong>MatSeg synthetic Dataset Samples&nbsp;</strong></h2> <p>Synthethic dataset of images of materials spread on object surfaces and their segmentation map.</p> <p>The synthetic dataset is a very big, sample of the dataset as been uploaded.</p> <p>Files:&nbsp; &nbsp; &nbsp; &nbsp;MatSegSynthehticDataSample*.zip</p> <p>The full dataset can be found in this URLS:</p> <p><a href="https://e.pcloud.link/publink/show?code=kZHCcnZOfzqInb3anSl7xzFBoqCDmkr2JKV">https://e.pcloud.link/publink/show?code=kZHCcnZOfzqInb3anSl7xzFBoqCDmkr2JKV</a></p> <p><a href="https://icedrive.net/s/SBb3g9WzQ5wZuxX9892Z3R4bW8jw">https://icedrive.net/s/SBb3g9WzQ5wZuxX9892Z3R4bW8jw</a></p> <p>&nbsp;</p> <p>Generation Script for the synthetic data:</p> <p><a href="https://github.com/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script">https://github.com/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script</a></p> <p><a href="../records/10822596/files/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script-3.zip?download=1">https://zenodo.org/records/10822596</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-zeroMar 2024View details →
zenodo40/100

SIMpat: a synthetic benchmark for similarity metrics on patient representations

<div> <h2>Introduction</h2> <p>We used Synthea to generate six cohorts of patients with certain specified disease. Please refer to Synthea documentation for the generation process.</p> <p>We selected 6 different diseases that could be generated by Synthea, that were deemed by a medical professional as &ldquo;different enough&rdquo;. The goal of this simulation is to find a metric that can differentiate between patients.</p> <p>The six conditions are:</p> <ul> <li>Cerebral Palsy (SNOMED-CT code : 128188000 - Cerebral palsy (disorder))</li> <li>Colorectal Cancer (SNOMED-CT code : 93761005 - Primary malignant neoplasm of colon (disorder))</li> <li>Dialisys (SNOMED-CT code : 265764009 - Renal dialysis (procedure))</li> <li>Hypertension (SNOMED-CT code : 59621000 - Essential hypertension (disorder))</li> <li>Breast Cancer (SNOMED-CT code : 254837009 - Malignant neoplasm of breast (disorder))</li> <li>Prostate Cancer (SNOMED-CT code : 126906006 - Neoplasm of prostate (disorder))</li> </ul> <p>NB:</p> <ul> <li>Dialisys is not a disorder, but a condition, but is used here as a proxy for renal issue</li> <li>Synthea doesn&rsquo;t have a module to generate prostate cancer in men, but only prostate cancer in veteran, hence this is the module used here (all men with prostate cancer are veterans)</li> </ul> <p>We use those cohort to compare the ability of 12 different distance metrics to separate patients.</p> <p>Those 12 metrics are split in three groups :</p> <p>Sementic based metrics:</p> <ul> <li>AvgEmb* method encodes text by averaging the pre-trained word embeddings of all the words present in it.</li> <li>BERT* uses bidirectional transformer based neural model to solve the task of masked language modeling.</li> <li>Universal Sentence Encoders (USE)* use transformer based encoders to encode sentences into embedding vectors.</li> <li>Embeddings from Language Models (ELMo)* uses bi-directional LSTM based encoders to encode a sentence into a fixed size representation</li> </ul> <p>Graph based metrics:</p> <ul> <li>DeepWalk* uses random walks to generate sequences of vertices (vertex sentences) which are subsequently fed to a skip-gram model to learn the embeddings corresponding to the vertices.</li> <li>Node2Vec* uses biased random walks to optimize a neighborhood preserving objective function such that the nodes which are highly interconnected and the nodes with similar roles in the graph are closer in the embedding space.</li> <li>LINE* tries to directly optimize the vertex embeddings based on one hop and two hop random walk probabilities.</li> <li>HARP* proposes a meta-strategy for embedding vertices of a graph such that they preserve the higher-order structural features.</li> <li>Bags of findings^</li> <li>Average Links^</li> <li>Average Links Weighted by Information Content (IC)^</li> <li>Path Distance weighted by IC^</li> </ul> <p>Concept followed by a * are extracted from&nbsp;<a href="https://proceedings.mlr.press/v116/pattisapu20a/pattisapu20a.pdf">this paper</a>&nbsp;and can be downloaded&nbsp;<a href="../records/3842143">here</a></p> <p>Concept followed by a ^ were develloped by Jean-Virgile Voegeli (SIMED)</p> </div> <div> <div>&nbsp;</div> <h2>Descriptive analysis of the sample</h2> <p>We will first look at the cohort that were created by Synthea. The cohorts were created using the seed 123456789 for reproducibility.</p> <p>For this first experiment, Synthea was asked to generate 100 alive individuals for each specific disease. We asked Synthea to keep only 10 years of history. Each individual was set to be between the age of 18 and 80 years old. Except for specific sex-disease such as breast cancer and prostate cancer, all cohorts contains both male and female individuals. We used the default location, which is Massachussetts.</p> <p>One important note on age. The Synthea modules sometimes specify a minimum age to onset a certain condition / disease. For example, colorectal cancer can only onset after 50 years old, and prostate cancer after 60 years old.</p> <p>Each Synthea run was set to run 10.000 times. If after 10.000 tries, the software didn&rsquo;t manage to generate a patient that fit the criterion (here, a specific snomed code), the run would fail. Synthea can also generate patients that dies before the &ldquo;run date&rdquo;, and if this happens will simulate another patient.</p> <p>This explains why we have cohorts of more than 100 individuals but less than 100 alive individuals. We can also have in certain cases a little above 100 individuals. This is due to the fact that the synthea generator is multicore, and patients are generated simultaneously.</p> </div>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Zebra finch dataset for the paper: Benchmarking nearest neighbor retrieval of zebra finch vocalizations across development

<p>This is the dataset created in the paper "Benchmarking nearest neighbor retrieval of zebra finch vocalizations across development".</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

QH9: A Quantum Hamiltonian Prediction Benchmark for QM9 Molecules

<p>This is the official QH9 datasets from paper 'QH9: A Quantum Hamiltonian Prediction Benchmark for QM9 Molecules'.&nbsp; QH9 is a new&nbsp; Quantum Hamiltonian dataset providing precise Hamiltonian matrices for 130,831 stable molecular geometries, based on the QM9 dataset. Here is the QH9Stable dataset which is used in QH-Stable-iid and QH-Stable-ood.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record