Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,075

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,075 results for “ml”

Learn how ShareScore rates datasets ↗
zenodo40/100

Global ML-ready dataset for mining areas in satellite images

<p>This dataset is a global resource for machine learning applications in mining area detection and semantic segmentation on satellite imagery. It contains Sentinel-2 satellite images and corresponding mining area masks + bounding boxes for 1,210 sites worldwide. Ground-truth masks are derived from&nbsp;<a href="https://doi.org/10.1594/PANGAEA.942325" target="_blank" rel="noopener">Maus et al. (2022)</a> and <a href="https://doi.org/10.5281/zenodo.6806817" target="_blank" rel="noopener">Tang et al. (2023)</a>, and validated through manual verification to ensure accurate alignment with Sentinel-2 imagery from specific timestamps.&nbsp;</p> <p>The dataset includes three mask variants:</p> <ul> <li>Masks exclusively from Maus et al. (n=1,090)</li> <li>Masks exclusively from Tang et al. (n=817)</li> <li>A preferred mask selected from either Maus or Tang based on alignment quality determined during manual review (n=1,210).</li> </ul> <p>Each tile corresponds to a 2048x2048 pixel Sentinel-2 image, with metadata on mine type (surface, placer, underground, brine &amp; evaporation) and scale (artisanal, industrial). For convenience, the preferred mask dataset is already split into training (75%), validation (15%), and test (10%) sets.&nbsp;</p> <p>Furthermore, dataset quality was validated by re-validating test set tiles manually and correcting any mismatches between mining polygons and visually observed true mining area in the images, resulting in the following estimated quality metrics:&nbsp;</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Combined</td> <td>Maus</td> <td>Tang</td> </tr> <tr> <td>Accuracy</td> <td>99.78</td> <td>99.74</td> <td>99.83</td> </tr> <tr> <td>Precision</td> <td>99.22</td> <td>99.20</td> <td>99.24</td> </tr> <tr> <td>Recall</td> <td>95.71</td> <td>96.34</td> <td>95.10</td> </tr> </tbody> </table> <p>Note that the dataset does not contain the Sentinel-2 images themselves but contains a reference to specific Sentinel-2 images. Thus, for any ML applications, the images must be persisted first. For example, Sentinel-2 imagery is available from Microsoft's Planetary Computer and filterable via STAC API: <a href="https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a" target="_blank" rel="noopener">https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a</a>. Additionally, the temporal specificity of the data allows integration with other imagery sources from the indicated timestamp, such as Landsat or other high-resolution imagery.</p> <p>Source code used to generate this dataset and to use it for ML model training is available at&nbsp;<a href="https://github.com/SimonJasansky/mine-segmentation" target="_blank" rel="noopener">https://github.com/SimonJasansky/mine-segmentation</a>. It includes useful Python scripts, e.g. to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/05_persist_pixels_masks.py">download Sentinel-2 images via STAC API</a>, or to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/06_make_chips.py">divide tile images (2048x2048px) into smaller chips (e.g. 512x512px)</a>.&nbsp;</p> <p>A database schema, a schematic depiction of the dataset generation process, and a map of the global distribution of tiles are provided in the accompanying images.&nbsp;</p>

opencc-by-sa-4.0Nov 2024View details →
zenodo40/100

Hubble eXtreme Deep Field that has been clipped and processed for ML applications

<p>This is a set of FITS files downloaded from https://archive.stsci.edu/ .</p> <p>&nbsp;</p> <p>The *.npy file has been clipped to the 99.99th percentile and then minmax normalised along each channel.</p>

openmit-licenseFeb 2022View details →
zenodo40/100

Training and Testing Data, Associated Code, and WRF Code for ML-based nonhydrostatic alternative scheme in dynamical core of atmosphere

<p>Data and codes for a nonhydrostatic alternative scheme (NAS) in dynamical core of atmosphere based on machine learning.</p> <p>In this new version, the&nbsp;randomly sampled&nbsp;training data samples testing data samples from nonhydrostatic simulations in WRF baraclinic wave test&nbsp;are provided. They are processed&nbsp;into a new data structure, which can be directly utilized in training and testing.&nbsp;</p> <p>Follow the instructions in README.txt and download the training and testing data, and the associated codes.</p> <p>Here we provide 3 parts of data and codes:</p> <p>1, Training and testing data from WRF;</p> <p>2, Training and testing codes for two machine learning emulators:&nbsp;machine learning and neural network</p> <p>3, WRF application.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Replication Package for Maintainability Challenges in ML: A Systematic Literature Reveiw

<p>In order to ensure transparency and reproducibility, we have&nbsp;made all study artefacts publicly here. Replication package for <strong>Maintainability Challenges in ML : A Systematic Literature Review</strong>. This package contains the data used&nbsp;for this Systematic Literature Review process and additional data synthesised from this study.</p> <ul> <li>Coding Summary Report Generated from Nvivo project. (Coding Summary By Code Report.pdf)</li> <li>Detailed List of Selected papers for Literature Review papers(Literature Review Papers.xlsx)</li> <li>Papers relating to implications for developers and Researchers(Implication for Developer and Researchers.pdf)</li> <li>Search query from Different Database.pdf</li> <li>No of papers used in to answer RQ using the&nbsp;&nbsp;Literature review paper by categories (No of Literature review papers.pdf)</li> <li>Maintainability challenges in ML workflow from SLR.pdf</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

ML_DP4_data

<p>Data and model&nbsp;needed to run the weather and climate Machine Learning Design Pattern 4 example.&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Science through ML: Post-storm Cooling Data and Programs

<p>These programs and data are associated with the AGU Space Weather Journal article &quot;Science through Machine Learning: Quantification of Post-storm Thermospheric Cooling&quot;.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Global AI, ML, Data salaries

<p><strong>Dataset de Salarios en el Ámbito de Inteligencia Artificial y Ciencia de Datos</strong></p><p>Este conjunto de datos recopila información salarial de diversos puestos de trabajo vinculados a inteligencia artificial y a la ciencia de datos en todo el mundo. Es una adaptación del conjunto original <strong>salaries</strong> que pertenece a <strong>ai-jobs.net</strong>.</p><p>Se ha ampliado con nuevas variables para ofrecer una visión más completa de las condiciones laborales. Las variables originales son:</p><ul><li><strong>work_year</strong>: Año en que se pagó el salario (<i>Categórica</i>)</li><li><strong>experience_level</strong>: Nivel de experiencia (EN: Junior, MI: Medio, SE: Senior, EX: Director/Ejecutivo) (<i>Categórica</i>)</li><li><strong>employment_type</strong>: Tipo de contrato (PT: Tiempo parcial, FT: Tiempo completo, CT: Contrato de servicio, FL: Freelance) (<i>Categórica</i>)</li><li><strong>job_title</strong>: Puesto de trabajo (<i>Categórica</i>)</li><li><strong>salary</strong>: Salario bruto anual (<i>Cuantitativa</i>)</li><li><strong>salary_currency</strong>: Divisa del salario bruto (<i>Categórica</i>)</li><li><strong>salary_in_usd</strong>: Salario en USD (<i>Cuantitativa</i>)</li><li><strong>employee_residence</strong>: País de residencia del empleado (ISO3166) (<i>Categórica</i>)</li><li><strong>remote_ratio</strong>: Cantidad de teletrabajo (0: Presencial, 50: Híbrido, 100: Remoto) (<i>Categórica</i>)</li><li><strong>company_location</strong>: País de la empresa (ISO3166) (<i>Categórica</i>)</li><li><strong>company_size</strong>: Tamaño de la empresa (S: Menos de 50 empleados, M: Entre 50 y 250 empleados, L: Más de 250 empleados) (<i>Categórica</i>)</li></ul><p>Las nuevas variables incluidas son:</p><ul><li><strong>company_continent</strong>: Continente donde se encuentra la empresa (<i>Categórica</i>)</li><li><strong>employee_continent</strong>: Continente donde reside el empleado (<i>Categórica</i>)</li><li><strong>company_continent_region</strong>: Región continental donde está ubicada la empresa (<i>Categórica</i>)</li><li><strong>employee_continent_region</strong>: Región continental donde reside el empleado (<i>Categórica</i>)</li><li><strong>salary_rounded_in_eur</strong>: Salario en euros redondeado (calculado con la media de cada año) (<i>Cuantitativa</i>)</li><li><strong>salary_k_in_eur</strong>: Salario en miles de euros (<i>Cuantitativa</i>)</li><li><strong>salary_type</strong>: Tipo de salario (VL: Very Low, L: Low, M: Medium, H: High, VH: Very High) (<i>Categórica</i>)</li><li><strong>same_country</strong>: ¿Residencia y lugar de trabajo en el mismo país? (<i>Categórica</i>)</li></ul>

opencc-zeroNov 2023View details →
zenodo40/100

Accelerogram (raw, msd), GLT4 seismic station (Galati, Romania), Vrancea (Romania) earthquake, 2020-04-25 01:04:18 ML=5.0 h=21.6 km

<p>Accelerogram (raw, msd) recorded at GLT4 seismic station (Galati, Romania).</p> <p>Vrancea (Romania) earthquake, 2020-04-25 01:04:18 (local time), ML=5.0, h=21.6 km.</p> <p>Recorded on a GeoSIG GMS-18 instrument.</p>

opencc-by-nd-4.0Apr 2024View details →
zenodo40/100

Index based dataset for training ML classification models

<p>This dataset contains 202122 rows of data containing 61 unique indices from different world urban areas.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West) in Bumps on the back: An unusual morphology in phylogenetically distinct Peridinium aff. cinctum (= Peridinium tuberosum; Peridiniales, Dinophyceae)

◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West)

opencc-by-4.0Jan 2024View details →
zenodo40/100

◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae) in Morphological and molecular variability of Peridinium volzii Lemmerm. (Peridiniaceae, Dinophyceae) and its relevance for infraspecific taxonomy

◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae)

opencc-by-4.0Oct 2021View details →
zenodo40/100

Brazilian Cohort for Predicting Cardiovascular Events Using Machine Learning (PRE-CARE ML project)

<p>The project PRE-CARE ML addresses the development and internal and external validation of predictive models for the assessment of risks of major adverse cardiovascular events.&nbsp;</p> <p>Global and local interpretability analyses of predictions were conducted towards improving model reliability and tailoring preventive interventions.&nbsp;</p> <p>The models were trained and validated in a retrospective cohort with the use of data from Hospital das Cl&iacute;nicas da Faculdade de Medicina de Ribeirao Preto, Brazil</p> <p>The International Classification of Diseases (ICD-10) classified patients with MACE (case group). (see Table I for the ICD-10 codes that defined MACE).</p> <p>Table 1: ICD-10 Codees for MACE definition</p> <table> <tbody> <tr> <td><strong>ICD-10</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>I20</td> <td>Angina pectoris</td> </tr> <tr> <td>I21</td> <td>Acute myocardial infarction</td> </tr> <tr> <td>I24</td> <td>Other acute ischaemic heart diseases</td> </tr> <tr> <td>I46</td> <td>Cardiac arrest</td> </tr> <tr> <td>I63</td> <td>Cerebral infarction</td> </tr> <tr> <td>I64</td> <td>Stroke, not specified as hemorrhage or infarction</td> </tr> <tr> <td>I71</td> <td>Aortic aneurysm and dissection</td> </tr> <tr> <td>I74</td> <td>Arterial embolism and thrombosis</td> </tr> </tbody> </table> <p>Only a patient&rsquo;s first MACE was considered, and all previous hospitalizations within a 5-year window were considered MACE. The control group (non-MACE) involved hospitalizations with no MACE and no death within a 5-year window (between 2017 and 2022). A sample of 6,000 MACE (labeled as 1) cases and 12,000 non-MACE (labeled as 0) ones was constructed with data from HCFMRP for training and internal validation purposes. Another balanced MIMIC IV sample of 8,000 MACE cases and 8,000 nonMACE cases was used for external validation.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019) in Variability of the gene cyt b in the Korean field mouse Apodemus peninsulae Thomas, 1906 - a reservoir host of AMRV in the Khasansky District of Primorsky Krai

Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019)

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)

Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain

opencc-by-4.0Jul 2024View details →
zenodo40/100

Bio-ML: Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching

<p>&nbsp;</p> <blockquote> <p><strong>This version is used in the Bio-ML track of the OAEI 2024; the only change compared to the OAEI 2023 is the deletion of certain training subsumption mappings.</strong></p> </blockquote> <p>&nbsp;</p> <h3><strong>Overview</strong></h3> <p>The purpose of these datasets is to support&nbsp;<em>equivalence</em> and <em>subsumption</em> ontology matching.</p> <p>There are five ontology pairs extracted from MONDO and UMLS:</p> <table> <tbody> <tr> <td>Source</td> <td>Task</td> <td>Category</td> <td>#SrcCls</td> <td>#TgtCls</td> <td>#Ref (equiv)</td> <td>#Ref (subs)</td> </tr> <tr> <td>Mondo</td> <td>OMIM-ORDO</td> <td>Disease</td> <td>9,648</td> <td>9,275</td> <td>3,721</td> <td>103</td> </tr> <tr> <td>Mondo</td> <td>NCIT-DOID</td> <td>Disease</td> <td>15,762</td> <td>8,465</td> <td>4,686</td> <td>3,338 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-FMA</td> <td>Body</td> <td>34,418</td> <td>88,955</td> <td>7,256</td> <td>5,453 (-53)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Pharm</td> <td>29,500</td> <td>22,136</td> <td>5,803</td> <td>4,224 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Neoplas</td> <td>22,971</td> <td>20,247</td> <td>3,804</td> <td>213</td> </tr> </tbody> </table> <p>The "-" numbers reflect the changes due to lthe deletion of certain training subsumption mappings.</p> <p>The main track is available at "bio-ml", where each pair is associated with a task folder, containing the source and target ontologies, reference equivalence mappings (in "refs_equiv"), reference subsumption mappings ("refs_subs").&nbsp;</p> <p>The special sub-track is available at "bio-llm", where each pair is associated with a task folder, containing the source and target ontologies, and the test candidate mappings.&nbsp;</p> <p>&nbsp;</p> <h3><strong>Citation</strong></h3> <p><strong>Bio-ML (Main Track)</strong></p> <pre>```<br>@inproceedings{he2022machine, title={Machine learning-friendly biomedical datasets for equivalence and subsumption ontology matching}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Jim{\'e}nez-Ruiz, Ernesto and Hadian, Ali and Horrocks, Ian}, booktitle={International Semantic Web Conference}, pages={575--591}, year={2022}, organization={Springer} }<br>```</pre> <p><strong>Bio-LLM (Sub-track)</strong></p> <pre>```<br>@article{he2023exploring, title={Exploring large language models for ontology alignment}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Horrocks, Ian}, journal={arXiv preprint arXiv:2309.07172}, year={2023} }<br>```</pre> <p>&nbsp;</p> <h3><strong>Important Links</strong></h3> <ul> <li>See detailed documentation at:&nbsp;<a href="https://krr-oxford.github.io/DeepOnto/bio-ml">https://krr-oxford.github.io/DeepOnto/bio-ml</a>.</li> <li>See the OAEI Bio-ML track at:&nbsp;<a href="https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/">https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/</a></li> <li>See our resource paper for the original Bio-ML at&nbsp;<a href="https://arxiv.org/abs/2205.03447">arxiv</a>&nbsp;or <a href="https://link.springer.com/chapter/10.1007/978-3-031-19433-7_33">springer</a>&nbsp;(accepted at&nbsp;<em>ISWC-2022</em> and nominated as the <em>best resource paper candidate</em>). See our poster paper for the Bio-LLM sub-track at&nbsp;<a href="https://arxiv.org/abs/2309.07172">arxiv </a>(accepted at <em>ISWC-2023 Posters &amp; Demos</em>).</li> </ul> <p>&nbsp;</p> <h3><strong>Changelog</strong></h3> <p>The only change in this version compared to the OAEI 2023 is the deletion of certain training subsumption mappings that can be directly exploited through deductive reasoning.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

RAIRS spectrum of a H2O:CS2 (400:60 ML) ice mixture at 6 K

<p>The title is self-explanatory. First column contains wavenumbers in cm-1. Second column contains absorbance. &nbsp;</p> <p>The spectrum was collected in the SPACE TIGER setup at the Center for Astrophysics | Harvard &amp; Smithsonian. We assume that the spectrum is the same as in transmittance mode except for the absorbance, that is a factor 2.32 higher (note that this factor is setup-dependent and was measured for the SPACE TIGER setup in Mart&iacute;n-Dom&eacute;nech et al. 2020).&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Рис. 7–16. Carabus (Ophiocarabus) spp., генитаΛии самцов, виà сбоку (7–11 – меÃиаΛьная ÃоΛя эÃеагуса; 12–16 – энÃофаΛΛус). 7–8, 12–13 – C. ernsti ulastaiensis subsp. n., паратипы; 9–10, 14–15 – C. ernsti ernsti Kabak, 2002: 9, 14 – с гор к запаÃу от реки Сарык, 10, 15 – из ÃоΛины правого притока реки Сарык; 11, 16 – C. ernsti nilkiensis Kabak, 2014: 11 – гоAотип, 12 – паратип. ml – меÃиаΛьный бугор; pp – препуциаΛьный бугор; rbll – правый базоΛатераΛьный бугор. Figs 7–16. Carabus (Ophiocarabus) spp., male genitalia, lateral view (7–11 – medial lobe of the aedeagus; 12–16 – endophallus). 7–8, 12–13 – C. ernsti ulastaiensis subsp. n., paratypes; 9–10, 14–15 – C. ernsti ernsti Kabak, 2002: 9, 14 – from the locality in mountains to the West of the Saryk River, 10, 15 – from the locality in right tributary of the Saryk River; 11, 16 – C. ernsti nilkiensis Kabak, 2014: 11 – holotype, 12 – paratype. ml – median lobe; pp – praeputial pad; rbll – right basolateral lobe. in New data on the taxonomy of the genus Carabus Linnaeus, 1758 (Coleoptera: Carabidae) from the Ili River basin (China)

Рис. 7–16. Carabus (Ophiocarabus) spp., генитаΛии самцов, виà сбоку (7–11 – меÃиаΛьная ÃоΛя эÃеагуса; 12–16 – энÃофаΛΛус). 7–8, 12–13 – C. ernsti ulastaiensis subsp. n., паратипы; 9–10, 14–15 – C. ernsti ernsti Kabak, 2002: 9, 14 – с гор к запаÃу от реки Сарык, 10, 15 – из ÃоΛины правого притока реки Сарык; 11, 16 – C. ernsti nilkiensis Kabak, 2014: 11 – гоAотип, 12 – паратип. ml – меÃиаΛьный бугор; pp – препуциаΛьный бугор; rbll – правый базоΛатераΛьный бугор. Figs 7–16. Carabus (Ophiocarabus) spp., male genitalia, lateral view (7–11 – medial lobe of the aedeagus; 12–16 – endophallus). 7–8, 12–13 – C. ernsti ulastaiensis subsp. n., paratypes; 9–10, 14–15 – C. ernsti ernsti Kabak, 2002: 9, 14 – from the locality in mountains to the West of the Saryk River, 10, 15 – from the locality in right tributary of the Saryk River; 11, 16 – C. ernsti nilkiensis Kabak, 2014: 11 – holotype, 12 – paratype. ml – median lobe; pp – praeputial pad; rbll – right basolateral lobe.

opencc-by-4.0Dec 2019View details →
zenodo40/100

Actors and Satellites in the African Earth Observations Sector: Insights from the 2021 Radiant Earth ML for EO Market Map and the Union of Concerned Scientists Database

<p>The database of organizational actors, "Actors and Satellites in the African Earth Observations Sector: Insights from the 2021 Radiant Earth ML for EO Market Map and the Union of Concerned Scientists Database" analyzed in "<span>Whose Priorities? Examining Inequities in Earth </span><span>Observation Advancements Across Africa" </span>this study, is available on Zenodo, an open-access repository developed under the European OpenAIRE program. The dataset comprises information on 310 space-centric earth observation organizations, including headquarters locations.&nbsp;For the 31 organizations in our sample, we provide additional details including the African countries where their projects are active, the type of initiative or program, other focus areas, organizational classification (commercial, government, or nongovernmental), funding source (public or private), organizational type (research, startup, or established industry), capabilities (data analysis, data storage, image labeling, competition platforms), involvement in early warning systems, data accessibility, availability of global products, and whether they build commercial satellites.</p> <p>This open sharing of the compiled organizational data aims to promote transparency, reproducibility, and additional investigations into the evolving landscape of earth observation activities globally and across Africa. Analyses of this dataset's relationships, funding flows, and priorities can provide further insights to guide equitable advancement of earth observation capabilities.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Generic and ML Workloads in an HPC Datacenter

<p>Updated Version of the <a title="previous upload" href="../records/13625495">previous upload</a>, adjusts node timestamps lacking behind at the beginning of the data collection.</p> <p>This archive contains hardware and workload traces from SURF Lisa, a Dutch datacenter consisting of 338 nodes, used by universities and researchers for various jobs. Around 85% of the nodes are equipped only with CPUs, handling generic compute-heavy workloads, the other 15% come with additional GPUs, serving as accelerators for Machine Learning (ML) jobs. Individual node hardware configurations are listed in `node_hardware_info.parquet`.</p> <p>Jobs within Lisa are submitted over the SLURM scheduler, where we logged job start and end time, resource allocation, and exit state for roughly 10 months (December 2021 to November 2022). This data saved in `slurm_table_cleaned.parquet`.</p> <p>Addidionally, we provide detailed Prometheus monitoring logs from all nodes over a timespan of 5 months (June 2022 to November 2022) in `prom_table_cleaned.parquet`. These logs contain over 90 attributes, including CPU/GPU power and temperatures, network I/O, memory and storage usage, and many more. These metrics are sampled at 30s intervals, resulting in a total of almost 130 million records across all nodes.</p> <p>Finally, job and node data are provided as a joined dataset in `prom_slurm_joined.parquet` for their 4 months of overlapping timespan. This combined data can provide more insights into the resource consumption and performance patterns of jobs.</p> <p>We conducted detailed analysis of this data where we specifically looked at the different characteristics of generic vs. ML workloads in a heterogeneous HPC environment. The pre-print of our analysis work can be found on <a href="https://arxiv.org/abs/2409.08949">arXiv</a>. Our code used for evaluation can be found on <a href="https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization">GitHub</a>.<br><br></p> <table> <tbody> <tr> <th>Dataset Name</th> <th>Explanation</th> </tr> <tr> <td>slurm_table_cleaned.parquet</td> <td>Job data collected by SLURM</td> </tr> <tr> <td>prom_table_cleaned.parquet</td> <td>Node data collected by Prometheus</td> </tr> <tr> <td>prom_slurm_joined.parquet</td> <td>Joined Job and Node dataset</td> </tr> <tr> <td>node_hardware_info.parquet</td> <td>Hardware configurations of each node</td> </tr> </tbody> </table>

opencc-by-4.0Aug 2024View details →
zenodo40/100

ECOCLIMAP-SG-ML: an ensemble land cover map for numerical weather prediction

<p>This dataset contains ensemble land cover maps for numerical weather prediction at 60 m resolution over Europe. As they were<br>generated thanks to machine learning, the weights and the training data are also provided.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record