Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

195

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

195 results for “biomedical”

Learn how ShareScore rates datasets ↗
zenodo52/100

BioASQ-QA: A manually curated corpus for Biomedical Question Answering

<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>

opencc-by-2.5Dec 2022View details →
zenodo48/100

MiRoR15-P2-Development of ARCADIA: a tool for assessing the quality of peer-review reports in biomedical research

<p>Survey questionnaire, anonymised survey data, and codebook related to: Superchi C, Hren D, Blanco D, Rius R, Recchioni A, Boutron I, Gonz&aacute;lez JA. Development of ARCADIA: a tool for assessing the quality of peer-review reports in biomedical research. BMJ Open 2020;0:e035604. doi:10.1136/bmjopen-2019-035604</p>

opencc-by-4.0Aug 2020View details →
zenodo48/100

Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes

<p>This contains the merged dataset as described in the work "<strong>Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes"</strong>.</p> <p>This dataset consists of 4 seperate datasets:</p> <ul> <li><a href="../records/8224056" target="_blank" rel="noopener">MedProcNer</a></li> <li><a href="../records/7614764" target="_blank" rel="noopener">DisTEMIST</a></li> <li><a href="../records/4270158" target="_blank" rel="noopener">PharmaCoNER</a></li> <li><a href="../records/10635215" target="_blank" rel="noopener">SympTEMIST</a></li> </ul> <p>The dataset contains two tasks:</p> <p><strong>Task 1:</strong> This task is related to multi-class Named Entity Recognition. This dataset contains 5 possible classes: SYMPTOM, PROCEDURE, DISEASE, CHEMICAL and PROTEIN.</p> <p><strong>Task 2:</strong> This task is related to Named Entity Linking, where each code corresponds to a code within the SNOMED-CT corpus. The exact corpus used can be obtained <a href="https://download.nlm.nih.gov/umls/kss/IHTSDO20190131/SnomedCT_SpanishRelease-es_PRODUCTION_20190430T120000Z.zip" target="_blank" rel="noopener">here</a>. Further for the MedProcNER, SympTEMIST and DisTEMIST datasets, a gazetteer is provided in the original datasets.&nbsp;</p> <p>For more information on the construction of the dataset, aswell as dataloaders, we refer you to our <a href="https://github.com/ieeta-pt/Multi-Head-CRF" target="_blank" rel="noopener">GitHub repository</a>.<br><br>Further this also contains the embeddings from the <a href="https://huggingface.co/cambridgeltl/SapBERT-UMLS-2020AB-all-lang-from-XLMR-large" target="_blank" rel="noopener">SapBERT</a> model.</p> <p><strong>Please, cite:</strong></p> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <blockquote> <div>@article{jonker2024a, title = {Multi-head {{CRF}} classifier for biomedical multi-class named entity recognition on {{Spanish}} clinical notes}, author = {Jonker, Richard A. A. and Almeida, Tiago and Antunes, Rui and Almeida, Jo{\~a}o R. and Matos, S{\'e}rgio}, year = {2024}, journal = {Database}, publisher = {Oxford University Press} }</div> </blockquote> <div>Jonker, R. A. A., Almeida, T., Antunes, R., Almeida, J. R., &amp; Matos, S. (2024). Multi-head CRF classifier for biomedical multi-class named entity recognition on Spanish clinical notes. (Submitted.)&nbsp;</div> <div>&nbsp;</div> <div> <p><strong>License</strong></p> <p>This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>

opencc-by-4.0May 2024View details →
zenodo48/100

Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions

<p>This data set corresponds to the paper: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions [1] (Experiment: Comprehensive reporting).</p> <p>The key research questions corresponding to this data set were:</p> <p>RQ1: What is the role of challenges for the field of biomedical image analysis (e.g. How many challenges conducted to date? In which fields? For which algorithm categories? Based on which modalities?)</p> <p>RQ2: What is common practice related to challenge design (e.g. choice of metric(s) and ranking methods, number of training/test images, annotation practice etc.)? Are there common standards?</p> <p>RQ3: Does common practice related to challenge reporting allow for reproducibility and adequate interpretation of results?</p> <p>To address these research questions, we aimed to capture all biomedical image analysis challenges that have been conducted up to 2016. To acquire the data, we analyzed the websites hosting/representing biomedical image analysis challenges, namely grand-challenge.org, dreamchallenges.org and kaggle.com as well as websites of main conferences in the field of biomedical image analysis, namely Medical Image Computing and Computer Assisted Intervention (MICCAI), International Symposium on Biomedical Imaging (ISBI), International Society for Optics and Photonics (SPIE) Medical Imaging, Cross Language Evaluation Forum (CLEF), International Conference on Pattern Recognition (ICPR), The American Association of Physicists in Medicine (AAPM), the Single Molecule Localization Microscopy Symposium (SMLMS) and the BioImage Informatics Conference (BII). This yielded a list of 150 challenges with 549 tasks.</p> <p>Next, a tool for instantiating the challenge parameter list introduced in [1] was used by some of the authors (engineers and medical student) to formalize all challenges that met our inclusion criteria as follows: (1) Initially, each challenge was independently formalized by two different observers. (2) The formalization results were automatically compared. In ambiguous cases, when the observers could not agree on the instantiation of a parameter - a third observer was consulted, and a decision was made. When refinements to the parameter list were made, the process was repeated for missing values. Based on the formalized challenge data set, a descriptive statistical analysis was performed to characterize common practice related to challenge design and reporting.</p> <p>[1] Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., M&auml;rz, K., Maier, O., Maier-Hein, K., Menze, B. H., M&uuml;ller, H., Neher, P. F., Niessen, W., Rajpoot, N., Sharp, G. C., Sirinukunwattana, K., Speidel, S., Stock, C., Stoyanov, D., Aziz Taha, A., van der Sommen, F., Wang, C.-W., Weber, M.-A., Zheng, G., Jannin, P., Kopp-Schneider, A.: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. arXiv preprint arXiv:1806.02051 (2018).</p>

opencc-by-4.0Jun 2018View details →
zenodo48/100

Workflow for detecting biomedical articles with openly available underlying datasets - Datasets and extraction forms

<p>The open data screening datasets contain both automatically detected (TRUE) Open Data statements by <a href="https://github.com/quest-bih/oddpub">ODDPub</a>, and its manual validation using <a href="https://github.com/bgcarlisle/Numbat">Numbat</a> extraction tool. Furthermore, extraction forms for both screenings &ndash; 2020 and 2021 &ndash; are included. The manually processed dataset for the calculation of the inter-rater reliability of manual validation can be also found here.&nbsp;&nbsp;</p> <p>(i) Data from articles published in 2020 (file &lsquo;<em>charite_open_data_2020.csv</em>&rsquo;) have been collected applying a slightly different sequence of questions in the extraction workflow than the articles published in 2021 (file &lsquo;<em>charite_open_data_2021.csv</em>&rsquo;). Both datasets were cleaned for any personal data or internal comments. Thus, they do not contain the default columns which in the raw export from Numbat contained commentaries regarding different question. Also, in another regard these files do not represent raw outputs of the Numbat extraction tool, but a processed version. This means that articles validated by more than two raters were first reconciled in Numbat, resulting in one final decision (output of extractions <strong>after reconciliation</strong>). Then from the output of extractions <strong>before reconciliation</strong> those articles validated by only 1 rater (and thus not part of the inter-rater reliability calculation) were selected, which were afterwards joined with the already reconciled dataset.&nbsp;&nbsp;</p> <p>The actual decision about Openness of validated dataset can be analysed in various ways:&nbsp;</p> <ol> <li>Column &lsquo;<em>open_data_assessment</em>&rsquo;/&rsquo;<em>assessment</em>&rsquo; shows a binary decision between Open Data TRUE and FALSE.&nbsp;</li> <li>If that column indicates &lsquo;<em>NULL</em>&rsquo;, the dataset was classified into &lsquo;non&rsquo;-open category, and the result can be found on one of the following ways:&nbsp; <ul> <li>Column &lsquo;<em>reference_to_data</em>&rsquo; as &lsquo;<em>n_a</em>&rsquo; for excluded articles, e.g. not producing any data.</li> <li>Column &lsquo;<em>data_access</em>&rsquo; as &lsquo;<em>restricted</em>&rsquo;.&nbsp;</li> <li>Column &lsquo;<em>own_or_reuse_data</em>&rsquo; as &lsquo;<em>open_data_reuse</em>&rsquo;.&nbsp;</li> </ul> </li> </ol> <p>The original extraction form contains an option &lsquo;unsure_open_data&rsquo; besides &lsquo;<em>open_data</em>&rsquo;/&rsquo;<em>no_open_data</em>&rsquo; which was resolved either during reconciliation between multiple raters or by case-related consultation with a second rater in case of doubt, and is not included here.&nbsp;</p> <p>(ii) The inter-rater reliability calculation was made on randomly selected 100 articles for 2 raters. The third rater screened 20 articles sample, which is part of 100 sample. The tables provided here include both article-level data, and dataset-level data.&nbsp;</p> <p>(iii) The Numbat extarction forms used for the screenings in 2020 and 2021 are included in two formats - JSON and Markdown.</p> <p>(iv) &lsquo;<em>data_dictionary_open_data.csv</em>&rsquo; table documents all variables of each data file containing here.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Transmission ultrasound data simulated using the k-Wave toolbox as a benchmark for biomedical quantitative ultrasound tomography using a ray approximation to Green's function

<p><strong>Transmission ultrasound data simulated using the k-Wave toolbox as a benchmark for biomedical quantitative ultrasound tomography using a ray approximation to&nbsp;Green&#39;s function&nbsp;</strong></p> <p>&nbsp;</p> <p>The folder &lsquo;&rsquo;simulation<em>&rsquo;&rsquo; </em>includes the transmission ultrasound data sets used in the project:<a href="https://github.com/Ash1362/ray-based-quantitative-ultrasound-tomography">https://github.com/Ash1362/ray-based-quantitative-ultrasound-tomography</a>. In the Github link, the associated project can be found in the branch master in the folder r-Wave #V1.1. (The folder &lsquo;&rsquo;data_ust_kWave_transmission.zip<em>&rsquo;&rsquo; </em>is deprecated.)</p> <p>...........................................................................................</p> <p>The ultrasound data were simulated using the k-Wave toolbox (version 1.3.)&nbsp; [5] and using a digital breast phantom [4]. In k-Wave version 1.4., no changes have been reported that affects the simulations. The simulations were done assuming isotropic point sources.</p> <p>The&nbsp;folder&nbsp;&lsquo;&rsquo;simulation<em>&rsquo;&rsquo;&nbsp;</em>&nbsp;must be added to the path:</p> <p><em>&#39;&#39;&hellip;r-Wave/data/simulation/&hellip;&#39;&#39;</em></p> <p>For running the Matlab example scripts in the project in the github, the user has two choices:&nbsp;</p> <ol> <li>Simulate the k-Wave ultrasound data by setting <em>data_sim=true;</em> in the examples in the project.</li> <li>Upload the already simulated k-Wave ultrasound data according to the description below and load them by setting &nbsp;<em>data_sim=false;</em>&nbsp;in the examples in the project.</li> </ol> <p>Please read the description in the example scripts!</p> <p>&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;</p> <p>The folder simulation includes 2 subfolders, &lsquo;&rsquo;phantom<em>&rsquo;&rsquo;&nbsp;</em>and&nbsp;&lsquo;&rsquo;data_ust_kWave_transmission<em>&rsquo;&rsquo;.</em></p> <p>1) The subfolder&nbsp;&lsquo;&rsquo;simulation/phantom<em>&rsquo;&rsquo;&nbsp;</em>&nbsp;includes&nbsp;&lsquo;&rsquo;OA-BREAST<em>&rsquo;&rsquo;.&nbsp;</em></p> <p>In the project: https://anastasio.bioengineering.illinois.edu/downloadable-content/oa-breast-database/,</p> <p>the user must upload the folder&nbsp;&lsquo;&rsquo;Neg_47_Left<em>&rsquo;&rsquo;&nbsp;</em>, and add it as&nbsp;&nbsp;&lsquo;&rsquo;r-wave/data/simulation/phantom/OA-BREAST/Neg_47_Left/<em>&rsquo;&rsquo;.</em></p> <p><em>.......................................................................................................................................................................</em></p> <p>2) The&nbsp;subfolder &lsquo;&rsquo;simulation/data_ust_kWave_transmission&rsquo;<em>&rsquo;&nbsp; </em>includes 2 subfolders, &lsquo;&rsquo;2D<em>&rsquo;&rsquo;&nbsp;</em> and &lsquo;&rsquo;3D<em>&rsquo;&rsquo;&nbsp;</em>.</p> <p>The subfolder&nbsp;&lsquo;&rsquo;2D<em>&rsquo;&rsquo;&nbsp;</em> includes:</p> <p><strong>data_ust_kWave_transmission/2D/PulsePammoth_1_dx4_cfl1_Nr256_Ne64_Interpoffgrid_Transgeompoint_Absorption1_CodeMatlab/data4_sphere_nonsmooth.mat</strong></p> <p>Two transmission ultrasound data sets were simulated using the k-wave for only water and breast in water according to section <em>&lsquo;&rsquo;6.1. data simulation&rsquo;&rsquo;</em> in [1]. 64 emitters and 256 receivers are simulated as off-grid points which are placed on a 2D circular ring. (The characters&nbsp;&lsquo;&rsquo;_sphere_&rsquo;&rsquo;&nbsp; are added to indicate that the transducers are placed on a ring.) To simulate the data, each emitter was individually driven by an excitation pulse, and the induced acoustic pressure time series were recorded on all the receivers. The k-Wave simulation was performed on a grid with grid spacing 0.4 mm, and the time spacing was set using a CFL number 0.1. The acoustic absorption and dispersion were accounted for based on the frequency power law. This data set is used for the purpose of image reconstruction, and therefore, the sound speed and absorption coefficients maps are not smoothed, i.e., the original maps are used for simulations. This data set can be used for image reconstruction using the time-of-flight-based approach and then the Green&#39;s approach.</p> <p><strong>data_ust_kWave_transmission/2D/PulsePammoth_1_dx4_cfl1_Nr256_Ne64_Interpoffgrid_Transgeompoint_Absorption1_CodeMatlab/data4_plane_nonsmooth.mat</strong></p> <p>Two transmission ultrasound data sets were simulated using the k-wave for only water and breast in water. 64 emitters and 256 receivers are simulated as off-grid points which are placed on 16 planar arrays which are all aligned with a circle. Each planar array includes 4 emitters and 16 receivers. Therefore, in contrast with&nbsp;the data mentioned above, the ray linking is performed using the line equations defining the 2D geometry of the linear arrays. (The characters&nbsp;&lsquo;&rsquo;_plane_&rsquo;&rsquo;&nbsp; are added to indicate that the transducers are placed on line.)&nbsp;To simulate the data, each emitter was individually driven by an excitation pulse, and the induced acoustic pressure time series were recorded on all the receivers. The k-Wave simulation was performed on a grid with grid spacing 0.4 mm, and the time spacing was set using a CFL number 0.1. The acoustic absorption and dispersion were accounted for based on the frequency power law. This data set is used for the purpose of image reconstruction, and therefore, the sound speed and absorption coefficients maps are not smoothed, i.e., the original maps are used for simulations. This data set can be used for image reconstruction using the time-of-flight-based approach, but ahs&nbsp;not been extended to the Green&#39;s approach yet. The image reconstruction should be slower than the circular array. the reason is&nbsp;for circular array,&nbsp;for each emitter, the raylinking problem is solved for all receivers once using the equation of circle. However, for this data set, for each emitter, the ray linking problem is solved for each receiver array&nbsp;separately, because receiver arrays are defined with different line equations.</p> <p><strong>data_ust_kWave_transmission/2D/PulsePammoth_1_dx4_cfl1_Nr256_Ne64_Interpoffgrid_Transgeompoint_Absorption1_CodeMatlab/data4_sphere_smooth_17_1.mat</strong></p> <p>Two transmission ultrasound data sets were simulated using the k-Wave for only water and breast in water &nbsp;as the benchmark for validation of ray approximation to&nbsp;Green&rsquo;s function in homogeneous&nbsp;and heterogenous media, respectively. The simulation was performed&nbsp;according to section <em>&lsquo;&rsquo;6.2. Numerical validation of the ray approximation to the Green&rsquo;s function&rsquo;&rsquo;</em> in [1].</p> <p>64 emitters and 256 receivers are simulated as off-grid points which are placed on a 2D circular ring. (The characters&nbsp;&lsquo;&rsquo;_sphere_&rsquo;&rsquo;&nbsp; are added to indicate that the transducers are placed on a ring.) The pressure field was produced by emitter 1 (of&nbsp;the 64 emitters) and was recorded in time on all 256 receivers. The k-Wave simulation was performed on a grid with grid spacing 0.4 mm, and the time spacing was set using a CFL number&nbsp;0.1. The acoustic absorption and dispersion were accounted for based on the frequency power law. The sound speed and absorption coefficient maps were smoothed by an averaging window of size 17 grid points. This data set is used as the benchmark for measuring accuracy of ray approximation to Green&rsquo;s function for&nbsp;computing phase and amplitude of the pressure field on the receivers.</p> <p><strong>data_ust_kWave_transmission/2D/PulsePammoth_1_dx4_cfl1_Nr256_Ne64_Interpoffgrid_Transgeompoint_Absorption1_CodeMatlab/data4_sphere_smooth_17_20.mat</strong></p> <p>&nbsp;This data set is the same as data4_smooth_17_1&nbsp;except&nbsp;the pressure field is produced by emitter 20.</p> <p>&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;.</p> <p>The subfolder &lsquo;&rsquo;3D<em>&rsquo;&rsquo;&nbsp;</em> includes:</p> <p><strong>data_ust_kWave_transmission/3D/PulsePammoth_1_dx5_cfl1_Nr4096_Ne1024_Interpnearest_Transgeompoint_Absorption0_CodeCUDA/data5_sphere_nonsmooth_tof_singram.mat</strong></p> <p>The discrepancy of time-of-flight data for two transmission ultrasound data sets simulated by the k-wave for breast in water and only water according to section 5.2 in [3]. The pressure fields were produced by 1024 emitters separately and were recorded on 4096 receivers. The emitters and receivers were simulated as points which are placed on a 3D hemispherical surface, and are interpolated onto the grid using a neighboring interpolation. &nbsp;The k-Wave simulations were performed on a grid with grid spacing 0.5 mm, and the time spacing was set using a CFL number 0.1. The time-of-flight data were computed and will be used for a refraction-corrected image reconstruction of the sound speed based on the inversion approach proposed in [3].</p> <p><strong>References</strong></p> <p>1 - A. Javaherian, ❝Hessian-inversion-free ray-born inversion for high-resolution quantitative ultrasound tomography❞, 2022, <a href="https://arxiv.org/abs/2211.00316/">https://arxiv.org/abs/2211.00316/</a> .</p> <p>2 - A. Javaherian and B. Cox, ❝Ray-based inversion accounting for scattering for biomedical ultrasound tomography❞, Inverse Problems vol. 37, no.11, 115003, 2021. &nbsp;<a href="https://iopscience.iop.org/article/10.1088/1361-6420/ac28ed/">https://iopscience.iop.org/article/10.1088/1361-6420/ac28ed/</a></p> <p>3- A. Javaherian, F. Lucka and B. T. Cox, ❝Refraction-corrected ray-based inversion for three-dimensional ultrasound tomography of the breast❞, Inverse Problems, 36 125010. &nbsp;<a href="https://iopscience.iop.org/article/10.1088/1361-6420/abc0fc/">https://iopscience.iop.org/article/10.1088/1361-6420/abc0fc/</a> &nbsp;</p> <p>4- Y. Lou, W. Zhou, T. P. Matthews, C. M. Appleton and M. A. Anastasio, ❝Generation of anatomically realistic numerical phantoms for photoacoustic and ultrasonic breast imaging❞, J. Biomed. Opt., vol. 22, no. 4, pp. 041015, 2017. <a href="https://anastasio.bioengineering.illinois.edu/downloadable-content/oa-breast-database/">https://anastasio.bioengineering.illinois.edu/downloadable-content/oa-breast-database/</a></p> <p>5 - B. E. Treeby and B. T. Cox, ❝k-Wave: MATLAB toolbox for the simulation and reconstruction of photoacoustic wave fields❞, J. Biomed. Opt. vol. 15, no. 2, 021314, 2010. <a href="http://www.k-wave.org/">http://www.k-wave.org/</a></p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

RDF Reification Benchmark (REF) using the Biomedical Knowledge Repository (BKR)

<p>This resource&nbsp;can be used for benchmarking different RDF modelling solutions for statement-level metadata, namely:&nbsp;</p> <p>- RDF Reification,</p> <p>- Singleton Property,</p> <p>- RDF* (RDF-star).&nbsp;</p> <p>&nbsp;</p> <p>More details about this resource can be found in the following publication:</p> <p>Fabrizio Orlandi, Damien Graux, Declan O&#39;Sullivan, &quot;Benchmarking RDF Metadata Representations: Reification, Singleton Property and RDF*&quot;,&nbsp;<em>15th IEEE International Conference on Semantic Computing (ICSC)</em>, 2021.</p> <p>Pre-print available at: http://fabriziorlandi.net/pdf/2021/ICSC2021_REF-Benchmark.pdf</p> <p>&nbsp;</p> <p>The&nbsp;dataset&nbsp;contains 3 different versions of the&nbsp;Biomedical Knowledge Repository (BKR) knowledge graph, as described in:</p> <p>Vinh Nguyen,&nbsp;Olivier Bodenreider,&nbsp;Amit Sheth. &quot;Don&#39;t Like RDF Reification? Making Statements About Statements Using Singleton Property&quot; WWW 2014,&nbsp;doi: 10.1145/2566486.2567973.</p> <p>and,</p> <p>Satya S. Sahoo, Olivier Bodenreider, Pascal Hitzler, Amit Sheth&nbsp;and&nbsp;Krishnaprasad Thirunarayan. &quot;Provenance Context Entity (PaCE): Scalable Provenance Tracking for Scientific RDF Data&quot; in Sci Stat Database Manag. 2010; 6187: 461&ndash;470. doi:&nbsp;10.1007/978-3-642-13818-8_32</p> <p>&nbsp;</p> <p>The 3 knowledge graphs&nbsp;dumps&nbsp;are packaged&nbsp;as Gzipped RDF files in Turtle (and Turtle*) syntax.&nbsp;</p> <p>BKR-R-fullKGdump.ttl.gz for the Reification method,</p> <p>BKR-S-fullKGdump.ttl.gz&nbsp;for the Singleton method,</p> <p>BKR-star-fullKGdump.ttls.gz&nbsp;for the RDF* (RDF-star) method.</p> <p>&nbsp;</p> <p>The RDF REiFication Benchmark&nbsp;(REF)&nbsp;includes also&nbsp;a set of SPARQL (and SPARQL*) queries that can be used to compare the performance of different triplestores.</p> <p>Details about the SPARQL queries, and the queries themselves, are included in the &quot;REF-Benchmark.tar.gz&quot; archive. The queries are named after the dataset they are designed for (BKR-R or BKR-S or BKR-star), plus they include a letter identifying&nbsp;the query set, and a query number.&nbsp;</p> <p>E.g. the query in the file &quot;BKR-R_F-Q3.rq&quot; is for the BKR-R (standard reification) dataset, it is part of the query set &quot;F&quot; and it is the number 3 of that set &quot;F&quot;. Hence, the same query, but translated for the RDF* dataset in SPARQL* syntax, is contained in &quot;BKR-star_F-Q3.rq&quot;.</p> <p>Sets &quot;A&quot; and &quot;B&quot; are derived from the queries introduced by V. Nguyen et al. in: &quot;Don&#39;t Like RDF Reification? Making Statements About Statements Using Singleton Property&quot; WWW 2014,&nbsp;doi: 10.1145/2566486.2567973. Set &quot;F&quot; has been designed more with RDF* in mind as part of this benchmark (see [Orlandi et al., ICSC 2021])&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

openapache2.0Oct 2020View details →
zenodo44/100

Practices and policies of preprint platforms for life and biomedical sciences

<p>Given the increase in the use and profile of preprint servers &ndash; and alternative publishing hybrid platforms such as F1000 Research &ndash; in the life sciences, it is increasingly important to identify how many such servers and hybrids exist, to describe their scope in terms of the scientific disciplines they cover, and to compare and contrast their characteristics and policies.</p> <p>We surveyed forty-four (44) platforms that host preprints relevant to life and biomedical sciences and that were active online and accepting submissions on 25 June 2019. Information on preprint platform policies, features and practices was collected through online research by the authors and by surveying preprint platform representatives directly.&nbsp;</p> <p>Full data sheets include an additional 5 platforms hosted on OSF Preprints&nbsp;(rows 49-53) to fulfil the wider scope for the ASAPbio project,&nbsp;not in disciplinary scope (biology and medical sciences) for the manuscript with Jamie Kirkham.</p> <p><strong>Tables 1-5: </strong>Data&nbsp;(44 platforms, manuscript) are separated into five main tables of information and a list of preprint platform websites for reference.</p> <p>Table 1: Scope and ownership of each server<br> Table 2: Content-specific characteristics and information relating to submission, journal transfer options,&nbsp;and external discoverability<br> Table 3: Screening, moderation, and permanence of content<br> Table 4: Usage metrics and other features<br> Table 5: Metadata<br> Preprint platform websites</p> <p>Data for each platform are listed as &lsquo;Verified&rsquo; in the tables if these tables (V1.0 or V2.0) were seen and approved by a platform representative between January 13 and January 27, 2020.</p> <p><strong>Original online survey:</strong>&nbsp;a blank copy of the original survey form used by online researchers (the authors) and supplied pre-filled (or empty, in some cases) to preprint platform representatives for verification (or completion, in some cases).&nbsp;</p> <p><strong>Final data:</strong>&nbsp;survey data is presented in .txt and .xlsx, as follows:</p> <ul> <li>Row 1: Heading (where field is included in manuscript tables, the heading presented here replaces any heading used in original survey. All columns are presented in the order the information was requested on the original survey form, with some supplementary columns added and columns removed (detailed below).</li> <li>Row 2: Schema or description of field</li> <li>Row 3: Whether and where included in manuscript tables. For supporting information for table data (e.g. source information, URLs), the table location for supported data is indicated in brackets, e.g. (Table 2) and supporting information is not included in tables. Data included in manuscript tables is presented in its final form, which in some cases is simplified from the original survey data. This simplified version of the data was presented to platform representatives for additional verification (v1.0/v2.0 verification). Data not included in manuscript tables is presented here as verified by platform representatives and/or found online. Some columns from the original survey have been removed due to the information not being informative or useful: specifically, Print ISSN (not reported for any platform); End date (no platforms have an end date; although two platforms stopped accepting submissions after survey completed; Personal contact information for platform representative(s) has been removed).</li> <li>Rows 4 onwards: data for each preprint platform (44 included in manuscript (rows 4-47), plus 5 additional OSF platforms (rows 48-52)</li> <li>Columns 3-6 (D-G) report online research and verification information and Column 13 (M) reports an additional data field (number of articles) &ndash; these are supplementary to the original survey columns</li> <li>Verification status: Released V1/V2 data applies to data included in manuscript tables (as indicated in row 3); Online survey data applies to data used for manuscript tables and also to original survey data included here but not included in manuscript tables (&lsquo;Not included&rsquo; in row 3)</li> <li>Note that data fields are presented as individual columns in these sheets, while some entries in Tables 1-5 combine several data fields.</li> </ul> <p>These data were collected in collaboration and as part of:<br> i. An ASAPbio project, led by Dr Naomi Penfold, to develop an online directory of preprint platforms<br> ii. A research project led by Prof&nbsp;Jamie Kirkham<br> These data are supplementary outputs for both projects.</p> <p>Data v1.0 were presented during the ASAPbio January 2020 workshop &ndash; see Penfold, Naomi C, &amp; Polka, Jessica. (2020, January). ASAPbio Preprint Platform Directory: 2019 data (presentation) (Version 1.0). Zenodo. http://doi.org/10.5281/zenodo.3626770.<br> <br> <strong>Version 3.0 updates (December 14, 2020): added new files with updated information about servers from the ASAPbio preprint directory (https://asapbio.org/preprint-servers), provided by Jessica Polka (now included as author).</strong></p>

opencc-zeroJan 2019View details →
zenodo44/100

Link-prediction on Biomedical Knowledge Graphs

<p>Release of code and experimental data from the paper <em>Towards Linking Graph Topology to Model Performance for Biomedical Knowledge Graph Completion&nbsp;</em>(<em>Machine Learning for Life and Material Sciences</em> workshop @ ICML2024) and <a href="https://arxiv.org/abs/2409.04103" rel="nofollow">The Role of Graph Topology in the Performance of Biomedical Knowledge Graph Completion Models</a>.</p> <div> <div>Knowledge Graph Completion has been increasingly adopted as a useful method for several tasks in biomedical research, like drug repurposing or drug-target identification.&nbsp;To that end, a variety of datasets and Knowledge Graph Embedding models has been proposed over the years. However, little is known about the properties that render a dataset useful for a given task and, even though theoretical properties of Knowledge Graph Embedding models are well understood, their practical utility in this field remains controversial. We conduct a comprehensive investigation into the topological properties of publicly available biomedical Knowledge Graphs and establish links to the accuracy observed in real-world applications. By releasing all model predictions we invite the community to build upon our work and continue improving the understanding of these crucial applications.</div> <div>&nbsp;</div> <div>Experiments were conducted on six datasets: five from the biomedical domain (<a href="../records/268568">Hetionet</a>, <a href="https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/IXA7BM">PrimeKG</a>, <a href="../records/4077338">PharmKG</a>, <a href="../records/5361324">OpenBioLink2020 HQ</a>, <a href="../records/7011027">PharMeBINet</a>) and one trivia KG (<a href="https://aclanthology.org/W15-4007.pdf">FB15k-237</a>). All datasets were randomly split into training, validation and test set (80% / 10% / 10%; in the case of PharMeBINet, 99.3% / 0.35% / 0.35% to mitigate the increased inference cost on the larger dataset).</div> <div>On each dataset, five different KGE models were compared:&nbsp;<a href="https://dl.acm.org/doi/10.5555/2999792.2999923">TransE</a>, <a href="https://arxiv.org/abs/1412.6575">DistMult</a>, <a href="https://arxiv.org/abs/1902.10197">RotatE</a>, <a href="https://arxiv.org/abs/2209.08271">TripleRE</a>, <a href="https://dl.acm.org/doi/10.5555/3504035.3504256">ConvE</a>. Hyperparameters were tuned on the validation split (see final train configurations in <code>train/scripts</code>). We release results for tail predictions on the test split. In particular, each test query&nbsp;<code>(h,r,?)</code> is scored against all entities in the KG and we compute the rank of the score of the correct completion <code>(h,r,t)</code> , after masking out scores of other <code>(h,r,t')</code> triples contained in the graph.</div> <div>Note: the ranks provided are computed as the average between the optimistic and pessimistic ranks of triple scores.</div> <div>&nbsp;</div> <div>Inside <code>experimental_data.zip</code>, the following files are provided.</div> <div> <ul> <li><code>datasets/{dataset}</code>: a folder for each dataset, containing <ul> <li><code>{dataset}_preprocessing.ipynb</code>: a Jupyter notebook for downloading and preprocessing the datasets. In particular, this generates the custom label-&gt;ID mapping for entities and relations, and the numerical tensor of&nbsp;<code>(h_ID,r_ID,t_ID)</code> triples for all edges in the graph, which can be used to compute graph topological metrics (e.g., using <a href="https://github.com/graphcore-research/kg-topology-toolbox">kg-topology-toolbox</a>)&nbsp; and compare them with the edge prediction accuracy.</li> <li><code>test_ranks.csv</code>: csv table with columns <code>["h", "r", "t"]</code> specifying the head, relation, tail IDs of the test triples, and columns <code>["DistMult", "TransE", "RotatE", "TripleRE", "ConvE"]</code> with the rank of the ground-truth tail in the ordered list of predictions made by the five KGE models;</li> <li><code>entity_dict.csv</code>: list of entity labels, ordered by entity ID (as generated in the preprocessing notebook);</li> <li><code>relation_dict.csv</code>: list of relation labels, ordered by relation ID (as generated in the preprocessing notebook).</li> </ul> </li> <li><code>train</code>: code to reproduce training (and validation) of the five KGE models, using the <a href="https://github.com/graphcore-research/bess-kge">BESS-KGE</a> distribution framework. <ul> <li><code>train/scripts</code>: executable scripts, with specifications of the final hyperparameters for all models and datasets.</li> </ul> </li> <li><code>notebooks</code>: Jupyter notebooks for data analysis and generation of all the figures in the paper.</li> </ul> <p>The separate <code>top_100_tail_predictions.zip</code> archive contains, for each of the test queries in the corresponding <code>test_ranks.csv</code> table, the IDs of the top-100 tail predictions made by each of the five KGE models, ordered by decreasing likelihood. The predictions are released in a <code>.npz</code>&nbsp;archive of numpy arrays (one array of shape <code>(n_test_triples, 100)</code> for each of the KGE models).&nbsp;</p> </div> </div>

openmit-licenseJun 2024View details →
zenodo44/100

RRID Gold Set of annotations for software tools in the biomedical literature

<p>This data is a subset of a larger human curated, machine assisted gold standard data set of RRID citations within the text of the scientific literature. The full set is accessible via Hypothes.is at https://hypothes.is/users/scibot?q=group%3A__world__ &nbsp;and via individual RRID records such as RRID:SCR_016250 https://scicrunch.org/resolver/SCR_016250/mentions?q=&amp;i=rrid:scr_016250&nbsp;</p><p>The data is based on authors who added RRIDs into their manuscripts. The data was then extracted by SciBot (RRID:SCR_016250), into Hypothes.is (RRID:SCR_000430) and then manually checked by a curator to determine if the author and the database agreed. The list of annotators is available in Hypothes.is user group: SciBotCurationGroup.</p><p>There are 78,140 rows and each row contains an annotation (annotation id, URI), linked to a paper (paper identifiers: PMID, DOI, PMC) and linked to the RRID (scr_id, exact, text_quote_selector).&nbsp;</p><p>A second spreadsheet contains a list of 8,322 software tools from the SciCrunch Registry (available here https://scicrunch.org/resources/data/source/nlx_144509-1/search ), enhanced by additions by thousands of individual authors, and curated over 10 years (Ozyurt et al., 2016).&nbsp;</p><p>The third spreadsheet contains a data dictionary and links to related ontologies, and tagging sets.&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

A high-throughput 3D X-ray histology facility for biomedical research and preclinical applications - Supplementary Data

<p><strong>Videos</strong></p><ul><li><strong>Video 1</strong> A video going through the Z stack in single slices. This is a cross- sectional view of the XRH image stack along the XY plane. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li><li><strong>Video 2 </strong>A video going through the Y stack in single slices. This is a cross- sectional view of the XRH image stack along the XZ plane. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li><li><strong>Video 3 </strong>A video going through the X stack in single slices. This is a cross- sectional view of the XRH image stack along the YZ plane. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li><li><strong>Video 4 </strong>3D X-ray histology (XRH) is a µCT -based workflow tailored to fit seamlessly into current histology workflows in biomedical and pre-clinical research, as well as clinical histopathology. Microanatomical detail can be captured from standard (non-stained) formalin-fixed and paraffin-embedded (FFPE) tissue blocks.</li><li><strong>Video 5</strong> Average Intensity Projection (AIP) of the sample through the Histologically relevant plane. This is a 2D visualisation rendering the Average Intensity of 20x single XY slices along the z-axis of the stack. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li><li><strong>Video 6 </strong>Maximum Intensity Projection (MIP) of the sample through the Histologically relevant plane. This is a 2D visualisation rendering the Maximum Intensity of 20x single XY slices along the z-axis of the stack. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li><li><strong>Video 7 </strong>Standard deviation projection of the sample going through the histologically relevant plane. This is a 2D visualisation rendering the Standard Deviation of 20x single XY slices along the z- axis of the stack. XRH datasets are normally oriented (resliced) in a way that a scroll through the stack along the XY plane emulates the physical histology slicing of the tissue.</li></ul><p><i>* <strong>Videos 5 -7</strong> are also referred to as "thick-slice rolls" </i>-&nbsp;<i>Thick-slice rolling is a 2D thick-slice viewing that allows rolling of a pre-selected number of slices (n) along the z-axis of the 3D data. A single thick-slice roll forwards is accomplished by translating the thick-slice by one single slice forwards; that is moving forward by one (+1) slice from the first and nth element and reapplying the criteria or operations to the new slice sub-stack.</i><br>&nbsp;</p><p><strong>The questionnaire used to collect feedback about the needs of the XRH community.</strong></p><ul><li>Survey.docx</li><li>Survey.pdf</li></ul><p><br><strong>Exemplar report of a semi-automatically generated augmented PDF file</strong> that contain sample information, imaging settings, still images with descriptive figure legends, and links to corresponding online videos</p><ul><li>DEMO02019-FFPE_report_99EbPXG.pdf</li></ul><p>&nbsp;</p><p>= = = = = = = = = = = = = = = =&nbsp;<br><strong>System performance data ZIP</strong><br>= = = = = = = = = = = = = = = = &nbsp;</p><p>This ZIP file contains imaging data collected through different systems and setups at the XRH facility at the μ-VIS X-ray Imaging Centre at the University of Southampton for the purpose of acceptance and/or system performance characterisation. Below is an overview of the folder structure and its contents</p><p>The following files are X-ray imaging data collected on September 28, 2017, using the Med-X system and a Jima phantom at 55 kV peak and 7 Watts.&nbsp;</p><ul><li>20170928_MEDX_1642_JIMA_55kVp7W-2.tif</li><li>20170928_MEDX_1642_JIMA_55kVp7W.tif</li><li>20170928_MEDX_1642_JIMA_55kVp7W.tif.profile.xml</li></ul><p>This PDF document is related to a QRM MicroCT bar pattern phantom, and its specifications</p><ul><li>QRM-MicroCT-Barpattern-Phantom.pdf</li></ul><p>Graphs showing the calculated focal-spot size as a function of the X-ray power (W) for the Molybdenum rotating target calculated using Edge Modulation function testing. The performance is then compared with the performance of the Reflection target across the same range of powers. Raw data can be found in XRH_QRM_Refl-vs-Rot-TargetComparison_SingleReconSlices_5umPixelSize folder. Test performed in July 2021. &nbsp;</p><ul><li>XRH_202107_MoRot-testing_EdgeModFunction-QRMrecons+RotReflCompar.png</li></ul><p>&nbsp;</p><p><i><strong>/ XRH-XT-H-225-ST_FocalSpots</strong></i><br>This directory contains radiographic data collected using the XRH system with a JIMA phantom and MoRt (Molybdenum rotating), TT (Transmission), and Reflection targets.</p><ul><li>20200113_XRH_Jima test MoRT 55kV 15W.tif, 20200113_XRH_Jima test MoRT 55kV 30W.tif, etc.:&nbsp;<br>These files represent radiographs taken on January 13, 2020, using the XRH system, Jima phantom, MoRT target at 55 kVp and varying wattages.</li><li>20200207_XRH_JIMA 80kV TT1a.tif, 20200207_XRH_JIMA 80kV TT1b.tif, etc.<br>Similar to the above, these files are from February 7, 2020, and use 80 kVp with a TT target.</li><li>20231115_XRH_reflW_80kVp6W.tif, 20231115_XRH_reflW_80kVp6W_02.tif, etc.<br>These files are from November 15, 2023, and collected using the XRH system with a Reflection target at 80 kVp and 6 Watts.</li></ul><p><i><strong>/ XRH_QRM_Refl-vs-Rot-TargetComparison_SingleRadioFromCTs_5umPixelSize</strong></i><br>This directory contains single radiographs taken with a pixel size of 5 micrometers using the Molybdenum rotating (MoRt), and the Reflection target using tungsten (W) and Molybdenum (Mo) metals.</p><p><i><strong>/ XRH_QRM_Refl-vs-Rot-TargetComparison_SingleReconSlices_5umPixelSize</strong></i><br>This directory contains sinlge reconstruction slices of the setups mentioned above. Slices are exported from CT volumes and were used for the Edge Modulation function study. &nbsp;</p><p>For interpretation of the filenames in the folders listed above please see below and refer to specific files and folders for detailed information and results related to each imaging session:</p><ul><li><i>&lt;xx&gt;kVp or &lt;xx&gt;kV &nbsp;&nbsp;</i>:Imaging at a peak voltage of &lt;xx&gt; kVp.</li><li><i>&lt;y&gt;W</i> &nbsp; :Imaging at &lt;y&gt; Watts;<i>&nbsp; </i>"." is represented with "-"; i.e. 20210705_XRH_2766_PJB_TEST03552-EQPMT_W_6-9W is acquired using a power of 6.9 W</li><li><i>MoRt, TT, Refl&nbsp;</i> &nbsp;:Molybdenum, Transmission, and Reflection targets, respectively.</li><li><i>_W_ and _Mo_&nbsp;</i> &nbsp;:Tungsten and Molybdenum target materials.</li><li><i>_horiz</i> &nbsp; :Reconstruction slices in line with the X-ray beam's propagation direction.</li><li><i>_vert</i> &nbsp; :Reconstruction slices normal to the X-ray beam's propagation direction and parallel to the detector plane.</li></ul>

opencc-by-4.0Jun 2023View details →
zenodo44/100

[MedMNIST+] 18x Standardized Datasets for 2D and 3D Biomedical Image Classification with Multiple Size Options: 28 (MNIST-Like), 64, 128, and 224

<h2><strong>Code</strong>&nbsp;[<a href="https://github.com/MedMNIST/MedMNIST" target="_blank" rel="noopener">GitHub</a>]&nbsp;| <strong>Publication</strong>&nbsp;[<a href="https://doi.org/10.1038/s41597-022-01721-8" target="_blank" rel="noopener">Nature Scientific Data'23</a>&nbsp;/&nbsp;<a href="https://doi.org/10.1109/ISBI48211.2021.9434062" target="_blank" rel="noopener">ISBI'21</a>]&nbsp;| <strong>Preprint</strong>&nbsp;[<a href="https://arxiv.org/abs/2110.14795" target="_blank" rel="noopener">arXiv</a>]</h2> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>We introduce MedMNIST, a large-scale MNIST-like collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D. All images are pre-processed into 28x28 (2D) or 28x28x28 (3D) with the corresponding classification labels, so that no background knowledge is required for users. Covering primary data modalities in biomedical images, MedMNIST is designed to perform classification on lightweight 2D and 3D images with various data scales (from 100 to 100,000) and diverse tasks (binary/multi-class, ordinal regression and multi-label). The resulting dataset, consisting of approximately 708K 2D images and 10K 3D images in total, could support numerous research and educational purposes in biomedical image analysis, computer vision and machine learning. We benchmark several baseline methods on MedMNIST, including 2D / 3D neural networks and open-source / commercial AutoML tools. The data and code are publicly available at&nbsp;<a href="https://medmnist.com/">https://medmnist.com/</a>.</p> <p><em><strong>Disclaimer</strong></em>: The only official distribution link for the MedMNIST dataset is&nbsp;<a href="https://doi.org/10.5281/zenodo.10519652">Zenodo</a>. We kindly request users to refer to this original dataset link for accurate and up-to-date data.</p> <p><strong><em>Update</em>:</strong> We are thrilled to release&nbsp;<a href="https://github.com/MedMNIST/MedMNIST/blob/main/on_medmnist_plus.md">MedMNIST+</a> with larger sizes: 64x64, 128x128, and 224x224 for 2D, and 64x64x64 for 3D. As a complement to the previous 28-size MedMNIST, the large-size version could serve as a standardized benchmark for medical foundation models. Install the latest API to try it out!</p> <p>&nbsp;</p> <p><strong>Python Usage</strong></p> <p>We recommend our official <a href="https://github.com/MedMNIST/MedMNIST">code</a> to download, parse and use&nbsp;the MedMNIST dataset:</p> <blockquote> <pre>% pip install medmnist<br>% python</pre> <div> <div>To use the standard 28-size (MNIST-like) version utilizing the downloaded files:</div> <br> <div>&gt;&gt;&gt; from medmnist import PathMNIST</div> <div>&gt;&gt;&gt; train_dataset = PathMNIST(split="train")</div> <br> <div>To enable automatic downloading by setting `download=True`:</div> <br> <div>&gt;&gt;&gt; from medmnist import NoduleMNIST3D</div> <div>&gt;&gt;&gt; val_dataset = NoduleMNIST3D(split="val", download=True)</div> <br> <div>Alternatively, you can access MedMNIST+ with larger image sizes by specifying the `size` parameter:</div> <br> <div>&gt;&gt;&gt; from medmnist import ChestMNIST</div> <div>&gt;&gt;&gt; test_dataset = ChestMNIST(split="test", download=True, size=224)</div> </div> </blockquote> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you find this project useful, please cite both v1 and v2 paper as:</p> <blockquote> <p>Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, Bingbing Ni. Yang, Jiancheng, et al. "MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification." Scientific Data, 2023.</p> <p>Jiancheng Yang, Rui Shi, Bingbing Ni. "MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis". IEEE 18th International Symposium on Biomedical Imaging (ISBI), 2021.</p> </blockquote> <p>or using bibtex:</p> <blockquote> <pre>@article{medmnistv2, title={MedMNIST v2-A large-scale lightweight benchmark for 2D and 3D biomedical image classification}, author={Yang, Jiancheng and Shi, Rui and Wei, Donglai and Liu, Zequan and Zhao, Lin and Ke, Bilian and Pfister, Hanspeter and Ni, Bingbing}, journal={Scientific Data}, volume={10}, number={1}, pages={41}, year={2023}, publisher={Nature Publishing Group UK London} } @inproceedings{medmnistv1, title={MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis}, author={Yang, Jiancheng and Shi, Rui and Ni, Bingbing}, booktitle={IEEE 18th International Symposium on Biomedical Imaging (ISBI)}, pages={191--195}, year={2021} }</pre> </blockquote> <p>Please also cite the corresponding paper(s) of source data if you use any subset of MedMNIST&nbsp;as per the description on the&nbsp;<a href="https://medmnist.github.io/">project website</a>.</p> <p>&nbsp;</p> <p><strong>License</strong></p> <p>The MedMNIST dataset is licensed under&nbsp;<em>Creative Commons Attribution 4.0 International</em>&nbsp;(<a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>), except DermaMNIST under&nbsp;<em>Creative Commons Attribution-NonCommercial 4.0 International</em>&nbsp;(<a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC 4.0</a>).</p> <p>The code is under&nbsp;<a href="https://github.com/MedMNIST/MedMNIST/blob/main/LICENSE">Apache-2.0 License</a>.</p> <p>&nbsp;</p> <p><strong>Changelog</strong></p> <p><a href="https://doi.org/10.5281/zenodo.10519652">v3.0</a> (this repository): Released MedMNIST+ featuring larger sizes: 64x64, 128x128, and 224x224 for 2D, and 64x64x64 for 3D.</p> <p><a href="https://doi.org/10.5281/zenodo.10519195">v2.2</a>: Removed a small number of mistakenly included blank samples in OrganAMNIST, OrganCMNIST, OrganSMNIST, OrganMNIST3D, and VesselMNIST3D.&nbsp;</p> <p><a href="https://doi.org/10.5281/zenodo.6496656">v2.1</a>: Addressed an issue in the NoduleMNIST3D file (i.e., nodulemnist3d.npz). Further details can be found in this <a href="https://github.com/MedMNIST/MedMNIST/issues/22#issuecomment-1103438191">issue</a>.</p> <p><a href="https://doi.org/10.5281/zenodo.5208230">v2.0</a>: Launched the initial repository of MedMNIST v2, adding 6 datasets for 3D and 2 for 2D.</p> <p><a href="https://doi.org/10.5281/zenodo.4269852">v1.0</a>: Established the initial repository (in a separate repository) of MedMNIST v1, featuring 10 datasets for 2D.</p> <p>&nbsp;</p> <p><strong>Note</strong>: This dataset is&nbsp;<strong>NOT</strong> intended for clinical use.</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting

<p><span><span><span><span>We demonstrate a resilient workflow enabled by the LEXIS Platform, running a time- and safety-critical biomedical simulation of virtual stent placement in intracranial arteries using the HemoFlow application. The workflow, as captured on the video, gracefully handles failures of single computing steps or entire computing systems and thus lends itself to urgent computing applications. <br><br><span><span>The concept of this workflow has potential for realising ab-initio computational biomedical simulations which can provide live, targeted guidance to surgeons.</span></span></span></span></span></span></p>

opencc-by-nc-nd-4.0Apr 2024View details →
zenodo44/100

Additional Artifacts - Supplements to: A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting

<p>In this dataset, we have collected supplementary artifacts to support an understanding of the workflow presented in the submission cited (see related identifiers).</p> <p>These artifacts are (cf. README.md in the main folder of the tar.gz archive):</p> <p>A1: modified HemoFlow code (cf. https://github.com/gzavo/hemoflow) for our workflow experiments (subfolder "hemoflowcfd");<br>A2: workflow descriptions in python for Apache Airflow (subfolder "workflow");<br>A3: inputs (.xml/.npz) and output (.txt) for the example (subfolder "case").</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

TBGA: A Large-Scale Gene-Disease Association Dataset for Biomedical Relation Extraction

<p>This repository contains the TBGA dataset. TBGA is a large-scale, semi-automatically annotated&nbsp;dataset&nbsp;for Gene-Disease Association (GDA) extraction. The dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files&nbsp;corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong>&nbsp;sentence from which the GDA was extracted.</li> <li><strong>relation:</strong>&nbsp;relation name associated with the given GDA.</li> <li><strong>h:&nbsp;</strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id:&nbsp;</strong>NCBI Entrez ID associated with the gene entity.</li> <li><strong>name:</strong>&nbsp;NCBI official gene symbol associated with&nbsp;the gene entity.</li> <li><strong>pos:&nbsp;</strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong>&nbsp;JSON object representing the disease entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated with the disease entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred term associated with the disease entity.</li> <li><strong>pos:</strong>&nbsp;list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>TBGA contains over 200,000 instances and 100,000 bags.<br> The zip file consists of one folder, named TBGA,&nbsp;containing the files corresponding to the dataset.</p> <p>If you use or extend our work, please cite the following:&nbsp;https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-022-04646-6#citeas<br> TBGA paper can be found at:&nbsp;<a href="https://rdcu.be/cKkY2">https://rdcu.be/cKkY2</a><br> TBGA code is available at:&nbsp;https://github.com/GDAMining/gda-extraction</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Spanish Biomedical Crawled Corpus

<p>The largest Spanish biomedical and heath corpus to date gathered from a massive Spanish health domain crawler over more than 3,000 URLs were downloaded and preprocessed. All the collected data have been preprocessed to produce the CoWeSe (Corpus Web Salud Espa&ntilde;ol) resource, a large-scale and high-quality corpus intended for biomedical and health NLP in Spanish.</p> <p>Enlarged version with less restrictive document and sentence deduplication.</p> <p><strong>Citation</strong></p> <p>If you use this resource in your work, please cite our paper:</p> <pre>@misc{carrino2021spanish, title={Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models}, author={Casimiro Pio Carrino and Jordi Armengol-Estap&eacute; and Ona de Gibert Bonet and Asier Guti&eacute;rrez-Fandi&ntilde;o and Aitor Gonzalez-Agirre and Martin Krallinger and Marta Villegas}, year={2021}, eprint={2109.07765}, archivePrefix={arXiv}, primaryClass={cs.CL} } </pre> <p>Copyright (c) 2022 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Preprints in biology as a fraction of the biomedical literature

<p>These data and chart present an approximate&nbsp;calculation of the proportion of preprints in biology when compared to publications in PubMed, based on monthly figures and incorporating monthly preprint submissions (or counts) across a selection of servers relevant to biology.&nbsp;</p> <p>Version 1.0 of these data&nbsp;represents data from January 2007 until May 31, 2019 for preprint servers: arXiv q-bio, Nature Precedings, F1000Research*, PeerJ Preprints*, bioRxiv**, Winnower*,&nbsp;<a href="http://preprints.org">preprints.org</a>, Wellcome Open Research*.&nbsp;</p> <p>* Counts may not be specific to biology preprints only; ** Counts may include all versions posted that month, so may be an overestimate for version 1 submissions.</p> <p>From January 2019, data has been gathered manually by the authors, as per the methods described in the .csv here, and is included here in &#39;Preprints_per_month_direct_2019-01to05.csv&#39;.&nbsp;Until December 2018, monthly preprint submissions data are based on those contributed by Jordan Anaya (ORCID: <a href="https://orcid.org/0000-0002-6166-4113">https://orcid.org/0000-0002-6166-4113</a>) for PrePubMed, source:&nbsp;<a href="https://raw.githubusercontent.com/OmnesRes/prepub/master/analyses/preprint_data.txt">https://raw.githubusercontent.com/OmnesRes/prepub/master/analyses/preprint_data.txt</a>; Github repository:&nbsp;<a href="https://github.com/OmnesRes/prepub">https://github.com/OmnesRes/prepub</a>; website: <a href="http://www.prepubmed.org/">http://www.prepubmed.org</a>). Data are not included here, they are&nbsp;provided from the source linked above under MIT license associated with the website code: <a href="https://github.com/OmnesRes/prepub/blob/master/LICENSE">https://github.com/OmnesRes/prepub/blob/master/LICENSE</a>.</p> <p>A live version of these data and the chart are available from this GSheet:&nbsp;<a href="https://docs.google.com/spreadsheets/d/1bkGEcfQcL0LpIanVqNHci1ZFY6oVNGz7IQbEugzkqkU/edit?usp=sharing">https://docs.google.com/spreadsheets/d/1bkGEcfQcL0LpIanVqNHci1ZFY6oVNGz7IQbEugzkqkU/edit?usp=sharing</a>. Between version updates here, please refer to this sheet for updated counts and method updates e.g.&nbsp;to include more servers and ensure only version 1 submissions are counted.</p> <p>For more information, please contact naomi.penfold@asapbio.org.</p> <p>When presenting these data and/or chart, please attribute to ASAPbio (https://asapbio.org, twitter: @ASAPbio_).</p>

opencc-zeroJun 2019View details →
zenodo44/100

Extended data of the project "A survey exploring biomedical editors' perceptions of editorial interventions to improve adherence to reporting guidelines"

<p>Figure S1: Survey questionnaire</p> <p>Table S2:&nbsp;Barriers, facilitators and possible improvements of the&nbsp;interventions included in the survey</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

Building Large-Scale Gene-Disease Association Datasets for Biomedical Relation Extraction

<p>This repository contains the GDAb and GDAt datasets. GDAb and GDAt are large-scale, distantly supervised, and manually enhanced datasets for Gene-Disease Association (GDA) extraction. Each dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files&nbsp;corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong>&nbsp;sentence from which the GDA was extracted.</li> <li><strong>relation:</strong>&nbsp;relation name associated to the given GDA.</li> <li><strong>h:&nbsp;</strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated to the gene entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred name associated to the gene entity.</li> <li><strong>pos:&nbsp;</strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong>&nbsp;JSON object representing the disease entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated to the disease entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred name associated to the disease entity.</li> <li><strong>pos:</strong>&nbsp;list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>Both datasets contain over 2,500,000 sentences and 500,000 bags.<br> The zip file consists of two folders, GDAb and GDAt,&nbsp;&nbsp;containing the files corresponding to the two datasets, respectively.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

A study on biomedical researchers' perspectives on public engagement in Southeast Asia

<p>Survey data from biomedical researchers in Southeast Asia about their perceptions of public engagement. The survey used open and closed questions.</p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record