Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

195

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

195 results for “biomedical”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset: Talis Biomedical Corporation (TLIS) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Psyence Biomedical Ltd. (PBM) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Psyence Biomedical Ltd. (PBMWW) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Ocean Biomedical, Inc. (OCEAW) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Ocean Biomedical, Inc. (OCEA) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Lifecore Biomedical, Inc. (LFCR) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Multicontextual Phenotype Models: Biomedical Database and Literature Phenotype and Social Media Phenotype

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo40/100

Data from the paper "The landscape of biomedical research"

<p>Data from the paper "<a href="https://www.cell.com/patterns/fulltext/S2666-3899(24)00076-X">The landscape of biomedical research</a>".</p> <p>The paper used the PubMed 2020 baseline (download date: 26.01.2021, not available anymore) supplemented with additional files from the 2021 baseline (download date: 27.04.2022, not available anymore), both originally obtained from <a href="https://www.nlm.nih.gov/databases/download/pubmed_medline.html">https://www.nlm.nih.gov/databases/download/pubmed_medline.html</a>, courtesy of the U.S. National Library of Medicine. This data can be found in v2 of this repository (<a href="../records/7849020">https://zenodo.org/records/7849020</a>).</p> <p>In the latest version of this repository we provide the PubMed 2024 baseline (download date: 06.02.2024) including all papers until the end of 2023, which is&nbsp;<strong>not</strong> the main data we analyzed in the paper but an updated version including newer articles. The paper contains two supplementary figures (S9 and S10) with the updated embedding.</p> <p>The latest version provided here includes the following files:</p> <p>pubmed_landscape_data_2024_v2.zip, which includes:</p> <p>- from the PubMed database: article title, journal, PMID, and publication year.</p> <p>- produced by us: t-SNE embedding X and Y coordinates, label, color, whether the paper is retracted or not (combining PubMed and Retraction Watch information), affiliation country ( from the first affiliation of the first author), and inferred gender (of both first and last author).</p> <p>(Note: pubmed_landscape_data_2024_v2.zip is identical to pubmed_landscape_data_2024.zip from v3 of this repository, but includes inferred genders additionally.)</p> <p>&nbsp;</p> <p>pubmed_landscape_abstracts_2024.zip, which includes:</p> <p>- from the PubMed database: PMID, and paper abstracts.</p> <p>&nbsp;</p> <p>PubMedBERT_embeddings_float16_2024.npy, which includes:</p> <p>- produced by us: PubMedBERT embeddings of the paper abstracts (numpy.ndarray of shape&nbsp;23,389,083x768).</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Bio-ML: Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching

<p>&nbsp;</p> <blockquote> <p><strong>This version is used in the Bio-ML track of the OAEI 2024; the only change compared to the OAEI 2023 is the deletion of certain training subsumption mappings.</strong></p> </blockquote> <p>&nbsp;</p> <h3><strong>Overview</strong></h3> <p>The purpose of these datasets is to support&nbsp;<em>equivalence</em> and <em>subsumption</em> ontology matching.</p> <p>There are five ontology pairs extracted from MONDO and UMLS:</p> <table> <tbody> <tr> <td>Source</td> <td>Task</td> <td>Category</td> <td>#SrcCls</td> <td>#TgtCls</td> <td>#Ref (equiv)</td> <td>#Ref (subs)</td> </tr> <tr> <td>Mondo</td> <td>OMIM-ORDO</td> <td>Disease</td> <td>9,648</td> <td>9,275</td> <td>3,721</td> <td>103</td> </tr> <tr> <td>Mondo</td> <td>NCIT-DOID</td> <td>Disease</td> <td>15,762</td> <td>8,465</td> <td>4,686</td> <td>3,338 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-FMA</td> <td>Body</td> <td>34,418</td> <td>88,955</td> <td>7,256</td> <td>5,453 (-53)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Pharm</td> <td>29,500</td> <td>22,136</td> <td>5,803</td> <td>4,224 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Neoplas</td> <td>22,971</td> <td>20,247</td> <td>3,804</td> <td>213</td> </tr> </tbody> </table> <p>The "-" numbers reflect the changes due to lthe deletion of certain training subsumption mappings.</p> <p>The main track is available at "bio-ml", where each pair is associated with a task folder, containing the source and target ontologies, reference equivalence mappings (in "refs_equiv"), reference subsumption mappings ("refs_subs").&nbsp;</p> <p>The special sub-track is available at "bio-llm", where each pair is associated with a task folder, containing the source and target ontologies, and the test candidate mappings.&nbsp;</p> <p>&nbsp;</p> <h3><strong>Citation</strong></h3> <p><strong>Bio-ML (Main Track)</strong></p> <pre>```<br>@inproceedings{he2022machine, title={Machine learning-friendly biomedical datasets for equivalence and subsumption ontology matching}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Jim{\'e}nez-Ruiz, Ernesto and Hadian, Ali and Horrocks, Ian}, booktitle={International Semantic Web Conference}, pages={575--591}, year={2022}, organization={Springer} }<br>```</pre> <p><strong>Bio-LLM (Sub-track)</strong></p> <pre>```<br>@article{he2023exploring, title={Exploring large language models for ontology alignment}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Horrocks, Ian}, journal={arXiv preprint arXiv:2309.07172}, year={2023} }<br>```</pre> <p>&nbsp;</p> <h3><strong>Important Links</strong></h3> <ul> <li>See detailed documentation at:&nbsp;<a href="https://krr-oxford.github.io/DeepOnto/bio-ml">https://krr-oxford.github.io/DeepOnto/bio-ml</a>.</li> <li>See the OAEI Bio-ML track at:&nbsp;<a href="https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/">https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/</a></li> <li>See our resource paper for the original Bio-ML at&nbsp;<a href="https://arxiv.org/abs/2205.03447">arxiv</a>&nbsp;or <a href="https://link.springer.com/chapter/10.1007/978-3-031-19433-7_33">springer</a>&nbsp;(accepted at&nbsp;<em>ISWC-2022</em> and nominated as the <em>best resource paper candidate</em>). See our poster paper for the Bio-LLM sub-track at&nbsp;<a href="https://arxiv.org/abs/2309.07172">arxiv </a>(accepted at <em>ISWC-2023 Posters &amp; Demos</em>).</li> </ul> <p>&nbsp;</p> <h3><strong>Changelog</strong></h3> <p>The only change in this version compared to the OAEI 2023 is the deletion of certain training subsumption mappings that can be directly exploited through deductive reasoning.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data associated with "A collaborative filtering based approach to biomedical knowledge discovery"

<p>This is the data set associated with the publication: &quot;A collaborative filtering based approach to biomedical knowledge discovery&quot; published in Bioinformatics.</p> <p>The data are sets of cooccurrences of biomedical terms extracted from published abstracts and full text articles. The cooccurrences are then represented in sparse matrix form. There are three different splits of this data denoted by the prefix number on the files.</p> <p>1. All - All cooccurrences combined in a single file</p> <p>2. Training/Validation - All cooccurrences in publications before 2010 in training, all novel cooccurrences in publication in 2010 go in validation</p> <p>3. Training+Validation/Test - All cooccurrences in publication upto and including 2010 in training+validation. All novel cooccurrences after 2010 in year by year increments and also all combined together</p> <p>&nbsp;</p> <p>Furthermore there are subset files which are used in some experiments to deal with the computational cost of evaluating the full set. The associated cuids.txt file containing a link between the row/column in the matrix with the UMLS Metathesaurus CUIDs. Hence the first row of cuids.txt matches up to the 0th row/column in the matrix. Note that the matrix is square and symmetric. This work was done with UMLS Metathesaurus 2016AB.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Neural Machine Translation for the Biomedical Domain - WMT19

<p>This package contains the files needed to use the Neural Machine Translation (NMT) system for the Biomedical Domain.</p> <p>The available language directions for translation are:</p> <ul> <li>English to Spanish</li> <li>Spanish to English</li> <li>English to Portuguese</li> <li>Portuguese to English</li> <li>Spanish to Portuguese</li> <li>Portuguese to Spanish</li> </ul> <p>The code&nbsp;for using the translation files is in&nbsp;<a href="https://github.com/PlanTL-SANIDAD/Medical-Translator-WMT19">https://github.com/PlanTL-SANIDAD/Medical-Translator-WMT19</a></p> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

BVS Corpus: A Multilingual Parallel Corpus and Translation Experiments of Biomedical Scientific Texts

<p>The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME in&nbsp;agreement with the Pan American Health Organization (OPAS). Abstracts are available in English, Spanish, and Portuguese, with a subset in more than one language, thus being a possible source of parallel corpora. In this article, we present the development of parallel corpora from BVS in three languages: English, Portuguese, and Spanish. Sentences were automatically aligned using the Hunalign algorithm for EN/ES and EN/PT language pairs, and for a subset of trilingual articles also. We demonstrate the capabilities of our corpus by training a Neural Machine Translation (OpenNMT) system for each language pair, which outperformed related works on scientific biomedical articles. Sentence alignment was also manually evaluated, presenting an average 96\% of correctly aligned sentences across all languages. Our parallel corpus is freely available, with complementary information regarding article metadata.</p> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Underlying data of the project "A survey exploring biomedical editors' perceptions of editorial interventions to improve adherence to reporting guidelines"

<p><em><strong>Survey dataset.xlsx</strong></em>&nbsp;contains the anonymised&nbsp;responses to the survey.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

A global network of biomedical relationships derived from text

<p>This repository contains labeled, weighted networks of chemical-gene, gene-gene, gene-disease, and chemical-disease relationships based on single sentences in PubMed abstracts. All raw dependency paths are provided in addition to the labeled relationships.</p> <p>PART I: Connects dependency paths to labels, or &quot;themes&quot;. Each record contains a dependency path followed by its score for each theme, and indicators of whether or not the path is part of the flagship path set for each theme (meaning that it was manually reviewed and determined to reflect that theme). The themes themselves are listed below and are in our paper (reference below).</p> <p>PART II: Connects sentences to dependency paths. It consists of sentences and associated metadata, entity pairs found in the sentences, and dependency paths connecting those entity pairs. Each record contains the following information:</p> <ul> <li>PubMed ID</li> <li>Sentence number (0 = title)</li> <li>First entity name, formatted</li> <li>First entity name, location (characters from start of abstract)</li> <li>Second entity name, formatted</li> <li>Second entity name, location</li> <li>First entity name, raw string</li> <li>Second entity name, raw string</li> <li>First entity name, database ID(s)</li> <li>Second entity name, database ID(s)</li> <li>First entity type (Chemical, Gene, Disease)</li> <li>Second entity type (Chemical, Gene, Disease)</li> <li>Dependency path</li> <li>Sentence, tokenized</li> </ul> <p>The &quot;with-themes.txt&quot; files only contain dependency paths with corresponding theme assignments from Part I. The plain &quot;.txt&quot; files contain all dependency paths.</p> <p>This release contains the annotated network for the&nbsp;<strong>September 15, 2019&nbsp;version of PubTator</strong>. The version discussed in our paper, below, is an older one - from April 30, 2016. If you&#39;re interested in that network, it can be found in Version 1 of this repository.&nbsp;We will be releasing updated networks periodically, as the PubTator community continues to release new versions of named entity annotations for Medline each month or so.</p> <p>------------------------------------------------------------------------------------<br> REFERENCES</p> <p>Percha B, Altman RBA (2017) A global network of biomedical relationships derived from text. <em>Bioinformatics,&nbsp;</em>34(15): 2614-2624.<br> Percha B, Altman RBA (2015) Learning the structure of biomedical relationships from unstructured text. <em>PLoS Computational Biology,</em> 11(7): e1004216.</p> <p>This project depends on named entity annotations from the PubTator project:<br> https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/PubTator/</p> <p>Reference:<br> Wei CH et. al., PubTator: a Web-based text mining tool for assisting Biocuration, Nucleic acids research, 2013, 41 (W1): W518-W522.</p> <p>Dependency parsing was provided by the Stanford CoreNLP toolkit (<strong>version 3.9.1</strong>):<br> https://stanfordnlp.github.io/CoreNLP/index.html</p> <p>Reference:<br> Manning, Christopher D., Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP Natural Language Processing Toolkit In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 55-60.</p> <p>------------------------------------------------------------------------------------<br> THEMES</p> <p><strong>chemical-gene</strong><br> (A+) agonism, activation<br> (A-) antagonism, blocking<br> (B) binding, ligand (esp. receptors)<br> (E+) increases expression/production<br> (E-) decreases expression/production<br> (E) affects expression/production (neutral)<br> (N) inhibits</p> <p><strong>gene-chemical</strong><br> (O) transport, channels<br> (K) metabolism, pharmacokinetics<br> (Z) enzyme activity</p> <p><strong>chemical-disease</strong><br> (T) treatment/therapy (including investigatory)<br> (C) inhibits cell growth (esp. cancers)<br> (Sa) side effect/adverse event<br> (Pr) prevents, suppresses<br> (Pa) alleviates, reduces<br> (J) role in disease pathogenesis</p> <p><strong>disease-chemical</strong><br> (Mp) biomarkers (of disease progression)</p> <p><strong>gene-disease</strong><br> (U) causal mutations<br> (Ud) mutations affecting disease course<br> (D) drug targets<br> (J) role in pathogenesis<br> (Te) possible therapeutic effect<br> (Y) polymorphisms alter risk<br> (G) promotes progression</p> <p><strong>disease-gene</strong><br> (Md) biomarkers (diagnostic)<br> (X) overexpression in disease<br> (L) improper regulation linked to disease</p> <p><strong>gene-gene</strong><br> (B) binding, ligand (esp. receptors)<br> (W) enhances response<br> (V+) activates, stimulates<br> (E+) increases expression/production<br> (E) affects expression/production (neutral)<br> (I) signaling pathway<br> (H) same protein or complex<br> (Rg) regulation<br> (Q) production by cell population</p> <p>------------------------------------------------------------------------------------<br> FORMATTING NOTE</p> <p>A few users have mentioned that the dependency paths in the &quot;part-i&quot; files are all lowercase text, whereas those in the &quot;part-ii&quot; files maintain the case of the original sentence. This complicates mapping between the two sets of files.</p> <p>We kept the part-ii files in the same case as the original sentence to facilitate downstream debugging - it&#39;s easier to tell which words in a particular sentence are contributing to the dependency path if their original case is maintained. When working with the part-ii &quot;with-themes&quot; files, if you simply convert the dependency path to lowercase, it is guaranteed to match to one of the paths in the corresponding part-i file and you&#39;ll be able to get the theme scores.</p> <p>Apologies for the additional complexity, and please reach out to us if you have any questions (see correspondence information in the&nbsp;<em>Bioinformatics</em> manuscript, above).</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

3-Sulfopropyl acrylate potassium-based polyelectrolyte hydrogels: Sterilizable synthetic material for biomedical application

<p>Hydrogels are extensively used in the biomedical field due to their highly valued properties, biocompatibility and antimicrobial activity and resistance to rheological stress. However, determining an efficient sterilization protocol that does not compromise the functional properties of hydrogels is one of the challenges researchers face when developing a material for a medical application. In this work, conventional sterilization methods (steam-, radiation- and gas sterilization) were investigated regarding the influence on the degree of swelling, mechanical performance and chemical effects on the poly 3-sulfopropyl acrylate potassium (pAESO<sub>3</sub>) hydrogel, which is a promising representative for biomedical engineering applications. In summary, no significant changes in the gel properties were observed after sterilization, showing the potential of the selected hydrogel for biomedical applications.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Biomedical Data-to-Text Generation via Fine-Tuning Transformers

<p>Biomedical Dataset (&rdquo;BioLeaflets&rdquo;) for the paper &quot;Biomedical Data2Text Generation via fine-tuning transformers&quot; (INLG&#39;21)</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Biomedical Spanish CBOW Word Embeddings in Floret

<p><strong>Biomedical Spanish CBOW Word Embeddings in Floret</strong></p> <p>The embeddings have been trained with a biomedical Spanish corpus&nbsp;using&nbsp;<a href="https://arxiv.org/pdf/1607.04606.pdf">floret</a>&nbsp;with the following&nbsp;&nbsp;hyperparameters:</p> <blockquote> <p>mode: str = &quot;floret&quot;,<br> model: str = &quot;cbow&quot;,<br> dim: int = 300,<br> mincount: int = 10,<br> minn: int = 5,<br> maxn: int = 6,<br> neg: int = 10,<br> hashcount: int = 2,<br> bucket: int = 50000,<br> thread: int = 128,</p> </blockquote> <p>The embeddings were trained on the concatenation of all corpora from the <strong>Spanish biomedical corpus</strong>&nbsp;that includes Spanish data from various sources for a total of 1.1B tokens across 2,5M documents.</p> <table> <thead> <tr> <th scope="col">Source</th> <th scope="col">No. tokens</th> </tr> </thead> <tbody> <tr> <td>Medical crawler</td> <td>903,558,136</td> </tr> <tr> <td>Clinical cases misc.</td> <td>102,855,267</td> </tr> <tr> <td>EHRs documents<strong>*</strong></td> <td>95,267,204</td> </tr> <tr> <td>Scielo</td> <td>60,007,289</td> </tr> <tr> <td>BARR2 Background</td> <td>24,516,442</td> </tr> <tr> <td>Wikipedia (Life Sciences)</td> <td>13,890,501</td> </tr> <tr> <td>Patents</td> <td>13,463,387</td> </tr> <tr> <td>EMEA</td> <td>5,377,448</td> </tr> <tr> <td>Mespen (MedlinePlus)</td> <td>4,166,077</td> </tr> <tr> <td>PubMed</td> <td>1,858,966</td> </tr> </tbody> </table> <p>More information about the corpus can be found here&nbsp;<a href="https://aclanthology.org/2022.bionlp-1.19/">https://aclanthology.org/2022.bionlp-1.19/</a> and here&nbsp;<a href="https://arxiv.org/abs/2109.07765">https://arxiv.org/abs/2109.07765</a></p> <p>The processing took place&nbsp;on an HPC <a href="https://www.bsc.es/innovation-and-services/technical-information-cte-amd">node</a> equipped&nbsp;with an AMD EPYC 7742 (@ 2.250GHz) processor with 128 threads.</p> <p><strong>How to use</strong></p> <p>First initialize the spacy&nbsp;vectors from the floret table (.floret file):</p> <pre><code class="language-bash">spacy init vectors es floret_embeddings_bio_es.floret floret_embeddings_bio_es --mode floret</code></pre> <pre><code>import spacy # Load the floret vectors floret_embeddings = spacy.load("floret_embeddings_bio_es") # Get the embeddings of some words diabetes = floret_embeddings.vocab["diabetes"] insulina = floret_embeddings.vocab["insulina"] radiografia = floret_embeddings.vocab["radiografia"] # Get some similarities print(diabetes.similarity(insulina)) print(diabetes.similarity(radiografia)) # diabetes should be more similar to insuline than radiografia </code></pre> <p>&nbsp;</p> <p><strong>Intended Uses and Limitations</strong></p> <p>At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model. However, we are well aware that our models may be biased since the corpora have been collected using crawling techniques on multiple web sources. We intend to conduct research in these areas in the future, and if completed, this&nbsp;card will be updated.</p> <p><strong>Authors</strong></p> <p>The Text Mining Unit from Barcelona Supercomputing Center.</p> <p><strong>Contact Information</strong></p> <p>For further information, send an email to <a href="http://plantl-gob-es@bsc.es">plantl-gob-es@bsc.es</a></p> <p><strong>Funding</strong></p> <p>This work was funded by the <a href="https://portal.mineco.gob.es/en-us/digitalizacionIA/Pages/sedia.aspx">Spanish State Secretariat for Digitalization and Artificial Intelligence (SEDIA)</a>&nbsp;within the framework of the Plan-TL.</p> <p><strong>Copyright </strong></p> <p>Copyright (c) 2022&nbsp;Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

MORFITT : A multi-label corpus of French scientific articles in the biomedical domain

<p>This article presents MORFITT, the first multi-label corpus in French annotated in specialties in the medical field. MORFITT is composed of 3~624 abstracts of scientific articles from PubMed, annotated in 12 specialties for a total of 5,116 annotations. We detail the corpus, the experiments and the preliminary results obtained using a classifier based on the pre-trained language model CamemBERT. These preliminary results demonstrate the difficulty of the task, with a weighted average F1-score of 61.78%.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications

<p>This repository contains the dataset for the study of <a href="https://doi.org/10.1093/gigascience/giad113">computational reproducibility of Jupyter notebooks from biomedical publications</a>. Our focus lies in evaluating the extent of reproducibility of Jupyter notebooks derived from GitHub repositories linked to publications present in the biomedical literature repository, PubMed Central. We analyzed the reproducibility of Jupyter notebooks from GitHub repositories associated with publications indexed in the biomedical literature repository PubMed Central. The dataset includes the metadata information of the journals, publications, the Github repositories mentioned in the publications and the notebooks present in the Github repositories.</p> <p><strong>Data Collection and Analysis</strong></p> <p>We use the code for reproducibility of Jupyter notebooks from the study done by <a href="../record/2592524">Pimentel et al., 2019</a> and adapted the code from <a href="https://github.com/fusion-jena/ReproduceMeGit">ReproduceMeGit</a>. We provide code for collecting the publication metadata from PubMed Central using <a href="https://biopython.org/docs/1.76/api/Bio.Entrez.html">NCBI Entrez utilities via Biopython</a>.</p> <p>Our approach involves searching PMC using the esearch function for Jupyter notebooks using the query: ``(ipynb OR jupyter OR ipython) AND github''. We meticulously retrieve data in XML format, capturing essential details about journals and articles. By systematically scanning the entire article, encompassing the abstract, body, data availability statement, and supplementary materials, we extract GitHub links. Additionally, we mine repositories for key information such as dependency declarations found in files like requirements.txt, setup.py, and pipfile. Leveraging the GitHub API, we enrich our data by incorporating repository creation dates, update histories, pushes, and programming languages.</p> <p>All the extracted information is stored in a SQLite database. After collecting and creating the database tables, we ran a pipeline to collect the Jupyter notebooks contained in the GitHub repositories based on the code from Pimentel et al., 2019.</p> <p>Our reproducibility pipeline was started on 27 March 2023.</p> <p><strong>Repository Structure</strong></p> <p>Our repository is organized into two main folders:</p> <ul> <li><strong>archaeology</strong>: This directory hosts scripts designed to download, parse, and extract metadata from PubMed Central publications and associated repositories. There are 24 database tables created which store the information on articles, journals, authors, repositories, notebooks, cells, modules, executions, etc. in the db.sqlite database file.</li> <li><strong>analyses</strong>: Here, you will find notebooks instrumental in the in-depth analysis of data related to our study. The db.sqlite file generated by running the archaelogy folder is stored in the analyses folder for further analysis. The path can however be configured in the config.py file. There are two sets of notebooks: one set (naming pattern N[0-9]*.ipynb) is focused on examining data pertaining to repositories and notebooks, while the other set (PMC[0-9]*.ipynb) is for analyzing data associated with publications in PubMed Central, i.e.\ for plots involving data about articles, journals, publication dates or research fields. The resultant figures from the these notebooks are stored in the 'outputs' folder.</li> <li><strong>MethodsWorkflow</strong>: The MethodsWorkflow file provides a conceptual overview of the workflow used in this study.</li> </ul> <p><strong>Accessing Data and Resources:</strong></p> <ul> <li>All the data generated during the initial study can be accessed at https://doi.org/10.5281/zenodo.6802158</li> <li>For the latest results and re-run data, refer to this link.</li> <li>The comprehensive SQLite database that encapsulates all the study's extracted data is stored in the db.sqlite file.</li> <li>The metadata in xml format extracted from PubMed Central which contains the information about the articles and journal can be accessed in pmc.xml file.</li> </ul> <p><strong>System Requirements:</strong></p> <ul> <li>Centos 7 (Documentation: https://www.centos.org/)</li> <li>Conda 4.9.4 (Installation Guide: https://docs.anaconda.com/anaconda/install/linux/)</li> <li>Python 3.7.6 (Download Link: https://www.python.org/downloads/)</li> <li>GitHub account (Get Started: https://github.com/, Requires GitHub Username and Token)</li> <li>gcc 7.3.0 (Installation Guide: https://gcc.gnu.org/install/)</li> <li>lbzip2 (Command: `conda install -c conda-forge lbzip2')</li> </ul> <p><strong>Running the pipeline:</strong></p> <ul> <li>Clone the computational-reproducibility-pmc repository using Git:<br>git clone https://github.com/fusion-jena/computational-reproducibility-pmc.git<br>&nbsp;</li> <li>Navigate to the computational-reproducibility-pmc directory:<br>cd computational-reproducibility-pmc/computational-reproducibility-pmc</li> <li>Configure environment variables in the config.py file:<br>GITHUB_USERNAME = os.environ.get("JUP_GITHUB_USERNAME", "add your github username here")<br>GITHUB_TOKEN = os.environ.get("JUP_GITHUB_PASSWORD", "add your github token here")</li> <li>Other environment variables can also be set in the config.py file.<br>BASE_DIR = Path(os.environ.get("JUP_BASE_DIR", "./")).expanduser() # Add the path of directory where the GitHub repositories will be saved<br>DB_CONNECTION = os.environ.get("JUP_DB_CONNECTION", "sqlite:///db.sqlite") # Add the path where the database is stored.</li> <li>To set up conda environments for each python versions, upgrade pip, install pipenv, and install the archaeology package in each environment, execute:<br>source conda-setup.sh</li> <li>Change to the archaeology directory<br>cd archaeology</li> <li>Activate conda environment. We used py36 to run the pipeline.<br>conda activate py36</li> <li>Execute the main pipeline script (r0_main.py):<br>python r0_main.py</li> </ul> <p><strong>Running the analysis:</strong></p> <ul> <li>Navigate to the analysis directory.<br>cd analyses</li> <li>Activate conda environment. We use raw38 for the analysis of the metadata collected in the study.<br>conda activate raw38</li> <li>Install the required packages using the requirements.txt file.<br>pip install -r requirements.txt</li> <li>Launch Jupyterlab<br>jupyter lab</li> <li>Refer to the Index.ipynb notebook for the execution order and guidance.</li> </ul> <p><strong>References:</strong></p> <ul> <li>Sheeba Samuel, Daniel Mietchen. (2024). Computational reproducibility of Jupyter notebooks from biomedical publications, https://doi.org/10.1093/gigascience/giad113, GigaScience</li> <li>Sheeba Samuel, Daniel Mietchen. (2022). Computational reproducibility of Jupyter notebooks from biomedical publications, https://arxiv.org/pdf/2209.04308.pdf, CoRR abs/2209.04308</li> <li>Sheeba Samuel, &amp; Daniel Mietchen. (2022). Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6802158</li> </ul> <p>&nbsp;</p>

opencc-zeroJul 2022View details →
dryad40/100

The new normal? Redaction bias in biomedical science

Open the record for dataset details and reuse information.

publicDec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record