Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

764

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

764 results for “Reproducibility”

Learn how ShareScore rates datasets ↗
zenodo40/100

Design and Development of a Smartphone-Based Geolocalized Exposure Therapy Software for Anxiety Disorders: SyMptOMS-ET -- Reproducibility Package

<p>R Notebook and datasets for the submitted paper "<em>Towards a self-applied, mobile-based&nbsp;geolocated exposure therapy software&nbsp;for anxiety disorders: SyMptOMS-ET app</em>"</p> <blockquote> <p>Alberto Gonz&aacute;lez-P&eacute;rez, Laura Diaz-Sanahuja, Miguel Matey-Sanz, Jorge Osma, Carlos Granell, Juana Bret&oacute;n-L&oacute;pez, Sven Casteleyn.&nbsp;Towards a self-applied, mobile-based&nbsp;geolocated exposure therapy software&nbsp;for anxiety disorders: SyMptOMS-ET app.&nbsp;<a href="https://journals.sagepub.com/home/dhj">Digital Health Journal</a> [Submitted]</p> </blockquote> <p>Experiments were conducted using the v1.2.0 version of the SyMptOMS-ET open-source app, which can be found <a href="https://github.com/GeoTecINIT/symptoms-mobile-app/releases/tag/v1.2.0">here</a>.</p>

openother-openDec 2022View details →
zenodo40/100

Human CD34 bone marrow SCE data set to reproduce Totem protocols

<p>The data set <code>human_cd34_bm_rep1.rds</code> was parsed with the R script <code>download_h5ad_to_SCE_rds_script.R</code> (see github repository <a href="https://github.com/elolab/Totem-protocol">elolab/Totem-protocol</a>). It is a parsed <code>SingleCellExperiment</code> <code>RDS</code> object corresponding to the anndata h5ad <code>human_cd34_bm_rep1.h5ad</code> available on <a href="https://github.com/elolab/Totem-protocol/blob/main">HCA Portal</a> and published by <a href="https://www.nature.com/articles/s41587-019-0068-4">Setty et al., 2019</a>.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Reproducibility use cases from ORKG

<p>The dataset offers a selection of use cases from ORKG that serve as noteworthy examples for the reproducibility score. We previously showcased the concept of the reproducibility score during the inaugural symposium organized by <a href="https://zenodo.org/communities/grn/?page=1&amp;size=20">The German Reproducibility Network (GRN)</a>. For detailed information, please refer to our presentation, accessible through the following <a href="https://doi.org/10.5281/zenodo.7974324">link</a>.</p> <p>&nbsp;</p> <p><strong>How to use:</strong></p> <ol> <li>The column &quot;Paper ID&quot; corresponds to a resolve paper within ORKG. To reach the specific resource, please navigate to the following URL:&nbsp;<a href="https://orkg.org/paper/XXX/">https://orkg.org/paper/XXX/</a> , where you should replace the &quot;XXX&quot; with the actual Paper ID.</li> <li>The column &quot;Research Field&quot; symbolizes the domain to which the paper is assigned.</li> <li>The column &quot;Template ID&quot; functions likewise to the &quot;Paper ID,&quot; but in this case, you will need to visit the URL:&nbsp;<a href="https://orkg.org/template/XXX/">https://orkg.org/template/XXX/</a> &nbsp;and replace the &quot;XXX&quot; accordingly. Please be aware that a single paper may have multiple templates associated with it.</li> </ol>

opencc-by-4.0May 2023View details →
zenodo40/100

Data and code to reproduce analysis in Sacchi et al. 2023: Sex-specific fitness consequences of mate change in Scopoli's shearwaters (Calonectris diomedea)

<p>This data package contains data and code needed to reproduce the analysis reported in&nbsp;<strong>Sex-specific fitness consequences of mate change in Scopoli&rsquo;s shearwaters (Calonectris diomedea), doi&nbsp;</strong>10.1016/j.anbehav.2023.05.017</p> <p>Included are:</p> <p>1) Readme file - description of all other files and variables described in datasets</p> <p>2) Appendix.Rmd = R code file containing all the analysis reported in the paper</p> <p>3) breed.txt = this data table contains data for the analysis of fitness/breeding success</p> <p>4) skip.txt = this data table contains data for the analysis of skipping behaviour</p> <p>5) females.txt = this file contains data in headed format to build CR models for the female population. Covariate indicates if an individual&#39;s life history started with event 1 (partner known, first partner) or with event 3 (partner unknown).</p> <p>6) males.txt = this file contains data in headed format to build CR models for the male population. Covariate indicates if an individual&#39;s life history started with event 1 (partner known, first partner) or with event 3 (partner unknown).</p> <p>7) gepat.pat = this file contains the matrix design necessary to run CR models with e-surge (version 2.2.3)</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Dataset to reproduce firgures for the paper "A unifying method to study Respiratory Sinus Arrhythmia dynamics implemented in a new toolbox"

<p>Dataset provided to reproduce figures&nbsp;<br> for the paper &quot;A unifying method to study Respiratory Sinus Arrhythmia dynamics implemented in a new toolbox&quot;</p> <p>Jupyter notebooks are available here:<br> https://github.com/samuelgarcia/physio_benchmark</p> <p>Human dataset<br> =============</p> <p>Context: A research aimed to decipher the impact of respiration on brain oscillations</p> <p>Data collection methods: ECG and Respiration of 15 healthy adults subjects&nbsp;<br> (age : 30.9 +/- 9.5 yo).All participants gave informed consent to take part to the study, and all experiments<br> were approved by the national french committee (CPP number 4090). They were sitting quietly and instructed just<br> to relax. Recording lasted 5 minutes. Respiration signal was recorded from a nasal sensor<br> (Sensortechnics GmbH, Puchheim , Germany) at a sampling rate &nbsp;of 1000 Hz, amplified by actiCHamp<br> Plus amplifier (Brain Products GmbH, Gilching, Germany). ECG signal was recorded from 3 skin electrodes<br> (right forearm, left forearm, left iliac region), at a sampling rate of 1000 Hz (same amplifier).</p> <p>Structure of files: tabular separated values text files.<br> The first columns correspond to the ECG signal, the second is the respiratorysignal.<br> The sampling rate is 1000Hz</p> <p>Data manipulations: The original dataset has longer durationand and contain channels (EEG).<br> This sub-dataset was extracted from the original using the neo python package from the VHDR brain product format.<br> Signal tarces haven&#39;t been preprocessed they correspond to the &quot;raw&quot; signal.</p> <p><br> Data confidentiality and permissions: Experiments were approved by the national french committee (CPP number 4090)</p> <p><br> Animal dataset<br> ==============</p> <p>Context: The dataset was recorded to validate the device telemetric jacket from Etisense.</p> <p>Data collection methods: ECG and Respiration of 1 adult rat were recorded. Recording lasted 30 seconds during<br> freely behaving. Respiration signal and ECG were recorded from a thoraco-abdominal telemetric<br> jacket at which it was habituated before. Recorded were done at a sampling rate of 500 Hz, amplified by<br> Etisense acquisition unit (Etisense, MedTech company, Lyon, France).</p> <p>Structure of files: tabular separated values text files.<br> The first columns correspond to the ECG signal, the second is the respirator signal.<br> The sampling rate is 500Hz</p> <p><br> Data manipulations:<br> The dataset was extracted from the original HDF5 structure.<br> The ECG signal correcpond to the &quot;raw&quot; traces from the HDF5 files.<br> The respiratory signal was originaly sample at 200Hz on the device and resample with linear interpolation<br> to 500Hz to be easy aligned with the ECG signal.</p> <p>Data confidentiality and permissions: Experiments were carried according to the ethical guidelines of the<br> European Communities Council Directive of 24 November 1986 (86/609/EEC), as well as the approval 16979 of the<br> Lyon 1 University CEEA-55 ethical committee and of the Ministry of Higher Education, Research and Innovation.<br> &nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications

<p>This repository contains the dataset for the study of <a href="https://doi.org/10.1093/gigascience/giad113">computational reproducibility of Jupyter notebooks from biomedical publications</a>. Our focus lies in evaluating the extent of reproducibility of Jupyter notebooks derived from GitHub repositories linked to publications present in the biomedical literature repository, PubMed Central. We analyzed the reproducibility of Jupyter notebooks from GitHub repositories associated with publications indexed in the biomedical literature repository PubMed Central. The dataset includes the metadata information of the journals, publications, the Github repositories mentioned in the publications and the notebooks present in the Github repositories.</p> <p><strong>Data Collection and Analysis</strong></p> <p>We use the code for reproducibility of Jupyter notebooks from the study done by <a href="../record/2592524">Pimentel et al., 2019</a> and adapted the code from <a href="https://github.com/fusion-jena/ReproduceMeGit">ReproduceMeGit</a>. We provide code for collecting the publication metadata from PubMed Central using <a href="https://biopython.org/docs/1.76/api/Bio.Entrez.html">NCBI Entrez utilities via Biopython</a>.</p> <p>Our approach involves searching PMC using the esearch function for Jupyter notebooks using the query: ``(ipynb OR jupyter OR ipython) AND github''. We meticulously retrieve data in XML format, capturing essential details about journals and articles. By systematically scanning the entire article, encompassing the abstract, body, data availability statement, and supplementary materials, we extract GitHub links. Additionally, we mine repositories for key information such as dependency declarations found in files like requirements.txt, setup.py, and pipfile. Leveraging the GitHub API, we enrich our data by incorporating repository creation dates, update histories, pushes, and programming languages.</p> <p>All the extracted information is stored in a SQLite database. After collecting and creating the database tables, we ran a pipeline to collect the Jupyter notebooks contained in the GitHub repositories based on the code from Pimentel et al., 2019.</p> <p>Our reproducibility pipeline was started on 27 March 2023.</p> <p><strong>Repository Structure</strong></p> <p>Our repository is organized into two main folders:</p> <ul> <li><strong>archaeology</strong>: This directory hosts scripts designed to download, parse, and extract metadata from PubMed Central publications and associated repositories. There are 24 database tables created which store the information on articles, journals, authors, repositories, notebooks, cells, modules, executions, etc. in the db.sqlite database file.</li> <li><strong>analyses</strong>: Here, you will find notebooks instrumental in the in-depth analysis of data related to our study. The db.sqlite file generated by running the archaelogy folder is stored in the analyses folder for further analysis. The path can however be configured in the config.py file. There are two sets of notebooks: one set (naming pattern N[0-9]*.ipynb) is focused on examining data pertaining to repositories and notebooks, while the other set (PMC[0-9]*.ipynb) is for analyzing data associated with publications in PubMed Central, i.e.\ for plots involving data about articles, journals, publication dates or research fields. The resultant figures from the these notebooks are stored in the 'outputs' folder.</li> <li><strong>MethodsWorkflow</strong>: The MethodsWorkflow file provides a conceptual overview of the workflow used in this study.</li> </ul> <p><strong>Accessing Data and Resources:</strong></p> <ul> <li>All the data generated during the initial study can be accessed at https://doi.org/10.5281/zenodo.6802158</li> <li>For the latest results and re-run data, refer to this link.</li> <li>The comprehensive SQLite database that encapsulates all the study's extracted data is stored in the db.sqlite file.</li> <li>The metadata in xml format extracted from PubMed Central which contains the information about the articles and journal can be accessed in pmc.xml file.</li> </ul> <p><strong>System Requirements:</strong></p> <ul> <li>Centos 7 (Documentation: https://www.centos.org/)</li> <li>Conda 4.9.4 (Installation Guide: https://docs.anaconda.com/anaconda/install/linux/)</li> <li>Python 3.7.6 (Download Link: https://www.python.org/downloads/)</li> <li>GitHub account (Get Started: https://github.com/, Requires GitHub Username and Token)</li> <li>gcc 7.3.0 (Installation Guide: https://gcc.gnu.org/install/)</li> <li>lbzip2 (Command: `conda install -c conda-forge lbzip2')</li> </ul> <p><strong>Running the pipeline:</strong></p> <ul> <li>Clone the computational-reproducibility-pmc repository using Git:<br>git clone https://github.com/fusion-jena/computational-reproducibility-pmc.git<br>&nbsp;</li> <li>Navigate to the computational-reproducibility-pmc directory:<br>cd computational-reproducibility-pmc/computational-reproducibility-pmc</li> <li>Configure environment variables in the config.py file:<br>GITHUB_USERNAME = os.environ.get("JUP_GITHUB_USERNAME", "add your github username here")<br>GITHUB_TOKEN = os.environ.get("JUP_GITHUB_PASSWORD", "add your github token here")</li> <li>Other environment variables can also be set in the config.py file.<br>BASE_DIR = Path(os.environ.get("JUP_BASE_DIR", "./")).expanduser() # Add the path of directory where the GitHub repositories will be saved<br>DB_CONNECTION = os.environ.get("JUP_DB_CONNECTION", "sqlite:///db.sqlite") # Add the path where the database is stored.</li> <li>To set up conda environments for each python versions, upgrade pip, install pipenv, and install the archaeology package in each environment, execute:<br>source conda-setup.sh</li> <li>Change to the archaeology directory<br>cd archaeology</li> <li>Activate conda environment. We used py36 to run the pipeline.<br>conda activate py36</li> <li>Execute the main pipeline script (r0_main.py):<br>python r0_main.py</li> </ul> <p><strong>Running the analysis:</strong></p> <ul> <li>Navigate to the analysis directory.<br>cd analyses</li> <li>Activate conda environment. We use raw38 for the analysis of the metadata collected in the study.<br>conda activate raw38</li> <li>Install the required packages using the requirements.txt file.<br>pip install -r requirements.txt</li> <li>Launch Jupyterlab<br>jupyter lab</li> <li>Refer to the Index.ipynb notebook for the execution order and guidance.</li> </ul> <p><strong>References:</strong></p> <ul> <li>Sheeba Samuel, Daniel Mietchen. (2024). Computational reproducibility of Jupyter notebooks from biomedical publications, https://doi.org/10.1093/gigascience/giad113, GigaScience</li> <li>Sheeba Samuel, Daniel Mietchen. (2022). Computational reproducibility of Jupyter notebooks from biomedical publications, https://arxiv.org/pdf/2209.04308.pdf, CoRR abs/2209.04308</li> <li>Sheeba Samuel, &amp; Daniel Mietchen. (2022). Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6802158</li> </ul> <p>&nbsp;</p>

opencc-zeroJul 2022View details →
dryad40/100

Data and reproducible code for Honor et al: Direct and indirect fitness effects of competition limit evolution of allelopathy in an invading plant

<p><span>Upon introduction to new continents, invading species encounter novel communities of consumers, pathogens, and competitors. Both phenotypic plasticity and rapid evolution can facilitate adaptation across these heterogenous communities, facilitating further invasion. However, the rate and extent of adaptive evolution on contemporary timescales can be constrained by phenotypic plasticity and limits imposed by genetic co-variation for traits under selection.</span></p> <p><span>We measured phenotypic plasticity and quantified genetic co-variation for growth, competition, and fitness among </span>23 naturally inbred seed families <span>of <em>Alliaria petiolata</em> (garlic mustard) </span>collected across its invasive range in eastern North America. After growing a self-pollinated generation in a uniform common garden to reduce maternal effects, we reared second-generation plants in a <span>two-year greenhouse and field experiment with naïve soil from an uninvaded habitat. W</span>e measured selection gradients and lifetime fitness when reared alone, with an intraspecific competitor, and under interspecific competition with naïve <em>Acer saccharum </em>(sugar maple) saplings.</p> <p>Total glucosinolate production was strongly correlated with the production of chlorophyll a (Chl a) (<em>R<sup>2</sup></em> = 0.45) such that first principal component (PC1) accounted for 84% of variation in these two traits. Furthermore, PC1 exhibited high plasticity across growing environments (p &lt; 0.001) with limited broad-sense heritability (<em>H<sup>2</sup> </em>= 2.91; p = 0.08). In contrast, investment in glucosinolate production relative to Chl a (PC2) was significantly heritable (<em>H</em><sup><em>2</em> </sup>=16.91, p &lt; 0.001) with minimal plasticity across treatments. Causal analysis revealed that plastic variation for higher Chl a + glucosinolate production (PC1) had an indirect positive effect on A. petiolata fitness via a direct, negative effect on <em>A. saccharum</em> performance. In contrast, heritable variation for higher glucosinolate investment (PC2) had a direct, positive effect on <em>A. saccharum</em> performance and an indirect negative effect on A. petiolata fitness. </p> <p>Applying causal inference, we find that evolution of allelopathy in <em>A. petiolata</em> has been constrained by (i) a lack of genetic variation, (ii) selection against glucosinolate investment under interspecific competition, and (iii) phenotypic plasticity. These factors limit adaptive evolution but maintain fitness during population growth as plants switch from interspecific to intraspecific competition during invasion.</p>

opencc-zeroAug 2023View details →
zenodo40/100

Data to reproduce analysis in Convergent evolution of extrachromosomal DNA in mCRPC paper

<p>Targeted cancer therapies can prolong the lives of men with metastatic castration resistant prostate cancer (mCRPC). However, these treatments also selectively favor the growth of tumor cells that harbor therapy resistance, and mCRPC is currently lethal. It has been challenging to study factors influencing how therapy resistance develops in this setting because few autopsy studies of have been performed in the settings of DNA-repair deficient mCRPC. Here, we assessed how resistance to targeted cancer therapies evolved in an autopsy cohort of 53 mCRPC tumors from six such men using deep whole genome and transcriptome analysis, validating our observations in an independent cohort of 135 mCRPC tumors. We identified intra-patient heterogeneity in clinically actionable DNA repair deficiencies and transcriptionally-defined tumor subtypes. Identical polygenic DNA repair resistance mutations were present in physically distinct tumors within the same individual, suggesting that these mutations pre-exist selection by later targeted therapy. Extra-chromosomal DNA (ecDNA) was present in more than half of mCRPC biopsies and frequently amplified the androgen receptor (<em>AR</em>) and enhancers of <em>AR</em> and <em>MYC</em>. Individual ecDNA amplicons included multiple driver genes on different chromosomes, and arose multiple times within distinct tumors in a single patient. The presence of ecDNA was significantly associated with whole genome doubling, chromothripsis, and with inactivating <em>TP53</em> alterations. We conclude that ecDNA amplification is a major contributor to therapy resistance in mCRPC and that late-stage mCRPC develops intra-patient heterogeneity in response to targeted therapy.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Data and program codes to reproduce the results presented in "Crustal structure beneath Central Kamchatka inferred from ambient noise tomography"

<p>This archive contains the SURF_TOMO code for the surface-wave tomography (<em>Koulakov et al., 2016</em>) and all the necessary files and instructions for reproducing the results presented in the paper &ldquo;Crustal structure beneath Central Kamchatka inferred from ambient noise tomography&rdquo; by Igor Egorushkin, Ivan Koulakov, Andrey Jakovlev, Hsin-Hua Huang, Eugeny I. Gordeev, Ilyas Abkadyrov, and Danila Chebrov.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

"The sound comes from a meadow in the Sierra Nevada Mountains in California. The meadow is at an elevation of 2400 meters near a mountain named Olancha Peak, which is 3700 meters in altitude. Ihave a group of friends with which Ibackpack (trek) into the mountains. Our goal was to spend some time in the mountains and hike to the top of Olancha Peak (…) By the time we reached the meadow, we were in a forest and there was still snow on the ground in some places. We took the trip in June of 2006. The Sierra Nevada Mountains are a large mountain range. Much of the range is protected by national parks or preserved areas we call 'wilderness areas' (…) Ihave been backpacking for nearly 40 years and Iwill hopefully continue with this challenging activity for 40 years more! Many of my friends are much younger than Iam and it gives me much satisfaction to be able to have as much or more stamina for this activity than they have! When we are on these trips, we hike up peaks, catch fish, drink some whiskey around campfires and enjoy our time in the beautiful solitude. My memories of this trip were of the steep, hot hike from the desert to the cool meadow; the overall beauty of the nature, the absolute solitude of our campsite near the meadow; the strenuous hike to the top of Olancha Peak; the camaraderie of my friends; and, of course the sound of the frogs in the meadow. The frog sounds were astounding to me and Iwould listen in awe of the creature's instinctual desire to reproduce and continue the existence of their kind. Surely there were different species in the meadow for some of the frog sounds were different than others. The sounds only occurred after the Sun went down for the evening. Istood next to the creek in the meadow and recorded the sounds using my digital camera." [Peter/plentz1960]16 in Collecting Sounds. Online Sharing of Field Recordings as Cultural Practice

"The sound comes from a meadow in the Sierra Nevada Mountains in California. The meadow is at an elevation of 2400 meters near a mountain named Olancha Peak, which is 3700 meters in altitude. Ihave a group of friends with which Ibackpack (trek) into the mountains. Our goal was to spend some time in the mountains and hike to the top of Olancha Peak (…) By the time we reached the meadow, we were in a forest and there was still snow on the ground in some places. We took the trip in June of 2006. The Sierra Nevada Mountains are a large mountain range. Much of the range is protected by national parks or preserved areas we call 'wilderness areas' (…) Ihave been backpacking for nearly 40 years and Iwill hopefully continue with this challenging activity for 40 years more! Many of my friends are much younger than Iam and it gives me much satisfaction to be able to have as much or more stamina for this activity than they have! When we are on these trips, we hike up peaks, catch fish, drink some whiskey around campfires and enjoy our time in the beautiful solitude. My memories of this trip were of the steep, hot hike from the desert to the cool meadow; the overall beauty of the nature, the absolute solitude of our campsite near the meadow; the strenuous hike to the top of Olancha Peak; the camaraderie of my friends; and, of course the sound of the frogs in the meadow. The frog sounds were astounding to me and Iwould listen in awe of the creature's instinctual desire to reproduce and continue the existence of their kind. Surely there were different species in the meadow for some of the frog sounds were different than others. The sounds only occurred after the Sun went down for the evening. Istood next to the creek in the meadow and recorded the sounds using my digital camera." [Peter/plentz1960]16

opencc-by-4.0Dec 2019View details →
ClinicalTrials.gov40/100

Reproducibility of 6 Minute Walk Tests for Oxygen Desaturation

ClinicalTrials.gov study NCT00740220. IPD Sharing: UNDECIDED. Countries: 1. Publications: 34.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad40/100

Reproducible data, example subsets, and analysis pipeline for the extended TAaCGH study of breast cancer genomic and transcriptomic profiles

Open the record for dataset details and reuse information.

publicJan 2026View details →
dryad40/100

Data, analytical, and plotting codes needed to reproduce meta-analyses on within-season divorce in birds

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad40/100

Data and reproducible code for Honor et al: Direct and indirect fitness effects of competition limit evolution of allelopathy in an invading plant

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad40/100

Data and reproducible analysis files from: Latitudinal clines in floral display associated with adaptive evolution during a biological invasion

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad40/100

Code and data used for reproducing calculations of lunar crustal thermal evolution and zircon resetting during a tidal heating event

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad40/100

Supporting code and data to reproduce analysis for: Genomic signatures of past megafrugivore-mediated dispersal in Malagasy palms

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad40/100

Data to reproduce: Cultural diffusion dynamics depend on behavioral production rules

Open the record for dataset details and reuse information.

publicAug 2022View details →
OpenNeuro36/100

3AM straight reproducibility phantoms

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo36/100

Resources for reproducing experiments in "Novel Entity Discovery from Web Tables"

<p>This repository contains resources developed for the paper: &quot;S. Zhang, E. Meij, K. Balog, and R. Reinanda. Novel Entity Discovery from Web Tables. In: Proceeding of the The Web Conference 2020 (WWW &rsquo;20), April 2020&quot;.</p> <p>It includes the three test collections for novel entity discovery for Web tables, entity type and mention resolution, as well as the mention-entity and heading-property correspondences for 3M tables. The cited datasets were used in this work.</p> <p>Files to recreate the entity linking experiments:</p> <ul> <li>training_el.csv</li> <li>training_el_type.csv</li> <li>training_el_type_wiki.csv</li> <li>training_el_wiki.csv</li> <li>training_schema.csv</li> </ul> <p>Files to recreate the table matching experiments:</p> <ul> <li>me_corres.csv - textual cells algorithmically linked to Wikipedia entities</li> <li>hp_corres.csv - same but only table headings</li> </ul> <p>Files to recreate the entity resolution experiments:</p> <ul> <li>ec_golden.csv - 20K unlinked mentions textual cells, manually linked to Wikipedia</li> <li>er_sf_golden.csv - 1K cell values, manually clustered</li> <li>er_type_golden.csv - 1K cell values, manually linked to DBpedia types</li> </ul>

opencc-by-4.0Jan 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record