Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

151

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

151 results for “Reference dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq

<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts &ndash; BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5&rsquo; and 3&rsquo; UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

AirHeritage Datalake: Multi-site, Multi-season, Multi Unit dataset including Fixed and Mobile Citizen science data from networked Air Quality Low-Cost Multi-Sensors devices and reference stations

<p>This datalake comprises several datasets from <strong>37 networked low cost air quality multisensors</strong> (<strong>30</strong> <strong>mobile</strong> ENEA MONICA(tm) +&nbsp;<strong>7</strong> <strong>fixed</strong>) along with <strong>3</strong> (fixed) + <strong>1</strong> (mobile) <strong>reference stations</strong> operated by Campania Regional Envronmental Protection Agency. The datalake is organized in 3 main directories respectively related to fixed nodes, mobile nodes and nearby reference stations including a mobile laboratory used for colocation campaigns; each subdirectory include its own metadata description file.</p> <p>Data, curated by Energy and Data Science Laboratory of ENEA, include multi-weeks colocation periods when low cost devices have been colocated with reference stations as well as operational periods during which sensors are deployed for fixed or mobile monitoring campaigns. Data have been recorded during 2021 and 2022 in a<strong> pervasive, multi-site, multi-seasonal deployment</strong> in Portici, a densely populated small area city (4km2, 55k + inhabitants) located 7km south of Naples, Italy.</p> <p>The datalake consists in actual sensors and reference intrumentations timeseries along with metadata description files with&nbsp; &nbsp;deployment dates and location data. The dataset files include high sampling frequency raw sensor data of quality-controlled sensor network along with co-located reference stations data sets. Sensor data include electrochemical sensors data (intended target pollutants: NO2, O3, CO), Optical sensor data (PM2.5, PM10, PM1) readings along with meteorological parameters. .</p> <p>Further description of sensors and reference instruments are reported in the accompanying paper (see citation request).</p> <p>The dataset can be used for&nbsp;</p> <ul> <li>&nbsp;<strong>advanced (remote/universal/in field) data driven calibration strategies</strong> test or development including <strong>machine learning </strong>models</li> <li><strong>mobile opportunistic data fusion</strong> methods development</li> <li><strong>geomatics and data assimilation</strong> models studies</li> </ul> <p>as well as low cost sensor characterization performance studies.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Reference sheet for the dataset: BioCam ALR PLOCAN Trials (03/14/2024 – 03/16/2024)

<p>This reference document provides a brief description of the data and a link to the full data repository, which includes mapping and imagery data.</p> <p>For further information, please email <span><a href="mailto:adrian.bodenmann@soton.ac.uk">adrian.bodenmann@soton.ac.uk</a> or <a href="mailto:comms.techoceans@aquatt.ie">comms.techoceans@aquatt.ie</a>.</span></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

An open dataset of IPCC reports 'references (6th Assessment Cycle) – Version 1

<p>We provide a first version of an open dataset of publications cited by the IPCC reports of the 6<sup>th</sup> Assessment Cycle (<a href="https://www.ipcc.ch/reports/">Reports &mdash; IPCC</a>)</p> <p>The lists were extracted from the reference sections of three special reports and the first assessment report. [see Figure 1]. The data are presented in two formats: one on hand in text files with a list of references for each section of the reports (generally each chapter) and one other hand a structured format (json) with identifiers for the documents and the sections, the reference in string format as well as the extracted digital object identifiers (dois). In this first version, the dois extracted are mainly those which are provided in the references. The table 1 show the number of references and doi for each report.</p> <p>We plan, for subsequent releases, following enhancements:</p> <ul> <li>Further quality assurance of dois. We note that some entries are provided in references of the IPCC reports without dois although they are indexed in Crossref. The table 2 shows substantial differences in doi coverage among reference sections. Spot checks of the dataset suggest that those differences are mainly due to referencing behaviour of the section&rsquo;s authors rather than on type of documents cited. We aim in the next version to complete those missing dois and systematically verify the dois included in the references.</li> <li>Expand the references lists to reports from past assessment cycles</li> </ul> <p>In addition, we plan a more detailed documentation of our dataflow &amp; extraction process and to demonstrate how it can be used to create open, community curated, datasets of references of others non-scholarly documents.</p> <p>The json files have each two keys (1) schema with the structure of the table and (2) data: with the records.</p> <p>A simple way to read them into table is via a pandas dataframe</p> <p><em>import pandas as pd</em></p> <p><em>df = pd.read_json(file_name.json, orient = &lsquo;table&rsquo;)</em></p> <p>&nbsp;</p> <p><strong>Acknowledgment</strong></p> <p>We thank Valentin Hancu (EC/DG ECFIN) - for fruitful discussions on the data extraction process.</p> <p><strong>Disclaimer:&nbsp;</strong></p> <p>The views expressed in this paper are the author&rsquo;s. They do not reflect the views or official positions of the European Commission.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Reference dataset of multi-objective and multi-fidelity optimization in laser-plasma acceleration

<p>This repository contains a dataset used for the article &quot;<em>Multi-objective and multi-fidelity Bayesian optimization of laser-plasma acceleration</em>&quot; (<a href="https://arxiv.org/abs/2210.03484">arXiv:2210.03484</a>). The dataset consists of 2443 FBPIC particle-in-cell simulations of a laser wakefield accelerator that were selected using a Bayesian optimizer. The goal of the optimization was to perform multi-objective multi-fidelity optimization of electron beam parameters. The dataset contains simulations of different resolutions, accordingly with differing&nbsp;fidelities. The typical runtime at lowest (highest) resolution is approximately 1 (90) minutes.</p> <p>In the dataset we have <em>train_x </em>and <em>train_obj </em>numpy arrays with dimensions <em>(n,5)</em> and<em> (n,3)</em>, respectively. Here&nbsp;<em>n</em> is the number of FBPIC simulations. The five columns in <em>train_x </em>are [plasma density, upramp length, laser focus, downramp length, fidelity]. The fidelity parameter controls the resolution and hence the runtime of the simulation. The three columns in the <em>train_obj </em>are the [total charge, distance of median&nbsp;to target energy, bandwidth of electron beams]. For the distance, the&nbsp;target energy is fixed to 300 MeV&nbsp;and for the bandwidth is defined by the median absolute deviation around the median. The two columns have negative values since the optimizer assumes a maximization of all objectives while the distance and bandwidth in this study were being minimized.</p> <p>The different folders contain data of different kind of single and multi-objectives that were used to produce figures 2, 3, 5 in the associated paper.&nbsp;For more details please see the referred article. The folder &quot;combined&quot; contains the data of all simulations together and is most suitable for (5D x 3D)&nbsp;surrogate model generation.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Hominid Palaeoproteomic Reference Dataset

<p>This dataset contains the &#39;Hominid Palaeoproteomic Reference Dataset&#39;.We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline )&nbsp; to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 176 publicly available whole genomes from extant and extinct hominid groups.</p> <p>We also translated 8 ancient hominin genomes from VCF files, including those of 3 Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose&nbsp;to translate proteins that have previously been&nbsp;reported as present in either teeth or bones. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them.&nbsp;For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of around 10,058 protein sequences for each individual in the dataset.</p> <p>Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf ). The full list of the proteins translated can be found here:&nbsp;https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/Reference_Protein_List.txt and a table with information on each sample included in the dataset can be found here:&nbsp;https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/Reference_Sample_List.csv</p> <p>&nbsp;</p> <p>Content:&nbsp;</p> <p>The zipped file contains 5 files: one txt file,&nbsp;two fasta files as well as two additional folders:</p> <p>-&nbsp;&nbsp;PalaeoProPhyler_Publication_Data_for_Tree.fa contains all of the sequences used to generate the phylogenetic tree presented at PalaeoProPhylers manuscript.</p> <p>-&nbsp;ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Hominid Palaeoproteomic Reference Dataset described above, all in a single fasta.</p> <p>-&nbsp;PER_PROTEIN is a folder containing one fasta file for each protein within the&nbsp;Hominid Palaeoproteomic Reference Dataset. Each protein fasta file has the sequences of all individuals for that particular protein.</p> <p>-&nbsp;PER_SAMPLE is a folder containing one fasta file for each sample/individual within the&nbsp;Hominid Palaeoproteomic Reference Dataset. Each sample fasta file has the sequences of all proteins for that particular sample.</p> <p>-Reference_Protein_List.txt is a txt file containing two columns. The first column is a list of all the proteins selected to be translated. The second column describes where each of these proteins&nbsp;was mentioned or identified. If a protein was identified in a publication, the title of the publication is given. If a protein was&nbsp;identified in one of the publications of our group (E.Cappellini group) the identifier &#39;our samples&#39; is given. If multiple publications supported a protein they are all given and seperated by comma.</p> <p>&nbsp;</p> <p>~ NOTE (!) ~</p> <p>Depending on which samples you use we highly encourage you to cite the original publication(s) from which we got the DNA data and translate the proteins from:</p> <p>&nbsp;</p> <p>Modern Humans:</p> <p>M Byrska-Bishop et al. &ldquo;High Coverage Whole Genome Sequencing of the Expanded 1000 Genomes Project Cohort Including 602 Trios. bioRxiv. 2021&rdquo;.</p> <p>Non-human great apes:</p> <p>Javier Prado-Martinez et al. &ldquo;Great ape genetic diversity and population history&rdquo;. In: Nature 499.7459 (2013), pp. 471&ndash;475.</p> <p>Pongos:</p> <p>Alexander Nater et al. &ldquo;Morphometric, behavioral, and genomic evidence for a new orangutan species&rdquo;. In: Current Biology 27.22 (2017), pp. 3487&ndash;3498.</p> <p>Neanderthal, Denisovan and other ancient anatomically modern humans:</p> <p>Kay Pr&uml;ufer et al. &ldquo;A high-coverage Neandertal genome from Vindija Cave in Croatia&rdquo;. In: Science 358.6363 (2017), pp. 655&ndash;658.</p> <p>Fabrizio Mafessoni et al. &ldquo;A high-coverage Neandertal genome from Chagyrskaya Cave&rdquo;. In: Proceedings of the National Academy of Sciences 117.26 (2020), pp. 15132&ndash;15136.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Dataset from the study "Analysis of the accuracy of scientific literature references provided by ChatGPT"

<p>This dataset corresponds to the study carried out to analyse 10 bibliographic references of 10 Spanish authors in the field of Information Sciences requested to the ChatGPT chatbot.</p> <p>The file &quot;Bibliographic_references_ analysis&quot; contains the 10 references returned by ChatGPT for each of the 10 authors (a total of 100 references), together with the variables analysed to check their authenticity.</p> <p>The &quot;Keywords_analysis&quot; file contains the normalisation carried out on the words considered to be key words extracted from the titles of the works, according to which a word cloud showing the frequency of occurrence could be drawn up.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Proboscidean Palaeoproteomic Reference Dataset

<p>This entry contains the &#39;Proboscidean Palaeoproteomic Reference Dataset&#39;.</p> <p>We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline )&nbsp;to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day Proboscidae. Using the first two modules of PaleoProPhyler, we translated more than 35 publicly available whole genomes from extant and extinct species. Details on the processing of the sequences can be found below.</p> <p>&nbsp;</p> <p>Which individuals / species are included?</p> <p>The full list of individuals, the original fastq repository location and the species included in the dataset are contained within the tab seperated file &#39;METATADATA.txt&#39;, that also contains headers. Most individuals of the dataset&nbsp;belong to one of these 3 species: <em>Loxodonta africana</em>,<em> Elephas maximus</em>,&nbsp;<em>Mammuthus&nbsp; primigenius.&nbsp;</em></p> <p>&nbsp;</p> <p>Which Proteins are included?</p> <p>We compiled a small initial&nbsp;list of 262 proteins that had been indentified in either teeth, bone or items made out of ivory. For each protein, both the canonical and all alternative protein coding isoforms (based on the Loxodonta africana reference proteome of Ensembl)&nbsp;were translated, leading to more than 350&nbsp;unique protein sequences for each individual in the dataset. The protein list is available in the file &#39;proteins.txt&#39;</p> <p>&nbsp;</p> <p>How were the proteins translated/generated?</p> <p>All genetic data were downloaded from ENA (https://www.ebi.ac.uk/ena/browser/home) as fastq files and mapped onto LoxAfr3, which is the latest annotated African elephant genome in Ensembl. The scripts used for the mapping are available here:&nbsp;https://github.com/johnpatramanis/Mapping_Scripts . We used the resulting bam files as input for PaleoProPhyler&#39;s module 1 &amp; 2 , using LoxAfr3 as the reference proteome.</p> <p>Other files included in the zip folder:</p> <p>-&nbsp;ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Proboscidean Palaeoproteomic Reference Dataset described above</p> <p>-&nbsp;PER_PROTEIN is a folder containing one fasta file for each protein within the&nbsp;Proboscidean Palaeoproteomic Reference Dataset, each protein fasta file has the sequences of all individuals for that particular protein</p> <p>-&nbsp;PER_SAMPLE is a folder containing one fasta file for each sample/individual within the&nbsp;Proboscidean Palaeoproteomic Reference Dataset, each sample fasta file has the sequences of all proteins for that particular sample.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

TauBench: A Dynamic Benchmark for Graphics Rendering (Dataset Reference Frames)

<p>TauBench is a dynamic graphics rendering benchmark dataset, targeted especially towards&nbsp;rendering methods relying on the reuse of temporal data. The dataset&nbsp;is available at <a href="https://doi.org/10.5281/zenodo.5729573">https://doi.org/10.5281/zenodo.5729573</a>, and this upload provides path traced reference frames for it&nbsp;in PNG format. The images are rendered with <a href="https://github.com/vga-group/tauray">Tauray</a>&nbsp;at 16384 samples per pixel (spp), at both 1080p and 2160p resolutions. Frame indices start&nbsp;from 0 and are <em>not</em> padded with leading zeroes.</p> <p>More information about TauBench is also available at&nbsp;<a href="https://webpages.tuni.fi/vga/taubench">https://webpages.tuni.fi/vga/taubench</a>.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Wikidata Subsetting: Reference-based Subsetting Experiment Datasets

<p>Files in this dataset have been produced during Flexibility experiments of Wikidata subsetting practical tools: Subsetting based on references using WDSub.</p>

opencc-byJun 2023View details →
zenodo44/100

The Object Detection for Olfactory References (ODOR) Dataset.

<p>Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide instance-level<br> annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The proposed ODOR dataset fills this gap, offering 38,116 object-level annotations across 4,712 images, spanning an extensive set of 139 fine-grained categories. Conducting a statistical analysis, we showcase challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. Furthermore, we provide an extensive baseline analysis for object detection models and highlight the challenging properties of the dataset through a set of secondary studies. Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.</p> <p><strong>How to use</strong></p> <p>The annotations are provided in COCO JSON format. To represent the two-level hierarchy of the object classes, we make use of the supercategory field in the categories array as defined by COCO. In addition to the object-level annotations, we provide an additional CSV file with image-level metadata, which includes content-related fields, such as Iconclass codes [72 , 73]) or image<br> descriptions, as well as formal annotations, such as artist, license, or creation year. For the sake of license compliance, we do not publish the images directly (although most of the images are public domain). Instead, we provide links to their source&nbsp; collections in the metadata file (meta.csv) and a python script to download the artwork images (download_images.py).</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

mGEMS Enterococcus faecalis reference dataset

<p>This dataset contains the <em>E. faecalis</em> sequences (all available assemblies from the NCBI as of 2 February 2020 which could be assigned to a multilocus sequence type), their multilocus sequence types inferred with the mlst software (v2.18.1) using the MLST scheme described in <a href="https://doi.org/10.1128/JCM.02596-05">10.1128/JCM.02596-05</a>, and accession numbers used in the mGEMS publication as the reference dataset.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Reference dataset to check the Python package for the calculation of the OLR-based MJO index (OMI)

<p>The Python package (<a href="https://doi.org/10.5281/zenodo.3613752">10.5281/zenodo.3613752</a>) for the calculation of the OLR-based MJO (OMI) index comes with integration tests, with which the users can check the calculation results.</p> <p>However, in order to run some of these tests, input and reference datasets are needed. Specifically, a dataset of Outgoing Longwave Radiation (OLR) described by Liebmann and Smith (1996) is needed as the input and OMI values described by Kiladis et al. (2014) are needed as the reference.</p> <p>These files are provided by this upload and can be downloaded either as .tar.gz or as .zip file.</p> <p>Whereas exactly the files provided here are needed for the tests, updated versions of these datasets can be found on the websites of NOAA: <a href="https://psl.noaa.gov/data/gridded/data.interp_OLR.html">https://psl.noaa.gov/data/gridded/data.interp_OLR.html</a> and <a href="https://www.esrl.noaa.gov/psd/mjo/mjoindex/">https://www.esrl.noaa.gov/psd/mjo/mjoindex/</a>.</p>

opengpl-3.0-or-laterApr 2020View details →
zenodo40/100

dataset for paper Vanhaebost J, Faouzi M, Mangin P, Michaud K: New reference tables and user-friendly Internet application for predicted heart weights. Int J Legal Med 2014, 128(4):615-620.

<p>The heart weight is the most important parameter in the determination of cardiac hypertrophy. The obtained heart weight value should be compared against tables of normal weights by age, gender and body weight and height</p> <p>In the study by Vanhaebost<em> et al</em>. &nbsp;has been shown in the Swiss population that the heart weight increases along with the increase of the body weight, body height, BMI and body surface area (BSA). The mean heart weight is greater in men than in women at a similar body weight. The reference tables for predicted heart weights obtained from this study are presented as an user-friendly internet application (<a href="http://calc.chuv.ch/Heartweight">http://calc.chuv.ch/Heartweight</a>)&nbsp; enabling the comparison of heart weights observed at autopsy with the reference values.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

The Neogene 2020 reference dataset (De Nil et al., 2020)

<p>Table 2, Table 3 and Table 4 as the resulting dataset from &#39;De Nil, K., De Ceukelaire, M. &amp; Van Damme, M., 2020. A reference dataset for the Neogene lithostratigraphy in Flanders, Belgium. Geologica Belgica, 23/3-4. <a href="https://eur03.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdoi.org%2F10.20341%2Fgb.2020.021&amp;data=04%7C01%7Ckatrien.denil%40vlaanderen.be%7Cb6f4944eb0d54af5762308d896026f24%7C0c0338a695614ee8b8d64e89cbd520a0%7C0%7C0%7C637424284489463336%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C1000&amp;sdata=vZAilldByLDJ7KhLFQ2z8ajnapX53lea4Zfh2YAAFGo%3D&amp;reserved=0">https://doi.org/10.20341/gb.2020.021</a>&#39;</p> <p><strong>Table</strong><strong> 2.</strong> The general and different (sub)reference datasets resulting from the individual Neogene 2020 papers, with their DOV URL.</p> <p><strong>Table 3.</strong> List of the individual boreholes and (temporary) outcrops of the Neogene reference set, with reference to the different individual papers of this collection, sorted by location as mentioned in the papers.&nbsp; *Complete &lsquo;<a href="https://www.dov.vlaanderen.be/data/">https://www.dov.vlaanderen.be/data/</a>boring/&rsquo; with this unique code. **Complete <a href="http://collections.naturalsciences.be/ssh-geology-archives/arch/">http://collections.naturalsciences.be/ssh-geology-archives/arch/</a> with this unique code.</p> <p><strong>Table 4</strong>. List of the individual CPT&rsquo;s with reference to the different individual papers of this collection, sorted by DOV name.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Dataset for comparison of QuantumPower method to the reference power standard

<p>Dataset for comparison of QuantumPower method to the power standard Radian RD-22.</p> <p>The QuantumPower method was compared to a power standard Radian RD-22. As a device under test, a Fluke 6100 power calibrator was used.</p> <p>To obtain the data, QPSW software was used:</p> <p>https://github.com/KaeroDot/QPsw</p> <p>Author: Martin &Scaron;&iacute;ra</p> <p>Contact: Czech Metrology Institute, Okružn&iacute; 31, 638 00 Brno, msira@cmi.cz</p> <p>Part of project Quantum traceability for AC power standards, QuantumPower, Project Number: 19RPT01.<em> </em>This project (19RPT01) has received funding from the EMPIR programme co-financed by the Participating States and from the European Union's Horizon 2020 research and innovation programme.</p> <p>https://www.euramet.org/research-innovation/search-research-projects/details/project/quantum-traceability-for-ac-power-standards/</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

The Helicobacter pylori Genome Project (HpGP) Phase1 dataset and 255 H. pylori population reference dataset

<p>This repository holds the HpGP Phase 1 genomic dataset for Hp26695 and 1011 study samples. All 1012 genomic sequences were annotated using the NCBI Prokaryotic Genome Annotation Pipeline(PGAP). Also, it has 255 curated public available H. pylori genomic sequences used for population structure analysis in Thorell et al. Nature Communications, 14:8184 (2023).</p> <p>You can check the NCBI BioProject website for the latest annotation and sequence updates.</p> <p>https://www.ncbi.nlm.nih.gov/bioproject/?term=HpGP</p> <p>Please cite the above-mentioned paper if you use the data.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Public metagenome datasets annotated using SingleM, using a supplemented reference package.

<p>The SingleM package used for supplementing is available at 10.5281/zenodo.10360136</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Dataset accompanying "A new species of Mesolepis from the Late Carboniferous of Scotland, with especial reference to Mesolepis wardi Young"

<p>This dataset accompanies the manuscript "A new species of <em>Mesolepis </em>from the Late Carboniferous of Scotland, with especial reference to <em>Mesolepis wardi </em>Young" (<span><a href="https://doi.org/10.1017/S1755691024000094" target="_blank" rel="noopener">https://doi.org/10.1017/S1755691024000094</a>)</span> and comprises the following items:&nbsp;</p> <p>- <em>Mesolepis arabellae</em> GLAHM 163398/1 (part) raw data (TIFF stack, zipped)</p> <p>- <em>Mesolepis arabellae</em> GLAHM 163398/1 (part) .mcs file</p> <p><em>- Mesolepis arabellae</em> GLAHM 163398/1 (part) .ply files (zipped)</p> <p>- <em>Mesolepis arabellae</em> GLAHM 163398/2 (counterpart) raw data (TIFF stack, zipped)</p> <p><em>- Mesolepis arabellae</em> GLAHM 163398/2 (counterpart).mcs file</p> <p>- <em>Mesolepis arabellae</em> GLAHM 163398/2 (counterpart) .ply files (zipped)</p> <p>-&nbsp;<em>Mesolepis arabellae </em>GLAHM 163398/2 (fin region) raw data (TIFF stack, zipped)</p> <p><em>- Mesolepis arabellae</em> GLAHM 163398/2 (fin region) .mcs file</p> <p>- <em>Mesolepis arabellae </em>GLAHM 163398/2 (fin region)&nbsp;.ply file</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

TCR-DeepInsight Reference Datasets

<p>Reference single-cell TCR immune profiling datasets for the TCR-DeepInsight analysis. We have excluded two controlled-access datasets (1) AML dataset from Abbas et al. 2021 (EGAS00001004894) and (2) Kawasaki disease dataset from Wang et al., 2021 (OEP001162).</p> <ul> <li>human_gex_reference_v2.h5ad: Processed H5AD file for transcriptome features from the single-cell TCR immune profiling datasets.&nbsp;<strong><em>anndata.read_h5ad<br>For raw count H5AD files, please check the additional repo here: <a href="https://zenodo.org/records/17405143">https://zenodo.org/records/17405143</a></em></strong></li> <li>human_tcr_reference_v2.h5ad: Processed H5AD file for unique TCR clonotypes single-cell TCR immune profiling dataset.&nbsp;<strong><em>anndata.read_h5ad</em></strong></li> <li>Yi_2023_Ankylosing_Spondylitis.h5ad: Processed H5AD file from Yi&nbsp;<em>et al</em>., 2023 (GSE216885). <strong><em>anndata.read_h5ad</em></strong></li> <li>GSE272993_cd8_nn_labeled_FINAL.fl_tcr.match_v2_5.transfered.h5ad: Processed H5AD file from Wang et al., 2024 (GSE272993). <strong><em>anndata.read_h5ad</em></strong><em><strong><br><br></strong></em></li> <li>human_bulk_tcr_reference.parquet: PARQUET file for bulk TCR&beta; sequencing. <strong><em>pandas.read_parquet</em></strong> <ul> <li>human_bulk_tcr_reference.cd4.parquet: PARQUET file for bulk TCR&beta; sequencing for CD4-sorted T cells<br>human_bulk_tcr_reference.mait.parquet: PARQUET file for bulk TCR&beta; sequencing for MAIT cells<br>human_bulk_tcr_reference.treg.parquet: PARQUET file for bulk TCR&beta; sequencing for Treg cells<br>human_bulk_tcr_reference.cd8.parquet: PARQUET file for bulk TCR&beta; sequencing for CD8-sorted T cells<br><em><br></em></li> </ul> </li> <li>Yi_2023_Ankylosing_Spondylitis.h5ad. <strong><em>anndata.read_h5ad<br></em></strong>Processed dataset from K. Yi et al. Analysis of Single‐Cell Transcriptome and Surface Protein Expression in Ankylosing Spondylitis Identifies OX40 ‐Positive and Glucocorticoid‐Induced Tumor Necrosis Factor Receptor&ndash;Positive Pathogenic Th17 Cells. <em>Arthritis &amp; Rheumatology</em> <strong>75</strong>, 1176&ndash;1186 (2023).</li> <li>GSE272993_cd8_nn_labeled_FINAL.fl_tcr.match_v2_5.transfered.h5ad. <strong><em>anndata.read_h5ad<br></em></strong>Processed dataset from K. Wang et al. Combination anti-PD-1 and anti-CTLA-4 therapy generates waves of clonal responses that include progenitor-exhausted CD8+ T cells. <em>Cancer Cell</em>, S1535610824003064 (2024).</li> </ul> <ul> <li>human_gex_reference_v2.scatlasvae.ckpt: pretrained weight state dict from scAtlasVAE model for human_gex_reference_v2.h5ad. <strong><em>torch.load</em></strong></li> <li>human_bert_pseudosequence.tcr_v2.ckpt: pretrained weight state dict from BERT for human_tcr_reference_v2.h5ad. <em><strong>torch.load</strong><br></em></li> <li><em>human_bert_pseudosequence_pca.tcr_v2.pkl: pretrained PCA weight for&nbsp;</em>human_tcr_reference_v2.h5ad. <em><strong>pickle.load</strong></em></li> </ul>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record